Subdivision of grid sequence
By using the √3 segmentation method to subdivide the grid in video encoding technology, the problem of low compression efficiency of dynamic grid sequences in the existing technology is solved, and more efficient data compression and transmission is achieved.
Patent Information
- Application Number
- CN202480004410.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2024-05-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the existing video encoding technology, the compression efficiency of dynamic grid sequences is low, especially when processing attribute graphs and connectivity information that change over time, the mid-point subdivision method leads to excessive granularity and excessive number of displacement vectors.
The mesh is subdivided using the √3 subdivision method, generating new vertices and flipping the inner edges, reducing the increase of the number of triangles and reducing the number of displacement vectors.
It improves the fine granularity of grid segmentation, reduces the number of displacement vectors, and improves the compression efficiency and data transmission performance of video encoding.
Smart Images

Figure BDA0005359694140000101 
Figure BDA0005359694140000111 
Figure BDA0005359694140000112
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims the priority benefit of U.S. Provisional Application No. 63 / 468,756, filed on May 24, 2023, and U.S. Application No. 18 / 672,798, filed on May 23, 2024. The entire disclosure of the prior applications is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure relates to a set of advanced video coding techniques that include mesh sequences. Background Art
[0003] Advances in three - dimensional (3D) capture, modeling, and rendering have facilitated the ubiquity of 3D content across a variety of platforms and devices. Today, a baby's first steps can be captured on one continent while its grandparents on another continent can see (and perhaps interact) and enjoy a fully immersive experience with the child. However, to achieve this realism, the models have become increasingly complex, and a large amount of data is associated with the creation and use of these models. 3D meshes are widely used to represent such immersive content.
[0004] A mesh is composed of multiple polygons that describe the surface of a volumetric object. Each polygon is defined by the vertices of its corresponding polygon in 3D space and information on how those vertices are connected (referred to as connectivity information). Optionally, vertex attributes (such as color, normal, etc.) can be associated with the vertices of the mesh. Through the use of mapping information, attributes can also be associated with the surface of the mesh, where the mapping information parameterizes the mesh using a two - dimensional (2D) attribute map. This mapping is typically described by a set of parametric coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. The 2D attribute map is used to store high - resolution attribute information, such as texture, normal, displacement, etc. This information can be used for various purposes, such as texture mapping and shading.
[0005] Dynamic mesh sequences can require a large amount of data as they can contain a large amount of information that changes over time. Therefore, efficient compression techniques are needed to store and transmit such content. The Moving Picture Experts Group (MPEG) has previously developed mesh compression standards (such as Information and Communication IC Mesh Compression, MESHGRID, Frame-Based Animation Mesh Compression (FAMC)) for handling dynamic meshes with constant connectivity, geometry that changes over time, and vertex attributes. However, these standards do not consider attribute maps and connectivity information that change over time. DCC (Digital Content Creation) tools typically generate such dynamic meshes. Accordingly, it is challenging for volumetric acquisition techniques to generate dynamic meshes with constant connectivity, especially under real-time constraints. Existing standards do not support this type of content. MPEG is planning to develop a new mesh compression standard to directly handle dynamic meshes with connectivity information that changes over time and an optional attribute map that changes over time. The standard targets lossy and lossless compression for various applications such as real-time communication, storage, free viewpoint video, augmented reality (AR), and virtual reality (VR). Features such as random access and scalable / progressive coding are also considered.
[0006] In October 2011, MPEG issued a call for proposals for dynamic mesh coding and received multiple responses in April 2022. A reference software, namely the so-called V-Mesh Test Model (V-MeshTestModel), was developed based on the responses received for evaluating dynamic mesh coding tools during the standardization process.
[0007] In a mesh codec such as the MPEG V-Mesh test model, the mid-point subdivision method is used to refine the sampled mesh sequence because it is relatively simple. However, one drawback of mid-point subdivision is that the granularity is large, i.e., each iteration of mid-point subdivision quadruples the number of triangles. For example, three subdivision iterations increase the number of triangles to 64 times the original number of triangles.
[0008] In addition, displacement vectors (either 3D vectors or one-dimensional (1D) vectors along the normal direction) are typically used to offset the positions of mid-point vertices and original vertices. Since a mid-point vertex is generated for each edge of the original mesh, the number of additional displacement vectors will be the number of edges in the original mesh.
[0009] Due to these reasons, technical solutions are needed to address these problems that arise in video coding techniques. Summary of the Invention
[0010] The present disclosure includes a method and an apparatus. The apparatus includes: a memory configured to store computer program code, and one or more processors configured to access the computer program code and operate according to the instructions of the computer program code. The computer program is configured to cause the processor to implement obtaining code, and the obtaining code is configured to cause at least one processor to obtain, from a bitstream, a mesh of encoded volumetric data representing at least one three-dimensional (3D) visual content; and decode the encoded volumetric data based on displacement vectors of vertices of the mesh. The displacement vectors are based on √3 subdivision, in which the faces of the mesh are iteratively subdivided in the following manner: the number of subdivided faces generated by subdividing the mesh in one iteration is less than four times the number of faces of the mesh before subdivision in that iteration.
[0011] According to one aspect of the present disclosure, in the iteration, the √3 subdivision may include: inserting one vertex among a plurality of inserted vertices into a face of the mesh before subdivision in the iteration, and interconnecting the plurality of inserted vertices to form a plurality of subdivided faces of the mesh.
[0012] According to one aspect of the present disclosure, inserting one vertex among a plurality of inserted vertices into a face of the mesh before subdivision in the iteration may include: inserting only one vertex among the plurality of inserted vertices into each face of the mesh before subdivision in the iteration.
[0013] According to one aspect of the present disclosure, inserting at least one vertex among the inserted vertices into a face of the mesh before subdivision may be based on a linear combination of coordinates of vertices of the one face.
[0014] According to one aspect of the present disclosure, the linear combination may include: multiplying each coordinate of the vertices in the face by a corresponding non-negative real coefficient.
[0015] According to one aspect of the present disclosure, each of the corresponding non-negative real coefficients may be set to 1 / 3.
[0016] According to one aspect of the present disclosure, the √3 subdivision may include: subdividing the boundary edges of the mesh in a manner different from the internal edges of the mesh, such that at least two new inserted vertices are added to at least one of the boundary edges in the iteration. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0018] Figure 1 illustrates aspects of a communication system according to one or more embodiments of the present disclosure.
[0019] Figure 2 Illustrates aspects of a media processing system according to one or more embodiments of the present disclosure.
[0020] Figure 3 Illustrates aspects of a decoder according to one or more embodiments of the present disclosure.
[0021] Figure 4 Illustrates aspects of an encoder according to one or more embodiments of the present disclosure.
[0022] Figure 5 Illustrates aspects of the encoder side according to one or more embodiments of the present disclosure.
[0023] Figure 6 Illustrates aspects of the decoder side according to one or more embodiments of the present disclosure.
[0024] Figure 7 Illustrates aspects of an encoder and decoder system in the context of media grid features according to one or more embodiments of the present disclosure.
[0025] Figure 8 Illustrates aspects of grid processing according to one or more embodiments of the present disclosure.
[0026] Figure 9 Illustrates aspects of grid processing according to one or more embodiments of the present disclosure.
[0027] Figure 10 Shows a flowchart of aspects of grid processing according to one or more embodiments of the present disclosure.
[0028] Figure 11 Illustrates aspects of a grid decoder according to one or more embodiments of the present disclosure.
[0029] Figure 12 Illustrates aspects of a grid subdivision scheme according to one or more embodiments of the present disclosure.
[0030] Figure 13 Illustrates aspects of a grid subdivision process according to one or more embodiments of the present disclosure.
[0031] Figure 14 Illustrates aspects of a grid subdivision scheme according to one or more embodiments of the present disclosure.
[0032] Figure 15 Illustrates aspects of a boundary - considered grid subdivision scheme according to one or more embodiments of the present disclosure.
[0033] Figure 16Illustrations of aspects of a grid displacement vector according to one or more embodiments of the present disclosure.
[0034] Figure 17 Is a simplified illustration of a system according to one or more embodiments of the present disclosure. Detailed implementation
[0035] The proposed features discussed below can be used alone or in any combination. In addition, embodiments can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0036] Figure 1 Shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional data transmission, the first terminal 103 may encode video data at a local location for transmission via the network 105 to another terminal 102. The second terminal 102 may receive the encoded video data of another terminal from the network 105, decode the encoded data, and display the restored video data. Unidirectional data transmission is more common in applications such as media services.
[0037] Figure 1 Shows a second pair of terminals 101 and 104 for supporting two-way transmission of encoded video, which may occur, for example, during a video conference. For two-way data transmission, each terminal 101 and 104 may encode video data collected at a local location for transmission via the network 105 to another terminal. Each terminal 101 and 104 may also receive the encoded video data sent by another terminal, decode the encoded data, and display the restored video data on a local display device.
[0038] In Figure 1Among them, the terminal devices 101, 102, 103, and 104 may be shown as servers, personal computers, and smart phones, but the principles disclosed in this application are not limited thereto. The embodiments disclosed in this application are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 105 represents any number of networks that transfer encoded video data between the terminal devices 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. The communication network 105 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless otherwise explained below, the architecture and topology of the network 105 may be irrelevant to the operations disclosed in this application.
[0039] As an example of the application of the disclosed subject matter, Figure 2 shows how a video encoder and a video decoder are placed in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-supported applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0040] The streaming system may include an acquisition subsystem 203, which may include a video source 201 such as a digital camera that creates an uncompressed video sample stream 213, for example. Compared with the encoded video bitstream, this sample stream 213 can be emphasized as having a high data volume and can be processed by an encoder 202 coupled to the video source 201 (which may be a camera as described above). The encoder 202 may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the sample stream, the encoded video bitstream 204 is emphasized as having a lower data volume and can be stored on the streaming server 205 for future use. One or more streaming clients 212 and streaming client 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes the input copy of the encoded video bitstream 208 and creates an output video sample stream 210 that can be presented on a display 209 or other presentation device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video coding / compression standards. Examples of these standards were mentioned above and are further described herein.
[0041] Figure 3 may be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.
[0042] The receiver 302 may receive one or more encoded video sequences to be decoded by the decoder 300. In the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequences may be received from a channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective entities of use (not shown). The receiver 302 may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory 303 may be coupled between the receiver 302 and the entropy decoder / parser 304 (hereinafter referred to as "parser"). While it may not be necessary to configure the buffer 303 or the buffer may be made smaller when the receiver 302 receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network. For use on an optimal packet network such as the Internet, the buffer 303 may also be required, which may be relatively large and may advantageously have an adaptive size.
[0043] Video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy-coded video sequence. The categories of these symbols include information for managing the operation of decoder 300, as well as potential information for controlling a rendering device (e.g., display screen 312), which is not part of the decoder but can be coupled to the decoder. The control information for the rendering device may be in the form of Supplemental Enhancement Information (SEI messages) or a Video Usability Information (VUI) parameter set segment (not depicted). The parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be performed according to video coding techniques or standards and may follow principles well-known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser 304 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The entropy decoder / parser may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0044] The parser 304 may perform entropy decoding / parsing operations on the video sequence received from buffer 303 to create symbols 313. The parser 304 may receive encoded data and selectively decode specific symbols 313. Additionally, the parser 304 may determine whether a specific symbol 313 is to be provided to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0045] Depending on the type of the encoded video picture or a portion of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbols 313 may involve multiple different units. Which units are involved and the manner in which they are involved may be controlled by subgroup control information parsed by the parser 304 from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser 304 and the multiple units below are not depicted.
[0046] In addition to the functional blocks already mentioned, decoder 300 can conceptually be subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.
[0047] The first unit is the scaler / inverse transform unit 305. The scaler / inverse transform unit 305 receives, from parser 304, quantized transform coefficients as symbols 313 and control information, including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit can output a block including sample values, which can be input into aggregator 310.
[0048] In some cases, the output samples of the scaler / inverse transform unit 305 can belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but can use predictive information from a previously reconstructed part of the current picture. Such predictive information can be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the current (partially reconstructed) picture 309. In some cases, aggregator 310 adds the prediction information generated by the intra-picture prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305, based on each sample.
[0049] In other cases, the output samples of the scaler / inverse transform unit 305 can belong to an inter-coded and potentially motion-compensated block. In this case, the motion-compensation prediction unit 306 can access the reference picture memory 308 to extract samples for prediction. After motion-compensating the extracted samples according to symbol 313, these samples can be added by aggregator 310 to the output of the scaler / inverse transform unit (which is called the residual sample or residual signal in this case), thereby generating output sample information. The extraction of the prediction samples by the motion-compensation unit from an address within the reference picture memory can be controlled by a motion vector, and the motion vector is in the form of the symbol 313 and is provided for use by the motion-compensation unit. The symbol 313, for example, includes an x component, a Y component, and a reference picture component. Motion compensation can also include interpolation of the sample values extracted from the reference picture memory when using sub-sampled accurate motion vectors, a motion vector prediction mechanism, and so on.
[0050] The output samples of aggregator 310 can be employed by various loop filtering techniques in loop filter unit 311. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, and the parameters can be used for loop filter unit 311 as symbols 313 from parser 304. However, in other embodiments, video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0051] The output of loop filter unit 311 can be a sample stream that can be output to display device 312 and stored in reference picture memory 557 for subsequent inter-picture prediction.
[0052] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified (by, e.g., parser 304) as a reference picture, the current reference picture 309 can become part of reference picture buffer 308, and a new current picture memory can be reallocated before starting reconstruction of subsequent encoded pictures.
[0053] Video decoder 300 can perform decoding operations according to predetermined video compression techniques in standards such as ITU-T Recommendation H.265. An encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, e.g., megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.
[0054] In an embodiment, receiver 302 can receive additional (redundant) data along with the encoded video. The additional data can be part of the encoded video sequence. The additional data can be used by video decoder 300 to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can be in the form of, e.g., temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0055] Figure 4 It may be a functional block diagram of the video encoder 400 according to an embodiment of the present disclosure.
[0056] The encoder 400 may receive video samples from a video source 401 (not part of the encoder), and the video source may capture video images to be encoded by the video encoder 400.
[0057] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder 303. The digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit...), any color space (e.g., BT.601 YCrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera capturing local image information as a video sequence. The video data may be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0058] According to an embodiment, the encoder 400 may encode and compress pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the couplings are not labeled in the figure. The parameters set by the controller may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 402 as they may involve optimizing the video encoder 400 for a specific system design.
[0059] Some video encoders operate in an "encoding loop" that is readily apparent to those skilled in the art. As a simple description, the encoding loop can include an encoding portion of encoder 400 (hereinafter referred to as the "source encoder") (responsible for creating symbols based on the input image and reference images to be encoded), and a (local) decoder 406 embedded in encoder 400 that reconstructs the symbols to create sample data that the (remote) decoder will also create (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input into reference picture memory 405. Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture buffer is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.
[0060] The operation of the "local" decoder 406 can be the same as the operation of the "remote" decoder 300 described above in conjunction with Figure 3 the detailed description. Additionally, briefly referring to Figure 4 , however, when the symbols are available and the entropy encoder 408 and parser 304 can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding portion of decoder 300, including channel 301, receiver 302, buffer 303, and parser 304, may not be fully implementable in local decoder 406.
[0061] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. More detailed description is only required in certain areas and is provided below.
[0062] As part of its operation, source encoder 403 can perform motion compensated predictive coding that predicts and encodes an input frame with reference to one or more encoded frames in a video sequence previously designated as a "reference frame". In this way, encoding engine 407 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, and the reference picture can be selected as the prediction reference for the input frame.
[0063] The local video decoder 406 may decode the encoded video data of a frame that can be designated as a reference frame based on the symbols created by the source encoder 403. The operation of the encoding engine 407 may advantageously be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 4 not shown herein), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that may be performed by the video decoder on the reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory 405 (which may be a cache, for example). In this way, the encoder 400 may locally store a copy of the reconstructed reference frame that has the same content (without transmission errors) as the reconstructed reference frame that will be obtained by the remote video decoder.
[0064] The predictor 404 may perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 may search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor 404 may operate block by block based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it may be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory 405.
[0065] The controller 402 may manage the encoding operations of the source encoder 403 (which may be a video encoder, for example), including setting parameters and subgroup parameters for encoding the video data.
[0066] The outputs of all the above functional units may be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by various functional units according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0067] The transmitter 409 may buffer the encoded video sequence created by the entropy encoder 408 to prepare for transmission over the communication channel 411, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter 409 may merge the encoded video data from the source encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0068] The controller 402 can manage the operation of the encoder 400. During encoding, the controller 402 can assign a certain type of encoded picture to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can generally be assigned to any of the following frame types:
[0069] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics.
[0070] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0071] A bi - predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0072] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block - by - block. These blocks can be predictively encoded with reference to other (encoded) blocks, which are determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be non - predictively encoded, or the block can be predictively encoded with reference to already - encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non - predictively encoded with reference to a previously - encoded reference picture through spatial prediction or through temporal prediction. Blocks of a B picture can be non - predictively encoded with reference to one or two previously - encoded reference pictures through spatial prediction or through temporal prediction.
[0073] The encoder 400 (which can be a video encoder for example) can perform encoding operations according to a predetermined video encoding technique or standard such as the ITU - T H.265 recommendation. In operation, the encoder 400 can perform various compression operations, including predictive encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard used.
[0074] In an embodiment, the transmitter 409 may transmit additional data when transmitting the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set segments, etc.
[0075] Figure 5 FIG. 500 shows a simplified block - style workflow diagram of an exemplary viewport - related process in the Omnidirectional Media Application Format (OMAF), which may allow 360 - degree virtual reality (VR360) streams described in OMAF.
[0076] In the acquisition block 501, when the image data can represent a scene in VR360, video data A is acquired, such as data of multiple images and audio data at the same time instance. In the processing block 503, the following one or more processes are performed on the image B i : stitching, mapping to a projected picture with respect to one or more virtual reality (VR) angles or other angles / viewpoints, and packing by region. Additionally, metadata may be created to indicate any such processed information and other information to assist in the transmission and rendering process.
[0077] For data D, in the image encoding block 505, the projected picture is encoded as data E in a viewport - independent streaming manner i and a media file is composed. In the video encoding block 504, the video picture is encoded as data E v , for example, as a single - layer bitstream. For data B a , in the audio encoding block 502, the audio data may also be encoded as data E a .
[0078] Data E a 、E v and E i 、the entire encoded bitstream F iAnd / or F can be stored on a (Content Delivery Network (CDN) / cloud) server and can generally be fully transmitted to the OMAF player 520, such as at delivery block 507 or elsewhere, and can also be fully decoded by the decoder such that at display block 516, at least one region of the decoded picture corresponding to the current viewport is presented to the user with respect to various metadata, file playback, and orientation / viewport metadata from the head / eye tracking block 508, where the orientation / viewport metadata is, for example, the angle at which the user can view with respect to the viewport specification of the VR image device. A notable feature of VR360 is that only one viewport can be displayed at any given time, and this feature can be used to improve the performance of the omnidirectional video system by selective delivery according to the user's viewport (or any other criterion, such as recommended viewport timing metadata). For example, according to an exemplary embodiment, viewport-related delivery can be achieved through tile-based video coding.
[0079] Similar to the encoding blocks described above, the OMAF player 520 according to an exemplary embodiment can similarly unpack files / segments for one or more of the data F’ and / or F’i and metadata, reverse process one or more aspects of such encoding, decode the audio data E’i at the audio decoding block 510, decode the video data E’v at the video decoding block 513, and decode the image data E’i at the image decoding block 514 to continue with the audio presentation of the data B’a at the audio presentation block 511 and the image presentation of the data D’ at the image presentation block 515, thereby outputting in VR360 format according to various metadata (such as orientation / viewport metadata), displaying the data A, i at the display block 516, and outputting the audio data A’s at the speaker / headphone block 512. The various metadata can affect one of the data decoding and presentation processes according to various tracks, languages, qualities, views that can be selected by or for the user of the OMAF player 520, and it should be understood that the processing order described herein is presented for an exemplary embodiment and can be implemented in other orders according to other exemplary embodiments.
[0080] Figure 6FIG. 600 shows a simplified block content flow diagram for (encoded) point cloud data having view position and angle related processing (herein referred to as "V-PCC") for acquisition / generation / encoding or decoding / rendering / display of 6-degree-of-freedom media point cloud data. It should be understood that the described features can be used alone or in any combination, and in addition to those shown, elements such as for encoding and decoding can also be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits), and according to an exemplary embodiment, the one or more processors can execute a program stored in a non-transitory computer-readable medium.
[0081] FIG. 600 shows an exemplary embodiment for streaming encoded point cloud data according to V-PCC.
[0082] In the volume data acquisition block 601, a real visual scene or a computer-generated visual scene (or a combination thereof) can be acquired by a set of camera devices or synthesized by a computer as volume data, which can have any format. In the point cloud conversion block 602, the volume data can be converted into a (quantized) point cloud data format through image processing. For example, according to an exemplary embodiment, by extracting one or more values from the volume data and any associated data and converting them into a desired point cloud format, the data from the volume data can be converted into points in the point cloud region by region. According to an exemplary embodiment, the volume data can be a 3D data set formed by 2D images, such as slices, from which 2D projections of the 3D data set can be projected. According to an exemplary embodiment, the point cloud data format includes a representation of data points in one or more different spaces, which can be used to represent the volume data, and can provide improvements in sampling and data compression (e.g., with respect to temporal redundancy), and for example, the point cloud data in x, y, z format represents color values (e.g., RGB, etc.), luminance, intensity, etc. at each point in the cloud data. Also, the point cloud data can be used for progressive decoding, polygon meshing, direct rendering, and octree 3D representation of 2D quadtree data.
[0083] In the image projection block 603, the acquired point cloud data can be projected onto a 2D image and encoded as an image / video picture using video-based point cloud coding (V-PCC). The projected point cloud data can consist of attributes, geometry, occupancy maps, and other metadata for point cloud data reconstruction, such as using painter's algorithm, ray casting algorithm, (3D) binary space partitioning algorithm, etc. for point cloud data construction.
[0084] On the other hand, in the scene generator block 609, the scene generator can generate, for example, some metadata for presenting and displaying 6 degrees-of-freedom (DoF) media according to the director's intention or the user's preference. In addition to allowing additional dimensions for moving forward / backward, up / down, and left / right in a virtual experience within or at least based on the point cloud encoded data, such 6DoF media can also include 3D viewing scenes such as 360VR of a scene with rotational changes on the x-axis, y-axis, and z-axis of 3D. As Figure 6 As shown in and related descriptions, the scene description metadata defines one or more scenes composed of the encoded point cloud data and other media data (including VR360, light field, audio, etc.), and can be provided to one or more cloud servers and / or file / fragment encapsulation / de-encapsulation processing.
[0085] After the video encoding block 604 and the image encoding block 605 similar to the above-mentioned video and image encoding (and it will be understood that audio encoding can also be provided as described above), the file / fragment encapsulation block 606 processes such that the encoded point cloud data is combined into a media file for file playback or a sequence of initialization segments and media segments for streaming according to a specific media container file format, such as one or more video container formats, and can be used, for example, for DASH described below. Among other things, such descriptions represent exemplary embodiments. The file container can also include, for example, the scene description metadata from the scene generator block 1109 into the file or segment.
[0086] According to an exemplary embodiment, the file is encapsulated according to the scene description metadata to include at least one view position at one or more time points in the 6DoF media and at least one or more angular views at each of those view positions, such that such a file can be sent on request according to user or creator input. Additionally, according to an exemplary embodiment, a segment of such a file can include one or more portions of such a file, such as a portion of the 6DoF media indicating a single viewpoint and the angle at that viewpoint at one or more time points. However, these are merely exemplary embodiments and can vary according to various conditions such as network, user, creator capabilities, and input.
[0087] According to an exemplary embodiment, the point cloud data is divided into multiple 2D / 3D regions, which are independently encoded, for example, at one or more of the video encoding block 604 and the image encoding block 605. Then, each independently encoded division of the point cloud data can be encapsulated as a track in a file and / or segment in the file / fragment encapsulation block 606. According to an exemplary embodiment, each point cloud track and / or metadata track can include some useful metadata for view position / angle related processing.
[0088] According to an exemplary embodiment, metadata useful for view position / angle related processing, such as metadata included in a file and / or a segment encapsulated in a file / segment encapsulation block, includes one or more of the following: layout information of a 2D / 3D partition with an index, (dynamic) mapping information associating a 3D volume partition with one or more 2D partitions (e.g., any one of tiles / tile groups / slices / sub-pictures), the 3D position of each 3D partition on a 6DoF coordinate system, a list of representative view positions / angles, a list of selected view positions / angles corresponding to the 3D volume partition, indices of 2D / 3D partitions corresponding to the selected view positions / angles, quality (grade) information of each 2D / 3D partition, and rendering information of each 2D / 3D partition that depends on each view position / angle, for example. When requested, invoking such metadata, for example, by a user of a V-PCC player or according to an instruction of a content creator for the user of the V-PCC player, can enable more efficient processing of a specific part of 6DoF media required for this metadata, such that the V-PCC player can transmit an image focused on the 6DoF media part with higher quality than other parts, rather than transmitting unused parts of the media.
[0089] A file or one or more segments of a file from the file / segment encapsulation block 606 can be directly transmitted to any V-PCC player 625 and a cloud server, such as the cloud server block 607, using a transmission mechanism (e.g., HTTP-based Dynamic Adaptive Streaming over HTTP (DASH)). On the cloud server block 607, the cloud server can extract one or more tracks and / or one or more specific 2D / 3D partitions from the file and can merge multiple encoded point cloud data into one data.
[0090] According to data such as from the position / viewpoint tracking block 608, if a current viewing position and angle are defined on a 6DoF coordinate system in the client system, the viewing position / angle metadata can be transmitted from the file / segment encapsulation block 606 or otherwise processed at the cloud server block 607 from a file or segment already at the cloud server, such that the cloud server can extract an appropriate partition area from the stored file and merge these partition areas (if necessary) according to the metadata from a client system having, for example, a V-PCC player 625, and the extracted data can be transmitted to the client as a file or a segment.
[0091] Regarding such data, at the file / fragment de-encapsulation block 615, the file de-encapsulator processes the file or received fragment, extracts the encoded bitstream, and parses the metadata; at the video decoding block 610 and the image decoding block 611, the encoded point cloud data is then decoded and reconstructed into point cloud data at the point cloud reconstruction block 612. The reconstructed point cloud data can be displayed at the display block 614 and / or can first be combined at the scene composition block 613 relative to the scene description data according to the scene generator block 609 depending on one or more different scene descriptions.
[0092] In view of the above, such an exemplary V-PCC stream represents advantages over the V-PCC standard, which includes one or more of the following: the described partitioning ability for multiple 2D / 3D regions, the ability to assemble the compressed domain of the encoded 2D / 3D partitions into a single conforming encoded video bitstream, and the bit extraction ability to extract the encoded 2D / 3D of the encoded pictures into a conforming encoded bitstream. Such a V-PCC system is further improved by forming through a container including a VVC bitstream to support a mechanism for including metadata carrying one or more of the above metadata.
[0093] In view of this, according to the exemplary embodiments described further below, the term "mesh" represents a combination of one or more polygons that describe the surface of a volumetric object. Each polygon is defined by the vertices of its corresponding polygon in 3D space and information on how these vertices are connected (referred to as connectivity information). Optionally, vertex attributes (such as color, normal, etc.) can be associated with the mesh vertices. Through the use of mapping information, attributes can also be associated with the surface of the mesh, and the mapping information parameterizes the mesh using a two-dimensional (2D) attribute map. Such mapping can be described by a set of parametric coordinates associated with the mesh vertices (referred to as UV coordinates or texture coordinates). The 2D attribute map is used to store high-resolution attribute information, such as texture, normal, displacement, etc. According to the exemplary embodiments, the high-resolution attribute information can be used for various purposes, such as texture mapping and shading.
[0094] However, a dynamic mesh sequence may require a large amount of data as it may contain a large amount of information that changes over time. For example, compared to a "static mesh" or "static mesh sequence" where the information of the mesh may not change from one frame to another, a "dynamic mesh" or "dynamic mesh sequence" indicates that the vertices in the vertices represented by the mesh move and change from one frame to another. Therefore, effective compression techniques are needed to store and transmit such content. MPEG has previously developed mesh compression standards (such as IC, MESHGRID, FAMC) for handling dynamic meshes with constant connectivity and geometry and vertex attributes that change over time. However, these standards do not consider the attribute graph and connectivity information that change over time. DCC (Digital Content Creation) tools typically generate such dynamic meshes. Accordingly, for volumetric acquisition techniques, it is challenging to generate dynamic meshes with constant connectivity, especially under real-time constraints. Existing standards do not support this type of content. According to the exemplary embodiments of this document, various aspects of a new mesh compression standard are described to directly handle dynamic meshes with connectivity information that changes over time and optionally an attribute graph that changes over time, which is targeted at lossy and lossless compression for various applications (such as real-time communication, storage, free viewpoint video, AR, and VR). Features such as random access and scalable / progressive coding are also considered.
[0095] Figure 7 An example framework 700 for dynamic mesh compression is shown, for example, for a method based on 2D atlas sampling. Each frame of the input mesh 701 can be preprocessed through a series of operations (such as tracking, remeshing, parameterization, voxelization). Note that these operations can be encoder-only, meaning they may not be part of the decoding process, and this possibility can be written into the metadata with a flag, for example, 0 for encoder-only and 1 for others. After that, a mesh 702 with a 2D UV atlas can be obtained, where each vertex of the mesh has one or more associated UV coordinates on the 2D atlas. Then, by sampling on the 2D atlas, the mesh is converted into multiple graphs, including a geometry graph and an attribute graph. Then, these 2D graphs can be encoded by a video / image codec (such as HEVC, VVC, AV1, AVS3, etc.). On the decoder 703 side, the mesh can be reconstructed based on the decoded 2D graphs. Any post-processing and filtering can also be applied to the reconstructed mesh 704. Note that for the purpose of 3D mesh reconstruction, other metadata may be sent to the decoder side. Note that the boundary information of the graph (including the uv and xyz coordinates of the boundary vertices) can be predicted, quantized, and entropy encoded in the bitstream. The quantization step size can be configured on the encoder side to trade off quality and bitrate.
[0096] In some embodiments, a 3D mesh may be divided into several segments (or patches / charts). According to an exemplary embodiment, one or more 3D mesh segments may be considered as a "3D mesh". Each segment consists of a set of connected vertices associated with their geometric shape, attributes, and connectivity information. As shown in the example 800 of volumetric data Figure 8 as shown, the UV parameterization process 802 that maps from a 3D mesh segment to a 2D map (e.g., maps to the 2D UV atlas block 702 mentioned above) maps one or more mesh segments 801 onto a 2D map 803 in the 2D UV atlas 804. Each vertex (v n ) in the mesh segment will be assigned a 2D UV coordinate in the 2D UV atlas. Note that the vertices (v n ) in the 2D map form a connected component as their 3D counterparts. The geometric shape, attributes, and connectivity information of each vertex can also be obtained from their 3D counterparts. For example, it can be indicated that vertex v 4 is directly connected to vertices v 0 , v 5 , v 1 and v 3 , and similar information for each other vertex can be indicated in the same way. Additionally, according to an exemplary embodiment, the 2D texture mesh will further indicate information fragment by fragment, for example, by each triangle (e.g., the triangle formed by v 2 , v 5 , v 3 as a fragment), such as color information.
[0097] For example, further considering the features of the example 800 Figure 8 , refer to the example 900 Figure 9 , where the 3D mesh segment 801 can also be mapped to multiple separate 2D maps 901 and 902. In this case, a vertex in the 3D map may correspond to multiple vertices in the 2D UV atlas. As shown in Figure 9 , in the 2D UV atlas, the same 3D mesh segment is mapped to multiple 2D maps, rather than a single map as shown in Figure 8 . For example, 3D vertices v 1 and v 4 respectively have two corresponding 2D vertices v 1 , v 1 ' and v 4 , v 4 '. Therefore, the general 2D UV atlas of the 3D mesh can consist of multiple maps, as shown in Figure 14As shown, each graph may include multiple (usually greater than or equal to 3) vertices associated with their 3D geometry, attributes, and connectivity information.
[0098] Figure 9 An example 903 of a derived triangle partition in a graph having boundary vertices B0, B1, B2, B3, B4, B5, B6, B7 is shown. When presenting such information, any triangle partition method can be applied to create connections between vertices (including boundary vertices and sampled vertices). For example, for each vertex, find the two closest vertices. Alternatively, for all vertices, triangles are continuously generated until the minimum number of triangles is obtained after a set number of attempts. As shown in example 903, there are various regularly shaped repeating triangles and various oddly shaped triangles, and these oddly shaped triangles are usually closest to the boundary vertices and have their own unique dimensions, which may or may not be shared with any other triangle. The connectivity information can also be reconstructed through explicit signaling. If the polygon cannot be recovered by implicit rules, according to an exemplary embodiment, the encoder can write the connectivity information into the bitstream.
[0099] The boundary vertices B0, B1, B2, B3, B4, B5, B6, B7 are defined in the 2D UV space. The edges of the boundary can be determined by checking if an edge appears in only one triangle. According to an exemplary embodiment, the following information of the boundary vertices is important and should be written into the bitstream: geometric information, such as 3D XYZ coordinates (even in the current 2D UV parameter form) and 2D UV coordinates.
[0100] For the case where one boundary vertex in the 3D graph corresponds to multiple vertices in the 2D UV atlas, as Figure 9 shown, the mapping from 3D XUZ to 2D UV can be one-to-many. Therefore, a UV-to-XYZ (or called UV2XYZ) index can be written to indicate the mapping function. UV2XYZ can be a 1D array of indices that map each 2D UV vertex to a 3D XYZ vertex.
[0101] According to an exemplary embodiment, in order to effectively represent the mesh signal, a subset of the mesh vertices and their connectivity information can be encoded first. In the original mesh, the connections between these vertices may not exist because they are subsampled from the original mesh. There are different ways to write the connectivity information between vertices, so such a subset is called the base mesh or basic vertices.
[0102] According to an exemplary embodiment, a number of methods for dynamic mesh compression are implemented, and these methods are part of the above-mentioned edge-based vertex prediction framework, in which the base mesh is first encoded, and then more additional vertices are predicted based on the connectivity information of the edges from the base mesh. Note that these methods can be applied individually or in any form of combination.
[0103] For example, consider Figure 10 the example flowchart 1001 of vertex grouping for a prediction pattern. At S101, vertices within the mesh can be obtained, and at S102, the vertices can be segmented into different groups for prediction, such as Figure 9 . In one example, at S104, the segmentation is completed using fragmentation / tile division. In another example, at S105, each fragment / tile is divided. Whether to proceed from S103 to S104 or S105 can be indicated by a flag or the like. In the case of proceeding to S105, several vertices of the same fragment / tile form a prediction group and will share the same prediction pattern, while several other vertices of the same fragment / tile can use another prediction pattern. Herein, a "prediction pattern" can be considered as a specific pattern used by the decoder to predict video content including fragments. The prediction pattern can be explicitly divided into an intra prediction pattern and an inter prediction pattern, and within each category, the decoder can select different specific patterns. According to an exemplary embodiment, each group, i.e., a "prediction group" can share the same specific pattern (e.g., an angular pattern at a specific angle) or the same classified prediction pattern (e.g., all intra prediction patterns, but can be predicted at different angles). This grouping of S106 can be assigned at different levels by determining according to the number of vertices involved in each group. For example, according to an exemplary embodiment, every 64, 32, or 16 vertices following the scan order within a fragment / tile will be assigned the same prediction pattern, and other vertices can be assigned differently. For each group, the prediction pattern can be an intra prediction pattern or an inter prediction pattern. This can be signaled or assigned. According to the example flowchart 1000, at S107, if the mesh frame or mesh slice is determined to be of the intra type, for example, by checking whether the flag of the mesh frame or mesh slice indicates the intra type, then all vertex groups within the mesh frame or mesh slice will use the intra prediction pattern; otherwise, at S108, for all vertices in each group, an intra prediction or inter prediction pattern can be selected.
[0104] In addition, for a set of mesh vertices using an intra prediction mode, these vertices can only be predicted by using previously encoded vertices within the same sub-partition of the current mesh. Sometimes, according to an exemplary embodiment, the sub-partition can be the current mesh itself, and for a set of mesh vertices using an intra prediction mode, according to an exemplary embodiment, these vertices can only be predicted by using previously encoded vertices from another mesh frame. Each of the above information can be determined and signaled by a flag or the like. At S110, feature prediction can be performed. At S111, the result of the prediction can be obtained and written into the bitstream.
[0105] According to an exemplary embodiment, for each vertex in a set of vertices in exemplary flowchart 1000 and flowchart 1100 described below, after prediction, the residual will be a 3D displacement vector indicating the offset from the current vertex to its predicted point. The residuals of a set of vertices need to be further compressed. In one example, before entropy coding, at S111, a transform can be applied to the residuals of the vertex grouping and written into the bitstream. The following methods can be implemented to handle the coding of a set of displacement vectors. For example, in one method, in order to properly signal a set of displacement vectors, some displacement vectors or their components have only zero values. In another embodiment, a flag is sent for each displacement vector indicating whether the displacement vector has any non-zero components; if not, the coding of all components of the displacement vector can be skipped. In addition, in another embodiment, a flag is sent for each set of displacement vectors indicating whether the set of displacement vectors has any non-zero vectors; if not, the coding of all displacement vectors in the set of displacement vectors can be skipped. In addition, in another embodiment, a flag is sent for each component of a set of displacement vectors indicating whether the component of the set of displacement vectors has any non-zero vectors; if not, the coding of that component of all displacement vectors in the set of displacement vectors can be skipped. In addition, in another embodiment, there may be a case where a signal indicates that a set of displacement vectors or a component of the set of displacement vectors needs to be transformed, and if not, the transform can be skipped and quantization / entropy coding can be directly applied to the set of displacement vectors or the component of the set of displacement vectors. In addition, in another embodiment, a flag can be sent for each set of displacement vectors indicating whether a transform is required; if not, the transform coding of all displacement vectors in the set of displacement vectors can be skipped. In addition, in another embodiment, a flag is sent for each component of a set of displacement vectors indicating whether the component of the set of displacement vectors needs to be transformed; if not, the transform coding of that component of all displacement vectors in the set of displacement vectors can be skipped. The above embodiments regarding the processing of vertex prediction residuals in this paragraph can also be combined and implemented in parallel on different fragments respectively.
[0106] The mesh is composed of multiple polygons that describe the surface of a volumetric object. Each polygon is defined by the vertices of its corresponding polygon in 3D space and information on how these vertices are connected (referred to as connectivity information). Optionally, vertex attributes such as color, normal, etc. can be associated with the mesh vertices. By leveraging mapping information, attributes can also be associated with the surface of the mesh, where the mapping information uses a two-dimensional (2D) attribute map to parameterize the mesh. This mapping is typically described by a set of parametric coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. The 2D attribute map is used to store high-resolution attribute information such as texture, normal, displacement, etc. This information can be used for various purposes such as texture mapping and shading.
[0107] The mesh geometry information consists of vertex connectivity information, 3D coordinates, 2D texture coordinates, etc. The compression of vertex 3D coordinates (also known as vertex positions) is very important because in most cases, it occupies a large portion of the entire geometry-related data.
[0108] In Figure 11 is shown an example V-Mesh decoding process according to an exemplary embodiment. The input bitstream is demultiplexed into three parts, namely the base mesh bitstream, the displacement bitstream, and the attribute bitstream.
[0109] According to the encoding mode used in the base mesh bitstream, an intra decoding path or an inter decoding path is selected. According to an embodiment, for the intra decoding path, the base mesh is decoded by a static mesh decoder; for the inter decoding path, the motion vectors are decoded by a motion decoder. The vertices in the reference mesh stored in the mesh buffer and the decoded motion vectors are combined to form a reconstructed base mesh. If the subdivision process is enabled in the bitstream, the midpoint subdivision process is applied to generate midpoints using the vertices in the base mesh and the previously generated midpoints. The subdivided mesh is called m’(i) in Figure 11 After the inverse quantization process, m’(i) becomes m”(i).
[0110] According to an embodiment, to decode the displacement bitstream, a video decoder is applied, followed by image unpacking, inverse quantization, and inverse wavelet transform to generate the decoded displacements for the vertices in the base mesh and the midpoints generated in the subdivision process. The decoded displacements and m’(i) are combined to form a decoded mesh represented as M”(i).
[0111] According to an embodiment, to decode the attribute bitstream, if needed, a video decoder is employed, followed by a color space conversion.
[0112] However, as described above, in a mesh codec such as the MPEG V-Mesh test model, due to the simplicity of the midpoint subdivision method, this method is used to refine the extracted mesh sequence. However, one drawback of midpoint subdivision is its large granularity, i.e., each iteration of midpoint subdivision quadruples the number of triangles. For example, three subdivision iterations increase the number of triangles to 64 times the original number of triangles. Therefore, it has been found that it is feasible to introduce a subdivision with finer granularity as disclosed herein.
[0113] In addition, displacement vectors (either 3D vectors or 1D vectors along the normal direction) are typically used to offset the positions of midpoint vertices and original vertices. Since a midpoint vertex is generated for each edge of the original mesh, the number of additional displacement vectors will be the number of edges in the original mesh.
[0114] Similar to in a mesh codec such as the MPEG V-Mesh codec, according to an embodiment of the present disclosure, a midpoint subdivision scheme is used to subdivide a base mesh to recover the details of the original mesh after the extraction process. In this process, an edge is divided into two parts of equal length, such as Figure 12 in Example 1200.
[0115] In Figure 3 the original triangle is P 0 P 1 P 2 . The points P 3 、P 4 and P 5 are the midpoints of the edges P 2 P 0 、P 0 P 1 and P 1 P 2 . In this way, the original triangle is subdivided into 4 smaller triangles, namely P 0 P 4 P 3 、P 4 P 5 P 3 、P 4 P 1 P 5 、P 5 P 2 P 3 . Note that the midpoint subdivision process can be completed in multiple iterations. For example, after one more subdivision iteration, Figure 12 the triangles in will become 16 smaller triangles. In practice, according to an embodiment, 1 to 3 iterations are typically used.
[0116] According to an embodiment, the coordinates of the point Pi can be expressed as (xi , y i , z i ). Without loss of generality, taking the midpoint P3 as an example. The coordinates of P3 can be calculated by floating-point arithmetic as follows: where (x i , y i , z i ) is represented as a single-precision or double-precision number.
[0117] According to an embodiment, "√3 subdivision" is provided. The √3 subdivision is a uniform subdivision scheme in which new vertices are generated for each triangle of a given mesh. This process is shown in Example 1300 in Figure 13 and is described as the following three steps: (1) In S1302, after obtaining the original mesh M0 in S1301, a vertex is inserted into each triangle of the original mesh M0. (2) In S1303, the new vertices are connected to the three vertices of the surrounding triangles. (3) In S1304, except for those original edges at the boundary, each original edge connecting two old vertices is flipped. In the flipping operation, two new vertices in two adjacent triangles are connected, and the common edge in the two adjacent triangles is removed.
[0118] In the √3 subdivision, the number of faces increases by 3 times, rather than 4 times in the midpoint subdivision. In this sense, the √3 subdivision provides a finer granularity.
[0119] Furthermore, since a new vertex is generated for each triangle, the number of additional displacement vectors will be the number of triangles in the original mesh, which is less than the number of edges in the original mesh.
[0120] As shown in Example 1400 in Figure 4 , a triangle in the mesh M0 is represented as P 0 P 1 P 2 , and the new vertex inside the triangle is P 3 . That is to say, P 3 can be considered as the new vertex in the triangle.
[0121] According to one or more embodiments, the coordinates of the new vertex P 3 are represented as a linear combination of the coordinates of the three vertices as follows: P 3 = c 0 P 0 + c 1 P1 +c 2 P 2 , equation (2) where c 0 、c 1 、c 2 are non - negative real coefficients, and c 0 +c 1 +c 2 = 1. Note that for i = 0, 1, 2, 3, P i =(x i , y i , z i ) represents the 3D coordinates of point P i . In practice, c 0 、c 1 、c 2 are often set to be equal to In another embodiment, a non - linear combination can be employed.
[0122] As Figure 13 shown in S1304 of, starting from the original mesh, all internal edges are flipped. The boundary edges cannot be flipped because it has only one associated triangle. Thus, in the first subdivision iteration, the boundary edges are not affected.
[0123] For illustration, a small part of S1304 of Figure 13 is enlarged and shows the step S1501 of example 1500 of Figure 15 . That is, in the second or subsequent subdivision iterations, vertices are inserted into each internal triangle; for boundary triangles (such as ABD), by inserting two vertices (such as E and F), the boundary edge is subdivided into 3 equal - length edges. The inserted vertices are connected to the original vertices in the triangles around them. According to the embodiment, as Figure 15 shown in step S1502 of, for the boundary triangle ABD, ED and FD are connected. Finally, as shown in step S1503, the internal edges are flipped to complete the second iteration. Note that the triangle ABC in the original mesh is subdivided into 9 triangles after 2 iterations, just like all the internal triangles in the original mesh.
[0124] In mesh compression, displacement vectors are usually used to adjust the vertex coordinates of the subdivided mesh. See example 1600 of Figure 16 , which shows the displacement vectors of vertices P 0 , P 1 , P 2 and the new vertex P 3 .
[0125] In Figure 16 , D 0 、D 1 、D2 and D 3 are the displacement vectors of corresponding vertices. According to an embodiment, to facilitate compression, a lifted wavelet transform can be used to transform the displacement vectors, and the lifting includes prediction and update operations. Note that other types of transforms are also possible.
[0126] In one embodiment of the present invention, the displacement vector of a new vertex is predicted using the displacement vectors of the vertices of the surrounding triangles of the new vertex. Taking Figure 16 as an example, D 0 , D 1 , D 2 are used to predict D 3 . In one embodiment, linear prediction is used, as shown below: where is the predictor of D 3 ; α 0 , α 1 , α 2 are non - negative real coefficients, and α 0 +α 1 +α 2 = 1. In fact, the coefficients α 0 , α 1 , α 2 can be set to Note that D i , i = 0, 1, 2, 3 can be 3 - D vectors, or, if only the displacement along the normal direction of the associated vertices is used, D i , i = 0, 1, 2, 3 can be 1 - D vectors. Additionally, D i , i = 0, 1, 2, 3 can be 3 - D vectors in 444 sampling format, 420 sampling format, 400 sampling format, or some other sampling format. In another embodiment, non - linear prediction can be employed, i.e., where f(D 0 , D 1 , D 2 ) is a non - linear function of D 0 , D 1 , D 2 .
[0127] After prediction, the prediction residual can be calculated by subtracting its predicted value from the displacement vector. For example, the prediction residual of D 3 is as shown below:
[0128] In one embodiment of the present invention, an update operation is used, i.e., the prediction residual of the displacement vector is used to adjust the prediction residual of the original mesh. The details of this process are disclosed below.
[0129] In the encoder, it is assumed that K subdivision iterations have been completed. Before any subdivision, for distinction, the mesh is called the base mesh, and the vertices of the mesh are called the vertices in LOD 0 ; after the first subdivision, the new vertices generated inside each triangle in the subdivision are called the vertices in LOD 1 ; similarly, the new vertices generated in the k-th subdivision are called the vertices in LOD k . In the encoder, starting from LOD K to LOD 1 , a prediction / update step is performed for each LOD level.
[0130] In the prediction step, w represents the vertex in LOD k ; v i represents the vertex when the LOD level is lower than k. Without loss of generality, v 0 , v 1 and v 2 represent the three vertices of the peripheral triangle of w. D(p) and represent the displacement vector and its prediction residual of vertex p at any LOD level. According to Equations 2 and 3, we obtain the following formula: or
[0131] In the update step, v represents the vertex with an LOD level lower than k; W * represents the set of vertices in LOD k that are directly adjacent to vertex v. In one embodiment, the displacement vector of v is updated as follows: where β is a real coefficient. If β = 0, the update step is not applied. In fact, β can be set to a positive number close to 0.1.
[0132] In order to indicate the use of √3 subdivision in the encoder and the displacement vector is associated with this type of subdivision, certain signaling is required in the compressed bitstream.
[0133] Without loss of generality, taking the MPEG-V-Mesh standard as an example. The table of subdivision methods can be expanded as follows: asps_vmc_extsubdivision_method Name of the subdivision method 0 None 1 Midpoint subdivision 2 SQRT3 where SQRT3 represents √3 subdivision.
[0134] According to the embodiment, the coefficients related to the subdivision (i.e., c 0 , c 1 , c 2) and the coefficients associated with the lifted wavelet transform (i.e., α 0 , α 1 , α 2 and β) can be fixed in the standard or sent in the parameter set.
[0135] Thus, in a mesh codec such as the MPEG V-Mesh test model, since the midpoint subdivision method is relatively simple, this method is used to refine the extracted mesh sequence. However, one drawback of midpoint subdivision is that the granularity is large, i.e., the number of triangles increases fourfold with each iteration of midpoint subdivision. For example, three subdivision iterations increase the number of triangles to 64 times the original number of triangles. Therefore, as described in the embodiments herein, a subdivision with a finer granularity needs to be introduced.
[0136] In addition, displacement vectors (whether 3D vectors or 1D vectors along the normal direction) are typically used to offset the positions of midpoint vertices and original vertices. Since a midpoint vertex is generated for each edge of the original mesh, the number of additional displacement vectors will be the number of edges in the original mesh.
[0137] Through the embodiments herein, a square root of three (i.e., √3) subdivision is provided, where the number of triangles only triples after each iteration. First, the method of √3 subdivision is introduced. As an example, the related changes to the syntax and the decoding process of the MPEG V-Mesh standard are described above according to the exemplary embodiments.
[0138] The above techniques can be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media or implemented by one or more specially configured hardware processors. For example, Figure 17 FIG. 1700 shows a computer system suitable for implementing certain embodiments of the disclosed subject matter.
[0139] The computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subject to mechanisms such as assembly, compilation, linking, or the like to create code that includes instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.
[0140] The instructions can be executed on various types of computers or their components, such as personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0141] Figure 17The components of the computer system 1700 shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system 1700.
[0142] The computer system 1700 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users through inputs such as the following: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., speech, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious inputs, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.
[0143] The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard 1701, mouse 1702, touchpad 1703, touch screen 1710, joystick 1705, microphone 1706, scanner 1708, camera 1707.
[0144] The computer system 1700 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users, for example, through tactile outputs, sounds, lights, and smells / tastes. Such human-machine interface output devices may include tactile output devices (e.g., the touch screen 1710, or the tactile feedback of the joystick 1705, but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers 1709, headphones (not shown)), visual output devices (e.g., the screen 1710 including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each screen having or not having a touch screen input function, each screen having or not having a tactile feedback function, some of which are capable of outputting two-dimensional visual outputs or more than three-dimensional outputs through devices such as stereoscopic image outputs, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted).
[0145] The computer system 1700 may also include human-accessible storage devices and their associated media: for example, optical media including CD / DVD ROM / RW 1720 with media such as CD / DVD 1711, thumb drives 1722, removable hard disk drives or solid state drives 1723, conventional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.
[0146] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0147] The computer system 1700 may also include an interface 1799 to one or more communication networks 1798. The network 1798 may be, for example, a wireless network, a wired network, an optical network. The network 1798 may further be a local area network, a wide area network, a metropolitan area network, vehicle and industrial networks, a real-time network, a delay-tolerant network, etc. Examples of the network 1798 include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial television broadcasting, vehicle and industrial television including CANBus, and so on. Certain networks 1798 typically require an external network interface adapter (such as the USB port of the computer system 1700) connected to certain common data ports or peripheral buses (1750 and 1751); as described below, other network interfaces are typically integrated into the core of the computer system 1700 by connecting to the system bus (for example, an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system 1700 may communicate with other entities using any of these networks 1798. Such communication may be one-way reception only (for example, television broadcasting), one-way transmission only (for example, CANbus connected to certain CANbus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces described above.
[0148] The above-described human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core 1740 of the computer system 1700.
[0149] The kernel 1740 may include one or more central processing units (CPUs) 1741, a graphics processing unit (GPU) 1742, a graphics adapter 1717, a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) 1743, a hardware accelerator 1744 for certain tasks, etc. These devices, as well as a read-only memory (ROM) 1745, a random access memory 1746, and an internal mass storage 1747 such as an internal hard disk drive, SSD, etc. that is not user-accessible, may be connected via a system bus 1748. In some computer systems, the system bus 1748 may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly connected to or connected to the system bus 1748 of the kernel via a peripheral bus 1749. The architecture of the peripheral bus includes PCI, USB, etc.
[0150] The CPU 1741, GPU 1742, FPGA 1743, and accelerator 1744 may execute certain instructions, which may be combined to form the aforementioned computer code. The computer code may be stored in the ROM 1745 or the RAM 1746. Transitional data may also be stored in the RAM 1746, while permanent data may be stored, for example, in the internal mass storage 1747. Fast storage and retrieval of any storage device may be performed by using a cache, which may be closely associated with one or more CPUs 1741, GPUs 1742, mass storage 1747, ROM 1745, RAM 1746, etc.
[0151] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code that are specially designed and constructed for the purposes of the present disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the field of computer software.
[0152] As a non-limiting example, because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software included in one or more tangible computer-readable media, a computer system having architecture 1700, particularly having a core 1740, can provide functionality. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-transitory memories of the core 1740, such as the on-core mass memory 1747 or the ROM 1745. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 1740. Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core 1740, particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific parts of specific processes described herein, including defining data structures 1746 stored in the RAM and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., accelerator 1744), which can replace the software or operate in conjunction with the software to execute specific processes or specific parts of specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or include both. The present disclosure encompasses any suitable combination of hardware and software.
[0153] Although the present disclosure has described some exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.
Claims
1. A method for video decoding, characterized in that: The method is executed by at least one processor and includes: Obtaining a grid from a bitstream, the grid representing encoded volume data of at least one three-dimensional 3D visual content; and decoding the encoded volume data based on displacement vectors of vertices of the mesh, The displacement vector is based on a √3 subdivision in which the faces of the mesh are iteratively subdivided in such a way that the number of faces produced by subdividing the mesh in one iteration is less than four times the number of faces before the mesh is subdivided in the iteration.
2. The method for video decoding according to claim 1, characterized in that In the iteration, the √3 subdivision includes: inserting a vertex from among a plurality of inserted vertices into a face of the mesh before subdividing it in the iteration, and The plurality of inserted vertices are interconnected to form a plurality of subdivided faces of the mesh.
3. The video decoding method according to claim 2, characterized in that: The inserting one of the plurality of inserted vertices into the face before subdividing the mesh in the iteration comprises: Only one vertex of the plurality of inserted vertices is inserted into each face of the mesh before subdividing it in the iteration.
4. The video decoding method according to claim 3, characterized in that: At least one of the inserted vertices inserted in one of the faces before subdividing the mesh is based on a linear combination of coordinates of the vertices of the one face.
5. The method for video decoding according to claim 4, characterized in that: The linear combination includes: multiplying each coordinate of the vertex in the face by a corresponding non-negative real coefficient.
6. The method for video decoding according to claim 5, characterized in that: Each of the corresponding non-negative real coefficients is set to 1 / 3.
7. The method for video decoding according to claim 2, characterized in that: The √3 segmentation includes: Border edges of the mesh are subdivided differently than internal edges of the mesh such that at least two new inserted vertices are added to at least one of the border edges in the iteration.
8. A device for video decoding, characterized in that: The device comprises: at least one memory configured to store computer program code; At least one processor is configured to access the computer program code and operate according to the instructions of the computer program code, wherein the computer program code comprises: Obtaining code configured to cause the at least one processor to obtain a grid from a bitstream, the grid representing encoded volume data of at least one three-dimensional 3D visual content; and decoding code configured to cause the at least one processor to decode the encoded volume data based on displacement vectors of vertices of the mesh, The displacement vector is based on a √3 subdivision in which the faces of the mesh are iteratively subdivided in such a way that the number of faces produced by subdividing the mesh in one iteration is less than four times the number of faces before the mesh is subdivided in the iteration.
9. The apparatus for video decoding according to claim 8, characterized in that: In the iteration, the √3 subdivision includes: inserting a vertex from among a plurality of inserted vertices into a face of the mesh before subdividing it in the iteration, and The plurality of inserted vertices are interconnected to form a plurality of subdivided faces of the mesh.
10. The apparatus for video decoding according to claim 9, wherein: The inserting one of the plurality of inserted vertices into the face before subdividing the mesh in the iteration comprises: Only one vertex of the plurality of inserted vertices is inserted into each face of the mesh before subdividing it in the iteration.
11. The device for video decoding according to claim 10, characterized in that At least one of the inserted vertices inserted in one of the faces before subdividing the mesh is based on a linear combination of coordinates of the vertices of the one face.
12. The apparatus for video decoding according to claim 11, characterized in that The linear combination includes: Each coordinate of a vertex in the face is multiplied by a corresponding non-negative real coefficient.
13. The apparatus for video decoding according to claim 12, characterized in that Each of the corresponding non-negative real coefficients is set to 1 / 3.
14. The apparatus for video decoding according to claim 8, characterized in that The √3 segmentation includes: Border edges of the mesh are subdivided differently than internal edges of the mesh such that at least two new inserted vertices are added to at least one of the border edges in the iteration.
15. A non-transitory computer-readable medium, characterized in that A program is stored thereon, the program causing the computer to: Obtaining a grid from a bitstream, the grid representing encoded volume data of at least one three-dimensional 3D visual content; and decoding the encoded volume data based on displacement vectors of vertices of the mesh, The displacement vector is based on a √3 subdivision in which the faces of the mesh are iteratively subdivided in such a way that the number of faces produced by subdividing the mesh in one iteration is less than four times the number of faces before the mesh is subdivided in the iteration.
16. The non-transitory computer readable medium of claim 15, wherein: In the iteration, the √3 subdivision includes: inserting a vertex from among a plurality of inserted vertices into a face of the mesh before subdividing it in the iteration, and The plurality of inserted vertices are interconnected to form a plurality of subdivided faces of the mesh.
17. The non-transitory computer readable medium of claim 16, wherein: The inserting one of the plurality of inserted vertices into the face before subdividing the mesh in the iteration comprises: Only one vertex of the plurality of inserted vertices is inserted into each face of the mesh before subdividing it in the iteration.
18. The non-transitory computer readable medium of claim 17, wherein: At least one of the inserted vertices inserted in one of the faces before subdividing the mesh is based on a linear combination of coordinates of the vertices of the one face.
19. The non-transitory computer readable medium of claim 18, wherein: The linear combination includes: multiplying each coordinate of the vertex in the face by a corresponding non-negative real coefficient; Wherein, each of the corresponding non-negative real coefficients is set to 1 / 3.
20. The non-transitory computer readable medium of claim 15, wherein: The √3 subdivision includes subdividing border edges of the mesh differently than internal edges of the mesh, such that at least two new inserted vertices are added to at least one of the border edges in the iteration.