Video encoding and decoding method and device

By dividing and merging the vertices of the dynamic 3D mesh, the problem of difficulty in efficient compression of dynamic mesh in the existing technology is solved, and the reduction of the number of faces and the efficiency of data compression is achieved, and real-time applications are supported.

CN120050427APending Publication Date: 2025-05-27TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411561364.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-30
Filing Date
2024-11-04
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently compress dynamic 3D mesh containing time-change connectivity information and optional time-change attribute maps, especially in real-time applications.

Method used

By obtaining multiple vertices of the mesh, dividing them into groups, classifying the groups based on the orientation of the surface normal, and merging adjacent faces of common edges, thereby reducing the number of faces and decoding the encoded volume data based on these groups.

Benefits of technology

It effectively reduces the number of surfaces of the mesh, improves data compression efficiency, and supports real-time communication and other high-demand applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050427A_ABST
    Figure CN120050427A_ABST
Patent Text Reader

Abstract

The invention relates to a video encoding and decoding method and device. A method and apparatus including computer code configured to cause one or more processors to retrieve a grid from a code stream, the grid representing encoded volumetric data of at least one three-dimensional (3D) visual content; dividing a plurality of vertexes of the grid into a plurality of groups by determining a surface normal of each surface in the grid, classifying the plurality of groups according to orientations of the plurality of groups relative to the surface normal, and combining adjacent surfaces sharing edges among the plurality of groups; and decoding the encoded volume data based on the plurality of groups.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 603,020, filed on November 27, 2023, and U.S. Application 18 / 820,800, filed on August 30, 2024. The entire contents of the above - mentioned prior applications are incorporated herein by reference. Technical Field

[0003] This disclosure relates to reducing the number of faces in lossy mesh compression with and without subdivision and displacement encoding techniques. Background Art

[0004] Advances in three - dimensional (3D) capture, modeling, and rendering technologies have facilitated the popularity of 3D content across multiple platforms and devices. Today, the first steps of a baby can be captured on one continent and viewed (and perhaps even interacted with) by grandparents on another continent, enjoying a fully immersive experience with the child. However, to achieve this level of realism, models have become increasingly complex, and a large amount of data is associated with the creation and use of these models. 3D meshes are widely used to represent such immersive content.

[0005] A mesh consists of multiple polygons that describe the surface of a volumetric object. Each polygon is defined by its vertices in 3D space and information on how the vertices are connected, called connectivity information. Optionally, vertex attributes (such as color, normal, etc.) can be associated with mesh vertices. By leveraging mapping information that parameterizes the mesh into a 2D attribute map, attributes can also be associated with the surface of the mesh. This mapping is typically described by a set of parametric coordinates, called UV coordinates or texture coordinates, which are associated with mesh vertices. The 2D attribute map is used to store high - resolution attribute information such as texture, normal, displacement, etc. Such information can be used for various purposes, such as texture mapping and shading.

[0006] Dynamic mesh sequences may require a large amount of data because they may contain a large amount of information that changes over time. Therefore, efficient compression techniques are needed to store and transmit such content. MPEG has previously developed mesh compression standards such as IC, MESHGRID, and FAMC to handle dynamic meshes with constant connectivity and time-varying geometry and vertex attributes. However, these standards do not consider time-varying attribute graphs and connectivity information. Digital content creation (DCC) tools typically generate such dynamic meshes. On the other hand, it is challenging to generate dynamic meshes with constant connectivity using volume acquisition techniques, especially under real-time constraints. This type of content is not supported by existing standards. MPEG plans to develop a new mesh compression standard to directly handle dynamic meshes with time-varying connectivity information and an optional time-varying attribute graph. The standard targets lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, augmented reality (AR), and virtual reality (VR). Features such as random access and scalable / progressive coding are also taken into consideration.

[0007] Therefore, for any of the above reasons, it is desirable to have corresponding technical solutions to the problems that arise in video coding techniques. Summary of the Invention

[0008] This application includes a method and an apparatus, including a memory and at least one processor. The memory is configured to store computer program code, and the at least one processor is configured to access the computer program code and operate according to the instructions of the computer program code. The computer program is configured to cause the processor to implement: obtaining code, configured to cause the at least one processor to obtain a mesh from a bitstream, the mesh representing encoded volume data of at least one three-dimensional (3D) visual content; partitioning code, configured to cause the at least one processor to divide a plurality of vertices of the mesh into a plurality of groups by determining the face normals of each face in the mesh, classify the plurality of groups according to the orientation of the plurality of groups relative to the face normals, and merge adjacent faces sharing edges between the plurality of groups; and decoding code, configured to cause the at least one processor to decode the encoded volume data based on the plurality of groups.

[0009] Dividing the plurality of vertices of the mesh into a plurality of groups is further based on identifying faces sharing edges, and merging the faces sharing edges into a larger polygon face than before the faces are merged.

[0010] Decoding the encoded volume data based on the plurality of groups further includes splitting the merged adjacent faces into target m-sided polygon faces.

[0011] The target m-sided polygon face can be a triangular face.

[0012] Dividing the merged adjacent faces into target m-sided faces can be based on the decoded displacements.

[0013] At least one of the decoded displacements in the decoded displacements can be based on a non-zero displacement syntax that indicates the sum of values.

[0014] Dividing the merged adjacent faces into target m-sided faces can be further based on a non-zero flag nzFlag syntax. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0016] Figure 1 is a schematic diagram of a computer environment according to an embodiment;

[0017] Figure 2 is a simplified block diagram of media processing according to an embodiment;

[0018] Figure 3 is a simplified schematic diagram of decoding according to an embodiment;

[0019] Figure 4 is a simplified schematic diagram of encoding according to an embodiment;

[0020] Figure 5 is a simplified schematic diagram of media processing according to an embodiment;

[0021] Figure 6 is a simplified schematic diagram of media processing according to an embodiment;

[0022] Figure 7 is a simplified schematic diagram of media processing according to an embodiment;

[0023] Figure 8 is a simplified schematic diagram of grid features according to an embodiment;

[0024] Figure 9 is a simplified schematic diagram of grid features according to an embodiment;

[0025] Figure 10 is a simplified flowchart of media processing according to an embodiment;

[0026] Figure 11 is a simplified flowchart of media processing according to an embodiment;

[0027] Figure 12 is a simplified flowchart of media processing according to an embodiment;

[0028] Figure 13 is a simplified schematic diagram of grid features according to an embodiment;

[0029] Figure 14 is a simplified schematic diagram of the grid features according to an embodiment;

[0030] Figure 15 is a simplified schematic diagram of the grid features according to an embodiment;

[0031] Figure 16 is a simplified schematic diagram of the grid features according to an embodiment;

[0032] Figure 17 is a simplified schematic diagram of the grid features according to an embodiment; and

[0033] Figure 18 is a simplified diagram of the computer features according to an embodiment. Detailed implementation manners

[0034] The proposed features discussed below can be used alone or in any combination. In addition, the embodiments can be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-volatile computer-readable medium.

[0035] Figure 1 FIG. shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 may include at least two terminals 102 and 103 interconnected by a network 105. In a one-way data transmission, the first terminal 103 can encode video data at a local location for transmission to another terminal 102 through the network 105. The second terminal 102 can receive the encoded video data of another terminal from the network 105, decode the encoded data, and display the restored video data. One-way data transmission may be common in media service applications and the like.

[0036] Figure 1 FIG. shows a second pair of terminals 101 and 104 provided to support two-way transmission of encoded video that may occur, for example, during a video conference. In a two-way data transmission, each of the terminals 101 and 104 can encode video data captured at a local location for transmission to another terminal through the network 105. Each of the terminals 101 and 104 can also receive the encoded video data sent by another terminal, decode the encoded data, and display the restored video data on a local display device.

[0037] In Figure 1Among them, the terminals 101, 102, 103, and 104 may be shown as servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applied to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network 105 represents any number of networks for transmitting encoded video data among the terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. The communication network 105 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise explained below, the architecture and topology of the network 105 may be irrelevant to the operation of the present disclosure.

[0038] Figure 2 The placement of video encoders and decoders in a streaming environment is shown as an example of the application of the disclosed subject matter. The disclosed subject matter may equally apply to other video-supported applications, including, for example, video conferencing, digital television, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0039] The streaming system may include an acquisition subsystem 203, which may include a video source 201 such as a digital camera that creates, for example, an uncompressed video sample stream 213. Compared with the encoded video bitstream, this sample stream 213 can be emphasized as having a high data volume and can be processed by an encoder 202 coupled to the video source 201, and the encoder may be, for example, a camera as described above. The encoder 202 may include hardware, software, or a combination thereof to implement or carry out various aspects of the disclosed subject matter described in more detail below. Compared with this sample stream, the encoded video bitstream 204 can be emphasized as having a low data volume, and the encoded video bitstream 204 can be stored on the streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes an input copy of the encoded video bitstream 208 and creates an output video sample stream 210 that can be presented on a display 209 or other presentation device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video coding / compression standards. Examples of these standards are mentioned above and are further described herein.

[0040] Figure 3 It may be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.

[0041] The receiver 302 may receive one or more codec video sequences to be decoded by the decoder 300; in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel 301, which may be a hardware / software link to a storage device storing the encoded video data. The receiver 302 may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver 302 may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory 303 may be coupled between the receiver 302 and the entropy decoder / parser 304 (hereinafter referred to as "parser"). When the receiver 302 receives data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous synchronous network, it may also not be necessary to configure the buffer memory 303, or the buffer memory may be made smaller. For use on a best-effort traffic packet network such as the Internet, a buffer memory 303 may also be required, which may be relatively large and may advantageously have an adaptive size.

[0042] Video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy-coded video sequence. The categories of these symbols include information for managing the operation of decoder 300 and information that may be used to control a rendering device (e.g., display 312), which is not part of the decoder but may be coupled to the decoder. The control information for the rendering device may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not labeled). The parser 304 may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard and may follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract subgroup parameters for at least one subgroup among subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. Subgroups may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0043] The parser 304 may perform entropy decoding / parsing operations on the video sequence received from the buffer memory 303 to create symbols 313. The parser 304 may receive the encoded data and selectively decode specific symbols 313. In addition, the parser 304 may determine whether to provide a specific symbol 313 to the motion compensation prediction unit 306, scaler / inverse transform unit 305, intra prediction unit 307, or loop filter 311.

[0044] Depending on the type of the encoded video image or a part of the encoded video image (e.g., inter-frame image and intra-frame image, inter-frame block and intra-frame block) and other factors, the reconstruction of symbols 313 may involve multiple different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by the parser 304 from the encoded video sequence. For clarity, such subgroup control information flows between the parser 304 and the multiple units below are not described.

[0045] In addition to the functional blocks already mentioned, decoder 300 can conceptually be subdivided into a number of functional units as described below. In a practical implementation running under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.

[0046] The first unit is scaler / inverse transform unit 305. Scaler / inverse transform unit 305 receives, from parser 304, quantized transform coefficients as symbols 313 and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. Scaler / inverse transform unit 305 can output a block including sample values, which can be input into aggregator 310.

[0047] In some cases, the output samples of scaler / inverse transform unit 305 can belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previous reconstructed image, but can use prediction information from a previously reconstructed portion of the current image. Such prediction information can be provided by intra-image prediction unit 307. In some cases, intra-image prediction unit 307 uses the already reconstructed surrounding information extracted from the current (partially reconstructed) image 309 to generate a block having the same size and shape as the block being reconstructed. In some cases, aggregator 310 adds, on a per-sample basis, the prediction information generated by intra-prediction unit 307 to the output sample information provided by scaler / inverse transform unit 305.

[0048] In other cases, the output samples of scaler / inverse transform unit 305 can belong to an inter-coded and possibly motion-compensated block. In such a case, motion compensation prediction unit 306 can access reference image memory 308 to extract samples for prediction. After motion compensating the extracted samples according to symbols 313 related to the block, these samples can be added by aggregator 310 to the output of scaler / inverse transform unit (referred to as residual samples or residual signal in this case), thereby generating output sample information. The address in the reference image memory from which the motion compensation unit extracts the prediction samples can be controlled by a motion vector, and the motion vector is in the form of symbols 313 for use by the motion compensation unit, and symbols 313 can have, for example, X, Y, and reference image components. Motion compensation can also include interpolation of sample values extracted from the reference image memory when using sub-sample accurate motion vectors, a motion vector prediction mechanism, etc.

[0049] The output samples of aggregator 310 can be subjected to various loop filtering techniques in loop filter unit 311. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, and the parameters are available as symbols 313 from parser 304 to loop filter unit 311, and the video compression techniques can also respond to meta-information obtained during the decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and in response to previously reconstructed and loop-filtered sample values.

[0050] The output of loop filter unit 311 can be a sample stream that can be output to rendering device 312 and stored in reference picture buffer 308 for subsequent inter-picture prediction.

[0051] Once fully reconstructed, some encoded pictures can be used as reference pictures for subsequent prediction. Once an encoded picture is fully reconstructed and the encoded picture (e.g., by parser 304) is identified as a reference picture, the current reference picture 309 can become part of reference picture buffer 308, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0052] Video decoder 300 can perform decoding operations according to predetermined video compression techniques that can be recorded in a standard (e.g., ITU-T Recommendation H.265). In the sense that the encoded video sequence follows the syntax of the video compression technique or standard, the encoded video sequence can conform to the syntax specified by the video compression technique or standard used, as specified in the video compression technique document or standard and specifically in the profiles therein. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the levels of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, e.g., megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.

[0053] In one embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The video decoder 300 may use the additional data to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0054] Figure 4 May be a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.

[0055] The encoder 400 may receive video samples from a video source 401 (not part of the encoder), and the video source 401 may capture video images to be encoded by the encoder 400.

[0056] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder 303. The digital video sample stream may be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 YCrCb, RGB...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0057] According to one embodiment, the encoder 400 may encode and compress the pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls the other functional units as described below and is functionally coupled to the other functional units. For clarity, the couplings are not labeled in the figure. The parameters set by the controller may include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 402 as these functions may be related to the video encoder 400 optimized for a specific system design.

[0058] Some video encoders operate in a situation that is readily recognizable to those skilled in the art as an "encoding loop". In simplified terms, the encoding loop can include the encoding part of encoder 400 (hereinafter referred to as the "source encoder") (responsible for creating symbols based on the input image to be encoded and one or more reference images) and the (local) decoder 406 embedded in encoder 400, which reconstructs the symbols to create the sample data that the (remote) decoder will also create (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input into the reference image memory 405. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference image buffer is also bit-exact between the local encoder and the remote encoder. In other words, the reference image samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. The basic principle of this reference image synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.

[0059] The operation of the "local" decoder 406 can be the same as that of the "remote" decoder 300 already described in detail above. However, briefly referring also to Figure 3 Since the symbols are available and the entropy encoder 408 and the parser 304 can encode / decode the symbols into the encoded video sequence losslessly, the entropy decoding part of decoder 300 (including channel 301, receiver 302, buffer 303, and parser 304) may not be fully implementable in the local decoder 406. Figure 4

[0060] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also necessarily exist in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the fully described decoder technology. More detailed descriptions are only needed in certain places and are provided below.

[0061] As part of its operation, the source encoder 403 can perform motion-compensated predictive coding. Referring to one or more previously encoded frames designated as "reference frames" in the video sequence, this motion-compensated predictive coding performs predictive coding on the input frame. In this way, the coding engine 407 encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, which can be selected as the prediction reference for this input frame.

[0062] ​The local video decoder 406 can decode the encoded video data of the frames that can be designated as reference frames based on the symbols created by the source encoder 403. The operation of the encoding engine 407 can advantageously be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 4 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that can be performed by the video decoder on the reference frames and can cause the reconstructed reference frames to be stored in the reference image memory 405 (e.g., a cache memory). In this way, the encoder 400 can locally store a copy of the reconstructed reference frames, which has common content (without transmission errors) with the reconstructed reference frames that will be obtained by the remote video decoder.

[0063] The predictor 404 can perform a prediction search on the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 can search the reference image memory 405 for sample data (as candidate reference pixel blocks) or some metadata, such as reference image motion vectors, block shapes, etc., which can be used as a suitable prediction reference for the new image. The predictor 404 can operate on a sample block-by-pixel block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor 404, the input image can have prediction references taken from multiple reference images stored in the reference image memory 405.

[0064] The controller 402 can manage the encoding operations of the source encoder 403 (e.g., a video encoder), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0065] The outputs of all the above functional units can be entropy encoded in the entropy encoder 408. The entropy encoder performs lossless compression on the symbols generated by various functional units according to techniques well known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting these symbols into an encoded video sequence.

[0066] The transmitter 409 can buffer the encoded video sequence created by the entropy encoder 408, thereby preparing for transmission through the communication channel 411, which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter 409 can merge the encoded video data from the source encoder 403 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0067] The controller 402 can manage the operation of the encoder 400. During encoding, the controller 402 can assign a certain encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding image. For example, an image can generally be assigned to one of the following frame types:

[0068] An intra picture (I picture), which can be an image that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics.

[0069] A predictive picture (P picture), which can be an image that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0070] A bi - predictive picture (B picture), which can be an image that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference images and associated metadata for reconstructing a single block.

[0071] The source image can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment applied to the corresponding image of the block. For example, the blocks of an I picture can be non - predictively encoded, or these blocks can be predictively encoded with reference to the encoded blocks of the same picture (spatial prediction or intra prediction). The pixel blocks of a P picture can be non - predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference image. The blocks of a B picture can be non - predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference images.

[0072] The encoder 400 can perform encoding operations according to a predetermined video encoding technique or standard such as ITU - T Recommendation H.265. In operation, the encoder 400 can perform various compression operations, including predictive encoding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the used video encoding technique or standard.

[0073] In one embodiment, the transmitter 409 may transmit additional data when transmitting the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, SEI messages, VUI parameter set segments, etc.

[0074] Figure 5 FIG. 500 shows a simplified block - type flowchart of an exemplary view - point - dependent processing in the Omnidirectional Media Application Format (OMAF), which may allow 360 - degree virtual reality (VR360) streaming as described in OMAF.

[0075] At the acquisition block 501, in the case where the image data may represent a scene in VR360, video data A is acquired, such as data of multiple images and audio at the same time instance. At the processing block 503, the images Bi at the same time instance are subjected to one or more processes such as stitching, mapping to a projected image with respect to one or more virtual reality (VR) angles or other angles / viewpoints, and region packing. In addition, metadata may be created to indicate such processing information and other information to assist the transmission and presentation processes.

[0076] Regarding the data D, at the image encoding block 505, the projected image is encoded into data E i and combined into a media file. In viewport - independent streaming, at the video encoding block 504, the video images are encoded into data E v As, for example, a single - layer bitstream, regarding the data B a , the audio data may also be encoded into data Ea at the audio encoding block 502.

[0077] Data E a 、Data E v 、Data E i 、The entire encoded bitstream F iF and / or F can be stored on a (Content Delivery Network (CDN) / cloud) server and can generally be fully transmitted, for example, at distribution block 507 or otherwise, to the OMAF player 520 and can be fully decoded by a decoder such that at display block 516, at least one region of the decoded image corresponding to the current viewport is presented to the user with respect to various metadata, file playback, and orientation / viewport metadata, for example, the angle at which the user views through a VR image device, which is obtained from the head / eye tracking block 508 according to the viewport specification of the device. A notable feature of VR360 is that only one viewport can be displayed at any given time, and this feature can improve the performance of an omnidirectional video system through selective distribution according to the user's viewport (or any other criterion such as recommended viewport timing metadata). For example, according to an exemplary embodiment, tile-based video coding can enable viewport-dependent distribution.

[0078] Similar to the above encoding blocks, the OMAF player 520 according to an exemplary embodiment can similarly reverse one or more aspects of such encoding for file / fragment de-encapsulation of one or more of the data F' and / or F'i and metadata, decode the audio data E'i at the audio decoding block 510, decode the video data E'v at the video decoding block 513, and decode the image data E'i at the image decoding block 514 to continue the audio rendering of the data B'a at the audio rendering block 511 and the image rendering of the data D' at the image rendering block 515, thereby outputting the display data A'i at the display block 516 in VR360 format and outputting the audio data A's at the speaker / headphone block 512 according to various metadata (such as orientation / viewport metadata). Various metadata may affect the selection of the data decoding and rendering processes, and these selections may be made by or for the user of the OMAF player 520, depending on different tracks, languages, qualities, views, etc. It should be understood that the processing order described herein is for illustrative purposes of the exemplary embodiment, and in other exemplary embodiments, it may be implemented in other orders.

[0079] Figure 6 A simplified block diagram of a content stream process 600 of (encoded) point cloud data with view position and angle-dependent processing of six degrees of freedom media for acquisition / generation / encoding or decoding / rendering / display (hereinafter simply referred to as "V-PCC") is shown. It should be understood that the described features can be used alone or in combination in any order, and elements such as encoding and decoding can be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits), and according to an exemplary embodiment, the one or more processors can execute a program stored in a non-volatile computer-readable medium.

[0080] Figure 600 shows an exemplary embodiment of streaming encoded point cloud data according to V-PCC.

[0081] At the volume data acquisition block 601, a real-world visual scene or a computer-generated visual scene (or a combination of both) can be acquired by a set of camera devices or synthesized by a computer into volume data, and volume data in any format can be converted to (quantized) point cloud data format through image processing at the point cloud block 602. For example, according to an exemplary embodiment, by extracting one or more values described below from the volume data and its associated data and converting to the required point cloud format, the data in the volume data can be converted region-by-region data to points in the point cloud. According to an exemplary embodiment, the volume data can be a 3D data set of 2D images, such as slices, from which 2D projections of the 3D data set can be projected. According to an exemplary embodiment, the point cloud data format includes representations of data points in one or more various spaces and can be used to represent volume data, which can provide improvements in sampling and data compression, such as in terms of temporal redundancy. For example, at each point in a plurality of points of cloud data, point cloud data in the x, y, z format representing color values (e.g., RGB, etc.), luminance, intensity, etc., and can be used in conjunction with progressive decoding, polygon meshing, direct rendering, octree 3D representation of 2D quadtree data.

[0082] At the projection onto image block 603, the acquired point cloud data can be projected onto a 2D image, and the acquired point cloud data is encoded into an image / video image using video-based point cloud coding (V-PCC). The projected point cloud data can include attributes, geometry, occupancy maps, and other metadata for point cloud data reconstruction, such as the painter's algorithm, ray casting algorithm, (3D) binary space partitioning algorithm, etc.

[0083] On the other hand, at the scene generator block 609, the scene generator can generate some metadata to be used for rendering and displaying six-degree-of-freedom (DoF) media, for example, according to the director's intention or user preferences. Such 6DoF media can include 3D scene views similar to 360VR obtained from rotational changes on the 3D axes X, Y, Z, and in addition, there are additional dimensions that allow forward / backward, up / down, and left / right movement relative to a virtual experience within or at least according to the point cloud encoded data. The scene description metadata defines one or more scenes, which are composed of the encoded point cloud data and other media data including VR360, light field, audio, etc., and the scene description metadata can be provided to one or more cloud servers and / or file / fragment encapsulation / de-encapsulation processing, asFigure 6 as indicated in and related descriptions.

[0084] After video encoding block 604 and image encoding block 605 similar to the above video and image encoding (it can be understood that audio encoding can also be provided as described above), file / fragment encapsulation block 606 processes such that the encoded point cloud data is combined into a media file for file playback, or combined into a sequence of initialization fragments and media fragments for streaming according to a specific media container file format (e.g., one or more video container formats), and can be used for DASH, for example. These descriptions represent exemplary embodiments. The file container may also include scene description metadata, such as metadata from scene generator block 609, which is added to the file or fragment.

[0085] According to an exemplary embodiment, the file is encapsulated according to scene description metadata to include at least one view position in 6DoF media and at least one or more angular views at one or more time points at each of the one or more view positions, so that the file can be transmitted on demand according to user or creator input. Additionally, according to an exemplary embodiment, a fragment of such a file may include one or more parts of the file, such as a part of 6DoF media at a single viewpoint and angle at one or more time indications. However, these are merely exemplary embodiments and can vary according to various conditions such as network, user, creator capabilities, and input.

[0086] According to an exemplary embodiment, the point cloud data is divided into a plurality of 2D / 3D regions, which are independently encoded, for example, at one or more of video encoding block 604 and image encoding block 605. Then, each independently encoded partition of the point cloud data can be encapsulated as a track in the file and / or fragment at file / fragment encapsulation block 606. According to an exemplary embodiment, each point cloud track and / or metadata track may include some useful metadata for view position / angle-dependent processing.

[0087] According to an exemplary embodiment, metadata useful for view position / angle dependent processing, such as included in a file and / or a segment encapsulated in a file / segment encapsulation block, includes one or more of the following: layout information of a 2D / 3D partition with an index, (dynamic) mapping information associating a 3D volume partition with one or more 2D partitions (e.g., any tile / tile group / slice / sub-image), 3D position of each 3D partition in a 6DoF coordinate system, a list of representative view positions / angles, a list of selected view positions / angles corresponding to a 3D volume partition, an index of a 2D / 3D partition corresponding to a selected view position / angle, quality (grade) information of each 2D / 3D partition, and rendering information of each 2D / 3D partition according to each view position / angle, for example. Invoking such metadata upon request, such as by a user request of a V-PCC player or by a content creator for user guidance of a V-PCC player, can allow for more efficient processing of a specific portion of 6DoF media required by such metadata, so that the V-PCC player can provide a higher quality image for a focused portion of 6DoF media than other portions, rather than distributing unused portions of the media.

[0088] A file or one or more segments of a file can be directly distributed from the file / segment encapsulation block 606 to any one of the V-PCC player 625 and the cloud server through a distribution mechanism (e.g., through Dynamic Adaptive Streaming over HTTP (DASH) based on HTTP), such as at the cloud server block 607. The cloud server can extract one or more tracks and / or one or more specific 2D / 3D partitions from the file and can merge multiple encoded point cloud data into one data.

[0089] According to data such as the position / viewpoint tracking block 608, if a current viewing position and one or more angles are defined in a 6DoF coordinate system at the client system, view position / angle metadata can be distributed from the file / segment encapsulation block 606, or the view position / angle metadata can be additionally processed from a file or a segment already on the cloud server, such as at the cloud server block 607, so that the cloud server can extract one or more appropriate partitions from one or more stored files according to the metadata from a client system having, for example, a V-PCC player 625 and merge the partitions if necessary, and the extracted data can be distributed to the client as a file or a segment.

[0090] Regarding such data, at the file / fragment de-encapsulation block 615, the file de-encapsulator processes the file or received fragment, extracts the encoded bitstream, and parses the metadata. And at the video decoding block 610 and the image decoding block 611, the encoded point cloud data is then decoded into decoded point cloud data and reconstructed into point cloud data at the point cloud reconstruction block 612. And the reconstructed point cloud data can be displayed at the display block 614 and / or can first be synthesized at the scene synthesis block 613 according to one or more various scene descriptions of the scene description data in the scene generator block 609.

[0091] In view of the above, this exemplary V-PCC stream embodies advantages over the V-PCC standard, which include one or more of the following: the partitioning ability of multiple 2D / 3D regions, the ability to assemble the encoded 2D / 3D partitions into a single compliant encoded video bitstream in the compression domain, and the ability to extract the bitstream of the encoded 2D / 3D of the encoded image into a compliant encoded bitstream, where this V-PCC system support is further improved by forming a container including a VVC bitstream to support a mechanism for carrying one or more of the above metadata.

[0092] With this understanding and according to the exemplary embodiments described further below, the term "mesh" represents the composition of one or more polygons that describe the surface of a volumetric object. Each polygon is defined by its vertices in 3D space and information on how these vertices are connected, called connectivity information. Optionally, vertex attributes (such as color, normal, etc.) can be associated with the mesh vertices. By using mapping information that parameterizes the mesh into a 2D attribute map, attributes can also be associated with the surface of the mesh. Such a mapping is typically described by a set of parametric coordinates, called UV coordinates or texture coordinates, which are associated with the mesh vertices. The 2D attribute map is used to store high-resolution attribute information, such as texture, normal, displacement, etc. According to the exemplary embodiments, such information can be used for various purposes, such as texture mapping and shading.

[0093] However, a dynamic mesh sequence may require a large amount of data because it may contain a large amount of information that changes over time. For example, unlike a "static mesh" or "static mesh sequence" where the mesh information does not change from one frame to another, a "dynamic mesh" or "dynamic mesh sequence" represents the movement of vertices represented by the mesh changing from one frame to another. Therefore, effective compression techniques are needed to store and transmit such content. MPEG previously developed IC, MESHGRID, FAMC mesh compression standards to handle dynamic meshes with constant connectivity and time-varying geometry and vertex attributes. However, these standards did not consider time-varying attribute graphs and connectivity information. DCC tools typically generate such dynamic meshes. On the other hand, it is challenging to generate dynamic meshes with constant connectivity using volume acquisition techniques, especially under real-time constraints. This type of content is not supported by existing standards. According to an exemplary embodiment of the present disclosure, various aspects of a new mesh compression standard are described that can directly handle dynamic meshes with time-varying connectivity information and optionally time-varying attribute graphs, and the standard is for lossy and lossless compression for various applications such as real-time communication, storage, free viewpoint video, AR, and VR. Features such as random access and scalable / progressive coding are also taken into consideration.

[0094] Figure 7 An exemplary framework 700 for dynamic mesh compression is shown, such as a method based on 2D atlas sampling. Each frame of the input mesh 701 can be preprocessed through a series of operations such as tracing, remeshing, parameterization, voxelization. It should be noted that these operations can be limited to the encoder, that is, these operations may not be part of the decoding process, and this possibility can be signaled in the metadata through a flag, for example, 0 indicates limited to the encoder and 1 indicates otherwise. After preprocessing, a mesh with a 2D UV atlas 702 can be obtained, where each vertex of the mesh has one or more associated UV coordinates on the 2D atlas. Then, by sampling on the 2D atlas, the mesh can be converted into multiple graphs, including a geometry graph and an attribute graph. These 2D graphs can then be encoded by a video / image codec, such as HEVC, VVC, AV1, AVS3, etc. On the decoder 703 side, the mesh can be reconstructed from the decoded 2D graphs. Any post-processing and filtering can also be applied to the reconstructed mesh 704. It should be noted that for the purpose of 3D mesh reconstruction, other metadata can be signaled to the decoder side. It should be noted that the chart boundary information (including the uv and xyz coordinates of the boundary vertices) can be predicted, quantized, and entropy encoded in the bitstream. The quantization step size can be configured on the encoder side to trade off quality and bitrate.

[0095] In some implementations, a 3D mesh can be divided into several segments (or patches / charts). According to an exemplary embodiment, one or more 3D mesh segments can be considered as a "3D mesh". Each segment consists of a set of connected vertices associated with its geometric, property, and connectivity information. As Figure 8 shown in the example 800 of volumetric data in Figure 8 , the UV parameterization process 802 that maps from a 3D mesh segment to a 2D chart (such as to the above-mentioned 2D UV atlas tile 702) maps one or more mesh segments 801 onto a 2D chart 803 in a 2D UV atlas 804. Each vertex (vn) in the mesh segment will specify a 2D UV coordinate in the 2D UV atlas. It should be noted that the vertices (vn) in the 2D chart form a connected component as its 3D counterpart. The geometric, property, and connectivity information of each vertex can also be inherited from its 3D counterpart. For example, the information can indicate that vertex v4 is directly connected to vertices v0, v5, v1, and v3, and similarly for the information of each other vertex. Additionally, according to an exemplary embodiment, such a 2D texture mesh will further indicate information, such as color information, on a per-patch basis (such as each triangular patch, for example, v2, v5, v3 as one "patch").

[0096] For example, further regarding the features of the example 800 in Figure 8 , refer to the example 900 in Figure 9 where the 3D mesh segment 801 can also be mapped to multiple separate 2D charts 901 and 902. In this case, a vertex in 3D can correspond to multiple vertices in the 2D UV atlas. As Figure 9 shown, in the 2D UV atlas, the same 3D mesh segment is mapped to multiple 2D charts instead of a single chart as in Figure 8 . For example, 3D vertex v1 has two 2D correspondences v1 and v1', and 3D vertex v4 has two 2D correspondences v4 and v4'. Thus, a general 2D UV atlas of a 3D mesh can consist of multiple charts as shown in Figure 14 where each chart can contain multiple (usually greater than or equal to 3) vertices associated with its 3D geometric, property, and connectivity information.

[0097] Figure 9Example 903 is shown, which shows the derived triangulation in a chart having boundary vertices B0, boundary vertex B1, boundary vertex B2, boundary vertex B3, boundary vertex B4, boundary vertex B5, boundary vertex B6, and boundary vertex B7. When presenting such information, any triangulation method can be applied to create the connectivity between vertices (including boundary vertices and sampled vertices). For example, for each vertex, find the two closest vertices. Or for all vertices, continuously generate triangles until the minimum number of triangles is reached after a set number of attempts. As shown in Example 903, there are various regularly shaped repeating triangles and various oddly shaped triangles that are generally closest to the boundary vertices, each with its unique dimensions, which may or may not be shared with any other triangle. The connectivity information can also be reconstructed through explicit signaling. According to an exemplary embodiment, if the polygon cannot be recovered by implicit rules, the encoder can signal the connectivity information in the bitstream.

[0098] The boundary vertices B0, boundary vertex B1, boundary vertex B2, boundary vertex B3, boundary vertex B4, B boundary vertex 5, boundary vertex B6, and boundary vertex B7 are defined in the 2D UV space. The boundary edges can be determined by checking whether an edge appears in only one triangle. According to an exemplary embodiment, the following information of the boundary vertices is important and should be signaled in the bitstream: geometric information (e.g., 3D XYZ coordinates even if currently in 2D UV parameter form), and 2D UV coordinates.

[0099] For the case where the boundary vertices in 3D correspond to multiple vertices in the 2D UV atlas, as Figure 9 shown, the mapping from 3D XYZ to 2D UV can be one-to-many. Therefore, the UV-to-XYZ (or called UV2XYZ) index can be signaled to indicate the mapping function. UV2XYZ can be a 1D index array that maps each 2D UV vertex to a 3D XYZ vertex.

[0100] According to an exemplary embodiment, to efficiently represent the mesh signal, a subset of the mesh vertices and the connectivity information between them can be encoded first. In the original mesh, the connections between these vertices may not exist because the connections between these vertices are subsampled from the original mesh. There are different ways to signal the connectivity information between these vertices, so such a subset is called the base mesh or base vertices.

[0101] According to an exemplary embodiment, many methods for dynamic mesh compression are implemented, which are part of the above-mentioned edge-based vertex prediction framework, where the base mesh is first encoded, and then more additional vertices are predicted based on the connectivity information of the edges from the base mesh. It should be noted that these methods can be applied individually or in any form of combination.

[0102] For example, consider Figure 10 Exemplary flowchart 1001 of vertex grouping for prediction mode. At step S101, vertices within the mesh can be obtained, and at step S102, the vertices within the mesh can be divided into different groups for prediction purposes, for example, see Figure 9 . In one example, the division is completed using patch / chart partitioning at step S104. In another example, at step S105, the division is completed under each patch / chart. Whether step S103 proceeds to step S104 or step S105 can be signaled by a flag or the like. In the case of step S105, several vertices of the same patch / chart form a prediction group and will share the same prediction mode, while several other vertices of the same patch / chart can use another prediction mode. Among them, the "prediction mode" can be considered as a specific mode used by the decoder to predict video content including the patch. The prediction mode can be divided into two major categories: intra prediction mode and inter prediction mode. Within each category, the decoder can select different specific modes. According to an exemplary embodiment, each group, that is, the "prediction group", can share the same specific mode (for example, an angular mode at a specific angle) or the same category of prediction mode (for example, all intra prediction modes, but can be predicted at different angles) according to an exemplary embodiment. At step S106, this grouping can be assigned at different levels by determining the corresponding number of vertices involved in each group. For example, according to an exemplary embodiment, every 64, 32, or 16 vertices in the patch / chart in scan order will be assigned the same prediction mode, while other vertices may be assigned differently. For each group, the prediction mode can be an intra prediction mode or an inter prediction mode. This can be signaled or assigned. According to exemplary flowchart 1000, at step S107, if it is determined that the mesh frame or mesh slice is of intra type, for example, by checking whether the flag of the mesh frame or mesh slice indicates intra type, then all vertex groups within the mesh frame or mesh slice should use the intra prediction mode; otherwise, at step S108, an intra prediction or inter prediction mode can be selected for each group and applied to all vertices within the group.

[0103] In addition, for a group of mesh vertices using an intra prediction mode, its vertices can only be predicted by using previously encoded vertices within the same sub-partition of the current mesh. Sometimes, according to an exemplary embodiment, the sub-partition can be the current mesh itself. According to an exemplary embodiment, for a group of mesh vertices using an inter prediction mode, its vertices can only be predicted by using previously encoded vertices of another mesh frame. Each of the above pieces of information can be determined and signaled by a flag or the like. The prediction feature can occur at step S110, and the result of the prediction and signaling can occur at step S111.

[0104] According to an exemplary embodiment, in exemplary flowchart 1000 and flowchart 1100 described below, for each vertex in a group of vertices, after prediction, the residual will be a 3D displacement vector representing the offset from the current vertex to its predicted value. The residuals of the group of vertices need to be further compressed. In one example, before entropy coding, the transformation and signaling at step S111 can be applied to the residuals of the group of vertices. The following methods can be used to process the encoding of a group of displacement vectors. For example, in one method, the cases where a group of displacement vectors, some displacement vectors, or their components have only zero values are signaled appropriately. In another embodiment, a flag is signaled for each displacement vector to indicate whether the vector has any non-zero components, and if not, the encoding of all components of the displacement vector can be skipped. In addition, in another embodiment, a flag is signaled for each group of displacement vectors to indicate whether the group has any non-zero vectors, and if not, the encoding of all displacement vectors in the group can be skipped. In addition, in another embodiment, a flag is signaled for each component of each group of displacement vectors to indicate whether the component of the group has any non-zero vectors, and if not, the encoding of that component of all displacement vectors in the group can be skipped. In addition, in another embodiment, there can be signaling for whether a group of displacement vectors or a group of components of displacement vectors needs to be transformed. If not, the transformation can be skipped, and quantization / entropy coding can be directly applied to the group or the components of the group. In addition, in another embodiment, a flag can be signaled for each group of displacement vectors to indicate whether the group needs to be transformed. If not, the transformation encoding of all displacement vectors in the group can be skipped. In addition, in another embodiment, a flag is signaled for each component of each group of displacement vectors to indicate whether the component of the group needs to be transformed. If not, the transformation encoding of that component of all displacement vectors in the group can be skipped. The embodiments described in this paragraph regarding processing vertex prediction residuals can also be combined and implemented separately and in parallel on different patches.

[0105] Figure 11An exemplary flowchart 1100 is shown, where, at step S121, a mesh frame encoded as an entire data unit can be obtained, which means that there can be associations between all vertices or attributes of the mesh frame. Alternatively, according to the determination result at step S122, the mesh frame can be divided into smaller independent sub - partitions at step S123, where the sub - partitions are conceptually similar to slices or tiles in 2D video or images. At step S124, a prediction type can be assigned to the encoded mesh frame or the encoded mesh sub - partition. Possible prediction types include intra - frame coding type and inter - frame coding type. For the intra - frame coding type, at step S125, prediction is only allowed from the reconstructed parts of the same frame or slice. On the other hand, in addition to intra - mesh - frame prediction, the inter - frame prediction type will allow prediction from previously encoded mesh frames at step S125. Furthermore, the inter - frame prediction type can be classified into more subtypes, such as P - type or B - type. In the P - type, only one predictor can be used for prediction purposes, while in the B - type, two predictors from two previously encoded mesh frames can be used to generate a predictor. The weighted average of the two predictors can be taken as an example. When the mesh frame is encoded as a whole, the frame can be regarded as an intra - frame or inter - frame encoded mesh frame. In the case of an inter - frame mesh frame, the P - type or B - type can be further identified by signaling. Or, if the mesh frame is encoded after being further segmented intra - frame, a prediction type is assigned to each sub - partition at step S124. The above - mentioned information can be determined and signaled by flags, etc., similar to Figure 10 steps S110 and S111, the prediction feature can occur at step S126, and the results of the prediction and signaling can occur at step S127.

[0106] Therefore, although a dynamic mesh sequence may require a large amount of data because it may contain a large amount of information that changes over time, an efficient compression technique is needed to store and transmit such content, and the features described herein improve the 3D position prediction of mesh vertices by allowing the use of previously decoded vertices (intra - frame prediction) in the same mesh frame or from previously encoded mesh frames (inter - frame prediction), thus representing this increased efficiency.

[0107] In addition, an exemplary embodiment can generate a displacement vector for the third layer 1303 of the mesh based on one or more reconstructed vertices of one or more previous layers (such as the second layer 1302 and the first layer 1301). Assuming the index of the second layer 1302 is T, a predictor for the vertices of the third layer 1303T + 1 is generated based on at least the reconstructed vertices of the current layer or the second layer 1302. An example of such a layer - based prediction structure is as Figure 13As shown in Example 1300, this example illustrates vertex prediction based on reconstruction: progressive vertex prediction using edge-based interpolation, where the predictor is generated based on previously decoded vertices rather than predictor vertices. The first layer 1301 can be a mesh bounded by the first polygon 1340, and the vertices of the first polygon 1340 are decoded vertices at the boundary and interpolated vertices along the line between those decoded vertices. When progressive encoding proceeds from the first layer 1301 to the second layer 1302, an additional polygon 1341 can be formed by the displacement vectors from the interpolated vertices of the first layer to the additional vertices of the second layer 1302. Thus, the total number of vertices in the second layer 1302 may be greater than the total number of vertices in the first layer 1301. Similarly, when proceeding to the third layer 1303, the additional vertices of the second layer 1302 together with the decoded vertices from the first layer 1301 can serve the encoding in a manner similar to the decoded vertices served when proceeding from the first layer 1301 to the second layer 1302; that is, multiple additional polygons can be formed. Note that, refer to Figure 14 Example 1400 showing such progressive encoding, different from Figure 13 the situation in, Example 1400 shows that when proceeding from the first layer 1401 to the second layer 1402 and then to the third layer 1403, each additional formed polygon can be entirely within the polygon formed by the boundary of the first layer 1401.

[0108] For Example 1300 and / or Example 1400, according to an exemplary embodiment, refer to Figure 12 Exemplary flowchart 1200 in. Since the interpolated vertices on the current layer are predicted values, these predicted values need to be reconstructed before the predictor for generating the vertices of the next layer. This is done by encoding the base mesh in step S131, implementing vertex prediction in step S132, and then adding the decoded displacement vectors of the current layer to the predictor of the vertices (e.g., of layer 1302) in step S133. Then, the reconstructed vertices of this layer together with all the decoded vertices of one or more previous layers (such as checking the additional vertex values of these layers at step S134) can be used to generate and signal the predicted vertices of the next layer 1303, as shown in step S135. This process can also be summarized as follows: Let P[t](Vi) denote the predicted value of vertex Vi on layer t; Let R[t](Vi) denote the reconstructed vertex Vi on layer t; Let D[t](Vi) denote the displacement vector of vertex Vi on layer t; Let f(*) denote the predicted value generation function, specifically, it can be the average of two existing vertices. Then according to an exemplary embodiment, for each layer t, there is the following equation:

[0109] P[t](Vi) = f(R[s|s<t](Vj),R[m|m<t](Vk)), where

[0110] Vj and Vk are reconstructed vertices of the previous layer

[0111] R[t](Vi) = P[t](Vi) + D[t](Vi) Equation (1)

[0112] Then, all vertices in a mesh frame are divided into layer 0 (base mesh), layer 1, layer 2, …. Then, the reconstruction of vertices on one layer depends on the reconstruction of vertices on one or more previous layers. In the above, each of P, R, and D represents a 3D vector in the 3D mesh representation. D is the decoded displacement vector, and quantization may or may not be applied to this vector.

[0113] According to an exemplary embodiment, vertex prediction using reconstructed vertices can be applied only to certain layers. For example, layer 0 and layer 1. For other layers, vertex prediction can still use adjacent predicted vertices without adding a displacement vector for reconstruction. This allows these other layers to be processed simultaneously without waiting for a previous layer to be reconstructed. According to an exemplary embodiment, for each layer, it can be signaled whether to select vertex prediction based on reconstructed vertices or vertex prediction based on predicted values, or it can be signaled which layers (and subsequent layers) do not use vertex prediction based on reconstructed vertices.

[0114] For those displacement vectors whose vertex prediction values are generated by reconstructed vertices, quantization can be applied to them without further performing a transform, such as a wavelet transform, etc. For those displacement vectors whose vertex prediction values are generated by other predicted vertices, a transform may be required, and quantization can be applied to the transform coefficients of those displacement vectors.

[0115] Therefore, a dynamic mesh sequence may require a large amount of data because a dynamic mesh sequence may contain a large amount of information that changes over time, so an effective compression technique is needed to store and transmit such content. In the framework of the above interpolation-based vertex prediction method, an important process is to compress the displacement vector, which occupies the main part in the encoded bitstream and is also the focus of the present disclosure. The present disclosure alleviates this problem by providing such compression.

[0116] In addition, similar to the above other examples, even in those embodiments, a dynamic mesh sequence may still require a large amount of data because a dynamic mesh sequence may contain a large amount of information that changes over time, so an effective compression technique is needed to store and transmit such content. In the framework of the above 2D atlas sampling method, important advantages can be obtained by inferring connectivity information from the sampled vertices plus boundary vertices on the decoder side. This is the main part in the decoding process and is also the focus of other examples described below.

[0117] According to an exemplary embodiment, the connectivity information of the base mesh can be inferred (derived) from the decoded boundary vertices and sampled vertices of each graph on both the encoder and decoder sides.

[0118] Similarly as described above, any triangulation method can be applied to create the connectivity between vertices (including boundary vertices and sampled vertices). According to an exemplary embodiment, the connectivity type can be signaled in the high-level syntax, such as in the sequence header, slice header.

[0119] As described above, the connectivity information can also be reconstructed by explicit signaling, such as for an irregular-shaped triangular mesh. That is, if it is determined that the polygon cannot be recovered by implicit rules, the encoder can signal the connectivity information in the bitstream. According to an exemplary embodiment, the overhead of such explicit signaling can be reduced based on the boundary of the polygon.

[0120] According to an embodiment, only the connectivity information between the boundary vertices and the sampling positions is signaled, while the connectivity information between the sampling positions themselves is inferred.

[0121] Furthermore, in any embodiment, the connectivity information can be signaled by prediction, such that only the difference in the inferred connectivity (as a prediction) from one mesh to another is signaled in the bitstream.

[0122] It should be noted that, according to an exemplary embodiment, the orientation of the inferred triangles (such as each triangle being inferred in a clockwise or counterclockwise manner) can be signaled in the high-level syntax (such as sequence header, slice header, etc.) for all graphs, or the orientation of the inferred triangles can be fixed (assumed) by the encoder and decoder. For each graph, the orientation of the inferred triangles can also be signaled differently.

[0123] Furthermore, any reconstructed mesh can have a different connectivity from the original mesh. For example, the original mesh can be a triangular mesh, while the reconstructed mesh can be a polygon mesh (e.g., a quadrilateral mesh).

[0124] According to an exemplary embodiment, the connectivity information of any base vertices can be not signaled, but instead the edges between the base vertices can be derived using the same algorithm on both the encoder and decoder sides. According to an exemplary embodiment, the interpolation of the predicted vertices of the additional mesh vertices can be based on the derived edges of the base mesh.

[0125] According to an exemplary embodiment, a flag can be used to indicate whether the connectivity information of the base vertices is signaled or needs to be derived, and this flag can be signaled at different levels of the bitstream (such as at the sequence level, frame level, etc.).

[0126] According to an exemplary embodiment, first, the same algorithm is used on both sides of the encoder and the decoder to derive the edges between the base vertices. Then, it is compared with the original connectivity of the base mesh vertices, and the difference between the derived edges and the actual edges will be signaled. Therefore, after decoding this difference, the original connectivity of the base vertices can be restored.

[0127] In one example, for the derived edges, if they are determined to be incorrect compared to the original edges, this information can be signaled in the bitstream (by indicating the pair of vertices forming the edge); for the original edges, if they are not derived, it can be signaled in the bitstream (by indicating the pair of vertices forming the edge). Additionally, the connectivity on the boundary edges and the vertex interpolation involving the boundary edges can be performed separately from the internal vertices and internal edges.

[0128] Therefore, through the exemplary embodiments described herein, the above technical problems can be advantageously improved by one or more technical solutions. For example, since a dynamic mesh sequence may require a large amount of data because it may contain a large amount of information that changes over time, the exemplary embodiments described herein at least represent an efficient compression technique for storing and transmitting such content.

[0129] The above embodiments can further be applied to instance-based mesh coding, where the instance can be a mesh of an object or a mesh of a part of an object. For example, Figure 15 Illustrative example 1500 shows mesh example 1501, where there are different instances 1502 (mesh representing a cup), 1503 (mesh representing a spoon), and 1504 (mesh representing a plate), and these instances can be separated and encoded separately. Each of instances 1501, 1502, 1503, and 1504 is shown within its respective bounding box, but it should be noted that instance 1501 can be considered to be bounded by a "mesh-based bounding box", while each of instances 1502, 1503, and 1504 can be considered to be bounded by its respective "instance-based bounding box".

[0130] Furthermore, looking at Figure 16Examples 1600 and the other figures. According to an embodiment, considering that "bidegree" mesh encoding is a specific technique aimed at efficiently encoding the connectivity of a polygon mesh. By applying the mathematical principle of duality, this method encodes the connectivity data by constructing two separate sequences: one sequence characterizes the degrees of vertices, and the other sequence depicts the degrees of faces (see the illustration of Example 1601, which is an example illustration of a bidegree traversal). That is to say, the encoding process involves a simultaneous traversal around both faces and vertices. Specifically, the traversal starts from an arbitrary seed face, and the degree of this face is recorded (for example, F3 represents the valence of the vertex and the degree of the face). Subsequently, as shown in Example 1601, the degrees of adjacent vertices are also recorded (for example, V5, V5, V5). Then, the algorithm selects a so-called "pivot vertex", which is identified by having the minimum degree of freedom - representing the count of non-traversed adjacent faces. The traversal continues around this pivot vertex, adding new faces and vertices and recording their respective degrees. To handle unique scenarios involving vertex splitting or merging, supplementary symbols are adopted according to an embodiment. Although this design highlights the excellent efficiency of bidegree mesh encoding in polygon meshes, including those with high irregularity or worst-case polygon meshes, the performance of bidegree encoding depends to a large extent on the regularity of face degrees and vertex valences.

[0131] In terms of position attribute encoding, in addition to connectivity, a mesh typically includes other attributes, such as vertex positions, texture coordinates, normal vectors, and associated texture maps. Therefore, considering that connectivity usually accounts for only a small part of the mesh data, it may contribute less than 10% or even 1% of the total bitstream, while position attributes may contribute more than half of the bitstream.

[0132] Regarding the position prediction using parallelograms, see Example 1602, where among the attributes, the 3D positions of vertices usually account for the majority of the bits required for geometric attributes. To efficiently encode vertex positions, a predictive coding scheme is adopted according to an embodiment, and a prominent example is parallelogram prediction. In the context of polygon meshes, it has been found that parallelogram prediction performs best in quadrilateral meshes.

[0133] That is to say, Example 1602 represents an illustration of "cross" parallelogram prediction and parallelogram "inner" prediction for vertex encoding in a polygon mesh, where A, B, and C are three reference positions for these two types of parallelogram predictions. According to an embodiment, for a given polygon mesh, as shown in Example 1602, the position (V) of the vertex to be predicted uses three previously encoded vertices (A, B, C) as references to estimate its position (V) using the following equation:

[0134] V = w 1 A + w 2 B + w 3 C Equation (2)

[0135] Among them, the weighting factor is usually selected as w 1 = w 3 = 1, w 2 = -1. These weights can be further adjusted to adapt to different polygon structures. In the context of a polygon mesh, parallelogram prediction can be divided into two types: "intra-prediction" and "cross-prediction", as shown in Example 1602. In "intra-prediction", all three reference vertices are on the same face, while in "cross-prediction", vertex C is obtained from the opposite face.

[0136] According to an embodiment, refer to Example 1603 regarding position prediction using multiple parallelograms. That is, as an extension of parallelogram prediction, a method called multi-parallelogram prediction has been introduced to enhance the position prediction of a triangle mesh. As shown in Example 1603, multi-parallelogram prediction "a" adopts the average position derived from two or more parallelogram predictions when feasible. Prediction "b" shows an instance of double-parallelogram prediction. The number of available prediction candidates determines how many predicted values are averaged.

[0137] These proposed methods can be used alone or in any combination in any order. In addition, these proposed methods can be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). Although some methods are explicitly applied to a triangle mesh, they are also applicable to any polygon mesh. In the present disclosure, a method is proposed for adaptively merging triangular faces into quadrilateral faces in a traversal order to create an intra-prediction of parallelogram prediction.

[0138] For example, Example 1604 shows adaptive parallelogram prediction, where, in the context of a triangle mesh, as in Example 1604, following the traversal order of pivot vertex A, the next face to be traversed is F0, and the vertex to be encoded is D.

[0139] According to these embodiments of Example 1604, parallelogram prediction uses reference vertex A, reference vertex B, and reference vertex C 0 to predict D. Although vertex A and vertex B are on the same face (F 0 ) as D, vertex C 0 is selected from the opposite face (F 1 ) on edge A - B, as shown by C 1 in Example 1604. However, in many cases, this prediction may not be accurate enough, especially if the quadrilateral face formed by vertices A, C, B, and D is concave, or if the position of vertex D is irregular.

[0140] To solve this problem, an alternative reference vertex C is provided in the embodiments of the present disclosure 1 , where C 1 is the vertex connected to A in the face opposite to F 1 on the edge A-C 1 (i.e., the face F in Example 1604 2 ). In this case, the encoder can select C 0 or C 1 as a reference to predict D

[0141] In one embodiment, the selection criterion for the reference C in Example 1604 is based on the sum of the absolute differences (Sum of Absolute Differences, SAD) of the residual vector between the predicted position P i = A + B - C i and the position to be encoded, calculated as R i = C i - P i , where i = 0, 1

[0142] In another embodiment, regarding Example 1604, the selection between C 0 or C 1 is based on the estimated bits required for encoding the residual, called the bit cost (R i ). This cost calculation also includes the bits for signaling the adaptive flag

[0143] In another embodiment, regarding Example 1604, an additional reference vertex C 0 is selected from the face opposite to F 1 on the edge B-C 2

[0144] In yet another embodiment, regarding Example 1604, the number of reference vertices for predicting D is extended to include all adjacent vertices of A and B, which increases the signaling bits

[0145] In yet another embodiment, regarding Example 1604, the number of reference vertices for predicting D is extended to include the first N available adjacent vertices of A and B, where N can be specified at the header syntax (such as the frame-level or sequence-level header), or can be assumed by both the encoder and the decoder. Assuming that all available adjacent vertices are greater than the number N, the first N vertices are selected as candidates in the given order to reduce the signaling cost

[0146] The present disclosure further provides signaling for adaptive reference selection, including using a flag to signal the adaptive reference and hiding the adaptive reference flag by modifying the degree of the face

[0147] ​According to an embodiment, for signaling an adaptive reference using a flag, to indicate which reference vertex C (i.e., C 0 or C 1 ) is used, a straightforward approach is to signal a binary flag for each face. Specifically, each face will have a bit to indicate whether to use vertex C 0 (flag value 0) or vertex C 1 (flag value 1). According to an embodiment, binary arithmetic coding is used to encode these flags. Also, to further optimize at least the entropy coding of the adaptive reference flags, the flag context of the faces adjacent to F1 is used to estimate the probability of the current adaptive reference flag.

[0148] For hiding the adaptive reference flags by modifying the face degree, given that C 0 is typically preferred in vertex prediction, signaling the adaptive reference flags may incur a significant bit overhead. To address this issue, embodiments of the present disclosure employ a simple face merging algorithm to hide the adaptive reference flags. That is, for example, for a triangular mesh, the face degree is always 3. On the encoder side, the traversal follows the double-degree connectivity coding, and the face degree is modified to 3 + i only when i > 0, where i represents the corresponding prediction mode. On the decoder side, the vertex predictor will select the corresponding reference based on the current face degree.

[0149] Furthermore, for hiding the adaptive reference flags by modifying the face degree, embodiments of the present disclosure can limit each face to have only one reference flag to ensure decodability. By doing so, these embodiments avoid signaling the most common case (i.e., i = 0) and hide other adaptive reference flags within the face symbol of the double-degree connectivity coding. Thus, although this results in an increase in the bits for connectivity, it saves the bits for position coding. According to an embodiment, in cases where more than two reference modes need to be handled, the face degree can be increased accordingly. The less preferred reference selections are added at a higher cost. For example, if the reference candidate list includes C 0 , C 1 , and C 2 , the degrees will be incremented by +0, +1, and +2 respectively.

[0150] VMesh is an ongoing MPEG standard for compressing dynamic meshes. According to an embodiment, see Figure 17Example 1700 shows the application of the current VMesh reference software, which compresses a mesh based on a reduced base mesh (encoded by Draco), displacement vectors, and a motion field (if applicable). The displacement is calculated by searching for the nearest point on the input mesh for each vertex of the subdivision-based mesh. To encode the displacement, the displacement vectors are transformed into wavelet coefficients through a linear lifting scheme, and then these coefficients are quantized and encoded through a video codec or an arithmetic codec. This process also refines the base mesh to minimize the displacement. According to an embodiment, texture transfer is performed to match the texture with the reparameterized geometry and UVs, as well as an optimized texture for image compression.

[0151] According to an embodiment, the displacement is the residual vector between the interpolated midpoint and the corresponding position on the surface mesh, expressed as:

[0152] d 0 = pos - pos 0 Equation (3)

[0153] Then, the displacement is decorrelated according to the surface normal n, the tangent vector t, and the binormal vector b.

[0154] d = [d 0 *n, d 0 *t, d 0 *b] Equation (4)

[0155] Then, the displacement (disp) vector is quantized with the corresponding quantizers [lqp 1 , lqp 2 , lqp 3 , where the quantization generally satisfies lqp 0 ≥ lqp 0 = lqp 0 .

[0156]

[0157] where, θ i represents the quantization offset, which is typically set to 1 / 3.

[0158] At the decoder, the displacement is decoded as:

[0159]

[0160] Then, it is normalized back to the original coordinates:

[0161]

[0162] On the other hand, the displacement is encoded by signal-encoding the first non-zero displacement and then performing entropy encoding, separately for each dimension.

[0163] However, even with these techniques, there are still problems when using displacement encoding in combination with subdivision, which may lead to a significant increase in the number of faces, and thus an increase in rendering complexity. Moreover, for flat surfaces, subdivision is not required as it does not enhance any geometric details. Additionally, the rendering technique is based on triangular meshes.

[0164] The disclosed methods can be used independently or in combination with each other in any order to generate any displacement. The disclosed methods can be applied to any polygon mesh, not limited to triangular meshes. In this article, the number of triangles is used only as an example.

[0165] For example, a method for reducing the number of faces on the decoder side is provided. This embodiment introduces a bottom-up strategy to minimize the number of faces on the decoder side. This method is applicable regardless of whether subdivision and displacement encoding are used. The two-step method is as follows:

[0166] In the first step, merge the faces on the flat surface. This merging involves combining adjacent faces that share an edge and have the same face normal into a single polygon face. The process is as follows: (1) Determine the face normal of each face: Calculate the face normal of each face in the input mesh. (2) Group the faces with the same normal orientation: Classify the faces into groups according to the orientation of the normal. (3) Merge the adjacent faces that share the same edge:

[0167] - Identify the adjacent faces that share a common edge in the same group.

[0168] - Merge these adjacent faces to form a larger polygon face.

[0169] In the second step, after the first step, divide the merged faces into target m-sided faces. In this step, divide the higher-order polygon faces merged in the first step into target m-sided faces. Taking a triangular mesh as an example, dividing an N-sided face into triangles will generate N - 2 triangular faces. In the worst case, the number of triangles after division is the same as the number before merging.

[0170] According to the embodiment, a top-down face number reduction strategy for mesh encoding using subdivision and displacement encoding techniques is also disclosed. The decision of whether to further subdivide a given face is made adaptively on the decoder side based on the decoded displacement.

[0171] Such embodiments assume that the encoder uses n levels of subdivision and displacement encoding for each base face, where a "base face" refers to the initial input face. The criterion for determining whether an additional level of subdivision is required depends on the total energy of all displacement vectors in the subsequent levels. If the displacement energy is zero, no further subdivision is required. The detailed pseudocode is given in Table 1 below:

[0172] Table 1

[0173]

[0174] According to an exemplary embodiment, a method for reducing the number of encoder sides is also provided. Such an embodiment provides a method for reducing the complexity and the number of sides of a decoder by signaling an adaptive subdivision level. When a targeted subdivision level for a given side is received, the decoder terminates the subdivision process accordingly in advance.

[0175] Initially, an adaptive subdivision level is determined for each base side at the encoder. In one instance, similar to how using displacement coding in combination with subdivision can lead to a significant increase in the number of sides and thus an increase in rendering complexity, these adaptive subdivision levels are determined based on the total energy of subsequence displacements. If the total energy of all subsequent displacements from level k + 1 to n max is zero, then the side has a subdivision level k < n max . According to an exemplary embodiment, the method for signaling the adaptive subdivision level is shown in Table 2 below:

[0176] Table 2

[0177]

[0178] where CABAC is context adaptive arithmetic coding. Subdiv_level represents the level of detail of the current side. In one embodiment, the optimal group size (groupSize) is estimated and signaled to the decoder. In one embodiment, the method is further combined with using displacement coding in combination with subdivision, which can cause a significant increase in the number of sides, to reduce the signaling cost and further reduce the number of sides.

[0179] These proposed methods can be used independently or in combination with each other in any order to generate any displacement. These proposed methods can be applied to any polygon mesh, not limited to triangular meshes. In this article, we use the number of triangles as an example.

[0180] The above techniques can be implemented as computer software via computer-readable instructions and physically stored in one or more computer-readable media, or implemented by one or more specially configured hardware processors. For example, Figure 18 FIG. 1800 shows a computer system suitable for implementing certain embodiments of the disclosed subject matter.

[0181] Computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subject to assembly, compilation, linking, or similar mechanisms to create code including instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through interpretation, microcode, etc.

[0182] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0183] Figure 18 The components shown for the computer system 1800 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system 1800.

[0184] The computer system 1800 can include certain human-machine interface input devices. Such human-machine interface input devices can respond to the input of one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices can also be used to capture certain media not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0185] The input human-machine interface devices can include one or more of the following (only one of each is shown): keyboard 1801, mouse 1802, touchpad 1803, touch screen 1810, joystick 1805, microphone 1806, scanner 1808, camera 1807.

[0186] The computer system 1800 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through the touch screen 1810 or the joystick 1805, but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers 1809, headphones (not depicted)), visual output devices (e.g., the screen 1810 including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each screen having or not having touch screen input functionality, each screen having or not having tactile feedback functionality, where some screens are capable of outputting two-dimensional visual output or output beyond three dimensions through means such as stereoscopic output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).

[0187] The computer system 1800 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 1820 with CD / DVD 1811 or similar media, thumb drives 1822, removable hard disk drives or solid state drives 1823, traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0188] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other volatile signals.

[0189] The computer system 1800 may also include an interface 1899 leading to one or more communication networks 1898. The network 1898 may be, for example, a wireless network, a wired network, or an optical fiber network. The network 1898 may also be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, and so on. Examples of the network 1898 include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks 1898 generally require an external network interface adapter (such as the USB port of the computer system 1800) to be connected to certain general data ports or peripheral buses (1850 and 1851); as described below, other network interfaces are generally integrated into the kernel of the computer system 1800 by connecting to the system bus (for example, an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system 1800 may use any of these networks 1898 to communicate with other entities. Such communication may be only one-way reception (such as broadcast television), only one-way transmission (such as CANBus connected to certain CANbus devices), or two-way, such as using a local area network or a wide area digital network to connect to other computer systems. Certain protocols and protocol stacks may be used on each of the networks and network interfaces described above.

[0190] The above-mentioned human-machine interface devices, human-machine accessible storage devices, and network interfaces may be attached to the kernel 1840 of the computer system 1800.

[0191] The kernel 1840 may include one or more central processing units (CPUs) 1841, a graphics processing unit (GPU) 1842, a graphics adapter 1817, a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) 1843, a hardware accelerator 1844 for certain tasks, and so on. These devices, as well as a read-only memory (ROM) 1845, a random access memory 1846, and an internal mass storage 1847 such as an internal non-user accessible hard disk drive, SSD, etc., may be connected through a system bus 1848. In some computer systems, the system bus 1848 may be accessed in the form of one or more physical plugs to enable expansion through additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the system bus 1848 of the kernel or connected to the system bus 1848 of the kernel through a peripheral bus 1849. The architecture of the peripheral bus includes PCI, USB, etc.

[0192] The CPU 1841, GPU 1842, FPGA 1843, and accelerator 1844 can execute certain instructions that, when combined, can constitute the aforementioned computer code. This computer code can be stored in the ROM 1845 or the RAM 1846. Transitional data can also be stored in the RAM 1846, while permanent data can be stored, for example, in the internal mass storage 1847. Fast storage and retrieval of any storage device can be achieved by using a cache, which can be closely associated with the following: one or more CPUs 1841, GPUs 1842, mass storage 1847, ROM 1845, RAM 1846, etc.

[0193] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code can be of the type well known and available to those skilled in the field of computer software.

[0194] As a non-limiting example, a computer system having the architecture 1800, particularly the core 1840, can provide functionality due to software executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-volatile memories of the core 1840, such as the core internal mass storage 1847 or the ROM 1845. The software implementing the various embodiments of this disclosure can be stored in such devices and executed by the core 1840. Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core 1840, particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in the RAM 1846 and modifying such data structures according to the software-defined processes. Additionally or alternatively, a computer system can provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., accelerator 1844), which can replace the software or operate in conjunction with the software to execute the specific processes or specific parts of the specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or include both. This disclosure encompasses any suitable combination of hardware and software.

[0195] Although the present disclosure has described multiple exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.

Claims

1. A video decoding method, characterized in that: The method comprises: Obtaining a grid from a bitstream, the grid representing encoded volume data of at least one three-dimensional 3D visual content; Dividing a plurality of vertices of the mesh into a plurality of groups by determining a face normal for each face in the mesh, sorting the plurality of groups according to orientations of the plurality of groups relative to the face normals, and merging adjacent faces that share edges between the plurality of groups; and The encoded volume data is decoded based on the plurality of groups.

2. The method according to claim 1, characterized in that The dividing the plurality of vertices of the mesh into a plurality of groups further comprises: The faces that share the edge are identified, and the faces that share the edge are merged into a polygonal face that is larger than before the faces were merged.

3. The method according to claim 1, characterized in that The decoding of the encoded volume data based on the plurality of groups further includes segmenting the merged adjacent faces into target m-gonal faces.

4. The method according to claim 3, characterized in that The target m-gonal face is a triangular face.

5. The method according to claim 3, characterized in that: The segmenting of the merged adjacent faces into target m-gonal faces is based on the decoded displacements.

6. The method according to claim 5, characterized in that At least one of the decoded displacements is based on a non-zero displacement syntax indicating a sum of values.

7. The method according to claim 5, characterized in that The segmenting of the merged adjacent faces into target m-gonal faces is further based on a non-zero flag nzFlag syntax.

8. A video decoding device, characterized in that: include: at least one memory configured to store computer program code; At least one processor is configured to access the computer program code and perform the operation according to any one of claims 1 to 7 according to instructions of the computer program code.

9. A video encoding method, characterized in that: The method comprises: Obtaining a grid of volumetric data representing at least one three-dimensional 3D visual content; Dividing a plurality of vertices of the mesh into a plurality of groups by determining a face normal for each face in the mesh, sorting the plurality of groups according to orientations of the plurality of groups relative to the face normals, and merging adjacent faces that share edges between the plurality of groups; and The volume data is encoded based on the plurality of groups.

10. A video encoding device, characterized in that: include: at least one memory configured to store computer program code; At least one processor is configured to access the computer program code and perform the operation of claim 9 according to instructions of the computer program code. 11 . A non-volatile computer-readable medium having a program stored thereon, the program causing a computer to execute the method according to any one of claims 1 to 8 or the method according to claim 9.

12. A method for processing a video code stream, characterized in that: The video code stream is generated according to the encoding method according to claim 9, or is decoded based on the decoding method according to any one of claims 1 to 8.