Efficient coding and decoding of texture coordinate connectivity in polygonal meshes

By receiving and processing the encoded information of the grid, generating and reconstructing the grid and predicting the seam edges, cutting the grid into small block components, solving the problem of low encoding and decoding efficiency of polygon mesh texture coordinates, achieving more efficient data processing and a better immersive experience.

CN120112952APending Publication Date: 2025-06-06TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004582.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-21
Filing Date
2024-08-22
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is inefficient when dealing with the texture coordinate connectivity of polygon mesh, making it difficult to effectively code, and affecting the immersive experience.

Method used

By receiving the encoded information of the grid, a reconstructed grid is generated, the seam vertices and seam edges are determined, the seam edges are predicted, and the grid is cut into small pieces of components to determine the UV connectivity of the UV vertices.

Benefits of technology

It improves the texture coordinate connectivity encoding and decoding efficiency of polygon mesh, reduces the use of data transmission and storage resources, and improves the quality of immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112952A_ABST
    Figure CN120112952A_ABST
Patent Text Reader

Abstract

A grid processing method includes: receiving encoded information of a grid, the encoded information including position connectivity of a plurality of three-dimensional (3D) vertexes of the grid in a 3D space, and correspondence between the plurality of 3D vertexes and a plurality of UV vertexes of the grid in a UV space; generating a reconstructed grid in the 3D space based on the plurality of 3D vertexes according to the position connectivity of the plurality of 3D vertexes; determining a plurality of seam vertices based on the plurality of 3D vertices, the plurality of seam vertices corresponding to two or more UV vertices in the UV space; predicting a plurality of seam edges in the 3D space based at least on the plurality of seam vertices, where at least one edge between two seam vertices is predicted as a seam edge, the seam edge corresponding to two or more UV edges in the UV space; cutting the reconstructed grid into a plurality of small block components according to the plurality of joint edges; and, according to the plurality of small block components, determining the UV connectivity of the plurality of UV vertices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by Reference

[0002] This application claims priority to U.S. application No. 18 / 811,591, filed on August 21, 2024, entitled “Efficient Encoding and Decoding of Texture Coordinate Connectivity in Polygonal Meshes,” and U.S. provisional application No. 63 / 534,109, filed on August 22, 2023, entitled “Efficient Encoding and Decoding of Texture Coordinate Connectivity in Polygonal Meshes,” the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The embodiments of the present application relate to video encoding and decoding. Background Art

[0004] The background description provided herein is intended to present the background of the present application as a whole. The extent to which the work of the presently named inventors described in the background section and various aspects of this specification is performed does not indicate that it is prior art at the time of filing this application, and it is never explicitly or implicitly admitted that it is prior art for this application.

[0005] Various technologies have been developed to capture and represent the world, such as objects in the world in three-dimensional (3D) space, environments in the world, etc. 3D representations of the world can enable more immersive forms of interaction and communication. For example, the development of 3D media processing technology, such as advances in three-dimensional (3D) capture, 3D modeling, and 3D rendering, has promoted the ubiquity of 3D media content on multiple platforms and devices. In one example, a baby's first steps can be captured in one area, and media technology allows grandparents to watch (and possibly interact) with the baby in another area and enjoy an immersive experience. According to one aspect of the present application, in order to improve the immersive experience, 3D models have become increasingly complex, and the creation and consumption of 3D models occupy a large amount of data resources, such as data storage and data transmission resources. In some examples, a 3D mesh can be used as a 3D representation of the world. Summary of the invention

[0006] Various aspects of the present application provide code streams, methods and devices for trellis coding and decoding. In some examples, the trellis coding and decoding device includes a processing circuit.

[0007] In some examples, the mesh processing method includes:

[0008] Receiving encoded information of a mesh, the encoded information including position connectivity of a plurality of three-dimensional 3D vertices of the mesh in a 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in a UV space;

[0009] generating a reconstructed mesh in the 3D space based on the plurality of 3D vertices according to positional connectivity of the plurality of 3D vertices;

[0010] Based on the plurality of 3D vertices, determining a plurality of seam vertices, the plurality of seam vertices corresponding to two or more UV vertices in the UV space;

[0011] Based at least on the plurality of seam vertices, predict a plurality of seam edges in the 3D space, wherein at least one edge between two seam vertices is predicted as a seam edge, and the seam edge corresponds to two or more UV edges in the UV space;

[0012] According to the plurality of seam edges, the reconstructed mesh is cut into a plurality of small components; and

[0013] UV connectivity of the plurality of UV vertices is determined based on the plurality of patch components.

[0014] Some aspects of the present application provide a grid processing method, including:

[0015] Encoding, in the encoded information of the mesh, position connectivity of a plurality of three-dimensional 3D vertices of the mesh in the 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in the UV space;

[0016] Based on the plurality of 3D vertices, determining a plurality of seam vertices, the plurality of seam vertices corresponding to two or more UV vertices in the UV space;

[0017] Predicting a plurality of predicted seam edges in the 3D space based on the plurality of seam vertices, wherein at least one edge between two seam vertices is predicted as a seam edge, and the seam edge corresponds to two or more UV edges in the UV space; and,

[0018] The adjustment information is encoded in the encoded information of the grid, and the adjustment information is used to generate a plurality of real seam edges from the plurality of predicted seam edges.

[0019] Some aspects of the present application provide a grid data processing method, including:

[0020] Process the code stream of the grid data according to the format rules, where:

[0021] The bitstream includes encoded information of the mesh, the encoded information including position connectivity of a plurality of three-dimensional 3D vertices of the mesh in a 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in a UV space;

[0022] The format rules specify:

[0023] generating a reconstructed mesh in the 3D space based on the plurality of 3D vertices according to positional connectivity of the plurality of 3D vertices;

[0024] Based on the plurality of 3D vertices, determining a plurality of seam vertices, the plurality of seam vertices corresponding to two or more UV vertices in the UV space;

[0025] Based at least on the plurality of seam vertices, predict a plurality of seam edges in the 3D space, wherein at least one edge between two seam vertices is predicted as a seam edge, and the seam edge corresponds to two or more UV edges in the UV space;

[0026] According to the plurality of seam edges, the reconstructed mesh is cut into a plurality of small components; and

[0027] UV connectivity of the plurality of UV vertices is determined based on the plurality of patch components.

[0028] Some aspects of the present application provide a non-volatile computer-readable storage medium having instructions stored thereon, which, when executed by a computer, enables the computer to implement the above-mentioned grid processing method. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Other features, properties and various advantages of the disclosed subject matter will become further apparent from the following detailed description and accompanying drawings, in which:

[0030] Figure 1 is an exemplary block diagram of a flow system in some examples;

[0031] Figure 2 is a schematic diagram of an exemplary block diagram of a decoder;

[0032] Figure 3 is a schematic diagram of an exemplary block diagram of an encoder;

[0033] Figure 4 An example of an encoding process (400) for grid processing according to an embodiment of the present application is shown;

[0034] Figure 5 An example of a decoding process (500) for grid processing according to an embodiment of the present application is shown;

[0035] Figure 6 shows a schematic diagram illustrating a mapping from a 3D grid (610) to a 2D atlas (620) in some examples;

[0036] Figure 7 An example of a diagram (700) showing a grid in some examples;

[0037] Figure 8 An example of a grid diagram (800) is shown showing some examples;

[0038] Fig. 9 An example of a graph (900) in 2D UV space is shown in some examples;

[0039] Fig.10 An example of a UV map (1000) in 2D UV space is shown in some examples;

[0040] Fig.11 A diagram showing a cutting grid in an example;

[0041] Fig.12 A diagram showing a cutting grid in another example;

[0042] Fig.13 A flowchart of a decoding process according to an embodiment of the present application is shown;

[0043] Fig.14 A method flow chart of an encoding process according to an embodiment of the present application is shown;

[0044] Fig.15 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION

[0045] Various aspects of the present application provide techniques in the field of grid processing.

[0046] A mesh (also called a mesh model) includes several polygons (also called faces) that describe the surface of a volumetric object. Each polygon can be defined by vertices in a three-dimensional (3D) space and information about how the vertices are connected, referred to as connectivity information. In some examples, the mesh also includes vertex attributes associated with the mesh vertices, such as color, normal, displacement, etc. In addition, in some examples, the mesh can include attributes associated with the mesh surface by utilizing mapping information that parameterizes the mesh using a two-dimensional (2D) attribute map. This mapping is typically described by a set of parametric coordinates associated with the mesh vertices, referred to as UV coordinates or texture coordinates. 2D attribute maps are used to store high-resolution attribute information, such as textures, normals, displacements, etc. 2D attribute maps can be used for various purposes, such as texture mapping, shading, and mesh reconstruction.

[0047] Figure 1A block diagram of a video processing system (100) is shown in some examples. The video processing system (100) is an example of an application of the subject matter disclosed in this application, a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.

[0048] The streaming system (100) includes a capture subsystem (113) that may include a 3D source (101), such as a light detection and ranging (LIDAR) system, a 3D camera, a 3D scanner, a graphics generation component, etc., for creating an uncompressed 3D data stream (102). In one example, the 3D data stream (102) includes samples collected by a 3D camera system. The 3D data stream (102) is depicted as thick lines to emphasize the high data volume compared to the encoded 3D data (104) (or encoded code stream), and can be processed by an electronic device (120) including a 3D encoder (103) coupled to the 3D source (101). The 3D encoder (103) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter as described in more detail below. The encoded 3D data (104) (or encoded code stream) is depicted as thin lines to emphasize the lower data volume compared to the 3D data stream (102), and can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 The client subsystems (106) and (108) in the streaming server (105) can access the streaming server (105) to retrieve copies (107) and (109) of the encoded 3D data (104). The client subsystem (106) can include, for example, a 3D decoder (110) in the electronic device (130). The 3D decoder (110) decodes the incoming copy (107) of the encoded 3D data and creates an output stream (111) of a 3D representation that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded 3D data (104), (107) and (109) (e.g., video streams) can be encoded according to some 3D encoding / compression standard, such as a grid encoding / compression standard.

[0049] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a 3D decoder (not shown), and the electronic device (130) may also include a 3D encoder (not shown).

[0050] It should also be noted that in some examples, the 3D encoder and / or 3D decoder may use 2D encoding / decoder technology.For example, the 3D encoder and / or 3D decoder may include a video decoder or a video encoder.

[0051] Figure 2 An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) may be disposed in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) is shown in an example.

[0052] The receiver (231) may receive at least one encoded video sequence to be decoded by the video decoder (210); in the same or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be provided external to the video decoder (210) (not shown). In other cases, a buffer memory (not shown) is provided outside the video decoder (210) to prevent network jitter, for example, and another buffer memory (215) may be configured inside the video decoder (210) to handle broadcast timing, for example. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (215), or the buffer memory may be made smaller. Of course, in order to use on a service packet network such as the Internet, a buffer memory (215) may also be required, and the buffer memory may be relatively large and have an adaptive size, and may be at least partially implemented in an operating system or a similar element (not shown) outside the video decoder (210).

[0053] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The types of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device such as a display device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown in . The control information for the display device may be a parameter set segment (not indicated) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (220) may extract a subgroup parameter set for at least one subgroup of a pixel subgroup in a video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a small block, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0054] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215), thereby creating symbols (221).

[0055] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by parser (220) from the coded video sequence. For the sake of brevity, such subgroup control information flow between parser (220) and the multiple units below is not described.

[0056] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In a practical embodiment operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0057] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) can output a block including sample values, which can be input into an aggregator (255).

[0058] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates surrounding blocks of the same size and shape as the block being reconstructed using reconstructed information extracted from a current picture buffer (258). For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0059] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221), these samples may be added to the output of the scaler / inverse transform unit (251) (in this case referred to as residual value samples or residual value signals) by the aggregator (255) to generate output sample information. The acquisition of the prediction samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit (253) in the form of the symbols (221), for example, including X, Y and reference picture components. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0060] The output samples of the aggregator (255) may be used by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and that are available to the loop filter unit (256) as symbols (221) from the parser (220). However, in other embodiments, the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to previously reconstructed and loop filtered sample values.

[0061] The output of the loop filter unit (256) may be a sample stream that may be output to a display device (212) and stored in a reference picture memory (257) for subsequent inter-picture prediction.

[0062] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0063] The video decoder (210) may perform decoding operations according to a predetermined video compression technique, such as in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0064] In one example, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0065] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used to replace Figure 1 A 3D encoder (103) in an embodiment.

[0066] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In another embodiment, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source can collect video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is a part of the electronic device (320).

[0067] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), wherein the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.301 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include at least one sample depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.

[0068] According to an embodiment, the video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (350). In some embodiments, the controller (350) controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, couplings are not shown in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) can be used to have other suitable functions that are related to the video encoder (303) optimized for a certain system design.

[0069] In some embodiments, the video encoder (303) operates in a coding loop. As a simple description, in one example, the coding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (because in the video compression technology considered in this application, any compression between the symbols and the encoded video code stream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, eg due to channel error values) is also used in some related techniques.

[0070] The operation of the "local" decoder (333) may be combined with, for example, Figure 5 The "remote" decoder described in detail for the video decoder (210) is identical. However, additional brief reference is made to Figure 5 , when symbols are available and the entropy encoder (345) and parser (220) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).

[0071] In one embodiment, any decoder technology except the parsing / entropy decoding present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. For this reason, the application focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually reversed with the decoder technology described comprehensively. A more detailed description is only needed in certain areas and is provided below.

[0072] During operation, in some embodiments, the source encoder (330) may perform motion compensated predictive coding. The motion compensated predictive coding predictively encodes an input picture with reference to at least one previously encoded picture from a video sequence designated as a "reference picture." In this manner, the encoding engine (332) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture that may be selected as a prediction reference for the input picture.

[0073] The local video decoder (333) may decode the encoded video data that may be designated as a reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may be a lossy process. When the encoded video data is available at the video decoder ( Figure 3 When the video sequence is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some error values. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (334). In this way, the video encoder (303) may locally store a copy of the reconstructed reference picture that has common content (without the transmission error values) with the reconstructed reference picture to be obtained by the remote video decoder.

[0074] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as appropriate prediction references for the new picture. The predictor (335) may perform operations on a sample block-by-pixel block basis to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it may be determined that the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory (334).

[0075] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0076] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0077] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (660), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0078] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0079] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and features.

[0080] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values ​​for each block.

[0081] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.

[0082] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block-wise. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or the blocks may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0083] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0084] In one example, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set segments, etc.

[0085] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0086] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order, respectively). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block may be predicted by a combination of the first reference block and the second reference block.

[0087] In addition, merge mode technology can be used in inter-picture prediction to improve encoding and decoding efficiency.

[0088] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Furthermore, each CTU can be split into at least one coding unit (CU) using a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into at least one prediction unit (PU). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one example, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.

[0089] It should be noted that the encoders (103), (303) and decoders (110), (210) may be implemented using any suitable technology. In one example, the 3D encoders (103), (303) and decoders (110), (210) may be implemented using at least one integrated circuit. In another embodiment, the encoders (103), (303) and decoders (110), (210) may be implemented using at least one processor that executes software instructions.

[0090] In some examples, the 3D data includes a mesh model, the 3D encoder (103) may include a mesh encoder, and the 3D decoder (110) may include a mesh decoder.

[0091] According to one aspect of the present application, a dynamic mesh is a mesh in which at least one of the components (geometric information, connectivity information, mapping information, vertex attributes, and attribute graph) changes over time. A dynamic mesh can be described by a series of meshes (also referred to as mesh frames). In some examples, a mesh frame in a dynamic mesh can be a representation of the surface of an object at different times, and each mesh frame is a representation of the surface of the object at a specific time (also referred to as a time instance). A dynamic mesh may require a large amount of data because the dynamic mesh may include a large amount of information that changes over time. Compression techniques for meshes can allow efficient storage and transmission of media content represented by meshes.

[0092] Dynamic mesh sequences can require a lot of data, since dynamic meshes can include a lot of information that changes over time. Therefore, efficient compression techniques can be used to store and transmit such content.

[0093] Figure 4 An example of an encoding process (400) for mesh processing according to one aspect of the present application is shown. Figure 4 As shown, the encoding process (400) includes a preprocessing step (410) and an encoding step (420). The preprocessing step (410) is configured to generate a base grid m(i) of the current frame and a displacement field d(i) of the current frame including a displacement vector according to an input grid M(i) of the current frame. The encoding step (420) is configured to encode the base grid m(i), the displacement field d(i), and the texture information of the base grid m(i). The displacement field d(i) of the current frame includes the displacement vector. The index i is used to refer to the current frame. In one aspect, a mode decision method can be performed in the encoding process (400) to determine whether to apply inter-frame coding (also known as inter-frame prediction or inter-frame mode), intra-frame coding (also known as intra-frame prediction or intra-frame mode), etc. to the current frame. For example, the mode decision method can compare the cost of the intra-frame mode and the cost of the inter-frame mode, and decide the encoding mode of the base grid m(i) of the current frame based on which cost is smaller. In some examples, the base grid m(i) is encoded using a skip mode. In the example, the skip mode is a special mode of the inter mode. For example, the base grid m(i) may be intra-coded, or inter-coded, or coded in the skip mode.

[0094] Still reference Figure 4, the preprocessing step (410) may include a mesh decimation process (412), a parameterization process such as an atlas parameterization process (414), and a subdivision surface fitting process (416). The mesh decimation process (412) is configured to downsample the vertices of the input mesh M(i) to generate a decimated mesh dm(i) that may include a plurality of decimated (or downsampled) vertices. In an example, the number of the plurality of decimated vertices is less than the number of vertices of the input mesh M(i). The parameterization process such as the atlas parameterization process (414) is configured to map the decimated mesh dm(i) onto a planar domain, such as onto a UV atlas (or UV map), to generate a re-parameterized mesh pm(i). In an example, the atlas parameterization may be performed based on a video processing tool (such as a UV Atlas tool). The subdivision surface fitting process (416) is configured to take the reparameterized mesh pm(i) and the input mesh M(i) as input and generate a base mesh m(i) and a displacement field d(i) including a displacement vector or a set of displacements. In an example of the subdivision surface fitting process (416), pm(i) is subdivided using a subdivision scheme such as iterative interpolation to obtain a subdivided mesh. Iterative interpolation includes inserting a new point in the middle of each edge of the reparameterized mesh pm(i) at each iteration. Any suitable subdivision scheme can be applied to subdivide pm(i). The displacement field d(i) is calculated by determining the nearest point on the surface of the input mesh M(i) for each vertex of the subdivided mesh.

[0095] Advantages of the subdivided mesh may include that the subdivided mesh has a subdivision structure that allows efficient compression while providing a faithful approximation of the input mesh. Increased compression efficiency may be obtained due to the following properties. The decimated mesh dm(i) may have a low number of vertices and may be encoded and transmitted using fewer bits than the input mesh M(i) or the subdivided mesh. Figure 4 , the base mesh m(i) can be generated by the extracted mesh dm(i). In the example, the base mesh m(i) is the extracted mesh dm(i). Since the subdivided mesh can be generated based on the subdivision method, the subdivided mesh can be automatically generated by the decoder when decoding the base mesh or the extracted mesh (for example, without using any information other than the subdivision scheme and the subdivision iteration count). On the decoder side, the displacement field d(i) can be generated by decoding the displacement vectors associated with the vertices of the subdivided mesh. In addition to allowing spatial / quality scalability, the subdivision structure also enables efficient transformations such as wavelet decomposition, which can provide high compression performance.

[0096] exist Figure 4 In the example, the encoding step (420) includes base grid encoding (422), displacement encoding (424), texture encoding (426), etc. The base grid encoding (422) is configured to encode geometric information of the base grid m(i) associated with the current frame. In intra-frame coding, the base grid m(i) may be first quantized (e.g., using uniform quantization) and then encoded, for example, by using a coding mode determined by a mode decision method. The coding mode may be an inter-frame mode, an intra-frame mode, a skip mode, etc. An encoder for intra-coding the base grid m(i) may be referred to as a static grid encoder. In inter-frame coding, a reference base grid associated with a reference frame indicated by index j (e.g., a reconstructed quantized reference base grid m'(j)) may be used to predict the base grid m(i) associated with the current frame indicated by index i. The displacement encoding (424) is configured to encode the displacement field d(i) generated in the preprocessing step (410). The displacement field d(i) may include a set of displacement vectors (or displacements) associated with the subdivided mesh vertices. Texture encoding (426) is configured to encode attribute information of the base mesh m(i). The attribute information may include texture, normal, color, etc. The attribute information may be encoded based on a suitable codec, such as High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC).

[0097] On the one hand, reference Figure 4 , a mesh encoding process (such as encoding process (420)) begins with preprocessing (e.g., preprocessing step (410)). The preprocessing may convert an input mesh (e.g., an input dynamic mesh) M(i) into a base mesh m(i) and a displacement field d(i) including a set of displacements (or a set of displacement vectors). The encoding step (420) may compress the output from the preprocessing (e.g., m(i), d(i), etc.) and generate a compressed code stream b(i). The compressed code stream b(i) may include a compressed base mesh code stream, a compressed displacement field code stream, a compressed attribute code stream, etc.

[0098] Figure 5An example of a decoding process (500) for grid processing according to an aspect of the present application is shown. The decoding process (500) may include a decoding step (510) and a post-processing step (520). A compressed code stream b(i) may be fed to the decoding step (510). In an example, such as for lossless transmission, the compressed code stream b(i) is the output b(i) from the encoding process (400). The decoding step (510) may extract various sub-code streams, such as a compressed base grid sub-stream, a compressed displacement field sub-stream, a compressed attribute sub-stream, etc. The decoding step (510) may decode the sub-code streams to generate the following components: small block metadata indicated by metadata(i), a decoded base grid m"(i), a decoded displacement field (including displacement) d"(i), a decoded attribute map A"(i), etc.

[0099] On the one hand, the base grid substream can be fed to a grid decoder to generate a reconstructed quantized base grid m'(i). The decoded base grid (or reconstructed base grid) m"(i) can be obtained by applying inverse quantization to m'(i). The displacement field substream including the encoded packed and quantized wavelet coefficients can be decoded by a video and / or image decoder. Image unpacking and inverse quantization can be applied to the packed quantized wavelet coefficients, and the packed quantized wavelet coefficients are reconstructed to obtain unpacked and unquantized transform coefficients (e.g., wavelet coefficients). An inverse wavelet transform can be applied to the unpacked and unquantized wavelet coefficients to generate a decoded displacement field (or reconstructed displacement) d"(i).

[0100] The decoded components (e.g., including metadata(i), m"(i), d"(i), A"(i), etc.) can be fed to a post-processing step (520). A mesh (also referred to as a decoded / reconstructed mesh) M"(i) can be generated by the post-processing step (520) based on m"(i) and d"(i). In an example, the mesh M"(i) (also referred to as a reconstructed deformed mesh DM(i)) can be obtained by subdividing m"(i) using a subdivision scheme and applying the reconstructed displacements d"(i) to the vertices of the subdivided mesh. In an example, DM(i) can include a curve of the displacement. In an example, when the encoding process (400), the decoding process (500), and the transmission are lossless, the mesh M"(i) can be the same as the input mesh M(i). When one of the encoding process (400), the decoding process (500), and the transmission is lossy, M"(i) is different from M(i). In various examples, the difference between M"(i) and M(i), if any, can be relatively small. In the example, an attribute graph A”(i) is also generated by a post-processing step (520).

[0101] In some examples, the mesh may also include attributes associated with the vertices, such as color, normals, etc. Attributes may be associated with the surface of the mesh using mapping information that parameterizes the mesh using a 2D attribute map. The mapping information is typically described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. The 2D attribute map (referred to as a texture map in some examples) is used to store high-resolution attribute information, such as texture, normals, displacement, etc. Such information may be used for various purposes, such as texture mapping and shading.

[0102] In some embodiments, a mesh may include components referred to as geometric information, connectivity information, mapping information, vertex attributes, and property graphs. In some examples, the geometric information is described by a set of 3D positions associated with the vertices of the mesh. In the example, (x, y, z) coordinates can be used to describe the 3D position of the vertex and are also referred to as 3D coordinates. In some examples, connectivity information includes a set of vertex indices describing how to connect vertices to create a 3D surface. In some examples, mapping information describes how to map a mesh surface to a 2D region of a plane. In the example, mapping information is described by a set of UV parameters / texture coordinates (u, v) associated with mesh vertices and connectivity information. In some examples, vertex attributes include scalar or vector attribute values ​​associated with mesh vertices. In some examples, the property graph includes attributes associated with the mesh surface and stored as a 2D image / video. In the example, the mapping between a video (e.g., a 2D image / video) and a mesh surface is defined by mapping information.

[0103] According to one aspect of the present application, some techniques referred to as UV mapping or mesh parameterization are used to map the surface of a mesh in a 3D domain to a 2D domain. In some examples, the mesh is cut into patches (also referred to as patch components) in the 3D domain. A patch is a continuous subset of a mesh whose boundaries are formed by boundary edges. The boundary edges of a patch are edges that belong to only one polygon of the patch and are not shared by two adjacent polygons in the patch. In some examples, the vertices of the boundary edges in a patch are referred to as the boundary vertices of the patch, and the non-boundary vertices in the patch can be referred to as the internal vertices of the patch.

[0104] According to one aspect of the present application, in some examples, tiles are parameterized as 2D shapes (also referred to as UV tiles, 2D tiles, or UV graphs). In some examples, the 2D shapes can be packed (e.g., oriented and placed) into a graph, which is also referred to as a UV atlas. In some examples, the graph can be further processed using 2D image or video processing techniques.

[0105] In an example, the UV mapping technique generates a UV atlas (also referred to as a UV map) and one or more texture atlases (also referred to as texture maps) in 2D corresponding to a small block of a 3D mesh. The UV atlas includes assigning 3D vertices of a 3D mesh to 2D points in a 2D domain (e.g., a rectangle). The UV atlas is a mapping between coordinates of a 3D surface to coordinates of a 2D domain. In an example, a point at a 2D coordinate (u, v) in the UV atlas has a value formed by the coordinates (x, y, z) of the vertex in the 3D domain. In an example, the texture atlas includes color information of the 3D mesh. For example, a point at a 2D coordinate (u, v) in the texture atlas (which has a 3D value (x, y, z) in the UV atlas) has a color that specifies the color attribute of the point at (x, y, z) in the 3D domain. In some examples, the coordinates (x, y, z) in the 3D domain are referred to as 3D coordinates or xyz coordinates, and the 2D coordinates (u, v) are referred to as uv coordinates or UV coordinates.

[0106] According to some aspects of the present application, mesh compression may be performed by representing the mesh using one or more 2D graphs (also referred to as 2D atlases in some examples), and then encoding the 2D graphs using an image or video codec. Different techniques may be used to generate the 2D graphs.

[0107] Figure 6 A schematic diagram illustrating the mapping from a 3D grid (610) to a 2D atlas (620) in some examples is shown. Figure 6 In the example, the 3D mesh (610) includes four vertices 1 to 4 that form four tiles A to D. Each of the tiles has a set of vertices and associated attribute information. For example, tile A is formed by vertices 1, 2, and 3 connected to form a triangle; tile B is formed by vertices 1, 3, and 4 connected to form a triangle; tile C is formed by vertices 1, 2, and 4 connected to form a triangle; and tile D is formed by vertices 2, 3, and 4 connected to form a triangle. In some examples, vertices 1, 2, 3, and 4 may have corresponding attributes, and the triangle formed by vertices 1, 2, 3, and 4 may have corresponding attributes.

[0108] In an example, tiles A, B, C, and D in 3D are mapped to a 2D domain, such as a 2D atlas (620), which is also referred to as a UV atlas (620) or a graph (620). For example, tile A is mapped to a 2D shape (also referred to as a UV tile) A' in the graph (620), tile B is mapped to a 2D shape (also referred to as a UV tile) B' in the graph (620), tile C is mapped to a 2D shape (also referred to as a UV tile) C' in the graph (620), and tile D is mapped to a 2D shape (also referred to as a UV tile) D' in the graph (620). In some examples, coordinates in the 3D domain are referred to as (x, y, z) coordinates, and coordinates in a 2D domain (such as the graph (620)) are referred to as UV coordinates. Vertices in a 3D mesh may have corresponding UV coordinates in the graph (620).

[0109] The map (620) may be a geometry map having geometry information, or may be a texture map having color, normal, fabric or other attribute information, or may be an occupancy map having occupancy information.

[0110] Although in Figure 6 Each tile is represented by a triangle in the examples, but note that a tile can include any suitable number of vertices connected to form a continuous subset of a mesh. In some examples, the vertices in a tile are connected into triangles. Note that other suitable shapes can be used to connect the vertices in a tile.

[0111] In an example, the geometric information of the vertices may be stored in a 2D geometric map. For example, the 2D geometric map stores the (x, y, z) coordinates of the sample points at corresponding points in the 2D geometric map. For example, a point at a (u, v) position in the 2D geometric map has a vector value of 3 components corresponding to the x, y, and z values ​​of the corresponding sample point in the 3D grid, respectively.

[0112] According to one aspect of the present application, the area in the figure may not be fully occupied. Figure 6 In the decoded image, the area outside the 2D shapes A', B', C' and D' is undefined. After decoding, the sample values ​​of the area outside the 2D shapes A', B', C' and D' can be discarded. In some cases, the occupancy map is used to store some additional information for each pixel, such as storing a binary value to identify whether the pixel belongs to a small block or is undefined.

[0113] Figure 7 An example of a map (700) of a mesh in some examples is shown. The map (700) is a UV map (also referred to as a UV atlas) that includes texture / UV coordinates and UV connectivity. The map (700) includes a plurality of UV charts that may correspond to tiles of the mesh in the 3D domain, the UV charts including UV coordinates of points and connectivity of points. In some examples, when the mesh includes a UV chart, such as Figure 7 As shown, the mesh codec needs to encode the UV coordinates and corresponding connectivity (hereinafter referred to as UV connectivity) of the UV chart.

[0114] In a first related example, when all properties of a mesh share the same connectivity (a single connectivity mesh), such as when positional connectivity and UV connectivity in 3D are exactly the same, the single connectivity is encoded and the UV connectivity does not need to be encoded separately.

[0115] In some examples, UV connectivity is different from position connectivity. For example, when a mesh is cut, some vertices and / or edges are split, and UV connectivity is different from position connectivity in 3D. In a second related example, UV connectivity is encoded as a separate mesh in a direct encoding method, and the correspondence between the corners of the UV coordinates and the 3D positions can be signaled. However, signaling the corner correspondence requires a large number of bits, and the redundancy between position connectivity and UV connectivity is significant, so encoding UV coordinate connectivity using a direct encoding method may be inefficient.

[0116] In a third related example, seam edges are signaled. When an edge between two 3D vertices in 3D space is split into two edges in 2D UV space, the edge is called a seam edge. In a third related example, signaling of angle correspondences can be avoided, and UV connectivity can be inferred from position connectivity and seam edges.

[0117] According to some aspects of the present application, signaling seam edges is also not efficient because the number of edges is much larger than the number of vertices and faces. Some aspects of the present application provide more efficient techniques to encode UV connectivity for polygon mesh compression. These techniques for encoding UV connectivity for polygon mesh compression can be applied individually or in any combination.

[0118] In some examples, when a vertex in 3D space is split into two or more vertices in 2D UV space, the vertex is called a seam vertex. According to one aspect of the present application, the seam vertex can be used to determine the seam edge, so that the direct signaling of the seam edge to encode the UV connectivity can be avoided. The seam edge is then used to cut the reconstructed 3D mesh into connected components (also referred to as small patches) corresponding to the UV chart. In addition, the UV connectivity within the UV chart can be obtained based on position connectivity. For example, the encoder / decoder can determine the seam vertex from multiple 3D vertices, and the seam vertex corresponds to two or more UV vertices in the UV space. In addition, the encoder / decoder can perform prediction of the seam edge in the 3D space based at least on the seam vertex, at least predicting the edge between the two seam vertices as a seam edge corresponding to two or more UV edges in the UV space. The encoder can encode the predicted adjustment information based on the true seam edge. The decoder may determine a true seam edge based on the adjustment information, cut the reconstructed mesh into small patch components according to the true seam edge, and determine UV connectivity of UV vertices according to the small patch components.

[0119] Figure 8 A schematic diagram (800) of a mesh in some examples is shown. The schematic diagram (800) shows a mesh in 3D space as shown in (810), and a mesh in 2D UV space as shown in (820). For example, the mesh in 3D space (810) is cut based on seam edges to obtain a mesh in 2D UV space (820). Figure 8 In the example, 3D vertices "pos1" and "pos4" are seam vertices. Seam vertex "pos1" has two corresponding UV vertices "uv1" and "uv6", and seam vertex "pos4" has two corresponding UV vertices "uv4" and "uv7". Figure 8 In the example, other vertices in 3D space have only one corresponding UV vertex and are non-seam vertices. For example, "pos0" in 3D space has a corresponding vertex "uv0" in 2D UV space, "pos2" in 3D space has a corresponding vertex "uv2" in 2D UV space, "pos3" in 3D space has a corresponding vertex "uv3" in 2D UV space, and "pos5" in 3D space has a corresponding vertex "uv5" in 2D UV space. Because they have only one corresponding UV vertex. Figure 8 In this example, the edge between "pos1" and "pos4" in 3D space is the seam edge that is split into two edges in 2D UV space.

[0120] According to some aspects of the present application, seam vertices and seam edges may be determined based on a combination of prediction and signaling. Encoding of UV (also known as texture) coordinate connectivity for polygonal mesh compression may be determined based on cutting a 3D mesh into a UV chart by seam edges and positional connectivity in 3D.

[0121] According to one aspect of the present application, encoding of UV coordinate connectivity for polygonal mesh compression may be performed in three steps.

[0122] In the first step, the seam vertices are determined. In an example, the correspondence between 3D vertices and UV vertices is available and can be used to determine the seam vertices. For each 3D vertex, when the 3D vertex has multiple corresponding UV vertices, the 3D vertex is a seam vertex.

[0123] In another example, the valence of the vertex can be used to determine the seam vertex. For example, for each 3D vertex, the valence of the 3D vertex can be determined, which is defined as the number of incident faces of the 3D vertex. In addition, for each 2D vertex, the valence of the 2D vertex can be determined. When the valence of a 3D vertex is different from the valence of one of the corresponding UV vertices, then the 3D vertex is a seam vertex.

[0124] In the second step, a seam edge is determined (eg, predicted) based on the seam vertices. The seam vertices determined in the first step may be used to predict the seam edge. In an example, when an edge is between two seam vertices in 3D space, the edge is predicted to be a seam edge.

[0125] In another example, a mesh has a border, and a virtual face is added to the mesh to generate a closed mesh including a closed surface without a border. The virtual face may include the added UV vertices and faces. In the example, the added UV vertices and faces are disconnected from the original UV vertices, so the border vertices in the original mesh become seam vertices in the closed mesh after the virtual face is added. Therefore, in the example, each edge in the virtual face is a seam edge in the closed mesh.

[0126] In some examples, the predicted seam edges are incorrect.

[0127] Fig. 9 A schematic diagram of a map (900) in a 2D UV space in some examples is shown. The map (900) is a UV map and includes UV charts corresponding to small tiles of a mesh in a 3D space. In an example, the mesh in the 3D space is cut into small tiles from seam edges.

[0128] exist Fig. 9In the example, the graph (900) includes UV graphs (910), (920), and (930). The UV graph (910) includes UV vertices uv0, uv7, uv8, uv9, and uv10; the UV graph (920) includes UV vertices uv1, uv5, uv6, uv11, and uv12; and the UV graph (930) includes UV vertices uv2, uv3, uv4, uv13, and uv14. Fig. 9 In the example, UV vertices uv0, uv1, and uv2 correspond to the first seam vertex (901) in the 3D space; UV vertices uv4 and uv5 correspond to the second seam vertex (902) in the 3D space; UV vertices uv6 and uv7 correspond to the third seam vertex (903) in the 3D space, UV vertices uv10 and uv11 correspond to the fourth seam vertex (904) in the 3D space, and UV vertices uv12 and uv13 correspond to the fifth seam vertex (905) in the 3D space. In the example, using the prediction rule of predicting an edge as a seam edge when the edge is between two seam vertices in the 3D space, the edge between the seam vertices corresponding to the second seam vertex (902) and the third seam vertex (903) is a seam edge. Note that the edge between seam vertices (902) and (903) is not a seam edge, and the prediction is incorrect.

[0129] Fig.10 A schematic diagram of a UV map (1000) in a 2D UV space in some examples is shown. The map (1000) is a UV map corresponding to a mesh of a cube in a 3D space. The cube in the 3D space is cut along the seam edges, and a UV chart (1010) is formed in the UV map (1000). In the example, the 3D vertices of the cube are split into Fig.10 UV vertices uv1 and uv3 are shown.

[0130] In some examples, when a 3D vertex has multiple corresponding UV vertices, the 3D vertex is determined to be a seam vertex. Since the 3D vertex corresponding to UV vertex uv2 has only one corresponding UV vertex, the 3D vertex corresponding to UV vertex uv2 is not a seam vertex. In the example, using a prediction rule that predicts an edge as a seam edge when the edge is between two seam vertices in 3D space, the prediction rule cannot predict the edge between the 3D vertices corresponding to UV vertices uv2 and uv1 / uv3 as a seam edge. However, the edge between the 3D vertices corresponding to UV vertices uv2 and uv1 / uv3 is actually a seam edge that is split into a first edge (1001) between UV vertices uv1 and uv2 and a second edge (1002) between UV vertices uv2 and uv3 in 2D UV space.

[0131] According to one aspect of the present application, signaling can be used to handle incorrect predictions of seam edges. In some examples, different symbols are used to signal different types of vertices determined in the first step. In an example, three symbols such as 0, 1, and 2 are used to signal non-seam vertices (e.g., symbol 0), seam vertices adjacent to the seam edge correctly predicted (e.g., symbol 1), and seam vertices adjacent to the seam edge incorrectly predicted (e.g., symbol 2), respectively.

[0132] In some examples, for non-seam vertices (e.g., symbol 0), additional symbols are used to signal whether the non-seam vertex has a different connectivity than the corresponding UV vertex. In some examples, non-seam vertices with different connectivity than the corresponding UV vertex are referred to as semi-seam vertices. For example, additional symbols are used to signal whether the non-seam vertex has a different connectivity than the corresponding UV vertex. Fig.10 The 3D vertex of the UV vertex uv2 in is identified as a half-seam vertex. With additional notation, when a half-seam vertex is identified, the edge between the seam vertex and the half-seam vertex can be predicted as a seam edge. Note that the edge between two half-seam vertices cannot be a seam edge.

[0133] According to one aspect of the present application, at the decoder side, the decoder can correctly predict the seam edge connected to the seam vertex by correctly predicting the adjacent seam edge (e.g., signaling with, for example, a symbol 1). In some embodiments, in order for the decoder to correctly determine all seam edges, the real seam edges of the seam edges around the seam vertex that are incorrectly predicted (e.g., signaled with, for example, a symbol 2) can be signaled. For example, in Fig. 9In the example, the third seam vertex (903) corresponding to the UV vertices uv6 and uv7 is signaled as a seam vertex whose seam edge is incorrectly predicted with a symbol 2. For the third seam vertex (903), the associated edge between the third seam vertex (903) and the first seam vertex (901) is a real seam edge, the associated edge between the third seam vertex (903) and the fourth seam vertex (904) is a real seam edge, while the edge between the third seam vertex (903) and the second seam vertex (902) is not a real seam edge, and the edge between the third seam vertex (903) and the sixth 3D vertex (906) corresponding to the UV vertex uv8 is not a real seam edge. According to the prediction rule (e.g., the edge between two seam vertices is a seam edge), the associated edge between the third seam vertex (903) and the first seam vertex (901) is predicted as a seam edge, the associated edge between the third seam vertex (903) and the fourth seam vertex (904) is predicted as a seam edge, and the edge between the third seam vertex (903) and the second seam vertex (902) is predicted as a seam edge. The edge between the third seam vertex (903) and the sixth 3D vertex (906) is not predicted as a seam edge because the 3D vertex (906) is not a seam vertex. Therefore, the prediction of the seam edge between the third seam vertex (903) and the second seam vertex (902) is incorrect. In some examples, when the real seam edge around the seam vertex (903) is signaled, the decoder can correctly determine the real seam edge associated with the seam vertex (903).

[0134] In this example, by using the predicted seam edges, the real seam edges can be more efficiently predictively encoded. Fig. 9 In the example, the real seam edge (in the counterclockwise direction starting from the 3D vertex corresponding to uv8) around the 3D vertex (903) can be represented as 0, 1, 0, 1, where 0 represents a non-seam edge and 1 represents a real seam edge. In the example, the predicted seam edge around the 3D vertex (903) is represented as 0, 1, 1, 1 using a prediction rule that predicts an edge as a seam edge when the edge is between two seam vertices in 3D space.

[0135] In an embodiment, at the encoder side, the predicted seam edge symbol may be subtracted from the real seam edge symbol to obtain a prediction residual, and the prediction residual may be encoded using entropy coding. In another embodiment, the predicted seam edge symbol is used to select a context for entropy coding of the corresponding real seam edge symbol. For example, in Fig. 9 In the example, context 0 is used to encode the first associated edge symbol, and context 1 is used to encode the other three associated edge symbols.

[0136] According to one aspect of the present application, in order to save additional bits for signaling seam edges, the edges that have been signaled or predicted can be marked. Fig. 9 In the example, when the associated edge of the 3D vertex (903) is encoded, the associated edge between the 3D vertex (903) and the 3D vertex (902) is marked (as signaled or predicted). To encode the associated edge of the 3D vertex (902), it is not necessary to signal the edge between the 3D vertex (902) and the 3D vertex (903) because the edge has already been encoded.

[0137] In a third step, the 3D mesh is cut using the seam edges. In some examples, after the seam edges have been determined, the seam edges can be used to cut the reconstructed 3D mesh into connected components (also referred to as tiles or tile components) corresponding to the UV chart. In an example, the seam vertices can be traversed. At each seam vertex, the polygons or faces associated with the seam vertex are traversed in order, such as in a counter-clockwise direction, and the original index at the corner of each associated face at the seam vertex is replaced with the new vertex index. Note that the original vertex index can be used for all corners of non-seam vertices.

[0138] In some examples, a corner is a point where a polygon (also referred to as a face) connects to a vertex. In an example, a triangle has three corners, each of which has a different vertex. A vertex has the same number of corners as the number of polygons connected to the vertex. In some examples, the previous / front side or corner and the next / back side or corner are defined in counterclockwise order of the vertices.

[0139] Fig.11 and Fig.12 Schematic diagrams showing two examples for cutting a grid according to some aspects of the present application.

[0140] exist Fig.11 In the example of , the process starts with face 0 and checks the associated faces (eg, face 0, face 1, face 2, face 3, face 4, and face 5) of the seam vertex (1101) in a counterclockwise direction, for example. Fig.11In the example, from the previous face (e.g., the adjacent face in the clockwise direction, such as face5) to face face0, the seam edges (1111) are traversed, and then the corner in face face0 (e.g., at the seam vertex (1101)) is replaced with a new vertex with index "idx1" at the seam vertex. Since the edge (1112) between face face0 and the next face (e.g., face1) is not a seam edge, face1 and face0 share the same vertex at the corner of the seam vertex (1101), so the corner of face1 at the seam vertex can be maintained as vertex "idx1". The edge (1113) between face1 and face2 is a seam edge, so the corner of face2 at the seam vertex (1101) can be replaced with another new vertex with index "idx2". face2 shares a non-seam edge (1114) with face3 at the corner of the seam vertex (1101), so the corner of face3 can be maintained as vertex "idx2". The edge between face3 and face4 is the seam edge. In the example, since the original vertex index has not been used, face4 and face5 can use the original vertex index "idx0", such as Fig.11 shown.

[0141] exist Fig.12 In the example, the seam edge is Fig.11 The same as in , but the process starts with face1 and checks associated faces in a counter-clockwise direction (e.g., face1, face2, face3, face4, face5, and face0). Fig.12 In the example, from the previous face (e.g., the adjacent face in a clockwise direction, such as face0) to face1, the non-seam edge (1212) is traversed, so face1 can use the original vertex index "idx0" at the corner of the seam vertex (1201). The edge (1213) between face1 and face2 is a seam edge, and a new vertex with index "idx1" is added at the corner of the seam vertex (1201) of face2. face2 shares a non-seam edge with face3, and the corner of the seam vertex (1201) in face3 can maintain the same vertex with index "idx1" as face2. Face3 shares a seam edge with face4 (1215), and a new vertex with index "idx2" is added at the corner of the seam vertex (1201) of face4. Face4 shares a non-seam edge with face5 (1216), and the corner of the seam vertex (1201) of face5 can use the same vertex index "idx2" as face4. Face5 shares the seam edge (1211) with face0, so face5 needs to use different vertex indices for the corners of the seam vertex (1201) than face0. Fig.12In the example, face5 uses the vertex with index "idx2" and face0 uses the vertex with index "idx0".

[0142] Fig.13 A flow chart outlining a process (1300) according to an aspect of the present application is shown. The process (1300) may be used in a trellis decoder. In various aspects, the process (1300) is performed by a processing circuit, such as a processing circuit that performs the functions of a 3D decoder (110). In some aspects, the process (1300) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1300). The process begins at (S1301) and proceeds to (S1310).

[0143] At (S1310), encoded information of a mesh is received, the encoded information including position connectivity of a plurality of three-dimensional 3D vertices of the mesh in a 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in a UV space.

[0144] At (S1320), a reconstructed mesh in the 3D space is generated based on the plurality of 3D vertices according to position connectivity of the plurality of 3D vertices.

[0145] At (S1330), based on the plurality of 3D vertices, a plurality of seam vertices are determined, the plurality of seam vertices corresponding to two or more UV vertices in the UV space.

[0146] At (S1340), a plurality of seam edges in the 3D space are predicted based at least on the plurality of seam vertices, wherein at least one edge between two seam vertices is predicted as a seam edge, the seam edge corresponding to two or more UV edges in the UV space.

[0147] At (S1350), the reconstructed mesh is cut into a plurality of small components according to the plurality of seam edges.

[0148] At (S1360), UV connectivity of the plurality of UV vertices is determined based on the plurality of patch components.

[0149] In some examples, when a 3D vertex corresponds to two or more UV vertices in the UV space, the 3D vertex is determined to be a seam vertex. In some examples, when a first valence of a 3D vertex is different from a second valence of a UV vertex to which the 3D vertex corresponds, the 3D vertex is determined to be a seam vertex.

[0150] According to one aspect of the present application, a plurality of syntax elements are decoded from the encoded information, the plurality of syntax elements being respectively associated with the plurality of 3D vertices, wherein a syntax element associated with a 3D vertex has: a first potential value indicating a non-seam vertex type of the 3D vertex; a second potential value indicating a seam vertex type of the 3D vertex when a plurality of adjacent seam edges are correctly predicted; and a third potential value indicating a seam vertex type of the 3D vertex when a plurality of adjacent seam edges are not correctly predicted. The prediction of the seam edge is also based on the plurality of syntax elements respectively associated with the plurality of 3D vertices.

[0151] In some examples, based on the syntax elements in the encoded information, a first 3D vertex is determined to be a half-seam vertex, the first 3D vertex corresponding to a first UV vertex in the UV space, and then, when the second 3D vertex is a seam vertex, a first edge between the first 3D vertex and the second 3D vertex is predicted as a seam edge.

[0152] According to one aspect of the present application, determining a first syntax element associated with a first seam vertex from the encoded information;

[0153] Based on the first syntax element, one or more first real seam edges associated with the first seam vertex are determined. In some examples, the first syntax element includes a plurality of bits associated with a plurality of edges associated with the first seam vertex, wherein a bit associated with an edge indicates whether the edge is a real seam edge or an incorrectly predicted seam edge.

[0154] In some examples, a plurality of prediction residuals associated with the plurality of edges are decoded from the encoded information, and the one or more first real seam edges are determined by combining the plurality of prediction residuals and a plurality of predictions associated with the plurality of edges. In an example, the plurality of prediction residuals associated with the plurality of edges are decoded from the encoded information based on a plurality of contexts, wherein the plurality of contexts are selected based on the plurality of predictions associated with the plurality of edges.

[0155] In some examples, when the one or more first real seam edges are determined, at least a first edge associated with the first seam vertex is marked for prediction or signaling.

[0156] In an example, to cut the reconstructed mesh, when a current associated face of a seam vertex shares a seam edge connected to the seam vertex with a previous associated face, the seam vertex in the current associated face is replaced with a new vertex with a new index. In another example, when a current associated face of a seam vertex shares a non-seam edge with a previous associated face, the index of the seam vertex used in the previous associated face is maintained in the current associated face. In another example, the index of the non-seam vertex remains unchanged.

[0157] Then, the process proceeds to (S1399) and terminates.

[0158] The process (1300) may be adapted as appropriate. One or more steps in the process (1300) may be modified and / or omitted. One or more additional steps may be added. Any suitable order of implementation may be used.

[0159] Fig.14 A flow chart outlining a process (1400) according to an aspect of the present application is shown. The process (1400) may be used in a trellis encoder. In various aspects, the process (1400) is performed by a processing circuit (such as a processing circuit that performs the functions of the 3D encoder (103)). In some aspects, the process (1400) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1400). The process begins at (S1401) and proceeds to (S1410).

[0160] At (S1410), position connectivity of a plurality of three-dimensional 3D vertices of a mesh in a 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in a UV space, are encoded in the encoded information of the mesh.

[0161] At (S1420), based on the plurality of 3D vertices, a plurality of seam vertices are determined, the plurality of seam vertices corresponding to two or more UV vertices in the UV space.

[0162] At (S1430), a plurality of predicted seam edges in the 3D space are predicted based on the plurality of seam vertices, wherein at least one edge between two seam vertices is predicted as a seam edge, and the seam edge corresponds to two or more UV edges in the UV space.

[0163] At (S1440), adjustment information is encoded in the encoded information of the grid, wherein the adjustment information is used to generate a plurality of real seam edges from the plurality of predicted seam edges.

[0164] In some examples, multiple syntax elements respectively associated with the multiple 3D vertices are encoded in the encoded information of the mesh, wherein a syntax element associated with a 3D vertex has: a first potential value, the first potential value indicating a non-seam vertex type of the 3D vertex; a second potential value, the second potential value indicating a seam vertex type of the 3D vertex when multiple adjacent seam edges are correctly predicted; and a third potential value, the third potential value indicating a seam vertex type of the 3D vertex when multiple adjacent seam edges are not correctly predicted.

[0165] In some examples, when a first 3D vertex corresponds to a first UV vertex in the UV space and a first edge between the first 3D vertex and a second 3D vertex is a seam edge, a syntax element indicating that the first 3D vertex is a half-seam vertex is encoded in the encoded information of the mesh.

[0166] In some examples, a first syntax element associated with a first seam vertex is encoded in the encoded information of the mesh, wherein the first syntax element indicates a plurality of real seam edges associated with the first seam vertex. In an example, the first syntax element includes a plurality of bits associated with a plurality of edges associated with the first seam vertex, wherein a bit associated with an edge indicates whether the edge is a real seam edge or an incorrectly predicted seam edge.

[0167] In some examples, a plurality of prediction residuals associated with the plurality of edges are encoded in the encoded information of the grid, the plurality of prediction residuals being differences between a plurality of predictions associated with the plurality of edges and the plurality of true seam edges. In an example, the plurality of prediction residuals associated with the plurality of edges are encoded based on a plurality of contexts, wherein the plurality of contexts are selected based on the plurality of predictions.

[0168] In some examples, when the one or more first real seam edges associated with the first seam vertex are determined, at least a first edge associated with the first seam vertex is marked for prediction or signaling, wherein the first edge is between the first seam vertex and the second seam vertex. In addition, in an example, signaling the first edge associated with the second seam vertex may be omitted.

[0169] Then, the process proceeds to (S1499) and terminates.

[0170] The process (1400) may be modified as appropriate. At least one step in the process (1400) may be modified and / or omitted. At least one additional step may be added. Any suitable implementation order may be used.

[0171] According to one aspect of the application, a grid compression method is provided. In the method, conversion between a grid file and a compressed grid code stream is performed according to a format rule. For example, the code stream can be a code stream decoded / encoded by any decoding and / or encoding method described herein. The format rule can specify one or more constraints of the code stream and / or one or more processes to be performed by a decoder and / or an encoder.

[0172] In one example, a code stream includes encoded information of a mesh, the encoded information including position connectivity of a plurality of three-dimensional 3D vertices of the mesh in a 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in a UV space. The format rule further specifies that, based at least on a plurality of seam vertices, a plurality of seam edges in the 3D space are predicted. For example, at least one edge between two seam vertices is predicted as a seam edge, and the seam edge corresponds to two or more UV edges in the UV space. In some examples, the format rule further specifies that, based on the plurality of seam edges, a reconstructed mesh is cut into a plurality of small-piece components; and, based on the plurality of small-piece components, UV connectivity of the plurality of UV vertices is determined.

[0173] The above techniques may be implemented as computer software via computer-readable instructions and physically stored in at least one computer-readable storage medium. Fig.15 A computer system (1500) is shown that is suitable for implementing certain embodiments of the disclosed subject matter.

[0174] The computer software may be encoded in any suitable machine code or computer language, and may be assembled, compiled, linked, or the like to create a code comprising instructions, which may be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), or the like, or executed by decoding, microcode, or the like.

[0175] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablets, servers, smartphones, gaming devices, IoT devices, etc.

[0176] Fig.15 The components shown for the computer system (1500) are exemplary in nature and are not intended to limit the scope of use or functionality of the computer software implementing the embodiments of the present application. The configuration of the components should not be interpreted as having any dependency or requirement on any component or combination of components shown in the exemplary embodiment of the computer system (1500).

[0177] The computer system (1500) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from at least one human user through tactile input (e.g., keyboard input, sliding, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-computer interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0178] The human-computer interface input device may include at least one of the following (only one of which is drawn): keyboard (1501), mouse (1502), touchpad (1503), touch screen (1510), data gloves (not shown), joystick (1505), microphone (1506), scanner (1507), camera (1508).

[0179] The computer system (1500) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate at least one sense of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (1510), a data glove (not shown), or a joystick (1505), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (1509), headphones (not shown)), visual output devices (e.g., screens (1510) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light emitting diode screens, each of which has or does not have a touch screen input function, each of which has or does not have a tactile feedback function - some of which may output two-dimensional visual output or output of more than three dimensions by means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)) and printers (not shown).

[0180] The computer system (1500) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical disks (CD / DVD ROM / RW) (1520) with CD / DVD or similar media (1521), thumb drives (1522), removable hard disk drives or solid state drives (1523), traditional magnetic media such as tapes and floppy disks (not shown), special-purpose ROM / ASIC / PLD-based devices such as security software protectors (not shown), and the like.

[0181] Those skilled in the art should also understand that the term "computer-readable storage media" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0182] The computer system (1500) may also include an interface (1554) to at least one communication network (1555). For example, the network may be wireless, wired, or optical. The network may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle network, an industrial network, a real-time network, a delay-tolerant network, and the like. The network also includes local area networks such as Ethernet, wireless local area networks, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), in-vehicle and industrial networks (including CANBus), and the like. Some networks typically require an external network interface adapter for connecting to some universal data port or peripheral bus (1549) (e.g., a USB port of the computer system (1500)); other systems are typically integrated into the core of the computer system (1500) by connecting to a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smart phone computer system). By using any of these networks, the computer system (1500) can communicate with other entities. The communication can be one-way, for receiving only (e.g., wireless television), one-way for sending only (e.g., CAN bus to certain CAN bus devices), or two-way, such as to other computer systems via a local or wide area digital network. Each of the above networks and network interfaces can use certain protocols and protocol stacks.

[0183] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core (1540) of the computer system (1500).

[0184] The core (1540) may include at least one central processing unit (CPU) (1541), a graphics processing unit (GPU) (1542), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1543), a hardware accelerator for a specific task (1544), a graphics adapter (1550), etc. These devices, as well as a read-only memory (ROM) (1545), a random access memory (1546), an internal mass storage (e.g., an internal non-user accessible hard disk drive, a solid state drive, etc.) (1547), etc., may be connected via a system bus (1548). In some computer systems, the system bus (1548) may be accessed in the form of at least one physical plug so that it may be expanded by additional central processing units, graphics processing units, etc. Peripheral devices may be directly attached to the system bus (1548) of the core, or connected via a peripheral bus (1549). In one example, the screen (1510) may be connected to the graphics adapter (1550). The architecture of the peripheral bus includes a peripheral controller interface PCI, a universal serial bus USB, etc.

[0185] The CPU (1541), GPU (1542), FPGA (1543) and accelerator (1544) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (1545) or RAM (1546). Transition data can also be stored in RAM (1546), while permanent data can be stored in, for example, internal mass storage (1547). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with at least one CPU (1541), GPU (1542), mass storage (1547), ROM (1545), RAM (1546), etc.

[0186] The computer readable storage medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purpose of this application, or may be medium and code well known and available to those skilled in the art of computer software.

[0187] As an example and not a limitation, a computer system having an architecture (1500), in particular a core (1540), can provide the function of executing software contained in at least one tangible computer-readable storage medium as a processor (including a CPU, a GPU, an FPGA, an accelerator, etc.). Such a computer-readable storage medium can be a medium associated with the above-mentioned user-accessible mass storage, as well as a specific memory of the core (1540) having non-volatility, such as a core internal mass storage (1547) or a ROM (1545). Software implementing various embodiments of the present application can be stored in such a device and executed by the core (1540). Depending on specific needs, the computer-readable storage medium may include one or more storage devices or chips. The software can enable the core (1540), in particular the processor therein (including a CPU, a GPU, an FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in the RAM (1546) and modifying such a data structure according to a software-defined process. Additionally or alternatively, the computer system may provide functionality hardwired in logic or otherwise contained in circuitry (e.g., accelerator (1544)) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable storage media may include circuitry (e.g., an integrated circuit (IC)) storing execution software, circuitry containing execution logic, or both. The present application includes any suitable combination of hardware and software.

[0188] The use of "at least one" in this application is intended to include any one or combination of the elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include only A, only B, only C, or any combination thereof.

[0189] Although the present application has described a number of exemplary embodiments, various changes, arrangements and various equivalent substitutions of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design a variety of systems and methods, which, although not explicitly shown or described herein, embody the principles of the present application and are therefore within the spirit and scope of the present application.

[0190] The present application also includes the following features. These features can be combined in various ways and are not limited to the combinations mentioned below.

[0191] (1) A mesh processing method, comprising: receiving encoded information of a mesh, the encoded information comprising position connectivity of a plurality of three-dimensional (3D) vertices of the mesh in a 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in a UV space; generating a reconstructed mesh in the 3D space based on the plurality of 3D vertices according to the position connectivity of the plurality of 3D vertices; determining a plurality of seam vertices based on the plurality of 3D vertices, the plurality of seam vertices corresponding to two or more UV vertices in the UV space; predicting a plurality of seam edges in the 3D space based at least on the plurality of seam vertices, wherein at least one edge between two seam vertices is predicted to be a seam edge, the seam edge corresponding to two or more UV edges in the UV space; cutting the reconstructed mesh into a plurality of small-piece components according to the plurality of seam edges; and determining UV connectivity of the plurality of UV vertices according to the plurality of small-piece components.

[0192] (2) According to the method described in claim (1), the determining of multiple seam vertices includes at least one of the following: when a 3D vertex corresponds to two or more UV vertices in the UV space, the 3D vertex is determined as a seam vertex; and / or, when a first valence of a 3D vertex is different from a second valence of a UV vertex corresponding to the 3D vertex, the 3D vertex is determined as a seam vertex.

[0193] (3) The method according to claim (1) or (2) further includes: decoding multiple syntax elements from the encoded information, wherein the multiple syntax elements are respectively associated with the multiple 3D vertices, wherein a syntax element associated with a 3D vertex has: a first potential value, wherein the first potential value indicates a non-seam vertex type of the 3D vertex; a second potential value, wherein the second potential value indicates a seam vertex type of the 3D vertex when multiple adjacent seam edges are correctly predicted; and a third potential value, wherein the third potential value indicates a seam vertex type of the 3D vertex when multiple adjacent seam edges are not correctly predicted; and adjusting the prediction of the multiple seam edges based on the multiple syntax elements respectively associated with the multiple 3D vertices.

[0194] (4) The method according to any one of claims (1)-(3) further includes: determining, based on the syntax elements in the encoded information, that a first 3D vertex is a half-seam vertex, the first 3D vertex corresponding to a first UV vertex in the UV space; and when a second 3D vertex is a seam vertex, predicting a first edge between the first 3D vertex and the second 3D vertex as a seam edge.

[0195] (5) The method according to any one of claims (1)-(4) further includes: determining a first syntax element associated with a first seam vertex from the encoded information; and determining one or more first real seam edges associated with the first seam vertex based on the first syntax element.

[0196] (6) The method of any one of claims (1) to (5), wherein the first syntax element comprises a plurality of bits associated with a plurality of edges associated with the first seam vertex, wherein a bit associated with an edge indicates whether the edge is a true seam edge or an incorrectly predicted seam edge.

[0197] (7) The method according to any one of claims (1)-(6) further includes: decoding multiple prediction residuals associated with the multiple edges from the encoded information; and determining the one or more first real seam edges by merging the multiple prediction residuals and the multiple predictions associated with the multiple edges.

[0198] (8) The method according to any one of claims (1) to (7), wherein the decoding of the multiple prediction residuals associated with the multiple edges from the encoded information further includes: decoding the multiple prediction residuals associated with the multiple edges from the encoded information based on multiple contexts, wherein the multiple contexts are selected based on the multiple predictions associated with the multiple edges.

[0199] (9) The method according to any one of claims (1)-(8) further includes: when the one or more first real seam edges are determined, marking at least a first edge associated with the first seam vertex as predicted or signaled.

[0200] (10) The method according to any one of claims (1) to (9), wherein the step of cutting the reconstructed mesh into a plurality of small components comprises: when a current associated face of a seam vertex shares a seam edge connected to the seam vertex with a previous associated face, replacing the seam vertex in the current associated face with a new vertex having a new index.

[0201] (11) The method according to any one of claims (1) to (10), wherein the step of cutting the reconstructed mesh into a plurality of small components comprises:

[0202] When the current associated face of a seam vertex shares a non-seam edge with the previous associated face, the index of the seam vertex used in the previous associated face is retained.

[0203] (12) A mesh processing method, comprising: encoding the position connectivity of multiple three-dimensional (3D) vertices of a mesh in a 3D space, and the correspondence between the multiple 3D vertices and the multiple UV vertices of the mesh in a UV space, in the encoded information of the mesh; determining multiple seam vertices based on the multiple 3D vertices, the multiple seam vertices corresponding to two or more UV vertices in the UV space; predicting multiple predicted seam edges in the 3D space based on the multiple seam vertices, wherein at least one edge between two seam vertices is predicted to be a seam edge, the seam edge corresponding to two or more UV edges in the UV space; and encoding adjustment information in the encoded information of the mesh, the adjustment information being used to generate multiple real seam edges from the predicted multiple seam edges.

[0204] (13) The method according to claim (12), wherein the encoding of the adjustment information further includes: encoding multiple syntax elements respectively associated with the multiple 3D vertices in the encoded information of the mesh, wherein a syntax element associated with a 3D vertex has: a first potential value, wherein the first potential value indicates a non-seam vertex type of the 3D vertex; a second potential value, wherein the second potential value indicates a seam vertex type of the 3D vertex when multiple adjacent seam edges are correctly predicted; and a third potential value, wherein the third potential value indicates a seam vertex type of the 3D vertex when multiple adjacent seam edges are not correctly predicted.

[0205] (14) The method according to claim (12) or (13), wherein the encoding of the adjustment information further includes: when a first 3D vertex corresponds to a first UV vertex in the UV space and a first edge between the first 3D vertex and a second 3D vertex is a seam edge, encoding a syntax element indicating that the first 3D vertex is a half-seam vertex in the encoded information of the mesh.

[0206] (15) The method according to any one of claims (12)-(14), wherein the encoding of the adjustment information further includes: encoding a first syntax element associated with the first seam vertex in the encoded information of the mesh, wherein the first syntax element indicates a plurality of real seam edges associated with the first seam vertex.

[0207] (16) The method of any one of claims (12)-(15), wherein the first syntax element comprises a plurality of bits associated with a plurality of edges associated with the first seam vertex, wherein a bit associated with an edge indicates whether the edge is a true seam edge or an incorrectly predicted seam edge.

[0208] (17) The method according to any one of claims (12)-(16), wherein the encoding of the adjustment information further includes: encoding a plurality of prediction residuals associated with the plurality of edges in the encoded information of the grid, the plurality of prediction residuals referring to the differences between a plurality of predictions associated with the plurality of edges and the plurality of actual seam edges.

[0209] (18) A method according to any one of claims (12)-(17), wherein the multiple prediction residuals associated with the multiple edges are encoded according to multiple contexts, wherein the multiple contexts are selected based on the multiple predictions.

[0210] (19) The method according to any one of claims (12) to (18), wherein, when the one or more first real seam edges associated with the first seam vertex are determined, at least the first edge associated with the first seam vertex is marked for prediction or signaling, wherein the first edge is between the first seam vertex and the second seam vertex; and when the first edge is marked, signaling of the first edge associated with the second seam vertex is omitted.

[0211] (20) A mesh data processing method, comprising: processing a code stream of mesh data according to a format rule, wherein: the code stream includes encoded information of the mesh, the encoded information includes position connectivity of multiple three-dimensional (3D) vertices of the mesh in a 3D space, and a correspondence between the multiple 3D vertices and multiple UV vertices of the mesh in a UV space; the format rule specifies: generating a reconstructed mesh in the 3D space based on the multiple 3D vertices according to the position connectivity of the multiple 3D vertices; determining multiple seam vertices based on the multiple 3D vertices, the multiple seam vertices corresponding to two or more UV vertices in the UV space; predicting multiple seam edges in the 3D space based on at least the multiple seam vertices, wherein at least one edge between two seam vertices is predicted to be a seam edge, the seam edge corresponding to two or more UV edges in the UV space; cutting the reconstructed mesh into multiple small-piece components according to the multiple seam edges; and determining UV connectivity of the multiple UV vertices according to the multiple small-piece components.

[0212] (21) A device for grid processing, comprising a processing circuit configured to perform the method of any one of features (1) to (11).

[0213] (22) An apparatus for grid processing, comprising a processing circuit configured to perform the method of any one of features (12) to (19).

[0214] (23) A non-volatile computer-readable storage medium storing instructions which, when executed by at least one processor, cause the at least one processor to perform the method of any one of features (1) to (19).

Claims

1. A grid processing method, characterized in that: The method comprises: Receiving encoded information of a mesh, the encoded information including position connectivity of a plurality of three-dimensional 3D vertices of the mesh in a 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in a UV space; generating a reconstructed mesh in the 3D space based on the plurality of 3D vertices according to positional connectivity of the plurality of 3D vertices; Based on the plurality of 3D vertices, determining a plurality of seam vertices, the plurality of seam vertices corresponding to two or more UV vertices in the UV space; Based at least on the plurality of seam vertices, predict a plurality of seam edges in the 3D space, wherein at least one edge between two seam vertices is predicted as a seam edge, and the seam edge corresponds to two or more UV edges in the UV space; According to the plurality of seam edges, the reconstructed mesh is cut into a plurality of small components; and UV connectivity of the plurality of UV vertices is determined based on the plurality of patch components.

2. The method according to claim 1, characterized in that The determining of the plurality of seam vertices comprises at least one of the following: When a 3D vertex corresponds to two or more UV vertices in the UV space, determining the 3D vertex as a seam vertex; and / or, When a first valence of a 3D vertex is different from a second valence of a UV vertex corresponding to the 3D vertex, the 3D vertex is determined as a seam vertex.

3. The method according to claim 1, further comprising: A plurality of syntax elements are decoded from the encoded information, wherein the plurality of syntax elements are respectively associated with the plurality of 3D vertices, wherein a syntax element associated with a 3D vertex has: a first potential value, the first potential value indicating a non-seam vertex type of the 3D vertex; a second potential value indicating a seam vertex type of the 3D vertex when a plurality of adjacent seam edges are correctly predicted; a third potential value indicating a seam vertex type of the 3D vertex when a plurality of adjacent seam edges are not correctly predicted; Prediction of the plurality of seam edges is adjusted based on the plurality of syntax elements respectively associated with the plurality of 3D vertices.

4. The method according to claim 1, further comprising: Based on syntax elements in the encoded information, determining a first 3D vertex as a half-seam vertex, the first 3D vertex corresponding to a first UV vertex in the UV space; When the second 3D vertex is a seam vertex, a first edge between the first 3D vertex and the second 3D vertex is predicted as a seam edge.

5. The method according to claim 1, further comprising: determining, from the encoded information, a first syntax element associated with a first seam vertex; Based on the first syntax element, one or more first real seam edges associated with the first seam vertex are determined.

6. The method according to claim 5, wherein: The first syntax element includes a plurality of bits associated with a plurality of edges associated to the first seam vertex, wherein a bit associated with an edge indicates whether the edge is a true seam edge or an incorrectly predicted seam edge.

7. The method according to claim 6, further comprising: decoding, from the encoded information, a plurality of prediction residuals associated with the plurality of edges; The one or more first true seam edges are determined by combining the plurality of prediction residuals and the plurality of predictions associated with the plurality of edges.

8. The method according to claim 7, wherein: The step of decoding a plurality of prediction residuals associated with the plurality of edges from the encoded information further comprises: The plurality of prediction residuals associated with the plurality of edges are decoded from the encoded information according to a plurality of contexts, wherein the plurality of contexts are selected according to the plurality of predictions associated with the plurality of edges.

9. The method according to claim 5, further comprising: When the one or more first real seam edges are determined, at least a first edge associated with the first seam vertex is marked for prediction or signaling.

10. The method according to claim 1, wherein: The step of cutting the reconstructed grid into a plurality of small components comprises: When a current associated face of a seam vertex shares a seam edge connected to the seam vertex with a previous associated face, the seam vertex in the current associated face is replaced with a new vertex with a new index.

11. The method according to claim 1, wherein: The step of cutting the reconstructed grid into a plurality of small components comprises: When the current associated face of a seam vertex shares a non-seam edge with the previous associated face, the index of the seam vertex used in the previous associated face is retained.

12. A grid processing method, characterized in that: The method comprises: Encoding, in the encoded information of the mesh, position connectivity of a plurality of three-dimensional 3D vertices of the mesh in the 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in the UV space; Based on the plurality of 3D vertices, determining a plurality of seam vertices, the plurality of seam vertices corresponding to two or more UV vertices in the UV space; Predicting a plurality of predicted seam edges in the 3D space based on the plurality of seam vertices, wherein at least one edge between two seam vertices is predicted as a seam edge, and the seam edge corresponds to two or more UV edges in the UV space; and, The adjustment information is encoded in the encoded information of the grid, and the adjustment information is used to generate a plurality of real seam edges from the plurality of predicted seam edges.

13. The method according to claim 12, wherein: The encoding of the adjustment information further includes: A plurality of syntax elements respectively associated with the plurality of 3D vertices are encoded in the encoded information of the mesh, wherein a syntax element associated with a 3D vertex has: a first potential value, the first potential value indicating a non-seam vertex type of the 3D vertex; a second potential value indicating a seam vertex type of the 3D vertex when a plurality of adjacent seam edges are correctly predicted; A third potential value indicates a seam vertex type of the 3D vertex when a plurality of adjacent seam edges are not correctly predicted.

14. The method according to claim 12, wherein: The encoding of the adjustment information further includes: When a first 3D vertex corresponds to a first UV vertex in the UV space and a first edge between the first 3D vertex and a second 3D vertex is a seam edge, a syntax element indicating that the first 3D vertex is a half-seam vertex is encoded in the encoded information of the mesh.

15. The method according to claim 12, wherein: The encoding of the adjustment information further includes: A first syntax element associated with a first seam vertex is encoded in the encoded information of the mesh, wherein the first syntax element indicates a plurality of real seam edges associated with the first seam vertex.

16. The method according to claim 15, wherein: The first syntax element includes a plurality of bits associated with a plurality of edges associated to the first seam vertex, wherein a bit associated with an edge indicates whether the edge is a true seam edge or an incorrectly predicted seam edge.

17. The method according to claim 16, wherein: The encoding of the adjustment information further includes: A plurality of prediction residuals associated with the plurality of edges are encoded in the encoded information of the grid, wherein the plurality of prediction residuals refer to differences between a plurality of predictions associated with the plurality of edges and the plurality of true seam edges.

18. The method according to claim 17, wherein: The plurality of prediction residuals associated with the plurality of edges are encoded according to a plurality of contexts, wherein the plurality of contexts are selected according to the plurality of predictions.

19. The method according to claim 15, wherein: When the one or more first real seam edges associated with the first seam vertex are determined, marking at least a first edge associated with the first seam vertex for prediction or signaling, wherein the first edge is between the first seam vertex and the second seam vertex; When the first edge is marked, signaling the first edge associated with the second seam vertex is omitted.

20. A grid data processing method, characterized in that: include: Process the code stream of the grid data according to the format rules, where: The bitstream includes encoded information of the mesh, the encoded information including position connectivity of a plurality of three-dimensional 3D vertices of the mesh in a 3D space, and a correspondence between the plurality of 3D vertices and a plurality of UV vertices of the mesh in a UV space; The format rules specify: generating a reconstructed mesh in the 3D space based on the plurality of 3D vertices according to positional connectivity of the plurality of 3D vertices; Based on the plurality of 3D vertices, determining a plurality of seam vertices, the plurality of seam vertices corresponding to two or more UV vertices in the UV space; Based at least on the plurality of seam vertices, predict a plurality of seam edges in the 3D space, wherein at least one edge between two seam vertices is predicted as a seam edge, and the seam edge corresponds to two or more UV edges in the UV space; According to the plurality of seam edges, the reconstructed mesh is cut into a plurality of small components; and UV connectivity of the plurality of UV vertices is determined based on the plurality of patch components.