Dynamic Mesh Coding Method and Apparatus Based on Texture Prediction

By reusing the texture information of the previous frame in video encoding and extracting the triangular motion information of inter-frame frames using bounding boxes, the problem of excessive storage requirements in dynamic mesh sequences is solved, and a more efficient compression effect is achieved.

CN122139364APending Publication Date: 2026-06-02JIAWEN GRP CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIAWEN GRP CO LTD
Filing Date
2024-10-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies in video coding suffer from excessive storage requirements due to the independent storage of texture images, which limits the practical application of dynamic mesh sequences.

Method used

By reusing texture information from the previous frame to encode dynamic mesh sequences, and using bounding boxes to extract triangular facet motion information between frames, a bitstream is generated, reducing storage requirements.

Benefits of technology

It significantly reduces storage requirements, improves resource utilization, and achieves higher compression ratios and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122139364A_ABST
    Figure CN122139364A_ABST
Patent Text Reader

Abstract

This invention relates to an encoding method. To achieve the above objectives, the invention may include: extracting mesh information and texture information from an input first frame; receiving the mesh information and texture information to determine a bounding box; extracting triangular facet motion information of inter-frame frames using the bounding box; and generating a bitstream using the triangular facet motion information, the mesh information, and the texture information of a second frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to video compression technology, and more specifically, to a method and apparatus for efficiently compressing data by reusing the grid and texture of the previous frame to reduce the data of the generated bitstream. Background Technology

[0002] The content described in this section is only to provide background information for this embodiment and does not constitute prior art.

[0003] Video-based Dynamic Mesh Coding (V-DMC) is a technique for video encoding. More specifically, V-DMC is used to encode dynamic mesh sequences based on video. Encoding dynamic mesh sequences via V-DMC requires both geometry coding and texture coding. Storage overhead issues can arise when using textures to represent dynamic mesh sequences. Storing and using texture images separately for each frame would result in considerable storage requirements. This could limit the practical application of V-DMC due to limited storage resources. Summary of the Invention

[0004] The problem the invention aims to solve

[0005] To address the aforementioned problems, this invention provides a method for efficient texture encoding in dynamic mesh sequences. The aim is to minimize the bitstream size by reusing the texture of the previous frame.

[0006] Problem Solving Methods

[0007] To achieve the objectives of this invention, the encoding method according to embodiments of this disclosure may include: the steps of extracting mesh information and texture information from an input first frame; receiving the mesh information and texture information to determine a bounding box; using the bounding box to extract triangular facet motion information of inter-frame frames; and using the triangular facet motion information, the mesh information, and the texture information of a second frame to generate a bitstream.

[0008] In the encoding method, the step of extracting triangular facet motion information of inter-frame frames using bounding boxes may include: extracting coordinates by comparing the bounding box of the first frame with the bounding box of the second frame; estimating the triangular facet motion of inter-frame frames by comparing the coordinates; and extracting the triangular facet motion information by estimating the triangular facet motion of inter-frame frames.

[0009] The step of receiving mesh information and texture information to determine the bounding box may further include the step of calculating the minimum bounding box using the aforementioned bounding box determining factors.

[0010] The step of extracting coordinates by comparing the bounding box of the first frame with the bounding box of the second frame may further include: searching for the bounding box of the second frame to be compared using a search algorithm.

[0011] The encoding method may also include a step of determining whether to use triangular face motion information.

[0012] The step of estimating the triangular motion of inter-frame frames by comparing the coordinates may include: estimating the triangular motion using a triangular motion candidate selection algorithm.

[0013] In the step of determining whether to use the triangular face motion information, the encoding method uses a cost function to determine whether to use the aforementioned triangular face motion information.

[0014] To achieve the objectives of this invention, an encoding apparatus according to an embodiment of this disclosure includes: a memory storing at least one or more instructions; and a processor executing the at least one or more instructions; wherein the processor, by executing the at least one or more instructions, can perform the following operations: extracting mesh information and texture information from an input first frame; receiving the mesh information and texture information to determine a bounding box; extracting triangular facet motion information of inter-frame frames using the bounding box; and generating a bitstream using the triangular facet motion information, the mesh information, and the texture information of a second frame.

[0015] When the encoding device extracts triangular facet motion information of inter-frame frames using the aforementioned bounding boxes, it can extract coordinates by comparing the bounding boxes of the first frame and the second frame; estimate the triangular facet motion of inter-frame frames by comparing the aforementioned coordinates; and extract the triangular facet motion information by estimating the triangular facet motion of inter-frame frames.

[0016] The processor can perform the following operations by executing at least one of the above instructions: when receiving the above mesh information and texture information to determine the bounding box, calculate the minimum bounding box using the above-mentioned bounding box determination factors.

[0017] The processor can perform the following operations by executing at least one of the above instructions: when extracting coordinates by comparing the bounding box of the first frame with the bounding box of the second frame, a search algorithm is used to search for the bounding box of the second frame to be compared.

[0018] The encoding device can determine whether to use the triangular facet motion information.

[0019] The processor can perform the following operations by executing at least one of the above instructions: when estimating the triangular motion of inter-frame frames by comparing the above coordinates, the processor can estimate the triangular motion by using a triangular motion candidate selection algorithm.

[0020] The processor can perform the following operations by executing at least one of the above instructions: when determining whether to use the above triangular face motion information, it uses a cost function to determine whether to use the above triangular face motion information.

[0021] Invention Effects

[0022] This invention provides an efficient texture encoding method for dynamic mesh sequences. By reusing textures from the previous frame, storage requirements can be significantly reduced. This improves resource utilization in scenarios with limited storage capacity or bandwidth. Texture reuse reduces the repeated storage of similar textures, enabling higher compression ratios. Furthermore, excellent compression efficiency is achieved by efficiently representing mesh motion and deformation. Attached Figure Description

[0023] Figure 1 This is a block diagram illustrating a dynamic mesh coding method according to a first embodiment of the present disclosure.

[0024] Figure 2 This is a block diagram illustrating a dynamic mesh coding method according to a second embodiment of the present disclosure.

[0025] Figure 3 This is a block diagram illustrating a dynamic mesh coding method according to a third embodiment of the present disclosure.

[0026] Figure 4 This is a block diagram illustrating a motion estimation method according to a first embodiment of the present disclosure.

[0027] Figure 5 This is a block diagram illustrating a motion estimation method according to a second embodiment of the present disclosure.

[0028] Figure 6 This is a block diagram illustrating a dynamic mesh decoding method according to an embodiment of the present disclosure. Detailed Implementation

[0029] This invention can be modified and has various embodiments, with specific embodiments illustrated in the accompanying drawings and described in detail. However, this is not intended to limit the invention to specific implementations, but rather to encompass all modifications, equivalents, and substitutions that fall within the scope of the invention's ideas and techniques.

[0030] While terms such as "first" and "second" may be used to describe various components, the components described above should not be limited by these terms. The purpose of these terms is to distinguish one component from another. For example, without departing from the scope of this invention, a first component may be named a second component, and similarly, a second component may be named a first component. The term "and / or" includes a combination of multiple related descriptions or one of multiple related descriptions.

[0031] In embodiments of this application, "at least one of A and B" can refer to "at least one of A or B" or "at least one of a combination of more than one of A and B". Additionally, in embodiments of this application, "more than one of A and B" can refer to "more than one of A or B" or "more than one of a combination of more than one of A and B".

[0032] When a component is described as "connected" or "accessed" to another component, it should be understood that it may be directly connected to or accessed by another component, or there may be other components in between. Conversely, when a component is described as "directly connected" or "directly accessed" to another component, it should be understood that there are no other components in between.

[0033] The terminology used in this application is merely illustrative of specific embodiments and not limiting of the invention. Where there is no explicit indication of singularity in the context, the singular designation includes the meaning of plural. In this application, terms such as "comprising" or "possessing" indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, rather than precluding the presence or additional possibility of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0034] Unless otherwise specified, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms that are generally used identically to those defined in dictionaries have the same meaning as in the context of the relevant art, and unless explicitly defined, do not have an ideal or excessive meaning in this application.

[0035] Furthermore, even technologies known prior to the date of this application may be included as part of the invention if necessary, and will be described herein without departing from the scope of the invention. However, in describing the features of the invention, to avoid obscuring the purpose of the invention, overly detailed descriptions of matters obvious to those skilled in the art prior to the date of this application have been omitted.

[0036] However, the purpose of this invention is not to claim the rights to these known technologies, and the content of the known technologies may be included as part of this invention without departing from the purpose of this invention.

[0037] The preferred embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. To facilitate a comprehensive understanding of the invention, the same components will be denoted by the same reference numerals, and repeated descriptions of the same components will be omitted.

[0038] Hereinafter, preferred embodiments will be described in detail with reference to the accompanying drawings, so that those skilled in the art can easily implement the invention.

[0039] Figure 1 This is a block diagram illustrating a dynamic mesh encoding method according to a first embodiment of the present disclosure.

[0040] The dynamic mesh sequence may include frame 100. Frame 100 may be one of the frames included in the dynamic frame sequence. Frame 100 may include mesh 111 and texture 121, etc.

[0041] A mesh can include attributes such as color and normals associated with its vertices. These attributes can be associated with the mesh's surface using mapping information that parameterizes the mesh as a 2D attribute map. The mapping information can typically be described by a set of parametric coordinates, such as UV coordinates or texture coordinates, associated with the mesh vertices. The 2D attribute map can be used to store high-resolution attribute information, such as texture, normals, and displacement. The mesh can include components called geometric information, accessibility information, mapping information, vertex attributes, and attribute maps. Geometric information can be described by a set of 3D positions associated with the mesh vertices. The 3D positions of vertices can be described using (x, y, z) coordinates.

[0042] Accessibility information may include a set of vertex indices describing how vertices are connected to generate a 3D surface. Mapping information describes how the mesh surface is mapped to a planar 2D region. Mapping information (also known as UV mapping or texture mapping) may be described together with accessibility information by a set of UV parameters / texture coordinates (u,v) associated with the mesh vertices. Vertex attributes may include scalar or vector attribute values ​​associated with the mesh vertices. Attribute graphs may include attributes associated with the mesh surface and stored as 2D images / videos. The mapping between video (e.g., 2D images / videos) and the mesh surface may be defined by mapping information.

[0043] A mesh can represent a static mesh comprising a single frame. Therefore, meshes and frames are used interchangeably. A dynamic mesh comprises multiple frames and may include at least one preset keyframe. A keyframe is a frame that does not reference other frames during encoding / decoding. The remaining frames, other than keyframes, are called non-keyframes.

[0044] A dynamic mesh can be at least one of the components (geometric information, accessibility information, mapping information, vertex attributes, and attribute graphs) that changes over time. A dynamic mesh can be described by a sequence of meshes. A dynamic mesh can be referred to as a mesh frame. Because dynamic meshes can include a large amount of information that changes over time, they may require a large amount of data.

[0045] Dynamic meshes can possess accessibility information, time-varying geometry, and time-varying vertex properties. Dynamic meshes can contain time-varying accessibility information. Digital content generation tools can typically generate dynamic meshes with time-varying attribute maps and time-varying accessibility information. Volumetric acquisition techniques can be used to generate dynamic meshes. Volumetric acquisition techniques are particularly effective at generating dynamic meshes with time-varying accessibility information under real-time constraints. However, due to the presence of this information, the mesh needs to be compressed.

[0046] Mesh 111 can be generated based on frame 100. Mesh compression 112 offers various techniques. These techniques can be used for various mesh compression methods, including static mesh compression, dynamic mesh compression, dynamic mesh compression with accessibility information, dynamic mesh compression with time-varying accessibility information, and dynamic mesh compression with time-varying attribute graphs. These techniques can be used for lossy and lossless compression in various applications such as real-time communication, storage, free-viewpoint video, augmented reality (AR), and virtual reality (VR). These applications may include features such as random access and scalable / progressive coding.

[0047] A reconstructed mesh 113 can be generated by reconstructing a compressed mesh. A mesh bitstream 119 can be generated based on a compressed mesh. A mesh bitstream can refer to representing the information contained in the mesh in bits.

[0048] Texture 121 can be generated based on a frame. Textures can be combined with meshes to form frames. Texture Packing 122 refers to packing textures extracted from frames using a reconstructed mesh. Texture Atlas 123 can be generated from the packed textures.

[0049] The encoded texture coordinates 129 can be generated based on the packed texture. In other words, texture coordinates can be generated based on the packed texture. The generated texture coordinates can be encoded to be used to generate a bitstream.

[0050] Video compression (132) can refer to encoding filled textures (texture atlases) based on suitable video coding standards such as HEVC and VVC.

[0051] The encoded texture video bitstream 139 can be generated based on compressed video. The dynamic mesh bitstream generation module 169 can merge the encoded texture coordinates, the encoded texture video bitstream, and the mesh bitstream to generate the final output.

[0052] Figure 2 This is a block diagram illustrating a dynamic mesh coding method according to a second embodiment of the present disclosure.

[0053] The dynamic mesh sequence may include frames. Frame 200 may be one of the frames included in the dynamic frame sequence. Frame 200 may include mesh 211 and texture 221, etc.

[0054] Mesh 211 can be generated based on frame 200. Mesh compression 212 can be generated based on the mesh extracted from the frame. Reconstructed mesh 213 can be generated from the reconstructed compressed mesh. The reconstructed mesh can be used for texture packing. In addition, the reconstructed mesh (TriangleFace) can be used to generate triangular faces. Mesh bitstream 219 can be generated from the compressed mesh.

[0055] Texture 221 can be combined with a mesh to form a frame. Alternatively, textures can be generated based on frames. Texture packing 222 can refer to packing textures extracted from frames using a reconstructed mesh.

[0056] Texture atlas 223 can be generated from packaged textures. When compressing frames in a dynamic mesh sequence, if it is the first frame, texture atlas 223 can be stored in texture motion search buffer 243. Alternatively, when compressing frames in a dynamic mesh sequence, if it is the first frame, texture atlas 223 can be compressed into video. On the other hand, when compressing frames in a dynamic mesh sequence, if it is not the first frame, texture atlas 223 can be passed to the texture coordinate inter-prediction module.

[0057] Encoded texture coordinates 229 can be generated based on packed textures. The generated texture coordinates can be encoded to generate a bitstream. Texture coordinates can be generated based on packed textures. In this case, encoded texture coordinates for the first frame can only be generated based on texture packing if the compressed frame is the first frame.

[0058] Video compression (232) can refer to encoding filled textures (texture atlases) based on suitable video coding standards such as HEVC and VVC. Video compression can be performed only during the compression process of the first frame.

[0059] The encoded single-texture bitstream 239 may refer to a bitstream generated from the texture of the first frame using compressed video.

[0060] The texture coordinate inter-frame prediction module 240 can refer to predicting texture motion in order to reuse textures in subsequent frames of a dynamic mesh sequence. A subsequent frame can refer to a frame that is neither the first frame nor the current frame. Alternatively, a subsequent frame can be referred to as the second frame. The current frame can be referred to as the first frame.

[0061] The texture coordinate inter-frame prediction module 240 can store texture coordinates extracted from frames for texture reuse. Additionally, the texture coordinate inter-frame prediction module can store textures (images) extracted from frames. The texture coordinate inter-frame prediction module 240 can match texture coordinates with textures. This embodiment minimizes storage requirements while maintaining visual quality through the texture coordinate inter-frame prediction module.

[0062] The texture coordinate inter-frame prediction module 240 can perform at least one of the following processes: generating triangle faces 241, generating bounding boxes 242, generating texture motion search buffers 243, estimating triangle face motion 244, generating inter-frame texture coordinates 245, and generating inter-frame triangle face vectors 246. Furthermore, the texture coordinate inter-frame prediction module 240 may include a bounding box determination module, an inter-frame triangle face motion estimation module, and a texture coordinate determination module.

[0063] Triangle faces 241 can be generated using a reconstructed mesh. A triangle face may include coordinate information and vertex information. The bounding box determination module can determine a bounding box 242 for each triangle face based on its coordinates. The bounding box may be called a rectangular bounding block. The bounding box can be generated based on the triangle face information. Additionally, the bounding box may contain triangle faces. The bounding box can be determined by its defining factors. These factors may include mesh information (vertices, coordinates), the width of the triangle face, and the height of the triangle face. The bounding box can be the smallest rectangle including the triangle faces. The bounding box determination module can determine and output the bounding box using the coordinates of the top-left corner of the mesh and the dimensions of the triangle face. The dimensions of the triangle face may include its width and height.

[0064] The bounding box determination module can define specific regions using bounding boxes. This allows motion search to be focused on specific regions, thereby reducing computational complexity and improving efficiency. Furthermore, the bounding box determination module can perform a search by comparing only the specific region with the second frame, achieving efficient motion estimation.

[0065] The texture motion search buffer 243 refers to storing the texture atlas in the texture coordinate inter-frame prediction module when compressing a dynamic mesh sequence, if it is the first frame. The texture motion search buffer can utilize the stored texture atlas and use it for triangular face motion estimation 244.

[0066] Inter-frame triangular facet motion estimation 244 can estimate the motion of the triangular faces based on texture atlas 233, rectangular bounding box 242, and texture motion search buffer. Triangular facet motion estimation can output inter-frame texture coordinates 245. Additionally, triangular facet motion estimation can output inter-frame triangular facet vectors 246. Triangular facet motion estimation 244 can output at least one of inter-frame texture coordinates 245 and inter-frame triangular facet vectors 246.

[0067] Inter-frame texture coordinates 245 can refer to the texture coordinates of the second frame. Inter-frame triangle face vector 246 can refer to the motion vector between the triangle face of the second frame and the triangle face of the first frame. Encoded inter-frame texture coordinates 248 can be generated by encoding the inter-frame texture coordinates. Encoded inter-frame triangle face vector 247 can be generated by encoding the inter-frame triangle face vector.

[0068] On the other hand, the texture coordinate inter-frame prediction module 240 can transmit information used for mesh generation. The mesh can be generated based on the mesh information sent by the texture coordinate inter-frame prediction module.

[0069] The dynamic mesh bitstream generation 269 can be generated based on the texture coordinates of the first frame, the mesh bitstream, the single texture bitstream, and the texture coordinates of the second frame. Additionally, the dynamic mesh bitstream generation 269 can be generated based on the texture coordinates 229 of the first frame, the mesh bitstream 219, the single texture bitstream 239, and the triangular face motion information 249.

[0070] Figure 3 This is a block diagram illustrating a dynamic mesh coding method according to a third embodiment of the present disclosure.

[0071] The dynamic mesh sequence may include frames. Frame 300 may be one of the frames included in the dynamic frame sequence. Frame 300 may include mesh 311 and texture 321, etc.

[0072] Mesh 311 can be generated based on frame 300. Alternatively, mesh 311 can be generated based on information transmitted to the texture coordinate inter-frame prediction module.

[0073] Mesh compression 312 offers a variety of techniques. These techniques can be used for various mesh compression methods, including static mesh compression, dynamic mesh compression, dynamic mesh compression with accessibility information, dynamic mesh compression with time-varying accessibility information, and dynamic mesh compression with time-varying attribute graphs. These techniques can be used for lossy and lossless compression in various applications such as real-time communication, storage, free-viewpoint video, augmented reality (AR), and virtual reality (VR). These applications may include features such as random access and scalable / progressive coding. The mesh bitstream 319 can be generated from the compressed mesh.

[0074] Reconstructed mesh 313 can be generated by reconstructing compressed meshes. Reconstructed meshes can be used for texture packing.

[0075] Texture 321 can be combined with a mesh to form a frame. Alternatively, textures can be extracted from frames. Texture packing 322 can refer to packing textures extracted from frames using a reconstructed mesh.

[0076] Texture atlas 323 can be generated from the packaged textures. When compressing frames in a dynamic mesh sequence, if it is not the first frame, the reconstructed mesh information and texture atlas can be transferred to the texture coordinate inter-prediction module. At this point, the reconstructed mesh information and texture atlas can be combined.

[0077] Encoded texture coordinates 329 can be generated based on the packed texture. The generated texture coordinates can be encoded to generate a bitstream. Texture coordinates can be generated based on the packed texture. In this case, when the compressed frame is the first frame, the encoded texture coordinates of the first frame can be generated based on the texture packing. In addition, when the texture intra-frame / inter-frame coordinate determination module uses intra-frame texture coordinates, the encoded texture coordinates of the current frame can be generated based on the texture packing.

[0078] Video compression 332 refers to encoding filled textures (texture atlases) based on suitable video coding standards such as HEVC and VVC. When compressing dynamic mesh sequences, video compression can be performed on the first frame. Additionally, when the texture intra-frame / inter-frame coordinate determination module uses intra-frame texture coordinates, video compression can be performed based on the texture atlas of the current frame.

[0079] The encoded single-texture bitstream 339 can refer to a bitstream generated using the texture of the first frame or the current frame from compressed video.

[0080] The texture coordinate inter-frame prediction module predicts and / or estimates the motion of triangular faces to enable the reuse of textures in the second frame of a dynamic mesh sequence. This module can store texture coordinates within a frame for texture reuse. Additionally, it can store the texture (image) within a frame. The module matches texture coordinates to textures. This approach minimizes storage requirements while maintaining visual quality.

[0081] The texture coordinate inter-frame prediction module 340 may correspond to the texture coordinate inter-frame prediction module 240. The texture coordinate inter-frame prediction module 340 may receive reconstructed mesh information, texture atlas, and at least one of the texture atlases of the second frame.

[0082] The texture intra-frame / inter-frame coordinate determination module 350 compares the current frame with the full-frame cost function evaluation. This determines whether to use inter-frame predicted coordinates or intra-frame coordinate encoding. When using this method, the encoder can use either intra-frame coordinates or inter-frame predicted coordinates simultaneously.

[0083] The texture intra-frame / inter-frame coordinate determination module 350 can determine whether to use either inter-frame predicted coordinates or intra-frame texture coordinates. The texture intra-frame / inter-frame coordinate determination module 350 can use the first frame to compare the full-frame cost function evaluation. Alternatively, the texture intra-frame / inter-frame coordinate determination module allows the encoder to use both intra-frame coordinates and inter-frame predicted coordinates simultaneously. Inter-frame predicted coordinates can refer to triangulation motion information. On the other hand, the texture intra-frame / inter-frame coordinate determination module 350 can use a cost function to determine either intra-frame coordinates or inter-frame predicted coordinates. Furthermore, the texture intra-frame / inter-frame coordinate determination module can use a loss threshold to determine the coordinates.

[0084] When the texture intra-frame / inter-frame coordinate determination module determines to use intra-frame texture coordinates, it can transmit the texture atlas of the current frame to the texture coordinate inter-frame prediction module. The texture atlas of the current frame transmitted to the texture coordinate inter-frame prediction module can be used to compress the next frame.

[0085] Additionally, to use intra-frame texture coordinates, a process called intra-frame texture coordinate module 330 can be performed. Intra-frame texture coordinate module 330 can compress textures or texture coordinates intra-frame. Intra-frame texture coordinate module 330 can consist of encoded current texture coordinates 329, a grid bitstream 319, and a video compression module 332. When the intra-frame / inter-frame texture coordinate determination module uses intra-frame coordinates, dynamic grid bitstream generation 369 can be generated based on the texture coordinates, grid bitstream, and texture bitstream of the current frame. When the intra-frame / inter-frame texture coordinate determination module uses inter-frame coordinates, dynamic grid bitstream generation 369 can be generated based on the texture coordinates, grid bitstream, and triangular facet motion information 349 of the first frame. Triangular facet motion information 349 can include the texture coordinates and triangular facet vectors of the second frame.

[0086] Figure 4 This is a block diagram illustrating a motion estimation method according to a first embodiment of the present disclosure.

[0087] The inter-frame triangular facet motion estimation module 400 may include a triangular facet motion estimation module. The inter-frame triangular facet motion estimation module 400 can extract triangular facet motion information using bounding box data.

[0088] The inter-frame triangulation motion estimation module 410 can combine texture atlases and bounding boxes. The inter-frame triangulation motion estimation module 410 can use bounding box data. The inter-frame triangulation motion estimation module 410 can use pixels of the texture in the bounding box of the first frame as reference points. The current frame can be referred to as the first frame. The reference frame can be referred to as the second frame. The second frame can refer to the first frame or the frame preceding the current frame. For example, if the first frame is the fifth frame, then the second frame can refer to the fourth frame or the first frame. The reference frame can refer to the frame used to predict the bounding box of the current frame.

[0089] The inter-frame triangulation motion estimation module 410 can compare the bounding boxes of the first frame and the second frame. The inter-frame triangulation motion estimation module 410 can perform the bounding box comparison process by moving a predetermined search region to the next coordinate to be compared using various search algorithms. Furthermore, the inter-frame triangulation motion estimation module 410 can perform the bounding box comparison process by moving the entire region of the frame to the next coordinate to be compared using various search algorithms. The search algorithms may include at least one of full search, hierarchical search, diamond search, and logarithmic search.

[0090] The triangular facet motion estimation module can use the atlas from the first frame. Additionally, it can use the atlas from the second frame. The first frame may contain bounding boxes. Each bounding box may contain triangular faces. The module can search within the first frame for regions corresponding to bounding boxes containing triangular faces within that frame. Triangular faces can be compared within these corresponding regions.

[0091] The inter-frame triangular surface motion estimation module 410 can extract coordinates by comparing bounding boxes using a search algorithm.

[0092] The inter-frame triangular facet motion estimation module 400 may include a motion candidate selection algorithm. Triangular facet motion can be estimated using extracted coordinates. Triangular facet motion estimation can utilize texture atlases to estimate motion. After extracting coordinates through a search, the motion candidate selection algorithm can compare specific coordinates of the first frame with the searched coordinates, finding the minimum SAD or MSE. The motion candidate selection algorithm can compare the bounding box of the first frame with the previous bounding box as pixel data. By comparing the coordinates of the bounding boxes in the first frame with the coordinates of the bounding boxes in the second frame, the motion vector can be calculated. The triangular facet motion estimation module can output the bounding box coordinates of the second frame. The motion candidate selection algorithm can accurately identify the motion representing the movement of the bounding box in the second frame.

[0093] Additionally, the motion candidate selection algorithm can compare the bounding boxes of the first and second frames within the entire frame using the Sum of Absolute Difference (SAD). The motion candidate selection algorithm can also compare the bounding boxes of the first and second frames within the entire frame using the Mean Square Error (MSE).

[0094] The motion candidate selection algorithm compares the textures of the first and second frames within the entire frame using the pixels of the triangular faces. Alternatively, it can evaluate and compare the textures of the first and second frames within the entire frame using the pixels of the bounding boxes.

[0095] This allows for the calculation of inter-frame texture coordinates. The texture coordinates of the first frame can be obtained. Additionally, the triangular face vectors of the first frame can be obtained from this. A triangular face vector can refer to the vector between a triangular face in the first frame and the corresponding triangular face in the first frame. After searching for coordinates, the triangular face motion estimation module can use a motion selection algorithm to obtain the motion vector of the bounding box. The inter-frame triangular face motion estimation module can accurately estimate the motion or deformation of the frame mesh.

[0096] On the other hand, triangular facet motion information can refer to texture coordinates or triangular facet vectors.

[0097] Figure 5 This is a block diagram illustrating a motion estimation method according to a second embodiment of the present disclosure.

[0098] The current frame 510 may include a bounding box. The bounding box may include triangles. The triangles may include a search point 511. The search point may refer to the top-left vertex.

[0099] The second frame may include a bounding box. The bounding box of the second frame may correspond to the bounding box of the first frame. The bounding box of the second frame may include triangular faces. The triangular faces of the second frame may include triangular candidates 531, deformable candidates 532, and rotated candidates 533. A deformable candidate may refer to a triangular candidate that has been deformed. A rotated candidate 533 may refer to a triangular candidate rotated by 0 degrees, 90 degrees, 180 degrees, or 270 degrees. The bounding box of the second frame may include a candidate point 521. A candidate point may refer to the top-left vertex of a triangular face.

[0100] Figure 6 This is a block diagram illustrating a dynamic mesh decoding method according to an embodiment of the present disclosure.

[0101] The dynamic mesh bitstream 610 can refer to the encoded dynamic mesh sequence. The dynamic mesh bitstream may include texture bitstream, mesh bitstream, and intra-frame / inter-frame coordinate syntax.

[0102] The encoded texture bitstream 621 can be generated from the dynamic mesh bitstream. Video decoding 622 can be generated from the decoded texture bitstream. The texture atlas 623 can be generated from the decoded video information.

[0103] The Intra-Inter Coordinating Syntax (ITC) 631 can be generated from a dynamic mesh bitstream. It generates encoded intra-texture coordinates, triangular facet motion information, and triangular facet vectors.

[0104] Triangle Face Motion (632) can be generated using intra-frame / inter-frame coordinate syntax. Texture coordinates can refer to the texture coordinates of the second frame. The second frame can refer to a subsequent frame. The texture coordinates of the second frame can generate encoded inter-frame texture coordinates.

[0105] Encoded inter-frame texture coordinates 633 can be generated based on triangular facet motion information. These encoded inter-frame texture coordinates can be used for texture mapping. When used for texture mapping, the latest texture can be used as a reference. Encoded intra-frame texture coordinates 635 can be generated based on intra-inter-frame coordinate syntax. These encoded intra-frame texture coordinates can be used for texture mapping. When mapping using encoded intra-frame texture coordinates, a texture buffer refresh can be requested.

[0106] Mesh bitstream 641 can be generated based on dynamic mesh bitstream. Mesh bitstream can be used in the mesh decoding process. Mesh decoding 642 refers to generating a mesh using the mesh bitstream. Reconstructing the mesh 643 can be based on the mesh bitstream.

[0107] Triangle Texture Mapping (650) can be generated based on texture atlases, encoded inter-frame texture coordinates, and encoded intra-frame texture coordinates. Re-Coloring (651) can refer to changing the color of the texture in the texture map.

[0108] The terms and words used in this specification and claims described above should not be limited to their usual or dictionary meanings. Based on the principle that inventors may appropriately define terms and concepts in order to best explain their own invention, they should be interpreted as meanings and concepts consistent with the technical ideas of this invention.

[0109] Therefore, the features shown in the accompanying drawings and embodiments described in this specification are only one preferred embodiment of the present invention and do not represent all the technical ideas of the present invention. It should be understood that various equivalents and modifications that can replace these may exist at the time of filing this application.

Claims

1. An encoding method, comprising: The steps to extract mesh and texture information from the first input frame; The step of receiving the mesh information and texture information to determine the bounding box; The step of extracting triangular face motion information between frames using the bounding box; as well as The step of generating a bitstream using the triangular face motion information, the mesh information, and the texture information of the second frame.

2. The encoding method according to claim 1, wherein the step of extracting the triangular facet motion information of inter-frame frames using the bounding box includes: The step of extracting coordinates by comparing the bounding box of the first frame with the bounding box of the second frame; The step of estimating the triangular surface motion between frames by comparing the coordinates; as well as The step of extracting the triangular face motion information by estimating the triangular face motion of the inter-frame frames.

3. The encoding method according to claim 1, the step of receiving the mesh information and texture information to determine the bounding box further includes: The step of calculating the minimum bounding box using the determining factors of the bounding box.

4. The encoding method according to claim 2, the step of extracting coordinates by comparing the bounding box of the first frame with the bounding box of the second frame, further includes: The step of searching for the bounding box of the second frame to be compared using a search algorithm.

5. The encoding method according to claim 2 further includes: The step of determining whether to use the triangular facet motion information.

6. The encoding method according to claim 2, wherein the step of estimating the triangular motion of inter-frame frames by comparing the coordinates includes: The steps of estimating the triangular face motion using a triangular face motion candidate selection algorithm.

7. In the encoding method according to claim 5, in the step of determining whether to use the triangular facet motion information, The cost function is used to determine whether to use the triangular facet motion information.

8. An encoding device, comprising: Memory that stores at least one instruction; as well as The processor executes at least one of the instructions; and, The processor performs the following operations by executing at least one or more instructions: Extract mesh and texture information from the first input frame; Receive the mesh information and texture information to determine the bounding box; The triangular motion information of inter-frame frames is extracted using the bounding box; A bitstream is generated using the triangular face motion information, the mesh information, and the texture information of the second frame.

9. The encoding apparatus of claim 8, wherein the processor performs the following operations by executing the at least one or more instructions: When extracting triangular facet motion information between frames using the bounding box, Coordinates are extracted by comparing the bounding box of the first frame with the bounding box of the second frame. The triangular motion between frames is estimated by comparing the coordinates. The triangular facet motion information is extracted by estimating the triangular facet motion of the inter-frame frames.

10. The encoding apparatus of claim 9, wherein the processor performs the following operations by executing the at least one or more instructions: When receiving the mesh information and texture information to determine the bounding box, the determining factors of the bounding box are used to calculate the minimum bounding box.

11. The encoding apparatus of claim 9, wherein the processor performs the following operations by executing the at least one or more instructions: When extracting coordinates by comparing the bounding box of the first frame with the bounding box of the second frame, a search algorithm is used to search for the bounding box of the second frame to be compared.

12. The encoding apparatus of claim 9, wherein the processor performs the following operations by executing the at least one or more instructions: Determine whether to use the triangular facet motion information.

13. The encoding apparatus of claim 9, wherein the processor performs the following operations by executing the at least one or more instructions: When estimating the triangular motion between frames by comparing the coordinates, the triangular motion is estimated using a triangular motion candidate selection algorithm.

14. The encoding apparatus of claim 12, wherein the processor performs the following operations by executing the at least one or more instructions: When determining whether to use the triangular face motion information, a cost function is used to determine whether to use the triangular face motion information.

15. A decoding method, comprising: The steps for receiving a dynamic grid bitstream; The step of extracting grid information from the dynamic grid bitstream; The step of extracting texture information from the dynamic mesh bitstream; The step of extracting triangular face motion information from the dynamic mesh bitstream; The steps of reconstructing the mesh using the mesh information; The step of performing texture mapping using the texture information and the triangular facet motion information; as well as The step of generating the first frame using the reconstructed mesh and the result of the texture mapping.

16. The decoding method according to claim 15, further comprising: The steps for extracting intra-frame and inter-frame coordinate syntax from the dynamic mesh bitstream. The steps for extracting the motion information of the triangular face include: The step of extracting the motion information of the triangular face using the intra-frame / inter-frame coordinate syntax. The steps for performing the texture mapping include: The step of using the triangular face motion information to perform triangular face texture mapping.

17. The decoding method according to claim 16, characterized in that, The steps for extracting the motion information of the triangular face include: The steps for generating texture coordinates for the second frame using the intra-frame / inter-frame coordinate syntax; and The steps of generating inter-frame texture coordinates using the texture coordinates of the second frame; and The step of generating the triangular facet motion information containing the inter-frame texture coordinates; and, The inter-frame texture coordinates are used in the texture mapping step.

18. The decoding method according to claim 16, characterized in that, The steps for extracting the motion information of the triangular face include: The steps for generating intra-frame texture coordinates using the intra-frame / inter-frame coordinate syntax; and The step of generating the triangular facet motion information containing the intra-frame texture coordinates; and, The intra-frame texture coordinates are used in the texture mapping step.

19. The decoding method according to claim 15, characterized in that, Also includes: The step of generating a texture atlas using the texture information; and, The texture atlas is used in the texture mapping step.

20. A decoding device, comprising: Memory that stores at least one instruction; as well as The processor executes at least one of the instructions; and, The processor performs the following operations by executing at least one or more instructions: Receive dynamic grid bitstream; Extract grid information from the dynamic grid bitstream; Extract texture information from the dynamic mesh bitstream; Extract triangular face motion information from the dynamic grid bitstream; Reconstruct the mesh using the mesh information; Texture mapping is performed using the texture information and the triangular facet motion information; The first frame is generated by combining the reconstructed mesh and texture mapping results.