Method and device for dynamic mesh coding using texture prediction
The dynamic mesh coding method addresses storage overhead issues in video compression by reusing texture information from previous frames, achieving efficient compression and improved resource utilization.
Patent Information
- Application Number
- PCT/KR2024/016650
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-08
AI Technical Summary
Existing video compression technologies, such as V-DMC, face storage overhead issues due to the need for separate texture images for each frame in dynamic mesh sequences, limiting their application in scenarios with restricted storage or bandwidth.
A dynamic mesh coding method that efficiently codes texture information by reusing the texture of the previous frame, using triangular motion information and boundary box calculations to minimize the bitstream size.
This method significantly reduces storage requirements by reusing texture information, improves resource utilization in constrained environments, and achieves higher compression efficiency by effectively representing mesh movement and deformation.
Smart Images

Figure KR2024016650_08052025_PF_FP_ABST
Abstract
Description
Method and device for dynamic mesh coding using texture prediction
[0001] The present invention relates to a technology for compressing video, and more particularly, to a method and device for effectively compressing data by reducing data in a bitstream generated by reusing meshes and textures of previous frames.
[0002] The material described in this section merely provides background information for the present embodiment and does not constitute prior art.
[0003] Video-based Dynamic Mesh Coding (V-DMC) is a technology for coding video. More specifically, V-DMC is used to code video-based dynamic mesh sequences. Coding dynamic mesh sequences using V-DMC requires geometry coding and texture coding. Using textures to represent dynamic mesh sequences can result in storage overhead. Storing and using a separate texture image for each frame can result in significant storage requirements. Limited storage resources can limit the practical application of V-DMC.
[0004] To address the above issues, the present invention can propose a method for efficient texture coding in dynamic mesh sequences. The purpose may be to minimize the size of the bitstream by reusing textures from previous frames.
[0005] An encoding method according to an embodiment of the present disclosure for achieving the object of the present invention may include a step of extracting mesh information and texture information from an input first frame, a step of receiving the mesh information and texture information and determining a bounding box, a step of extracting triangular motion information of an inter-frame using the bounding box, and a step of generating a bitstream using the triangular motion information, the mesh information, and the texture information of a second frame.
[0006] The encoding method may include a step of extracting triangular motion information of an inter-frame using a bounding box, a step of extracting coordinates by comparing a bounding box of the first frame with a bounding box of the second frame, a step of estimating triangular motion of the inter-frame by comparing the coordinates, and a step of extracting the triangular motion information through triangular motion estimation of the inter-frame.
[0007] The step of determining a bounding box by inputting mesh information and texture information may further include a step of calculating a minimum bounding box using the determination elements of the bounding box.
[0008] The step of extracting coordinates by comparing the bounding box of the first frame with the bounding box of the second frame may further include the step of searching for the bounding box of the second frame to be compared through a search algorithm.
[0009] The encoding method may further include a step of determining whether to use triangular motion information.
[0010] The step of estimating the triangle motion of the inter-frame by comparing the coordinates may include the step of estimating the triangle motion through a triangle motion candidate selection algorithm.
[0011] An encoding method, wherein the encoding method uses a cost function in a step of determining whether to use the triangle motion information to determine whether to use the triangle motion information.
[0012] An encoding device according to an embodiment of the present disclosure for achieving the object of the present invention includes a memory storing at least one command; and a processor executing the at least one command, wherein the processor extracts mesh information and texture information from an input first frame by executing the at least one command, determines a bounding box by receiving the mesh information and texture information, extracts triangular motion information of an inter-frame using the bounding box, and generates a bitstream using the triangular motion information, the mesh information, and the texture information of a second frame.
[0013] The encoding device can extract triangular motion information of an inter-frame using the bounding box, compare the bounding box of the first frame with the bounding box of the second frame to extract coordinates, estimate triangular motion of the inter-frame by comparing the coordinates, and extract the triangular motion information through triangular motion estimation of the inter-frame.
[0014] The processor can calculate a minimum bounding box by using the determination elements of the bounding box when determining a bounding box by receiving the mesh information and texture information by executing at least one of the above commands.
[0015] The processor can extract coordinates by comparing the bounding box of the first frame with the bounding box of the second frame by executing at least one command, and can search for the bounding box of the second frame to be compared through a search algorithm.
[0016] The encoding device can determine whether to use triangular motion information.
[0017] The processor can estimate the triangle motion of the inter-frame by comparing the coordinates by executing at least one command, and can estimate the triangle motion through a triangle motion candidate selection algorithm.
[0018] The processor may determine whether to use the triangle motion information by using a cost function when determining whether to use the triangle motion information by executing at least one of the above instructions.
[0019] The present invention can be an efficient texture coding method for dynamic mesh sequences. By reusing textures from previous frames, storage requirements can be significantly reduced. This can improve resource utilization in scenarios with limited storage capacity or bandwidth. Texture reuse can reduce redundant storage for similar textures and enable higher compression ratios. Furthermore, mesh motion and deformation can be efficiently expressed, achieving excellent compression efficiency.
[0020] FIG. 1 is a block diagram illustrating a dynamic mesh encoding method according to a first embodiment of the present disclosure.
[0021] FIG. 2 is a block diagram illustrating a dynamic mesh encoding method according to a second embodiment of the present disclosure.
[0022] FIG. 3 is a block diagram illustrating a dynamic mesh encoding method according to a third embodiment of the present disclosure.
[0023] FIG. 4 is a block diagram illustrating a motion estimation method according to a first embodiment of the present disclosure.
[0024] FIG. 5 is a block diagram illustrating a motion estimation method according to a second embodiment of the present disclosure.
[0025] FIG. 6 is a block diagram illustrating a dynamic mesh decryption method according to one embodiment of the present disclosure.
[0026] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.
[0027] While terms such as first, second, etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component could be referred to as the second component, and similarly, the second component could also be referred to as the first component. The term "and / or" includes any combination of multiple related items described or any one of multiple related items described.
[0028] In the embodiments of the present application, “at least one of A and B” may mean “at least one of A or B” or “at least one of combinations of one or more of A and B.” Furthermore, in the embodiments of the present application, “at least one of A and B” may mean “at least one of A or B” or “at least one of combinations of one or more of A and B.”
[0029] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0030] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0031] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0032] Meanwhile, even if a technology was known prior to the filing date of this application, it may be included as part of the composition of the invention of this application, if necessary, and such technology will be described in this specification to the extent that it does not obscure the spirit of the invention. However, in describing the composition of the invention of this application, a detailed description of matters that were known prior to the filing date of this application and would be clearly understood by those skilled in the art may obscure the spirit of the invention, and therefore, an excessively detailed description of the known technology will be omitted.
[0033] However, the purpose of the present invention is not to claim rights to these known technologies, and the contents of the known technologies may be included as part of the present invention within a scope that does not deviate from the purpose of the present invention.
[0034] Hereinafter, with reference to the attached drawings, preferred embodiments of the present invention will be described in more detail. In order to facilitate an overall understanding in describing the present invention, identical reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted.
[0035] Hereinafter, a preferred embodiment of the present invention will be described in detail with reference to the attached drawings so that a person having ordinary knowledge in the same technical field can easily practice the present invention.
[0036]
[0037] FIG. 1 is a block diagram illustrating a dynamic mesh encoding method according to a first embodiment of the present disclosure.
[0038] A dynamic mesh sequence may include frames (100). The frame (100) may be one of the frames included in the dynamic frame sequence. The frame (100) may include a mesh (111), a texture (121), etc.
[0039] A mesh may contain attributes, such as color, normals, etc., associated with vertices. The attributes may be associated with the surface of the mesh by using mapping information that parameterizes the mesh with 2D attribute maps. The mapping information may be generally described by a set of parameter coordinates, referred to as UV coordinates or texture coordinates, associated with mesh vertices. The 2D attribute maps may be used to store high-resolution attribute information, such as textures, normals, displacements, etc. A mesh (111) may contain components referred to as geometry information, connectivity information, mapping information, vertex attributes, and attribute maps. The geometry information may be described by a set of 3D positions associated with the vertices of the mesh. (x,y,z) coordinates may be used to describe the 3D positions of the vertices.
[0040] Connectivity information may include a set of vertex indices that describe how to connect vertices to create a 3D surface. Mapping information may describe how to map the mesh surface to 2D regions of a plane. Mapping information (also referred to as UV mapping, texture mapping) may be described by a set of UV parameters / texture coordinates (u, v) associated with mesh vertices along with the connectivity information. Vertex attributes may include scalar or vector attribute values associated with mesh vertices. Attribute maps may include attributes associated with the mesh surface and stored as 2D images / videos. Mapping between videos (e.g., 2D images / videos) and the mesh surface may be defined by the mapping information.
[0041] A mesh can represent a static mesh containing a single frame. Therefore, mesh and frame can be used interchangeably. A dynamic mesh can contain multiple frames, including at least one predefined keyframe. A keyframe can refer to a frame that does not reference other frames during encoding / decoding. Frames other than keyframes can be referred to as non-keyframes.
[0042] A dynamic mesh can be a mesh in which at least one of its components (geometry information, connectivity information, mapping information, vertex attributes, and attribute maps) changes over time. A dynamic mesh can be described by a sequence of meshes. A dynamic mesh can be referred to as a mesh frame. Because a dynamic mesh can contain a significant amount of information that changes over time, a dynamic mesh can require a large amount of data.
[0043] Dynamic meshes can have constant connectivity information, time-varying geometry, and time-varying vertex properties. Dynamic meshes can have time-varying connectivity information. Digital content creation tools can typically generate dynamic meshes with time-varying property maps and time-varying connectivity information. Volumetric acquisition techniques can be used to generate dynamic meshes. Volumetric acquisition techniques can generate dynamic meshes with time-varying connectivity information, particularly under real-time constraints. However, because these meshes contain this information, they need to be compressed.
[0044] A mesh (111) can be generated based on a frame (100). Mesh compression (112) can be provided using various techniques. These techniques can be used for various mesh compressions, static mesh compression, dynamic mesh compression, compression of dynamic meshes with constant connectivity information, compression of dynamic meshes with time-varying connectivity information, compression of dynamic meshes with time-varying attribute maps, etc. These techniques can be used in lossy and lossless compression for various applications, such as real-time communication, storage, free-view video, augmented reality (AR), virtual reality (VR), etc. These applications can include functionalities such as random access and scalable / progressive coding.
[0045] A reconstructed mesh (113) can be generated by reconstructing a compressed mesh. A mesh bitstream (119) can be generated based on the compressed mesh. A mesh bitstream can mean information contained in a mesh expressed in bits.
[0046] A texture (121) can be generated based on a frame. The texture can be combined with a mesh to form a frame. Texture packing (122) can mean packing a texture extracted from a frame using a reconstructed mesh. A texture atlas (123) can be generated using the packed texture.
[0047] The coded texture coordinates (129) can be generated based on the packed texture. In other words, the texture coordinates can be generated based on the packed texture. The generated texture coordinates can be coded for use in generating a bitstream.
[0048] Video Compression (132) may mean encoding a padded texture (texture atlas) based on a suitable video coding standard such as HEVC or VVC.
[0049] A coded texture video bitstream (139) can be generated based on compressed video. A dynamic mesh bitstream generation module (169) can be generated by merging coded texture coordinates, a coded texture video bitstream, and a mesh bitstream.
[0050] FIG. 2 is a block diagram illustrating a dynamic mesh encoding method according to a second embodiment of the present disclosure.
[0051] A dynamic mesh sequence may include frames. A frame (200) may be one of the frames included in the dynamic frame sequence. The frame (200) may include a mesh (211), a texture (221), and the like.
[0052] A mesh (211) can be generated based on a frame (200). Mesh compression (212) can be generated based on a mesh extracted from a frame. A reconstructed mesh (213) can be generated by reconstructing the compressed mesh. The reconstructed mesh can be used for texture packing. In addition, the reconstructed mesh (Triangle Face) can be used to generate a triangle face. A mesh bitstream (219) can be generated by the compressed mesh.
[0053] A texture (221) can be combined with a mesh to form a frame. Additionally, the texture can be generated based on the frame. Texture packing (222) can mean packing a texture extracted from a frame using a reconstructed mesh.
[0054] The texture atlas (223) can be generated through a packed texture. When compressing a frame in a dynamic mesh sequence, if it is the first frame, the texture atlas (223) can be stored in the texture motion lookup buffer (243). Also, when compressing a frame in a dynamic mesh sequence, if it is the first frame, the texture atlas (223) can be compressed into a video. Meanwhile, when compressing a frame in a dynamic mesh sequence, if it is not the first frame, the texture atlas (223) can be transmitted to a texture coordinate inter-prediction module.
[0055] The coded texture coordinates (229) can be generated based on the packed texture. The generated texture coordinates can be coded for use in generating a bitstream. The texture coordinates can be generated based on the packed texture. In this case, only when the compressed frame is the first frame, the coded texture coordinates of the first frame can be generated based on texture packing.
[0056] Video compression (232) may mean encoding a padded texture (texture atlas) based on a suitable video coding standard such as HEVC, VVC, etc. Video compression may only be performed when the compression process is for the first frame.
[0057] A coded single texture bitstream (239) may mean a bitstream for the texture of the first frame using compressed video.
[0058] The texture coordinate inter prediction module (240) may be configured to predict the motion of a texture in order to enable reuse of the texture within a subsequent frame of a dynamic mesh sequence. The subsequent frame may refer to a frame other than the first frame or the current frame. The subsequent frame may also be referred to as the second frame. The current frame may be referred to as the first frame.
[0059] The texture coordinate inter prediction module (240) can store texture coordinates extracted from a frame in order to reuse the texture. In addition, the texture coordinate inter prediction module can store the texture (image) extracted from the frame. The texture coordinate inter prediction module (240) can match texture coordinates and textures. The present embodiment can minimize storage requirements while maintaining visual quality through the texture coordinate inter prediction module.
[0060] The texture coordinate inter prediction module (240) may perform at least one of a procedure for generating a triangle (241), a procedure for generating a bounding box (242), a procedure for generating a texture motion search buffer (243), a procedure for estimating triangle motion (244), a procedure for generating inter-frame texture coordinates (245), and a procedure for generating an inter-frame triangle vector (246). In addition, the texture coordinate inter prediction module (240) may include a bounding box determination module, an inter-frame triangle face motion estimation module, and a texture coordinate determination module.
[0061] A triangle face (241) can be generated using a reconstructed mesh. The triangle face can include coordinate information and vertex information. The bounding box determination module can determine a bounding box (242) for each triangle face based on the coordinates. The bounding box can be referred to as a rectangular boundary block. The bounding box can be generated based on triangle face information. Additionally, the bounding box can include the triangle face. The bounding box can be determined by the determination elements of the bounding box. The determination elements of the bounding box can mean mesh information (vertex, coordinate), the width of the triangle face, the height of the triangle face, etc. The bounding box can be the smallest rectangle that includes the triangle face. The bounding box determination module can determine and output the bounding box using the upper left coordinate of the mesh and the size of the triangle face. The size of the triangle face can include the width and height.
[0062] The bounding box determination module can determine a specific region using a bounding box. This module can reduce computational complexity and improve efficiency by focusing motion searches within a specific region. Furthermore, the bounding box determination module can enable efficient motion estimation by comparing only a specific region with the second frame.
[0063] The texture motion lookup buffer (243) may mean storing a texture atlas in a texture coordinate inter prediction module when compressing a dynamic mesh sequence, if it is the first frame. The texture motion lookup buffer may be used for triangle motion estimation (244) using the stored texture atlas.
[0064] Inter-frame triangle motion estimation (244) can estimate the motion of a triangle based on a texture atlas (233), a rectangular bounding box (242), and a texture motion lookup buffer. The triangle motion estimation can output inter-frame texture coordinates (245). Additionally, the triangle motion estimation can output an inter-frame triangle vector (246). The triangle motion estimation (244) can output at least one of the inter-frame texture coordinates (245) or the inter-frame triangle vector (246).
[0065] The inter-frame texture coordinates (245) may refer to the texture coordinates of the second frame. The inter-frame triangle vector (246) may refer to a motion vector between the triangle of the second frame and the triangle of the first frame. The coded inter-frame texture coordinates (248) may be generated by encoding the inter-frame texture coordinates. The coded inter-frame triangle vector (247) may be generated by encoding the inter-frame triangle vector.
[0066] Meanwhile, the texture coordinate inter prediction module (240) can transmit information used when generating a mesh. The mesh can be generated based on the mesh information sent from the texture coordinate inter prediction module.
[0067] Dynamic mesh bitstream generation (269) can be generated based on the texture coordinates of the first frame, the mesh bitstream, the single texture bitstream, and the second frame texture coordinates. Additionally, dynamic mesh bitstream generation (269) can be generated based on the texture coordinates (229), the mesh bitstream (219), the single texture bitstream (239), and the triangle motion information (249) of the first frame.
[0068] FIG. 3 is a block diagram illustrating a dynamic mesh encoding method according to a third embodiment of the present disclosure.
[0069] A dynamic mesh sequence may include frames. A frame (300) may be one of the frames included in the dynamic frame sequence. The frame (300) may include a mesh (311), a texture (321), and the like.
[0070] The mesh (311) can be generated based on the frame (300). Additionally, the mesh (311) can be generated based on information transmitted to the texture coordinate inter prediction module.
[0071] Mesh compression (312) may be provided using various techniques. These techniques may be used for various mesh compressions, static mesh compression, dynamic mesh compression, compression of dynamic meshes with constant connectivity information, compression of dynamic meshes with time-varying connectivity information, compression of dynamic meshes with time-varying attribute maps, etc. These techniques may be used in lossy and lossless compression for various applications, such as real-time communication, storage, free-view video, augmented reality (AR), virtual reality (VR), etc. These applications may include functionalities such as random access and scalable / progressive coding. A mesh bitstream (319) may be generated by a compressed mesh.
[0072] A reconstructed mesh (313) can be generated by reconstructing a compressed mesh. The reconstructed mesh can be used for texture packing.
[0073] A texture (321) can be combined with a mesh to form a frame. Additionally, the texture can be extracted from the frame. Texture packing (322) can mean packing the texture extracted from the frame using the reconstructed mesh.
[0074] The texture atlas (323) can be generated using a packed texture. When compressing a frame in a dynamic mesh sequence, if it is not the first frame, the reconstructed mesh information and the texture atlas can be transmitted to a texture coordinate inter-prediction module. At this time, the reconstructed mesh information and the texture atlas can be combined.
[0075] The coded texture coordinates (329) can be generated based on the packed texture. The generated texture coordinates can be coded to generate a bitstream. The texture coordinates can be generated based on the packed texture. At this time, if the frame to be compressed is the first frame, the coded texture coordinates of the first frame can be generated based on texture packing. In addition, if the texture intra-inter coordinate determination module uses intra texture coordinates, the coded texture coordinates of the current frame can be generated based on texture packing.
[0076] Video compression (332) may refer to encoding a padded texture (texture atlas) based on a suitable video coding standard such as HEVC, VVC, etc. Video compression may be performed when compressing a dynamic mesh sequence, if it is the first frame. Additionally, if the texture intra-inter coordinate determination module uses intra texture coordinates, video compression may be performed based on the texture atlas for the current frame.
[0077] A coded single texture bitstream (339) may mean a bitstream for the texture of the first or current frame using compressed video.
[0078] The texture coordinate inter prediction module may be configured to predict and / or estimate the motion of a triangle in order to enable reuse of a texture within a second frame of a dynamic mesh sequence. The texture coordinate inter prediction module may store texture coordinates within the frame in order to reuse the texture. Additionally, the texture coordinate inter prediction module may store a texture (image) within the frame. The texture coordinate inter prediction module may match texture coordinates with textures. The texture coordinate inter prediction module may minimize storage requirements while maintaining visual quality.
[0079] The texture coordinate inter prediction module (340) may correspond to the texture coordinate inter prediction module (240). The texture coordinate inter prediction module (340) may receive at least one of the reconstructed mesh information, the texture atlas, and the texture atlas of the second frame.
[0080] The texture intra-inter coordinate determination module (350) can compare the current frame and the full frame cost function evaluations. This can be used to determine whether to use an inter-prediction coordinate or an intra-prediction coordinate coding method. Using this method, the encoder can use either intra-prediction coordinates or inter-prediction coordinates.
[0081] The texture intra-inter coordinate determination module (350) can determine to use either the inter-prediction coordinates or the intra-texture coordinates. The texture intra-inter coordinate determination module (350) can compare the full-frame cost function evaluation using the first frame. In addition, the texture intra-inter coordinate determination module can enable the encoder to use both the intra-coordinates and the inter-prediction coordinates. The inter-prediction coordinates can mean triangle motion information. Meanwhile, the texture intra-inter coordinate determination module (350) can use a cost function to determine the intra-coordinates or the inter-prediction coordinates. In addition, the texture intra-inter coordinate determination module can determine using a loss threshold.
[0082] If the texture intra-inter coordinate determination module decides to use intra texture coordinates, it can send the texture atlas of the current frame to the texture coordinate inter prediction module. The texture atlas of the current frame sent to the texture coordinate inter prediction module can be used when compressing the next frame.
[0083] In addition, in order to use intra texture coordinates, a procedure by a texture coordinate intra module (330) can be performed. The texture coordinate intra module (330) can compress textures or texture coordinates in an intra manner. The texture coordinate intra module (330) can be composed of a coded current texture coordinate (329), a mesh bitstream (319), and a video compression (332) module. When the texture intra-inter coordinate determination module uses intra coordinates, the motion mesh bitstream generation (369) can be generated based on the texture coordinates, mesh bitstream, and texture bitstream of the current frame. When the texture intra-inter coordinate determination module uses inter coordinates, the dynamic mesh bitstream generation (369) can be generated based on the texture coordinates, mesh bitstream, and triangle motion information (349) of the first frame. The triangle motion information (349) can include the texture coordinates and triangle vectors of the second frame.
[0084] FIG. 4 is a block diagram illustrating a motion estimation method according to a first embodiment of the present disclosure.
[0085] The inter-frame triangle motion estimation module (400) may include a triangle motion estimation module. The inter-frame triangle motion estimation module (400) may extract triangle motion information using bounding box data.
[0086] The inter-frame triangle motion estimation module (410) can be used by combining a texture atlas and a bounding box. The inter-frame triangle motion estimation module (410) can use data of the bounding box. The inter-frame triangle motion estimation module (410) can use pixels of the texture in the bounding box for the first frame as a reference point. The current frame may be referred to as the first frame. The reference frame may be referred to as the second frame. The second frame may mean the first frame or a frame preceding the current frame. For example, if the first frame is the fifth frame, the second frame may mean the fourth frame or the first frame. The reference frame may mean a frame used to predict the bounding box for the current frame.
[0087] The inter-frame triangle motion estimation module (410) can compare the bounding box of the first frame with the bounding box of the second frame. The inter-frame triangle motion estimation module (410) can perform a bounding box comparison procedure by moving a search area predetermined by the bounding box to the next coordinate to be compared using various search algorithms. In addition, the inter-frame triangle motion estimation module (410) can perform a bounding box comparison procedure by moving the entire area of the frame to the next coordinate to be compared using various search algorithms. The search algorithm may include at least one of a full search, a hierarchical search, a diamond search, and a logarithmic search.
[0088] The triangle motion estimation module can utilize the atlas of the first frame. Additionally, the triangle motion estimation module can utilize the atlas of the second frame. The first frame may include a bounding box. The bounding box may include a triangle. The module can search for an area corresponding to a bounding box including a triangle in the first frame within the first frame. The triangles can be compared within the corresponding area.
[0089] The inter-frame triangle motion estimation module (410) can extract coordinates by comparing bounding boxes through a search algorithm.
[0090] The inter-frame triangle motion estimation module (400) may include a motion candidate selection algorithm. The triangle motion may be estimated through the extracted coordinates. The triangle motion estimation may estimate the motion using a texture atlas. After extracting the coordinates through the search, the motion candidate selection algorithm may compare the specific coordinates of the first frame with the searched coordinates to find the smallest SAD or MSE among them. The motion candidate selection algorithm may compare the bounding box of the first frame and the previous bounding box within the bounding box as pixel data. A motion vector may be calculated by comparing the coordinates of the bounding box within the first frame with the coordinates of the bounding box within the second frame. The triangle motion estimation module may output the bounding box coordinates of the second frame. The motion candidate selection algorithm may accurately identify the motion representing the movement of the bounding box in the second frame.
[0091] Additionally, the motion candidate selection algorithm can compare the bounding boxes of the first frame and the second frame within the entire frame using the Sum of Absolute Difference (SAD). The motion candidate selection algorithm can compare the bounding boxes of the first frame and the second frame within the entire frame using the Mean Square Error (MSE).
[0092] The motion candidate selection algorithm can compare the texture of the first frame and the texture of the second frame within the entire frame by evaluating them as pixels of a bounding box.
[0093] Through this, inter-frame texture coordinates can be calculated. Texture coordinates for the first frame can be obtained. In addition, a triangle vector for the first frame can be obtained through this. The triangle vector can mean a vector between the triangle of the first frame and the triangle of the first frame corresponding to the triangle of the first frame. After searching for coordinates, the triangle motion estimation module can obtain the motion vector of the bounding box using a motion selection algorithm. The inter-frame triangle motion estimation module can accurately estimate the motion or deformation of the frame mesh.
[0094] Meanwhile, triangle motion information can mean texture coordinates or triangle vectors.
[0095] FIG. 5 is a block diagram illustrating a motion estimation method according to a second embodiment of the present disclosure.
[0096] The current frame (510) may include a bounding box. The bounding box may include a triangle. The triangle may include a search point (511). The search point may refer to the upper left vertex.
[0097] The second frame may include a bounding box. The bounding box of the second frame may correspond to the bounding box of the first frame. The bounding box of the second frame may include a triangle. The triangle of the second frame may include a triangle candidate (531), a transformation candidate (532), and a rotation candidate (533). The transformation candidate may refer to a transformed version of the triangle candidate. The rotation candidate (533) may refer to one of the triangle candidates rotated by 0 degrees, 90 degrees, 180 degrees, or 270 degrees. The bounding box of the second frame may include a candidate point (521). The candidate point may refer to the upper left vertex of the triangle face.
[0098] FIG. 6 is a block diagram illustrating a dynamic mesh decryption method according to one embodiment of the present disclosure.
[0099] A dynamic mesh bitstream (610) may mean an encoded dynamic mesh sequence. The dynamic mesh bitstream may include a texture bitstream, a mesh bitstream, and intra-inter coordinate syntax.
[0100] A coded texture bitstream (621) can be generated via a dynamic mesh bitstream. Video decoding (622) can be generated by decoding the texture bitstream. A texture atlas (623) can be generated via decoded video information.
[0101] Intra-Inter Coordinate Syntax (631) can be generated via a dynamic mesh bitstream. The intra-inter coordinate syntax can generate coded intra texture coordinates. The intra-inter coordinate syntax can generate triangle motion information. Additionally, the intra-inter coordinate syntax can generate triangle vectors.
[0102] Triangle Face Motion information (632) can be generated by intra-inter coordinate syntax. The texture coordinates can refer to texture coordinates for a second frame. The second frame can refer to a subsequent frame. The texture coordinates of the second frame can generate coded inter-texture coordinates.
[0103] The coded inter-texture coordinates (633) can be generated based on triangle motion information. The coded inter-texture coordinates can be used for texture mapping. When the coded inter-texture coordinates are used for texture mapping, the latest texture can be used as a reference. The coded intra-texture coordinates (635) can be generated based on the intra-inter coordinate syntax. The coded intra-texture coordinates can be used for texture mapping. When the coded intra-texture coordinates are used for mapping, a texture buffer refresh can be requested.
[0104] A mesh bitstream (641) may be generated based on a dynamic mesh bitstream. The mesh bitstream may be used in a mesh decoding procedure. Mesh decoding (642) may mean generating a mesh using the mesh bitstream. A reconstructed mesh (643) may be generated based on the mesh bitstream.
[0105] Triangle Texture Mapping (650) can be generated based on a texture atlas, coded inter-texture coordinates, and coded intra-texture coordinates. Re-Coloring (651) can mean changing the color of a texture in the texture mapping again.
[0106] The terms and words used in the present specification and claims described above should not be interpreted as limited to their usual or dictionary meanings, but should be interpreted as meanings and concepts that conform to the technical idea of the present invention based on the principle that the inventor can appropriately define the concept of the term in order to explain his own invention in the best way.
[0107] Accordingly, the configurations depicted in the drawings and embodiments described in this specification are only one of the most preferred embodiments of the present invention, and do not represent all of the technical ideas of the present invention. Therefore, it should be understood that there may be various equivalents and modified examples that can replace them at the time of filing this application.
Claims
1. In the encoding method, A step of extracting mesh information and texture information from the first input frame; A step of determining a bounding box by inputting the above mesh information and texture information; A step of extracting triangular motion information of an inter-frame using the above bounding box; and A step of generating a bitstream using the above triangle motion information, the above mesh information, and the texture information of the second frame; An encoding method including:
2. In paragraph 1, In the step of extracting triangular motion information of an inter-frame using the above bounding box, A step of extracting coordinates by comparing the bounding box of the first frame with the bounding box of the second frame; A step of estimating the triangular motion of the inter-frame by comparing the above coordinates; and A step of extracting the triangle motion information through triangle motion estimation of the inter-frame; An encoding method including:
3. In paragraph 1, The step of determining a bounding box by inputting the above mesh information and texture information is as follows: A step of calculating a minimum bounding box using the decision elements of the above bounding box; An encoding method that further includes .
4. In paragraph 2, The step of extracting coordinates by comparing the bounding box of the first frame and the bounding box of the second frame is as follows: A step of searching for a bounding box of the second frame to be compared through a search algorithm; An encoding method that further includes .
5. In paragraph 2, A step of determining whether to use the above triangle motion information; An encoding method that further includes .
6. In paragraph 2, The step of estimating the triangular motion of the inter-frame by comparing the above coordinates is: A step of estimating the triangle motion through a triangle motion candidate selection algorithm; An encoding method including:
7. In paragraph 5, In the step of determining whether to use the above triangle motion information, Using a cost function to determine whether to use the above triangle motion information, Encoding method.
8. As an encoding device, memory for storing at least one command; and A processor for executing at least one of the above instructions; Including, The processor executes at least one instruction: Extract mesh information and texture information from the first input frame, Determine the bounding box by inputting the above mesh information and texture information, Using the above bounding box, the triangular motion information of the inter-frame is extracted, Generating a bitstream using the above triangle motion information, the mesh information, and the texture information of the second frame, Encoding device.
9. In paragraph 8, The processor executes at least one instruction: In extracting the triangular motion information of the inter-frame using the above bounding box, Extract coordinates by comparing the bounding box of the first frame with the bounding box of the second frame, Estimate the triangular motion of the inter-frame by comparing the above coordinates, Extracting the triangle motion information through the triangle motion estimation of the above inter-frame, Encoding device.
10. In paragraph 9, The processor executes at least one instruction: In determining the bounding box by inputting the above mesh information and texture information, Compute the minimum bounding box using the decision elements of the above bounding box, Encoding device.
11. In paragraph 9, The processor executes at least one instruction: In extracting coordinates by comparing the bounding box of the first frame and the bounding box of the second frame, Searching for the bounding box of the second frame to be compared through a search algorithm, Encoding device.
12. In paragraph 9, The processor executes at least one instruction: Determining whether to use the above triangle motion information, Encoding device.
13. In paragraph 9, The processor executes at least one instruction: In estimating the triangular motion of the inter-frame by comparing the above coordinates, Estimating the above triangle motion through a triangle motion candidate selection algorithm, Encoding device.
14. In paragraph 12, The processor executes at least one instruction: In determining whether to use the above triangular motion information, Using a cost function to determine whether to use the above triangle motion information, Encoding device.
15. In the decoding method, A step of receiving a dynamic mesh bitstream; A step of extracting mesh information from the above dynamic mesh bitstream; A step of extracting texture information from the above dynamic mesh bitstream; A step of extracting triangle motion information from the above dynamic mesh bitstream; A step of reconstructing a mesh using the above mesh information; A step of performing texture mapping using the above texture information and the above triangle motion information; and A step of generating a first frame using the reconstructed mesh and the result of the texture mapping; A decoding method including:
16. In paragraph 15, A step of extracting intra-inter coordinate syntax from the above dynamic mesh bitstream; Including more, The step of extracting the above triangular motion information is: A step of extracting the triangular motion information using the intra-inter coordinate syntax; Including, The steps for performing the above texture mapping are: A step of performing triangle texture mapping using the above triangle motion information; A decoding method including:
17. In paragraph 16, The step of extracting the above triangular motion information is: A step of generating texture coordinates of a second frame using the intra-inter coordinate syntax; and A step of generating inter texture coordinates using the texture coordinates of the second frame; and A step of generating the triangular motion information including the inter-texture coordinates; The step of performing the above texture mapping comprises: using the inter texture coordinates; A decoding method characterized by .
18. In paragraph 16, The step of extracting the above triangular motion information is: A step of generating intra texture coordinates using the intra-inter coordinate syntax; and A step of generating the triangle motion information including the intra texture coordinates; The step of performing the above texture mapping comprises: using the intra texture coordinates; A decoding method characterized by .
19. In paragraph 15, A step of generating a texture atlas using the above texture information; Including more, The step of performing the above texture mapping comprises: using the above texture atlas; A decoding method characterized by .
20. As a decoding device, memory for storing at least one command; and A processor for executing at least one of the above instructions; Including, The processor executes at least one instruction: Receive a dynamic mesh bitstream, Extracting mesh information from the above dynamic mesh bitstream, Extracting texture information from the above dynamic mesh bitstream, Extracting triangle motion information from the above dynamic mesh bitstream, Reconstruct the mesh using the above mesh information, Texture mapping is performed using the above texture information and the above triangle motion information, Generating the first frame by combining the reconstructed mesh and texture mapping results, Decoding device.
Citation Information
Patent Citations
System and method to stabilize display of an object tracking box
KR1020160102248A
Probe for detecting near field and near field detecting system including the same
KR1020220034494A
Functional pure cotton non-woven fabric and manufacturing method thereof
KR1020240057159A
Display device
KR102724448B1