Method for encoding and decoding improved video-based dynamic mesh data

By stacking frames into a merged displacement frame and using AI compression, the method enhances video-based dynamic mesh coding efficiency, reducing compression time and maintaining image quality for real-time applications.

WO2025147042A1PCT designated stage expired Publication Date: 2025-07-10INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/021486
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-30
Filing Date
2024-12-30
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing video-based dynamic mesh coding technologies face inefficiencies in compression time and image quality, particularly as the number of encoded displacement YUV frames increases, leading to prolonged processing times and potential distortion.

Method used

A method for encoding and decoding dynamic mesh data by stacking frames into a merged displacement frame, applying an All Intra (AI) compression method, and optimizing pixel indexing and resolution to enhance compression efficiency while maintaining lossless processing.

Benefits of technology

The proposed method significantly reduces compression time by 47% and maintains excellent image quality with minimal decoding delay, optimizing resource use for real-time applications like augmented and virtual reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024021486_10072025_PF_FP_ABST
    Figure KR2024021486_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to an encoding and decoding field for video-based dynamic mesh coding (video dynamic mesh coding, V-DMC) and a method for encoding displacement data of a dynamic mesh by a computer device according to an embodiment of the present invention may include a method in which the computer device acquires multiple frames in which displacement data is stored, the computer device stacks the multiple frames to generate a merged displacement frame, and the computer device performs dynamic mesh displacement encoding on the basis of the merged displacement frame.
Need to check novelty before this filing date? Find Prior Art

Description

Method for encoding and decoding improved video-based dynamic mesh data

[0001] The present invention relates to the field of encoding and decoding of video dynamic mesh coding (V-DMC), and relates to a method for such encoding and decoding, a method for recording such data, and components, devices, and systems for realizing such a method.

[0002] The present invention may be a technical field including at least one of digital video compression technology standards known by the names of standards such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG, or a technical field for improving the inherent efficiency of the standards, or a technical field for improving or replacing the standards.

[0003] Digital video encoding and decoding are widely used in various digital video applications. For example, digital television broadcasting, video transmission over communication networks, video calls / video conversations / video chats, recording and providing video content using optical media including video compact discs (VCDs), digital versatile discs (DVDs), and Blu-Rays, all processes for producing, editing, collecting, and distributing video content, and devices such as video recording devices and camcorders for various purposes, including personal, commercial, industrial, and security purposes, all depend on video encoding and decoding technologies.

[0004] Accordingly, implementations that may be referred to as digital video encoders and decoders may form part of a wide range of devices related to the generation, recording, and provision of digital video, including digital televisions, digital broadcasting systems, wireless broadcasting systems, computers in the form of notebooks / desktops / tablets, e-book readers, digital cameras, digital recording devices, digital multimedia playback devices, video game devices / terminals / consoles, mobile phones (including smartphones) with multimedia playback capabilities, equipment for video conferencing, and other devices.

[0005] The above digital video encoders and decoders can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art. The digital video compression standard may include at least one of compression standards known by a standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.

[0006] Video encoders and decoders can be implemented to more efficiently encode or decode digital video information while complying with the above standards, or by improving or modifying the above standards. Attempts to modify the above standards can also lead to the development of new standards. A well-known example is the so-called enhanced compression model (ECM), an attempt to improve and replace the existing H.266 / VVC standard, currently being developed by the Joint Video Experts Team (JVET), a joint international standardization group of ISO, IEC, and ITU-T.

[0007] Image / video-based compression technologies, such as ISO / IEC 23090-5 Visual Volumetric Video-Based Coding (V3C), were developed to efficiently compress 3D volumetric data, such as point clouds (i.e., V3C / V-PCC) or 3DoF+ content (V3C / MIV). The V3C standard enables the compression of 3D data, such as static and dynamic point clouds, by combining existing image / video coding technologies with metadata through a well-defined syntax structure and processing steps. Image / video coding technologies are used to compress 3D projection data on a 2D plane, such as shape and properties, and the metadata contains information on how to extract and reconstruct a 3D representation from that 2D projection. However, the 3D representation may also include 3D textured static and / or dynamic meshes, which are not currently supported by V3C.

[0008] The present invention proposes a novel compression solution for 3D dynamic meshes, thereby providing a method for achieving efficient compression performance and saving compression time by improving the format of input files.

[0009] A method for encoding displacement data of a dynamic mesh by a computer device according to one embodiment of the present invention for solving the above-described problem of the present invention may include a method in which the computer device obtains a plurality of frames in which displacement data is stored, the computer device stacks the plurality of frames to generate a merged displacement frame, and the computer device performs dynamic mesh displacement encoding based on the merged displacement frame.

[0010] Multiple frames in which the above displacement data is stored can be acquired in units of GOP (Group of Pictures).

[0011] The plurality of frames in which the above displacement data is stored may include displacement data stored in YUV color format.

[0012] In performing dynamic mesh displacement encoding based on the above merged displacement frame, compression using an All Intra (AI) method can be performed on the above merged displacement frame.

[0013] In generating a merged displacement frame by stacking the plurality of frames, the method may include checking whether the resolution of each of the plurality of frames is a multiple of the encoding unit unit, and if the resolution of the frame is not a multiple of the encoding unit unit, modifying the resolution of the frame to correspond to the multiple.

[0014] The above coding unit may be a unit based on a coding tree unit (CTU).

[0015] In creating a merged displacement frame by stacking the plurality of frames, the plurality of frames can be stacked vertically.

[0016] In generating a merged displacement frame by stacking the above plurality of frames, a pixel index that increases progressively according to the stacking order can be assigned to each pixel constituting the merged displacement frame.

[0017] The method may further include the computer device comparing the plurality of frames and the merged displacement frame to verify whether lossless processing has been performed.

[0018] In verifying whether lossless processing is performed by comparing the plurality of frames and the merged displacement frame with each other, the method may include separating the merged displacement frame, matching the resolution of each of the separated displacement frames with that of the plurality of frames, and calculating a signal-to-noise ratio (PSNR) by comparing the separated displacement frame with each of the plurality of frames.

[0019] A method for decoding displacement data of a dynamic mesh by a computer device according to another embodiment of the present invention for solving the above-described problem of the present invention may include a method in which the computer device obtains a merged displacement frame through dynamic mesh displacement decoding, the computer device separates the merged displacement frame to obtain a plurality of separated frames, and the computer device restores displacement data of the dynamic mesh from the plurality of separated frames.

[0020] The above merged displacement frame may include multiple frames merged in units of GOP (Group of Pictures).

[0021] In restoring displacement data of a dynamic mesh from the above plurality of separated frames, displacement data saved in YUV color format can be restored.

[0022] In obtaining a plurality of separated frames by separating the above merged displacement frame, the method may include calculating the number of separated frames included in the merged frame and dividing the merged displacement frame into a plurality of separated frames of the same size according to the calculated number.

[0023] The above-mentioned merged displacement frame includes a plurality of vertically stacked frames, and the vertical stacking can be released to obtain the plurality of separated frames by dividing the above-mentioned merged displacement frame into a plurality of separated frames of the same size.

[0024] The above plurality of separated frames may have a resolution determined as a multiple of the encoding unit unit.

[0025] In obtaining a plurality of separated frames by separating the above merged displacement frame, it may include restoring the original resolution by identifying and removing a background area from the separated frames.

[0026] The above coding unit may be a unit based on a coding tree unit (CTU).

[0027] In restoring displacement data of a dynamic mesh from the plurality of separation frames, the method may include determining a temporal order for each displacement frame included in the plurality of separation frames, and restoring the displacement data based on the plurality of separation frames according to the determined temporal order.

[0028] According to another embodiment of the present invention for solving the above-described problem of the present invention, a decoder device for decoding displacement data of a dynamic mesh may include a processor and a memory for storing instructions for controlling the operation of the processor, and the processor may include a device configured to obtain a merged displacement frame through dynamic mesh displacement decoding by executing the instructions stored in the memory, obtain a plurality of separated frames by separating the merged displacement frame, and restore displacement data of the dynamic mesh from the plurality of separated frames.

[0029] According to the present invention, the problem of compression time gradually increasing as the number of encoded displacement YUV frames increases can be solved. By combining multi-frame displacement YUV into a single frame by combining them into a group of pictures (GOP) through the displacement data reconstruction method proposed in the present invention, reading and writing time can be eliminated, so that the time required for the compression process can be significantly shortened. According to the experimental results, the compression time of the original data consisting of 32 frames was measured to be 5.86 seconds on average, while when the vertical stacking method was applied, the compression time was measured to be 2.75 seconds on average, confirming an encoding speed improvement effect of approximately 47%.

[0030] The present invention, in particular, maximizes compression efficiency while minimizing distortion in image / video processing by reconstructing displacement data stored in YUV space using the CTU method and efficiently adding and reducing background space. Furthermore, by applying intra-frame prediction-only (all intra; AI) compression to vertically stacked frames using a lossless method, the time delay occurring during the decoding process can be limited to 9% while maintaining excellent image quality.

[0031] The present invention optimizes the resources required for storing and transmitting displacement data in dynamic mesh encoding, making it particularly useful in applications requiring real-time processing of large amounts of data, such as real-time communications or augmented / virtual reality. Furthermore, the method of the present invention offers enhanced performance while maintaining compatibility with existing image / video codec standards, making it highly applicable to industrial applications.

[0032] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention;

[0033] FIG. 2 is a conceptual diagram of the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention.

[0034] Figure 3 is a functional unit conceptual diagram of a video decoder according to one embodiment of the present invention;

[0035] Figure 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention;

[0036] Figure 5 is a conceptual diagram of a frame type according to one embodiment of the present invention;

[0037] Figure 6 is a conceptual diagram showing the structure of a video encoder according to another embodiment of the present invention.

[0038] FIG. 7 is a block diagram illustrating a dynamic mesh encoding method according to one embodiment of the present invention;

[0039] FIG. 8 is a block diagram showing a dynamic mesh encoding method according to another embodiment of the present invention;

[0040] FIG. 9 is a block diagram showing a dynamic mesh encoding method according to another embodiment of the present invention;

[0041] FIG. 10 is a block diagram showing a dynamic mesh decryption method according to one embodiment of the present invention;

[0042] Figure 11 is a high-level conceptual diagram of a dynamic mesh encoding procedure according to one embodiment of the present invention;

[0043] Figure 12 is a high-level conceptual diagram of a dynamic mesh decryption procedure according to one embodiment of the present invention;

[0044] Figure 13 is an exemplary diagram showing a method of preprocessing according to one embodiment of the present invention;

[0045] FIG. 14 is a block diagram detailing an on-screen encoding procedure according to one embodiment of the present invention;

[0046] Figure 15 is a conceptual diagram of a mesh segmentation method according to one embodiment of the present invention.

[0047] FIG. 16 is a block diagram detailing an inter-screen encoding procedure according to one embodiment of the present invention;

[0048] FIG. 17 is a block diagram detailing an in-screen decryption procedure according to one embodiment of the present invention;

[0049] FIG. 18 is a block diagram detailing an inter-screen decryption procedure according to one embodiment of the present invention;

[0050] Figure 19 is a block diagram of a remeshing procedure according to one embodiment of the present invention.

[0051] FIG. 20 is a conceptual diagram summarizing the entire encoding pipeline in a V-DMC encoding procedure according to one embodiment of the present invention.

[0052] Figure 21 is a conceptual diagram showing a method of expressing displacement information using YUV frames according to one embodiment of the present invention.

[0053] FIG. 22 is a conceptual diagram for a progressive increase in index information for multiple frames according to one embodiment of the present invention;

[0054] Figure 23 is a conceptual diagram for processing displacement information based on CTU-based processing according to one embodiment of the present invention.

[0055] Figure 24 is a conceptual diagram of a layering process of a frame according to one embodiment of the present invention, and

[0056] Figure 25 is a conceptual diagram of a lamination procedure for lossless processing according to one embodiment of the present invention.

[0057] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.

[0058] Although terms such as “first,” “second,” etc. may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component could be referred to as the second component, and similarly, the second component could also be referred to as the first component. The term “and / or” includes any combination of multiple related listed items or any of multiple related listed items, and is non-exclusive unless otherwise indicated. The listing of items in this specification is merely an exemplary description to easily explain the spirit and possible implementation methods of the invention herein, and therefore is not intended to limit the scope of embodiments of the present invention.

[0059] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."

[0060] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."

[0061] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".

[0062] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”

[0063] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0064] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0065] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0066] In describing the invention herein, embodiments may be described or illustrated in terms of unit blocks that perform the described function or functions. The blocks may be expressed herein as one or more devices, units, modules, parts, etc. The blocks may be implemented in hardware by one or more logic gates, integrated circuits, processors, controllers, memories, electronic components, or information processing hardware implementation methods, but not limited thereto. Alternatively, the blocks may be implemented in software by application software, operating system software, firmware, or information processing software implementation methods, but not limited thereto. A single block may be implemented by being separated into multiple blocks that perform the same function, or conversely, a single block may be implemented to perform the functions of multiple blocks simultaneously. The blocks may also be implemented by being physically separated or combined according to any criteria. The blocks may be implemented to operate in an environment where their physical locations are not specified and are separated from each other by a communication network, the Internet, a cloud service, or a communication method, but not limited thereto. All of the above implementation methods are within the scope of various embodiments that can be taken by a person skilled in the field of information and communication technology to implement the same technical idea, and therefore, any detailed implementation method should be interpreted as being included within the scope of the technical idea of ​​the invention in this specification.

[0067] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. To facilitate a comprehensive understanding of the present invention, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted. Furthermore, the multiple embodiments are not mutually exclusive, and it is assumed that some embodiments may be combined with one or more other embodiments to form new embodiments.

[0068]

[0069] digital video codec

[0070] Figure 1 is a conceptual diagram of a video communication system according to one embodiment of the present invention. The video communication system (100) may be configured to include at least two terminals (110, 120) connected to each other via a network (105).

[0071] In one embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a one-way video communication network. Among the terminals, a first terminal (110) may encode video data in order to transmit (111) the video data via a network (105). Among the terminals, a second terminal (120) may be configured to receive (121) the encoded video data via a network and decode and display the same.

[0072] In another embodiment of the present invention, the above-described FIG. 1 may refer to a block diagram for configuring a two-way video communication network. For the two-way video communication, each terminal (110, 120) may be configured to encode video data acquired by itself for video transmission (112, 122) to each other terminal via the network. Each terminal may also be configured to receive (113, 123) video data transmitted by another terminal via the network, decode the same, and display the decoded video data.

[0073] The terminals (110, 120) shown in Fig. 1 may be exemplified as devices such as server computers, personal computers, portable computers, and smartphones, depending on the embodiment, but are not limited thereto. The present invention is applicable to all environments for establishing a one-way or two-way video communication network, and it should be understood that the network (105) may be established by any means for transporting encoded video data between the terminals (110, 120).

[0074] In one embodiment of the present invention, the network (105) may refer to a wired or wireless communication network. Depending on the embodiment, the network may be configured to communicate information using any communication standard, which may include packet-based communication. The packet communication may be understood to include packets, for example, known as TCP or UDP.

[0075] However, in another embodiment of the present invention, the network (105) may be understood to include a process of information transmission using a recording medium. In this case, the configuration of the network is not limited to a communication medium, and should be understood to include a process of temporarily storing and physically transporting information on a hard disk, a solid state disk (SSD), a flash memory, a compact disc (CD), a digital versatile disc (DVD), a Blu-ray disc, and other mechanical, electronic, or optical recording media.

[0076] Any other means of information communication or transport, regardless of the method employed, can be considered within the scope of embodiments of the present invention as long as it has a structure that supports the transmission and decoding of video data in an encoded state. Therefore, in addition to the examples listed above, any means of information communication or transport, whether known in the past or newly available, can fall within the scope of application of the present invention.

[0077] Figure 2 is a conceptual diagram illustrating the arrangement of an encoder and decoder in a real-time video streaming environment according to one embodiment of the present invention. The streaming system (200) illustrated in Figure 2 can be applied to video data communication networks, including, for example, digital broadcasting, video telephony, and video conferencing. However, it should be noted that technical structures identical or similar to the streaming system can be equally applied even when information is transmitted via a recording medium, as described above.

[0078] According to one embodiment of the present invention, the streaming system may include a video source (210) that generates a video stream. The video source may include a digital video acquisition means (212), which may be configured as a digital camera or other device, for acquiring uncompressed raw video. The raw video stream (215) may have a large capacity and may therefore be compressed by a video encoder (217) coupled or connected to the video source.

[0079] The above encoder (217) may be configured as a means including hardware, software, or a combination of the two configured to implement an image encoding method and / or an implementation method thereof according to one embodiment of the present invention.

[0080] Through the encoder (217), an encoded bitstream (219) having a reduced capacity compared to the original video stream can be output. The bitstream (219) can be provided in real time via a relay device, which may be referred to as a streaming server (220), and / or can be stored in a recording medium (225) of the streaming server (220) for subsequent use.

[0081] The streaming system (200) may include at least one streaming client (230, 240) that connects to the streaming server (220) to receive the encoded bit string (229) in real time or obtain it later. The streaming client may include a video decoder (232) that obtains the encoded bit string (229) (which may also be regarded as a copy of the bit string (219) received by the streaming server), decodes the bit string (229), and outputs the resulting video data as video data in a form that can be displayed by a display (235) or other visual, auditory, or other sensory display means.

[0082] As described above, the functions for encoding and decoding video data are collectively called a coder-and-decoder, or video codec.

[0083] FIG. 3 is a conceptual diagram of a functional unit of a video decoder according to an embodiment of the present invention. As shown in FIG. 3, a receiving unit (310) can receive at least one encoded video data to be decoded by a decoder (305). In an embodiment of the present invention, the encoded video data may be independent for each reception, and the decoding procedure of each independent video data may be independent from the decoding procedure of other video data. The encoded video data may be received by the receiving unit (310) through a hardware or software connection (315) to a device storing the same, and as described above, the storing device may be a type of streaming server located at the other end of a communication network, or may mean a physical recording medium, but is not limited thereto.

[0084] The above-described receiving unit (310) can receive the encoded video data together with other data accompanying it, such as encoded audio data or other auxiliary data, and each of the data can be separated from the video data and provided to an appropriate processing function unit (312) other than the video decoder.

[0085] When the video data is provided through a communication network, a buffer memory (320) may be coupled between the receiving unit (310) and the decoder (305) to minimize delay and disconnection according to the network environment. The buffer memory (320) may refer to a computer-readable recording medium that temporarily stores the received video data and stably supplies it to a parser (330) corresponding to the input terminal of the decoder (305). However, if the bandwidth of the communication network is sufficient, if the video data is read from a recording medium in a local location that is not physically separated, or if the possibility of communication delay is not predicted in other environments, the buffer memory may be unnecessary.

[0086] The video decoder (305) may include the parser (330) as its input terminal to interpret the encoded video data. The parser may perform a function of separating (parsing) a plurality of pieces of information stored in the form of a bit string in the encoded video data according to a predetermined rule, and, if necessary, performing an entropy decoding (335) of entropy-coded video data, thereby performing a function of reconstructing symbols (338), which are paragraphs of video encoding information. The symbols (338) may include all information for controlling the operation of the decoder (305), and / or may further include information for controlling a device that may be attached to and operate the decoder (305), such as a display device. Control information for controlling the above display device may include information in a format called supplementary enhancement information (SEI) or video usability information (VUI).

[0087] As described above, the parser (330) may be configured to perform entropy decoding (335) of the encoded video data. The entropy encoding method of the encoded video data may vary depending on the encoding standard, and decoding may be performed accordingly. Representative examples of the entropy encoding standard may include variable length coding, Huffman coding, and arithmetic coding, and each of the encoding methods may be a context-adaptive or context-sensitive method depending on the standard, and may also be based on principles widely known to those skilled in the art.

[0088] The parser (330) may be configured to extract at least one screen from the encoded video data. The definition of the screen may vary depending on the encoding standard, and depending on the standard, one or more of the examples listed below may correspond simultaneously and overlappingly. The screen may be defined as, for example, an image, a picture, a frame, etc., or may be divided into encoding / decoding units such as tiles, slices, macroblocks, blocks, subblocks, transform units (TUs), and prediction units (PUs), or may be grouped into encoding / decoding units such as a group of pictures (GOP), a network abstraction layer, a layered video, and a scalable video.

[0089] The parser (330) may be configured to extract encoding information, such as transform coefficients, quantization parameters (QPs), and / or motion vectors, from the encoded video data. The parser (330) may be configured to perform entropy decoding (335) and parsing operations on the video data received from the buffer memory, and to selectively decode symbols (338) representing the encoding information. In addition, the parser (330) may be configured to selectively supply a specific symbol (338) to a specific decoding function unit within the decoder (305), such as an inverse quantization and inverse transform unit (340), an intra prediction unit (350), an inter prediction unit (355), or a loop filter unit (360). Control of such information supply can be determined by the information sequence included in the encoded video, and may vary depending on the encoding standard, and is not limited within the scope of the embodiments of the present invention, and is not described in detail in this conceptual diagram.

[0090] The decoder (305) may be comprised of a number of conceptual functional units that receive and process the encoded information from the parser (330). It is apparent that these conceptual functional units may be combined or further subdivided, depending on implementation needs. For example, they may be further separated for ease of implementation, or integrated into one for operational efficiency. In any case, each functional unit may be configured to perform close interaction with each other. However, despite the possibility of such integration or separation, the following description will be given as a combination of conceptual functional units to illustrate the decoding procedure of video data applied as an embodiment of the present invention.

[0091] The decoder may include an inverse quantization and inverse transformation unit (340). The inverse quantization and inverse transformation unit (340) may be configured to receive encoding information including a method to be used for numerical transformation (transform), a block size, quantization coefficients for recovering quantized information, and distinction information of a quantization matrix that simplifies and represents the quantized coefficients from the parser (330), and may be configured to output block values ​​(341) that can be input to an aggregator (370) as a result of processing the encoding information.

[0092] In one embodiment of the present invention, the output values ​​of the inverse quantization and inverse transformation unit (340) may include a predicted encoded block value within the screen. The predicted block value within the screen may mean a value that can be decoded using prediction information within the screen currently being decoded, for example, the current frame, but without using prediction information from a previously decoded screen, for example, a previous frame.

[0093] The prediction information within the current screen may be provided by the within-screen prediction unit (350). According to an embodiment of the present invention, the within-screen prediction unit (350) generates a block value of the same form as the block being decoded as the prediction information by using the screen information of a spatially adjacent area derived from a screen currently being decoded and of which decoding has been partially completed. The screen information may be provided (381) from a buffer for the current screen, a so-called line buffer (380). The merging unit (370), according to an embodiment, may be configured to merge the prediction information (351) generated by the within-screen prediction unit (350) with the block values ​​(341) provided by the inverse quantization and inverse transformation unit (340).

[0094] In another embodiment, the output values ​​of the inverse quantization and inverse transformation unit (340) may include block values ​​that have undergone inter-screen prediction encoding, and in some cases, block values ​​that have undergone motion compensation. In this case, the inter-screen prediction unit (355) may extract and use sample information (386) used for motion-based prediction from a reference picture buffer (385). The information (356) derived by performing motion compensation on the sample information based on the symbols (338) included in the block values ​​as the output values ​​may be configured to be merged with the block values ​​(341) provided by the inverse quantization and inverse transformation unit (340) by the merger unit (370). In this case, the block values ​​(341) may be referred to as so-called differential or residual values.

[0095] The position information within the memory used by the inter-screen prediction unit (355) to extract the sample information from the reference screen may be determined by a motion vector provided to the inter-screen prediction unit (355) and composed of a combination of symbols (338) for representing, for example, X, Y, and other specific points of the reference screen. The inter-screen prediction unit (355) may also include a function for interpolating and using the sample values ​​when a so-called subsampling-capable motion vector is provided, and may further include a function for predicting and reinforcing the value of the motion vector.

[0096] The output values ​​(371) of the above merging unit (370) may be provided to the loop filter unit (360) and processed by various loop filtering methods. The loop filter unit (360) may be configured to receive not only the block unit output (371) of the merging unit (370) but also the symbol (338) provided from the parser (330) to control its operation. The output of the loop filter unit (360) may be output to an external display means such as the display device through an output connection (390), but may be stored (361) in a line buffer (380) for use in prediction for interpreting the encoded block value within or between screens later, and may also be stored in a reference screen buffer (385) through this.

[0097] Certain screens, such as frames, after their decoding is completed can be utilized as reference screens for performing predictive decoding in a subsequent decoding process. One screen (i.e., picture / frame) can be gradually accumulated in a line buffer (380) and decoded, and when one frame is decoded, the contents of the line buffer (380) are transferred (383) to the reference screen buffer (385), and a new line buffer (380) can be allocated for decoding the new frame.

[0098] The above video decoder (305) may be configured to perform a decoding operation according to a predetermined video compression technology that may be documented by various international standards or commercial standards. The standards may include, for example, international standard recommendations such as H.264, H.265, and H.266 defined by the International Telecommunication Union Standardization Sub-Division (ITU-T). A person skilled in the art will understand that each of the above recommendations is equivalent to an international standard jointly defined by the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC). The encoded video data may comply with a specific bitstream syntax defined by the relevant standards, as defined in the video compression standard documents and standard documents, and specifically, by the profiles and levels specified within such documents. In addition, the complexity of the encoded video data may be limited to a certain level to comply with the profiles and levels. For example, a profile or level may be configured to limit a maximum screen size, a maximum decoding speed, and a maximum reference screen size. These limitations may, in some embodiments, also be further restricted via metadata signals for a hypothetical reference decoder (HRD) and the HRD buffer management contained in the encoded video data.

[0099] According to one embodiment of the present invention, the receiver (310) may receive additional redundant data together with the encoded video. The additional data may be considered as part of the encoded video data. The additional data may include information that may be used by the decoder (305) to properly decode the data or to more accurately reconstruct an image that approximates the original image. The additional data may be provided in the form of, for example, layers for temporal, spatial, or signal-to-noise ratio (SNR) enhancement, redundant slices, redundant scenes, and forward error correction codes.

[0100] Figure 4 is a functional unit conceptual diagram of a video encoder according to one embodiment of the present invention. The encoder (405) may be configured to receive original video information (402) from a video source (401) and perform encoding.

[0101] The original video information (402) may have any suitable bit depth, for example, 8 bits, 10 bits, 12 bits, etc. In addition, the original video information (402) may have any suitable color space, for example, R / G / B, Y / U / V, Y / Cb / Cr, etc. In addition, the original video information (402) may have any suitable sampling structure corresponding to the color space, for example, may have a format such as Y / Cb / Cr 4:2:0, Y / Cb / Cr 4:4:4. The original video information (402) having such a predetermined format may be provided to the encoder in the form of a digital video stream.

[0102] In a one-way video communication network, the original video information (402) can be obtained from a recording medium storing a previously prepared original video. In a two-way video communication network, the original video information (402) can be obtained from a video acquisition device, such as a camera, that generates at least one video transmission stream included in the two-way video communication.

[0103] The video data including the above original video information (402) may be configured as a plurality of screens configured to simulate motion by playing them in chronological order. The screens may also be expressed using concepts such as frames. The screens may include one or more samples depending on the type of sampling structure, color space, etc. being used. Those skilled in the art will understand that the terms "samples" and "pixels" in digital images are closely related. The following describes the operation of the encoder focusing on these samples.

[0104] According to one embodiment of the present invention, the encoder (405) may be configured to encode and compress the screens (and / or the information grouped or segmented therefrom) constituting the original video information (402) in the form of encoded video information in real time (or according to other temporal requirements required according to the implementation method).

[0105] In the encoder (405), the control unit (450) may be a functional unit configured to control an appropriate encoding speed. The control unit (450) may be configured to control other functional units and be functionally coupled to the following functional units as described below. The parameters set by the control unit (450) may include parameters related to bitrate control, such as screen skip, quantizer, and variable values ​​for applying image quality optimization techniques, and may also include values ​​such as screen size, structure of a group of pictures (GOP), and maximum search range of a motion vector. A person skilled in the art will be able to understand various other functions that the control unit (450) may have, and such other functions may be added or removed according to the design of a video encoder optimized for an individual system design.

[0106] According to an embodiment of the present invention, the encoder (405) may be configured to operate in a structure such as a "coding loop" well known to those skilled in the art. To briefly explain by way of example, the encoding loop may be configured with an internal encoder (so-called "source coder") (410) responsible for receiving a picture to be encoded and generating symbols based on at least one reference picture that has been encoded in the past, and a local decoder (420) configured to be connected to the internal encoder. The local decoder (420) may be configured to perform an operation to reproduce sample data to be generated by a decoder (490) located at an actual remote location that will receive encoded video information from the encoder (405) by receiving an output of the internal encoder (410).

[0107] The video data composed of sample data reconstructed by the internal decoder (420) may be configured to be input into the reference picture buffer of the encoder (405). As described above, the internal decoder (420) is implemented to reproduce the result output by the encoder (405) and to be decoded by a remote decoder, so the video data recorded in the reference picture buffer may also be identical in bit units to the information of the reference picture buffer of the remote decoder. That is, the prediction function unit that may be included in the encoder (405) may read the same values ​​as the sample values ​​of the previous frame that the decoder will later refer to in the decoding process from the reference picture buffer of the encoder (405).

[0108] As described above, the principle of achieving matching of the reference screen buffers between the encoder (405) and the decoder (490) by means of the internal decoder (420) on the encoder (405) side is well known to those skilled in the art, and a method of responding to an environment in which such an environment is not guaranteed (e.g., information loss due to communication failure, etc.) can also follow what is known to those skilled in the art.

[0109] An embodiment of the operation method of the internal decoder (420) has been described in detail above with reference to FIG. 3. The decoder of FIG. 3 may be regarded as the aforementioned "remote" decoder (490). The internal decoder (420) may be implemented excluding lossless encoding and decoding sections such as the parser (330) or entropy decoding (335). This is because the internal encoder (405) is implemented to simply reproduce the operation of a decoder located at a remote location, and thus may directly decode symbols without requiring a process of compressing and then decompressing symbols. Accordingly, the functional units preceding the parser and entropy decoder as shown in FIG. 3 may not be provided or may be implemented at least partially.

[0110] As described above, according to a preferred embodiment of the present invention, any decoder function (excluding a parser and an entropy decoder) present in the decoder can naturally exist as a substantially identical function in the corresponding encoder (405).

[0111] The operation of the encoding function unit that may be included in the above encoder (405) can be considered as the reverse operation of the decoder function unit. Therefore, the embodiment can be explained by performing the operation of the decoder function unit in reverse. For example, a quantization and transform function unit corresponding to the inverse quantization and inverse transform unit may be provided, and an inter-screen prediction encoding unit corresponding to the inter-screen prediction unit may be provided. In addition, some additional explanations will be added.

[0112] The internal encoder (410) may be configured to perform encoding on input screen information, for example, an input frame, by a predictive encoding method executed by a predictive encoding unit (440) that operates by referencing at least one temporally previous encoded screen (i.e., picture / frame) from a reference screen buffer (430) from at least one reference screen information, for example, video data designated as a reference frame. In this case, the encoder (405) may be configured to encode a differential between blocks of samples constituting the input screen and blocks of samples constituting the reference screen.

[0113] The internal decoder (420) can decode video data that can be designated as the reference screen from symbols generated by the internal encoder (410). As described above, since the video data is subject to the same decoding operation as that performed by a remote decoder, the video data used as the reference screen may be provided to the encoder (405) in a form that has undergone lossy compression and has been partially damaged, and this operation may be intended to ensure operational consistency with the decoder.

[0114] The prediction encoding unit (440) may be configured to perform a prediction search operation within the encoder (405). The prediction search operation may refer to an operation corresponding to the inter-screen prediction or intra-screen prediction described in the description of the decoder. For input screen information that is scheduled to be newly encoded, the prediction unit may access the reference screen buffer (430) to retrieve information such as a motion vector, a block shape, and metadata that may include the same, which are information indicating a point of a reference screen that can function as prediction reference information suitable for the new screen information, and a sample block to be actually referenced. The above prediction encoding unit (440) may operate on the basis of the so-called "sample block by pixel block" criteria in order to obtain appropriate prediction reference information. According to one embodiment of the present invention, at least one prediction reference information may be designated for the input screen, which designates at least one reference screen information stored in the reference screen buffer (430), as determined based on the search results obtained by the prediction encoding unit (440).

[0115] In one embodiment of the present invention, the control unit (450) may be configured to manage the overall encoding operation of the internal encoder (410), including setting parameters used to encode video data.

[0116] All outputs of the above-described functional units may be subjected to entropy encoding (460) in order to be finally output. The entropy encoding (460) may include various entropy encoding techniques, such as variable length coding, Huffman coding, and arithmetic coding, for the symbols generated by the various functional units as described above, and each encoding method may be a context-adaptive or context-sensitive method according to the standard, or may be based on principles widely known to those skilled in the art. Such entropy encoding (460) can typically achieve lossless compression, and thus can be configured to convert at least one symbol generated by the functional units into encoded video data.

[0117] The control unit (450) may, when controlling the operation of the encoder (405), apply to each screen (i.e., picture / frame) the type of encoding a specific screen is to be encoded during the encoding period. Depending on the type, the method by which the screen is encoded may be affected. Depending on the embodiment, the type may include what is categorized as the following "frame type."

[0118] Fig. 5 is a conceptual diagram of a frame type according to one embodiment of the present invention. The following description will be made with reference to Fig. 5.

[0119] An intra-picture encoded ("I") picture (510) may refer to a picture that can be encoded and decoded only with its own information without referring to other picture information in the video data through predictive encoding. The "I" picture may be designated by names such as a key frame, an independent / instantaneous decoder refresh (IDR) frame, and a clean random-access (CRA) frame, depending on the video encoding standard, and the "I" pictures designated by the various names as described above may have various modifications and application methods as permitted by each standard and may be partially different from each other. In addition to those listed above, various application methods for implementing the "I" picture may be by various methods that are already known to those skilled in the art or may be newly provided.

[0120] An inter-picture prediction coded (“P”) picture (520) may refer to a picture that can be encoded and decoded through intra-picture or inter-picture prediction based on at least one prediction information and / or a motion vector that designates at least one reference picture to predict sample values ​​of blocks constituting the picture. The “P” picture may be configured to reference only one reference frame, or may be configured to reference one or more reference frames, according to a video encoding standard. When referencing more than one reference frame, sample information and / or associated metadata derived from multiple reference pictures may be used to reconstruct a single block. However, in common cases, a picture designated as a “P” picture may be understood as a picture that performs reference only to a temporally preceding picture.

[0121] A bidirectional prediction ("B") picture (530) may refer to a picture that can be encoded and decoded through intra- or inter-picture prediction based on at least one piece of prediction information and / or a motion vector that designates at least two reference pictures in order to predict sample values ​​of blocks constituting the picture. In a common case, a picture designated as the "B" picture is distinct from a picture designated as the "P" picture, and may be understood as a picture that performs a reference without being limited to a temporally preceding picture.

[0122] Video data may be spatially divided into a plurality of sample blocks during the encoding and decoding process, and encoding may be performed in units of the blocks. The block units may include, but are not limited to, sizes such as 4x4, 8x8, 4x8, or 16x16 in units of horizontal / vertical pixels, as is widely known. The block may be encoded using a predictive encoding method with reference to any other (already encoded) blocks, as permitted and / or restricted by the type specified for each picture in which the block is included. For example, the blocks of the "I" picture (510) may not use a predictive encoding method, or may be encoded with reference to blocks that have already been encoded within the same partial picture. That is, only the so-called intra-picture prediction method may be used. In contrast, the "P" picture (520) may further reference a reference picture encoded in at least one previous time unit, and thus, inter-picture prediction as well as intra-picture prediction may be used for encoding. In the case of the "B" screen (530), reference can be made not only to a previously encoded picture in the encoding order but also to a subsequent reference picture in terms of time unit. However, it is widely known that there may be blocks encoded without relying on predictive encoding within the "P" screen or the "B" screen.

[0123] The above video encoder (405) may be configured to perform encoding operations according to a predetermined video compression technique that may be documented by various international standards or commercial standards. Examples of the above standards may include all those described in the above decoder.

[0124] According to one embodiment of the present invention, the transmitter (470) may buffer the encoded video data generated by the entropy encoding in order to provide / transmit the video data (ultimately to a remote decoder (490)) to a device storing the encoded video data via a hardware or software connection (495). According to an embodiment, when the transmitter (470) provides / transmits the encoded video data from the video encoder (405), it may receive and merge other data accompanying the encoded video data, for example, encoded audio data or other auxiliary data, from a separate source (480).

[0125] According to one embodiment of the present invention, the transmitter (470) may be configured to transmit additional data along with the encoded video. The additional data may be considered part of the encoded video data. The additional data may include information that can be used by a decoder to properly decode the data or to more accurately reconstruct an image that approximates the original image. Examples of the additional data may include all of the examples previously presented with respect to the receiver (310) of the decoder.

[0126] The present invention can be implemented by a digital video compression standard that is widely used and understood by those skilled in the art as described above. The digital video compression standard may include at least one of compression standards known by the standard name such as MPEG-2, MPEG-4 Video, H.263, H.264 / AVC, H.265 / HEVC, H.266 / VVC, VC-1, AV1, QuickTime, VP-9, VP-10, and Motion JPEG.

[0127] Fig. 6 is a conceptual diagram illustrating the structure of a video encoder according to another embodiment of the present invention. What is depicted in Fig. 6 may be a rough structure of a video encoder widely known as a standard code such as ITU-T H.266 and ISO / IEC 23090-3, and also known as MPEG-I Part 3 or versatile video coding (VVC).

[0128] According to FIG. 6, a video encoder (605) may be configured to receive raw video data (601) that has not been compressed or encoded as input and output an encoded bit string (602). The video data (601) may be directly supplied to a luma mapping unit (610a) when encoded within a screen, or may be supplied to a luma mapping unit (610b) via an inter-screen prediction unit (620) including motion vector extraction. In the case of intra-screen encoding, the mapped luma signal may be supplied to an output merger (606) by selecting (608) at least one of an intra-screen prediction encoding signal via an intra-screen prediction unit (625) or an inter-screen prediction encoding signal output from the luma mapping unit (610b) via the inter-screen prediction unit (620). The result of the above output merger can be applied to a chroma scaling unit (615). (The operation of the luminance signal mapping unit (610) and the operation of the chroma scaling unit (615) are collectively referred to as a luma mapping / chroma scalaing (LMCS) process.) The reduced chroma signal can be provided to a transform unit (630), and the transform unit (630) can perform an adaptive color transform, particularly on the chroma signal. The coefficients derived as a result of the transform are applied to a quantization unit (640) and quantized. As a result, lossy compression is achieved, and the result of the lossy compression can be output as a bit string (602) through a multi-hypothesis CABAC (650), which is a lossless compression method.

[0129] Meanwhile, the result of the lossy compression may actually enter the decoding process by going through the processes of inverse quantization (645), inverse transform (635), and luminance signal expansion (617) to generate an encoding loop. The result of the luminance signal expansion may be supplied to the internal merger (607) together with the result of selecting (608) at least one of the previously generated intra-screen prediction encoding signal or inter-screen prediction encoding signal. The result of the internal merger may go through inverse luma mapping (617), and then may go through processing such as a deblocking filter (660), sample adaptive offset (SAO) (670), and an adaptive loop filter (ALF) to reproduce the image quality improvement process in the decoder. The result of reproducing the operation in the decoder as described above is applied to the reference screen buffer (690) and can be reused for prediction encoding by the inter-screen prediction unit (620).

[0130] The present invention can also be utilized by or incorporated into an enhanced compression model (ECM), which is an implementation of a next-generation video codec currently being developed by the Joint Video Experts Team (JVET), an international standardization expert organization. The enhanced compression model can include an enhanced intra-picture prediction coding method, an enhanced inter-picture prediction coding method, an enhanced transform and transform coefficient coding method, an enhanced adaptive loop filtering method, a bilateral filtering method, a new sample adaptive offset (SAO) method for improving image quality, an extended entropy coding method, and an improved gradual decoding refresh (GDR) technique.

[0131]

[0132] Encoding / decoding method of dynamic mesh

[0133] FIG. 7 is a block diagram illustrating a dynamic mesh encoding method according to one embodiment of the present invention. According to one embodiment of the present invention, a dynamic mesh sequence (701) may include frames (700) corresponding to “screens” in the description of the general digital video codec described above. The frame (700) may be one of at least one individual frame included in the dynamic frame sequence (701). The frame (700) may include a mesh (711), a texture (721), and the like.

[0134] The above dynamic mesh sequence (701) may be understood as a static mesh if, according to an embodiment, it contains only one frame (700). Therefore, it can be understood that all configurations described below regarding a "dynamic mesh" are essentially equally applicable to a static mesh within a technically acceptable range, as will be apparent to those skilled in the art.

[0135] The mesh may include properties, such as color, normal, etc., associated with vertices. The properties may be associated with the surface of the mesh by using mapping information that parameterizes the mesh with 2D property maps. The mapping information may be generally described by a set of parameter coordinates, referred to as UV coordinates or texture coordinates, associated with the vertices of the mesh. The 2D property maps may be used to store high-resolution property information, such as textures, normals, displacements, etc. The mesh (711) may include components referred to as geometry information, connectivity information, mapping information, vertex properties, and property maps. The geometry information may be described by a set of 3D positions associated with the vertices of the mesh. (x,y,z) coordinates may be used to describe the 3D positions of the vertices.

[0136] The connectivity information may include a set of vertex indices describing how to connect the vertices to create a 3D surface. Mapping information may describe how to map the mesh surface to 2D regions of a plane. Mapping information (which may also be referred to as UV coordinate mapping, texture mapping) may be described by a set of UV coordinate parameters / texture coordinates (u, v) associated with the mesh vertices together with the connectivity information. The vertex attributes may include scalar or vector attribute values ​​associated with the mesh vertices. Attribute maps may include attributes that are associated with the mesh surface and stored as a 2D image, i.e., as a still image and / or a moving image. The mapping between the 2D image and the mesh surface may be defined by the mapping information.

[0137] The above mesh may represent a static mesh containing a single frame. Therefore, mesh and frame may be used interchangeably. A dynamic mesh may contain multiple frames, and may include at least one pre-set keyframe. A keyframe may refer to a frame that does not reference other frames during encoding / decoding. Frames other than keyframes may be referred to as non-keyframes.

[0138] A dynamic mesh may refer to a mesh in which at least one of its components (geometry information, connectivity information, mapping information, vertex attributes, and attribute maps) changes over time. The dynamic mesh may be described by a sequence of meshes. The dynamic mesh may be referred to as a mesh frame. Since the dynamic mesh may contain a significant amount of information that changes over time, it may require a large amount of data.

[0139] The above dynamic mesh can have constant connectivity information, time-varying geometry, and time-varying vertex properties. The dynamic mesh can have time-varying connectivity information. Digital content creation tools can generally generate dynamic meshes with time-varying property maps and time-varying connectivity information. Volumetric acquisition techniques can be used to generate dynamic meshes. Volumetric acquisition techniques can generate dynamic meshes with time-varying connectivity information, especially under real-time constraints. However, because the mesh contains this information, it is necessary to compress the mesh.

[0140] A mesh (711) can be generated based on a frame (700). Mesh compression (712) can be provided using various techniques. These techniques can be used for various mesh compressions, static mesh compression, dynamic mesh compression, compression of dynamic meshes with constant connectivity information, compression of dynamic meshes with time-varying connectivity information, compression of dynamic meshes with time-varying attribute maps, etc. These techniques can be used in lossy and lossless compression for various applications, such as real-time communication, storage, free-view video, augmented reality (AR), virtual reality (VR), etc. These applications can include functions such as random access and scalable / progressive encoding.

[0141] A reconstructed mesh (713) can be generated by reconstructing a compressed mesh. A mesh bitstream (719) can be generated based on the compressed mesh. The mesh bitstream can mean information contained in the mesh expressed in bits.

[0142] A texture (721) may be generated based on a frame. The texture may be combined with a mesh to form a frame. Texture packing (722) may mean packing a texture extracted from a frame using a reconstructed mesh. A texture atlas (723) may be generated using the packed texture.

[0143] Encoded texture coordinates (729) can be generated based on a packed texture. In other words, texture coordinates can be generated based on a packed texture. The generated texture coordinates can be encoded for use in generating a bit string.

[0144] Video compression (732) may mean encoding a padded texture (i.e., the texture atlas) based on a suitable video encoding method / standard, such as HEVC, VVC, etc., including the above-described digital codec.

[0145] An encoded texture image bitstream (739) can be generated based on a compressed image. A dynamic mesh bitstream generation module (769) can be generated by merging the encoded texture coordinates, the encoded texture image bitstream, and the mesh bitstream.

[0146] FIG. 8 is a block diagram illustrating a dynamic mesh encoding method according to another embodiment of the present invention. According to one embodiment of the present invention, a dynamic mesh sequence (801) may include frames (800) corresponding to "screens" in the description of the general digital video codec described above. The frame (800) may be one of at least one individual frame included in the dynamic frame sequence (801). The frame (800) may include a mesh (811), a texture (821), and the like.

[0147] A mesh (811) can be generated based on a frame (800). Mesh compression (812) can be generated based on a mesh extracted from a frame. A reconstructed mesh (813) can be generated by reconstructing the compressed mesh. The reconstructed mesh can be used for texture packing. Additionally, the reconstructed mesh can be used to generate a triangular face. A mesh bitstream (819) can be generated by the compressed mesh.

[0148] Textures (821) can be combined with meshes to form frames. Additionally, textures can be generated based on frames. Texture packing (822) can mean packing textures extracted from frames using reconstructed meshes.

[0149] The texture atlas (823) can be generated through a packed texture. When compressing a frame in a dynamic mesh sequence, if it is the first frame, the texture atlas (823) can be stored in a texture motion lookup buffer (843). Also, when compressing a frame in a dynamic mesh sequence, if it is the first frame, the texture atlas (823) can be compressed into an image (video). Meanwhile, when compressing a frame in a dynamic mesh sequence, if it is not the first frame, the texture atlas (823) can be transmitted to an inter-screen texture coordinate prediction module (840).

[0150] Encoded texture coordinates (829) can be generated based on a packed texture. The generated texture coordinates can be encoded for use in generating a bitstream. The texture coordinates can be generated based on the packed texture. In this case, only when the compressed frame is the first frame, the encoded texture coordinates of the first frame can be generated based on texture packing.

[0151] Image compression (832) may mean encoding a padded texture (i.e., the texture atlas) based on a suitable image encoding method / standard, such as HEVC, VVC, etc., including the above-described digital codec. The image compression may be performed only when the compression process is for the first frame.

[0152] The encoded single texture bitstream (839) may mean a bitstream for the texture of the first frame using a compressed image.

[0153] The inter-screen texture coordinate prediction module (840) may be configured to predict the motion of a texture in order to enable reuse of the texture within a subsequent frame of a dynamic mesh sequence. The subsequent frame may be a frame other than the first frame or the current frame. The subsequent frame may also be referred to as the second frame. The current frame may be referred to as the first frame.

[0154] The inter-screen texture coordinate prediction module (840) can store texture coordinates extracted from a frame in order to reuse the texture. In addition, the inter-screen texture coordinate prediction module can store textures (2D images / videos) extracted from the frame. The inter-screen texture coordinate prediction module (840) can match texture coordinates with textures. The present embodiment can minimize storage requirements while maintaining visual quality through the inter-screen texture coordinate prediction module.

[0155] The inter-screen texture coordinate prediction module (840) may perform at least one of a procedure for generating a triangle (841), a procedure for generating a bounding box (842), a procedure for texture motion search buffering (843), a procedure for triangle motion estimation (844), a procedure for generating inter-screen frame texture coordinates (845), and a procedure for generating an inter-screen frame triangle vector (846). In addition, the inter-screen texture coordinate prediction module (840) may include a bounding box determination module, an inter-screen frame triangle face motion estimation module, and a texture coordinate determination module.

[0156] A triangle (841) can be generated using a reconstructed mesh. The triangle can include coordinate information and vertex information. A bounding box determination module can determine a bounding box (842) for each triangle based on the coordinates. The bounding box can be referred to as a rectangular bounding block. The bounding box can be generated based on triangle information. Additionally, the bounding box can include the triangle. The bounding box can be determined by the determination elements of the bounding box. The determination elements of the bounding box can mean mesh information (vertex, coordinate), the width of the triangle, the height of the triangle, etc. The bounding box can be the smallest rectangle that includes the triangle. The bounding box determination module can determine and output a bounding box using the upper left coordinate of the mesh and the size of the triangle. The size of the triangle can include the width and height.

[0157] The bounding box determination module can determine a specific region using a bounding box. This module can reduce computational complexity and improve efficiency by focusing motion search within a specific region. Furthermore, the bounding box determination module can enable efficient motion estimation by comparing only a specific region with the second frame.

[0158] The texture motion lookup buffer (843) may mean storing a texture atlas in an inter-screen texture coordinate prediction module when compressing a dynamic mesh sequence, if it is the first frame. The texture motion lookup buffer may be used for triangle motion estimation (844) using the stored texture atlas.

[0159] Inter-screen frame triangle motion estimation (844) can estimate the motion of a triangle based on a texture atlas (833), a rectangular bounding box (842), and a texture motion lookup buffer. The triangle motion estimation can output inter-screen frame texture coordinates (845). Additionally, the triangle motion estimation can output an inter-screen frame triangle vector (846). The triangle motion estimation (844) can output at least one of the inter-screen frame texture coordinates (845) or the inter-screen frame triangle vector (846).

[0160] The inter-screen frame texture coordinates (845) may refer to the texture coordinates of the second frame. The inter-screen frame triangle vector (846) may refer to a motion vector between the triangle of the second frame and the triangle of the first frame. The encoded inter-screen frame texture coordinates (848) may be generated by encoding the inter-screen frame texture coordinates. The encoded inter-screen frame triangle vector (847) may be generated by encoding the inter-screen frame triangle vector.

[0161] Meanwhile, the inter-screen texture coordinate prediction module (840) can transmit information used when generating a mesh. The mesh can be generated based on the mesh information sent from the inter-screen texture coordinate prediction module.

[0162] Dynamic mesh bitstream generation (869) can be generated based on the texture coordinates, mesh bitstream, single texture bitstream, and second frame texture coordinates of the first frame. In addition, dynamic mesh bitstream generation (869) can be generated based on the texture coordinates (829), mesh bitstream (819), single texture bitstream (839), and triangle motion information (849) of the first frame.

[0163] FIG. 9 is a block diagram illustrating a dynamic mesh encoding method according to another embodiment of the present invention. According to one embodiment of the present invention, a dynamic mesh sequence (901) may include frames (900) corresponding to "screens" in the description of the general digital video codec described above. The frame (900) may be one of the frames included in the dynamic frame sequence (901). The frame (900) may include a mesh (911), a texture (921), and the like.

[0164] The mesh (911) can be generated based on the frame (900). Additionally, the mesh (911) can be generated based on information transmitted to the inter-screen texture coordinate prediction module.

[0165] Mesh compression (912) can be provided using various techniques. These techniques can be used for various mesh compression, static mesh compression, dynamic mesh compression, compression of dynamic meshes with constant connectivity information, compression of dynamic meshes with time-varying connectivity information, compression of dynamic meshes with time-varying attribute maps, etc. These techniques can be used for lossy and lossless compression for various application functions, such as real-time communication, storage, free-view image, AR, VR, etc. These application functions can include functionalities such as random access and scalable / progressive encoding. The mesh bitstream (919) can be generated by the compressed mesh.

[0166] A reconstructed mesh (913) can be generated by reconstructing a compressed mesh. The reconstructed mesh can be used for texture packing.

[0167] A texture (921) can be combined with a mesh to form a frame. Additionally, the texture can be extracted from the frame. Texture packing (922) can mean packing the texture extracted from the frame using the reconstructed mesh.

[0168] The texture atlas (923) can be generated using a packed texture. When compressing a frame in a dynamic mesh sequence, if it is not the first frame, the reconstructed mesh information and texture atlas can be transmitted to the inter-screen texture coordinate prediction module. At this time, the reconstructed mesh information and texture atlas can be combined.

[0169] Encoded texture coordinates (929) can be generated based on a packed texture. The generated texture coordinates can be encoded to generate a bit string. The texture coordinates can be generated based on the packed texture. In this case, when the frame to be compressed is the first frame, the encoded texture coordinates of the first frame can be generated based on texture packing. In addition, when the intra-screen texture coordinate determination module uses intra-screen texture coordinates, the encoded texture coordinates of the current frame can be generated based on texture packing.

[0170] Image compression (932) may refer to encoding a padded texture (i.e., the texture atlas) based on a suitable still image or video encoding method / standard, such as HEVC, VVC, etc., including the aforementioned digital video codec. Image compression may be performed when compressing a dynamic mesh sequence, if it is the first frame. Additionally, if the intra-screen texture coordinate determination module uses intra-screen texture coordinates, image compression may be performed based on the texture atlas for the current frame.

[0171] The encoded single texture bitstream (939) may mean a bitstream for the texture of the first or current frame using a compressed image.

[0172] The inter-screen texture coordinate prediction module may be configured to predict and / or estimate the motion of a triangle in order to enable reuse of textures within a second frame of a dynamic mesh sequence. The inter-screen texture coordinate prediction module may store texture coordinates within a frame in order to reuse the textures. Furthermore, the inter-screen texture coordinate prediction module may store textures (2D images / videos) within the frame. The inter-screen texture coordinate prediction module may match texture coordinates with textures. The inter-screen texture coordinate prediction module may minimize storage requirements while maintaining visual quality.

[0173] The inter-screen texture coordinate prediction module (940) of FIG. 9 may correspond to the inter-screen texture coordinate prediction module (840) described above with reference to FIG. 8. The inter-screen texture coordinate prediction module (940) may receive at least one of reconstructed mesh information, a texture atlas, and a texture atlas of a second frame.

[0174] The intra-screen-to-inter-screen texture coordinate determination module (950) can compare the current frame and the full-frame cost function evaluations. This can be used to determine whether to use an inter-screen predicted coordinate or an intra-screen coordinate encoding method. This method allows the encoder to use either intra-screen or inter-screen predicted coordinates.

[0175] The intra-screen texture coordinate determination module (950) can determine to use either the intra-screen predicted coordinates or the intra-screen texture coordinates. The intra-screen texture coordinate determination module (950) can compare the full-frame cost function evaluation using the first frame. In addition, the intra-screen texture coordinate determination module can enable the encoder to use both the intra-screen coordinates and the inter-screen predicted coordinates. The inter-screen predicted coordinates can mean triangle motion information. Meanwhile, the intra-screen texture coordinate determination module (950) can use a cost function to determine the intra-screen coordinates or the inter-screen predicted coordinates. In addition, the intra-screen texture coordinate determination module can make the determination using a loss threshold.

[0176] If the intra-screen texture coordinate determination module decides to use intra-screen texture coordinates, it can send the texture atlas of the current frame to the inter-screen texture coordinate prediction module. The texture atlas of the current frame sent to the inter-screen texture coordinate prediction module can be used when compressing the next frame.

[0177] In addition, in order to use the texture coordinates within the screen, a procedure by the texture coordinates within the screen module (930) can be performed. The texture coordinates within the screen module (930) can compress the texture or texture coordinates in an within-screen manner. The texture coordinates within the screen module (930) can be composed of an encoded current texture coordinate (929), a mesh bitstream (919), and an image compression (932) module. When the within-screen to between-screen texture coordinate determination module uses the within-screen coordinates, the motion mesh bitstream generation (969) can be generated based on the texture coordinates, the mesh bitstream, and the texture bitstream of the current frame. When the within-screen to between-screen texture coordinate determination module uses the between-screen coordinates, the dynamic mesh bitstream generation (969) can be generated based on the texture coordinates, the mesh bitstream, and the triangle motion information (949) of the first frame. The triangle motion information (949) can include the texture coordinates and the triangle vector of the second frame.

[0178] Fig. 10 is a block diagram illustrating a dynamic mesh decoding method according to one embodiment of the present invention. According to one embodiment of the present invention, a merged dynamic mesh bit stream (1010) may mean that a dynamic mesh sequence (1001) including at least one frame (1000) is encoded. The merged dynamic mesh bit stream (1010) may include an encoded texture bit stream (1021), an encoded mesh bit stream (1041), and intra-screen coordinate syntax information (syntax) (1031).

[0179] An encoded texture bitstream (1021) may be generated from the merged dynamic mesh bitstream (1010). Image decoding (1022) may refer to a process of decoding an image / video contained in the texture bitstream. A texture atlas (1023) may be generated based on the decoded image / video information.

[0180] The intra-screen coordinate syntax information (1031) can be generated from the merged dynamic mesh bit string. The intra-screen coordinate syntax information (1031) can generate encoded intra-screen texture coordinates (1035). The intra-screen coordinate syntax information (1031) can generate triangular face motion information (1032). In addition, the intra-screen coordinate syntax information (1031) can generate a triangular face vector.

[0181] The above triangular motion information (1032) can be generated by the intra-screen coordinate syntax information (1031). The texture coordinates can refer to texture coordinates for the second frame. The second frame can refer to a subsequent frame. The texture coordinates of the second frame can generate encoded inter-screen texture coordinates (1033).

[0182] Encoded inter-screen texture coordinates (1033) can be generated based on the triangular motion information (1032). The encoded inter-screen texture coordinates (1033) can be used for texture mapping (1050). When the encoded inter-screen texture coordinates (1033) are used for texture mapping (1050), the latest texture can be used as a reference. The encoded intra-screen texture coordinates (1035) can be generated based on the intra-screen coordinate syntax information (1031). The encoded intra-screen texture coordinates (1035) can be used for texture mapping (1050). When the encoded intra-screen texture coordinates (1035) are used for the mapping (1050), a refresh for the texture buffer can be requested.

[0183] A mesh bitstream (1041) may be generated from the merged dynamic mesh bitstream (1010). The mesh bitstream (1042) may be used in a mesh decryption (1042) procedure. The mesh decryption (1042) may refer to a procedure of decrypting an encoded mesh bitstream to generate mesh information. A reconstructed mesh (1043) may be generated as a result of the mesh decryption (1042).

[0184] Texture mapping (1050) may refer to a process of triangular texture mapping based on the texture atlas (1023), the encoded inter-screen texture coordinates (1033), and the encoded intra-screen texture coordinates (1035). Re-coloring (1051) may refer to a process of changing the color of the texture included in the texture mapping again. Based on the decoding procedure as described above, a dynamic mesh sequence (1001) including at least one frame (1000) as described above may be restored.

[0185]

[0186] Image / video-based mesh compression structure

[0187] FIG. 11 is a high-level conceptual diagram of a dynamic mesh encoding procedure according to one embodiment of the present invention, and FIG. 12 is a high-level conceptual diagram of a dynamic mesh decoding procedure according to one embodiment of the present invention.

[0188] According to another representation method of the present invention, a dynamic mesh sequence can be represented as a set of frames F(0), F(1), F(2), …, F(n). Each of the frames can include at least one 3D mesh, i.e., M(0), M(1), M(2), …, M(n). Each of the meshes M(i) can be defined as a combination of individual pieces of information including connectivity C(i), geometric information G(i), texture coordinates T(i), and texture connectivity CT(i). In addition, each of the meshes M(i) can be associated with one or more attribute maps A(i, 0), A(i, 1), …, A(i, D-1). The attribute maps can include 2D images / videos as described above, and can include information configured to describe a set of attributes associated with a surface of the mesh. Examples of the above attributes may include texture information, and may also include information including attributes for vertices, such as color, normal, transparency, etc. Accordingly, as described above, it can be understood that each frame F(i) constituting the dynamic mesh sequence may include mesh information M(i) and attribute map (texture) information A(i,j).

[0189] As described above, the geometry and attribute information of a 3D mesh can be efficiently compressed by mapping it back to a 2D image using conventional image / video encoding techniques. However, there is a limitation in that connectivity information cannot be encoded efficiently when using a similar method, and therefore, an encoding method optimized for such information may be required. Referring to FIG. 11, the encoding process may include a preprocessor (1103) that receives a static or dynamic mesh M(i) and an attribute map A(i). The preprocessor generates a base mesh m(i) and a displacement d(i), which can be provided to an encoder (1102), and the encoder can generate a compressed bitstream b(i) therefrom. The encoder (1102) may also directly receive the attribute map A(i). The encoder (1102) may correspond to any one of the dynamic mesh encoders exemplified in FIGS. 7 to 9, and may also include other encoder implementation methods. The feedback loop (1101) enables the encoder (1102) to guide the preprocessor (1103) and change its parameters to achieve the best compromise for encoding the bitstream b(i) based on various criteria such as rate distortion, encoding / decoding complexity, random access, reconstruction complexity, terminal performance, encoding / decoding power consumption, and / or network bandwidth and latency, and not limited to the examples described above.

[0190] Referring to FIG. 12, at the decoder side, a compressed bitstream b(i) may be received by a decoder (1202) that decodes the bitstream to generate METADATA(i) associated with the bitstream and a decoded mesh, a decoded mesh m'(i), a decoded displacement d'(i), and a decoded attribute map A'(i). The decoder (1102) may correspond to the dynamic mesh decoder exemplified in FIG. 10, and may also include other decoder implementations. Each of the outputs of the decoder (1202) may be provided to a postprocessor (1203) that may perform various postprocessing steps, such as adaptive tessellation. The postprocessor (1203) may generate a postprocessed mesh M"(i) and a postprocessed attribute map A"(i) corresponding to the input mesh M(i) and the input attribute map A(i) provided to the encoder. (However, it is obvious that the output is not identical to the input due to the lossy nature of compression caused by quantization and other encoding effects.) The application function (1201) that consumes the content can provide feedback (1201a) to the decoder (1202) to guide the decoding process, and can also provide feedback (1201b) to the post-processor (1203).

[0191] FIG. 13 is an exemplary diagram illustrating a method of preprocessing according to one embodiment of the present invention. Referring to FIG. 13, an exemplary preprocessing method that can be applied by the preprocessor (1103) of FIG. 11 is illustrated. Although the example of FIG. 11 uses the case of a 2D curve to simplify the explanation, the same concept can also be applied to an input static or dynamic 3D mesh M(i)=(C(i), G(i), T(i), CT(i)) for generating a base mesh m(i) and a displacement field d(i), as shown in FIG. 11. Referring to FIG. 13, for example, an input curve (1301) (represented as a 2D polyline), referred to as an original curve, is first downsampled to generate a base curve (1302), also referred to as a decimated curve. Various subdivision methods known in the art can be applied to the abbreviated curve (1302) to generate subdivided curves (1303). As an example, as illustrated in FIG. 13, a subdivision method using an iterative interpolation method can be applied. This may involve inserting a new point in the middle of each edge of the curve at each iteration. For example, in the example illustrated in FIG. 13, two subdivision iterations are applied, and by subdividing once and then subdividing again, a line segment is divided into four pieces in total, and three new points are inserted.

[0192] According to one embodiment of the present invention, the method of the present invention can be independent of the selected subdivision method and can be combined with any subdivision method that is known in the art or may be newly provided. The subdivided curve (1303) can then be deformed to better approximate the original curve. More precisely, a displacement vector can be calculated for each vertex of the subdivided curve (1303) (as indicated by the arrows of the displaced curve (1304) in FIG. 13 ), such that the shape of the displaced curve can be sufficiently close to the shape of the original curve, as in the example shown in FIG. 9 . One advantage of the subdivided curve (mesh) (1303) is that it can have a subdivision structure that allows for more efficient compression while providing a faithful approximation of the original curve (i.e., mesh). This improved compression efficiency can be achieved due to various properties, including but not limited to:

[0193] In one embodiment, the base (condensed) curve (1302) may have fewer vertices than the original curve (1301), and thus may require fewer bits to encode and / or transmit.

[0194] In one embodiment, when the base curve (1302) is decoded, the subdivided curve (1303) may be automatically generated by the decoder. That is, no information other than the type of subdivided scheme to be encoded and / or transmitted and the number of subdivided iterations may be required, thereby reducing the amount of information.

[0195] In one embodiment, the displaced curve (1304) may be generated by decoding a displacement vector associated with each vertex of the segmented curve (1303). In particular, depending on the embodiment, the segmentation structure may provide higher compression performance by enabling information component decomposition by efficient wavelet transform in addition to allowing spatial / quality scalability.

[0196] When applying the same concept as illustrated in Figure 11 to an input mesh M(i), a base (reduced) mesh can be generated using various mesh reduction techniques, both known and newly available. Furthermore, a refined mesh can be generated using various subdivision methods, both known and newly available. The corresponding displacement field d(i) can be computed using various methods, both known and newly available.

[0197] A re-sampling procedure can compute a new parameterization atlas that may be more suitable for compression. For dynamic meshes, this can be achieved using a temporally consistent remeshing procedure that can generate the same subdivision structure shared by the current mesh M'(i) and the reference mesh M'(j). This temporally consistent remeshing procedure can enable skipping the encoding of the base mesh m(i) and reusing the base mesh m() associated with the reference frame M(). This can also enable better temporal prediction of both attribute and geometry information. More precisely, a motion field f(i), which describes how to move the vertices of m() to match the positions of m(i), can be computed and encoded, as described in more detail below.

[0198] FIG. 14 is a block diagram detailing an on-screen encoding procedure according to one embodiment of the present invention. FIG. 14 may include all or part of the dynamic mesh encoder shown in FIGS. 7 to 9 described above. Furthermore, FIG. 14 may also represent a detailed configuration of the encoder (1102) shown in FIG. 11 described above. However, FIG. 14 is not necessarily to be described or understood in a continuous manner with the other drawings described above.

[0199] According to one embodiment of the present invention, base mesh encoding may be performed by the following procedure. The base mesh m(i) associated with the current frame may first be quantized (1401) (e.g., using uniform quantization) and encoded using a static mesh encoder (1402). In the present invention, the encoding method and apparatus are independent of which codec for mesh compression is used. That is, any suitable one of various mesh compression codecs may be combined or fused with the technology described in the present invention and used together. The mesh compression codec used may be explicitly specified in the bit string by encoding the identifier (ID) / ID of the corresponding codec, or may be implicitly defined / specified in advance by a fixed specification of an application function, etc.

[0200] Since the quantization step and / or the mesh compression module may be lossy, taking into account the loss on the decoder side, a reconstructed quantized version of m(i), denoted as m'(i), can be derived by a static mesh decoder (1403) within the frame encoder within the screen. It can be appreciated that the above structure corresponds to a structure similar to the encoding loop shown in the structure of the general video codec. In some embodiments, if the mesh information is losslessly encoded and / or the quantization step is omitted, m(i) will exactly match m'(i).

[0201] Displacement coding can be performed by the following procedure. Depending on functional requirements, target bit rate, and / or visual quality, the encoder can optionally encode a set of displacement vectors associated with the refined mesh vertices, referred to as displacement field d(i). The reconstructed quantized basis mesh m'(i) can be used to update the displacement field d(i) by displacement update (1404) to generate an updated displacement field d'(i) that takes into account the difference between the reconstructed basis mesh m'(i) and the original basis mesh m(i). By utilizing the structure of the refined surface mesh, a wavelet transform (1405) can be applied to d'(i) to generate a set of wavelet coefficients e(i). The wavelet coefficients e(i) can then be quantized (1406) to generate quantized wavelet coefficients e'(i). Next, the above coefficients can be packed into a 2D image / video by image packing (1407) and compressed using an image / video encoder (1408).

[0202] The encoding of the above wavelet coefficients may be lossless or lossy. Therefore, a reconstructed version of the wavelet coefficients can be obtained by applying image unpacking (1409) and inverse quantization (1410) to the reconstructed packed and quantized wavelet coefficients (i.e., the image / video encoding result based on the wavelet coefficients) generated in the above-described encoding procedure. The reconstructed displacement d"(i) can be calculated by applying an inverse wavelet transform (1411) to the reconstructed wavelet coefficients. The reconstructed base mesh m"(i) can be obtained by applying inverse quantization (1412) to the reconstructed quantized base mesh m'(i). The reconstructed warped mesh DM(i) can be obtained by subdividing the m"(i) and applying the reconstructed displacement d"(i) to each vertex by a warped mesh reconstruction process (1413).

[0203] Fig. 15 is a conceptual diagram of a mesh segmentation method according to an embodiment of the present invention. Depending on the embodiment, various segmentation methods may be used for segmenting the mesh. According to an embodiment of the present invention, a midpoint segmentation method may be applied. Referring to Fig. 15, the midpoint segmentation method may mean a method of dividing each side of the triangle shown in each segmentation iteration into four sub-triangles by bisecting each triangle. For example, an initial condition S having two triangles (1501 and 1502) 0 Starting with two triangles, the first iteration S 1 can generate four sub-triangles (1501a-1501d) for triangle (1501) and four sub-triangles (1502a-1502d) for triangle (1502). Each sub-triangle is generated by subsequent iterations S 2 can be further split in . The new vertex (1503) is repeated S 1 can be introduced in the middle of each edge in S, and a new vertex (1504) is repeated 2 is introduced in the middle of each edge. Since the connectivity to geometry and texture coordinates may be different, the subdivision procedure can be applied independently to geometry and texture coordinates.

[0204] In some embodiments, the operation of the segmentation method may be adaptively modified based on various implicit and / or explicit criteria (e.g., to preserve sharp edges of the mesh). In one embodiment of the present invention, the modification may be based on face / edge / vertex-specific attribute information associated with the base mesh and explicitly encoded as mesh properties by the mesh codec. In another example, the method of the segmentation operation may be dynamically modified based on the results of analyzing the base mesh or mesh in a previous encoding / decoding process.

[0205] Depending on the embodiment, various wavelet transforms, both conventionally known and newly available, may be applied. According to one embodiment of the present invention, a low-complexity wavelet transform implementation using a lifting method may be used. In this implementation, first, prediction weights controlling the prediction step may be used, and second, update weights controlling the update step may be used.

[0206] The displacement field d(i) can be defined in the same Cartesian coordinate system as the input mesh. In some cases, it may be a feasible optimization to transform d(i) from this standard coordinate system to a local coordinate system. This can be defined as a normal to the refined mesh at each vertex. The normal vector associated with the refined mesh can be directly decoded or computed based on the quantized geometry, or the normal vector associated with the vertex introduced during the refinement process can be computed as described above.

[0207] One potential advantage of local coordinate systems for displacement is that the tangential component of the displacement can be quantized more than the normal component. In many cases, the normal component of the displacement can have a greater impact on the quality of the reconstructed mesh than the two tangential components.

[0208] The decision to use either a regular or local coordinate system can be made at the sequence, frame, patch group, or patch level. This decision can be made explicitly by encoding additional properties associated with the base mesh vertices, edges, or faces, or it can be made implicitly by analyzing the connectivity / geometry / property information of the base mesh.

[0209] Depending on the embodiment, various strategies may be used to quantize the wavelet coefficients of the displacement for the present invention. One embodiment includes using a uniform quantizer with a dead zone and adjusting the quantization step so that high-frequency coefficients are quantized more. Instead of directly defining the quantization step, individual quantization parameters may be used. As a more sophisticated adaptive quantization method, trellis quantization may be used in one embodiment. In another embodiment, a method may be used to simultaneously optimize the quantization parameters for three components to minimize the distance between the reconstructed mesh and the original. In another embodiment, an adaptive quantization rounding method may be used.

[0210] Depending on the embodiment, various strategies may be used to pack wavelet coefficients for the present invention into a 2D image. In one embodiment, coefficients may first be searched from low frequency to high frequency. Then, for each coefficient, an index of an N×M pixel block (e.g., N=M=16) that should be stored according to the raster order for the block may be determined. Finally, the location within the N×M pixel block may be computed using Morton order to maximize locality. In one embodiment, the values ​​of N and M may be selected as powers of 2.

[0211] In accordance with an embodiment, the attribute transfer module for the present invention may be configured to compute a new attribute map based on an input mesh M(i) and an input texture map A(i). The new attribute map may be computed to better fit the reconstructed deformed mesh MD(i).

[0212] In the image / video encoding for the present invention, which image / video encoder or standard is used is irrelevant to the implementation of the technical idea of ​​the present invention, which may mean that various image / video codecs that are known in the art or may be newly provided are applicable. According to one embodiment, when encoding the displacement wavelet coefficients, quantization may be applied to a separate module, so that a lossless approach may be used. According to another embodiment, the image / video encoder may be configured to compress the coefficients in a lossy manner and apply quantization to the original domain (i.e., spatial domain) or the transform domain (i.e., frequency domain).

[0213] As with traditional 2D image / video encoding, it can be configured to optionally apply color space transformation and chroma subsampling to achieve better rate distortion performance (e.g., converting RGB 4:4:4 to YUV 4:2:0). When applying the color space transformation and chroma subsampling procedures, it may be advantageous to consider surface discontinuities of the texture image. For example, this can be done by considering only samples belonging to the same patch and excluding potentially empty areas.

[0214] FIG. 16 is a block diagram detailing an inter-screen encoding procedure according to one embodiment of the present invention. FIG. 16 may include all or part of the dynamic mesh encoder shown in FIGS. 7 to 9 described above. Furthermore, FIG. 16 may also represent a detailed configuration of the encoder (1102) shown in FIG. 11 described above. However, FIG. 16 is not necessarily to be described or understood in a continuous manner with the other drawings described above.

[0215] According to one embodiment of the present invention, FIG. 16 may be understood to illustrate a block diagram of an inter-frame encoding procedure, i.e., an encoding procedure in which encoding relies on temporally separated (e.g., previous) versions of a mesh. According to one embodiment of the present invention, a reconstructed quantized reference basis mesh m'(j) may be used to predict a current frame basis mesh m(i). The preprocessing module described above may be configured to perform the preprocessing operation such that m(i) and m'(j) share the same number of vertices, connectivity, texture coordinates, and texture connectivity. Therefore, only the positions of vertices may differ in m(i) and m'(j).

[0216] The motion field f(i) corresponding to the vertex displacement between m(i) and m'(j) can be computed by the motion encoder (1601) considering the quantized version (1602) and the reconstructed version of m(i). In one embodiment, vertices can be merged / removed, so that m'(j) can have a different number of vertices than m(j). Therefore, the mesh encoder can be configured to track the transformation applied to obtain m'(j) from m(j). Next, the mesh encoder can be configured to ensure a one-to-one correlation between m(j) and the transformed and quantized version m*(i) of m(i) by applying the same transformation to m(i). Next, the motion field f(i) can be computed by the motion encoder (1601) by subtracting the quantized positions p(i,v) of the vertex v of m*(i) from the positions p(j, v) of the vertex v of m'(j). Next, the motion field can be further predicted using the connection information of m'(j), and the result can be entropy encoded using, for example, context-adaptive binary arithmetic coding.

[0217] Since the motion field compression procedure may be lossy, considering the loss on the decoder side, the reconstructed motion field denoted as f'(i) can be computed by applying the motion decoder module (1603). It can be understood that the above structure corresponds to a structure similar to the encoding loop shown in the structure of the general video codec. The reconstructed quantized basic mesh m'(i) can be computed (1604) by adding the motion field f'(i) to the position of m'(j). The remaining encoding process may be the same as or similar to the frame encoding process within the screen described above.

[0218] FIG. 17 is a block diagram detailing an in-screen decryption procedure according to one embodiment of the present invention. FIG. 17 may include all or part of the dynamic mesh decoder shown in FIG. 10 described above. FIG. 17 may also represent a detailed configuration of the decoder (1202) shown in FIG. 12 described above. However, FIG. 17 is not necessarily to be described or understood in a continuous manner with the other drawings described above.

[0219] According to one embodiment of the present invention, in-screen decoding may be performed by the following procedure. Referring to FIG. 17, first, the bitstream b(i) may be demultiplexed into three or more individual sub-bitstreams, namely (1) a mesh sub-bitstream, (2) a position-versus-displacement and potential additional information sub-bitstream for each vertex, and (3) an attribute map sub-bitstream for each attribute map (1701). According to another embodiment of the present invention, an atlas sub-bitstream including patch information may be included, and in this case, a method based on various visual volumetric video-based coding (V3C) and / or video-based point cloud coding (V-PCC) standards that may be provided in the past or in the future may be applied in the same or similar manner.

[0220] The mesh sub-bitstream can be supplied to a static mesh decoder (1702) corresponding to a mesh encoder used to encode the sub-bitstream to generate a reconstructed quantized base mesh m'(i). Next, a decoded base mesh m"" can be obtained by applying inverse quantization (1703) to the reconstructed m'(i). In the present invention, the decoding method and apparatus are independent of which codec for mesh compression is used. That is, any suitable one of various codecs for mesh compression can be combined or fused with the technology described in the present invention and used together. The mesh compression codec used can be explicitly specified in the bitstream, or can be implicitly defined / specified in advance by a fixed specification of an application function, etc.

[0221] The displacement sub-bit stream can be decoded by an image / video decoder (1704) corresponding to the image / video encoder used to encode the corresponding sub-bit stream. The generated image / video can be unpacked (1705), and inverse quantization (1706) can be applied to the resulting wavelet coefficients. For the implementation of the image / video decoder (1704), various image / video codecs that are known in the art or may be newly provided can be used, for example, image / video codecs such as VVC / H.266, HEVC / H.265 AVC / H.264, AV1, AV2, JPEG, JPEG2000, etc. can be utilized. The use of such codecs allows the mesh encoding and decoding techniques of the present invention to utilize well-developed encoding and decoding algorithms implemented in hardware of various platforms, thereby providing high performance and high power efficiency.

[0222] According to another embodiment of the present invention, the displacement information can be decoded by a decoder dedicated to displacement data. According to an embodiment, a motion decoder or a dictionary-based decoder such as ZIP used to decode mesh motion information can be used as the decoder dedicated to displacement data. In this case, the decoded displacement d"(i) can be generated by applying an inverse wavelet transform (1707) to the unquantized wavelet coefficients. The final decoded mesh M"(i) can be generated by applying a deformation mesh reconstruction procedure (1708) to the decoded base mesh m"(i) and adding information of the decoded displacement field d"(i).

[0223] The attribute sub-bit stream can be decoded by an image / video decoder (1709) corresponding to the image / video encoder used to encode the corresponding sub-bit stream. The decoded attribute map A"(i) can be generated directly as an output of the decoder and / or through an appropriate color format / color space conversion (i.e., a recoloring procedure) (1710) for the output. As with the displacement sub-bit stream, various image / video codecs known in the art or newly provided can be used, for example, VVC / H.266, HEVC / H.265, AVC / H.264, AV1, AV2, JPEG, JPEG2000, etc. can be utilized, and the advantages obtained thereby can also be similar.

[0224] According to another embodiment of the present invention, the attribute sub-bit string can be decoded using a decoder that is independent of the image / video, for example, a dictionary-based decoder such as ZIP. In one embodiment, multiple sub-bit strings associated with different attribute maps can be decoded by different decoders. Furthermore, in some embodiments, each sub-bit string can use a decoder based on a different codec.

[0225] FIG. 18 is a block diagram illustrating details of an inter-frame decoding procedure according to an embodiment of the present invention. FIG. 18 may include all or part of the dynamic mesh decoder shown in FIG. 10 described above. In addition, FIG. 18 may refer to a detailed configuration of a decoder (1202) shown in FIG. 12 described above. However, FIG. 18 is not necessarily to be described or understood in a continuous line with other drawings described above. According to an embodiment of the present invention, intra-frame decoding may proceed by the following procedure. Referring to FIG. 18, first, a bit stream may be separated (1801) into three individual sub-bit streams: (1) a motion sub-bit stream, (2) a position-to-displacement and potential side information sub-bit stream for each vertex, and (3) an attribute sub-bit stream for each attribute map. According to another embodiment of the present invention, an atlas sub-bit sequence including patch information may be included, and in this case, a method based on various V3C and / or V-PCC standards that may be provided in the past or in the future may be applied in the same or similar manner to the specific configuration and inclusion method of the sub-bit sequence.

[0226] The motion sub-bit stream can be decoded by applying a motion decoder (1802) corresponding to the motion encoder used to encode the sub-bit stream. As described above in connection with the present invention, various motion codecs / standards can be used to decode the motion information. In other words, various motion decoding methods can be used. Thereafter, the decoded motion information can be optionally added to the decoded reference quantized base mesh m'(j) through a base mesh reconstruction procedure (1803) to generate a reconstructed quantized base mesh m'(i). That is, the already decoded mesh in instance j can be used to predict the mesh in instance i (together with the motion information). Thereafter, the decoded base mesh m""(i) can be generated by applying inverse quantization (1804) to m'(i).

[0227] The displacement and attribute sub-bitstreams can be decoded in the same or similar manner as described above with respect to the decoding procedure of the frames within the screen. The decoded mesh M" (i) can also be reconstructed in the same or similar manner. The details of the dequantization and reconstruction procedure are not limited by the specification of the present invention and are not normative, and can be implemented in various ways known in the art or provided in the future or combined with the rendering procedure of the mesh.

[0228] According to one embodiment of the present invention, an additional post-processing module may be applied to improve the visual / objective quality of the decoded mesh and attribute map or to adapt the resolution / quality of the decoded mesh and attribute map to the viewpoint or terminal performance. The post-processing may include, depending on the embodiment, at least one of color format / color space conversion, utilization of patch information and occupancy maps to assist chroma upsampling, geometry smoothing, attribute smoothing, image / video smoothing / filtering algorithms, and adaptive tessellation.

[0229] According to one embodiment of the present invention, in some application methods, it may be advantageous to subdivide a mesh into sets of patches (i.e., sub-parts) and optionally group the patches as patch groups / tile sets. In some cases, various parameters (segmentation, quantization, wavelet transform, coordinate system, etc.) may be used to compress each patch or patch group. In such cases, it may be desirable to encode the patch information into separate sub-bitstreams. In this case, the sub-bitstreams may be encoded using the V3C and / or V-PCC methods described above. This technique may be advantageous in handling cracks at patch boundaries and may be configured to utilize, for example, lossless encoding of boundary vertices, verification of matching of position / vertex attributes after applying displacement, use of local coordinate systems, and optionally disabling quantization of wavelet coefficients.

[0230] The encoder / decoder arrangement described herein can also support various levels of scalability. For example, various transformations can be applied, including temporal scalability achieved through temporal subsampling and frame reordering. Similarly, quality and spatial scalability can be achieved using different mechanisms for shape / vertex attribute data and attribute map data. As one example, shape scalability can be achieved by utilizing a subdivision structure, and mesh resolution can be changed by moving from one level of detail to the next. Displacement information can be segmented and stored into multiple image / video sub-bitstreams. The multiple image / video sub-bitstreams can be hierarchized according to the frequency bands of wavelet coefficients. A base layer can constitute a first refinement layer as a sub-bitstream for low-frequency wavelet coefficients. Wavelet coefficients in a higher frequency band than the base layer can be constituted as a sub-bitstream of a second refinement layer, and wavelet coefficients in sequentially higher frequency bands can be constituted as sub-bitstreams of subsequent refinement layers. The wavelet coefficients of the highest frequency band can be composed of sub-bitstreams of the final refinement layer. For example, detail level m-1 can be combined with detail level m-1 to generate detail level m. In addition, the attribute map can be scalably encoded using scalable image / video coding techniques, such as other approaches that support quality / spatial scalability in image / video codecs such as VVC / H.266, HEVC / H.265, AVC / H.264, VVC, AV1, AV2, JPEG, and JPEG2000.

[0231] According to one embodiment of the present invention, region of interest (ROI) encoding can be provided by configuring the encoding procedure described above to perform encoding at a higher resolution and / or higher quality for at least a portion of the geometry, vertex attributes, and / or attribute map data corresponding to the region of interest. Such a configuration can be useful for providing content of higher visual quality under narrow bandwidth and complexity constraints. According to one embodiment, when encoding a mesh representing a person, a higher quality can be used for the face as opposed to the rest of the body. Priority / importance / space / bounding box information can be associated with patches, patch groups, tiles, NAL units, and / or low-order bitstreams in a manner that allows a decoder to adaptively decode a subset of the mesh. It will be appreciated that any combination of such encoding units can be used together to achieve such functionality. For example, NAL units and low-order bitstreams can be used together.

[0232] According to one embodiment of the present invention, temporal and spatial random access functions may be provided. The temporal random access function may be implemented by introducing intra-random access points (IRAPs) into various sub-bitstreams (e.g., atlas, image / video, mesh, motion, and displacement sub-bitstreams). In some embodiments, spatial random access functions may be supported through the definition and use of tiles, image segmentation units, patch groups, and / or patches, or combinations of such coding units. To facilitate the random access, in some embodiments, metadata describing the relationship between the layouts of different units may be generated and included in the bitstream to assist a decoder in determining which unit to decode.

[0233] According to one embodiment of the present invention, lossless geometry / vertex attribute encoding may be provided. The lossless encoding may be supported by various applied methods, such as remeshing, applying a type of refinement that sets the base mesh to be the same as the input mesh by setting the subdivision level to 0, quantizing the base mesh, disabling one or more of the displacement sub-bitstream calculations, and other various methods. According to another embodiment of the present invention, a simplified version of the base mesh (e.g., a quantized low-quality version) may be configured to be encoded together with the displacement information, such that a decoder may be configured to obtain a higher-quality version that includes a level of consistency with the original mesh information. Furthermore, according to an embodiment, lossless attribute map encoding may be supported by configuring an image / video encoder to compress the attribute map using an inherently lossless compression method, such as pulse code modulation (PCM) mode.

[0234] According to one embodiment of the present invention, as a method for maintaining high-quality texture coordinates, a separate displacement sub-bitstream may be sent for texture coordinates. Similarly, a motion sub-stream may also be used for texture coordinates.

[0235] According to one embodiment of the present invention, per-vertex attributes may be compressed in the same manner as geometric information. For example, a mesh compression codec may be used to encode vertex attributes associated with base mesh vertices. Wavelet-based encoding may be used for attributes associated with high-resolution meshes, which may be stored / transmitted as separate vertex attribute sub-bitstreams. Equivalent procedures applied on the decoder side may decode / restore / reconstruct vertex attribute information.

[0236] According to one embodiment of the present invention, the same encoding / decoding method can be configured to be applied even to quad / polygonal meshes. In this case, methods such as using a mesh compression codec capable of encoding quad / polygonal meshes, or selecting a mesh subdivision method suitable for non-triangular meshes, such as Catmull-Clark or Doo-Sabin, can be used, and various other application methods can be combined.

[0237] According to one embodiment of the present invention, in the array notation shown in the above-described implementation method, texture coordinates for a base mesh can be explicitly specified and encoded in a bitstream by a mesh encoder. According to another embodiment of the present invention, implicit texture coordinates derived from position information can be used through a projection method identical or similar to that used in V-PCC or MPEG immersive video (MIV). Alternatively, other models (e.g., B-spline surfaces or polynomial functions, etc.) can be configured to be considered in the representation and / or encoding of the coordinates.

[0238]

[0239] Remeshing for efficient compression

[0240] Figure 19 is a block diagram of a remeshing procedure according to one embodiment of the present invention. Referring to Figure 19, the input mesh M(i) may be an irregular mesh. The output may be a base mesh m(i) having a refined version of m(i) and an associated displacement set d(i).

[0241] According to one embodiment of the present invention, duplicate vertices (i.e., vertices having the same location) or vertices having close 3D locations (e.g., vertices having a distance less than a user-defined threshold) can be merged through duplicate vertex removal (1901). Depending on the embodiment, the duplicate vertex removal procedure can be accelerated by utilizing various application data structures including a hash table, a kd-tree, and an octree. By removing duplicate vertices, cracks in the mesh can be suppressed in subsequent processing steps including decimation of the mesh. Additionally, duplicate vertex removal can be configured to improve encoding efficiency and / or encoding / decoding complexity by eliminating calculations that use or are based on unnecessary data.

[0242] According to one embodiment of the present invention, a mesh can be compressed through mesh reduction (1902). As described above, the reduction may mean simplifying the mesh by reducing the number of vertices / faces while substantially maintaining the shape of the original mesh. In this case, "substantially preserving the shape of the original mesh" may include preserving the shape of the input mesh sufficiently to achieve the desired encoder and / or decoder performance while achieving a desired level of accuracy or fidelity in the resulting mesh representation. Accordingly, the definition may be defined differently depending on the embodiment and implementation aspects of the present invention, and specifically, may be set differently depending on the functions of the available encoder and decoder and their driving equipment, the functions of the display or other output equipment, and / or the requirements of specific application functions / programs.

[0243] According to one embodiment of the present invention, the mesh reduction (1902) procedure may be configured to apply a mesh reduction algorithm that extends any mesh reduction algorithm known in the art or newly provided by tracking a mapping between a full-resolution input mesh and a reduced mesh. More specifically, in each iteration of the reduction procedure, the mesh reduction (1902) procedure may be configured to project the removed vertices onto a reduced version of the mesh.

[0244] According to one embodiment of the present invention, the mesh reduction (1902) procedure may be configured to project the removed points onto the closest counterpart of the simplified mesh. In some embodiments, the "closest" counterpart may mean a vertex with the shortest L2 distance in 3D space, etc. In some embodiments, other criteria may be used to define the projection procedure. For example, instead of the L2 distance in 3D space, another distance measure in 3D space may be used (e.g., L1, Lp, L_inf, etc.). Alternatively, distances in a lower dimensional space may be used by projecting onto a 2D local plane. In this case, orthogonal and / or non-orthogonal projections may be used as the projection method. Other projection procedures may be appropriately used depending on the given use case. Furthermore, in some embodiments, a modified simplification algorithm may be applied to prevent errors such as triangle flipping in the reduced mesh and / or projected mesh. Such an implementation can help create a better mapping between the condensed and projected meshes.

[0245] According to one embodiment of the present invention, duplicate triangles (i.e., triangles referencing the same vertex) can be detected and removed from a reduced mesh dm(i) through a duplicate triangle removal (1903) procedure. This can improve compression efficiency and encoding / decoding complexity. However, depending on the embodiment, the duplicate triangle removal (1903) procedure may be selectively applied or not applied.

[0246] According to one embodiment of the present invention, a small connected element removal procedure (1904) may be configured to remove some connected elements. The connected elements may refer to a set of vertices that are connected to each other but not to the rest of the mesh. When selecting small connected elements to be removed from among the connected elements, at least one small connected element may be selected that has fewer triangles or vertices than a user-defined threshold (e.g., 8) and / or occupies an area below a user-defined threshold ratio (e.g., 0.1% or less of the original mesh area). These small connected elements, while requiring high encoding costs, have only a limited impact on the final visual quality of the model. Therefore, removing them based on the aforementioned criteria may help improve compression efficiency and complexity during encoding.

[0247] The criteria for selecting and / or removing small connected elements may be selected to be fixed for the entire mesh, or may be selected adaptively based on user-provided information describing local surface properties, the importance of individual meshes, and / or saliency sub-parts. For example, for a mesh containing a representation of a person, a stronger removal criterion may be used for the region depicting the head so that fewer small connected elements are removed to preserve facial details, while a more relaxed removal criterion may be used for the region depicting the body so that relatively more small connected elements are removed. Additionally or alternatively, the threshold for selecting / removing small connected elements may be adjusted adaptively and / or variably based on encoding / decoding parameters such as speed / distortion criteria, complexity criteria, power consumption criteria, and output target bit rate. According to one embodiment of the present invention, such thresholds may be provided as or derived from feedback (1101) from an encoder module as illustrated in FIG. 11 .

[0248] The parameter information associated with an input mesh M(i) may be suboptimal in that it may define an unnecessarily large number of small patches, making reduction, remeshing, and compression difficult. Therefore, according to one embodiment of the present invention, instead of attempting to preserve the initial parameter information during the reduction process, the parameters may be selectively recomputed via atlas parameterization (1905) for a reduced mesh dm(i), or a mesh cm(i) with duplicate triangles and small connected elements removed.

[0249] A remeshing method according to an embodiment of the present invention may include a mesh refinement (1906) procedure implemented using various mesh refinement techniques, which may be conventionally known or newly provided. The remeshing method according to an embodiment of the present invention may be used with any of various mesh refinement techniques, which may or may not be conventionally known or newly provided, described or not described herein. Depending on the embodiment, for triangular meshes, mid-edge interpolation, loop, butterfly, and Catmull-Clark refinement techniques may be applied. The exemplary methods described above may offer various compromises in terms of computational complexity, generality (e.g., whether they are applicable to quadrangle / polygon-based meshes in the same / similar manner as triangle-based meshes), and approximation and smoothness of the generated surface, which may affect the rate distortion performance of the encoder.

[0250] According to one embodiment of the present invention, the vertices of the mesh S(i) segmented by the initial mesh deformation (1907) procedure can be moved to deform it to have a shape close to the input mesh M(i). At this time, the quality of the approximation through the deformation can directly affect the speed distortion performance of the encoder. According to one embodiment, the process can be performed by the following sequence of algorithms. (1) For each vertex v of the segmented mesh S(i), Pos(v) representing its initial 3D position and N(v) representing its normal vector are obtained. (2) For each initial 3D position Pos(v), the nearest point H(v) on the representation of the projected mesh P(i) is obtained. At this time, the angular difference between the normals N(v) and H(v) may be required to be less than a user-defined threshold. Various distance designation methods can be used to obtain the proximity distance, including, but not limited to, L1, L2, Lp, and Linfo. The threshold may be fixed for the entire mesh or may be adaptively changed based on user-supplied information describing the importance or saliency of local surface features and / or mesh sub-parts (e.g., face vs. body). Additionally or alternatively, the threshold may be based on a rate distortion criterion or other criteria (e.g., complexity, power consumption, bit rate, etc.) provided as feedback from the encoder module. The H(v) may be identified by the index of the triangle it belongs to (tindex(v)) and the barycentric coordinates (a, b, c) for that triangle. Since the projected mesh P(i) and the mesh UM(i) have a 1:1 mapping between their vertices (i.e., have the same connectivity), the point H'(v) located in UM(i) can be computed. The barycentric coordinates (a, b, c) can be used with respect to the triangle with index tindex(v) in the mesh UM(i).

[0251] According to one embodiment of the present invention, the iterative mesh warping (1908) procedure may have as input an initial deformed mesh F0(i) and may generate a final deformed mesh F(i) therefrom. The iterative mesh warping (1908) procedure may be configured to iteratively apply an algorithm, including, according to an embodiment: (1) Recalculate normal vectors associated with mesh vertices. (2) For each vertex v of the deformed mesh having a position Pos(v), find the closest point H(v) in the input mesh M(i). Here, the angle between the normal vectors associated with Pos(v) and H(v) may be required to be less than a user-defined threshold. As mentioned above, the closest point may mean the point with the smallest distance, and various distances such as L1, L2, Lp, Linf, etc. are used. Also, as mentioned above, the threshold can be fixed for the entire mesh and adaptively changed based on user-provided information describing the importance or saliency of local surface features and / or sub-parts of the mesh (e.g., face vs. body). It can also be adaptively changed based on rate distortion criteria or other criteria (e.g., complexity, power consumption, bit rate) provided as feedback from the encoder module. (3) The vertex v can be moved to a new location determined by: Optionally, check if the triangle is not flipped in the previous step (i.e., its normal vector is not inverted), otherwise, do not move the vertex v and mark it as a missing vertex. This step can help ensure better remeshing results.

[0252] According to one embodiment of the present invention, a mesh smoothing algorithm can be applied to missing vertices while simultaneously considering updated positions for other vertices. According to one embodiment of the present invention, a mesh smoothing algorithm can be applied to all vertices and parameters can be adjusted to reduce the smoothing intensity based on a fitting iteration index and other criteria. Smoothing can be applied to vertex positions and / or displacement vectors for the initial mesh. According to one embodiment of the present invention, the number of deformation iterations, i.e., the number of iterations through the algorithm described above, can be a user-provided parameter or can be automatically determined based on a convergence criterion. For example, the displacement applied in the last iteration can be configured to fall below a user-defined threshold.

[0253] According to one embodiment of the present invention, the final deformed mesh F(i) and the reduced mesh pm(i) that have undergone refinement processing may be taken as input through a basic mesh optimization (1909) procedure. In some embodiments, if the iterative mesh deformation (1908) procedure is omitted, the initial deformed mesh F0(i) may be replaced with the final deformed mesh F(i). Next, the position of pm(i) may be updated through the basic mesh optimization (1909) procedure to minimize the difference between pm(i) and the refined version F(i) (or F0(i)). According to one embodiment, the minimization of the difference based on the position update may be achieved by applying a sparse linear scheme. According to one embodiment, a conjugate gradient method may be used as a method of efficiently applying the sparse linear scheme. Of course, various conventional or newly provided methods may be applied within a scope that does not depart from the spirit and scope of the present invention.

[0254] According to one embodiment of the present invention, the displacement d(i) can be calculated by taking the difference between F(i) (or F0(i)) and pm(i) through the displacement calculation (1910) procedure, thereby exploiting the correlation between the two meshes and generating a more compressible representation. The resulting displacement field d(i) can be fed as input to a dynamic mesh encoder (along with the base mesh m(i) as described above).

[0255] The remeshing procedure according to one embodiment of the present invention described above can be configured to independently process each frame M(i). While this implementation is optimal for intra-frame encoding, using temporally consistent remeshing may enable better temporal prediction for both the mesh and image data.

[0256] According to one embodiment of the present invention, for temporally consistent remeshing, a base mesh pm(j) associated with a reference frame M() can be reused for a base mesh pm(i) having the same connectivity. By ensuring that there is a 1:1 mapping between pm(i) and pm(j), and also ensuring that pm(i) and pm(j) have the same number of vertices, the same number of triangles (or quads / polygons), the same number of texture coordinates, and the same number of texture coordinate triangles (or polygons), pm(i) and pm(j) can be configured to differ only in their vertex positions. In this case, two cases can be assumed: (1) when the input meshes M(i) and M(j) themselves are temporally consistent, and (2) when the input meshes M(i) and M(j) are not temporally consistent.

[0257] In the first case, i.e., when the input meshes M(i) and M(j) are temporally consistent, only the subdivision surface fitting procedure (1911) can be applied. That is, there may be no need to simplify or remove duplicate vertices and connected elements. In this case, the inputs of the subdivision surface fitting (1911) (which may be constructed by each of the procedures 1906-1910 described above) may be the input mesh M(i), the projected mesh P(j) (from the reference frame), and the reduced mesh pm(j) (also from the reference frame), rather than M(i), P(i), and pm(i).

[0258] In the second case, when the input meshes M(i) and M(j) are temporally inconsistent, a deformed version of M(j), M'(j), can be generated with the same shape as M(i). M'(j) can be generated by various techniques, either known or newly available. Then, as described above, only the subdivision surface fitting procedure (1911) can be applied, and instead of M(i), P(i), and pm(i), M'(j), P(j), pm(j) can be provided as input.

[0259]

[0260] Efficient encoding method of displacement information

[0261] Referring back to FIG. 11, the dynamic mesh encoding process in the video dynamic mesh coding (V-DMC) according to one embodiment of the present invention can be divided into a processing process of a preprocessor (1103) and a processing process of a practical encoder (1102). That is, as described above, the operation of the preprocessor (1103) provides various advantages, and can be configured to support, for example, better rate distortion (RD) performance and optionally applicable progressive transmission. The preprocessor (1103) can be configured to receive the ith frame input M(i) of the dynamic mesh as input and output a base mesh m(i) and a series of displacements d(i) for the base mesh m(i). As described above, the related attribute map A(i) can also be adjusted accordingly. The encoder (1102) can be configured to generate a compressed bitstream b(i) by encoding the output values. As described above, it is self-evident that the bit string can be composed of at least one sub-bit string.

[0262] FIG. 20 is a conceptual diagram summarizing the entire encoding pipeline in the V-DMC encoding procedure according to one embodiment of the present invention. FIG. 20 can be understood as illustrating the dynamic mesh encoding procedure according to each of the drawings described above, focusing on the encoding part of the V-DMC encoding displacement. The displacement information after preprocessing can be encoded using an encoder based on the standards of an image / video codec, such as VVC / H.266, HEVC / H.265 AVC / H.264, AV1, AV2, JPEG, JPEG2000, etc., as described above, and can undergo an encoding process together with the basic mesh. The displacement value can be converted and aligned from low frequency to high frequency using wavelet transform, and depending on the standard and operation method of the image / video codec, it can be ultimately output as a frame in YUV color format. In the prior art, a structure identical to / similar to general image / video encoding was adopted for the prediction frame within the screen. In this case, since predictive encoding is performed on a block-by-block basis within a frame, the characteristics of displacement data cannot be utilized, leading to a problem of reduced compression efficiency. Therefore, the present invention proposes a method of compressing displacement images between multiple frames into a single frame by stacking them in a form based on a group of pictures (GOP) unit, thereby reducing compression time and improving compression efficiency.

[0263] The encoding process of displacement data according to one embodiment of the present invention may include the following procedures. After the processing by the preprocessor is completed, the displacement data may be stored as a frame in YUV color format. Within the YUV frame, each pixel block containing displacement information may be assigned a pixel-by-pixel index according to the raster scan order, and a YUV frame storing displacement information may be generated by this index assignment method.

[0264] FIG. 21 is a conceptual diagram illustrating a method of expressing displacement information using YUV frames according to an embodiment of the present invention. In the conventional art, YUV files storing displacement data for an arbitrary source model were handled as GOP units. This may typically mean extracting and encoding 32 frames from a video representing the displacement data. As illustrated in FIG. 21, when the pixel blocks are stacked into an image, rearrangement of the YUV sequence may be required. In one embodiment of the present invention, each YUV frame may have a size of 256×160 pixels, and in this case, the YUV pixel index for displacement data within one frame may have a value from 1 to 40,959 (256×160). Therefore, when defining pixel indices for multiple frames, these values ​​may increase progressively. FIG. 22 is a conceptual diagram for a progressive increase in index information for multiple frames according to an embodiment of the present invention.

[0265] The present invention can use a coding tree unit (CTU) as a basic processing unit to improve compatibility with image / video codecs used for compressing the frame, particularly high-resolution codec standards such as VVC / H.266 or HEVC / H.265, and to improve encoding performance. In addition, in order to maximize spatial consistency, the pixel blocks can be arranged with a high frequency in the upper left area of ​​the frame, and pixel blocks having a zero (0) value can be rearranged to be concentrated in the lower right area. The improved spatial consistency obtained through the pixel rearrangement as described above can significantly improve compression efficiency.

[0266] FIG. 23 is a conceptual diagram for processing displacement information based on CTU-based processing according to one embodiment of the present invention. For example, let's assume that three frames are processed, each frame is processed in YUV format with a size of 256×160 pixels and a 10-bit depth, and the CTU size is 64×64 pixels. In this case, each frame must be divided into blocks of 64×64 pixels. However, since the assumed frame height of 160 pixels is not a multiple of 64, the frame height needs to be modified to a multiple of 64. According to one embodiment, as shown in FIG. 23, a background area (2310) with a height of 32 pixels is added to make the frame height a multiple of 64, so that the resolution can be expanded and modified to a size of 256×192 pixels. According to another embodiment, the background area may be deleted instead of added. For example, if the information value of data in the lower area of ​​the frame is low and can be lost, the resolution may be reduced to 256×128 pixels by deleting the lower 32-pixel section to make the height a multiple of 64. In addition, it is obvious that, depending on the embodiment, the modification of the resolution may be equally applied to the horizontal width, and further, the direction in which addition / deletion to the frame is made may be located in at least one direction among the periphery of the frame, such as up, down, left, and right, and, depending on the embodiment, a modification may be made to add a background area for one of the width and height and delete the other.

[0267] FIG. 24 is a conceptual diagram illustrating a frame stacking process according to an embodiment of the present invention. Following the background area addition process illustrated in FIG. 23, as illustrated in FIG. 24, the frames with the added background can be stacked in GOP units. FIG. 24 illustrates a case of vertical stacking, whereby a GOP composed of three frames can be processed as a single frame with a vertical length of 576 pixels. However, the direction of stacking is not limited in the present invention. For example, the frames can be stacked horizontally, in different stacking directions, or in a grid format. The frames stacked as described above can be output as a single frame. In this process, the pixel index may also need to be rearranged according to the stacking order. This stacking method enables efficient compression while maintaining the spatial relationship between frames.

[0268] Decoding displacement data according to one embodiment of the present invention can be performed to restore a file compressed by the encoding process and return it to its original state. The purpose of the decoding is to reconstruct the structure and content of the original data by reversing the transformation steps performed in the encoding process, i.e., stacking, CTU-based segmentation, and background region addition. In other words, depending on the embodiment, it can mean a series of processes of separating stacked frames, extracting displacement information in units of CTUs from each frame, rearranging the displacement information, and finally reducing it to displacement sub-bit strings and respective processing reference times (frames). Through this decoding process, the file can be effectively restored to its original form, and the integrity and accuracy of the data can be guaranteed.

[0269] In V-DMC, when encoding a dynamic mesh, one of the all-intra (AI) method and the low-delay (LD) method can be selected as a compression condition. In the all-intra (AI) method, only encoding by intra-prediction is enforced, and in the low-delay method, inter-prediction is also allowed. In the prior art, it was common for image / video-based encoders to perform compression by using the AI ​​method for single-frame prediction and selecting the LD method for multi-frame prediction. However, even in multi-frame prediction, the AI ​​method has been used for the first frame to secure better image quality. However, according to the present invention, even when the original mesh of a multi-frame group is compressed by the LD method, it can be seen that superior compression performance can be achieved by directly encoding the single-frame displacement YUV data combined by stacking as described above by applying the AI ​​method.

[0270] FIG. 25 is a conceptual diagram of a stacking procedure for lossless processing according to one embodiment of the present invention. According to one embodiment of the present invention, a comparative analysis may be performed to verify that no data loss occurred during the stacking process of frames as described above. As illustrated in FIG. 25, encoding and decoding processes using an image / video codec may be performed on each of the original frame and the stacked frame. At this time, the efficiency of the merging process may be evaluated by comparing the bit rate and encoding time for the processing results of both sides. In particular, when the background is removed from the stacked and merged frame data and the signal-to-noise ratio (PSNR) is compared after restoring it to a multi-frame format, if the ratio is infinite (INF), it may be determined that lossless processing without any loss in the intermediate process has been achieved.

[0271]

[0272] Effects of the present invention

[0273] In order to verify the effectiveness of the present invention, the results of a performance evaluation on the improvement in encoding time are described below.

[0274] In evaluating the mesh encoding time, it was observed that the compression time gradually increased as the number of encoded displacement YUV frames increased. However, by compressing the multi-frame displacement YUV into a single frame by vertically stacking them by group of pictures (GOP) using the displacement data reconstruction method according to the present invention, it was confirmed that the reading and writing time could be eliminated, reducing the time required for the compression process by 47%. In addition, considering that a single experiment may not have enough statistical validity, 10,000 repeated experiments were performed under the same conditions to obtain more definitive experimental results. Specifically, displacement data compression of the original mesh and single frame compression after vertical stacking using the AI ​​method were each performed 10,000 times.

[0275] According to the encoder experiment results, for the original data consisting of 32 frames, a file of 290,032 bytes (2,175.240 kbps) was generated, and the average encoding time for 10,000 repeated experiments was measured to be 58,687,955 milliseconds, or approximately 5.86 seconds. In contrast, when the vertical stacking method was applied, the 10,000 repeated experiments resulted in a total execution time of 27,507.509 seconds, and the average encoding time was measured to be 2.75 seconds.

[0276] In the decoder experiment results, when the original data was repeated 10,000 times, the total execution time was measured as 643,991 milliseconds, and the average decryption time was measured as 0.064 seconds. In contrast, for vertically stacked data, the total execution time was measured as 701,111 milliseconds, and the average decryption time was measured as 0.071 seconds.

[0277] Through the 10,000 repetition experiments described above, the processing time of a single frame YUV using the vertical stacking method was confirmed to be 2.75 seconds compared to 5.86 seconds for encoding the original YUV frame, showing an improvement of about 47% in encoding speed. In contrast, the decoding process showed a slight delay of about 9%, with the required time increasing from 0.064 seconds to 0.071 seconds. This is interpreted as a result proving that the vertical stacking method proposed by the present invention can significantly improve the overall processing efficiency by greatly improving the encoding efficiency while imposing a limited burden on the decoding process.

[0278]

[0279] Although the present invention has been described with reference to drawings and embodiments, as already mentioned above, it does not mean that the scope of protection of the present invention is limited to the drawings or embodiments presented above, and it will be understood that a person skilled in the relevant technical field can modify and change the present invention in various ways within a scope that does not depart from the spirit and scope of the present invention described in the claims of the present invention patent.

Claims

1. A method for encoding a dynamic mesh by a computer device, A step of the computer device obtaining a plurality of frames in which displacement data is stored; The computer device generates a merged displacement frame by stacking the plurality of frames; and An encoding method, comprising: a step of the computer device performing dynamic mesh displacement encoding based on the merged displacement frame.

2. In paragraph 1, An encoding method, characterized in that a plurality of frames in which the displacement data is stored are acquired in units of GOP (Group of Pictures).

3. In paragraph 1, An encoding method, characterized in that a plurality of frames in which the displacement data is stored include displacement data stored in a YUV color format.

4. In paragraph 1, The step of performing dynamic mesh displacement encoding based on the above merged displacement frame is: An encoding method characterized by performing compression using an All Intra (AI) method for the above-mentioned merged displacement frame.

5. In paragraph 1, The step of creating a merged displacement frame by stacking the above multiple frames is: For each of the above multiple frames, a step of checking whether the resolution of the frame is a multiple of the encoding unit unit; and An encoding method, comprising: a step of modifying the resolution of the frame to correspond to the multiple when the resolution of the frame is not a multiple of the encoding unit unit.

6. In paragraph 5, An encoding method, characterized in that the above encoding unit is a unit based on a coding tree unit (CTU).

7. In paragraph 1, An encoding method, characterized in that the step of stacking the plurality of frames to generate a merged displacement frame comprises: stacking the plurality of frames vertically.

8. In paragraph 1, The step of creating a merged displacement frame by stacking the above multiple frames is: An encoding method characterized by: assigning a pixel index that progressively increases according to the order of the stacking to each pixel constituting the above-mentioned merged displacement frame.

9. In paragraph 1, An encoding method, further comprising: a step of the computer device verifying whether lossless processing is performed by comparing the plurality of frames with the merged displacement frame.

10. In paragraph 9, The step of verifying whether lossless processing is performed by comparing the plurality of frames and the merged displacement frame is as follows. A step of separating the above merge displacement frame; a step of matching the resolution of each of the separated displacement frames with the plurality of frames; and An encoding method, comprising: a step of calculating a signal-to-noise ratio (PSNR) by comparing the separated displacement frame with each of the plurality of frames.

11. A method for decoding displacement data of a dynamic mesh by a computer device, A step of the computer device obtaining a merged displacement frame through dynamic mesh displacement decoding; The step of the computer device separating the merged displacement frame to obtain a plurality of separated frames; and A decoding method, comprising: a step of the computer device restoring displacement data of the dynamic mesh from the plurality of separated frames.

12. In paragraph 11, A decoding method, characterized in that the above merged displacement frame includes a plurality of frames merged in units of GOP (Group of Pictures).

13. In paragraph 11, The step of restoring displacement data of the dynamic mesh from the above multiple separation frames is: A decoding method characterized by restoring displacement data stored in a YUV color format.

14. In paragraph 11, The step of separating the above merged displacement frame to obtain multiple separated frames is: a step of calculating the number of separated frames included in the above merged frame; and A decoding method, comprising: a step of dividing the merged displacement frame into a plurality of separated frames of the same size according to the calculated number; 15. In paragraph 14, The above-mentioned merge displacement frame comprises a plurality of vertically stacked frames, A decoding method, characterized in that the step of dividing the above-mentioned merged displacement frame into a plurality of separate frames of the same size comprises: obtaining the plurality of separate frames by releasing the vertical stacking.

16. In paragraph 11, A decoding method, characterized in that the plurality of separated frames have a resolution determined as a multiple of an encoding unit unit.

17. In paragraph 16, The step of separating the above merged displacement frame to obtain multiple separated frames is: A decoding method, comprising: a step of restoring the original resolution by identifying and removing a background area from the separated frame; 18. In paragraph 16, A decoding method, characterized in that the above coding unit unit is a unit based on a coding tree unit (CTU).

19. In paragraph 11, The step of restoring displacement data of the dynamic mesh from the above multiple separation frames is: A step of determining a temporal order for each displacement frame included in the plurality of separation frames; and A decoding method, comprising: a step of restoring the displacement data based on the plurality of separated frames according to the determined temporal order; 20. In a decoder device for decoding displacement data of a dynamic mesh, processor; and A memory storing instructions that control the operation of the above processor; A decoder device, characterized in that the processor is configured to obtain a merged displacement frame through dynamic mesh displacement decoding by executing a command stored in the memory, separate the merged displacement frame to obtain a plurality of separated frames, and restore displacement data of the dynamic mesh from the plurality of separated frames.

Citation Information

Patent Citations

  • Circulation Sterilization Humidifier

    KR1020250060493A

  • Base anchored models and inference for the compression and upsampling of video and multiview imagery

    US20200021824A1

  • Method and apparatus for point cloud compression

    US20200394822A1

  • Data compression for multidimensional time series data

    US20220207778A1

  • Network based image filtering for video coding

    WO2022235595A1