Encoding method and device, decoding method and device, encoding end and decoding end
By concurrently decoding geometry information during connection relationship reconstruction, the method addresses decoding delays in VDMC by allowing parallel processing, thus reducing latency.
Patent Information
- Application Number
- CN202410048774.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-15
AI Technical Summary
In the existing video processing technology, the video-based dynamic mesh compression method immediately performs multi-parallelogram prediction after determining the connection relationship mode, resulting in a large delay in the decoding process.
By obtaining the connection relationship code stream and geometric information code stream in the basic grid code stream at the decoding end, decoding the geometric information code stream during the connection relationship reconstruction process of vertices is used to realize parallel processing of connection relationships and geometric information.
Reduce the decoding delay and improve the efficiency of the decoding process.
Smart Images

Figure CN120321395A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of video processing, and particularly relates to an encoding and decoding method, apparatus, encoding end, and decoding end. Background Art
[0002] Regarding the processing part of connection relationships and geometric information, in the existing Video-based dynamic mesh coding (VDMC) solution, after the connection relationship mode is determined to be the C mode (or referred to as mode C), the geometric coordinates of vertices are immediately predicted using multiple parallelograms.
[0003] Determination of connection relationship mode (which can also be called connection mode): The Edgebreaker algorithm is used to traverse the mesh, and the connection relationship of the mesh is simply represented as a sequence of five modes (C, L, E, R, S).
[0004] The existing vertex prediction is carried out on the premise that all connection relationships of vertices are known. In order to be consistent with the prediction at the encoding end, the decoding end must complete the reconstruction of all connection relationships before starting the geometric prediction of vertices, and the decoding end needs to perform two traversals.
[0005] It can be seen from this that the existing decoding of connection relationships and geometric relationships requires multiple traversals of network vertices, resulting in a relatively large delay in the decoding process. Summary of the Invention
[0006] Embodiments of this application provide an encoding and decoding method, apparatus, encoding end, and decoding end to reduce the decoding delay.
[0007] In a first aspect, a decoding method is provided. The method includes:
[0008] The decoding end obtains a basic mesh bitstream, where the basic mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the basic mesh belongs;
[0009] During the process of reconstructing the connection relationship of vertices according to the connection relationship sequence, the decoding end decodes the geometric information bitstream.
[0010] In a second aspect, a decoding apparatus is provided, including:
[0011] An obtaining module, configured to obtain a basic mesh bitstream, where the basic mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the basic mesh belongs;
[0012] A decoding module, configured to decode the geometric information bitstream during the process of reconstructing the connection relationship of vertices according to the connection relationship sequence.
[0013] In a third aspect, an encoding method is provided, which includes:
[0014] The encoding end determines the connection mode of the basic mesh. When it is determined that the connection mode of the triangle corresponding to the second target vertex is the first mode, at least one parallelogram is used to predict the second target vertex to obtain a prediction result, and the connection modes of the triangles corresponding to the vertices other than the second target vertex in the at least one parallelogram have been determined;
[0015] The encoding end encodes based on the prediction result to obtain a geometric information bitstream and encodes the connection relationship sequence corresponding to the basic mesh to obtain a connection relationship bitstream;
[0016] The encoding end sends a basic mesh bitstream to the decoding end. The basic mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode of the triangle corresponding to the vertex in the basic mesh.
[0017] In a fourth aspect, an encoding device is provided, including:
[0018] A prediction module, configured to determine the connection mode of the basic mesh. When it is determined that the connection mode of the triangle corresponding to the second target vertex is the first mode, at least one parallelogram is used to predict the second target vertex to obtain a prediction result, and the connection modes of the triangles corresponding to the vertices other than the second target vertex in the at least one parallelogram have been determined;
[0019] An encoding module, configured to encode based on the prediction result to obtain a geometric information bitstream and encode the connection relationship sequence corresponding to the basic mesh to obtain a connection relationship bitstream;
[0020] A first sending module, configured to send a basic mesh bitstream to the decoding end. The basic mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode of the triangle corresponding to the vertex in the basic mesh.
[0021] In a fifth aspect, a decoding end is provided, including a processor and a memory. The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0022] In a sixth aspect, a decoding end is provided, including a processor and a communication interface. The processor is configured to obtain a base mesh bitstream, where the base mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the base mesh belongs.
[0023] During the process of reconstructing the connection relationship of the vertex according to the connection relationship sequence, the geometric information bitstream is decoded.
[0024] In a seventh aspect, an encoding end is provided, including a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the third aspect are implemented.
[0025] In an eighth aspect, an encoding end is provided, including a processor and a communication interface. The processor is configured to determine the connection mode of the base mesh. When it is determined that the connection mode corresponding to the triangle to which the second target vertex belongs is the first mode, at least one parallelogram is used to predict the second target vertex to obtain a prediction result. For the vertices of the at least one parallelogram other than the second target vertex, the connection modes of the triangles to which they belong have been determined.
[0026] Encoding is performed based on the prediction result to obtain a geometric information bitstream and encoding is performed on the connection relationship sequence corresponding to the base mesh to obtain a connection relationship bitstream.
[0027] A base mesh bitstream is sent to the decoding end. The base mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the base mesh belongs.
[0028] In a ninth aspect, a coding and decoding system is provided, including: an encoding end and a decoding end. The encoding end can be used to execute the steps of the method described in the third aspect, and the decoding end can be used to execute the steps of the method described in the first aspect.
[0029] In a tenth aspect, a readable storage medium is provided. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of the method described in the first aspect or the third aspect are implemented.
[0030] In an eleventh aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run a program or instruction to implement the steps of the method described in the first aspect or the third aspect.
[0031] In a twelfth aspect, a computer program / program product is provided. The computer program / program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method described in the first aspect or the third aspect.
[0032] In the embodiments of the present application, during the process of reconstructing the connection relationship of vertices according to the connection relationship sequence, the geometric information bitstream is decoded to achieve the reconstruction of geometric information. When reconstructing the connection relationship in the embodiments of the present application, the decoding of the geometric information bitstream is performed simultaneously to achieve the reconstruction of geometric information, thereby realizing the parallel processing of the connection relationship and geometric coordinates, and thus being able to shorten the decoding time and achieve the purpose of reducing the decoding delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a schematic diagram of the encoding and decoding system provided by the embodiments of the present application;
[0034] Figure 2 is a schematic structural diagram of the encoder provided by the embodiments of the present application;
[0035] Figure 3 is a schematic structural diagram of the decoder provided by the embodiments of the present application;
[0036] Figure 4 is a schematic flowchart of the encoding method provided by the embodiments of the present application;
[0037] Figure 5 is a schematic diagram of Encoding Method 1 at the encoding end;
[0038] Figure 6 is a schematic diagram of the single parallelogram prediction method at the encoding end;
[0039] Figure 7 is a schematic diagram of Encoding Method 2 at the encoding end;
[0040] Figure 8 is a schematic diagram of the multi - parallelogram prediction method at the encoding end;
[0041] Figure 9 is a schematic diagram of the overall encoding process;
[0042] Figure 10 is a schematic diagram of mesh simplification;
[0043] Figure 11 is a schematic diagram of the geometric displacement vector calculation method;
[0044] Figure 12 is a schematic diagram of subdivision;
[0045] Figure 13 is a schematic flowchart of the decoding method provided by the embodiments of the present application;
[0046] Figure 14 It is a schematic diagram of the first decoding method at the decoding end;
[0047] Figure 15 It is a schematic diagram of the single parallelogram prediction method at the decoding end;
[0048] Figure 16 It is a schematic diagram of the second decoding method at the decoding end;
[0049] Figure 17 It is a schematic diagram of the multi - parallelogram prediction method at the decoding end;
[0050] Figure 18 It is a schematic diagram of the overall decoding process;
[0051] Figure 19 It is a schematic diagram of the modules of the encoding device according to the embodiments of the present application;
[0052] Figure 20 It is a schematic diagram of the structure of the encoding end according to the embodiments of the present application;
[0053] Figure 21 It is a schematic diagram of the modules of the decoding device according to the embodiments of the present application;
[0054] Figure 22 It is a schematic diagram of the structure of the electronic device according to the embodiments of the present application. Detailed implementation manners
[0055] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0056] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same kind, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "or" in the present application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0057] Figure 1It is a schematic diagram of the codec system provided by the embodiments of the present application. The technical solution of the embodiments of the present application relates to encoding and decoding (CODEC) (including encoding or decoding) video data. Among them, the video data includes original unencoded video, encoded video, decoded (e.g., reconstructed) video, or syntax elements, etc.
[0058] As Figure 1 shown, the codec system includes a source device 100, and the source device 100 provides encoded video data to be decoded and displayed by a destination device 110. Specifically, the source device 100 provides video data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile phone, a wearable device (such as a smart watch or a wearable camera), a television, a camera, a display device, a vehicle-mounted device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video game console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an airplane, a robot, a satellite, etc.
[0059] In Figure 1 the example, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of a video encoding device, and the destination device 110 represents an example of a video decoding device. In other examples, the source device 100 and the destination device 110 may not include Figure 1 some components in Figure 1 or may also include other components outside
[0060] Although Figure 1 the source device 100 and the destination device 110 are depicted as separate devices, in some examples, the two may also be integrated into one device. In such embodiments, the functions corresponding to the source device 100 and the functions corresponding to the destination device 110 may be implemented using the same hardware or software, or using separate hardware or software, or any combination thereof.
[0061] In some examples, the source device 100 and the destination device 110 can perform unidirectional video transmission or bidirectional video transmission. If it is bidirectional video transmission, the source device 100 and the destination device 110 can operate in a substantially symmetric manner, that is, each of the source device 100 and the destination device 110 includes an encoder and a decoder.
[0062] The data source 101 represents the source of video data (i.e., the original, unencoded video data) and provides consecutive pictures containing video data to the encoder 200, and the encoder 200 encodes the data of the pictures. The data source 101 of the source device 100 can include a video capture device (such as a video camera), a video archive containing previously captured original video, or a video feed interface for receiving video from a video content provider. Alternatively, the data source 101 can generate computer graphics-based data as the source video, or combine real-time video, archived video, and computer-generated video. In these cases, the encoder 200 encodes the captured, pre-captured, or computer-generated video data. The encoder 200 can rearrange the pictures from the received order (sometimes referred to as the "display order") into the encoding order. The encoder 200 can generate a bitstream including the encoded video data. The source device 100 can then output the encoded video data onto the communication medium 120 via the output interface 104 for reception or retrieval by, for example, the input interface 111 of the destination device 110.
[0063] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general-purpose memories. In some examples, the memory 102 can store the original video data from the data source 101, and the memory 113 can store the decoded video data from the decoder 300. Additionally or alternatively, the memories 102, 113 can store software instructions executable by, for example, the encoder 200 and the decoder 300. Although the memory 102 and the memory 113 are shown separately from the encoder 200 and the decoder 300 in this example, it should be understood that the encoder 200 and the decoder 300 can also include internal memories for functionally similar or equivalent purposes. If the encoder 200 and the decoder 300 are deployed on the same hardware device, the memory 102 and the memory 113 can be the same memory. Furthermore, the memories 102, 113 can store, for example, the encoded video data output from the encoder 200 and input to the decoder 300. In some examples, portions of the memories 102, 113 can be allocated as one or more video buffers, for example, for storing original, decoded, or encoded video data.
[0064] In some examples, the source device 100 may output the encoded data from the output interface 104 to the memory 113. Similarly, the destination device 110 may access the encoded data from the memory 113 via the input interface 111. The memory 113 or the memory 102 may include any one of various distributed or local access data storage media, such as hard drives, Blu-ray discs, Digital Versatile Discs (DVDs), Compact Disc Read-Only Memories (CD-ROMs), flash memories, volatile or non-volatile memories, or any other suitable digital storage media for storing encoded video data.
[0065] The output interface 104 may include any type of medium or device capable of sending the encoded video data from the source device 100 to the destination device 110. For example, the output interface 104 may include a transmitter or transceiver, such as an antenna, configured to send the encoded video data from the source device 100 directly in real time to the destination device 110. The encoded video data may be modulated according to the communication standard of a wireless communication protocol and sent to the destination device 110.
[0066] The communication medium 120 may include transient media, such as wireless broadcasts or wired network transmissions. For example, the communication medium 120 may include the radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 120 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard drive, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0067] In some embodiments, communication medium 120 may include a router, a switch, a base station, or any other device that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) may receive the encoded video from source device 100 and provide the encoded video data to destination device 110, e.g., via network transmission to destination device 110. The server may include, for example, a web server (for a website), a server configured to provide file transfer protocol services such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol, a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached storage (NAS) device, etc. The server may implement one or more HTTP streaming protocols such as MPEG Media Transport (MMT) protocol, Dynamic Adaptive Streaming over HTTP (DASH) protocol, HTTP Live Streaming (HLS) protocol, or Real Time Streaming Protocol (RTSP), etc.
[0068] Destination device 110 may access the encoded video data from the server, e.g., via a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., Digital subscriber line (DSL), cable modem, etc.) for accessing the encoded video data stored on the server.
[0069] The output interface 104 and the input interface 111 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 standard or the IEEE 802.15 standard (e.g., ZigBeeTM), the Bluetooth standard, etc., or other physical components. In an example where the output interface 104 and the input interface 111 include wireless components, the output interface 104 and the input interface 111 may be configured to transmit data, such as encoded video data, according to WIFI, Ethernet, a cellular network (such as 4G, LTE (Long Term Evolution), Advanced LTE, 5G, 6G, etc.).
[0070] The technology provided by the embodiments of the present application can be applied to support video encoding and decoding in one or more of the following multimedia applications: video conferencing, over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission, digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0071] The input interface 111 of the destination device 110 receives the encoded video bitstream from the communication medium 120. The encoded video bitstream may include syntax elements and encoded data units (e.g., sequences, groups of pictures, pictures, slices, blocks, etc.), where the syntax elements are used to decode the encoded data units to obtain the decoded video data. The display device 114 displays the decoded video data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid-crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0072] The encoder 200 and the decoder 300 may be implemented as one or more of various processing circuits, which may include a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. When the technology is implemented in whole or in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided by the embodiments of the present application.
[0073] The encoder 200 and the decoder 300 may perform processing based on the following video coding and decoding standards: H.263, H.264, H.265 (also known as High Efficiency Video Coding, HEVC), H.266 (also known as Versatile Video Coding, VVC), Moving Picture Experts Group 2 (MPEG-2), MPEG-4, VP8, VP9, Alliance for Open Media Video 1 (AV1), Audio Video Coding Standard 1 (AVS1), AVS2, AVS3, or the next-generation video standard protocol. The embodiments of the present application do not make specific limitations.
[0074] Generally, the encoder 200 and the decoder 300 may perform block-based coding and decoding of pictures. The term "block" generally refers to a structure including data to be processed (e.g., encoded, decoded, or otherwise used during the encoding or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance or chrominance data. For example, the encoder 200 and the decoder 300 may perform coding and decoding on video data represented in the YUV format.
[0075] See Figure 2 , which is a schematic structural diagram of the encoder 200 provided by the embodiments of the present application. The encoder 200 may be the Figure 1 encoder 200 in Figure 2 . In the example of
[0076] The memory 201 may store video data to be encoded. For example, the encoder 200 may receive and store video data from the data source 101 shown in Figure 1 . In some examples, the memory 201 may be on the same chip as other components of the encoder 200 (as shown in Figure 2 ), or may be independent of the chip where these components are located.
[0077] The coding parameter determination unit 210 includes a mode selection unit 211, an inter-frame prediction unit 212, and an intra-frame prediction unit 213. The inter-frame prediction unit 212 is configured to obtain a first prediction block of a current block by using an inter-frame prediction mode, the intra-frame prediction unit 213 is configured to obtain a second prediction block of the current block by using an intra-frame prediction mode, and the mode selection unit 211 is configured to obtain a target prediction block based on the first prediction block and the second prediction block, and determine a final prediction mode. In addition, the coding parameter determination unit 210 may further include other functional units, such as a functional unit for determining a partitioning manner of a coding unit (CU), a functional unit for determining a transform type of residual data of the CU, or a functional unit for determining quantization parameters of the residual data of the CU, etc.
[0078] For the convenience of description and understanding, in the embodiments of the present application, the CU to be processed in the current image is referred to as the current CU, and the image block to be processed in the current CU is referred to as the current block or the image block to be processed. For example, in encoding, it refers to the block currently being encoded; in decoding, it refers to the block currently being decoded.
[0079] The inter-frame prediction unit 212 may include a motion estimation unit and a motion compensation unit. For the inter-frame prediction of the current block, the motion estimation unit may perform a motion search to identify one or more matching reference blocks in one or more reference pictures (for example, one or more previously encoded and decoded pictures stored in the DPB 209).
[0080] The motion estimation unit may form one or more motion vectors (MVs) of the positions of the reference blocks in the reference pictures relative to the position of the current block in the current picture. The motion compensation unit may obtain a predicted value with the accuracy indicated by the motion vector through interpolation.
[0081] The coding parameter determination unit 210 may provide the target prediction block to the residual generation unit 202. The residual generation unit 202 receives the original unencoded video data of the current block from the memory 201, and calculates the residual between the current block and the target prediction block to obtain a residual block. In some examples, the function of the residual generation unit 202 may be implemented by using one or more subtractor circuits that perform binary subtraction.
[0082] As an example, the coding parameter determination unit 210 may provide syntax elements representing coding parameters to the entropy coding unit 220 for encoding. The coding parameters include one or more of a partitioning manner of the CU, a final prediction mode, a transform type of the residual data of the CU, or quantization parameters of the residual data of the CU, etc.
[0083] The transform processing unit 203 performs a transform on the residual blocks output by the residual generation unit 202 to obtain a transform coefficient block. The transform may include, for example, a Discrete Cosine Transform (DCT), an integer transform, a directional transform, or a Karhunen-Loeve transform. In some examples, the encoder 200 may not include the transform processing unit 203.
[0084] The quantization unit 204 may quantize the transform coefficients in the transform coefficient block according to the quantization parameter (QP) value associated with the current block to produce a quantized transform coefficient block.
[0085] The dequantization unit 205 and the inverse transform processing unit 206 may perform dequantization and inverse transform on the transform coefficient block respectively to obtain a reconstructed residual block. The reconstruction unit 207 may generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the target prediction block generated by the coding parameter determination unit 210.
[0086] The filter unit 208 may perform one or more filter operations on the reconstructed block. For example, the filter unit 208 may be a deblocking filter (DBF), an adaptive loop filter (ALF), a sample adaptive offset (SAO) filter, etc. In some examples, the encoder 200 may not include the filter unit 208.
[0087] The encoder 200 stores the reconstructed picture obtained from the reconstructed block in the DPB 209. For example, in an example where the operation of the filter unit 208 is not required, the reconstruction unit 207 may store the reconstructed block in the DPB 209. In an example where the operation of the filter unit 208 is required, the filter unit 208 may store the filtered reconstructed block in the DPB 209. The inter prediction unit 212 obtains the reconstructed picture from the DPB 209 to perform inter prediction on the blocks of the subsequent pictures to be encoded. In some examples, the DPB 209 may be replaced by other types of memories.
[0088] The entropy coding unit 220 may perform entropy coding on the syntax elements of other components in the encoder 200 and output the encoded video data. For example, the entropy coding unit 220 may perform entropy coding on the quantized transform coefficient block from the quantization unit 204. As another example, the entropy coding unit 220 may perform entropy coding on the syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from the coding parameter determination unit 210.
[0089] It can be understood that Figure 2 the composition of the encoder 200 described above is only illustrative and does not constitute a limitation on the embodiments of the present application.
[0090] Figure 3 FIG. 6 is a schematic structural diagram of a decoder 300 provided by an embodiment of the present application. The decoder 300 may be Figure 1 the decoder 300 described above. In Figure 3 this example, the decoder 300 includes a coded picture buffer (CPB) 301, an entropy decoding unit 302, a prediction processing unit 310, an inverse quantization unit 303, an inverse transform processing unit 304, a reconstruction unit 305, a filter unit 306, and a DPB 307.
[0091] The entropy decoding unit 302 can receive the encoded video data from the CPB 301 and perform entropy decoding on the video data to obtain syntax elements, where the syntax elements indicate encoding parameters, and the encoding parameters include one or more of the partitioning manner of the CU, the final prediction mode, the transform type of the residual data of the CU, or the quantization parameter of the residual data of the CU.
[0092] When the syntax element includes the final prediction mode, the prediction processing unit 310 obtains the final prediction mode. If the final prediction mode is an inter prediction mode, the prediction block of the current CU can be obtained through the inter prediction unit 311 of the prediction processing unit 310; if the final prediction mode is an intra prediction mode, the prediction block of the current CU can be obtained through the intra prediction unit 312 of the prediction processing unit 310. In some examples, the prediction processing unit 310 may further include a unit for performing a prediction function according to other prediction modes.
[0093] The CPB 301 can obtain and store the encoded video data from the communication medium 120 as shown in Figure 1 FIG. 21. The DPB 307 is used to store the decoded pictures. Optionally, the CPB 301 and the DPB 307 may also be replaced with other types of memories, and the present application does not make specific limitations. In some examples, the CPB 301 may be on the same chip as other components of the decoder 300 (as shown in the figure), or may be independent of the chip where these components are located.
[0094] The decoder 300 can perform reconstruction operations on each block separately. The entropy decoding unit 302 can perform entropy decoding on the syntax elements of the quantized transform coefficients and the transform information (such as QP or transform mode indication) to obtain the quantized transform coefficients. The quantized transform coefficients are dequantized by the dequantization unit 303 to obtain a transform coefficient block including transform coefficients. The transform coefficient block is inverse-transformed by the inverse transform processing unit 304 to generate a residual block corresponding to the current block, and this inverse transform is the reverse operation of the above-mentioned transform.
[0095] The reconstruction unit 305 can reconstruct the current block based on the prediction block and the residual block. For example, the reconstruction unit 305 can add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.
[0096] The filter unit 306 can perform one or more filter operations on the reconstructed block. For example, the type of the filter unit 306 can refer to the type of the filter unit 208, which will not be elaborated here. In some examples, the operation of the filter unit 306 can be skipped.
[0097] The decoder 300 can store the reconstructed picture obtained from the reconstructed block in the DPB 307. For example, in an example where the operation of the filter unit 306 is not performed, the reconstruction unit 305 can store the reconstructed block in the DPB 307. In an example where the operation of the filter unit 306 is performed, the filter unit 306 can store the filtered reconstructed block in the DPB 307. The decoder 300 can output the decoded picture (such as decoded video) from the DPB 307 for subsequent presentation to a display device (such as Figure 1 the display device 114).
[0098] The encoding and decoding methods, apparatuses, encoding end, and decoding end provided in the embodiments of the present application will be introduced below in conjunction with the accompanying drawings. The encoding method provided in the embodiments of the present application can be executed by the encoding end, such as Figure 1 or Figure 2 the encoder 200 shown. The decoding method provided in the embodiments of the present application can be executed by the decoding end, such as Figure 1 or Figure 3 the decoder 300 described. Among them, the encoding end and the decoding end can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding end device or a video encoding device, and the decoding end can be referred to as a decoding end device or a video decoding device.
[0099] The related technologies related to the embodiments of the present application will be described as follows first.
[0100] In recent years, with the rapid development of multimedia technology, relevant research results have been quickly industrialized and become an essential part of people's lives. 3D models have become a new generation of digital media following audio, images, and videos. 3D meshes are a commonly used representation of 3D models. Compared with traditional multimedia such as images and videos, 3D mesh models have stronger interactivity and realism, making them increasingly widely used in various fields such as commerce, manufacturing, construction, education, medicine, entertainment, art, and the military.
[0101] Although there are many representation methods for 3D meshes currently, triangular meshes are still the most common representation method. A 3D mesh can be regarded as composed of three basic elements: vertices, edges, and faces. Vertices are the most basic elements in the mesh, which define positions in a three-dimensional space. Edges are line segments connecting two vertices in the mesh. Faces can be regarded as polygons formed by closed paths of edges. For triangular meshes, each face is a triangle.
[0102] The information contained in the mesh is usually divided into three categories: geometric information, connectivity information, and attribute information. Geometric information is the position of each vertex of the mesh in three-dimensional space. Connectivity information describes the association relationships between elements in the mesh, that is, the connection relationships between vertices. Attribute information is optional, and it can associate attributes to the corresponding mesh elements (such as vertex colors, normal vectors, etc. can be associated with mesh vertices). Mesh parameterization can also be used to map the mesh from three-dimensional space to a two-dimensional planar region. This mapping relationship is usually described by a set of parametric coordinates, called UV coordinates or texture coordinates, and is associated with mesh vertices. This two-dimensional mapping can be used to represent high-resolution attribute information such as textures and normal vectors.
[0103] In almost all application fields using 3D meshes (such as computational simulation, entertainment, medical imaging, digital cultural relics, computer design, e-commerce, etc.), as people's requirements for the visual effects of 3D mesh models are getting higher and higher, the models are becoming increasingly complex and the accuracy of the models is also getting higher. Therefore, the amount of data required to represent 3D meshes is correspondingly increasing. The above problems have led to the processing, visualization, transmission, and storage of 3D meshes becoming increasingly complex. 3D mesh compression can be regarded as a way to solve the above problems. It reduces the size of model data, which is beneficial to the processing, storage, and transmission of 3D meshes. Therefore, it is necessary to propose an efficient and general 3D mesh compression algorithm.
[0104] Recently, the international standardization organization MPEG in the field of audio and video coding and compression has started to develop a compression standard for three-dimensional meshes, VDMC (Video-based dynamic mesh coding). This standard is specified based on the existing V3C (Visual Volumetric Video-based Coding) standard. The V3C standard provides a general method for compressing three-dimensional models, which can be presented in the form of point clouds, meshes, or panoramic videos, etc. Making the compression method of the three-dimensional mesh model compatible with this standard helps to promote the method and its applicability. Therefore, it is of great significance to optimize the three-dimensional mesh encoding and decoding method in VDMC and combine the optimized method with the V3C standard. A possible optimization method is to optimize the intra-frame base mesh encoding and decoding. In the existing framework, the information of the base mesh is divided into connection relationships, geometric information, and attribute information for processing. For the geometric information encoding and decoding of the base mesh in three-dimensional mesh encoding and decoding, a geometric information encoding and decoding method that only depends on partial connection relationships is proposed, which can support a parallel scheme for connection relationship and geometric information encoding and decoding. The encoding and decoding of connection relationships and geometric information can be completed only through one traversal, which can reduce the time complexity.
[0105] V3C standard
[0106] The V3C standard provides a method for encoding and decoding various three-dimensional media through video or image coding techniques. Specifically, before encoding, the three-dimensional media content is converted from a three-dimensional representation to multiple two-dimensional representations (called V3C components) through projection and other means, and then the existing video or image coding techniques are used to encode the two-dimensional representations. The V3C components mainly include occupancy components, geometric components, and attribute components. The occupancy component can represent which regions in the two-dimensional representation are associated with the data of the three-dimensional representation; the geometric component represents the information related to the position of the three-dimensional data in space, and the attribute component can provide the attribute information corresponding to the vertices, such as materials, textures, etc. In addition, the components also contain the information on how to reconstruct the three-dimensional model through these components, which is called atlas information.
[0107] The atlas information is used to associate all components, and the additional information for reconstructing from two dimensions back to three dimensions is also included in the atlas component. The atlas consists of multiple basic units, and the basic units are called patches. Each patch represents a region in the available two-dimensional components and contains the information required to project that region back into three-dimensional space.
[0108] VDMC
[0109] VDMC is a standard for compressing 3D meshes developed by MPEG. Its main idea is to compress 3D meshes by leveraging the existing V3C standard. Since 3D meshes have connectivity information that needs to be encoded, the specific encoding process is slightly different from that of V3C, and it is necessary to extend the syntax semantics and decoding operations of the V3C standard decoder to support the decoding and reconstruction of 3D meshes.
[0110] At the encoding end, for the input mesh, first, it is simplified through a simplification module, then new texture coordinates are generated for the mesh through mesh parameterization. Subsequently, the parameterized mesh is subdivided and deformed, that is, new vertices are inserted into the mesh according to a specific subdivision method and the distance from the vertices of the subdivided mesh to the nearest neighbor points of the input mesh is calculated, which is called displacement information. Subsequently, the parameterized mesh is adjusted according to the displacement information, that is, the vertex positions of the mesh before subdivision and deformation, and the adjusted mesh is called the base mesh, which is sent to the base mesh encoding module for compression. When encoding the base mesh, it is divided into 3 types of sub-meshes for independent encoding. After the base mesh is encoded, it is reconstructed, and then the order of the displacements is adjusted according to the vertex order of the reconstructed base mesh. Subsequently, the vertex displacement information with the adjusted order is first subjected to wavelet transform, the transformed coefficients are quantized, and then the quantized coefficients are arranged into a 2D image according to a specific scan order, and a video encoder is used to encode the 2D image. Then, the reconstructed displacement information is applied to the subdivided base mesh to obtain the reconstructed subdivided and deformed mesh. The mesh and the original input mesh and its corresponding texture map are input into the corresponding texture map conversion module to obtain the texture map corresponding to the reconstructed mesh, and this texture map is also encoded using a video encoder. For the parameters used in the encoding process, such as the type of video encoder used, the type of mesh encoder, transformation parameters, quantization parameters, etc., they are passed to the decoding end through auxiliary information.
[0111] At the decoding end, for the received bitstream, the decoding end first demultiplexes each part of the bitstream to obtain the base mesh bitstream, the displacement video bitstream, the texture map video bitstream, and the auxiliary information bitstream respectively. For the base mesh bitstream, the base mesh is decoded using the mesh decoder indicated by the auxiliary information. The displacement video bitstream and the texture map video bitstream are decoded through the video decoder. For the displacement part, after video decoding, the displacement needs to be extracted from the image through the displacement decoding module and subjected to steps such as inverse quantization and inverse transformation, and then it is applied to the subdivided base mesh to obtain the deformed mesh reconstructed at the decoding end. The texture map after decoding is the texture map corresponding to the reconstructed deformed mesh. The subsequent application or rendering module processes the reconstructed deformed mesh and the decoded texture map as inputs.
[0112] The VDMC generally encodes the base mesh in two modes: intra-frame mode and inter-frame mode. In the intra-frame mode, the processing of the base mesh at the encoding end mainly includes: for the input base mesh, through the connection relation encoding module, geometric information encoding module, and attribute information encoding module, it is encoded into three bitstreams: connection information bitstream, geometric information bitstream, and attribute information bitstream, and then mixed and output. Among them, the encoding of geometric information and attribute information both refer to the traversal order of the connection relation.
[0113] Corresponding to the encoding end, in the intra-frame mode, the decompression processing of the base mesh at the decoding end mainly includes: for the input bitstream, it is first de-streamed into three bitstream information: connection information bitstream, geometric information bitstream, and attribute information bitstream, and then respectively sent to the connection relation decoding module, geometric information decoding module, and attribute information decoding module to decode and obtain the connection relation, geometric information, and attribute information of the base mesh, and finally reconstruct the base mesh. Similar to the encoding end, the decoding of geometric information and attribute information at the decoding end also needs to refer to the traversal order of their connection relations.
[0114] At the encoding end, for the connection relation, the Edgebreaker algorithm is used to traverse the 3D mesh and represent its connection relation in five modes (C, L, E, R, S) to achieve efficient encoding of the connection relation. For geometric information, referring to the traversal order of the connection relation, the vertex geometric coordinates traversed are predicted, and finally the residuals between the predicted coordinates and the real coordinates are encoded.
[0115] At the decoding end, for the connection relation, the mesh connection relation is reconstructed by traversing the connection relation sequence transmitted from the encoding end. For geometric information, the mesh is traversed according to the fully reconstructed connection relation, and the same prediction scheme as the encoding end is used to predict the geometric positions of the vertices, and finally added to the geometric coordinate residuals transmitted from the encoding end to obtain the real geometric coordinates of the vertices.
[0116] Next, in combination with the accompanying drawings, the encoding and decoding methods, devices, encoding end, and decoding end provided by the embodiments of the present application will be described in detail through some embodiments and their application scenarios.
[0117] As Figure 4 shown, the embodiments of the present application provide an encoding method, including:
[0118] Step 401, the encoding end determines the connection mode of the base mesh. When it is determined that the connection mode of the triangle corresponding to the second target vertex is the first mode, at least one parallelogram is used to predict the second target vertex to obtain a prediction result, and the triangles corresponding to the vertices other than the second target vertex in the at least one parallelogram have all been determined for the connection mode;
[0119] It should be noted that when performing vertex prediction in the embodiments of the present application, the triangles to which the vertices constituting the parallelogram belong have all been determined for the connection mode, so that triangles to which vertices without connection mode determination are not used for parallelogram construction; correspondingly, the decoding end will not use vertices without connection relationship reconstruction for geometric information reconstruction, so that the decoding end can decode geometric information during the connection relationship reconstruction process; it should be clear here that the improvement of the encoding end is an adaptive adjustment made to keep consistent with the decoding end.
[0120] It should be noted that since the connection mode is determined based on triangles, the triangles to which the vertices mentioned in the embodiments of the present application belong refer to the triangles on which the connection mode determination is based. The triangle to which the second target vertex belongs refers to the triangle formed by the two vertices where the active edge is located and the second target vertex. This active edge (which can also be called a traversal edge) can be understood as the edge used for connection mode determination.
[0121] Step 402, the encoding end encodes based on the prediction result to obtain a geometric information bitstream and encodes the connection relationship sequence corresponding to the base mesh to obtain a connection relationship bitstream;
[0122] Step 403, the encoding end sends a base mesh bitstream to the decoding end. The base mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangles to which the vertices in the base mesh belong.
[0123] It should be noted that the connection relationship sequence includes the connection modes corresponding to all the triangles to which the vertices in the base mesh belong. This connection mode is determined when traversing the base mesh using the Edgebreaker algorithm, and can also be called a connection relationship mode.
[0124] It should be noted that in the embodiments of the present application, in order to cooperate with the implementation of the decoding end, the implementation of the encoding end has also been improved accordingly, so as to construct a parallelogram based on the triangles to which the vertices other than the second target vertex that have all been determined for the connection mode belong, so as to cooperate with the implementation method of the decoding end and achieve the purpose of reducing the latency of the decoding end.
[0125] Optionally, this base network can be understood as being obtained after simplifying, parameterizing the mesh, and performing subdivision deformation on the input mesh. In the embodiments of the present application, the main focus is on the process of compressing the base mesh. By shortening the time-consuming for compressing the base mesh, the encoding latency can be reduced.
[0126] It should be noted that the encoding method in the embodiments of the present application can be understood as a method of parallel encoding of connection relationships and geometric information.
[0127] It should be noted that the connection modes involved in the embodiments of the present application mainly include mode C, mode L, mode R, mode S, and mode E.
[0128] Among them, for mode C, the triangles connected by the other two edges except the traversed edge are not traversed; for mode L, the triangle connected by the left edge except the traversed edge has been traversed; for mode R, the triangle connected by the right edge except the traversed edge has been traversed; for mode S, the triangles connected by the incoming edge except the traversed edge are not traversed, but the vertices of the triangles have been traversed; for mode E, the triangles connected by the other two edges except the traversed edge are all traversed.
[0129] Optionally, the first mode mentioned in the embodiments of the present application may be mode C.
[0130] Optionally, in the embodiments of the present application, the unencoded vertices corresponding to the triangles corresponding to mode C are encoded. That is to say, if the parallel encoding of connection relationship and geometric information is required, the encoding end performs the parallel encoding of connection relationship and geometric information for the unencoded vertices corresponding to the triangles corresponding to mode C in the basic grid.
[0131] Furthermore, if the parallel encoding of connection relationship and geometric information is performed, when encoding the second target vertex in the embodiments of the present application, there are two different implementation manners. One implementation manner is that as long as it is determined that the connection mode corresponding to the triangle to which the second target vertex belongs is mode C, the second target vertex is immediately encoded. This manner can be understood as a non-offset encoding manner (i.e., no offset encoding). Non-offset encoding can be understood as: without waiting for the determination of the connection mode corresponding to the triangle to which the vertex after the second target vertex belongs. Another implementation manner is the offset encoding manner (i.e., offset encoding). Offset encoding can be understood as: waiting for the determination of the connection mode corresponding to the triangle to which the vertex after the second target vertex belongs. As for when to encode the second target vertex, optionally, in one implementation manner, the specific implementation of predicting the second target vertex with at least one parallelogram and obtaining the prediction result includes:
[0132] The encoding end predicts the second target vertex with at least one parallelogram and obtains the prediction result under the condition of meeting the second target condition;
[0133] Among them, the second target condition includes one of the following:
[0134] A11. The number of connection modes determined after the second target vertex reaches the maximum waiting number;
[0135] This situation can be understood as that after the connection mode corresponding to the triangle to which the second target vertex belongs is determined, N connection modes are determined, where N is equal to the maximum waiting number.
[0136] It should be noted that the maximum waiting number can be understood as the maximum waiting connection mode number or the maximum misalignment number; for example, if the misalignment number is 2, it means that after the connection modes of two vertices after the second target vertex are determined, the second target vertex is encoded.
[0137] A12. There is a first mode within the maximum waiting number of connection mode determinations after the second target vertex;
[0138] This situation can be understood as: there is a next first mode within the maximum waiting number of connection mode determinations after the second target vertex; it can also be understood that after the connection mode corresponding to the triangle to which the second target vertex belongs is determined, a next first mode is encountered during the subsequent connection mode determination process, and the next first mode appears before the maximum waiting number is reached.
[0139] Optionally, this situation can be understood as that when the next mode C is encountered, the second target vertex needs to be encoded.
[0140] A13. Reaching the last second mode in the connected component to which the second target vertex belongs;
[0141] Optionally, the second mode mentioned in the embodiments of the present application may be mode E.
[0142] It should be noted that each basic grid may include multiple traversals. Each traversal edge corresponds to a connected component, and each connected component ends with mode E, that is, the last mode of each connected component is mode E.
[0143] It should be noted that as long as one of the above A11 - A13 is satisfied, the encoding end needs to encode the second target vertex. For example, if the condition of A12 is reached first, the encoding end encodes the second target vertex; if the condition of A13 is reached first, the encoding end encodes the second target vertex; if the conditions of A12 or A13 are not reached, but the condition of A11 is reached, the encoding end encodes the second target vertex.
[0144] Optionally, the specific implementation of the encoding end obtaining the geometric information bitstream based on the prediction result includes:
[0145] Step b1. The encoding end obtains the residual value corresponding to the second target vertex according to the prediction result and the reference coordinates of the second target vertex;
[0146] It should be noted that the reference coordinate refers to the true coordinate of the second target vertex; usually, the residual value can be obtained by subtracting the reference coordinate from the prediction result.
[0147] Step b2: The encoding end encodes the second target vertex according to the residual value.
[0148] It should be noted that in the above implementation method, without dislocation encoding, as long as the encoding end determines that the connection mode corresponding to the triangle to which the second target vertex belongs is mode C, the encoding end immediately predicts the vertex coordinates based on a parallelogram, and then obtains the residual value; in the case of dislocation encoding, the encoding end needs to determine whether the second target condition is met. If the second target condition is met, after waiting for several vertex connection mode determinations, the encoding end predicts the vertex coordinates based on one or more parallelograms; it should be noted here that usually in the case of dislocation encoding, the encoding end needs to predict the coordinates based on multiple parallelograms, but there may be a situation where the number of waiting dislocations is small and not enough to form multiple parallelograms. Therefore, in the case of dislocation encoding, it may ultimately use a single parallelogram for coordinate prediction.
[0149] When using multiple parallelograms for coordinate prediction, a prediction value needs to be obtained based on each parallelogram respectively, and then these prediction values are averaged to finally obtain the predicted coordinates.
[0150] The following specifically describes two encoding implementation methods as follows.
[0151] I. Parallel encoding method without dislocation
[0152] To enable the decoding end to complete the connection relationship reconstruction and vertex geometry prediction through one traversal and keep the predictions of the encoding and decoding ends consistent, corresponding restrictions are imposed on the encoding end in the embodiments of the present application, and an implementation method without dislocation encoding is proposed. In this method, during the process of traversing the grid, if the connection mode is determined to be mode C, the geometric coordinates of the vertex can be immediately encoded, such as performing parallelogram prediction encoding, and its schematic diagram is as Figure 5 shown.
[0153] The determination of the connection relationship mode uses the Edgebreaker algorithm. Based on the idea of the above parallel encoding method without dislocation, the geometric coordinates of the vertex only use the triangles that have been traversed currently during prediction. That is, for each vertex to be predicted, only single parallelogram prediction can be used.
[0154] Such as Figure 6As shown in the figure, for the vertex D to be predicted, only the vertices A, B, and C of the triangle traversed immediately before it are used to perform parallelogram prediction on it. Then, the predicted result (i.e., the predicted coordinates) is subtracted from the true coordinates of the vertex to obtain the geometric coordinate residual (residual value) of the vertex.
[0155] II. Parallel encoding method with misalignment
[0156] Since the geometric prediction of vertices in the parallel encoding method without misalignment can only use single parallelogram prediction, the prediction effect may not be ideal. To improve the prediction effect, a parallel encoding method with misalignment is proposed. In this implementation method, after the connection mode is determined to be mode C, the geometric prediction of the vertex is postponed. After traversing the connection modes corresponding to the triangles of several subsequent vertices, the geometric coordinates of the vertex are predicted. At this time, since some subsequent triangles have been traversed, multi-parallelogram prediction can be used to predict the geometric coordinates of the vertex. The schematic diagram is as Figure 7 shown.
[0157] The determination of the connection relationship mode still uses the Edgebreaker algorithm. The geometric prediction of vertices in mode C is postponed, and a maximum waiting number (MAX_COUNT) for misalignment is specified, which can also be understood as the maximum waiting connection relationship number or the maximum misalignment number. When the second target condition is met, multi-parallelogram prediction is performed on the vertex to be predicted.
[0158] It should be noted that the triangles traversed during the waiting process can also be used for the geometric coordinate prediction of the triangle vertices in mode C. In this way, for most vertices to be predicted, multi-parallelogram prediction can be used to obtain a better prediction effect.
[0159] As Figure 8 shown, due to misalignment, a certain number of triangles after mode C are traversed. Rotating counterclockwise around the vertex D to be predicted, parallelogram 1 and parallelogram 2 are searched. Then, vertices A, B, C and C, E, F can all be used to perform single parallelogram prediction on point D, and then the average value of the two prediction results is the final prediction result.
[0160] It should be noted that in order to ensure that the decoding end can accurately decode the geometric information bitstream, optionally, in one implementation method, the method further includes:
[0161] The encoding end sends the encoding method of the vertices in the basic grid to the decoding end;
[0162] wherein, the encoding method includes one of the following:
[0163] B11. The method without using misalignment encoding;
[0164] This situation can be understood as that the encoding end uses a parallel encoding method without dislocation.
[0165] B12. The method of using dislocation encoding and the maximum waiting number of dislocations;
[0166] This situation can be understood as that the encoding end uses a parallel encoding method with dislocation, and at the same time, the maximum waiting number is informed to the decoding end, so that the decoding end can make a decoding judgment based on the same conditions as the encoding end to ensure the accuracy of decoding.
[0167] Optionally, this encoding method can be placed in the basic grid sequence parameter set and transmitted to the decoding end.
[0168] Specifically, in the intra-frame mode, in order to ensure the consistency of the connection relationship of the basic grid and the geometric information processing method between the encoding and decoding ends, it is necessary to transmit the encoding scheme adopted by it. Therefore, the syntax structure needs to specify what method the decoding end uses for the connection relationship and geometric information, and the relevant information about the maximum waiting number of dislocations if the dislocation encoding method is used. This syntax structure is designed based on the VDMC syntax structure, and the relevant parameter information should be placed in the basic grid sequence parameter set, and its syntax structure is as follows.
[0169]
[0170] Among them, bmsps_max_malposition_count is used to identify the encoding method adopted for the connection relationship of the current basic grid and the vertex geometry prediction and whether the dislocation encoding method is used. For example, when the value is -1, it means that the parallel encoding method of the embodiment of the present application is not used, and the decoding of the vertex geometry is traversed only after all the connection relationships are reconstructed; when the value is 0, it means that the parallel encoding method without dislocation of the embodiment of the present application is used, and at this time, the reconstruction of the connection relationship and the vertex geometry prediction are strictly parallel; when the value is a positive integer, it means that the parallel encoding method with dislocation of the embodiment of the present application is used, and its value represents the maximum waiting number of dislocations between the reconstruction of the connection relationship and the vertex geometry prediction.
[0171] As Figure 9 shown, the specific implementation process of applying the encoding method of the embodiment of the present application is as follows:
[0172] At the encoding end, for the input mesh, it is first simplified by a simplification module, and then new texture coordinates are generated for the mesh through mesh parameterization. Subsequently, the parameterized mesh is subdivided and deformed, that is, new vertices are inserted into the mesh according to a specific subdivision method and the distances from the vertices of the subdivided mesh to the nearest neighbor points of the input mesh are calculated, which is called displacement information. Subsequently, the parameterized mesh is adjusted according to the displacement information, that is, the vertex positions of the mesh before subdivision and deformation, and the adjusted mesh is called the base mesh, which is sent to the base mesh encoding module for compression. When encoding the base mesh, it will be divided into connection relationship, geometric information, and attribute information for separate encoding. After encoding, the base mesh is reconstructed, and then the order of the displacements is adjusted according to the vertex order of the reconstructed base mesh. Subsequently, the vertex displacement information with the adjusted order is first subjected to wavelet transform, the transformed coefficients are quantized, and then the quantized coefficients are arranged into a two-dimensional image according to a specific scanning order, and the two-dimensional image is encoded using a video encoder. Then, the reconstructed displacement information is applied to the subdivided base mesh to obtain the reconstructed subdivided and deformed mesh, and the mesh, the original input mesh, and its corresponding texture map are input into the corresponding texture map conversion module to obtain the texture map corresponding to the reconstructed mesh, and the texture map is also encoded using a video encoder. For the parameters used in the encoding process, such as the type of video encoder used, the type of mesh encoder, transformation parameters, quantization parameters, etc., they are transmitted to the decoding end through auxiliary information.
[0173] Among them, in the base mesh compression part in the intra-frame mode, for the input mesh connection relationship and vertex geometric coordinates, traverse the mesh to determine the connection relationship mode, and the connection relationship of the mesh can be represented as a connection relationship sequence (CLERS sequence); at the same time, after partial connection relationship mode determination, predict the geometric coordinates of the vertices, obtain the predicted coordinates of the vertices, and subtract them from their true coordinates to obtain the geometric coordinate residuals of the vertices. At this time, the mode determination of the connection relationship and the prediction of the geometric coordinates of the vertices can be completed with only one mesh traversal. Finally, the connection relationship sequence and the geometric coordinate residuals are entropy encoded and then multiplexed with the encoded bitstream of the attribute information to form the final output bitstream.
[0174] The following is a brief description of each implementation process as follows.
[0175] 1. Mesh Simplification
[0176] Mesh simplification is to simplify the currently input mesh to a base mesh with relatively fewer points and faces, and as much as possible maintain the shape of the original mesh. The key point of mesh simplification lies in the simplification operation and the corresponding error metric. A feasible mesh simplification operation is as Figure 10 shown, merging the vertices at both ends of an edge into one vertex and deleting the connection between these two vertices. Repeat this process in the whole mesh according to a certain rule to reduce the number of faces and vertices of the mesh to the target value.
[0177] During the simplification process, a certain error metric can be selected to optimize the simplification result. For example, the sum of the equation coefficients of all adjacent faces of a vertex can be chosen as the error metric for that vertex, and the error metric for the corresponding edge is the sum of the error metrics of the two vertices on the edge. In other words, the error generated by merging an edge is the sum of the distances from the merged vertex to all the planes adjacent to the original two vertices of the edge.
[0178] After determining the simplification operation and the corresponding error metric, the mesh simplification is iteratively carried out. First, the vertex errors of the initial mesh are calculated to obtain the errors of each edge. Then, each edge is sorted in ascending order of error, and the edge with the smallest error is selected for merging each time. Meanwhile, the position of the merged vertex is calculated, and the errors of all the edges related to the merged vertex are updated. That is, the order of the edge arrangement is updated to ensure that each iteration is based on the global error metric. Through iteration, the faces of the mesh are simplified to the number required for lossy coding.
[0179] 2. Mesh Parameterization
[0180] This step requires regenerating texture coordinates for the simplified mesh. Currently, there are many algorithms for parameterizing the mesh, such as the Isocharts algorithm, which uses spectral analysis to achieve stretch-driven 3D mesh parameterization, unfolds, fragments, and packs the 3D mesh into a 2D texture domain.
[0181] 3. Base Mesh Compression
[0182] The input base mesh is divided into three types of information: connectivity, geometry, and attribute information. These three types of information are respectively processed and compressed into bitstreams and output.
[0183] Any of the above implementation methods can be used for encoding the geometry information, which will not be elaborated here.
[0184] 4. Subdivision Deformation
[0185] The basic idea of the subdivision and deformation module is as Figure 11 shown. The same concept is applied to the input 3D mesh to generate displacement vector information. In Figure 11 , the input 2D curve (represented by a 2D polyline), called the "original" curve, is first downsampled to generate a basic curve / polyline, called the "simplified" curve. Then, the subdivision scheme is applied to the simplified polyline to generate the "subdivided" curve. Subsequently, the subdivided polyline is deformed to obtain a better approximation of the original curve. That is, a geometric displacement vector ( Figure 11As shown by the red arrows in [Figure 0], the shape of the subdivided curve is made to approximate the shape of the original curve as closely as possible. These geometric displacement vectors are the geometric displacement vector information output by this module. The same deformation process is also applied to the attribute information corresponding to the vertices, thereby obtaining the corresponding attribute displacement vectors.
[0186] The subdivision deformation module takes the parameterized sub-mesh as input. First, the input mesh is subdivided, and the subdivision scheme can be arbitrarily selected. One possible scheme is the midpoint subdivision scheme, which divides each triangle into four sub-triangles in each subdivision iteration, as Figure 12 shown. New vertices are introduced in the middle of each edge. The subdivision of geometric information and attribute information is carried out independently because the connection relationships of geometric information and attribute information are usually different.
[0187] This scheme calculates the position Pos(v 12 ) of the midpoint v of the newly introduced edge (v1, v2) as shown in Formula 1. 12
[0188] Formula 1:
[0189]
[0190] where Pos(v1) and Pos(v2) are the geometric coordinates of vertices v1 and v2 respectively.
[0191] For the subdivided mesh, find the nearest neighbor points (including points on the original mesh surface) of each of its points on the original input mesh. The search can be accelerated through data structures such as kdTree. The displacement vectors of the geometric coordinates of each vertex of the subdivided mesh are obtained by calculating the distances between each vertex on the subdivided mesh and the geometric coordinates of its nearest neighbor points on the original input mesh. This module passes the generated displacement vectors to the subsequent module for encoding.
[0192] At the same time, for the generated displacement vectors, they are in the same global coordinate system as the input mesh. One possible optimization method is to transform them into the local coordinate system. The local coordinate system of each vertex is defined by the normal vector of the vertex on the subdivided mesh. The advantage of this method is that the normal component of the geometric displacement vector has a more significant impact on the quality of the reconstructed mesh than the two tangential components. Therefore, larger quantization parameters can be set for the tangential components.
[0193] 5. Wavelet Transform
[0194] A transform can be applied to the displacement vectors to reduce the correlation between their data. An optional transform is the linear wavelet transform, and its prediction process is defined as shown in Formula 2.
[0195] Formula 2:
[0196]
[0197] Among them, v is the newly inserted midpoint on the edge (v1, v2), and Signal(v), Signal(v1), and Signal(v2) are the displacement vectors corresponding to the vertices v, v1, and v2 respectively. After predicting the displacement vector of vertex v, it is updated, and the update process is defined as shown in Equation 3.
[0198] Equation 3:
[0199]
[0200] where v * is the set of all vertices adjacent to vertex v. The transformed displacement vector is called the wavelet coefficient.
[0201] (6) Coefficient quantization
[0202] The transformed displacement vector, that is, the wavelet coefficient, can be quantized, and there are various quantization methods. One method is shown in Equations 4 and 5.
[0203] Equation 4:
[0204] disp[v].d[k] = floor(disp[v].d[k] * scale[k])
[0205] Equation 5:
[0206]
[0207] Among them, disp[v] represents the value after transformation of the displacement vector of the v-th vertex, d[k] represents the k-th value of the displacement vector, floor represents rounding down. bitDepthPosition represents the bit depth of the geometric position of the current mesh vertex, and qp[k] represents the quantization parameter of the k-th coefficient. As mentioned above, after converting the coordinate system of the displacement vector, its normal component has a more significant effect on the quality than the tangential component. Therefore, a larger quantization parameter can be used for the tangential component.
[0208] At the same time, according to the characteristics of wavelet transform, different quantization parameters can also be used for the newly generated vertices and the original vertices after subdivision. That is, for the vertices after subdivision, the quantization parameter is updated as shown in Equation 6.
[0209] Equation 6:
[0210] scale[k] = scale[k] * lodScale[k]
[0211] Among them, lodScale[k] represents the coefficient of the quantization parameter at the current subdivision level.
[0212] 7. Displacement Encoding
[0213] The displacement encoding part takes the quantized wavelet coefficients as input for video encoding. The quantized wavelet coefficients need to be arranged in a two-dimensional image. One arrangement method is as follows:
[0214] Traverse the wavelet coefficients in the order from low frequency to high frequency.
[0215] For each coefficient, determine the index of the NxM pixel block it belongs to (for example, N = M = 16), and it should be stored in it according to the raster scan order of the block.
[0216] Calculate the position of the corresponding NxM pixel block on the image according to the Morton order.
[0217] Here, the arrangement method is not limited, and other arrangement schemes can also be used, such as zigzag order, raster order, etc. The encoder can explicitly specify the corresponding arrangement scheme in the bitstream.
[0218] After arranging the wavelet coefficients on the two-dimensional image, the video encoder can be directly used to encode them. The proposed scheme can use any existing video encoder, and the type information of the video encoder needs to be encoded in the auxiliary information.
[0219] 8. Deformed Mesh Reconstruction
[0220] In the displacement encoding module, obtain the reconstructed displacement, that is, obtain the displacement vector consistent with the decoding end through inverse quantization and inverse transformation. After obtaining the reconstructed geometric displacement vector, subdivide the reconstructed basic mesh and obtain the reconstructed subdivided deformed mesh according to the corresponding displacement vector, and transfer it to the texture map conversion module.
[0221] 9. Texture Map Conversion
[0222] The texture map conversion module performs texture map conversion according to the input original mesh, input original texture map, and reconstructed deformed mesh.
[0223] The steps of texture map conversion are as follows:
[0224] S11. Calculate the texture coordinates of each pixel on the texture map to be generated. For example, the texture coordinates corresponding to pixel A(i, j) are P(u, v).
[0225] S12. Determine whether the texture coordinates are within a certain triangular face after parameterization of the subdivided deformed mesh.
[0226] S13. If the texture coordinates do not belong to any triangular face, mark the pixel as an empty pixel, and then it can be filled with a filling algorithm.
[0227] S14. If the texture coordinate belongs to a triangular face, then:
[0228] S141. Mark the pixel as filled.
[0229] S142. Calculate the barycentric coordinates of the texture coordinate in the current triangular face according to the texture coordinate.
[0230] S143. Map the two-dimensional texture coordinate to the three-dimensional geometric coordinate according to the barycentric coordinates and the corresponding triangular face, that is, map to the point on the subdivided and deformed mesh corresponding to the texture coordinate, as shown by M(x, y, z) in the figure.
[0231] S144. Find the point closest to the three-dimensional coordinate on the input original mesh, as shown by M′(x, y, z) in the figure.
[0232] S145. Calculate the barycentric coordinates of the three-dimensional coordinate according to the triangular face where it is located and map it to two dimensions, calculate its texture coordinate, that is, P′(u′, v′).
[0233] S146. Sample on the input original texture map through the texture coordinate to obtain the value A′(i′, j′) at the corresponding pixel position.
[0234] S147. Assign this value to the corresponding pixel A(i, j) on the texture map to be generated.
[0235] 10. Texture map compression
[0236] After obtaining the converted texture map, for the empty pixels in it, existing filling algorithms (such as the Push-Pull algorithm) can be used to fill these empty pixels. Then, existing video encoders, such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc., can be used to encode it to obtain the bitstream of the output texture map. In addition, operations such as color space conversion and chrominance subsampling can be selectively applied to make the video encoding obtain better rate-distortion performance, such as the color space conversion from RGB 444 to YUV420.
[0237] 11. Auxiliary information
[0238] During the encoding process, there are some alternative schemes in each module, such as the type of mesh encoder, the type of video encoder, the mesh subdivision scheme, the displacement vector transformation scheme, and the encoding method of displacement, etc. The proposed framework allows the use of different schemes. Therefore, the selected scheme needs to be passed to the decoding end to guide correct decoding.
[0239] After all modules are coded, the base mesh part bitstream, texture coordinate part bitstream, displacement vector video bitstream, attribute map video bitstream, and auxiliary information bitstream are mixed to obtain the final output bitstream of the coding end.
[0240] In summary, in the embodiment of the present application, when traversing the connection relationship, the geometric coordinates are coded simultaneously to achieve parallel processing of the connection relationship and geometric coordinates. When decoding, parallel processing is also used, so as to shorten the coding time and achieve the purpose of reducing the coding delay.
[0241] As Figure 13 shown, the embodiment of the present application provides a decoding method, including:
[0242] Step 1301, the decoding end obtains a base mesh bitstream, where the base mesh bitstream includes a connection relationship bitstream and a geometric information bitstream, the connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the base mesh belongs;
[0243] It should be noted that the connection relationship sequence includes the connection modes corresponding to all triangles to which the vertices in the base mesh belong. The connection mode is determined when traversing the base mesh using the Edgebreaker algorithm, and can also be referred to as the connection relationship mode.
[0244] It should be noted that the connection modes involved in the embodiment of the present application mainly involve mode C, mode L, mode R, mode S, and mode E.
[0245] Among them, mode C means that the triangles connected by the other two sides except the traversed side have not been traversed; mode L means that the triangle connected by the left side except the traversed side has been traversed; mode R means that the triangle connected by the right side except the traversed side has been traversed; mode S means that the triangles connected by the incoming side except the traversed side have not been traversed, but the vertices of the triangle have been traversed; mode E means that the triangles connected by the other two sides except the traversed side have all been traversed.
[0246] It should be noted that since the connection mode is determined based on the triangle, the triangle to which the vertex belongs mentioned in the embodiment of the present application refers to the triangle based on which the connection mode determination is made.
[0247] Step 1302, during the process of the decoding end reconstructing the connection relationship of the vertices according to the connection relationship sequence, the geometric information bitstream is decoded.
[0248] It should be noted that, when reconstructing the connection relationship of vertices in the embodiments of the present application, the decoding of the geometric information bitstream is performed simultaneously, so as to achieve the parallel processing of the connection relationship and geometric coordinates, which can shorten the decoding time and achieve the purpose of reducing the decoding delay.
[0249] Optionally, in one implementation, the specific implementation of decoding the geometric information bitstream during the process of reconstructing the connection relationship of vertices according to the connection relationship sequence includes:
[0250] Step c1: The decoding end obtains the connection mode corresponding to the triangle to which the first target vertex belongs according to the connection relationship sequence;
[0251] Optionally, the triangle to which the first target vertex belongs refers to the triangle formed by the two vertices where the active edge is located and the first target vertex. This active edge (which can also be referred to as the traversed edge) can be understood as the edge used for reconstructing the connection relationship.
[0252] Step c2: If the connection mode is the first mode, the decoding end decodes the first target vertex;
[0253] Wherein, the first target vertex is any un-decoded vertex in the geometric information bitstream.
[0254] Optionally, the first mode mentioned in the embodiments of the present application can be mode C.
[0255] It should be noted that, in this implementation, the un-decoded vertices corresponding to the triangles corresponding to mode C are decoded. That is to say, if the parallel decoding method of the connection relationship and geometric information is required, the decoding end performs the parallel decoding method of the connection relationship and geometric information for each un-decoded vertex corresponding to the triangle corresponding to mode C in the geometric information bitstream.
[0256] Furthermore, if the parallel decoding method of the connection relationship and geometric information is executed, when decoding the first target vertex in the embodiments of the present application, there are two different implementation methods. One implementation method is that as long as it is determined that the connection mode corresponding to the triangle to which the first target vertex belongs is mode C, the first target vertex is immediately decoded. This method can be understood as a non-misaligned decoding method (i.e., no misaligned decoding). Non-misaligned decoding can be understood as: not waiting for the reconstruction of the connection relationship of the vertices after the first target vertex. Another implementation method is the misaligned decoding method (i.e., performing misaligned decoding). Misaligned decoding can be understood as: waiting for the reconstruction of the connection relationship of the vertices after the first target vertex; To determine which decoding method to specifically use, optionally, in one implementation, the step of if the connection mode is the first mode, the decoding end decodes the first target vertex includes:
[0257] Step d1: The decoding end obtains the encoding method of the vertices in the base grid.
[0258] Step d2: The decoding end decodes the first target vertex according to the encoding method.
[0259] Among them, the encoding method includes one of the following:
[0260] P11: The misalignment encoding method is not adopted.
[0261] This situation can be understood as that the encoding end uses a parallel encoding method without misalignment, and the corresponding decoding end needs to use a parallel decoding method without misalignment.
[0262] P12: The misalignment encoding method and the maximum waiting number of misalignments are used.
[0263] This situation can be understood as that the encoding end uses a parallel encoding method with misalignment, and the corresponding decoding end needs to use a parallel decoding method with misalignment.
[0264] It should be noted that if the decoding end wants to accurately decode the geometric information code stream, it must first know the encoding method of the encoding end. Only after determining the encoding method can the decoding end use the corresponding decoding method of the encoding end for decoding.
[0265] This encoding method is sent from the encoding end to the decoding end. Usually, this encoding method can be transmitted to the decoding end in the base grid sequence parameter set.
[0266] Optionally, in one implementation, the decoding of the first target vertex includes:
[0267] The decoding end uses at least one parallelogram to predict the first target vertex to obtain a prediction result.
[0268] The decoding end determines the reconstructed value of the first target vertex according to the prediction result and the residual value corresponding to the first target vertex.
[0269] Generally, subtracting the residual value from the prediction result can obtain the reconstructed value. This residual value is obtained by decoding the geometric information code; this reconstructed value can be understood as the true coordinates of the first target vertex.
[0270] It should be noted that if the coding method is not the misalignment coding method, it means that after the coding end determines that the connection mode corresponding to the triangle to which the first target vertex belongs is mode C, it immediately encodes the first target vertex. During decoding, after the decoding end determines that the connection mode corresponding to the triangle to which the first target vertex belongs is mode C, it immediately decodes the first target vertex. If the coding method is the misalignment coding method, it means that after the coding end determines that the connection mode corresponding to the triangle to which the first target vertex belongs is mode C, it waits for the determination of several connection relationships before encoding the first target vertex. Then, after the decoding end determines that the connection mode corresponding to the triangle to which the first target vertex belongs is mode C, it also needs to wait for the reconstruction of several connection relationships before decoding the first target vertex.
[0271] Optionally, in one implementation manner, when the coding method is the misalignment coding method and the maximum waiting number of misalignments, the specific implementation of decoding the first target vertex includes:
[0272] The decoding end decodes the first target vertex when the first target condition is met;
[0273] Among them, the first target condition includes one of the following:
[0274] C11. The number of connection modes traversed after the first target vertex reaches the maximum waiting number;
[0275] This situation can be understood as that after traversing the connection mode corresponding to the triangle to which the first target vertex belongs, N connection modes are traversed again, and N is equal to the maximum waiting number.
[0276] It should be noted that the maximum waiting number can be understood as the maximum waiting connection mode number or the maximum misalignment number. For example, if the misalignment number is 2, it means that it is necessary to wait for the reconstruction of the connection relationships corresponding to the two connection modes after the first target vertex to complete before encoding the first target vertex.
[0277] C12. There is a first mode within the maximum waiting number after the first target vertex;
[0278] This situation can be understood as that it is determined that there is a next first mode within the maximum waiting number after the first target vertex; it can also be understood as that after traversing the connection mode corresponding to the triangle to which the first target vertex belongs, a next first mode is encountered during the subsequent connection mode traversal, and the next first mode appears before the maximum waiting number is reached.
[0279] This situation can be understood as that when encountering the next mode C, it is necessary to decode the first target vertex.
[0280] C13, reaching the last second mode in the connected component to which the first target vertex belongs;
[0281] Optionally, the second mode mentioned in the embodiments of the present application may be mode E.
[0282] It should be noted that each basic grid may involve multiple traversals. Each traversal edge corresponds to a connected component, and each connected component ends with mode E, that is, the last mode of each connected component is mode E.
[0283] It should be noted that the conditions adopted by the decoding end correspond to those adopted by the encoding end. For example, if A11 is used at the encoding end, then C11 is used correspondingly at the decoding end; if A12 is used at the encoding end, then C12 is used correspondingly at the decoding end; if A13 is used at the encoding end, then C13 is used correspondingly at the decoding end.
[0284] The following specifically describes two decoding implementation methods as follows.
[0285] I. The parallel decoding method without misalignment, corresponding to the parallel decoding method without misalignment encoding at the encoding end
[0286] In order to reduce the number of traversals at the decoding end and enable the reconstruction of the connection relationship and the prediction of the vertex geometric coordinates to be completed through one traversal, a parallel decoding method without misalignment is proposed. When the connection relationship sequence traverses to mode C, this decoding method first reconstructs its connection relationship, and then immediately performs a single parallelogram prediction on the newly created vertex, as Figure 14 shown.
[0287] The Edgebreaker algorithm is used to traverse the connection relationship sequence for the reconstruction of the connection relationship. Based on the idea of the above parallel decoding method without misalignment, the geometric coordinates of the vertex are predicted only using the known connection relationship. At this time, for each vertex to be predicted, only a single parallelogram prediction can be used.
[0288] As Figure 15 shown, for the vertex D to be predicted, only the vertices A, B, and C of the previous reconstructed triangle are used to perform a single parallelogram prediction on it. Then, the prediction result (i.e., the predicted coordinate) is added to the vertex geometric residual (i.e., the residual value) decoded from the bitstream to obtain the true coordinate of the vertex.
[0289] II. The parallel decoding method without misalignment, corresponding to the decoding method with misalignment encoding at the encoding end
[0290] In the parallel decoding method without misalignment, the geometric prediction of vertices can only use single parallelogram prediction, and the prediction effect may not be as ideal as that of multi-parallelogram prediction. Therefore, a misaligned parallel decoding method is proposed. After traversing to the C mode and reconstructing its connection relationship, the geometric prediction of its vertices is postponed. After traversing a certain connection relationship, the prediction of the vertex is performed. As Figure 16 shown.
[0291] The connection relationship reconstruction still uses the Edgebreaker algorithm to traverse the connection relationship sequence. The geometric prediction of the vertices of mode C is postponed, and a maximum waiting number (MAX_COUNT) of misalignment is specified, which can also be understood as the maximum waiting connection relationship number or the maximum misalignment number. When the first target condition is met, multi-parallelogram prediction is performed on the vertex to be predicted.
[0292] As Figure 17 shown, due to misalignment, a certain number of connection relationships after mode C are reconstructed. Rotating counterclockwise around the vertex D to be predicted, parallelogram 1 and parallelogram 2 are searched. Then the vertices A, B, C and C, E, F can be used for single parallelogram prediction of vertex D. The average value of the results of the two single parallelogram predictions is the final prediction result. Then, the prediction result (i.e., the predicted coordinates) is added to the vertex geometric residual (i.e., the residual value) decoded from the bitstream to obtain the true coordinates of the vertex.
[0293] As Figure 18 shown, the specific implementation process of applying the decoding method of the embodiment of the present application is as follows:
[0294] At the decoding end, for the received bitstream, the decoding end first demultiplexes each part of the bitstream to obtain the base mesh bitstream, the displacement video bitstream, the texture map video bitstream, and the auxiliary information bitstream respectively. For the base mesh bitstream, the base mesh is decoded using the mesh decoder indicated by the auxiliary information. The displacement video bitstream and the texture map video bitstream are decoded by the video decoder. For the displacement part, after video decoding, the displacement needs to be taken out of the image through the displacement decoding module, and steps such as inverse quantization and inverse transformation are performed, and then it is applied to the subdivided base mesh to obtain the deformed mesh reconstructed at the decoding end. The texture map is the texture map corresponding to the reconstructed deformed mesh after decoding. Subsequently, the application or rendering module processes the reconstructed deformed mesh and the decoded texture map as inputs.
[0295] Among them, in the basic mesh decompression part in the intra-frame mode, for the connection relationship sequence and the geometric coordinate residuals of the vertices obtained after entropy decoding, by traversing the connection relationship sequence, the connection relationship of the mesh can be reconstructed; at the same time, after partial connection relationship reconstruction, the geometric coordinates of the vertices are predicted to obtain their predicted coordinates, and the true geometric coordinates of the vertices are obtained after adding them to the input geometric coordinate residuals. At this time, the reconstruction of the mesh connection relationship and the prediction of the vertex geometric coordinates can be completed by traversing the connection relationship sequence only once. Finally, the basic mesh is reconstructed with the connection relationship, vertex geometric coordinates, and decoded attribute information.
[0296] The following briefly describes each implementation process as follows.
[0297] 1. Auxiliary information decoding
[0298] The decoding end first determines the decoding scheme according to the auxiliary information, which mainly includes the displacement coding method, indicating whether the displacement is encoded by the video encoder or by the entropy encoder; the static mesh encoder type, which guides the decoding end to use the corresponding static mesh decoder; the video encoder type, which guides the decoding end to use the corresponding video decoder; the subdivision scheme, that is, the scheme for subdividing the basic mesh in the reconstructed deformed mesh, and the subdivision schemes at the encoding and decoding ends should be consistent; and also optional spatial displacement transformation schemes, coefficient arrangement schemes, etc.
[0299] 2. Basic mesh decoding
[0300] Corresponding to the encoding end, for basic mesh decoding, the input bitstream is divided into three bitstreams: the connection relationship bitstream, the geometric information bitstream, and the attribute information bitstream, and the basic mesh is reconstructed after decoding them respectively.
[0301] The decoding of the geometric information bitstream can adopt any one of the above implementation methods, which will not be elaborated here.
[0302] 3. Displacement decoding module
[0303] In the displacement decoding module, it is necessary to determine the displacement decoding method according to the auxiliary information flag. If the auxiliary information indicates that the displacement information is encoded by the video encoder, the decoding end calls the corresponding video decoder to decode the displacement bitstream; if the auxiliary information indicates that the displacement information is encoded by the entropy encoder, the entropy decoder is directly used for decoding.
[0304] For the decoded displacement information, it is also necessary to obtain the displacement vector corresponding to the reconstructed subdivided mesh vertices through the displacement reconstruction module. The operations of the displacement reconstruction module are mainly to perform inverse quantization, inverse transformation, etc. on the decoded displacement information, that is, wavelet coefficients, according to the quantization parameters and transformation parameters indicated by the auxiliary information, and for the information decoded by the video, it is first necessary to extract the corresponding displacement information from the two-dimensional image according to the arrangement method at the encoding end.
[0305] 4. Subdivision
[0306] The operation of this subdivision module is the same as that of the subdivision operation at the encoding end. The subdivision method of the basic grid, the number of iterations, etc. are indicated by auxiliary information.
[0307] 5. Reconstruct the deformed grid
[0308] After the basic grid and the displacement vector decoding and reconstruction are completed, the deformed grid is reconstructed based on these two parts. Add the corresponding displacement vector to each vertex of the subdivided grid, as shown in Formula Seven.
[0309] Formula Seven:
[0310] deformedmesh[i].v[k] = subdivmesh[i].v[k] + displacement[k]
[0311] Where, subdivnmesh[i].v[k] is the geometric coordinate of the kth vertex after the subdivision of the basic grid in the current frame (indexed by i), displacement[k] is the spatial displacement vector corresponding to the kth vertex, and subdivnmesh[i].v[k] is the geometric coordinate of the kth vertex after the subdivision and deformation in the current frame.
[0312] 6. Texture map decoding
[0313] The texture map decoder is responsible for decoding the texture map bitstream. The texture map bitstream is decoded using the video decoder indicated in the auxiliary information. Optionally, perform a color space conversion on it to obtain an image format consistent with the input texture at the encoding end, and obtain the finally decoded output texture map. Figure 1 to obtain the same image format as the input texture at the encoding end, and obtain the finally decoded output texture map.
[0314] After the processing of each module is completed, finally, the deformed grid reconstructed at the decoding end and the corresponding attribute map are obtained. Subsequent applications use the reconstructed deformed grid and attribute map as inputs for processing.
[0315] It should be noted that in the embodiments of the present application, by using a corresponding method for decoding at the encoding end, parallel processing of the connection relationship and geometric coordinates is achieved, thereby shortening the decoding time and achieving the purpose of reducing the decoding delay.
[0316] The encoding method provided by the embodiments of the present application may have an encoding device as the execution subject. In the embodiments of the present application, taking the encoding device executing the encoding method as an example, the encoding device provided by the embodiments of the present application is described.
[0317] As Figure 19 shown, the encoding device 1900 of the embodiments of the present application includes:
[0318] A prediction module 1901 is configured to determine a connection mode of a basic mesh. When it is determined that the connection mode corresponding to the triangle to which a second target vertex belongs is a first mode, the second target vertex is predicted by using at least one parallelogram to obtain a prediction result, and the connection modes of the triangles to which the vertices other than the second target vertex in the at least one parallelogram belong have been determined;
[0319] An encoding module 1902 is configured to perform encoding based on the prediction result to obtain a geometric information bitstream and perform encoding on the connection relationship sequence corresponding to the basic mesh to obtain a connection relationship bitstream;
[0320] A first sending module 1903 is configured to send a basic mesh bitstream to a decoding end. The basic mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the basic mesh belongs.
[0321] Optionally, the prediction module 1901 is configured to:
[0322] When a second target condition is satisfied, the second target vertex is predicted by using at least one parallelogram to obtain a prediction result;
[0323] Wherein, the second target condition includes one of the following:
[0324] The number of connection modes determined after the second target vertex reaches a maximum waiting number;
[0325] There is a first mode within the maximum waiting number of connection mode determinations after the second target vertex;
[0326] Reach the last second mode in the connected component to which the second target vertex belongs.
[0327] Optionally, the device further includes:
[0328] A second sending module is configured to send the encoding method of the vertices in the basic mesh to the decoding end;
[0329] Wherein, the encoding method includes one of the following:
[0330] The method of non-use of staggered encoding;
[0331] The method of using staggered encoding and the maximum waiting number of staggers.
[0332] It should be noted that this device embodiment corresponds to the above method, and all implementation manners in the above method embodiment are applicable to this device embodiment and can achieve the same technical effects.
[0333] The encoding device in the embodiments of the present application may be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the terminal may include, but is not limited to, the types of the terminal 11 listed above, and other devices may be a server, a Network Attached Storage (NAS), etc., which are not specifically limited in the embodiments of the present application.
[0334] The encoding device provided in the embodiments of the present application can implement Figure 4 each process implemented by the method embodiment and achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0335] The embodiments of the present application further provide an encoding end, including a processor and a communication interface. The processor is used to determine the connection mode of the basic grid. When it is determined that the connection mode of the triangle corresponding to the second target vertex is the first mode, at least one parallelogram is used to predict the second target vertex to obtain a prediction result. For the triangles to which the vertices other than the second target vertex in the at least one parallelogram belong, the connection mode has been determined;
[0336] Based on the prediction result, encoding is performed to obtain a geometric information bitstream, and encoding is performed on the connection relationship sequence corresponding to the basic grid to obtain a connection relationship bitstream;
[0337] A basic grid bitstream is sent to the decoding end. The basic grid bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode of the triangle to which the vertex in the basic grid belongs.
[0338] Optionally, the processor is used to:
[0339] When a second target condition is satisfied, at least one parallelogram is used to predict the second target vertex to obtain a prediction result;
[0340] Wherein, the second target condition includes one of the following:
[0341] The number of connection modes determined after the second target vertex reaches the maximum waiting number;
[0342] There is a first mode within the maximum waiting number of connection mode determinations after the second target vertex;
[0343] Reach the last second mode in the connected component to which the second target vertex belongs.
[0344] Optionally, the communication interface is used for:
[0345] sending the encoding method of the vertices in the base grid to the decoding end;
[0346] wherein, the encoding method includes one of the following:
[0347] not adopting the misalignment encoding method;
[0348] using the misalignment encoding method and the maximum waiting number of misalignments.
[0349] This encoder embodiment corresponds to the above method embodiment. Each implementation process and implementation method of the above method embodiment can be applied to this electronic device embodiment and can achieve the same technical effect. Specifically, Figure 20 It is a schematic diagram of the hardware structure of an encoder for implementing an embodiment of the present application.
[0350] The encoder 2000 includes but is not limited to at least some components such as a radio frequency unit 2001, a network module 2002, an audio output unit 2003, an input unit 2004, a sensor 2005, a display unit 2006, a user input unit 2007, an interface unit 2008, a memory 2009, and a processor 2010.
[0351] Those skilled in the art can understand that the encoder 2000 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 2010 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 20 The structure of the electronic device shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine some components, or have different component arrangements, which will not be elaborated here.
[0352] It should be understood that in the embodiments of the present application, the input unit 2004 may include a Graphics Processing Unit (GPU) 20041 and a microphone 20042. The graphics processor 20041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 2006 may include a display panel 20061, and the display panel 20061 may be configured in the form of, for example, a liquid crystal display or an organic light emitting diode. The user input unit 2007 includes at least one of a touch panel 20071 and other input devices 20072. The touch panel 20071 is also referred to as a touch screen. The touch panel 20071 may include two parts: a touch detection device and a touch controller. The other input devices 20072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.
[0353] In the embodiments of the present application, after receiving downlink data from an access network device, the radio frequency unit 2001 may transmit it to the processor 2010 for processing; in addition, the radio frequency unit 2001 may send uplink data to a network-side device. Generally, the radio frequency unit 2001 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc.
[0354] The memory 2009 can be used to store software programs or instructions as well as various data. The memory 2009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 2009 can include volatile memory or non-volatile memory, or the memory 2009 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 2009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0355] The processor 2010 may include one or more processing units; optionally, the processor 2010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 2010 either.
[0356] Among them, the processor 2010 is used for:
[0357] Perform connection mode determination on the base mesh. When it is determined that the connection mode corresponding to the triangle to which the second target vertex belongs is the first mode, use at least one parallelogram to predict the second target vertex to obtain a prediction result. The connection modes of the triangles to which the vertices other than the second target vertex in the at least one parallelogram belong have all been determined;
[0358] Perform encoding based on the prediction result to obtain a geometric information bitstream and perform encoding on the connection relationship sequence corresponding to the base mesh to obtain a connection relationship bitstream;
[0359] Send the base mesh bitstream to the decoding end. The base mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the base mesh belongs.
[0360] Optionally, the processor is configured to:
[0361] When the second target condition is satisfied, use at least one parallelogram to predict the second target vertex to obtain a prediction result;
[0362] Wherein, the second target condition includes one of the following:
[0363] The number of connection modes determined after the second target vertex reaches the maximum waiting number;
[0364] There is a first mode within the maximum waiting number of connection mode determinations after the second target vertex;
[0365] Reach the last second mode in the connected component to which the second target vertex belongs.
[0366] Optionally, the radio frequency unit 2001 is configured to:
[0367] Send the encoding method of the vertices in the base mesh to the decoding end;
[0368] Wherein, the encoding method includes one of the following:
[0369] The method of non - staggered encoding is not adopted;
[0370] The method of staggered encoding and the maximum waiting number of staggering are used.
[0371] Preferably, an embodiment of the present application further provides an encoding end, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above - mentioned encoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0372] The embodiments of the present application further provide a readable storage medium. A program or instructions are stored on the computer-readable storage medium. When the program or instructions are executed by a processor, the various processes of the above-described encoding method embodiments are implemented, and the same technical effects can be achieved. To avoid repetition, details are not described herein again.
[0373] Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0374] Such as Figure 21 As shown, the decoding device 2100 of the embodiments of the present application includes:
[0375] An acquisition module 2101, configured to acquire a basic mesh code stream, where the basic mesh code stream includes a connection relationship code stream and a geometric information code stream, and the connection relationship code stream includes a connection relationship sequence, and the connection relationship sequence is used to indicate a connection mode corresponding to a triangle to which a vertex in the basic mesh belongs;
[0376] A decoding module 2102, configured to decode the geometric information code stream during the process of reconstructing the connection relationship of vertices according to the connection relationship sequence.
[0377] Optionally, the decoding module 2102 includes:
[0378] An acquisition unit, configured to acquire a connection mode corresponding to a triangle to which a first target vertex belongs according to the connection relationship sequence;
[0379] A decoding unit, configured to decode the first target vertex if the connection mode is a first mode;
[0380] Among them, the first target vertex is any undecoded vertex in the geometric information code stream.
[0381] Optionally, the decoding unit is configured to:
[0382] Acquire an encoding method of vertices in the basic mesh;
[0383] Decode the first target vertex according to the encoding method;
[0384] Among them, the encoding method includes one of the following:
[0385] A method without using staggered encoding;
[0386] A method using staggered encoding and a maximum waiting number of staggers.
[0387] Optionally, when the encoding method is the method using dislocation encoding and the maximum number of waiting times for dislocation, the decoding unit is configured to:
[0388] Decode the first target vertex when a first target condition is met;
[0389] Wherein, the first target condition includes one of the following:
[0390] The number of connection patterns traversed after the first target vertex reaches the maximum number of waiting times;
[0391] There is a first pattern within the maximum number of waiting times after the first target vertex;
[0392] Reach the last second pattern in the connected component to which the first target vertex belongs.
[0393] Optionally, the decoding unit is configured to:
[0394] Predict the first target vertex by using at least one parallelogram to obtain a prediction result;
[0395] Determine the reconstruction value of the first target vertex according to the prediction result and the residual value corresponding to the first target vertex.
[0396] It should be noted that this device embodiment corresponds to the above method. All implementation manners in the above method embodiment are applicable to this device embodiment and can achieve the same technical effects.
[0397] The decoding device in the embodiments of the present application may be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or a chip. This electronic device may be a terminal or other devices other than a terminal. Exemplarily, the terminal may include, but is not limited to, the types of the above-mentioned terminal 11, and other devices may be a server, a Network Attached Storage (NAS), etc., which are not specifically limited in the embodiments of the present application.
[0398] The decoding device provided in the embodiments of the present application can implement Figure 13 each process implemented by the method embodiment described above and achieve the same technical effects. To avoid repetition, details are not described here again.
[0399] The embodiments of the present application further provide a decoding end, including a processor and a communication interface. The processor is configured to obtain a basic grid bitstream, where the basic grid bitstream includes a connection relationship bitstream and a geometric information bitstream, and the connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection pattern corresponding to the triangle to which the vertex in the basic grid belongs;
[0400] During the process of reconstructing the connection relationship of vertices according to the sequence of connection relationships, decode the geometric information bitstream.
[0401] Optionally, the processor is configured to:
[0402] Obtain the connection mode corresponding to the triangle to which the first target vertex belongs according to the sequence of connection relationships;
[0403] If the connection mode is the first mode, decode the first target vertex;
[0404] Wherein, the first target vertex is any undecoded vertex in the geometric information bitstream.
[0405] Optionally, the processor is configured to:
[0406] Obtain the encoding method of the vertices in the basic mesh;
[0407] Decode the first target vertex according to the encoding method;
[0408] Wherein, the encoding method includes one of the following:
[0409] A method without using staggered encoding;
[0410] A method using staggered encoding and the maximum waiting number of staggers.
[0411] Optionally, when the encoding method is a method using staggered encoding and the maximum waiting number of staggers, the processor is configured to:
[0412] Decode the first target vertex when a first target condition is satisfied;
[0413] Wherein, the first target condition includes one of the following:
[0414] The number of connection modes traversed after the first target vertex reaches the maximum waiting number;
[0415] There is a first mode within the maximum waiting number after the first target vertex;
[0416] Reach the last second mode in the connected component to which the first target vertex belongs.
[0417] Optionally, the processor is configured to:
[0418] Predict the first target vertex by using at least one parallelogram to obtain a prediction result;
[0419] Determine the reconstruction value of the first target vertex according to the prediction result and the residual value corresponding to the first target vertex.
[0420] This decoding-end embodiment corresponds to the above method embodiment. Each implementation process and realization method of the above method embodiment can be applied to this electronic device embodiment and can achieve the same technical effect.
[0421] Optionally, an embodiment of the present application further provides a decoding end. The structure of the decoding end can be seen as Figure 20 shown and will not be elaborated here.
[0422] Among them, the processor is used for:
[0423] Obtain a basic mesh bitstream, where the basic mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the basic mesh belongs;
[0424] During the process of reconstructing the connection relationship of the vertices according to the connection relationship sequence, decode the geometric information bitstream.
[0425] Optionally, the processor is used for:
[0426] According to the connection relationship sequence, obtain the connection mode corresponding to the triangle to which the first target vertex belongs;
[0427] If the connection mode is the first mode, decode the first target vertex;
[0428] Among them, the first target vertex is any undecoded vertex in the geometric information bitstream.
[0429] Optionally, the processor is used for:
[0430] Obtain the encoding method of the vertices in the basic mesh;
[0431] According to the encoding method, decode the first target vertex;
[0432] Among them, the encoding method includes one of the following:
[0433] The method of not using staggered encoding;
[0434] The method of using staggered encoding and the maximum waiting number of staggers.
[0435] Optionally, when the encoding method is the method of using staggered encoding and the maximum waiting number of staggers, the processor is used for:
[0436] When the first target condition is satisfied, decode the first target vertex;
[0437] Wherein, the first target condition includes one of the following:
[0438] The number of connection patterns traversed after the first target vertex reaches the maximum waiting number;
[0439] There is a first pattern within the maximum waiting number after the first target vertex;
[0440] Reach the last second pattern in the connected component to which the first target vertex belongs.
[0441] Optionally, the processor is configured to:
[0442] Use at least one parallelogram to predict the first target vertex and obtain a prediction result;
[0443] Determine the reconstruction value of the first target vertex according to the prediction result and the residual value corresponding to the first target vertex.
[0444] Preferably, an embodiment of the present application further provides a decoding end, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0445] An embodiment of the present application further provides a readable storage medium. A program or instruction is stored on the computer-readable storage medium. When the program or instruction is executed by the processor, it implements each process of the above decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0446] Wherein, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0447] Optionally, such as Figure 22As shown in the figure, an embodiment of the present application further provides an electronic device 2200, including a processor 2201 and a memory 2202. A program or instruction that can run on the processor 2201 is stored on the memory 2202. When the electronic device is an encoding end, when the program or instruction is executed by the processor 2201, it implements each step of the above encoding method embodiment and can achieve the same technical effect. When the electronic device is a decoding end, when the program or instruction is executed by the processor 2201, it implements each step of the above decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0448] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above encoding method or decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0449] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.
[0450] Another embodiment of the present application provides a computer program / program product. The computer program / program product is stored in a storage medium. The computer program / program product is executed by at least one processor to implement each process of the above encoding method or decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0451] An embodiment of the present application further provides an encoding and decoding system, including: an encoding end and a decoding end. The encoding end can be used to execute the steps of the above encoding method, and the decoding end can be used to execute the steps of the above decoding method.
[0452] It should be noted that in this text, the term "including", "comprising" or any other variants thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also other elements not explicitly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0453] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0454] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A decoding method, characterized in that, Comprising: The decoding end obtains a basic mesh bitstream, where the basic mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the basic mesh belongs; During the process of reconstructing the connection relationship of the vertex according to the connection relationship sequence, the decoding end decodes the geometric information bitstream.
2. The method according to claim 1, wherein During the process of reconstructing the connection relationship of the vertex according to the connection relationship sequence, decoding the geometric information bitstream includes: The decoding end obtains the connection mode corresponding to the triangle to which the first target vertex belongs according to the connection relationship sequence; If the connection mode is the first mode, the decoding end decodes the first target vertex; Wherein, the first target vertex is any undecoded vertex in the geometric information bitstream.
3. The method according to claim 2, wherein If the connection mode is the first mode, the decoding end decodes the first target vertex, including: The decoding end obtains the encoding method of the vertices in the basic mesh; The decoding end decodes the first target vertex according to the encoding method; Wherein, the encoding method includes one of the following: The method of not using staggered encoding; The method of using staggered encoding and the maximum waiting number of staggers.
4. The method according to claim 3, wherein In the case where the encoding method is the method of using staggered encoding and the maximum waiting number of staggers, decoding the first target vertex includes: The decoding end decodes the first target vertex when the first target condition is satisfied; Wherein, the first target condition includes one of the following: The number of connection modes traversed after the first target vertex reaches the maximum waiting number; There is a first mode within the maximum waiting number after the first target vertex; Reaching the last second mode in the connected component to which the first target vertex belongs.
5. The method according to any one of claims 2-4, characterized in that, Decoding the first target vertex includes: The decoding end predicts the first target vertex by using at least one parallelogram to obtain a prediction result; The decoding end determines the reconstructed value of the first target vertex according to the prediction result and the residual value corresponding to the first target vertex.
6. A coding method, characterized in that, Comprising: The encoding end determines the connection mode of the basic mesh. When it is determined that the connection mode corresponding to the triangle to which the second target vertex belongs is the first mode, the encoding end predicts the second target vertex by using at least one parallelogram to obtain a prediction result. For the vertices other than the second target vertex in the at least one parallelogram, the connection mode has been determined; The encoding end encodes based on the prediction result to obtain a geometric information bitstream and encodes the connection relationship sequence corresponding to the basic mesh to obtain a connection relationship bitstream; The encoding end sends a basic mesh bitstream to the decoding end. The basic mesh bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the basic mesh belongs.
7. The method according to claim 6, characterized in that Predicting the second target vertex by using at least one parallelogram to obtain a prediction result, including: When the second target condition is satisfied, the encoding end predicts the second target vertex by using at least one parallelogram to obtain a prediction result; Wherein, the second target condition includes one of the following: The number of connection patterns determined after the second target vertex reaches the maximum waiting number; There is a first pattern within the maximum waiting number for connection pattern determination after the second target vertex; Reaching the last second pattern in the connected component to which the second target vertex belongs.
8. The method according to claim 6 or 7, characterized in that Further includes: The encoding end sends the encoding method of the vertices in the base grid to the decoding end; Wherein, the encoding method includes one of the following: Not using the misalignment encoding method; Using the misalignment encoding method and the maximum waiting number of misalignments.
9. A decoding device, characterized in that, Includes: An acquisition module, configured to acquire a base grid bitstream, where the base grid bitstream includes a connection relationship bitstream and a geometric information bitstream, the connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection pattern corresponding to the triangle to which the vertex in the base grid belongs; A decoding module, configured to decode the geometric information bitstream during the process of reconstructing the connection relationship of the vertices according to the connection relationship sequence.
10. The device according to claim 9, characterized in that, The decoding module includes: An acquisition unit, configured to acquire the connection pattern corresponding to the triangle to which the first target vertex belongs according to the connection relationship sequence; A decoding unit, configured to decode the first target vertex if the connection pattern is the first pattern; Wherein, the first target vertex is any undecoded vertex in the geometric information bitstream.
11. The device according to claim 10, characterized in that, The decoding unit is configured to: Acquire the encoding method of the vertices in the base grid; Decode the first target vertex according to the encoding method; Wherein, the encoding method includes one of the following: Not using the misalignment encoding method; Using the misalignment encoding method and the maximum waiting number of misalignments.
12. The device according to claim 11, characterized in that, When the encoding method is using the misalignment encoding method and the maximum waiting number of misalignments, the decoding unit is configured to: Decode the first target vertex when the first target condition is satisfied; Wherein, the first target condition includes one of the following: The number of connection patterns traversed after the first target vertex reaches the maximum waiting number; There is a first pattern within the maximum waiting number after the first target vertex; Reaching the last second pattern in the connected component to which the first target vertex belongs.
13. The device according to any one of claims 10-12, characterized in that, The decoding unit is configured to: Predict the first target vertex by using at least one parallelogram to obtain a prediction result; Determine the reconstruction value of the first target vertex according to the prediction result and the residual value corresponding to the first target vertex.
14. A decoding end, characterized in that, Includes a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the decoding method according to any one of claims 1 to 5 are implemented.
15. An encoding device, characterized in that, Includes: A prediction module, configured to determine a connection mode for a basic grid. When it is determined that the connection mode corresponding to the triangle to which the second target vertex belongs is the first mode, at least one parallelogram is used to predict the second target vertex to obtain a prediction result. The triangles to which the vertices other than the second target vertex in the at least one parallelogram belong have all been determined for the connection mode; An encoding module, configured to perform encoding based on the prediction result to obtain a geometric information bitstream and perform encoding on the connection relationship sequence corresponding to the basic grid to obtain a connection relationship bitstream; A first sending module, configured to send a basic grid bitstream to a decoding end. The basic grid bitstream includes a connection relationship bitstream and a geometric information bitstream. The connection relationship bitstream includes a connection relationship sequence, and the connection relationship sequence is used to indicate the connection mode corresponding to the triangle to which the vertex in the basic grid belongs.
16. The device according to claim 15, characterized in that, The prediction module is configured to: When a second target condition is satisfied, at least one parallelogram is used to predict the second target vertex to obtain a prediction result; Wherein, the second target condition includes one of the following: The number of connection modes determined after the second target vertex reaches the maximum waiting number; There is a first mode within the maximum waiting number for connection mode determination after the second target vertex; Reach the last second mode in the connected component to which the second target vertex belongs.
17. The device according to claim 15 or 16, characterized in that, It further includes: A second sending module, configured to send the encoding method of the vertices in the basic grid to the decoding end; Wherein, the encoding method includes one of the following: The method of not using staggered encoding; The method of using staggered encoding and the maximum waiting number of staggers.
18. An encoding end, characterized in that, It includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the encoding method according to any one of claims 6 to 8 are implemented.
19. A readable storage medium, characterized in that, The program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the decoding method according to any one of claims 1 to 5 or the steps of the encoding method according to any one of claims 6 to 8 are implemented.
20. A chip, characterized in that, The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run a program or instruction to implement the steps of the decoding method according to any one of claims 1 to 5 or the steps of the encoding method according to any one of claims 6 to 8.