Vertex motion vector encoding and decoding

Through vertex motion vector encoding and decoding technology, the 3D point cloud or grid is converted into 2D frames and compressed, solving the problem of inefficient dynamic 3D data transmission and achieving efficient data storage and transmission.

CN120500675APending Publication Date: 2025-08-15SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480007264.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-01-05
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently compress and transmit dynamic 3D point cloud and grid data, especially in scenarios with high bandwidth requirements, resulting in inefficient transmission efficiency.

Method used

Vertex motion vector encoding and decoding technology is used to convert 3D point clouds or grids into traditional 2D frames for compression, and vertex motion vector information is processed by selecting a combination of context encoding and bypass encoding to generate a compressed bit stream.

Benefits of technology

It improves the transmission efficiency of dynamic 3D point cloud and grid data, reduces dependence on dedicated hardware, and realizes efficient data storage and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120500675A_ABST
    Figure CN120500675A_ABST
Patent Text Reader

Abstract

An apparatus includes a communication interface configured to receive a compressed bitstream including vertex motion vector information. The vertex motion vector information includes one or more components. The apparatus includes a processor operably coupled to the communication interface. The processor is configured such that: a compressed bitstream including vertex motion vector information is parsed; a context is selected for a plurality of bins of a first component of the vertex motion vector information, where a first prefix bin of the first component is encoded based on a first context and a remaining prefix bin of the first component is encoded based on a second context, and the one or more suffix bins of the first component are encoded using bypass coding; and decoding a first component of vertex motion vector information based on the selected context.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 438,188, filed on January 10, 2023, entitled “VERTEX 3D MOTION VECTOR CODING”; U.S. Provisional Application No. 63 / 462,136, filed on April 26, 2023, entitled “VERTEX 3D MOTION VECTOR CODING”; and U.S. Application No. 18 / 399,323, filed on December 28, 2023, entitled “VERTEX MOTIONVECTOR CODING AND DECODING”, all of which are incorporated herein by reference in their entireties.

[0002] The present disclosure relates generally to video encoding and decoding, and more particularly, to vertex motion vector encoding and decoding, for example but not limited to, for dynamic meshes. Background Art

[0003] With the ready availability of powerful handheld devices such as smartphones, three-hundred-sixty-degree (360°) and three-dimensional (3D) stereoscopic video are emerging as new ways to experience immersive content. While 360° video offers consumers an immersive "real life" or "being there" experience by capturing a 360° outside-in view of the world, 3D stereoscopic video also provides a full six degrees of freedom (6DoF) experience of being and moving within the content. Users can interactively change their viewing angle and dynamically view any part of the captured scene or object they desire. Displays and navigation sensors can track the user's head movements in real time to determine the area of the 360° video or stereoscopic content that the user wants to view or interact with. Intrinsically 3D multimedia data, such as point clouds or 3D polygon meshes, can be used within immersive environments.

[0004] The description set forth in the Background section should not be assumed to qualify as prior art merely because it is set forth in the Background section.The Background section may describe aspects or embodiments of the present disclosure. Summary of the Invention

[0005] Technical Solution

[0006] One aspect of the present disclosure provides an apparatus comprising a communication interface configured to receive a compressed bitstream comprising vertex motion vector information. The vertex motion vector information may comprise one or more components. The apparatus may comprise a processor operably coupled to the communication interface. The processor may be configured to parse the compressed bitstream comprising the vertex motion vector information. The processor may be configured to select a context for a plurality of bins of a first component of the vertex motion vector information. A first prefix bin of the first component may be encoded based on a first context, and one or more remaining bins of the first component may be encoded using bypass coding. The processor may be configured to decode the first component of the vertex motion vector information based on the selected context.

[0007] One aspect of the present disclosure provides a method that includes receiving a compressed bitstream including vertex motion vector information. The vertex motion vector information may include one or more components. The method may include parsing the compressed bitstream including the vertex motion vector information. The method may include selecting a context for a plurality of bits of a first component of the vertex motion vector information. A first prefix bit of the first component may be encoded based on being a first context, and one or more remaining bits of the first component may be encoded using bypass coding. The method may include decoding the first component of the vertex motion vector information based on the selected context.

[0008] One aspect of the present disclosure provides an apparatus comprising a communication interface and a processor operably coupled to the communication interface. The processor may be configured to generate vertex motion vector information comprising one or more components. The processor may be configured to encode a first prefix bit of a first component of the vertex motion vector information based on a first context. The processor may be configured to encode one or more remaining bits of the first component using bypass coding. The processor may be configured to form a bitstream comprising the encoded vertex motion vector information. The processor may be configured to transmit the bitstream to a decoding device via the communication interface. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 An example communication system according to an embodiment is shown.

[0010] Figure 2 An example electronic device according to an embodiment of the present disclosure is shown.

[0011] Figure 3 An example electronic device according to an embodiment of the present disclosure is shown.

[0012] Figure 4An example block diagram of an inter-frame encoder according to an embodiment is shown.

[0013] Figure 5 An example block diagram of a decoder according to an embodiment is shown.

[0014] Figure 6 An example vertex set is shown, according to an embodiment.

[0015] Figure 7 An example set of codewords for encoding vertex motion vector information is shown according to an embodiment.

[0016] Figure 8 An example context set for encoding vertex motion vector information is shown in accordance with an embodiment.

[0017] Figure 9 An example of shared context among vertex motion vector information components according to an embodiment is shown.

[0018] Figure 10 An example of shared context among vertex motion vector information components according to an embodiment is shown.

[0019] Figure 11 A flowchart of an encoding process for vertex motion vector differences according to an embodiment is shown.

[0020] Figure 12 A flowchart of a decoding process for encoded vertex motion vector differences according to an embodiment is shown.

[0021] In one or more embodiments, not all of the depicted components in each figure may be required, and one or more embodiments may include additional components not shown in the figures. The arrangement and types of components may be varied without departing from the scope of this subject disclosure. Additional components, different components, or fewer components may be used within the scope of this subject disclosure. DETAILED DESCRIPTION

[0022] The detailed description set forth below in conjunction with the accompanying drawings is intended to serve as a description of various embodiments, rather than to represent the only embodiment in which the subject technology can be practiced. More specifically, in order to provide a thorough understanding of the subject matter of the present invention, the detailed description includes specific details. As will be appreciated by those skilled in the art, the described embodiments can be modified in various ways, all of which do not depart from the scope of this disclosure. Accordingly, the drawings and description are considered to be illustrative in nature, rather than restrictive. Similar reference numerals represent similar elements.

[0023] A point cloud is a collection of 3D points and properties such as color, normal, reflectivity, point size that represent the surface or volume of an object. Point clouds are commonly found in various applications such as gaming, 3D mapping, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view playback, and 6DoF immersive media. If uncompressed, point clouds generally require a large amount of bandwidth to transmit. Due to the large bitrate requirements, point clouds are often compressed before transmission. In order to compress 3D objects such as point clouds, specialized hardware is often required. To avoid specialized hardware to compress 3D point clouds, 3D point clouds can be transformed into traditional two-dimensional (2D) frames, where the traditional 2D frames can be compressed, reconstructed later, and viewed by the user.

[0024] Polygonal 3D meshes, particularly triangle meshes, are another popular format for representing 3D objects. A mesh typically consists of a collection of vertices, edges, and faces that represent the surface of a 3D object. A triangle mesh is a simple polygonal mesh where the faces are simple triangles that cover the surface of the 3D object. Typically, one or more attributes may be associated with a mesh. In one scenario, one or more attributes may be associated with each vertex in the mesh. For example, a texture attribute (RGB) may be associated with each vertex. In one scenario, each vertex may be associated with a pair of coordinates (u, v). The coordinates (u, v) may refer to a location in a texture map associated with the mesh. For example, the coordinates (u, v) may refer to a row index and a column index, respectively, in a texture map. A mesh can be thought of as a point cloud with additional connectivity information.

[0025] Point clouds or meshes can be dynamic, i.e., they can change over time. In these cases, the point cloud or mesh at a particular time can be referred to as a point cloud frame or mesh frame, respectively.

[0026] Since point clouds and meshes contain large amounts of data, they need to be compressed for efficient storage and transmission. This is especially true for dynamic point clouds and meshes, which may contain 60 frames per second or more.

[0027] The figures discussed below in this patent document and the various embodiments used to describe the principles of the present disclosure are illustrative only and should not be construed in any way to limit the scope of the present disclosure. Those skilled in the art will understand that the principles of the present disclosure can be implemented in any suitably arranged system or device.

[0028] Figure 1 An example communication system 100 is shown in accordance with an embodiment. Figure 1 The embodiment of the communication system 100 shown is for illustrative purposes only. The communication system 100 may be used in accordance with the embodiments without departing from the scope of the present disclosure.

[0029] Communication system 100 may include a network 102 that facilitates communication between various components in communication system 100. For example, network 102 may communicate Internet Protocol (IP) packets, Frame Relay frames, Asynchronous Transfer Mode (ATM) cells, or other information between network addresses. Network 102 may include all or part of one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), a global network such as the Internet, or any other communication system at one or more locations.

[0030] In the example, network 102 facilitates communication between a server 104 and various client devices 106-116. Client devices 106-116 may be, for example, smartphones, tablets, laptops, personal computers, TVs, interactive displays, wearable devices, head-mounted displays (HMDs), and the like. Server 104 may represent one or more servers. Each server 104 includes any suitable computing or processing device capable of providing computing services to one or more client devices, such as client devices 106-116. Each server 104 may include, for example, one or more processing devices, one or more memories for storing instructions and data, and one or more network interfaces for facilitating communication over network 102. As described in more detail below, server 104 may transmit a compressed bitstream representing a point cloud or mesh to one or more display devices, such as client devices 106-116. In embodiments, each server 104 may include an encoder.

[0031] Each client device 106-116 represents any suitable computing or processing device that interacts with at least one server (such as server 104) or other computing device via network 102. Client devices 106-116 include desktop computers 106, mobile phones or mobile devices 108 (such as smartphones), PDAs 110, laptops 112, tablet computers 114, and head-mounted display (HMD) 116. However, any other or additional client devices may be used in communication system 100. Smartphones represent a type of mobile device 108 that is a handheld device with a mobile operating system and integrated mobile broadband cellular network connectivity for voice, short message service (SMS), and internet data communications. HMD 116 can display a 360° scene including one or more dynamic or static 3D point clouds. In embodiments, any of client devices 106-116 may include an encoder, a decoder, or both. For example, mobile device 108 may record 3D stereoscopic video and then encode the video so that it can be transmitted to one of client devices 106-116. In an example, the laptop computer 112 may be used to generate a 3D point cloud or mesh that is then encoded and sent to one of the client devices 106 - 116 .

[0032] In the example, some client devices 108-116 communicate indirectly with the network 102. For example, mobile device 108 and PDA 110 communicate via one or more base stations 118, such as cellular base stations or eNodeBs (eNBs). Furthermore, laptop 112, tablet 114, and HMD 116 communicate via one or more wireless access points 120, such as IEEE 802.11 wireless access points. Note that this is for illustration only, and each client device 106-116 can communicate directly with the network 102 or indirectly via any suitable intermediary device or network. In an embodiment, server 104 or any client device 106-116 can be configured to compress a point cloud or mesh, generate a bitstream representing the point cloud or mesh, and send the bitstream to another client device, such as any client device 106-116.

[0033] In an embodiment, for example, any of the client devices 106-114 securely and efficiently transmits information to another device, such as the server 104. Furthermore, any of the client devices 106-116 can trigger information transfer between itself and the server 104. Any of the client devices 106-114 can function as a VR display when attached to a headset via a cradle and operate similarly to the HMD 116. For example, the mobile device 108 can operate similarly to the HMD 116 when attached to a cradle system and worn over the user's eyes. The mobile device 108 (or any of the other client devices 106-116) can trigger information transfer between itself and the server 104.

[0034] In an embodiment, any one of the client devices 106-116 or the server 104 may create a 3D point cloud or mesh, compress a 3D point cloud or mesh, transmit a 3D point cloud or mesh, receive a 3D point cloud or mesh, decode a 3D point cloud or mesh, render a 3D point cloud or mesh, or a combination thereof. For example, the server 104 may then compress the 3D point cloud or mesh to generate a bitstream and then transmit the bitstream to one or more of the client devices 106-116. For example, one of the client devices 106-116 may compress the 3D point cloud or mesh to generate a bitstream and then transmit the bitstream to another of the client devices 106-116 or the server 104.

[0035] although Figure 1 An example of a communication system 100 is shown, but Figure 1 Various changes may be made. For example, the communication system 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems come in a variety of configurations, and Figure 1 The scope of this disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may also be used in any other suitable system.

[0036] Figure 2 and Figure 3 An example electronic device according to an embodiment of the present disclosure is shown. Specifically, Figure 2 An example server 200 is shown and may represent Figure 1 The server 200 may represent one or more encoders, decoders, local servers, remote servers, cluster computers, and components that act as a single seamless resource pool, cloud-based servers, etc. The server 200 may be Figure 1 One or more of the client devices 106-116 or another server accesses the same.

[0037] Server 200 may represent one or more local servers, one or more compression servers, or one or more encoding servers (such as encoders). In an embodiment, the encoder may perform decoding. Figure 2 As shown, server 200 includes a bus system 205 that supports communication between at least one processing device (such as processor 210 ), at least one storage device 215 , at least one communication interface 220 , and at least one input / output (I / O) unit 225 .

[0038] The processor 210 executes instructions that may be stored in the memory 230. The processor 210 may include any suitable number and type of processors or other devices in any suitable arrangement. Example types of processors 210 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application specific integrated circuits, and discrete circuits.

[0039] In an embodiment, the processor 210 may encode a 3D point cloud or mesh stored in the storage device 215. In an embodiment, encoding the 3D point cloud also decodes the 3D point cloud or mesh to ensure that when the point cloud or mesh is reconstructed, the reconstructed 3D point cloud or mesh matches the 3D point cloud or mesh before encoding.

[0040] Memory 230 and persistent storage 235 are examples of storage devices 215 that represent any structure capable of storing and facilitating retrieval of information, such as data, program code, or other suitable information, whether temporary or permanent. Memory 230 may represent random access memory or any other suitable volatile or non-volatile storage device. For example, the instructions stored in memory 230 may include instructions for decomposing a point cloud into patches, instructions for packing the patches onto 2D frames, instructions for compressing the 2D frames, and instructions for encoding the 2D frames in a particular order to generate a bitstream. The instructions stored in memory 230 may also include instructions for rendering the point cloud onto an omnidirectional 360° scene, such as through a virtual reality (VR) headset such as a 360° scene. Figure 1 Persistent storage 235 may include one or more components or devices that support long-term storage of data, such as read-only memory, a hard drive, flash memory, or an optical disk.

[0041] The communication interface 220 supports communication with other systems or devices. For example, the communication interface 220 may include a Figure 1The communication interface 220 may include a network interface card or wireless transceiver for communicating with the network 102. The communication interface 220 may support communication via any suitable physical or wireless communication link. For example, the communication interface 220 may send a bitstream containing a 3D point cloud to another device (such as one of the client devices 106-116).

[0042] The I / O unit 225 allows for the input and output of data. For example, the I / O unit 225 can provide a connection for user input via a keyboard, mouse, keypad, touch screen, or other suitable input device. The I / O unit 225 can also send output to a display, printer, or other suitable output device. However, it is noted that the I / O unit 225 can be omitted, such as when I / O interaction with the server 200 is performed via a network connection.

[0043] Note that although Figure 2 Described as indicating Figure 1 The server 104 of FIG. 106 may be a server 104, but the same or similar structure may also be used for one or more of the various client devices 106-116. For example, a desktop computer 106 or a laptop computer 112 may have a server 104 that is configured to connect to the server 104. Figure 2 The same or similar structures shown.

[0044] Figure 3 An example electronic device 300 is shown and may represent Figure 1 For example, the electronic device 300 may be a mobile communication device such as a mobile station, a subscriber station, a wireless terminal, a desktop computer (similar to a Figure 1 Desktop computer 106), portable electronic device (similar to Figure 1 In an embodiment, Figure 1 One or more of the client devices 106-116 may include the same or similar configuration as the electronic device 300. In an embodiment, the electronic device 300 is an encoder, a decoder, or both. For example, the electronic device 300 may be used for data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.

[0045] like Figure 3As shown, electronic device 300 includes antenna 305, radio frequency (RF) transceiver 310, transmit (TX) processing circuitry 315, microphone 320, and receive (RX) processing circuitry 325. RF transceiver 310 may include, for example, an RF transceiver, a Bluetooth transceiver, a Wi-Fi transceiver, a Zigbee transceiver, an infrared transceiver, and various other wireless communication signals. Electronic device 300 also includes speaker 330, processor 340, input / output (I / O) interface (IF) 345, input 350, display 355, memory 360, and sensor 365. Memory 360 includes an operating system (OS) 361 and one or more applications 362.

[0046] The RF transceiver 310 receives incoming RF signals from the antenna 305, transmitted by an access point (such as a base station, a Wi-Fi router, or a Bluetooth device) or other devices on the network 102 (such as Wi-Fi, Bluetooth, cellular, 5G, LTE, LTE-A, WiMAX, or any other type of wireless network). The RF transceiver 310 downconverts the incoming RF signal to generate an intermediate frequency (IF) or baseband signal. The IF or baseband signal is sent to the RX processing circuitry 325, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or IF signal. The RX processing circuitry 325 sends the processed baseband signal to the speaker 330 (such as for voice data) or the processor 340 (such as for web browsing data) for further processing.

[0047] The TX processing circuitry 315 receives analog or digital voice data from the microphone 320, or other outgoing baseband data from the processor 340. The outgoing baseband data may include web data, email, or interactive electronic game data. The TX processing circuitry 315 encodes, multiplexes, and / or digitizes the outgoing baseband data to generate a processed baseband or intermediate frequency signal. The RF transceiver 310 receives the outgoing processed baseband or intermediate frequency signal from the TX processing circuitry 315 and up-converts the baseband or intermediate frequency signal into an RF signal that is transmitted via the antenna 305.

[0048] Processor 340 may include one or more processors or other processing devices. Processor 340 may execute instructions stored in memory 360 (such as OS 361) to control the overall operation of electronic device 300. For example, according to well-known principles, processor 340 may control the reception of forward channel signals and the transmission of reverse channel signals via RF transceiver 310, RX processing circuit 325, and TX processing circuit 315. Processor 340 may include any suitable number and type of processors or other devices in any suitable arrangement. For example, in an embodiment, processor 340 includes at least one microprocessor or microcontroller. Example types of processor 340 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application-specific integrated circuits, and discrete circuits.

[0049] The processor 340 is also capable of executing other processes and programs residing in the memory 360, such as operations for receiving and storing data. The processor 340 can move data into or out of the memory 360 as needed by the executed process. In an embodiment, the processor 340 is configured to execute one or more applications 362 based on the OS 361 or in response to signals received from an external source or operator. For example, the applications 362 may include encoders, decoders, VR or AR applications, camera applications (for still images and video), video phone calling applications, email clients, social media clients, SMS messaging clients, virtual assistants, etc. In an embodiment, the processor 340 is configured to receive and send media content. In an embodiment, the processor 340 may utilize improved vertex motion vector encoding or decoding as described in the present disclosure.

[0050] Processor 340 is also coupled to I / O interface 345 , which provides electronic device 300 with the ability to connect to other devices, such as client devices 106 - 114 . I / O interface 345 is the communication path between these accessories and processor 340 .

[0051] Processor 340 is also coupled to input 350 and display 355. An operator of electronic device 300 can use input 350 to enter data or input into electronic device 300. Input 350 can be a keyboard, touch screen, mouse, trackball, voice input, or other device capable of serving as a user interface to allow the user to interact with electronic device 300. For example, input 350 can include voice recognition processing, allowing the user to enter voice commands. In some examples, input 350 can include a touch panel, a (digital) pen sensor, a keypad, or an ultrasonic input device. A touch panel can recognize touch input using at least one scheme, such as capacitive, pressure-sensitive, infrared, or ultrasonic. By providing additional inputs to processor 340, input 350 can be associated with sensors 365 and / or a camera. In embodiments, sensors 365 include one or more inertial measurement units (IMUs) (such as accelerometers, gyroscopes, and magnetometers), motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, and the like. Input 350 may also include control circuitry. In a capacitive scheme, the input 350 may recognize touch or proximity.

[0052] Display 355 can be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED), an active matrix OLED (AMOLED), or other display capable of rendering text and / or graphics (such as from a website, video, game, image, etc.). Display 355 can be sized to fit within the HMD. Display 355 can be a single display screen or multiple displays capable of creating a stereoscopic display. In an embodiment, display 355 is a heads-up display (HUD). Display 355 can display 3D objects, such as a 3D point cloud or mesh.

[0053] Memory 360 is coupled to processor 340. A portion of memory 360 may include RAM, and another portion of memory 360 may include flash memory or other ROM. Memory 360 may include a persistent storage device (not shown) representing any structure capable of storing and facilitating retrieval of information (such as data, program code, and / or other suitable information). Memory 360 may include one or more components or devices that support long-term storage of data, such as read-only memory, a hard drive, flash memory, or an optical disk. Memory 360 may also include media content. Media content may include various types of media, such as images, videos, three-dimensional content, VR content, AR content, 3D point clouds, meshes, and the like.

[0054] The electronic device 300 further includes one or more sensors 365, which can measure physical quantities or detect the activation state of the electronic device 300 and convert the measured or detected information into electrical signals. For example, the sensors 365 may include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor (such as a gyroscope or a gyroscopic sensor and an accelerometer), an eye tracking sensor, an air pressure sensor, a magnetic sensor or a magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illumination sensor, an ultraviolet (UV) sensor, an electromyography (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an IR sensor, an ultrasonic sensor, an iris sensor, a fingerprint sensor, a color sensor (such as a red, green, and blue (RGB) sensor), etc. The sensors 365 may further include control circuitry for controlling any of the sensors included therein.

[0055] As discussed in more detail below, one or more of these sensors 365 can be used to control a user interface (UI), detect UI input, determine the orientation and facing the direction of the user for three-dimensional content display recognition, etc. Any of these sensors 365 can be located within the electronic device 300, within an auxiliary device operably connected to the electronic device 300, within a headset configured to hold the electronic device 300, or within a single device including the electronic device 300 and the headset.

[0056] The electronic device 300 can create media content (such as a virtual object) or capture (or record) content through a camera. The electronic device 300 can encode the media content to generate a bit stream so that the bit stream can be sent directly to another electronic device or transmitted to another electronic device such as a video camera. Figure 1 The electronic device 300 may receive the bit stream directly from another electronic device, or such as through Figure 1 The network 102 receives the bit stream indirectly from another electronic device.

[0057] although Figure 2 and Figure 3 An example of an electronic device is shown, but it is also possible to Figure 2 and Figure 3 Make various changes. For example, you can combine, further subdivide or omit Figure 2 and Figure 3 As a specific example, processor 340 can be divided into multiple processors, such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs). In addition, just like computing and communications, electronic devices and servers can have a variety of configurations, and Figure 2 and Figure 3 The present disclosure is not limited to any particular electronic device or server.

[0058] Various standards for video-based compression of dynamic meshes have been proposed. For example, the first vertex mesh test model (vmesh-v1.0) was established in ISO / IEC SC29 WG07 in July 2022. The following documents are hereby incorporated by reference in their entirety into this disclosure as if fully set forth herein:

[0059] “V-Mesh Test Model v1, ISO / IEC SC29 WG07 N00404, July 2022”.

[0060] Figure 4 An example block diagram of an inter-frame encoder according to an embodiment is shown. Figure 4 The inter-frame encoder 400 is shown for illustrative purposes only. Figure 4The scope of this disclosure is not limited to any particular implementation of the inter-frame encoder and inter-frame encoding process. The basic idea is that a base mesh, typically having a smaller number of vertices than the original mesh, is created and compressed in a lossy or lossless manner. The reconstructed base mesh undergoes subdivision, and then the displacement field between the original mesh and the subdivided reconstructed base mesh is calculated. In inter-frame coding of mesh frames, the base mesh is encoded by transmitting vertex motions rather than directly compressing the base mesh.

[0061] like Figure 4 As shown, the inter-frame encoder 400 may include a quantizer 401, a motion encoder 403, a motion decoder 404, a basic grid reconstruction module 405, a displacement updater 407, a wavelet transformer 409, a quantizer 411, an image packer 413, a video encoder 415, an image unpacker 417, an inverse quantizer 419, an inverse wavelet transformer 421, an inverse quantizer 423, a deformed grid reconstructor 425, an attribute transfer module 427, a padding module 429, a color space converter 431, a video encoder 433, a multiplexer 435 and a controller 437.

[0062] The quantizer 401 may quantize the base grid m(i) to generate a quantized base grid. The motion encoder 403 may encode the quantized base grid to generate a compressed motion bitstream. The motion decoder 404 may decode the compressed motion bitstream to generate a reconstructed motion field, and the base grid reconstruction module 405 may generate a quantized base grid m'(i). The displacement updater 407 may update the displacement d(i) based on the base grid m(i) and the reconstructed quantized base grid m'(i) to generate an updated displacement d'(i). The reconstructed base grid may be subdivided, and then a displacement field between the original grid and the subdivided reconstructed base grid may be calculated.

[0063] The wavelet transformer 409 can perform a wavelet transform using the updated displacement d'(i) to generate a wavelet coefficient e(i). The quantizer 411 can quantize the wavelet coefficient e(i) to generate a quantized wavelet coefficient e'(i). The image packer 413 can pack the quantized wavelet coefficient e'(i) into a 2D image to generate a packed quantized wavelet coefficient. The video encoder 415 can encode the packed quantized wavelet coefficient to generate a compressed displacement bitstream. The image depacketizer 417 can depacketize the packed quantized wavelet coefficient to generate an unpacked quantized wavelet coefficient. The inverse quantizer 419 can dequantize the quantized wavelet coefficient to generate a wavelet coefficient. In an embodiment, an alternative entropy encoder (such as an encoder based on arithmetic coding) can also be used to encode the quantized wavelet coefficient. The inverse wavelet transformer 421 can perform an inverse wavelet transform using the wavelet coefficient to generate a reconstructed displacement D''(i). The inverse quantizer 423 can inverse quantize the reconstructed quantized base grid m'(i) to generate a reconstructed base grid m''(i). The deformed grid reconstructor 425 can generate a reconstructed deformed grid DM(i) based on the reconstructed displacement D''(i) and the reconstructed base grid m''(i). The attribute transfer module 427 can update the attribute map A(i) based on the static / dynamic grid m(i) and the reconstructed deformed grid DM(i) to generate an updated attribute map A'(i). The attribute map can be a texture map, but other attributes can also be sent. The padding module 429 can perform padding to fill in the blank areas in the updated attribute map A'(i) to remove high-frequency components. The color space converter 431 can perform color space conversion on the padded updated attribute map A'(i). The video encoder 433 can encode the output of the color space converter 431 to generate a compressed attribute bitstream. Multiplexer 435 can multiplex header information (tile information), compressed grid bitstream, compressed displacement bitstream, and compressed attribute bitstream to generate compressed bitstream b(i). Controller 437 can control the modules of encoder 400, including quantizer 401 and quantizer 411, image packer 413, video encoder 415, attribute transfer module 427, padding module 429, and video encoder 433.

[0064] Figure 5 An example block diagram of a decoder according to an embodiment is shown. Figure 5 Decoder 500 is shown for illustrative purposes only. Figure 5 The scope of this disclosure is not limited to any particular implementation of the decoder or decoding process.

[0065] like Figure 5As shown, the decoder 500 may include a demultiplexer 501, a base grid decoder 510, a video decoder 521, an image depacketizer 523, an inverse quantizer 525, an inverse wavelet transformer 527, a deformed grid reconstructor 529, a video decoder 531, and a color space converter 533. The base grid decoder 510 may include a switch 503, a static grid decoder 505, a grid buffer 507, a motion decoder 509, a base grid reconstructor 511, a switch 513, and an inverse quantizer 515.

[0066] The demultiplexer 501 can receive a compressed bitstream b(i) from the encoder 400 to extract a compressed motion bitstream, a compressed displacement bitstream, and a compressed attribute bitstream from the compressed bitstream b(i). The switch 503 can determine whether the compressed bitstream has inter-frame coded grid frame data or intra-frame coded grid frame data. If the compressed bitstream has inter-frame coded grid frame data, the switch 503 can pass the inter-frame coded grid frame data to the motion decoder 509. If the compressed base grid bitstream has intra-frame coded grid frame data, the switch 503 can pass the intra-frame coded grid frame data to the static grid decoder 505. The static grid decoder 505 can decode the intra-frame coded grid frame data to generate a reconstructed quantized base grid frame. The grid buffer 507 can store the reconstructed quantized base grid frame and the inter-frame coded grid frame data for future use in decoding subsequent inter-frame coded grid frames. The reconstructed quantized base grid frame can be used as a reference grid frame. The motion decoder 509 can obtain a motion vector for the current inter-coded trellis frame based on the data stored in the trellis buffer 507 and syntax elements in the bitstream for the current inter-coded trellis frame. In an embodiment, the syntax elements in the bitstream for the current inter-coded trellis frame can be motion vector differences. The base trellis reconstructor 511 can generate a reconstructed quantized base trellis frame based on the motion vector for the current inter-coded trellis frame using syntax elements in the bitstream for the current inter-coded trellis frame. The reconstructed quantized base trellis frame from the base trellis reconstructor 511 is stored in the trellis buffer 507. If the compressed base trellis bitstream contains intra-coded trellis frame data, the switch 513 can transmit the reconstructed quantized base trellis frame from the static trellis decoder 505 to the inverse quantizer 515. If the compressed base trellis bitstream contains inter-coded trellis frame data, the switch 513 can transmit the reconstructed quantized base trellis frame from the base trellis reconstructor 511 to the inverse quantizer 515. The inverse quantizer 515 may perform inverse quantization using the reconstructed quantized base grid frame to generate a reconstructed base grid frame m''(i).

[0067] The video decoder 521 can decode the displacement bitstream to generate packed quantized wavelet coefficients. The image unpacker 523 can unpack the packed quantized wavelet coefficients to generate quantized wavelet coefficients. In an embodiment, an entropy decoder (such as a decoder based on arithmetic coding) can decode the quantized wavelet coefficients that have been arithmetic encoded. The inverse quantizer 525 can perform inverse quantization using the quantized wavelet coefficients to generate wavelet coefficients. The inverse wavelet transformer 527 can perform an inverse wavelet transform using the wavelet coefficients to generate displacements. The deformed grid reconstructor 529 can reconstruct the deformed grid based on the displacement and the reconstructed base grid frame m''(i). The video decoder 531 can decode the attribute bitstream to generate an attribute map before color space conversion. The color space converter 533 can perform color space conversion on the attribute map from the video decoder 531 to reconstruct the attribute map.

[0068] Figure 6 An example vertex set is shown, according to an embodiment. Figure 6 The set of vertices shown is for illustrative purposes only and does not limit the scope of the present disclosure to any particular implementation.

[0069] See Figure 6 , an example vertex set includes five vertices, such as vertex A, vertex B, vertex C, vertex D, and vertex E. When inter-frame prediction is enabled, for Figure 6 For a given vertex A shown, a flag or other means can be used to indicate whether the vertex motion vector of A is sent or the delta difference between the vertex motion vector of A and its predicted value is sent. The predicted value of the vertex motion vector can be calculated as the average of the vertex motion vectors of the adjacent vertices, for example. Figure 6 In the example of , the predicted value of the vertex motion vector of A can be calculated as the average of the vertex motion vectors of the neighboring vertices B, C, D and E.

[0070] In this disclosure, the incremental difference between a vertex motion vector and its predicted value may be referred to as a "vertex motion vector difference," a "vertex motion vector residual," or a "vertex motion vector prediction residual." For simplicity of explanation, vertex motion vectors, vertex motion vector differences, vertex motion vector residuals, and vertex motion vector prediction residuals may be collectively referred to as "vertex motion vector information" in this disclosure.

[0071] Vertex motion vector information (e.g., vertex motion vectors or vertex motion vector differences) may be transmitted using a combination of unary codes and exponential-Golomb codes during arithmetic coding. In an embodiment, the vertex motion vector information may be binarized using a combination of unary codes and exponential-Golomb codes before arithmetic coding is performed.

[0072] Figure 7 An example set of codewords for encoding vertex motion vector information is shown according to an embodiment. Figure 7 The example codeword sets shown are for illustrative purposes only and do not limit the scope of the present disclosure to any particular implementation.

[0073] exist Figure 7 In the example of , vertex motion vector information is binarized using a combination of unary code and Exponential-Golomb code. Figure 7 , the first and second columns 710 represent the unary code portion, and the remaining columns 720 represent the Exponential-Golomb code portion. In this example, the sign bit is encoded separately. In an embodiment, multiple context memories are used for the bins (or bits) of the codeword. When using context-based coding, the probability of a bin (or bit) taking a particular value can be predicted based on a probability model, which can also be referred to as a context model. The context for a bin can depend on the value of the relevant bin of a previously encoded syntax element, such as a motion vector difference.

[0074] Figure 8 An example context set for encoding vertex motion vector information is shown in accordance with an embodiment. Figure 8 The example context sets shown are also for illustrative purposes only and do not limit the scope of the present disclosure to any particular implementation.

[0075] exist Figure 8 In the example, the vertex motion vector information includes three components (X=15, Y=12 and Z=9), which are Figure 78. The first and second bits 810 in the codeword represent the unary code portion, while the third through sixth bits 820 represent the prefix portion of the Exponential-Golomb code. The seventh through ninth bits 830 represent the suffix portion of the Exponential-Golomb code. In this example, each bit of the vertex motion vector information component (X=15, Y=12, and Z=9) is encoded using its dedicated context. Therefore, each component of the vertex motion vector information requires nine context memories. Specifically, component X uses contexts (C0-C8), component Y uses contexts (C9-C17), and component Z uses contexts (C18-C26). A total of 27 contexts (C0-C26) are used to encode all three components of the vertex motion vector information (X=15, Y=12, and Z=9).

[0076] In an embodiment, some context can be shared among the bits of vertex motion vector information components as a means of reducing the amount of context memory. Vertex motion vector information components are typically either all zero or predominantly non-zero. Therefore, there is correlation between the X, Y, and Z components, which becomes apparent in the most significant bits (e.g., prefix bits) of components with a probability greater than 0.5. This can improve compression efficiency by sharing context among the most significant bits of the X, Y, and Z components. On the other hand, the X, Y, and Z components of one vertex motion vector information are typically different from those of adjacent motion vector information. Therefore, there is typically less correlation among the least significant bits (e.g., suffix bits) of the vertex motion vector information, and using context memory for these bits does not improve compression efficiency. Therefore, bypass coding, which offers the benefit of reducing the amount of context memory, can be used for the least significant bits.

[0077] Figure 9 An example of shared context among vertex motion vector information components according to an embodiment is shown. Figure 9 The examples shown are also for illustrative purposes only and do not limit the scope of the present disclosure to any particular implementation.

[0078] exist Figure 9 In, with Figure 8Similarly, the three motion vector information components (X=15, Y=12, and Z=9) are depicted. In this example, the context (C2-C8) is shared among the exponential-Golomb coded portions of the vertex motion vector information components. Specifically, the first and second bins 810 representing the unary portion are encoded using their own contexts (C0, C1, and C9-C12), while the remaining bins 820 and 830 representing the exponential-Golomb coded portion share the context (C2-C8) among the vertex motion vector information components. Therefore, the context memory for C2-C8 can be shared among the exponential-Golomb coded portions of the three motion vector information components.

[0079] In an example, the encoder 400 may encode the unary code portion 810 of the vertex motion vector information component (X=15, Y=12, and Z=9) using contexts (C0-C1, C9-C10, and C11-C12). As for the Exponent-Golomb code portions 820 and 830, the encoder 400 may encode the third binary bit of the motion vector information component (X=15, Y=12, and Z=9) using context C2. Similarly, the encoder 400 may encode the fourth binary bit of the motion vector difference component using context C3, the fifth binary bit of the motion vector difference component using context C4, the sixth binary bit of the motion vector difference component using context C5, the seventh binary bit of the motion vector difference component using context C6, the eighth binary bit of the motion vector difference component using context C7, and the ninth binary bit of the motion vector difference component using context C8. A total of 13 contexts (C0-C12) are used to encode the vertex motion vector information, thereby encoding the vertex motion vector information. Figure 8 Compared with the example of , 14 contexts are reduced.

[0080] Figure 10 An example of shared context among vertex motion vector information components according to an embodiment is shown. Figure 10 The examples shown are also for illustrative purposes only and do not limit the scope of the present disclosure to any particular implementation.

[0081] exist Figure 10 In, with Figure 8 and Figure 9Similarly, three motion vector difference components (X=15, Y=12, and Z=9) are depicted in FIG. The first and second bins 810 represent the unary code portion, while the third through sixth bins 820 represent the prefix portion of the Exponential-Golomb code, and the seventh through ninth bins 830 represent the suffix portion of the Exponential-Golomb code. In this example, the unary code portion (i.e., the first and second bins) 810 are encoded using their own contexts (C0-C1, C4-C5, and C6-C7), the prefix portion (i.e., the third through sixth bins) 820 are encoded by sharing the context (C2-C3) with the prefix portion of other motion vector information components, and the suffix portion (i.e., the seventh through ninth bins) 830 are bypass-coded. Bypass coding does not require context memory. In other words, the context memory for C2-C3 is shared among the prefix portion of the vertex motion vector information components, while the suffix portion of the vertex motion vector information components is bypass-coded.

[0082] In the above example, the encoder 400 can use the contexts (C0-C1, C4-C5, and C6-C7) to encode the unary code portion 810 of the vertex motion vector information component (X=15, Y=12, and Z=9). As for the prefix portion 820 of the Exponential-Golomb code, the encoder 400 can use context C2 to encode the first prefix binary bit (the third binary bit in the codeword) of the vertex motion vector information component and use context C3 to encode the remaining prefix binary bits (the fourth binary bit to the sixth binary bit in the codeword). Then, the suffix binary bits (the seventh binary bit to the ninth binary bit in the codeword) 830 of the vertex motion vector information component are bypass-encoded. In the example, a total of 8 contexts (C0-C7) are used to encode the vertex motion vector information, thereby Figure 8 The number of contexts is reduced by 19 compared to the example.

[0083] Table 1 shows example values of context for a coded syntax element according to an embodiment. The coded syntax element may be coded vertex motion vector information, such as coded vertex motion vector differences. In an embodiment, the decoder 500 receives a compressed bitstream including coded syntax elements (e.g., coded vertex motion vector differences) and parses the compressed bitstream. To decode the coded syntax element, the decoder 500 may select a context for each binary bit of the coded syntax element based on the information in Table 1. Figure 10 For example, the syntax elements of Table 1 can be encoded by the encoder 400.

[0084] [Table 1]

[0085]

[0086] In Table 1, the syntax element "sismu_mv_residual_abs_rem[k]" (k = 0, 1, 2) represents the exponential-Golomb code portion of the vertex motion vector difference component. In some embodiments, it may indicate the absolute value of the k-th component of the vertex motion vector prediction residual associated with the vertex having the submesh.

[0087] "Contexts[ctxTbl][ctxIdx]" represents the context table for the coded syntax element. The values of CtxTbl and CtxIdx can be determined based on the entries related to the syntax element. The parameters "ctxTbl" and "ctxIdex" represent the context table and context index for the syntax element, respectively. In the example of Table 1, the context for "sismu_mv_residual_abs[k]" can be stored in, for example, ctxTbl=5.

[0088] In this example, Contexts[ctxTbl][0] (i.e., ctxIdx = 0) contains context information for the first binary bit (i.e., the first prefix binary bit) of the prefix of "sismu_mv_residual_abs_rem[k]" (k = 0, 1, 2). Contexts[ctxTbl][1] (i.e., ctxIdx = 1) contains context information for the remaining binary bits (i.e., the remaining prefix binary bits) of the prefix of "sismu_mv_residual_abs_rem[k]". The suffix of "sismu_mv_residual_abs_rem[k]" is encoded using bypass mode. Therefore, in each vertex motion vector differential component, the first prefix binary bit is encoded using its own context, while the remaining prefix binary bits are encoded using the same context. In addition, the suffix binary bits are bypass coded and do not use the context. As a result, for the syntax element "sismu_mv_residual_abs_rem[k]" (k = 0, 1, 2) in Table 1, only two (2) contexts are employed.

[0089] In some embodiments, the remaining bins of the prefix of "sismu_mv_residual_abs_rem[k]" may also be encoded using bypass coding to further reduce the number of contexts.

[0090] In some embodiments, up to the first n bins of the prefix part may be encoded using their dedicated context, while the remaining bins (>n) of the prefix part may be encoded using the same context, and the suffix bins are encoded using bypass coding.

[0091] Figure 11 A flow chart of an encoding process 1100 for vertex motion vector differences according to an embodiment is shown. Although one or more operations are described or shown in a particular sequential order, in embodiments, the operations may be rearranged in a different order, which may include performing multiple operations in at least partially overlapping time periods.

[0092] Process 1100 may begin at operation 1101. In operation 1101, process 1100 binarizes vertex motion vector differences comprising three components (k=0, 1, 2) using a combination of unary code and Exponential-Golomb code. Process 1100 then initiates the encoding process starting with the first component (k=0) of the vertex motion vector difference.

[0093] In operation 1103 , the process 1100 encodes the unary code bins of the vertex motion vector differential components (k=0) using their own contexts. The process 1100 then proceeds to operation 1105 .

[0094] In operation 1105, process 1100 encodes the first prefix n bits of the Exponential-Golomb code of the vertex motion vector differential component (k=0) using their own context. Figure 10 In the example of , the first prefix binary bits of the Exponent-Golomb code of the motion vector difference component (k=0) are encoded using the first context (C2).

[0095] In operation 1107, process 1100 encodes the remaining prefix bins of the Exponential-Golomb code of the vertex motion vector difference component (k=0) using the same context. The same context is shared among all remaining prefix bins of this component (k=0). Figure 10 In the example of , the remaining prefix bins of the component (k=0) are encoded using the same context C3. The process then proceeds to operation 1109.

[0096] In operation 1109, process 1100 encodes the suffix of the Exponential-Golomb code of the vertex motion vector differential component (k=0) using bypass coding. Figure 10 In the example of , the suffix bins of the component (k=0) are bypass-coded. Therefore, the context memory is not utilized. Then, the process 1100 proceeds to operation 1111.

[0097] In operation 1111 , process 1100 determines whether k is less than 2. When k is less than 2, process 1100 proceeds to operation 1113 .

[0098] In operation 1113, the process updates the k value.Then, the process 1100 repeats operations 1103 to 1109 for other vertex motion vector differential components (k=1, 2).

[0099] When the encoding process for all components (k=0, 1, 2) is completed, process 1100 proceeds to operation 1115 .

[0100] In operation 115 , process 1100 generates a compressed bitstream using the encoded motion vector differences including code components (k=0, 1, 2).

[0101] In some embodiments, all bins of the unary code portion may be encoded using different contexts without sharing the context among other components, while the context employed for the prefix bins may be shared among other components. Figure 10 In the example shown in Figure 2, all first prefix bits of the vertex motion vector differential components (X, Y, Z) are encoded using the first context (C2). Additionally, all remaining prefix bits of the motion vector differential components (X, Y, Z) are encoded using the second context (C3).

[0102] Figure 12 A flowchart of a decoding process 1200 for encoded vertex motion vector differences is shown in accordance with an embodiment. Although one or more operations are described or shown in a particular sequential order, in embodiments, the operations may be rearranged in a different order, which may include performing multiple operations in at least partially overlapping time periods.

[0103] Process 1200 may begin at operation 1201. In operation 1201, process 1200 receives a compressed bitstream and parses at least a portion of the compressed bitstream. Process 1200 identifies coded vertex motion vector differences comprising three components (k=0, 1, 2) from the received bitstream. Process 1200 then initiates a decoding process starting with the first component (k=0) of the coded vertex motion vector differences.

[0104] In operation 1203, process 1200 decodes the unary coded bits of the vertex motion vector differential component (k=0) using their own context. A context is selected for each unary coded bit of the encoded vertex motion vector differential component (k=0). Process 1200 then proceeds to operation 1205.

[0105] In operation 1205, the process 1200 decodes the first prefix n bits of the Exponent-Golomb code of the encoded vertex motion vector difference component (k=0) using their own context. A context is selected for each of the first prefix n bits of the encoded motion vector difference component (k=0). Figure 10 In the example of , the first prefix bin of the Exponential-Golomb code of the motion vector difference component (k=0) is encoded using the first context (C2). Then, the process 1200 proceeds to operation 1207.

[0106] In operation 1207, the process 1200 decodes the remaining prefix bins of the Exponential-Golomb code of the vertex motion vector difference component (k=0) using the same context. The same context is shared among all remaining prefix bins of the component (k=0). The context is selected for all remaining prefix bins of the coded motion vector difference component (k=0). Figure 10 For example, the remaining bins of the Exponential-Golomb code of the component (k=0) are encoded using the same context C3. Then, the process 1200 proceeds to operation 1209.

[0107] In operation 1209, the process 1200 decodes the suffix of the Exponential-Golomb code of the encoded vertex motion vector difference component (k=0). The suffix bits of the motion vector difference component (k=0) are bypass-coded. Therefore, no context needs to be selected for the suffix bits. Figure 10 In the example of , all suffix bits of the Exponential-Golomb code of the component (k=0) are bypass-coded.

[0108] In operation 1211 , process 1200 determines whether k is less than 2. When k is less than 2, process 1200 proceeds to operation 1213 .

[0109] In operation 1213, process 1200 updates the value of k. Process 1200 then repeats operations 1203 to 1209 for other encoded vertex motion vector differential components (k=1, 2).

[0110] When the decoding process for all components (k=0, 1, 2) is completed, the process proceeds to operation 1215 .

[0111] In operation 1215 , process 1200 forms decoded vertex motion vector differences including code components (k=0, 1, 2).

[0112] One aspect of the present disclosure provides an apparatus comprising a communication interface configured to receive a compressed bitstream comprising vertex motion vector information. The vertex motion vector information may comprise one or more components. The apparatus may comprise a processor operably coupled to the communication interface. The processor may be configured to parse the compressed bitstream comprising the vertex motion vector information. The processor may be configured to select a context for a plurality of bins of a first component of the vertex motion vector information. A first prefix bin of the first component may be encoded based on a first context, and one or more remaining bins of the first component may be encoded using bypass coding. The processor may be configured to decode the first component of the vertex motion vector information based on the selected context.

[0113] According to an embodiment, the remaining prefix bins of the first component may be encoded based on the second context.

[0114] According to an embodiment, the processor may be further configured to select a context for a plurality of bins of the one or more second components of the vertex motion vector information. First prefix bins of the one or more second components may be encoded based on the first context, and one or more remaining bins of the one or more second components may be encoded using bypass coding. The processor may be further configured to decode the one or more second components of the vertex motion vector information based on the selected context.

[0115] According to an embodiment, the remaining prefix bins of the first component and the one or more second components may be encoded based on the second context.

[0116] According to an embodiment, the first context may be shared among the first prefix bins of the first component and one or more second components.The second context may be shared among the remaining prefix bins of the first component and one or more second components.

[0117] According to an embodiment, one or more components of the vertex motion vector information may be binarized using a combination of a unary code including a prefix part and a suffix part and an Exponential-Golomb code.

[0118] According to an embodiment, all suffix bins of the first component and the one or more second components may be encoded using bypass coding.

[0119] One aspect of the present disclosure provides a method that includes receiving a compressed bitstream including vertex motion vector information. The vertex motion vector information may include one or more components. The method may include parsing the compressed bitstream including the vertex motion vector information. The method may include selecting a context for a plurality of bins of a first component of the vertex motion vector information. First prefix bins of the first component may be encoded based on a first context, and one or more remaining bins of the first component may be encoded using bypass coding. The method may include decoding the first component of the vertex motion vector information based on the selected context.

[0120] According to an embodiment, the remaining prefix bins of the first component may be encoded based on the second context.

[0121] According to an embodiment, the method may further include selecting a plurality of contexts for the one or more second components of the vertex motion vector information. First prefix bits of the one or more second components may be encoded based on the first context, and one or more remaining bits of the one or more second components may be encoded using bypass coding. The method may further include decoding the one or more second components of the vertex motion vector information based on the selected context.

[0122] According to an embodiment, the remaining prefix bins of the first component and the one or more second components may be encoded based on the second context.

[0123] According to an embodiment, the first context may be shared among the first prefix bins of the first component and the one or more second components, and the second context may be shared among the remaining prefix bins of the first component and the one or more second components.

[0124] According to an embodiment, one or more components of the vertex motion vector information may be binarized using a combination of a unary code including a prefix part and a suffix part and an Exponential-Golomb code.

[0125] According to an embodiment, all suffix bins of the first component and the one or more second components may be encoded using bypass coding.

[0126] One aspect of the present disclosure provides an apparatus comprising a communication interface and a processor operably coupled to the communication interface. The processor may be configured to generate vertex motion vector information comprising one or more components. The processor may be configured to encode a first prefix bit of a first component of the vertex motion vector information based on a first context. The processor may be configured to encode one or more remaining bits of the first component using bypass coding. The processor may be configured to form a bitstream comprising the encoded vertex motion vector information. The processor may be configured to transmit the bitstream to a decoding device via the communication interface.

[0127] According to an embodiment, the processor may be further configured to cause the remaining prefix bins of the first component to be encoded based on the second context.

[0128] According to an embodiment, the processor may be further configured to cause first prefix bins of one or more second components of the vertex motion vector information to be encoded based on the first context, and the processor may be further configured to cause one or more remaining bins of the one or more second components to be encoded using bypass coding.

[0129] According to an embodiment, the processor may be further configured to cause the remaining prefix bins of the first component and the one or more second components to be encoded based on the second context.

[0130] According to an embodiment, the first context may be shared among the first prefix bins of the first component and one or more second components.The second context may be shared among the remaining prefix bins of the first component and one or more second components.

[0131] According to an embodiment, one or more components of the vertex motion vector information may be binarized using a combination of a unary code including a prefix part and a suffix part and an Exponential-Golomb code.

[0132] According to an embodiment, all suffix bins of the first component and one or more components may be encoded using bypass coding.

[0133] In some embodiments, all bins of a unary code portion may be decoded using different contexts without sharing the context among other components, whereas the context selected for the prefix bins may be shared among the other components.

[0134] An element mentioned in the singular is not intended to mean one and only one (unless specifically stated otherwise), but rather to mean one or more. For example, "a" module may refer to one or more modules. Without further limitation, an element beginning with "a," "an," "the," or "said" does not exclude the presence of additional identical elements.

[0135] Headings and subheadings (if any) are used for convenience only and do not limit the invention. The word "exemplary" is used to mean serving as an example or illustration. To the extent that the terms "include," "have," and the like are used, such terms are intended to be inclusive in the same manner as the term "comprise" is interpreted when used as a transitional term in a claim. Relational terms such as first and second may be used to distinguish one entity or action from another without necessarily requiring or implying any actual such relationship or order between those entities or actions. The term "couple" and its derivatives refer to any direct or indirect communication between two or more elements, regardless of whether those elements are in physical contact with one another. The terms "transmit," "receive," and "communicate," and their derivatives, encompass both direct and indirect communication. The term "controller" means any device, system, or portion thereof that controls at least one operation. Such a controller may be implemented in hardware or a combination of hardware and software and / or firmware. The functionality associated with any particular controller may be centralized or distributed, whether locally or remotely.

[0136] Phrases such as on the one hand, this aspect, on the other hand, some aspects, one or more aspects, an embodiment, this embodiment, another embodiment, some embodiments, one or more embodiments, an example, this example, another example, an example, one or more examples, a configuration, this configuration, another configuration, some configurations, one or more configurations, the subject technology, the disclosure, the present disclosure, and variations thereof are used for convenience and do not imply that the disclosure associated with such phrases is essential to the subject technology or that such disclosure applies to all configurations of the subject technology. The disclosure associated with such phrases may apply to all configurations or one or more configurations. The disclosure associated with such phrases may provide one or more examples. Phrases such as an aspect or some aspects may refer to one or more aspects and vice versa, and this applies similarly to the other aforementioned phrases.

[0137] The phrase "at least one of" preceding a list of items, and the terms "and" or "or" used to separate any items, modifies the entire list, not each member of the list. The phrase "at least one of" does not require selection of at least one item; rather, the phrase allows for the inclusion of at least one of any one item, and / or at least one of any combination of items, and / or at least one of each item. For example, each of the phrases "at least one of A, B, and C" or "at least one of A, B, or C" means: only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0138] It should be understood that the specific order or hierarchy of steps, operations, or processes disclosed is an illustration of an exemplary method. Unless expressly stated otherwise, it should be understood that the specific order or hierarchy of steps, operations, or processes can be performed in a different order. Some of these steps, operations, or processes can be performed simultaneously or as part of one or more other steps, operations, or processes. The accompanying method claims (if any) present elements of the various steps, operations, or processes in a sample order and are not intended to be limited to the specific order or hierarchy presented. These can be performed serially, linearly, in parallel, or in a different order. It should be understood that the described instructions, operations, and systems can generally be integrated together in a single software / hardware product or packaged into multiple software / hardware products.

[0139] This disclosure is provided to enable anyone skilled in the art to practice the various aspects described herein. In some cases, to avoid blurring the concepts of the subject technology, well-known structures and components are shown in the form of block diagrams. This disclosure provides various examples of the subject technology, and the subject technology is not limited to these examples. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles described herein may be applied to other aspects.

[0140] All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later become known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is expressly recited in the claims.

[0141] The title, background technology, figure descriptions, abstract, and figures are all incorporated into this disclosure and are provided as illustrative examples of the disclosure, rather than as limiting descriptions. They are submitted with the understanding that they will not be used to limit the scope or meaning of the claims. In addition, it can be seen in the detailed description that the description provides illustrative examples and that various features are grouped together in various embodiments for the purpose of simplifying the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed subject matter requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive subject matter lies in less than all the features of a single disclosed configuration or operation. The following claims are hereby incorporated into the detailed description, with each claim standing on its own as separately claimed subject matter.

[0142] The claims are not intended to be limited to the aspects described herein, but are to be accorded the full scope consistent with the language of the claims and to encompass all legal equivalents. Nevertheless, no claim is intended to encompass subject matter that fails to meet the requirements of applicable patent law, nor should they be construed in such a manner.

Claims

1. A device comprising: a communication interface configured to receive a compressed bitstream including vertex motion vector information, the vertex motion vector information including one or more components; as well as a processor operatively coupled to the communication interface, the processor configured to: Parsing a compressed bitstream including vertex motion vector information; selecting a context for a plurality of bins of a first component of vertex motion vector information, wherein first prefix bins of the first component are encoded based on a first context and one or more remaining bins of the first component are encoded using bypass coding; as well as A first component of the vertex motion vector information is decoded based on the selected context.

2. The device according to claim 1, wherein The remaining prefix bins of the first component are encoded based on a second context.

3. The device according to any one of claims 1 to 2, wherein: The processor is further configured to: selecting a context for a plurality of bins of one or more second components of vertex motion vector information, wherein first prefix bins of the one or more second components are encoded based on a first context and one or more remaining bins of the one or more second components are encoded using bypass coding; and One or more second components of the vertex motion vector information are decoded based on the selected context.

4. The device according to claim 3, wherein The remaining prefix bins of the first component and the one or more second components are encoded based on a second context.

5. The device according to claim 4, wherein The first context is shared among first prefix bins of the first component and the one or more second components, and the second context is shared among remaining prefix bins of the first component and the one or more second components.

6. The device according to any one of claims 1 to 5, wherein: One or more components of the vertex motion vector information are binarized using a combination of a unary code including a prefix portion and a suffix portion and an Exponential-Golomb code.

7. The device according to any one of claims 3 to 6, wherein: All suffix bins of the first component and the one or more second components are encoded using bypass coding.

8. A method comprising: receiving a compressed bitstream including vertex motion vector information, the vertex motion vector information including one or more components; Parsing a compressed bitstream including vertex motion vector information; selecting a context for a plurality of bins of a first component of vertex motion vector information, wherein first prefix bins of the first component are encoded based on a first context and one or more remaining bins of the first component are encoded using bypass coding; as well as A first component of the vertex motion vector information is decoded based on the selected context.

9. A device comprising: Communication interface; a processor operatively coupled to the communication interface, the processor configured to: generating vertex motion vector information comprising one or more components; encoding a first prefix binary bit of a first component of vertex motion vector information based on a first context; encoding the one or more remaining bins of the first component using bypass coding; forming a bitstream including encoded vertex motion vector information; as well as The bitstream is sent to a decoding device via a communication interface.

10. The device according to claim 9, wherein The processor is further configured to cause remaining prefix bins of the first component to be encoded based on the second context.

11. The device according to any one of claims 9 to 10, wherein: The processor is further configured to: encoding first prefix bins of one or more second components of the vertex motion vector information based on the first context; as well as The one or more remaining bins of the one or more second components are encoded using bypass coding.

12. The device according to any one of claims 9 to 11, wherein The processor is further configured to cause encoding of remaining prefix bins of the first component and the one or more second components based on a second context.

13. The device according to claim 12, wherein The first context is shared among first prefix bins of the first component and the one or more second components, and the second context is shared among remaining prefix bins of the first component and the one or more second components.

14. The device according to any one of claims 9 to 13, wherein One or more components of the vertex motion vector information are binarized using a combination of a unary code including a prefix portion and a suffix portion and an Exponential-Golomb code.

15. The device according to any one of claims 11 to 14, wherein The first component and all suffix bins of the one or more components are encoded using bypass coding.