Packing displacement component samples in displacement video frames for dynamic grid coding

By packing displacement component samples in ascending or descending order in the displacement video frame and adding padding samples in the lower area, the problems of computational delay and low decoding efficiency in dynamic grid coding are solved, and efficient displacement data decoding is achieved.

CN120604514APending Publication Date: 2025-09-05SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480009702.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2024-04-22
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In dynamic grid coding, existing techniques fill samples at the top or bottom of a frame, resulting in unnecessary calculations and prolonged adaptation time of the arithmetic coding model, and low decoding efficiency of signals at different levels of detail.

Method used

By packing displacement component samples in ascending or descending order in the displacement video frame and adding padding samples in the lower area of ​​the frame, the displacement component samples are generated by combining inverse quantization and inverse wavelet transform, and the packing mechanism is optimized to independently decode signals at different detail levels.

Benefits of technology

The decoding efficiency of dynamic grid coding is improved, unnecessary calculations are reduced, the adaptation time of the arithmetic coding model is optimized, and partial or independent decoding of displacement data is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604514A_ABST
    Figure CN120604514A_ABST
Patent Text Reader

Abstract

A device receives a compressed bitstream including an encoded shifted bitstream and packing method information indicating whether shifted component samples are packed in an ascending order or a descending order. The apparatus video decodes the encoded shifted bitstream to generate shifted video frames where padding is added at the bottom of the shifted video frames irrespective of whether the shifted component samples are packed in ascending order or descending order. The device performs image unpacking on the displacement video frame to generate an array of quantized displacement wavelet coefficients, performs inverse quantization on the array of quantized displacement wavelet coefficients to generate displacement wavelet coefficients, and performs inverse wavelet transform on the displacement wavelet coefficients to generate displacement component sample points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to improvements to dynamic grid coding, and more particularly to improvements such as, but not limited to, improvements to packing displacement component samples in displacement video frames for dynamic grid coding and improvements to padding displacement video frames for dynamic grid coding. Background Art

[0002] In general, mesh encoding and decoding operations can be highly sequential. The base mesh can be divided into multiple sub-meshes. Each decoded sub-mesh can be subdivided, and the decoded displacement field can then be used to refine the positions of the subdivision points belonging to that sub-mesh.

[0003] The displacement component samples can be wavelet transformed into displacement wavelet coefficients. The displacement wavelet coefficients can be quantized, and the quantized displacement wavelet coefficients can be packed into displacement video frames. When the size of the displacement video frame varies in time, some frames need to be padded to make the frame size constant.

[0004] The descriptions set forth in the Background section should not be assumed to qualify as prior art merely because they are set forth in the Background section.The Background section may describe aspects or embodiments of the present disclosure. Summary of the Invention

[0005] Technical issues

[0006] When reversible packing is used, in some embodiments, padding rows of samples may be added at the top of the frame. When padding is added at the top of the frame, in order to decode only a certain level of detail signal, the padding samples must be decoded, resulting in unnecessary computations. In addition, the arithmetic coding model may take longer to adapt.

[0007] In some embodiments, padding rows may be added at the bottom of the frame regardless of whether reversible packing is used.

[0008] In some embodiments, a packing mechanism is proposed, which enables different level-of-detail signals to be included in different slices, thereby enabling partial or independent decoding of displacement data.

[0009] Technical Solution

[0010] In an embodiment, a device includes: a communication interface configured to receive a compressed bitstream, the compressed bitstream including an encoded displacement bitstream and packing method information indicating whether displacement component samples are packed in ascending order or descending order, wherein the displacement component samples are packed in ascending order when the packing method information is equal to a first value, and the displacement component samples are packed in descending order when the packing method information is equal to a second value, wherein the first value is 0 and the second value is 1; and a processor operably coupled to the communication interface. The processor is configured to: decode the packing method information; perform video decoding on the encoded displacement bitstream to generate a displacement video frame; perform image unpacking on the displacement video frame to generate an array of quantized displacement wavelet coefficients, wherein the displacement video frame includes padding at a lower area of ​​the displacement video frame regardless of whether the displacement component samples are packed in descending order; perform inverse quantization on the array of quantized displacement wavelet coefficients to generate displacement wavelet coefficients; and perform inverse wavelet transform on the displacement wavelet coefficients to generate the displacement component samples.

[0011] In an embodiment, regardless of whether the packing method information indicates that the displacement component samples are packed in descending order, the upper area of ​​the displacement video frame includes quantized displacement wavelet coefficients, and the lower area of ​​the displacement video frame does not include any quantized displacement wavelet coefficients and includes samples for filling.

[0012] In an embodiment, if the packing method information indicates that the displacement component samples are packed in descending order, the image unpacking is performed on the displacement video frame starting from the lower right block of the upper area.

[0013] In an embodiment, the lower right block of the upper region comprises quantized displaced wavelet coefficients belonging to the lowest level of detail (LOD).

[0014] In an embodiment, a plurality of packed blocks are placed in a displaced video frame.

[0015] In an embodiment, the plurality of packed blocks include one or more non-filled blocks and partially filled blocks, each of the one or more non-filled blocks including quantized displacement wavelet coefficients belonging to a first level of detail (LOD) and not including any samples for filling, and the partially filled blocks including one or more quantized displacement wavelet coefficients belonging to the first LOD and one or more samples for filling.

[0016] In an embodiment, when the packing method information indicates that the displacement component samples are packed in descending order, the one or more non-filling blocks are arranged before the partially filled block in a scanning order of blocks.

[0017] In an embodiment, the scanning order of the blocks is a reverse raster scan.

[0018] In an embodiment, the partially filled block is preceded by one or more fully filled blocks, each of the one or more fully filled blocks not comprising any quantized shifted wavelet coefficients to be decoded and comprising samples used for filling.

[0019] In an embodiment, each of the one or more non-filling blocks is not allowed to include quantized shifted wavelet coefficients belonging to a second LOD, the second LOD being different from the first LOD.

[0020] In an embodiment, a plurality of packing blocks are placed in a shifted video frame, and a size of the plurality of packing blocks is an integer multiple of a size of a unit block used to configure a partially decodable or independently decodable area for a video coding standard used by video decoding, the integer being equal to or greater than 1.

[0021] In an embodiment, partially decodable or independently decodable regions for the video coding standard used by the video decoding are not allowed to include quantized displaced wavelet coefficients belonging to two or more levels of detail.

[0022] In an embodiment, when video decoding uses High Efficiency Video Coding (HEVC), the size of the unit block is the size of a coding tree unit specified in video decoding.

[0023] In an embodiment, when video decoding uses Advanced Video Coding (AVC), the size of the unit block is the size of a macroblock specified in video decoding.

[0024] In an embodiment, a method includes: receiving a compressed bitstream, the compressed bitstream including an encoded displacement bitstream and packing method information indicating whether the displacement component samples are packed in ascending order or in descending order, wherein the displacement component samples are packed in ascending order when the packing method information is equal to a first value, and the displacement component samples are packed in descending order when the packing method information is equal to a second value, wherein the first value is 0 and the second value is 1; decoding the packing method information; performing video decoding on the encoded displacement bitstream to generate a displacement video frame; performing image unpacking on the displacement video frame to generate an array of quantized displacement wavelet coefficients, wherein the displacement video frame includes padding at a lower area of ​​the displacement video frame, regardless of whether the displacement component samples are packed in descending order; dequantizing the array of quantized displacement wavelet coefficients to generate displacement wavelet coefficients; and performing inverse wavelet transform on the displacement wavelet coefficients to generate displacement component samples.

[0025] In an embodiment, a device includes: a communication interface; and a processor operably coupled to the communication interface. The processor is configured to: encode packing method information indicating whether displacement component samples are packed in ascending order or descending order, wherein when the packing method information is equal to a first value, the displacement component samples are packed in ascending order, and when the packing method information is equal to a second value, the displacement component samples are packed in descending order, wherein the first value is 0 and the second value is 1; perform wavelet transform on the displacement component samples to generate displacement wavelet coefficients; quantize the displacement wavelet coefficients to generate an array of quantized displacement wavelet coefficients; perform image packing on the array of quantized displacement wavelet coefficients to generate a displacement video frame, wherein padding is added at a lower area of ​​the displacement video frame regardless of whether the displacement component samples are packed in descending order; perform video encoding on the displacement video frame to generate an encoded displacement bitstream; and combine the encoded packing method information and the encoded displacement bitstream into a compressed bitstream.

[0026] In an embodiment, when the packing method information indicates that the displacement component samples are packed in descending order, the upper area of ​​the displacement video frame includes quantized displacement wavelet coefficients, and the lower area of ​​the displacement video frame does not include any quantized displacement wavelet coefficients and includes samples for filling.

[0027] In an embodiment, if the packing method information indicates that the displacement component samples are packed in descending order, image packing is performed on the displacement video frame starting from the lower right block of the upper area.

[0028] In an embodiment, the lower right block of the upper region comprises quantized displaced wavelet coefficients belonging to the lowest level of detail (LOD).

[0029] In an embodiment, a plurality of packing blocks are placed in a shifted video frame, and a size of the plurality of packing blocks is an integer multiple of a size of a unit block used to configure a partially decodable or independently decodable area for a video coding standard used by video decoding, the integer being equal to or greater than 1. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 An example communication system according to an embodiment of the present disclosure is shown.

[0031] Figure 2 and Figure 3 An example electronic device according to an embodiment of the present disclosure is shown.

[0032] Figure 4 A block diagram of an encoder for encoding intra frames according to an embodiment is shown.

[0033] Figure 5 A block diagram for a decoder according to an embodiment is shown.

[0034] Figure 6 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0035] Figure 7 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0036] Figure 8 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0037] Figure 9 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0038] Figure 10 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0039] Figure 11 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is not used, according to an embodiment.

[0040] Figure 12 The placement of displaced component samples and padding data in a displaced video frame when reversible packing is used according to an embodiment is shown.

[0041] 13A , 13B, and 13C illustrate placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0042] Figure 14 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0043] Figure 15 1 shows the placement of displacement component samples in a displacement video frame when reversible packing is enabled according to an embodiment.

[0044] Figure 16 is a flowchart illustrating the operation of an image depacketizer according to an embodiment.

[0045] Figure 17 is a flowchart illustrating the operation of an image depacketizer according to an embodiment.

[0046] In one or more implementations, not all components depicted in each figure may be required, and one or more implementations may include additional components not shown in the figures. The arrangement and types of components may be varied without departing from the scope of the present disclosure. Additional components, different components, or fewer components may be utilized within the scope of the present disclosure. DETAILED DESCRIPTION

[0047] The detailed description set forth below in conjunction with the accompanying drawings is intended to serve as a description of various embodiments and is not intended to represent the only embodiment in which the subject technology can be practiced. On the contrary, the detailed description includes specific details for the purpose of providing a thorough understanding of the subject of the present invention. As will be appreciated by those skilled in the art, the embodiments described can be modified in various ways without departing from the scope of this disclosure. Therefore, the drawings and description are to be considered illustrative and not restrictive in nature. The same reference numerals represent the same elements.

[0048] Thanks to the ready availability of powerful handheld devices such as smartphones, three hundred and sixty degree (360°) video and 3D volumetric video are becoming new ways to experience immersive content. While 360° video enables consumers to have an immersive, "real life," "there" experience by capturing a 360° outside-in perspective of the world, 3D volumetric video can provide a full 6DoF experience of being within and moving within the content. Users can interactively change their viewpoint and dynamically view any part of the captured scene or object they desire. Display and navigation sensors can track the user's head movements in real time to determine the area of ​​the 360° video or volumetric content that the user wants to view or interact with. Multimedia data that is inherently three-dimensional (3D), such as point clouds or 3D polygon meshes, can be used in immersive environments.

[0049] A point cloud is a collection of 3D points with properties (such as color, normal, reflectivity, point size, etc.) that represent the surface or volume of an object. Point clouds are common in various applications such as gaming, 3D mapping, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view replay, 6DoF immersive media, to name a few. Point clouds, if uncompressed, typically require a large amount of bandwidth for transmission. Due to the large bitrate requirements, point clouds are typically compressed before transmission. Compressing 3D objects such as point clouds typically requires dedicated hardware. To avoid dedicated hardware to compress 3D point clouds, 3D point clouds can be transformed into traditional two-dimensional (2D) frames and can be compressed and later reconstructed and made visible to the user.

[0050] Polygonal 3D meshes, particularly triangle meshes, are another popular format for representing 3D objects. A mesh is typically composed of a set of vertices, edges, and faces that are used to represent the surface of a 3D object. A triangle mesh is a simple polygonal mesh where the faces are simple triangles that cover the surface of the 3D object. Typically, there may be one or more attributes associated with a mesh. In one scenario, one or more attributes may be associated with each vertex in the mesh. For example, a texture attribute (RGB) may be associated with each vertex. In another scenario, each vertex may be associated with a pair of coordinates (u,v). The (u,v) coordinates may point to a location in a texture map associated with the mesh. For example, the (u,v) coordinates may refer to a row and column index, respectively, in a texture map. A mesh can be thought of as a point cloud with additional connectivity information.

[0051] Point clouds or meshes can be dynamic, i.e., they can change over time. In these cases, the point cloud or mesh at a particular moment in time can be referred to as a point cloud frame or mesh frame, respectively.

[0052] Since point clouds and meshes contain large amounts of data, they need to be compressed for efficient storage and transmission. This is especially true for dynamic point clouds and meshes that can contain 60 frames per second or higher.

[0053] The drawings discussed below and the various embodiments used to describe the principles of the present disclosure in this patent document are intended to be illustrative only and should not be construed in any way to limit the scope of the present disclosure. Those skilled in the art will understand that the principles of the present disclosure can be implemented in any suitably arranged system or device.

[0054] Figure 1 An example communication system according to an embodiment of the present disclosure is shown. Figure 1 The embodiment of the communication system 100 shown in FIGURE 1 is for illustration only. Other embodiments of the communication system 100 may be used without departing from the scope of this disclosure.

[0055] The communication system 100 includes a network 102 that facilitates communication between various components in the communication system 100. For example, the network 102 can communicate IP packets, frame relay frames, asynchronous transfer mode (ATM) cells, or other information between network addresses. The network 102 includes all or part of one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), a global network such as the Internet, or any other communication system(s) at one or more locations.

[0056] In this example, network 102 facilitates communication between server 104 and various client devices 106-116. Client devices 106-116 may be, for example, smartphones, tablets, laptops, personal computers, TVs, interactive displays, wearable devices, HMDs, etc. Server 104 may represent one or more servers. Each server 104 includes any suitable computing or processing device that can provide computing services to one or more client devices (such as client devices 106-116). Each server 104 may, for example, include one or more processing devices, one or more memories for storing instructions and data, and one or more network interfaces for facilitating communication over network 102. As described in more detail below, server 104 may send a compressed bitstream representing a point cloud or mesh to one or more display devices (such as client devices 106-116). In embodiments, each server 104 may include an encoder.

[0057] Each client device 106-116 represents any suitable computing or processing device that interacts with at least one server (such as server 104) or other computing device(s) via network 102. Client devices 106-116 include a desktop computer 106, a mobile phone or mobile device 108 (such as a smartphone), a PDA 110, a laptop computer 112, a tablet computer 114, and an HMD 116. However, any other or additional client devices may be used in communication system 100. A smartphone represents a type of mobile device 108 that is a handheld device with a mobile operating system and an integrated mobile broadband cellular network connection for voice, short message service (SMS), and internet data communications. The HMD 116 can display a 360° scene including one or more dynamic or static 3D point clouds. In embodiments, any of the client devices 106-116 may include an encoder, a decoder, or both. For example, the mobile device 108 may record 3D volumetric video and then encode the video so that it can be transmitted to one of the client devices 106-116. In another example, the laptop computer 112 may be used to generate a 3D point cloud or mesh, which is then encoded and sent to one of the client devices 106 - 116 .

[0058] In this example, some client devices 108-116 communicate indirectly with the network 102. For example, the mobile device 108 and the PDA 110 communicate via one or more base stations 118, such as cellular base stations or eNodeBs (eNBs). Additionally, the laptop 112, the tablet 114, and the HMD 116 communicate via one or more wireless access points 120, such as IEEE 802.11 wireless access points. Note that these are for illustration purposes only, and each client device 106-116 may communicate directly with the network 102 or indirectly with the network 102 via any suitable intermediary device(s) or network(s). In an embodiment, the server 104 or any client device 106-116 may be used to compress a point cloud or mesh, generate a bitstream representing the point cloud or mesh, and send the bitstream to another client device, such as any client device 106-116.

[0059] In an embodiment, any of the client devices 106-114 securely and efficiently transmits information to another device, such as, for example, the server 104. In addition, any of the client devices 106-116 can trigger information transmission between itself and the server 104. Any of the client devices 106-114 can be used as a VR display when attached to a head-mounted device via a bracket and function similarly to the HMD 116. For example, the mobile device 108 can function similarly to the HMD 116 when attached to a bracket system and worn on the user's eyes. The mobile device 108 (or any other client device 106-116) can trigger information transmission between itself and the server 104.

[0060] In an embodiment, any one of the client devices 106-116 or the server 104 may create a 3D point cloud or mesh, compress a 3D point cloud or mesh, transmit a 3D point cloud or mesh, receive a 3D point cloud or mesh, decode a 3D point cloud or mesh, render a 3D point cloud or mesh, or a combination thereof. For example, the server 104 may then compress the 3D point cloud or mesh to generate a bitstream, and then transmit the bitstream to one or more of the client devices 106-116. For another example, one of the client devices 106-116 may compress the 3D point cloud or mesh to generate a bitstream, and then transmit the bitstream to another of the client devices 106-116 or to the server 104. In other words, the server 104 and / or any one of the client devices 106-116 may be a device as described in the present disclosure.

[0061] although Figure 1 One example of a communication system 100 is shown, but may be used for Figure 1Various changes may be made. For example, the communication system 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems have a wide variety of configurations, and Figure 1 The scope of this disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may be used in any other suitable system.

[0062] Figure 2 and Figure 3 An example electronic device according to an embodiment of the present disclosure is shown. Specifically, Figure 2 An example server 200 is shown and may represent Figure 1 The server 200 may represent one or more encoders, decoders, local servers, remote servers, cluster computers, components that act as a single seamless resource pool, cloud-based servers, etc. The server 200 may be composed of Figure 1 One or more of the client devices 106-116 or another server accesses the server.

[0063] Server 200 may represent one or more local servers, one or more compression servers, or one or more encoding servers (such as encoders). In an embodiment, the encoder may perform decoding. Figure 2 As shown, server 200 includes a bus system 205 that supports communications between at least one processing device (such as processor 210 ), at least one storage device 215 , at least one communication interface 220 , and at least one input / output (I / O) unit 225 .

[0064] The processor 210 executes instructions that may be stored in the memory 230. The processor 210 may include any suitable number(s) and type(s) of processors or other devices in any suitable arrangement. Example types of processors 210 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application specific integrated circuits, and discrete circuits.

[0065] In an embodiment, the processor 210 may encode the 3D point cloud or mesh stored in the storage device 215. In an embodiment, the encoding operation on the 3D point cloud also decodes the 3D point cloud or mesh to ensure that when the point cloud or mesh is reconstructed, the reconstructed 3D point cloud or mesh matches the 3D point cloud or mesh before encoding.

[0066] Memory 230 and persistent storage 235 are examples of storage devices 215, which represent any structure(s) capable of storing and facilitating retrieval of information, such as data, program code, or other suitable information, on a temporary or permanent basis. Memory 230 may represent random access memory or any other suitable volatile or non-volatile storage device(s). For example, the instructions stored in memory 230 may include instructions for decomposing a point cloud into patches, instructions for packing patches on 2D frames, instructions for compressing 2D frames, and instructions for encoding 2D frames in a particular order to generate a bitstream. The instructions stored in memory 230 may also include instructions for displaying the point cloud in a manner such as through a VR headset, such as a Figure 1 Persistent storage 235 may include one or more components or devices that support long-term storage of data, such as read-only memory, a hard drive, flash memory, or an optical disk.

[0067] The communication interface 220 supports communication with other systems or devices. For example, the communication interface 220 may include a Figure 1 The communication interface 220 may include a network interface card or wireless transceiver for communication with the network 102. The communication interface 220 may support communication via any suitable physical or wireless communication link(s). For example, the communication interface 220 may send a bitstream containing a 3D point cloud to another device (such as one of the client devices 106-116).

[0068] The I / O unit 225 allows for the input and output of data. For example, the I / O unit 225 can provide a connection for user input via a keyboard, mouse, keypad, touch screen, or other suitable input device. The I / O unit 225 can also send output to a display, printer, or other suitable output device. However, it is noted that the I / O unit 225 can be omitted, such as when I / O interaction with the server 200 occurs via a network connection.

[0069] Note that although Figure 2 Described as indicating Figure 1 The server 104 of FIG. 106 may be configured as a server 104, but the same or similar architecture may be used in one or more of the various client devices 106-116. For example, a desktop computer 106 or a laptop computer 112 may have a server 104 configured as a server 104. Figure 2 The same or similar structure as shown in .

[0070] Figure 3 An example electronic device 300 is shown and may represent Figure 1The electronic device 300 may be a mobile communication device such as, for example, a mobile station, a subscriber station, a wireless terminal, a desktop computer (similar to a Figure 1 Desktop computer 106), portable electronic device (similar to Figure 1 In an embodiment, Figure 1 One or more of the client devices 106 to 116 may include the same or similar configuration as the electronic device 300. In an embodiment, the electronic device 300 is an encoder, a decoder, or both. For example, the electronic device 300 may be used for data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.

[0071] like Figure 3 As shown, the electronic device 300 includes an antenna 305, a radio frequency (RF) transceiver 310, a transmit (TX) processing circuit 315, a microphone 320, and a receive (RX) processing circuit 325. The RF transceiver 310 may include, for example, an RF transceiver, a Bluetooth transceiver, a WI-FI transceiver, a ZIGBEE transceiver, an infrared transceiver, and various other wireless communication signals. The electronic device 300 also includes a speaker 330, a processor 340, an input / output (I / O) interface (IF) 345, an input 350, a display 355, a memory 360, and (one or more) sensors 365. The memory 360 includes an operating system (OS) 361 and one or more applications 362.

[0072] The RF transceiver 310 receives incoming RF signals from the antenna 305, transmitted from an access point (such as a base station, a Wi-Fi router, or a Bluetooth device) or other devices of the network 102 (such as Wi-Fi, Bluetooth, cellular, 5G, LTE, LTE-A, WiMAX, or any other type of wireless network). The RF transceiver 310 downconverts the incoming RF signals to generate an intermediate frequency (IF) or baseband signal. The IF or baseband signal is sent to the RX processing circuit 325, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or IF signal. The RX processing circuit 325 sends the processed baseband signal to the speaker 330 (such as for voice data) or to the processor 340 for further processing (such as for web browsing data).

[0073] TX processing circuitry 315 receives analog or digital voice data from microphone 320 or other outgoing baseband data from processor 340. The outgoing baseband data may include web data, email, or interactive video game data. TX processing circuitry 315 encodes, multiplexes, and / or digitizes the outgoing baseband data to generate a processed baseband or intermediate frequency (IF) signal. RF transceiver 310 receives the outgoing processed baseband or IF signal from TX processing circuitry 315 and up-converts the baseband or IF signal into an RF signal that is transmitted via antenna 305.

[0074] The processor 340 may include one or more processors or other processing devices. The processor 340 may execute instructions stored in the memory 360 (such as the OS 361) to control the overall operation of the electronic device 300. For example, the processor 340 may control the reception of forward channel signals and the transmission of reverse channel signals by the RF transceiver 310, the RX processing circuit 325, and the TX processing circuit 315 according to well-known principles. The processor 340 may include any suitable number (one or more) and type (one or more) of processors or other devices of any suitable arrangement. For example, in an embodiment, the processor 340 includes at least one microprocessor or microcontroller. Example types of the processor 340 include a microprocessor, a microcontroller, a digital signal processor, a field programmable gate array, an application-specific integrated circuit, and a discrete circuit.

[0075] The processor 340 is also capable of executing other processes and programs residing in the memory 360, such as operations for receiving and storing data. The processor 340 can move data into or out of the memory 360 as needed for the execution process. In an embodiment, the processor 340 is configured to execute one or more applications 362 based on the OS 361 or in response to a signal received from (one or more) external sources or operators. For example, the application 362 may include an encoder, a decoder, a VR or AR application, a camera application (for still images and video), a video phone call application, an email client, a social media client, an SMS messaging client, a virtual assistant, etc. In an embodiment, the processor 340 is configured to receive and send media content.

[0076] The processor 340 is also coupled to an I / O interface 345 , which provides the electronic device 300 with the ability to connect to other devices, such as the client devices 106 - 114 . The I / O interface 345 is the communication path between these accessories and the processor 340 .

[0077] Processor 340 is also coupled to input 350 and display 355. The operator of electronic device 300 can use input 350 to enter data or input into electronic device 300. Input 350 can be a keyboard, touch screen, mouse, trackball, voice input, or other device capable of serving as a user interface to allow the user to interact with electronic device 300. For example, input 350 can include voice recognition processing, allowing the user to enter voice commands. In another example, input 350 can include a touch panel, a (digital) pen sensor, keys, or an ultrasonic input device. The touch panel can recognize touch input using at least one scheme (such as a capacitive scheme, a pressure-sensitive scheme, an infrared scheme, or an ultrasonic scheme). By providing additional inputs to processor 340, input 350 can be associated with (one or more) sensors 365 and / or cameras. In embodiments, sensors 365 include one or more inertial measurement units (IMUs) (such as accelerometers, gyroscopes, and magnetometers), motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, etc. Input 350 may also include control circuitry. In a capacitive solution, the input 350 may recognize touch or proximity.

[0078] Display 355 may be a liquid crystal display (LCD), a light emitting diode (LED) display, an organic LED (OLED), an active matrix OLED (AMOLED), or other display capable of rendering text and / or graphics (such as from a website, video, game, image, etc.). Display 355 may be sized to fit within the HMD. Display 355 may be a single display screen or multiple display screens capable of creating a stereoscopic display. In an embodiment, display 355 is a heads-up display (HUD). Display 355 may display 3D objects (such as a 3D point cloud or mesh).

[0079] Memory 360 is coupled to processor 340. A portion of memory 360 may include RAM, and another portion of memory 360 may include flash memory or other ROM. Memory 360 may include a persistent storage (not shown) representing any (one or more) structures capable of storing and facilitating the retrieval of information (such as data, program code, and / or other suitable information). Memory 360 may include one or more components or devices that support long-term storage of data (such as read-only memory, a hard drive, flash memory, or an optical disk). Memory 360 may also include media content. The media content may include various types of media (such as images, videos, three-dimensional content, VR content, AR content, 3D point clouds, meshes, etc.).

[0080] The electronic device 300 also includes one or more sensors 365 capable of measuring physical quantities or detecting the activation state of the electronic device 300 and converting the measured or detected information into electrical signals. For example, the sensor 365 may include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor (such as a gyroscope or gyro sensor and an accelerometer), an eye tracking sensor, an air pressure sensor, a magnetic sensor or a magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illumination sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an IR sensor, an ultrasonic sensor, an iris sensor, a fingerprint sensor, a color sensor (such as a red, green, and blue (RGB) sensor), etc. The sensor 365 may also include a control circuit for controlling any of the sensors included therein.

[0081] As discussed in more detail below, one or more of these sensor(s) 365 can be used to control a user interface (UI), detect UI input, determine the user's orientation and facing direction for three-dimensional content display recognition, etc. Any of these sensor(s) 365 can be located within the electronic device 300, within a secondary device operably connected to the electronic device 300, within a head-mounted device configured to hold the electronic device 300, or within a single device including the electronic device 300 and the head-mounted device.

[0082] The electronic device 300 may create media content, such as generating a virtual object or capturing (or recording) content through a camera. The electronic device 300 may encode the media content to generate a bitstream so that the bitstream can be sent directly to another electronic device, or for example, via Figure 1 The electronic device 300 may receive the bit stream directly from another electronic device, or indirectly from another electronic device such as through a Figure 1 The network 102 indirectly receives the bit stream.

[0083] although Figure 2 and Figure 3 An example of an electronic device is shown, but the Figure 2 and Figure 3 Make various changes. For example, Figure 2 and Figure 3 Various components in the can be combined, further subdivided, or omitted, and additional components can be added according to specific needs. As a specific example, processor 340 can be divided into multiple processors, such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs). In addition, as with computing and communications, electronic devices and servers can have a variety of configurations, and Figure 2 and Figure 3 The present disclosure is not limited to any particular electronic device or server.

[0084] Figure 4 A block diagram of an encoder for encoding intra frames according to an embodiment is shown.

[0085] like Figure 4 As shown, the encoder 400 for encoding intra frames according to an embodiment may include a quantizer 401, a static grid encoder 403, a static grid decoder 405, a displacement updater 407, a wavelet transformer 409, a quantizer 411, an image packer 413, a video encoder 415, an image unpacker 417, an inverse quantizer 419, an inverse wavelet transformer 421, an inverse quantizer 423, a deformed grid reconstructor 425, an attribute transfer module 427, a padding module 429, a color space converter 431, a video encoder 433, a multiplexer 435 and a controller 437.

[0086] The quantizer 401 may quantize the base mesh m(i) to generate a quantized base mesh. In an embodiment, the base mesh may have fewer vertices than the original mesh.

[0087] The static grid encoder 403 may encode and compress the quantized base grid to generate a compressed base grid bitstream. In an embodiment, the base grid may be compressed in a lossy or lossless manner. In an embodiment, an existing grid codec (such as Draco) may be used to compress the base grid.

[0088] The static grid decoder 405 may decode the compressed base grid bitstream to generate a reconstructed quantized base grid m′(i).

[0089] The displacement updater 407 can update the displacement d(i) based on the subdivided base mesh m(i) and the reconstructed quantized base mesh m'(i) to generate updated displacement d'(i). The reconstructed base mesh can be subdivided, and then the displacement field between the original mesh and the subdivided reconstructed base mesh can be calculated. In inter-frame coding of mesh frames, the base mesh can be encoded by transmitting vertex motion rather than directly compressing the base mesh. In either case, a displacement field can be created. The displacement field and the modified property map can be encoded using a video codec and also included as part of the V-DMC bitstream.

[0090] The wavelet transformer 409 may perform a wavelet transform using the updated displacement d'(i) to generate a displacement wavelet coefficient e(i). The wavelet transform may include a series of prediction and update lifting steps.

[0091] The quantizer 411 may quantize the shifted wavelet coefficients e(i) to generate quantized shifted wavelet coefficients e'(i). The quantized shifted wavelet coefficients may be represented by an array dispQuantCoeffArray.

[0092] The image packer 413 may pack the quantized displacement wavelet coefficients e'(i) into a 2D image including the packed quantized displacement wavelet coefficients dispQuantCoeffFrame. In the present disclosure, a 2D video frame may be referred to as a displacement frame or a displacement video frame.

[0093] The video encoder 415 may encode the packed quantized displacement wavelet coefficients dispQuantCoeffFrame to generate a compressed displacement bitstream.

[0094] The image unpacker 417 may unpack the packed quantized displacement wavelet coefficients dispQuantCoeffFrame to generate an array of quantized displacement wavelet coefficients dispQuantCoeffArray.

[0095] The dequantizer 419 may dequantize the array dispQuantCoeffArray of quantized shifted wavelet coefficients to generate shifted wavelet coefficients.

[0096] The inverse wavelet transformer 421 may perform inverse wavelet transform using the displacement wavelet coefficients to generate reconstructed displacements d”(i).

[0097] The inverse quantizer 423 may inverse quantize the reconstructed quantized base grid m′(i) to generate a reconstructed base grid m”(i).

[0098] The deformed mesh reconstructor 425 may generate a reconstructed deformed mesh DM(i) based on the reconstructed displacements D”(i) and the reconstructed base mesh m”(i).

[0099] The attribute transfer module 427 may update the attribute map A(i) based on the static / dynamic mesh m(i) and the reconstructed deformed mesh DM(i) to generate an updated attribute map A'(i). The attribute map may be a texture map, but other attributes may also be sent.

[0100] The padding module 429 may perform padding to fill blank areas in the updated attribute map A′(i) in order to remove high frequency components.

[0101] The color space converter 431 may perform color space conversion of the padded updated attribute map A′(i).

[0102] The video encoder 433 may encode the output of the color space converter 431 to generate a compressed attribute bitstream.

[0103] The multiplexer 435 may multiplex the compressed base grid bitstream, the compressed displacement bitstream, and the compressed attribute bitstream to generate a compressed bitstream b(i).

[0104] The controller 437 may control the modules of the encoder 400 .

[0105] Figure 5 A block diagram for a decoder according to an embodiment is shown.

[0106] like Figure 5 As shown, the decoder 500 may include a demultiplexer 501, a switch 503, a static grid decoder 505, a grid buffer 507, a motion decoder 509, a base grid reconstructor 511, a switch 513, an inverse quantizer 515, a video decoder 521, an image unpacker 523, an inverse quantizer 525, an inverse wavelet transformer 527, a deformed grid reconstructor 529, a video decoder 531 and a color space converter 533.

[0107] The demultiplexer 501 may receive the compressed bitstream b(i) from the encoder 400 to extract a compressed base grid bitstream, a compressed displacement bitstream, and a compressed attribute bitstream from the compressed bitstream b(i).

[0108] The switch 503 can determine whether the compressed base grid bit stream has inter-frame coded grid frame data or intra-frame coded grid frame data. If the compressed base grid bit stream has inter-frame coded grid frame data, the switch 503 can transmit the inter-frame coded grid frame data to the motion decoder 509. If the compressed base grid bit stream has intra-frame coded grid frame data, the switch 503 can transmit the intra-frame coded grid frame data to the static grid decoder 505.

[0109] The static grid decoder 505 may decode the intra-frame encoded grid frame data to generate a reconstructed quantized base grid frame.

[0110] The grid buffer 507 can store the reconstructed quantized base grid frame and inter-frame coded grid frame data for future use in decoding subsequent inter-frame coded grid frames. The reconstructed quantized base grid frame can be used as a reference grid frame.

[0111] The motion decoder 509 may obtain a motion vector for the current inter-frame coded grid frame based on the data stored in the grid buffer 507 and a syntax element in the bitstream for the current inter-frame coded grid frame. In an embodiment, the syntax element in the bitstream for the current inter-frame coded grid frame may be a motion vector difference.

[0112] The base grid reconstructor 511 may generate a reconstructed quantized base grid frame by using a syntax element in a bitstream of the grid frame for the current inter-frame encoding based on a motion vector of the grid frame for the current inter-frame encoding.

[0113] If the compressed base grid bitstream has intra-frame coded grid frame data, the switch 513 may transmit the reconstructed quantized base grid frame from the static grid decoder 505 to the inverse quantizer 515. If the compressed base grid bitstream has inter-frame coded grid frame data, the switch 513 may transmit the reconstructed quantized base grid frame from the base grid reconstructor 511 to the inverse quantizer 515.

[0114] The inverse quantizer 515 may perform inverse quantization using the reconstructed quantized base grid frame to generate a reconstructed base grid frame m”(i).

[0115] The video decoder 521 may decode the displacement bitstream to generate packed quantized displacement wavelet coefficients dispQuantCoeffFrame.

[0116] The image unpacker 523 may unpack the packed quantized displacement wavelet coefficients dispQuantCoeffFrame to generate an array of quantized displacement wavelet coefficients dispQuantCoeffArray.

[0117] The inverse quantizer 525 may perform inverse quantization using the array dispQuantCoeffArray of quantized shifted wavelet coefficients to generate shifted wavelet coefficients.

[0118] The inverse wavelet transformer 527 may perform inverse wavelet transform using the displacement wavelet coefficients to generate displacements.

[0119] The deformed mesh reconstructor 529 may reconstruct the deformed mesh based on the displacement and the reconstructed base mesh frame m”(i).

[0120] The video decoder 531 may decode the attribute bitstream to generate an attribute map before color space conversion.

[0121] The color space converter 533 may perform color space conversion on the property map from the video decoder 531 to reconstruct the property map.

[0122] Hereinafter, the encoder 400 will be described in detail.

[0123] As described above, a base mesh having a generally fewer number of vertices compared to the original mesh can be created and compressed either in a lossy or lossless manner. The reconstructed base mesh can undergo subdivision, and then the displacement field between the original mesh and the subdivided reconstructed base mesh can be calculated. In the inter-frame coding of mesh frames, the base mesh can be encoded by sending vertex motion instead of directly compressing the base mesh. In either case, a displacement field can be created. In some cases, each displacement sample can have 3 components represented by x, y, and z respectively. These can be relative to a canonical coordinate system or a local coordinate system, where x, y, and z represent displacements in the local normal, tangent, and binormal directions.

[0124] The displacement updater 407 of the decoder 400 can update the displacement field d(i) to generate an updated displacement field d'(i). The displacement field d(i) can be represented in Equation 1 below.

[0125] Equation 1

[0126] d(i) = [d x (i), d[[ID=!2]] y (i), d z (i)], 0 ≤ i < N

[0127] where N represents the number of 3-D displacement vectors in the displacement field of the mesh frames

[0128] The wavelet transformer 409 can perform one or more levels of wavelet transform using the updated displacement d'(i) to create a level of detail (LOD) signal d k (i), i = 0 ≤ i < N k , 0 ≤ k < numLOD, where k represents the index of the level of detail, N k represents the number of samples in the level of detail signal at level k, and numLOD represents the number of LODs. The quantizer 411 can perform scalar quantization on the LOD signal d k (i).

[0129] The image packer 413 can pack the quantized LOD signal into a 2D video frame, and the video encoder 415 can compress the 2D video frame by using a conventional video codec. The number of displacement component samples in a mesh frame can be different from the number of displacement component samples in another mesh frame. However, the video codec used in the video encoder 415 can expect the displacement frames to have the same dimensions. If needed, the image packer 413 can add padding in the displacement frame so that the displacement frame has a uniform dimension in time.

[0130] The image packer 413 can reversibly pack the quantized displacement component samples. For example, the number of LODs, numLOD, can be equal to (but not limited to) 3, and the LODs can be represented by lod0, lod1, and lod2, where lod0 contains the lowest frequency displacement wavelet coefficient and lod2 contains the highest frequency displacement wavelet coefficient. When reversible packing is enabled, when the image packer 413 creates a displacement video frame, the image packer 413 can arrange the LODs in a one-dimensional array from the lowest frequency or index of the LOD data to the highest frequency or index of the LOD data. For example, the one-dimensional array can include lod0, lod1 after lod0, and lod2 after lod1 in order of frequency or index. The image packer 413 can then divide the one-dimensional array of quantized displacement wavelet coefficients into blocks (typically 16×16) using the scanning order for the displacement wavelet coefficients, and then place the blocks into the displacement frame according to the scanning order used to pack the blocks. The scanning order for the displacement wavelet coefficients can be, but is not limited to, a reverse Morton scan. The scan order for packed blocks can be, but is not limited to, reverse raster scan when reverse packing is enabled. The scan order for packed blocks can be, but is not limited to, forward raster scan when forward packing is enabled. These blocks can be referred to as V-grid packed blocks, and their size can be represented by blockSize.

[0131] The width of the displaced frame can be fixed to a multiple of blockSize. The height of the displaced frame can vary depending on the number of points in the subdivided grid. Therefore, the image packer 413 can find the maximum height among all V-grid frames in a group of frames or the entire sequence. The image packer 413 can then set the height of the displaced video frame to the maximum height and add rows of padding samples to the displaced video frame so that the height of each displaced video frame is equal to the maximum height.

[0132] In an embodiment, the padding samples may be set equal to a default value. In an embodiment, the default value may be zero. In an embodiment, the default value may be the midpoint of the values ​​representable by the component bit depth of the displacement component samples. In an embodiment, the default value for the luma component of the displacement video may be 0, while the default value for the chroma components may be the midpoint of the values ​​representable by the component bit depth. In an embodiment, the padding samples may be set equal to any value. In an embodiment, the value of one padding sample may be equal to or different from the value of another padding sample. For example, if the bit depth of the padding samples is 8, then the padding samples may be set equal to any value between 0 and 255.

[0133] In the following, reference will be made to Figures 6 to 15 The placement of displacement component samples and padding in a displacement video frame according to various embodiments is described.

[0134] Figure 6 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0135] Reference Figure 6 , since reverse packing is enabled, the displacement component samples can be packed in descending order in the displacement video frame. The padding and packing blocks can be placed in the displacement video frame in reverse raster scan order, so that the packing block precedes the padding in reverse raster scan order, and the frequency or LOD index of the subsequent block is not lower than the frequency or LOD index of any block that precedes the subsequent block in reverse raster scan order in the displacement video frame.

[0136] In an embodiment, the displacement video frame includes an upper area and a lower area. The upper area may not include any displacement wavelet coefficients, but includes padding. The padding may include rows of padding samples. The number of rows of padding samples may be an integer multiple of blockSize, where the integer may be equal to or greater than 1. The lower area may include packed blocks, where the packed blocks contain displacement wavelet coefficients. The lower area may include or may not include padding samples. One or more packed blocks in the lower area may include or may not include padding samples. Since reversible packing is enabled, the packed blocks can be packed in the lower area in descending order of frequency or LOD index. For example, the packed blocks can be placed in the lower area in reverse raster scan order so that the frequency or LOD index of the subsequent block is not lower than the frequency or LOD index of any block in the lower area that precedes the subsequent block in reverse raster scan order.

[0137] Adding padding at the top of a displaced video frame has certain disadvantages. For example, if the decoder 500 wants to decode only a specific LOD (e.g., lod2), it may have to decode all the padding rows at the top. This can waste computational resources and therefore energy. It can also slow down the adaptation of the arithmetic coding model used by the video codec of the video encoder 415.

[0138] To mitigate these drawbacks, in an embodiment, padding rows may be placed at the bottom of the displaced video frame even when reversible packing is used. Figure 7 Describes the padding row placed at the bottom of the displacement video frame.

[0139] Figure 7 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0140] Reference Figure 7, since reversible packing is enabled, the displacement component samples can be packed in descending order in the displacement video frame. The packing block and padding can be placed in the displacement video frame in reverse raster scan order, so that the padding precedes the packing block in reverse raster scan order, and the frequency or LOD index of the subsequent block is not lower than the frequency or LOD index of any block that precedes the subsequent block in reverse raster scan order in the displacement video frame.

[0141] In an embodiment, the displacement video frame includes an upper area and a lower area. The upper area may include a packing block, wherein the packing block contains displacement wavelet coefficients belonging to the LOD. The upper area may or may not include padding samples. One or more packing blocks in the upper area may or may not include padding samples. Since reversible packing is enabled, the packing blocks can be packed in the upper area in descending order of frequency or LOD index. For example, the packing blocks can be placed in the upper area in a reverse raster scan order so that the frequency or LOD index of the subsequent block is not lower than the frequency or LOD index of any block in the upper area that precedes the subsequent block in the reverse raster scan order. The lower area may not include any displacement wavelet coefficients, but includes padding. The padding may include rows of padding samples. The number of rows of padding samples can be an integer multiple of blockSize.

[0142] In this case, decoder 500 may simply skip decoding the padding lines because they occur at the end of the frame and no valid data required by decoder 500 occurs after them.

[0143] In an embodiment, if the decoder 500 wishes to decode only lod2, it may start from the top left corner of the frame and only has to decode the block corresponding to lod2 without having to decode any other blocks.

[0144] Figure 8 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0145] As mentioned previously, the packed blocks belonging to lod2 may not completely fill the top row of blocks. In this case, some padding blocks at the beginning of the top row may be placed.

[0146] Reference Figure 8, the displacement video frame includes packed blocks. Hereinafter, a packed block with padding without any displacement wavelet coefficient may be referred to as a fully filled block or a filled block, a packed block with one or more displacement wavelet coefficients may be referred to as a displacement packed block, a displacement packed block containing one or more displacement wavelet coefficients and padding may be referred to as a partially filled block or a partially filled packed block, and a displacement packed block containing one or more displacement wavelet coefficients without padding may be referred to as a non-filled block or a non-filled packed block. The displacement video frame includes an upper area and a lower area. The upper area may include padding blocks and displacement packed blocks. The padding blocks may not include any displacement wavelet coefficients, but include padding samples. The displacement packed blocks may include displacement wavelet coefficients belonging to the LOD. Since reversible packing is enabled, the displacement component samples or displacement packed blocks can be packed in the upper area in descending order of frequency or LOD index. For example, padding blocks and displacement packing blocks can be placed in the upper region in reverse raster scan order such that the displacement packing block precedes the padding block in reverse raster scan order and the frequency or LOD index of the subsequent displacement packing block is not lower than the frequency or LOD index of any displacement packing block that precedes the subsequent displacement packing block in reverse raster scan order in the upper region. The lower region may not include any displacement wavelet coefficients but may include padding. The padding may include rows of padding samples. The number of rows of padding samples may be an integer multiple of blockSize.

[0147] One or more displacement packed blocks in the upper region may or may not include padding samples. Partially padded blocks having both displacement wavelet coefficients belonging to LOD and padding may not be allowed to precede non-padded packed blocks having displacement wavelet coefficients belonging to LOD but no padding in reverse raster scan order in the upper region. For example, referring to Figure 8 If four non-filled packed blocks with displacement wavelet coefficients belonging to LODlod2 but no filling and a partially filled block with displacement wavelet coefficients belonging to LODlod2 and filling are placed, the partially filled block is not allowed to be arranged before the four non-filled packed blocks in the upper area in the reverse raster scan order.

[0148] Reference Figure 8 , the displacement component samples are packed in the displacement video frame in descending order of frequency. The image packer 413 can pack the images belonging to multiple LODs ( Figure 8 The quantized shifted wavelet coefficients of lod0, lod1, and lod2 are arranged in a one-dimensional array in increasing order.

[0149] However, in this scenario, if the decoder 500 wishes to decode only lod2, it may decode the block in the beginning with the padding samples last.

[0150] Figure 9Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0151] In order to alleviate the Figure 8 To solve the problem described, in an embodiment, the start of reverse packing can be modified as follows. Instead of starting the placement of data from lod0 at the end of the last row of non-filled blocks in the upper area, the data is padded from the middle of the row so that all blocks in the top row of blocks are filled with actual data from LOD. For convenience, the displacement wavelet coefficients belonging to LOD can be referred to as data from LOD. Figure 9 , the upper area may include padding blocks and displacement packing blocks. Since reversible packing is enabled, the displacement component samples or displacement packing blocks can be packed in the upper area in descending order of frequency or LOD index. For example, the padding blocks and displacement packing blocks can be placed in the upper area in reverse raster scan order, so that the padding blocks are arranged before the displacement packing blocks in reverse raster scan order, and the frequency or LOD index of the subsequent displacement packing blocks is not lower than the frequency or LOD index of any displacement packing blocks that are arranged before the subsequent displacement packing blocks in reverse raster scan order in the upper area. The lower area may not include any displacement wavelet coefficients, but includes padding. The padding may include rows of padding samples.

[0152] One or more displacement packed blocks in the upper region may or may not include padding samples. A non-padded packed block having displacement wavelet coefficients belonging to the LOD but no padding may not be allowed to precede a partially padded block having both displacement wavelet coefficients belonging to the LOD and padding in reverse raster scan order in the upper region.

[0153] Reference Figure 9 , for the bottom non-padded row of blocks in the upper region, the last two blocks may correspond to padded blocks with padded samples and without any shifted wavelet coefficients. The third block from the right in the bottom non-padded row of blocks may correspond to a partially padded block with padded samples and at least one shifted wavelet coefficient belonging to LOD lod0. However, the leftmost block of blocks from the topmost row of blocks does not contain any padded samples for padding, but rather contains the highest frequency shifted wavelet coefficient belonging to lod2. Both the encoder and the decoder may calculate the start of the non-padded value in the reverse scan based on the number of samples in the LOD and the packed block size.

[0154] Figure 10 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0155] In an embodiment, in reverse packing, the lowest frequency-shifted wavelet coefficients belonging to lod0 may always start at the beginning of a new block in reverse packing order.

[0156] Reference Figure 10 , the upper area may include padding blocks and displacement packing blocks. Since reversible packing is enabled, the displacement component samples or displacement packing blocks can be packed in the upper area in descending order of frequency or LOD index. For example, the padding blocks and displacement packing blocks can be placed in the upper area in reverse raster scan order, so that the padding blocks are arranged before the displacement packing blocks in reverse raster scan order, and the frequency or LOD index of the subsequent displacement packing blocks is not lower than the frequency or LOD index of any displacement packing blocks that are arranged before the subsequent displacement packing blocks in reverse raster scan order in the upper area. The lower area may not include any displacement wavelet coefficients, but includes padding. The padding may include rows of padding samples.

[0157] One or more displacement packed blocks in the upper region may or may not include padding samples. Partially padded blocks having both displacement wavelet coefficients belonging to LOD and padding may not be allowed to precede packed blocks having displacement wavelet coefficients belonging to LOD without padding in reverse raster scan order in the upper region. For example, referring to Figure 8 If four non-filled packed blocks having displacement wavelet coefficients belonging to LODlod2 but no filling and a partially filled block having both displacement wavelet coefficients belonging to LOD lod2 and filling are placed, the partially filled block is not allowed to be arranged before the four packed blocks in the reverse raster scan order in the upper area.

[0158] Reference Figure 10 , lod0 starts at the third block from the right in the bottom non-padded row of blocks. The upper left block corresponds to the partially padded block containing the highest frequency shifted wavelet coefficients and padded samples belonging to lod2.

[0159] When more than one sub-grid is used, the above strategy for placing the starting point of the LOD data can be repeated for each sub-frame corresponding to the sub-grid. Similarly, the above strategy can be repeated for each of the X, Y, and Z components (or normal, tangent, bitangent components), regardless of whether they are packed into a single luma component of 4:2:0 format shifted video.

[0160] The placement method described above for LOD data can be applied when the padding is at the top of the displaced video frame or the packing order is forward. The placement method described above for LOD data can be applied when the padding is at the top of the displaced video frame and the packing order is forward. Simple modifications may be necessary to account for different scans and placement of padding rows. In an embodiment, the padding data can be placed into its own (one or more) stripes, and the stripe header is inserted at the end of the LOD data. In this way, the decoder 500 can fully decode the (one or more) stripes containing the LOD data that the decoder 500 wants to decode and skip decoding the last stripe containing the padding.

[0161] In an embodiment, if padding data is placed at the beginning of a frame, such as Figure 6 As shown, it can be put into a separate slice so that it can be skipped when decoding.

[0162] Figure 11 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is not used, according to an embodiment.

[0163] Reference Figure 11 , the displacement video frame includes an upper area and a lower area. The upper area includes blocks of multiple rows, and each of the multiple rows includes at least one block containing displacement wavelet coefficients belonging to the LOD. The upper area may or may not include padding samples. The displacement component samples are packed in the displacement video frame in ascending order of frequency. For example, the padding blocks and the displacement packed blocks can be placed in the upper area in raster scan order, so that the displacement packed blocks are arranged before the padding blocks in raster scan order, and the frequency or index of the subsequent block is not allowed to be lower than the frequency or index of any block in the upper area that is arranged before the subsequent block in raster scan order. The lower area does not include any displacement wavelet coefficients, but includes padding. The padding includes padding sample rows.

[0164] The decoder 500 may want to decode only some LODs and not others. In a more common scenario, when there are N LODs, indexed 0, 1, 2, ..., (N-1), the decoder 500 may want to decode all LODs from 0 to k, where k < (N-1), and discard the remaining LODs and padding rows indexed (k+1), ..., (N-1). As an example, the decoder 500 may want to decode lod0 and lod1, but not lod2 and the padding rows. However, when reverse scanning is enabled, the blocks belonging to lod0 are placed at the end of the displaced video frame in scan order, which makes partial and / or independent decoding more challenging.

[0165] To enable partial and / or independent decoding of displacement data, in embodiments, the encoder may use multiple slices to encode the displacement video frame. In embodiments, any slice contains data corresponding to only a single LOD or padding data. In embodiments, each slice may contain data belonging to multiple LODs and possibly padding rows. In embodiments, reference is made to Figure 7 , one stripe contains lod1 and lod2, and another stripe may contain lod0 and padding rows.

[0166] To ensure that different LODs can be separated into slices, in an embodiment, any basic block in the basic block structure used by the video codec of the video encoder 415 may contain data or padding data from only one LOD. The unit block used to define or configure a slice for the video coding standard used by the video encoder 415, or the largest block used to configure a slice for the video coding standard used by the video encoder 415, may be referred to as a basic block or codec basic block. For example, when the video encoder 415 uses HEVC, since a slice is defined as a plurality of coding tree units (CTUs), the size of the basic block may be set to be equal to the size of the coding tree unit (CTU) specified by the video encoder 415. When the video encoder 415 uses Advanced Video Coding (AVC), since a slice is defined as a plurality of macroblocks, the size of the basic block standard may be set to be equal to the size of the macroblock specified by the video encoder 415. In an embodiment, the new LOD may be aligned with the new CTU boundary. However, it is necessary to specify how to determine the next CTU. In an embodiment, a unit block that does not span multiple slices may be referred to as a basic block or codec basic block.

[0167] Figure 12 The placement of displaced component samples and padding data in a displaced video frame when reversible packing is used according to an embodiment is shown.

[0168] Reference Figure 12 , the displaced video frame includes 4 coding tree units CTU0, CTU1, CTU2 and CTU3. The displaced component samples belonging to 2 LODs are placed in the displaced video frame, the V-grid packing block size is 32×32, the HEVC CTU block size is 64×64, and reverse packing is enabled. There are no padding rows. lod0 ends in the fourth V-grid packing block in reverse raster scan order. With reverse raster scan, the next packing block will be the upper right block in coding tree unit CTU3. However, in order to be able to separate the data from lod0 and lod1 into different slices, in an embodiment, the start of lod1 may be placed in coding tree unit CTU1, as Figure 12As shown in FIG. 1 . In this embodiment, the row of V-grid packed blocks including the top right and top left blocks of coding tree unit CTU3 and the top right and top left blocks of coding tree unit CTU2 may contain padding data that may be discarded by the decoder 500. Therefore, when an LOD ends, data from the next LOD may be placed in the next CTU in scan order that does not contain any data from the previous LOD or padding row. Therefore, it may be necessary to skip several V-grid packed blocks to align with the next codec block, thereby increasing the overall size of the displaced video.

[0169] Since a displacement video codec like a wavelet-based codec may not have a concept of block size, it may be better not to define alignment based on the concept of codec block size. Instead, in an embodiment, a packed super block with dimensions that are integer multiples of blockSize may be used, where the integer may be equal to or greater than 1. In an embodiment, when an LOD ends, the data from the next LOD may be placed in the next packed super block in scan order, which does not contain any data from the previous LOD or padding rows, regardless of whether the scan order is forward or reverse. In an embodiment, the encoder may select the packed super block size to match the block size corresponding to the codec used by the video decoder 415 when possible. For example, if HEVC is being used, the encoder may match the packed super block size to the CTU size. In an embodiment, the packed super block size may be the same as the CTU size.

[0170] When the start of an LOD is aligned with the start of a new packed super block and the new packed super block matches a block in the basic block structure used by the video codec of the video encoder 415, in an embodiment, data from each LOD may be placed in a separate stripe, regardless of whether reverse packing is used and where the padding rows are placed. Padding rows may also be placed in separate stripes, regardless of whether reverse packing is used and where the padding rows are placed. In an embodiment, the number of padding rows may be increased to match the height of the basic block. For example, the number of padding rows may be increased to match the CTU height in HEVC. The v-trellis decoder 500 may decode only some stripes to access some LODs and discard other LODs.

[0171] In an embodiment, the v-grid packing block size may be aligned with the packed super block size or, equivalently, with the basic block size used by the displacement video codec of the video encoder 415. Alignment may then be performed so that data aligned with the new LOD is placed in the new packed super block, or equivalently, in the new codec block, in scan order, regardless of whether the scan order is forward or reverse. As previously described, different slices may be used for different LODs. This condition may be signaled using an SEI message or a flag in the v-grid bitstream as a syntax element. In an embodiment, the flag or SEI message may indicate whether each LOD and padding row is placed in a different slice. For example, a flag set to 1 may indicate that displacement component samples belonging to any two different LODs are not allowed to be placed in the same slice, and that displacement component samples and padding rows belonging to any different LODs are not allowed to be placed in the same slice. A flag set to 0 may indicate that displacement component samples belonging to any two different LODs are allowed to be placed in the same slice, and that displacement component samples and padding rows belonging to any different LODs are allowed to be placed in the same slice. When the value of the flag is 1, in an embodiment, the decoder 500 may infer the correspondence between the slice and the LOD and between the slice and the padding line data based on whether reverse scanning is used and whether the padding line is at the bottom of the video frame. In an embodiment, the v-grid packing block size may be an integer multiple of the codec basic block size, where the integer may be equal to or greater than 1. For example, the size of the v-grid packing block may be equal to the size of the codec basic block, or 2 times, 3 times, or 4 times the size of the codec basic block. The width of the v-grid packing block may be an integer multiple of the width of the codec basic block and the height of the v-grid packing block may be an integer multiple of the height of the codec basic block.

[0172] 13A , 13B, and 13C illustrate placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0173] 13A , the starting block for lod0 in the reverse packing case may be determined such that the top left block contains one or more samples from the highest LOD. There may be zero or more blocks containing padding samples to the right of the bottom most row of blocks containing data from lod0.

[0174] 13A , the upper region may include padding blocks and displacement packing blocks. Since reversible packing is enabled, the displacement component samples or displacement packing blocks may be packed in the upper region in descending order of frequency or LOD index. For example, the padding blocks and displacement packing blocks may be placed in the upper region in reverse raster scan order, so that the padding blocks precede the displacement packing blocks in reverse raster scan order, and the frequency or LOD index of the subsequent displacement packing blocks is not lower than the frequency or LOD index of any displacement packing blocks that precede the subsequent displacement packing blocks in reverse raster scan order in the upper region. The lower region may not include any displacement wavelet coefficients, but may include padding. The padding may include rows of padding samples.

[0175] One or more displacement packed blocks in the upper region may or may not include padding samples. A partially padded block having both displacement wavelet coefficients belonging to the LOD and padding may not be allowed to precede a non-padded packed block having displacement wavelet coefficients belonging to the LOD but no padding in reverse raster scan order in the upper region.

[0176] 13A , following the two padding blocks in reverse raster scan order are three non-padded packed blocks belonging to lod0. Following the three non-padded packed blocks belonging to lod0 in reverse raster scan order is a partially padded packed block belonging to lod0. Following the partially padded packed block belonging to lod0 in reverse raster scan order is three non-padded packed blocks belonging to lod1. Following the three non-padded packed blocks belonging to lod1 in reverse raster scan order is a partially padded packed block belonging to lod1. Following the partially padded packed block belonging to lod1 in reverse raster scan order is four non-padded packed blocks belonging to lod2. Following the four non-padded packed blocks belonging to lod2 in reverse raster scan order is a partially padded packed block belonging to lod2. Both the encoder and decoder can calculate the number of padded packed blocks required for each LOD based on the number of samples in each LOD (rounded to the next higher integer), and determine the number of padded packed blocks required at the beginning of the reverse scan by the packed blocks in the upper area.

[0177] In an embodiment, referring to Figure 13B, the starting block for lod0 in the reverse packing case may start at the lower right corner of the upper region.There may be zero or more padding blocks to the left of the topmost row of blocks in the upper region.

[0178] 13B , the upper region may include padding blocks and displacement packing blocks. Since reversible packing is enabled, the displacement component samples or displacement packing blocks may be packed in the upper region in descending order of frequency or LOD index. For example, the padding blocks and displacement packing blocks may be placed in the upper region in reverse raster scan order, so that the displacement packing block precedes the padding block in reverse raster scan order, and the frequency or LOD index of the subsequent displacement packing block is not lower than the frequency or LOD index of any displacement packing block that precedes the subsequent displacement packing block in reverse raster scan order in the upper region. The lower region may not include any displacement wavelet coefficients, but includes padding. The padding may include rows of padding samples.

[0179] One or more displacement packed blocks in the upper region may or may not include padding samples. A partially padded block having both displacement wavelet coefficients belonging to the LOD and padding may not be allowed to precede a non-padded packed block having displacement wavelet coefficients belonging to the LOD but no padding in reverse raster scan order in the upper region.

[0180] 13B , following the three non-filled packed blocks belonging to lod0 in reverse raster scan order is one partially filled packed block belonging to lod0. Following the one partially filled packed block belonging to lod0 in reverse raster scan order is three non-filled packed blocks belonging to lod1. Following the three non-filled packed blocks belonging to lod1 in reverse raster scan order is one partially filled packed block belonging to lod1. Following the one partially filled packed block belonging to lod1 in reverse raster scan order is four non-filled packed blocks belonging to lod2. Following the four non-filled packed blocks belonging to lod2 in reverse raster scan order is one partially filled packed block belonging to lod2. Following the one partially filled packed block belonging to lod2 in reverse raster scan order is two filling blocks. Both the encoder and decoder can calculate the number of packed blocks required for each LOD based on the number of samples in each LOD (rounded to the next higher integer), and determine the number of filling packed blocks required at the end of the reverse scan by the packed blocks in the upper area.

[0181] In order to achieve partial decoding of the shifted wavelet coefficients belonging to one or more LODs independently of the shifted wavelet coefficients belonging to other LODs, it may be necessary to align the boundaries of the slices with the boundaries of the codec basic blocks. For the boundaries of the slices that are aligned with the boundaries of the codec basic blocks, it may be necessary that the size of the packing block can be an integer multiple of the size of the codec basic block, where the integer can be equal to or greater than 1. These restrictions may apply not only to shifted video frames divided into slices, but also to shifted video frames with different frame structures. For example, restrictions may apply to shifted video frames divided into multiple rectangular areas. Multiple rectangular areas can be partially or independently decodable. In an embodiment, when the video encoder 415 uses HEVC as a video codec, the rectangular area can be a tile. In an embodiment, when the video encoder 415 uses Versatile Video Coding (VVC) as a video codec, the rectangular area can be a tile or a sub-picture. Below, a shifted video frame divided into tiles will be described with reference to FIG. 13C.

[0182] 13C , the starting block of lod0 in the reverse packing case may begin at the lower right corner of the upper region.There may be zero or more padding blocks to the left of the topmost row of blocks in the upper region.

[0183] 13C , the upper area may include padding blocks and displacement packing blocks. The displacement video frame is divided into multiple rectangular areas. Multiple LODs may be associated with corresponding ones of the multiple rectangular areas. A corresponding rectangular area contains displacement wavelet coefficients belonging to the LOD associated with the corresponding rectangular area, and does not contain displacement wavelet coefficients belonging to the LOD not associated with the corresponding rectangular area. Since reversible packing is enabled, the displacement component samples or displacement packing blocks can be packed in the upper area in descending order of frequency or LOD index. For example, padding blocks and displacement packing blocks may be placed in the upper area so that the packing blocks containing displacement wavelet coefficients belonging to the LOD with a lower LOD index are arranged in reverse raster scan order before the packing blocks containing displacement wavelet coefficients belonging to the LOD with a higher LOD index. The rectangular area may be placed in the upper area so that the rectangular area including the packing blocks containing displacement wavelet coefficients belonging to the LOD with a lower LOD index are arranged in reverse raster scan order before the rectangular area including the packing blocks containing displacement wavelet coefficients belonging to the LOD with a higher LOD index. Within each rectangular region, partially filled blocks may not be allowed to precede non-filled packed blocks in reverse raster scan order. Within each rectangular region, fully filled blocks may not be allowed to precede partially filled blocks in reverse raster scan order. The lower region may not include any shifted wavelet coefficients but may include padding. The padding may include rows of padding samples.

[0184] In order to achieve partial decoding of the shifted wavelet coefficients belonging to one or more LODs, in particular but not limited to, independent of the shifted wavelet coefficients belonging to other LODs, it may be necessary to align the boundaries of the rectangular area with the boundaries of the codec basic block. For the boundaries of the rectangular area to be aligned with the boundaries of the codec basic block, it may be necessary that the size of the packing block can be an integer multiple of the size of the codec basic block, where the integer can be equal to or greater than 1. The unit block used to define or configure a partially decodable or independently decodable area for the video coding standard used by the video encoder 415 or the largest block used to configure a partially decodable or independently decodable area for the video coding standard used by the video encoder 415 can be referred to as a basic block or a codec basic block. In an embodiment, when the video encoder 415 uses HEVC as the video codec, the partially decodable or independently decodable area can be a tile or a slice, and the unit block can be a coding tree unit (CTU). In an embodiment, when the video encoder 415 uses AVC as the video codec, the partially decodable or independently decodable area can be a slice, and the unit block can be a macroblock. In an embodiment, when the video encoder 415 uses VVC as a video codec, the partially decodable or independently decodable area may be a sub-picture, a tile, or a slice, and the unit block may be a coding tree unit (CTU).

[0185] Figure 14 Shown is the placement of displaced component samples and padding in a displaced video frame when reversible packing is enabled, according to an embodiment.

[0186] In the examples, reference Figure 14 , the start of the data for each LOD in the reverse scan can be adjusted so that each LOD ends at a block boundary.

[0187] Reference Figure 14 , the upper area may include padding blocks and packing blocks. Since reversible packing is enabled, the displacement component samples or packing blocks can be packed in the upper area in descending order of frequency or LOD index. For example, the padding blocks and packing blocks can be placed in the upper area in reverse raster scan order, so that the padding blocks are arranged before the packing blocks in reverse raster scan order, and the frequency or LOD index of the subsequent packing blocks is not lower than the frequency or LOD index of any packing blocks in the upper area that are arranged before the subsequent packing blocks in reverse raster scan order. The lower area may not include any displacement wavelet coefficients, but includes padding. The padding may include rows of padding samples.

[0188] One or more packed blocks in the upper region may or may not include padding samples. A non-padded packed block having shifted wavelet coefficients belonging to the LOD but no padding may not be allowed to precede a partially padded block having both shifted wavelet coefficients belonging to the LOD and padding in reverse raster scan order in the upper region.

[0189] Reference Figure 14 , following the two padding blocks in reverse raster scan order is a partially padded packed block belonging to lod0. Following the partially padded packed block belonging to lod0 in reverse raster scan order are three non-padded packed blocks belonging to lod0. Following the three non-padded packed blocks belonging to lod0 in reverse raster scan order is a partially padded packed block belonging to lod1. Following the partially padded packed block belonging to lod1 in reverse raster scan order are three non-padded packed blocks belonging to lod1. Following the three non-padded packed blocks belonging to lod1 in reverse raster scan order is a partially padded packed block belonging to lod2. Following the partially padded packed block belonging to lod2 in reverse raster scan order are four non-padded packed blocks belonging to lod2. When a packed block is a partially padded block, the encoder and decoder can determine where the non-padded data starts in the reverse scan based on the number of samples in the LOD.

[0190] The placement method described above for LOD data applies when padding is at the top of a displaced video frame or when forward packing is used instead of reverse packing. The placement method described above for LOD data applies when padding is at the top of a displaced video frame and the packing order is forward. Simple modifications may be necessary to account for different scan and padding row placements.

[0191] Figure 15 1 shows the placement of displacement component samples in a displacement video frame when reversible packing is enabled according to an embodiment.

[0192] In the examples, reference Figure 15 , the size of the v-grid packing block may be smaller than the packed super block size or smaller than the equivalent codec basic block size, and the codec basic block size may be an integer multiple of the v-grid packing block size (in terms of both rows and columns), where the integer may be equal to or greater than 1. In this scenario, a sub-block scan may be used from the v-grid packing block to the packed super block. In an embodiment, an inverse Morton scan may be used to place the LOD samples in the v-grid packing block. Then, an inverse raster scan may be used to place the packing block within the packed super block. Then, an inverse raster scan may be used to place the packed super block in the displaced video frame. Instead of inverse Morton and inverse raster scans, other scans may be used. Reference Figure 15, reversible packing is being used. In an embodiment, when reversible packing is disabled, forward Morton and raster scan may be used.

[0193] In the following, the syntax and semantics of packed blocks and packed super blocks will be described.

[0194] In an embodiment, as a requirement for bitstream compliance, the Packed Super Block size may be an integer multiple of the Packed Block size in both the horizontal and vertical directions, where the integer may be equal to or greater than 1. For example, the width of the Packed Super Block may be an integer multiple of the width of the Packed Block and the height of the Packed Super Block may be an integer multiple of the height of the Packed Block. In an embodiment, the encoder may select the Packed Super Block size to be equal to the codec block size. This condition ensures that the boundaries of the codec block and the V-grid Packed Block are aligned.

[0195] In WD2.0 of VDMC (WD2.0 of V-DMC, ISO / IEC JTC1 / SC29 / WG7 N00546, online, January 2023), the packing block size is not signaled. However, the following describes a syntax element for signaling the size of the packing block size in the v-grid bitstream. In an embodiment, the packing block size may be limited to a power of 2, with a minimum size of 4. In this case, the syntax is as shown in Table 1:

[0196]

Table 1

[0197] grammar Descriptor log2_packing_block_size_minus2 ue(v)

[0198] The syntax element log2_packing_block_size_minus2 plus 2 may indicate the size of the packing block size. The descriptor ue(v) may indicate an unsigned integer 0th order exponential Golomb coding syntax element with the left bit first. The packing block size blockSize may be derived as shown in Equation 2:

[0199] Equation 2

[0200] blockSize=1<<(log2_packing_block_size_minus2+2)

[0201] Instead of a minimum size of 4, other powers of 2 may be chosen as the minimum size. In an embodiment, when the minimum size is chosen to be 8, the syntax element log2_min_packing_block_size_minus3 may be signaled instead of the syntax element log2_packing_block_size_minus2.

[0202] Additionally, the packing super block size may be signaled by indicating the difference in power of 2 with respect to the packing block size. In an embodiment, the packing super block size may be signaled in the v-trellis bitstream by the syntax element log2_diff_packing_super_block_size as follows:

[0203]

Table 2

[0204] grammar Descriptor log2_diff_packing_super_block_size ue(v)

[0205] The syntax element log2_diff_packing_super_block_size may indicate the size of the packed super block size. The packed super block size, superBlockSize, may be derived as shown in Equation 3:

[0206] Equation 3

[0207] superBlockSize=(1<<(log2_diff_packing_super_block_size)*blockSize

[0208] In an embodiment, the number of displacement coefficients in each LOD can be signaled in the bitstream to the decoder 500. This can be signaled at the sequence level, grid frame level, sub-grid level, or as metadata in an SEI message, etc. In an embodiment, the decoder 500 can use the number of displacement coefficients in each LOD to determine the end of the LOD. This allows the decoder 500 to decode the LOD to any level it wants.

[0209] In an embodiment, the v-grid bitstream may include an Atlas sequence parameter set extended raw byte sequence payload (RBSP). The syntax of the Atlas sequence parameter set extended RBSP is shown in Table 3 below:

[0210]

Table 3

[0211]

[0212] The syntax element asps_vmc_ext_packing_method may specify whether the displacement component samples are packed in ascending or descending order. For example, the syntax element asps_vmc_ext_packing_method equal to 0 may specify that the displacement component samples are packed in ascending order. The syntax element asps_vmc_ext_packing_method equal to 1 may specify that the displacement component samples are packed in descending order.

[0213] When the syntax element asps_vmc_ext_packing_method indicates that the displacement component samples are packed in ascending order, the scanning order of the packed blocks may be forward. When the syntax element asps_vmc_ext_packing_method indicates that the displacement component samples are packed in descending order, the scanning order of the packed blocks may be reverse or backward.

[0214] The syntax element asps_vmc_ext_1D_displacement_flag may specify whether only a single component of the displacement is present in the compressed geometry video, or whether all three components of the displacement are present in the compressed geometry video. For example, the syntax element asps_vmc_ext_1D_displacement_flag equal to 1 may specify that only the normal (or x) component of the displacement is present in the compressed geometry video. The remaining two components may be inferred to be 0. The syntax element asps_vmc_ext_1D_displacement_flag equal to 0 may specify that all three components of the displacement are present in the compressed geometry video.

[0215] In the following, we will Figures 16 and 17 The operation of the image unpacker 523 which performs inverse image packing of wavelet coefficients is described.

[0216] Figure 16 is a source code illustrating the operation of an image depacketizer according to an embodiment. Figure 17 It shows that according to Figure 16 A flow chart illustrating the operation of an image depacketizer of the illustrated embodiment.

[0217] The image depacketizer 523 may receive as input a variable named width, a variable named height, a variable named bitDepth, a 3D array named dispQuantCoeffFrame, a variable named blockSize, and a variable named positionCount. The variable named width may indicate the width of the displaced video frame. The variable named height may indicate the height of the displaced video frame. The variable named bitDepth may indicate the bit depth of the displaced video frame. The 3D array named dispQuantCoeffFrame may be a 3D array of packed quantized displaced wavelet coefficients, with a size of width × height × 3. The variable named blockSize may indicate the size of the packed block. The variable named positionCount may indicate the number of positions in the subdivided subgrid.

[0218] The image depacketizer 523 may output a 2D array dispQuantCoeffArray. The 2D array dispQuantCoeffArray may be a 2D array indicating quantized displacement wavelet coefficients, and its size is positionCount×3.

[0219] The image depacketizer 523 may set the variable DisplacementDim based on the syntax element asps_vmc_ext_1D_displacement_flag. For example, if the syntax element asps_vmc_ext_1D_displacement_flag is equal to 1, the image depacketizer 523 may set the variable DisplacementDim to 1. Otherwise, if the syntax element asps_vmc_ext_1D_displacement_flag is equal to 0, the image depacketizer 523 may set the variable DisplacementDim to 3. The variable DisplacementDim may indicate the number of dimensions of the displacement field.

[0220] The image unpacker 523 may initialize the 2D array dispQuantCoeffArray to 0.

[0221] The image unpacker 523 may perform inverse image packing of the wavelet coefficients as shown in Table 4.

[0222]

Table 4

[0223]

[0224]

[0225] Reference Figure 16 , in operation 1601, the image unpacker 523 may calculate variables pixelsPerBlock, widthInBlocks, shift, blockCount, heightInBlocks, origHeight, and paddedHeight, as shown in Table 4. For example, the image unpacker 523 may calculate the variable pixelsPerBlock based on the variable blockSize. The variable pixelsPerBlock may indicate the number of pixels in a packed block. The image unpacker 523 may calculate the variable widthInBlocks based on the variables width and blockSize. The variable widthInBlocks may indicate the width of the shifted video frame in units of packed blocks. The image unpacker 523 may calculate the variable shift based on the variable bitDepth. The image unpacker 523 may calculate the variable blockCount based on the variables positionCount and pixelsPerBlock. The variable blockCount may indicate the number of packed blocks containing quantized shifted wavelet coefficients. The image unpacker 523 may calculate the variable heightInBlocks based on the variables blockCount and widthInBlocks. Figures 7 to 11 , the variable heightInBlocks may indicate the height of the upper area in units of packed blocks. The image depacketizer 523 may calculate the variable origHeight based on the variable heightInBlocks and blockSize. Figures 7 to 11 The variable origHeight may indicate the number of rows of samples in the lower area. The image depacketizer 523 may calculate the variable paddedHeight based on the variable origHeight. The variable paddedHeight may indicate the number of padding rows.

[0226] In operation 1603 , the image depacketizer 523 may determine whether the syntax element asps_vmc_ext_1D_displacement_flag is equal to 1.

[0227] If the syntax element asps_vmc_ext_1D_displacement_flag is equal to 0, that is, if the syntax element asps_vmc_ext_1D_displacement_flag indicates that three components of displacement exist in the compressed geometry video, in operation 1605 , the image depacketizer 523 may calculate the variable start based on the variables paddedHeight, origHeight, and width.

[0228] If the syntax element asps_vmc_ext_1D_displacement_flag is equal to 1, that is, if the syntax element asps_vmc_ext_1D_displacement_flag indicates that only a single component of displacement exists in the compressed geometry video, in operation 1607 , the image depacketizer 523 may calculate the variable start based on the variables width and height.

[0229] In operation 1608 , the image unpacker 523 may initialize a variable v to 0. The variable v may indicate the index of the current quantized displaced wavelet coefficient, specifically, the index in the array dispQuantCoeffArray of quantized displaced wavelet coefficients.

[0230] In operation 1609 , the image unpacker 523 may determine whether the variable v is less than the variable positionCount.

[0231] If the variable v is less than the variable positionCount, in operation 1611, the image depacketizer 523 may calculate the variables v0, blockIndex, indexWithinBlock, x0, y0, x, y, x1, and y1. The variables may be calculated as shown in Table 4. For example, the image depacketizer 523 may calculate the variable v0 based on the syntax element asps_vmc_ext_packing_method and the variables start and v. Considering the syntax element asps_vmc_ext_packing_method, the variable v0 may indicate the index of the current quantized displacement wavelet coefficient. If the syntax element asps_vmc_ext_packing_method is equal to 1, the image depacketizer 523 may determine the variable v0 by subtracting the value of the variable v from the value of the variable start. If the syntax element asps_vmc_ext_packing_method is equal to 0, the image depacketizer 523 may determine the value of the variable v as the variable v0. The image depacketizer 523 may calculate the variable blockIndex based on the variable v and pixelsPerBlock. The variable blockIndex may indicate the index of the current packed block containing the current quantized shifted wavelet coefficient. The image depacketizer 523 may calculate the variable indexWithinBlock based on the variables v and pixelsPerBlock. The variable indexWithinBlock may indicate the index of the current quantized shifted wavelet coefficient within the current packed block. The image depacketizer 523 may calculate the variable x0 based on the variables blockIndex, widthInBlocks, and blockSize. The image depacketizer 523 may calculate the variable y0 based on the variables blockIndex, widthInBlocks, and blockSize. The variables x0 and y0 may indicate the position of the top left sample point of the current packed block to which the current quantized shifted wavelet coefficient belongs. The image depacketizer 523 may calculate the variables x and y using the function computeMorton2D. The variables x and y may indicate the position of the current quantized shifted wavelet coefficient within the current packed block. The image depacketizer 523 may calculate the variable x1 based on the variables x0 and x. The image depacketizer 523 may calculate the variable y1 based on the variables y0 and y. The variables x1 and y1 may indicate the position of the current quantized shifted wavelet coefficient within the shifted video frame. The image unpacker 523 may initialize the variable d to 0. The variable d may indicate the dimension of the components of the current quantized shifted wavelet coefficient.

[0232] The function computeMorton2D(i) can be defined as shown in the following function 1.

[0233] Function 1

[0234] (x,y)=computeMorton2D(i){

[0235] x=extracOddBits(i>>1)

[0236] y = extracOddBits(i)

[0237] }

[0238] The function extracOddBits(x) may be defined as shown in Function 2 below.

[0239] Function 2

[0240] x=extracOddBits(x){

[0241] x=x&0x55555555

[0242] x=(x|(x>>1))&0x33333333

[0243] x=(x|(x>>2))&0x0F0F0F0F

[0244] x=(x|(x>>4))&0x00FF00FF

[0245] x=(x|(x>>8))&0x0000FFFF

[0246] }

[0247] In operation 1613 , the image depacketizer 523 may determine whether the variable d is less than the variable DisplacementDim.

[0248] If the variable d is less than the variable DisplacementDim, in operation 1614 , the image unpacker 523 may determine whether DecGeoChromaFormat is equal to 4:2:0.

[0249] If DecGeoChromaFormat is equal to 4:2:0, then in operation 1615, the image unpacker 523 may convert the 2D array dispQuantCoeffArray of quantized displacement wavelet coefficients into a displacement video frame dispQuantCoeffFrame based on the variables v, d, x1, y1, origHeight, and shift. In an embodiment, the image unpacker 523 may set the vth quantized displacement wavelet coefficient of the current dimension d in the 2D array dispQuantCoeffArray to be equal to the quantized displacement wavelet coefficient of the current dimension d at (x1, d*origHeight+y1) in the displacement video frame dispQuantCoeffFrame minus the variable shift.

[0250] If DecGeoChromaFormat is not equal to 4:2:0, the image unpacker 523 may convert the 2D array of quantized displacement wavelet coefficients, dispQuantCoeffArray, into a displacement video frame, dispQuantCoeffFrame, based on the variables v, d, x1, y1, and shift in operation 1617. In an embodiment, the image unpacker 523 may set the vth quantized displacement wavelet coefficient of the current dimension d in the 2D array dispQuantCoeffArray to be equal to the quantized displacement wavelet coefficient at (x1, y1) of the current dimension d in the displacement video frame dispQuantCoeffFrame minus the variable shift.

[0251] In operation 1619 , the image depacketizer 523 may increase the variable d by 1 and then proceed to operation 1613 .

[0252] If the variable d is not less than the variable DisplacementDim, in operation 1621 , the image depacketizer 523 may increase the variable v by 1 and then proceed to operation 1609 .

[0253] If the variable v is not less than the variable positionCount, in operation 1623 , the image unpacker 523 may output a 2D array dispQuantCoeffArray.

[0254] Figure 17 is a flowchart illustrating the operation of the image depacketizer 523 according to an embodiment.

[0255] Unless otherwise stated, Figure 16 The description will be applied to Figure 17 Example of .

[0256] The image unpacker 523 may perform inverse image packing of the wavelet coefficients as shown in Table 5.

[0257]

Table 5

[0258]

[0259]

[0260] Reference Figure 17, in operation 1701, the image unpacker 523 may calculate variables pixelsPerBlock, widthInBlocks, shift, blockCount, heightInBlocks, and origHeight, as shown in Table 5. For example, the image unpacker 523 may calculate the variable pixelsPerBlock based on the variable blockSize. The variable pixelsPerBlock may indicate the number of pixels in a packed block. The image unpacker 523 may calculate the variable widthInBlocks based on the variables width and blockSize. The variable widthInBlocks may indicate the width of the shifted video frame in units of packed blocks. The image unpacker 523 may calculate the variable shift based on the variable bitDepth. The image unpacker 523 may calculate the variable blockCount based on the variables positionCount and pixelsPerBlock. The variable blockCount may indicate the number of packed blocks containing quantized shifted wavelet coefficients. The image unpacker 523 may calculate the variable heightInBlocks based on the variables blockCount and widthInBlocks. Figures 7 to 11 , the variable heightInBlocks may indicate the height of the upper area in units of packed blocks. The image depacketizer 523 may calculate the variable origHeight based on the variable heightInBlocks and blockSize. Figures 7 to 11 , the variable origHeight can indicate the number of rows of sample points in the upper area.

[0261] In operation 1705 , the image unpacker 523 may calculate a variable start based on the variables origHeight and width.

[0262] In operation 1708 , the image unpacker 523 may initialize a variable v to 0. The variable v may indicate the index of the current quantized displaced wavelet coefficient, specifically, the index in the array dispQuantCoeffArray of quantized displaced wavelet coefficients.

[0263] In operation 1709 , the image unpacker 523 may determine whether the variable v is less than the variable positionCount.

[0264] If the variable v is less than the variable positionCount, in operation 1711, the image depacketizer 523 may calculate the variables v0, blockIndex, indexWithinBlock, x0, y0, x, y, x1, and y1. The variables may be calculated as shown in Table 5. For example, the image depacketizer 523 may calculate the variable v0 based on the syntax element asps_vmc_ext_packing_method and the variables start and v. Considering the syntax element asps_vmc_ext_packing_method, the variable v0 may indicate the index of the current quantized displacement wavelet coefficient. If the syntax element asps_vmc_ext_packing_method is equal to 1, the image depacketizer 523 may determine the variable v0 by subtracting the value of the variable v from the value of the variable start. If the syntax element asps_vmc_ext_packing_method is equal to 0, the image depacketizer 523 may determine the value of the variable v as the variable v0. The image depacketizer 523 may calculate the variable blockIndex based on the variable v and pixelsPerBlock. The variable blockIndex may indicate the index of the current packed block containing the current quantized shifted wavelet coefficient. The image depacketizer 523 may calculate the variable indexWithinBlock based on the variables v and pixelsPerBlock. The variable indexWithinBlock may indicate the index of the current quantized shifted wavelet coefficient within the current packed block. The image depacketizer 523 may calculate the variable x0 based on the variables blockIndex, widthInBlocks, and blockSize. The image depacketizer 523 may calculate the variable y0 based on the variables blockIndex, widthInBlocks, and blockSize. The variables x0 and y0 may indicate the position of the top left sample point of the current packed block to which the current quantized shifted wavelet coefficient belongs. The image depacketizer 523 may calculate the variables x and y using the function computeMorton2D. The variables x and y may indicate the position of the current quantized shifted wavelet coefficient within the current packed block. The image depacketizer 523 may calculate the variable x1 based on the variables x0 and x. The image depacketizer 523 may calculate the variable y1 based on the variables y0 and y. The variables x1 and y1 may indicate the position of the current quantized shifted wavelet coefficient within the shifted video frame. The image unpacker 523 may initialize the variable d to 0. The variable d may indicate the dimension of the components of the current quantized shifted wavelet coefficient.

[0265] In operation 1713 , the image depacketizer 523 may determine whether the variable d is less than the variable DisplacementDim.

[0266] If the variable d is less than the variable DisplacementDim, in operation 1714 , the image unpacker 523 may determine whether DecGeoChromaFormat is equal to 4:2:0.

[0267] If DecGeoChromaFormat is equal to 4:2:0, then in operation 1715, the image unpacker 523 may convert the 2D array dispQuantCoeffArray of quantized displacement wavelet coefficients into a displacement video frame dispQuantCoeffFrame based on the variables v, d, x1, y1, origHeight, and shift. In an embodiment, the image unpacker 523 may set the vth quantized displacement wavelet coefficient of the current dimension d in the 2D array dispQuantCoeffArray to be equal to the quantized displacement wavelet coefficient of the current dimension d at (x1, d*origHeight+y1) in the displacement video frame dispQuantCoeffFrame minus the variable shift.

[0268] If DecGeoChromaFormat is not equal to 4:2:0, the image unpacker 523 may convert the 2D array of quantized displacement wavelet coefficients, dispQuantCoeffArray, into a displacement video frame, dispQuantCoeffFrame, based on the variables v, d, x1, y1, and shift in operation 1717. In an embodiment, the image unpacker 523 may set the vth quantized displacement wavelet coefficient of the current dimension d in the 2D array dispQuantCoeffArray to be equal to the quantized displacement wavelet coefficient at (x1, y1) of the current dimension d in the displacement video frame dispQuantCoeffFrame minus the variable shift.

[0269] In operation 1719 , the image depacketizer 523 may increase the variable d by 1 and then proceed to operation 1713 .

[0270] If the variable d is not less than the variable DisplacementDim, in operation 1721 , the image depacketizer 523 may increase the variable v by 1 and then proceed to operation 1709 .

[0271] If the variable v is not less than the variable positionCount, in operation 1723 , the image unpacker 523 may output a 2D array dispQuantCoeffArray.

[0272] The various illustrative blocks, units, modules, components, methods, operations, instructions, terms, and algorithms may be implemented or performed with processing circuitry.

[0273] Reference to an element in the singular is not intended to mean one and only one, unless specifically stated otherwise, but rather to mean one or more. For example, "a" module may refer to one or more modules. Without further constraints, an element preceded by "a," "an," "the," or "said" does not exclude the presence of additional identical elements.

[0274] Titles and subtitles, if any, are used for convenience only and do not limit the subject technology. The term "exemplary" is used to mean serving as an example or illustration. To the extent that the terms "including," "having," "carrying," "comprising," and the like are used, such terms are intended to be inclusive in a manner similar to the term "comprising" (as interpreted when "comprising" is used as a transitional word in a claim). Relational terms such as first and second may be used to distinguish one entity or action from another without necessarily requiring or implying any actual such relationship or order between those entities or actions.

[0275] Phrases such as aspect, this aspect, another aspect, some aspects, one or more aspects, implementation, this implementation, another implementation, some implementations, one or more implementations, embodiment, this embodiment, another embodiment, some embodiments, one or more embodiments, configuration, this configuration, another configuration, some configurations, one or more configurations, subject technology, disclosure, the present disclosure, other variations thereof, etc. are for convenience and do not imply that the disclosure associated with such (one or more) phrases is essential to the subject technology or that such disclosure applies to all configurations of the subject technology. The disclosure associated with such (one or more) phrases may apply to all configurations or one or more configurations. The disclosure associated with such (one or more) phrases may provide one or more examples. Phrases such as aspect or some aspects may refer to one or more aspects and vice versa, and this applies similarly to the other aforementioned phrases.

[0276] The phrase "at least one of" preceding a series of items (with the terms "and" or "or" used to separate any items) modifies the list as a whole, rather than each member of the list. The phrase "at least one of" does not require selection of at least one item; rather, the phrase allows for a meaning that includes at least one of any one item, and / or at least one of any combination of items, and / or at least one of each item. For example, each of the phrases "at least one of A, B, and C" or "at least one of A, B, or C" means only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0277] It will be understood that the specific order or level of disclosed steps, operations or processes is an illustration of an exemplary method. Unless otherwise expressly stated, it will be understood that the specific order or level of steps, operations or processes can be performed in different orders. Some of the steps, operations or processes can be performed simultaneously, or can be performed as part of one or more other steps, operations or processes. The attached method claims (if any) present the elements of various steps, operations or processes in a sample order and are not intended to be limited to the specific order or level presented. These can be performed serially, linearly, in parallel or in different orders. It should be understood that the instructions, operations and systems described can usually be integrated together in a single software / hardware product or encapsulated in multiple software / hardware products.

[0278] This disclosure is provided to enable anyone skilled in the art to practice the various aspects described herein. In some instances, well-known structures and components are shown in block diagram form to avoid blurring the concepts of the subject technology. This disclosure provides various examples of the subject technology, and the subject technology is not limited to these examples. Various modifications to these aspects will be apparent to those skilled in the art, and the principles described herein may be applied to other aspects.

[0279] All structural and functional equivalents to the elements of various aspects described throughout this disclosure that are known or later become known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

[0280] The title, background technology, description of the figures, abstract and drawings are incorporated into this disclosure and are provided as illustrative examples of the present disclosure rather than as limiting descriptions. They are submitted with the understanding that they will not be used to limit the scope or meaning of the claims. In addition, in the detailed description, the description may provide illustrative examples, and various features may be combined together in various embodiments for the purpose of streamlining the present disclosure. The method of the present disclosure should not be interpreted as reflecting an intention that the claimed subject matter requires more features than those expressly recited in each claim. On the contrary, as reflected in the following claims, the inventive subject matter lies in less than all the features of a single disclosed configuration or operation. The appended claims are incorporated into the detailed description, and each claim exists independently as a separately claimed subject matter.

[0281] The embodiments are provided only as examples for understanding the present invention. They are not intended to and should not be interpreted as limiting the scope of the present invention in any way. Although embodiments and examples have been provided, it will be apparent to those skilled in the art based on the disclosure herein that the illustrated embodiments and examples may be modified without departing from the scope of the present invention.

[0282] The claims are not intended to be limited to the aspects described herein, but are to be accorded the full scope consistent with the language of the claims and encompass all legal equivalents. Nevertheless, the claims are not intended to encompass subject matter that does not satisfy applicable patent law requirements, nor should they be interpreted in such a manner.

Claims

1. A device comprising: a communication interface configured to receive a compressed bitstream, the compressed bitstream comprising an encoded displacement bitstream and packing method information indicating whether the displacement component samples are packed in ascending order or descending order, wherein the displacement component samples are packed in ascending order when the packing method information is equal to a first value, and the displacement component samples are packed in descending order when the packing method information is equal to a second value, wherein the first value is 0 and the second value is 1; and A processor operatively coupled to the communication interface, the processor being configured to: Decode the packaging method information; performing video decoding on the encoded displacement bitstream to generate a displacement video frame; performing image unpacking on the shifted video frame to generate an array of quantized shifted wavelet coefficients, wherein the shifted video frame includes padding at a lower region of the shifted video frame, regardless of whether the shifted component samples are packed in descending order; dequantizing the array of quantized shifted wavelet coefficients to generate shifted wavelet coefficients; and Perform inverse wavelet transform on the displacement wavelet coefficients to generate displacement component samples.

2. The device according to claim 1, wherein Regardless of whether the packing method information indicates that the displacement component samples are packed in descending order, the upper area of ​​the displacement video frame includes quantized displacement wavelet coefficients, and the lower area of ​​the displacement video frame does not include any quantized displacement wavelet coefficients and includes samples for padding.

3. The device according to claim 2, wherein If the packing method information indicates that the displacement component samples are packed in descending order, image unpacking is performed on the displacement video frame starting from the lower right block of the upper area.

4. The device according to claim 3, wherein The lower right block of the upper region comprises quantized displacement wavelet coefficients belonging to the lowest level of detail (LOD).

5. The apparatus according to claim 1, wherein Multiple packed blocks are placed in the shifted video frame.

6. The device according to claim 5, wherein The plurality of packed blocks include one or more non-filled blocks and partially filled blocks, Each of the one or more non-filling blocks includes quantized displaced wavelet coefficients belonging to a first level of detail (LOD) and does not include any samples used for filling, and The partially filled block includes one or more quantized shifted wavelet coefficients belonging to the first LOD and one or more sample points used for filling.

7. The apparatus according to claim 6, wherein When the packing method information indicates that the displacement component samples are packed in descending order, the one or more non-filling blocks are arranged before the partially filled block in a scanning order of blocks.

8. The apparatus according to claim 7, wherein The blocks are scanned in reverse raster order.

9. The apparatus according to claim 6, wherein The partially filled block precedes one or more fully filled blocks, each of which does not include any quantized shifted wavelet coefficients to be decoded and includes samples used for filling.

10. The apparatus according to claim 6, wherein Each of the one or more non-filling blocks is not allowed to include quantized shifted wavelet coefficients belonging to a second LOD, the second LOD being different from the first LOD.

11. The apparatus according to claim 1, wherein Multiple packing blocks are placed in the shifted video frame, and the sizes of the multiple packing blocks are integer multiples of the size of a unit block, which is used to configure a partially decodable or independently decodable area for a video coding standard used by video decoding, and the integer is equal to or greater than 1.

12. The apparatus according to claim 11, wherein Partially decodable or independently decodable regions for the video coding standard used by the video decoding are not allowed to include quantized shifted wavelet coefficients belonging to two or more levels of detail.

13. The apparatus according to claim 11, wherein When high efficiency video coding (HEVC) is used for video decoding, the size of the unit block is the size of a coding tree unit specified in the video decoding.

14. The apparatus according to claim 11, wherein When Advanced Video Coding (AVC) is used for video decoding, the size of the unit block is the size of a macroblock specified in the video decoding.

15. A method comprising: receiving a compressed bitstream, the compressed bitstream comprising an encoded displacement bitstream and packing method information indicating whether displacement component samples are packed in ascending order or descending order, wherein the displacement component samples are packed in ascending order when the packing method information is equal to a first value, and the displacement component samples are packed in descending order when the packing method information is equal to a second value, wherein the first value is 0 and the second value is 1; Decode the packaging method information; performing video decoding on the encoded displacement bitstream to generate a displacement video frame; performing image unpacking on the shifted video frame to generate an array of quantized shifted wavelet coefficients, wherein the shifted video frame includes padding at a lower region of the shifted video frame, regardless of whether the shifted component samples are packed in descending order; dequantizing the array of quantized shifted wavelet coefficients to generate shifted wavelet coefficients; and Perform inverse wavelet transform on the displacement wavelet coefficients to generate displacement component samples.