Packing of displacement data in video frames for dynamic grid coding
By identifying the video format and setting signaling elements, and determining the packaging layout of displacement data, the compatibility problem of displacement data storage in dynamic grid-encoded video frames is solved, and efficient storage and transmission in different video formats is achieved.
Patent Information
- Application Number
- CN202380071609.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-05
- Filing Date
- 2023-10-18
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to efficiently process and store displacement data in dynamic grid-encoded video frames, especially when compatible with different video formats.
By identifying the video format for compressed video and setting signaling elements, a packaging arrangement of displacement data is determined, such as storing the x, y, and z components of displacement data in different planes of the video frame.
It realizes efficient storage and transmission of displacement data in different video formats, enhances the compatibility of encoder and decoder, and avoids dependence on 4:4:4 video format.
Smart Images

Figure CN119999214A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to multimedia devices and processing. More specifically, the present disclosure relates to improved packing of displacement data in video frames for dynamic grid coding. Background Art
[0002] Due to the ready availability of powerful handheld devices (such as smartphones), three hundred sixty degree (360°) video and three-dimensional (3D) volumetric video are becoming new ways to experience immersive content. While 360° video enables an immersive "real life", "there" experience for consumers by capturing a 360° outside-in view of the world, 3D volumetric video can provide a full six degrees of freedom (DoF) experience of immersion and movement within the content. Users can interactively change their viewpoint and dynamically view any part of the captured scene or object they desire. Display and navigation sensors can track the user's head movement in real time to determine the area of 360° video or volumetric content that the user wants to view or interact with. Multimedia data that is 3D in nature (such as point clouds or 3D polygon meshes) can be used in immersive environments. This data can be stored in a video format and encoded and compressed for transmission to other devices as a bitstream. Summary of the invention
[0003] Technical Solution The present disclosure provides packing of displacement data in video frames for dynamic grid coding.
[0004] In an embodiment, a device for decoding a video may include a memory storing one or more instructions and at least one processor configured to execute the one or more instructions stored in the memory. The at least one processor may be configured to identify a video format for a compressed video. The at least one processor may be configured to determine a displacement data packing arrangement based on at least one signaling element and one or more of the identified video formats. The processor may be configured to retrieve displacement data based on the determined displacement data packing arrangement.
[0005] In an embodiment, a method for decoding a video may include identifying a video format for a compressed video from a format variable. The method may include determining a displacement data packing arrangement based on one or more of at least one signaling element and the identified video format. The method may include retrieving displacement data based on the determined displacement data packing arrangement.
[0006] In an embodiment, a device for encoding a video includes a memory storing one or more instructions and at least one processor configured to execute the one or more instructions stored in the memory. The at least one processor may be configured to determine a video format for the video. The at least one processor may be configured to set at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for displacement data. The at least one processor may be configured to encode the video into a bitstream according to the displacement data packing arrangement.
[0007] In an embodiment, a method for encoding a video may include determining a video format for the video. The method may include setting at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for displacement data. The method may include encoding the video into a bitstream according to the displacement data packing arrangement.
[0008] In an embodiment, a computer-readable storage medium storing a bitstream may be provided, the bitstream comprising a video format for a video and at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for displacement data.
[0009] Other technical features may be apparent to those skilled in the art from the following drawings, descriptions and claims.
[0010] Before carrying out the following specific embodiments, it may be advantageous to set forth the definitions of certain words and phrases used throughout this patent document.The term "coupling" and its derivatives refer to any direct or indirect communication between two or more elements, whether or not these elements are in physical contact with each other.The terms "send", "receive" and "communication" and their derivatives cover both direct and indirect communication.The terms "include" and "comprise" and their derivatives mean including but not limited to.The term "or" is inclusive, meaning and / or.The term "associated with..." and its derivatives mean including, included in, interconnected with, included in, connected to or connected with, coupled to or coupled with, communicable with, collaborative with, interlaced, juxtaposed, close to, bound to or bound with, have, have the property of, have to or with, etc., the relationship of.The term "controller" means any device, system or part thereof that controls at least one operation.Such a controller can be implemented in hardware or a combination of hardware and software and / or firmware.The function associated with any particular controller can be centralized or distributed, whether local or remote. When used with a list of items, the phrase "at least one of" means that different combinations of one or more of the listed items can be used, and that only one of the items in the list may be required. For example, "at least one of A, B, and C" includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A and B and C.
[0011] In addition, the various functions described below may be implemented or supported by one or more computer programs, each of which is formed by a computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or a portion thereof, suitable for implementation in a suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read-only memory (ROM), a random access memory (RAM), a hard drive, a compact disk (CD), a digital video disk (DVD), or any other type of memory. "Non-transitory" computer-readable media excludes wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Non-transitory computer-readable media include media in which data can be permanently stored and media in which data can be stored and rewritten later, such as rewritable optical disks or erasable memory devices.
[0012] Definitions for certain other words and phrases are provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most, instances, such definitions apply to prior, as well as future uses of such defined words and phrases. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein like reference numerals represent like parts: Figure 1 An example communication system according to the present disclosure is shown; Figure 2 and Figure 3 An example electronic device according to the present disclosure is shown; Figure 4 An example intra-coding process according to the present disclosure is shown; Figure 5A and Figure 5B An example packaging arrangement process based on a video format according to the present disclosure is shown; Fig. 6A and Figure 6B An example packaging arrangement process based on packaging type variables according to the present disclosure is shown; Figure 7 An example staggered packing arrangement process according to the present disclosure is shown; Figure 8 An example full packaging arrangement process according to the present disclosure is shown; Fig. 9 An example split packaging arrangement process according to the present disclosure is shown; Fig.10 An example split and interleaved packing arrangement process according to the present disclosure is shown; Fig.11A and Fig. 11B An example single component packaging arrangement process based on a video format according to the present disclosure is shown; Figures 12 to 15 Additional example packaging arrangements utilizing padding in accordance with the present disclosure are shown; Figures 16 to 20 An example of a sub-grid packing arrangement according to the present disclosure is shown; Fig.21 An example encoding method for improved packing of displacement data in video frames according to the present disclosure is shown; and Fig. 22 An example decoding method for improved packing of displacement data in video frames according to the present disclosure is shown. DETAILED DESCRIPTION
[0014] Described below Figures 1 to 22The embodiments used to describe the principles of the present disclosure are for illustration only and should not be interpreted in any way as limiting the scope of the present disclosure. It will be understood by those skilled in the art that the principles of the present disclosure can be implemented in any type of appropriately arranged device or system.
[0015] As described above, due to the ready availability of powerful handheld devices (such as smartphones), three hundred and sixty degree (360°) video and three-dimensional (3D) volumetric video are becoming new ways to experience immersive content. While 360° video enables an immersive "real life", "there" experience for consumers by capturing a 360° outside-in view of the world, 3D volumetric video can provide a full six degrees of freedom (DoF) experience of immersion and movement within the content. Users can interactively change their viewpoint and dynamically view any part of the captured scene or object they desire. Display and navigation sensors can track the user's head movement in real time to determine the area of 360° video or volumetric content that the user wants to view or interact with. Multimedia data that is essentially 3D (such as point clouds or 3D polygonal meshes) can be used in an immersive environment. The data can be stored in a video format and encoded and compressed for transmission to other devices as a bitstream.
[0016] A point cloud is a collection of 3D points with properties (such as color, normal direction, reflectivity, point size, etc.) that represent the surface or volume of an object. Point clouds are common in various applications such as gaming, 3D mapping, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view playback, and six degrees of freedom (DoF) immersive media, to name a few. Point clouds, if uncompressed, typically require a large amount of bandwidth for transmission. Due to the large bitrate requirements, point clouds are typically compressed before transmission. Compressing 3D objects such as point clouds typically requires dedicated hardware. In order to avoid dedicated hardware to compress 3D point clouds, 3D point clouds can be transformed into traditional two-dimensional (2D) frames and can be compressed and later reconstructed and visible to the user.
[0017] Polygonal 3D meshes, especially triangle meshes, are another popular format for representing 3D objects. A mesh is usually composed of a set of vertices, edges, and faces that are used to represent the surface of a 3D object. A triangle mesh is a simple polygonal mesh in which the faces are simple triangles that cover the surface of the 3D object. Typically, there may be one or more attributes associated with a mesh. In a scene, one or more attributes may be associated with each vertex in the mesh. For example, a texture attribute (RGB) may be associated with each vertex. In a scene, each vertex may be associated with a pair of coordinates (u, v). The (u, v) coordinates may point to a location in a texture map associated with the mesh. For example, the (u, v) coordinates may refer to a row and column index in a texture map, respectively. A mesh may be considered a point cloud with additional connectivity information.
[0018] Point clouds or meshes can be dynamic, i.e., they can change over time. In these cases, the point cloud or mesh at a particular moment in time can be referred to as a point cloud frame or mesh frame, respectively. Since point clouds and meshes contain a large amount of data, they need to be compressed for efficient storage and transmission. This is especially true for dynamic point clouds and meshes, which can contain 60 frames per second or more.
[0019] As part of the encoding process, a base mesh may be generated using an existing mesh, and a reconstructed base mesh may be constructed from the encoded base mesh. The base mesh typically contains a smaller number of vertices than the original mesh. The reconstructed base mesh may then be subdivided into one or more subdivided meshes, and a displacement field is created for each subdivided mesh. For example, if the reconstructed base mesh includes triangles covering the surface of a 3D object, the triangles are subdivided according to multiple subdivision levels, such as creating a first subdivision mesh of four triangles for each triangle of the reconstructed base mesh, a second subdivision mesh of sixteen triangles for each triangle of the reconstructed base mesh, and so on, depending on how many subdivision levels are applied. Each displacement field represents the difference between the vertex positions of the original mesh and the subdivided mesh associated with the displacement field. That is, each displacement of each subdivision is calculated for the additional vertices introduced by the subdivision process. Each displacement field is wavelet transformed to create a level of detail (LOD) signal, which is encoded as part of the compressed bitstream. During decoding, the displacement of each displacement field is added to their associated subdivided mesh to recreate the original mesh.
[0020] The quantized LOD signal can be packed into a 2D image / video and can be losslessly compressed using an image / video encoder such as HEVC. Alternatively, the unquantized LOD signal can be packed into a 2D image / video and then compressed in a lossy manner using an image / video encoder such as HEVC. Currently, the 4:4:4 video format is used to store x, y, and z components (normal components, tangent components, and bitangent components). For example, in a 4:4:4 video format, the x component will be stored in the Y plane, the y component will be stored in the Cb plane, and the z component will be stored in the Cr plane, where the Y, Cb, and Cr planes have the same width and height. In some cases, depending on the color space used, the x component is stored in the R plane, the y component is stored in the G plane, and the z component is stored in the B plane. However, video encoders and decoders that can operate on 4:4:4 format video are not widely available, especially in hardware, limiting the usefulness of this approach because many devices that implement encoders or decoders are not even compatible with using the 4:4:4 video format. Therefore, there is a need to store x, y, and z components of displacement data in video frames using different video formats and to efficiently and effectively determine how the components should be stored in video frames according to the video formats.
[0021] The present disclosure provides improved techniques for packing displacement data in video frames for dynamic grid coding. Depending on the video format used, the present disclosure provides methods for storing x, y, and z components of displacement data in video frames, so that encoders and decoders compatible with the 4:4:4 video format are not required, and encoders and decoders compatible with other more common formats (such as 4:2:0, 4:2:2, and 4:0:0 video formats) can be used. The present disclosure further provides techniques for identifying the video format of the video to be compressed, and determining the displacement data packing arrangement (such as different schemes for packing x, y, and z components in different planes of the video frame) based on the identified video format and based on other factors.
[0022] Figure 1 An example communication system 100 according to the present disclosure is shown. Figure 1 The embodiment of the communication system 100 shown in FIG. 1 is for illustration only. Other embodiments of the communication system 100 may be used without departing from the scope of the present disclosure.
[0023] like Figure 1As shown, the communication system 100 includes a network 102 that facilitates communication between various components in the communication system 100. For example, the network 102 can communicate IP packets, frame relay frames, asynchronous transfer mode (ATM) cells, or other information between network addresses. The network 102 includes all or part of one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), a global network (such as the Internet), or any other communication system(s) at one or more locations.
[0024] In this example, the network 102 facilitates communication between the server 104 and various client devices 106-116. The client devices 106-116 may be, for example, smart phones, tablet computers, laptop computers, personal computers, TVs, interactive displays, wearable devices, HMDs, etc. The server 104 may represent one or more servers. Each server 104 includes any suitable computing or processing device that can provide computing services to one or more client devices (such as client devices 106-116). Each server 104 may, for example, include one or more processing devices, one or more memories storing one or more instructions and data, and one or more network interfaces that facilitate communication through the network 102. As described in more detail below, the server 104 may send a compressed bitstream representing a point cloud or mesh to one or more display devices (such as client devices 106-116). In an embodiment, each server 104 may include an encoder. In an embodiment, the server 104 may utilize a displacement data packaging scheme based on the video format and / or other factors to improve the encoding of the displacement.
[0025] Each client device 106-116 represents any suitable computing or processing device that interacts with at least one server (such as server 104) or (one or more) other computing devices through network 102. Client devices 106-116 include desktop computers 106, mobile phones or mobile devices 108 (such as smart phones), PDAs 110, laptop computers 112, tablet computers 114, and HMDs 116. However, any other or additional client devices may be used in the communication system 100. Smart phones represent a type of mobile device 108 that is a handheld device with a mobile operating system and an integrated mobile broadband cellular network connection for voice, short message service (SMS), and Internet data communications. HMD 116 can display a 360° scene including one or more dynamic or static 3D point clouds. In embodiments, any of the client devices 106-116 may include an encoder, a decoder, or both. For example, mobile device 108 may record a 3D volumetric video and then encode the video so that the video can be sent to one of the client devices 106-116. In an example, laptop computer 112 may be used to generate a 3D point cloud or mesh, which is then encoded and sent to one of client devices 106 - 116 .
[0026] In this example, some of the client devices 108-116 communicate indirectly with the network 102. For example, the mobile device 108 and the PDA 110 communicate via one or more base stations 118, such as cellular base stations or eNodeBs (eNBs). In addition, the laptop 112, the tablet computer 114, and the HMD 116 communicate via one or more wireless access points 120, such as IEEE 802.11 wireless access points. Note that these are for illustration only, and each client device 106-116 may communicate directly with the network 102 or indirectly with the network 102 via any suitable intermediate device(s) or network(s). In an embodiment, the server 104 or any of the client devices 106-116 may be used to compress a point cloud or mesh, generate a bitstream representing the point cloud or mesh, and send the bitstream to another client device, such as any of the client devices 106-116.
[0027] In an embodiment, any of the client devices 106-114 securely and efficiently sends information to another device (such as, for example, the server 104). In addition, any of the client devices 106-116 can trigger the transmission of information between itself and the server 104. Any of the client devices 106-114 can be used as a VR display when attached to a head-mounted device via a bracket and function similarly to the HMD 116. For example, the mobile device 108 can function similarly to the HMD 116 when attached to the bracket system and worn on the user's eyes. The mobile device 108 (or any other client device 106-116) can trigger the transmission of information between itself and the server 104.
[0028] In an embodiment, any of the client devices 106-116 or the server 104 may create a 3D point cloud or mesh, compress a 3D point cloud or mesh, send a 3D point cloud or mesh, receive a 3D point cloud or mesh, decode a 3D point cloud or mesh, render a 3D point cloud or mesh, or a combination of these operations. For example, the server 104 may compress the 3D point cloud or mesh to generate a bitstream and then send the bitstream to one or more of the client devices 106-116. As an example, one of the client devices 106-116 may compress the 3D point cloud or mesh to generate a bitstream and then send the bitstream to another of the client devices 106-116 or the server 104. According to the present disclosure, the server 104 and / or the client devices 106-116 may utilize a displacement data packing scheme based on the video format and / or other factors to improve the encoding of the displacement.
[0029] although Figure 1 One example of a communication system 100 is shown, but may be used for Figure 1 Various changes may be made. For example, communication system 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems have a wide variety of configurations, and Figure 1 The scope of the present disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may be used in any other suitable system.
[0030] Figure 2 and Figure 3 An example electronic device according to the present disclosure is shown. Specifically, Figure 2 An example server 200 is shown, and the server 200 may represent Figure 1 Server 200 may represent one or more encoders, decoders, local servers, remote servers, cluster computers, and components that act as a single seamless resource pool, cloud-based servers, etc. Server 200 may be composed of Figure 1 One or more of the client devices 106-116 or another server accesses the client device 106-116.
[0031] like Figure 2 As shown, server 200 may represent one or more local servers, one or more compression servers, or one or more encoding servers (such as encoders). In an embodiment, the encoder may perform decoding. Figure 2 As shown, server 200 includes a bus system 205 that supports communications between at least one processing device (such as at least one processor 210 ), at least one storage device 215 , at least one communication interface 220 , and at least one input / output (I / O) unit 225 .
[0032] The processor 210 executes one or more instructions that may be stored in the memory 230. The processor 210 may include any suitable number(s) and type(s) of processors or other devices in any suitable arrangement. Example types of processors 210 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application specific integrated circuits, and discrete circuits.
[0033] In an embodiment, the processor 210 may encode the 3D point cloud or mesh stored in the storage device 215. In an embodiment, encoding the 3D point cloud also decodes the 3D point cloud or mesh to ensure that when the point cloud or mesh is reconstructed, the reconstructed 3D point cloud or mesh matches the 3D point cloud or mesh before encoding. In an embodiment, the processor 210 may improve the encoding of the displacement using a displacement data packing scheme based on the video format and / or other factors.
[0034] Memory 230 and persistent storage 235 are examples of storage 215, which represent any structure(s) capable of storing and facilitating retrieval of information, such as data, program code, or other suitable information, on a temporary or permanent basis. Memory 230 may represent random access memory or any other suitable volatile or non-volatile storage device(s). For example, the instructions stored in memory 230 may include instructions for decomposing a point cloud into patches, instructions for packing patches onto 2D frames, instructions for compressing 2D frames, and instructions for encoding 2D frames in a particular order to generate a bitstream. The instructions stored in memory 230 may also include instructions for rendering the point cloud onto an omnidirectional 360° scene, such as for viewing via a VR headset (such as a VR headset). Figure 1 Persistent storage 235 may include one or more components or devices that support long-term storage of data, such as read-only memory, a hard drive, flash memory, or an optical disk.
[0035] Communication interface 220 supports communication with other systems or devices. For example, communication interface 220 may include interfaces that facilitate communication via Figure 1 The communication interface 220 may include a network interface card or wireless transceiver for communication with the network 102. The communication interface 220 may support communication over any suitable (one or more) physical or wireless communication links. For example, the communication interface 220 may send a bitstream containing a 3D point cloud to another device (such as one of the client devices 106-116).
[0036] The I / O unit 225 allows for the input and output of data. For example, the I / O unit 225 can provide a connection for user input via a keyboard, mouse, keypad, touch screen, or other suitable input device. The I / O unit 225 can also send output to a display, printer, or other suitable output device. However, it is noted that the I / O unit 225 can be omitted, such as when I / O interaction with the server 200 occurs via a network connection.
[0037] Note that although Figure 2 Described as indicating Figure 1 104, but the same or similar structure may be used in one or more of the various client devices 106-116. For example, a desktop computer 106 or a laptop computer 112 may have a Figure 2 The same or similar structure as shown in .
[0038] Figure 3 An example electronic device 300 is shown and may represent Figure 1 The electronic device 300 may be a mobile communication device such as, for example, a mobile station, a subscriber station, a wireless terminal, a desktop computer (similar to a Figure 1 desktop computer 106), portable electronic device (similar to Figure 1 In an embodiment, the electronic device 300 may be a mobile device 108, a PDA 110, a laptop computer 112, a tablet computer 114, or an HMD 116). In an embodiment, the electronic device 300 may be an apparatus. In an embodiment, Figure 1 One or more of the client devices 106 to 116 may include the same or similar configuration as the electronic device 300. In an embodiment, the electronic device 300 is an encoder, a decoder, or both. For example, the electronic device 300 may be used for data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.
[0039] like Figure 3As shown, the electronic device 300 includes an antenna 305, a radio frequency (RF) transceiver 310, a transmission (TX) processing circuit 315, a microphone 320, and a reception (RX) processing circuit 325. The RF transceiver 310 may include, for example, an RF transceiver, a Bluetooth transceiver, a WI-FI transceiver, a ZIGBEE transceiver, an infrared transceiver, and various other wireless communication signals. The electronic device 300 also includes a speaker 330, a processor 340, an input / output (I / O) interface (IF) 345, an input 350, a display 355, a memory 360, and (one or more) sensors 365. The memory 360 includes an operating system (OS) 361 and one or more applications 362.
[0040] The RF transceiver 310 receives an incoming RF signal from an antenna 305 transmitted from an access point (such as a base station, a WI-FI router, or a Bluetooth device) or other device of the network 102 (such as WI-FI, Bluetooth, cellular, 5G, LTE, LTE-A, WiMAX, or any other type of wireless network). The RF transceiver 310 down-converts the incoming RF signal to generate an intermediate frequency or baseband signal. The intermediate frequency or baseband signal is sent to the RX processing circuit 325, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or intermediate frequency signal. The RX processing circuit 325 sends the processed baseband signal to the speaker 330 (such as for voice data) or to the processor 340 for further processing (such as for web browsing data).
[0041] The TX processing circuit 315 receives analog or digital voice data from the microphone 320 or other outgoing baseband data from the processor 340. The outgoing baseband data may include web data, email, or interactive video game data. The TX processing circuit 315 encodes, multiplexes, and / or digitizes the outgoing baseband data to generate a processed baseband or intermediate frequency signal. The RF transceiver 310 receives the outgoing processed baseband or intermediate frequency signal from the TX processing circuit 315 and up-converts the baseband or intermediate frequency signal to an RF signal that is transmitted via the antenna 305.
[0042] The processor 340 may include at least one processor or other processing device. The processor 340 may execute one or more instructions stored in the memory 360 (such as OS 361) to control the overall operation of the electronic device 300. For example, the processor 340 may control the reception of the forward channel signal and the transmission of the reverse channel signal by the RF transceiver 310, the RX processing circuit 325 and the TX processing circuit 315 according to well-known principles. The processor 340 may include any suitable (one or more) number and (one or more) type of processor or other device of any suitable arrangement. For example, in an embodiment, the processor 340 includes at least one microprocessor or microcontroller. Example types of the processor 340 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application-specific integrated circuits, and discrete circuits.
[0043] The processor 340 is also capable of executing other processes and programs resident in the memory 360, such as operations of receiving and storing data. The processor 340 may move data into or out of the memory 360 as required by the execution process. In an embodiment, the processor 340 is configured to execute one or more applications 362 based on the OS 361 or in response to a signal received from (one or more) external sources or operators. For example, the application 362 may include an encoder, a decoder, a VR or AR application, a camera application (for still images and videos), a video phone call application, an email client, a social media client, an SMS messaging client, a virtual assistant, etc. In an embodiment, the processor 340 is configured to receive and send media content. In an embodiment, the processor 340 may utilize a displacement data packaging scheme based on the video format and / or other factors to improve the encoding of the displacement.
[0044] Processor 340 is also coupled to I / O interface 345 , which provides electronic device 300 with the ability to connect to other devices, such as client devices 106 - 114 . I / O interface 345 is the communication path between these accessories and processor 340 .
[0045] The processor 340 is also coupled to an input 350 and a display 355. The operator of the electronic device 300 can use the input 350 to enter data or input into the electronic device 300. The input 350 can be a keyboard, a touch screen, a mouse, a trackball, a voice input, or other devices that can act as a user interface to allow the user to interact with the electronic device 300. For example, the input 350 may include a voice recognition process, thereby allowing the user to enter a voice command. In an example, the input 350 may include a touch panel, a (digital) pen sensor, a key, or an ultrasonic input device. The touch panel may recognize, for example, a touch input of at least one scheme (such as a capacitive scheme, a pressure-sensitive scheme, an infrared scheme, or an ultrasonic scheme). By providing additional inputs to the processor 340, the input 350 may be associated with (one or more) sensors 365 and / or cameras. In an embodiment, the sensor 365 includes one or more inertial measurement units (IMUs) (such as accelerometers, gyroscopes, and magnetometers), motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, etc. The input 350 may also include a control circuit. In a capacitive scheme, input 350 may recognize touch or proximity.
[0046] Display 355 may be a liquid crystal display (LCD), a light emitting diode (LED) display, an organic LED (OLED), an active matrix OLED (AMOLED), or other display capable of rendering text and / or graphics (such as from a website, video, game, image, etc.). Display 355 may be sized to fit within the HMD. Display 355 may be a single display screen or multiple display screens capable of creating a stereoscopic display. In an embodiment, display 355 is a heads-up display (HUD). Display 355 may display 3D objects (such as a 3D point cloud or mesh).
[0047] Memory 360 is coupled to processor 340. A portion of memory 360 may include RAM, and another portion of memory 360 may include flash memory or other ROM. Memory 360 may include a persistent storage (not shown) representing any (one or more) structures capable of storing and facilitating retrieval of information (such as data, program code, and / or other suitable information). Memory 360 may include one or more components or devices (such as read-only memory, hard drive, flash memory, or optical disk) that support long-term storage of data. Memory 360 may also include media content. Media content may include various types of media (such as images, videos, three-dimensional content, VR content, AR content, 3D point clouds, meshes, etc.).
[0048] The electronic device 300 also includes one or more sensors 365 that can measure physical quantities or detect the activation state of the electronic device 300 and convert the measured or detected information into electrical signals. For example, the sensor 365 may include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor (such as a gyroscope or gyro sensor and an accelerometer), an eye tracking sensor, an air pressure sensor, a magnetic sensor or a magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illumination sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an IR sensor, an ultrasonic sensor, an iris sensor, a fingerprint sensor. A color sensor (such as a red, green, and blue (RGB) sensor), etc. The sensor 365 may also include a control circuit for controlling any sensor included therein.
[0049] As discussed in more detail below, one or more of these sensor(s) 365 may be used to control a user interface (UI), detect UI input, determine the user's orientation and facing direction for three-dimensional content display recognition, etc. Any of these sensor(s) 365 may be located within the electronic device 300, within an auxiliary device operably connected to the electronic device 300, within a head mounted device configured to hold the electronic device 300, or within a single device in which the electronic device 300 includes the head mounted device.
[0050] The electronic device 300 may create media content, such as generating a virtual object or capturing (or recording) content through a camera. The electronic device 300 may encode the media content to generate a bitstream so that the bitstream can be directly transmitted to another electronic device, or transmitted to another electronic device, such as by Figure 1 The electronic device 300 may receive the bit stream directly from another electronic device, or may receive the bit stream indirectly from another electronic device, such as by Figure 1 The network 102 indirectly receives the bit stream.
[0051] although Figure 2 and Figure 3 An example of an electronic device is shown, but the Figure 2 and Figure 3 Make various changes. For example, Figure 2 and Figure 3 The various components in the embodiment may be combined, further subdivided, or omitted, and additional components may be added according to specific needs. As a specific example, processor 340 may be divided into multiple processors, such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs). In addition, as with computing and communications, electronic devices and servers may have a variety of configurations, and Figure 2 and Figure 3 The present disclosure is not limited to any particular electronic device or server.
[0052] Figure 4 An example intra-coding process 400 according to the present disclosure is shown. Figure 4 The intra-coding process 400 is shown for illustration only. Figure 4 The scope of the present disclosure is not limited to any particular implementation of the intra-coding process.
[0053] like Figure 4 As shown, the intra-frame encoding process 400 encodes the grid frame using an intra-frame encoder 402. The intra-frame encoder 402 may be composed of Figure 2 The server 200 or Figure 3 The electronic device 300 shown represents or performs a process of creating and quantizing a base mesh 404, which typically has a smaller number of vertices than the original mesh, and compressing the base mesh in a lossy or lossless manner, and then encoding the base mesh into a compressed base mesh bitstream. Figure 4 As shown, the static mesh decoder decodes and reconstructs the base mesh, providing a reconstructed base mesh 406. The reconstructed base mesh 406 then undergoes one or more levels of subdivision, and a displacement field is created for each subdivision representing the difference between the original mesh and the subdivided reconstructed base mesh. In the inter-frame coding of the mesh frame, the base mesh 404 is encoded by sending vertex motions rather than directly compressing the base mesh. In either case, a displacement field 408 is created. Each displacement of the displacement field 408 has three components, represented by x, y and z. These can be relative to a canonical coordinate system or a local coordinate system, where x, y and z represent displacements in the local normal direction, tangent direction and bitangent direction. It will be understood that multiple levels of subdivision can be applied, so that multiple subdivided mesh frames are created, and a displacement field for each subdivided mesh frame is also created.
[0054] Let the number of 3-D displacement vectors in the grid frame displacement 408 be N. Let the displacement field be The displacement field 408 undergoes one or more levels of wavelet transform 410 to create a level of detail (LOD) signal , where k represents the index of the detail level, represents the number of samples in the detail level signal at level k, and numLOD represents the number of LODs. Can be scalar quantized.
[0055] like Figure 4As shown, the quantized LOD signal corresponding to the displacement field 408 is encoded into a compressed bitstream. In an embodiment, the quantized LOD signal is packed into a 2D image / video using an image packing operation and is losslessly compressed using an image or video encoder. However, another entropy encoder (such as an asymmetric digital system (ANS) encoder or a binary arithmetic entropy encoder) may be used to encode the quantized LOD signal. There may be other dependencies based on previous samples, across components, and across LODs that can be utilized. It is also possible that, using an image packing operation, the unquantized LOD signal is packed into a 2D image / video and then compressed in a lossy manner by an image or video encoder.
[0056] Historically, the 4:4:4 video format has been used to store the x, y, and z components (normals, tangents, and bitangents) of the LOD signal. In the 4:4:4 video format, the x component will be stored in the Y plane, the y component will be stored in the Cb plane, and the z component will be stored in the Cr plane, where the Y, Cb, and Cr planes have the same width and height. In some cases, depending on the color space used, the x component is stored in the R plane, the y component is stored in the G plane, and the z component is stored in the B plane, where the R, G, and B planes have the same width and height. However, video encoders and decoders that can operate on 4:4:4 format video are not widely available, especially in hardware, limiting the usefulness of this approach because many devices that implement encoders or decoders are not even compatible with using the 4:4:4 video format. Therefore, as further described below, the present disclosure provides different video formats to be used and different component packing arrangements to be used based on the video format and / or other factors.
[0057] Also like Figure 4 As shown, image unpacking of the LOD signal is performed, and an inverse quantization operation and an inverse wavelet transform operation are performed to reconstruct the LOD signal. Another inverse quantization operation can be performed on the reconstructed base mesh 406, which is combined with the reconstructed LOD signal to reconstruct the deformed mesh. Attribute transfer operations are performed using the deformed mesh, static / dynamic mesh, and attribute graph. A point cloud is a collection of 3D points with attributes (such as color, normal, reflectivity, point size, etc.) representing the surface or volume of an object. These attributes are encoded as a compressed attribute bitstream. As shown Figure 4 As shown, the encoding of the compressed attribute bitstream may further include a padding operation, a color space conversion operation, and a video encoding operation. Figure 4 The various functions or operations shown in the figure may be controlled by the control process 412. The intra-frame encoding process 400 outputs a compressed bit stream, which may be sent to and decoded by an electronic device (such as the server 104 or the client devices 106-116), for example. Figure 4As shown, the output compressed bitstream may include a compressed base grid bitstream, a compressed displacement bitstream and a compressed attribute bitstream.
[0058] although Figure 4 A block diagram of an example intra-coding process 400 is shown, but may be used for Figure 4 Various changes may be made. For example, the number and placement of the various components of the intra-coding process 400 may vary as needed or desired. Furthermore, the intra-coding process 400 may be used in any other suitable process and is not limited to the specific process described above. In an embodiment, only the first (x) component of the displacement may be created and encoded, and the other two components (y and z) may be assumed to be zero. In such a case, a flag may be signaled in the bitstream to indicate that the bitstream contains only data corresponding to the first (x) component and that the other two components (y and z) should be assumed to be zero when decompressing and reconstructing the displacement field 408. As an example, Figure 4 The intra-coding process 400 may include using a packing technique based on the video format of the video being compressed and other factors, as described in this disclosure.
[0059] Figure 5A and Figure 5B An example packaging arrangement process 500 based on video format according to the present disclosure is shown. Figure 5A and Figure 5B The example packaging arrangement process 500 shown in FIG. is for illustration only. For ease of explanation, Figure 5A and Figure 5B The packaging arrangement process 500 can be described as using Figure 3 However, the example packaging arrangement process 500 may be used with any other suitable system and any other suitable electronic device (such as the server 200).
[0060] like Figure 5A and Figure 5B As shown in , the packing of x, y, z displacement data may depend on the video format of the video frame used to store the displacement data. Figure 5A As shown, in this example, when the video format is determined to be in a 4:4:4 video format, the x component of the displacement map is stored in the y (or luminance) plane, the y component of the displacement map is stored in the Cb plane, and the z component of the displacement map is stored in the Cr plane. Throughout this disclosure, the terms luma, luminance, and Y (as in the YcbCr format) are used interchangeably. Figure 5BAs shown, when the video format is determined to be a video format other than a 4:4:4 video format, the components are stored in the same plane, such as an x component of the displacement map is stored in a luma plane, a y component of the displacement map is stored in the luma plane after the x component, and a z component of the displacement map is stored in the luma plane after the y component. At least a portion of process 500 may be represented as follows: if (video format == 4:4:4) The x component of the displacement map is stored in the Y plane The y component of the displacement map is stored in the Cb plane The z component of the displacement map is stored in the Cr plane else if (video format == 4:2:0 OR video format == 4:2:2 OR video format == 4:0:0) The x component of the displacement map is first stored in the luma plane The y component of the displacement map is then stored in the luma plane after the x component The z component of the displacement map is then stored in the luma plane after the y component end In this way, displacement data may be efficiently and effectively stored in video frames depending on the video format, and so that video formats other than 4:4:4 format may be used to increase compatibility of encoders and decoders.
[0061] although Figure 5A and Figure 5B An example packaging arrangement process 500 is shown, but may be used for Figure 5A and Figure 5B For example, although the x, y, and z components are shown as being stored in certain planes of the video frame, the components may be stored in other ways without departing from the scope of the present disclosure, such as storing one of the x, y, and z components in a different plane than shown (e.g., storing the component in the Cr plane or the Cb plane) to match the Figure 5B The x, y, z components may be stored in an order different than that shown in , or otherwise stored in other ways such as described in this disclosure.
[0062] Fig. 6A and Figure 6B An example packaging arrangement process 600 based on packaging type variables according to the present disclosure is shown. Fig. 6A and Figure 6B The example packaging arrangement process 600 shown in FIG. is for illustration only. For ease of explanation, Fig. 6A and Figure 6B The packaging arrangement process 600 can be described as using Figure 3However, the example packaging arrangement process 600 may be used with any other suitable system and any other suitable electronic device (such as the server 200).
[0063] In an embodiment of the present disclosure, the packing type of x, y, z displacement data in a video frame may be signaled in the bitstream or by an external device. Fig. 6A and Figure 6B As shown in , the packing of x, y, z displacement data may depend on a syntax element or variable (e.g., "packing_type") that indicates how the displacement data is packed into a video frame. Fig. 6A As shown in , in this example, when the packing type indicates that separate planes are used to store x, y, and z components, the x component of the displacement map is stored in the R plane, the y component of the displacement map is stored in the G plane, and the z component of the displacement map is stored in the B plane of the 4:4:4 video frame. It will be understood that if the YcrCb color space is used, the x, y, and z components may be stored in the luminance, Cr, and Cb planes, respectively. In general, each video frame may have any suitable format, such as a red-green-blue (RGB) format or a luminance-chrominance (YUV or YcrCb) format. Each video frame may also have any suitable resolution.
[0064] like Figure 6B As shown, when the packing type indicates that the same plane is used to store the x, y, and z components, the x, y, and z components are stored one after another in the same plane of the video, such as Figure 6B 4:2:0 format shown in FIG. 6. In this example, the x component of the displacement map is stored in the luma plane, the y component of the displacement map is stored in the luma plane after the x component, and the z component of the displacement map is stored in the luma plane after the y component. At least a portion of process 600 may be represented as follows: if(packing_type == SEPARATE_PLANES) The x component of the displacement map is stored in the R plane The y component of the displacement map is stored in the G plane The z component of the displacement map is stored in the B plane else if(packing_type == SAME_PLANE) The x component of the displacement map is first stored in the luma plane The y component of the displacement map is then stored in the luma plane after the x component The z component of the displacement map is then stored in the luma plane after the y component end In this way, displacement data may be efficiently and effectively stored in video frames depending on the packing type, and video formats other than 4:4:4 format may be used to increase compatibility of encoders and decoders.
[0065] although Fig. 6A and Figure 6B An example packaging arrangement process 600 is shown, but may be used for Fig. 6A and Figure 6B For example, although the x, y, and z components are shown as being stored in certain planes of the video frame, the components may be stored in other ways without departing from the scope of the present disclosure, such as storing one of the x, y, and z components in a different plane than shown (e.g., storing the component in the Cr plane or the Cb plane) to match the Figure 6B The x, y, z components may be stored in an order different from that shown in FIG, or otherwise stored in other ways such as described in the present disclosure. Figure 6B The example shows the use of 4:2:0 video frames, but other video formats (such as 4:2:2, 4:0:0, etc.) may be used.
[0066] In an embodiment of the present disclosure, the resolution of the x, y, and z displacement data stored in the video frame is sent by signal in the bitstream or by an external device. The resolution can be the actual height and width of each of the components or a syntax element indicating the relative size with respect to the full resolution. For example, let scale_x_h be the relative resolution of the displacement data stored in the video frame in the vertical direction. Let scale_x_w be the relative resolution in the horizontal direction, where 0<=scale_x_h<=1, 0<=scale_x_w<=1. When either scale_x_h or scale_x_w is zero, the x displacement data is not stored in the video frame. In addition, scale_y_h, scale_y_w, scale_z_h, and scale_z_w are similar relative resolutions for the y component and the z component, respectively. Some example uses of this class include setting scale_x_h=1, scale_x_w=1, scale_y_h=0.5, scale_y_w=0.5, scale_z_h=0.5, scale_z_w=0.5, and setting packing_type = SEPARATE_PLANES to store the displacement data in a 4:2:0 video frame.
[0067] In an embodiment of the present disclosure, the video format (for example, as described in Figure 5A and Figure 5B described), packaging type (e.g., as described in Fig. 6A and Figure 6BA combination of one or more of the factors (described) and scaling factors (eg, as described above), as well as combinations of other factors also described in the disclosure, may be used to determine a packing scheme for displacement data in a video frame.
[0068] In an embodiment, the values of the scaling factors (scale_x_h, etc.) may be limited to 0, 0.5, and 1.0. In an embodiment, each of the scaling factors may be signaled using 3-symbol entropy coding (such as unary, exp-golomb, or arithmetic coding, etc.). In an embodiment, the scaling factors corresponding to the coordinates (i.e., x, y, or z) may be jointly encoded, as shown in example Table 1:
[0069] Table 1: Joint encoding of scaling factors In an embodiment, a 4-symbol entropy coding (such as unary, exponential Golomb or arithmetic coding) may be used to signal the jointly encoded scaling factor. The scheme shown in Table 1 may also be generalized to other values of the scaling factor. The subsampling indicated by the scaling factor may be achieved using interpolation or dropping samples or any other method.
[0070] In an embodiment, a flag indicating whether the y and z components (tangent and bitangent) are cleared to zero may be signaled. In an embodiment, on the encoder side, the displacement data (x, y and z (normal, tangent and bitangent)) is stored in a 4:2:0 format using a packing_type variable equal to SAME_PLANE, such as Figure 6B as shown. However, in an embodiment, the encoder may choose to zero the y and z (or tangent and bitangent) components of the displacement data. In such an instance, no additional flag may be signaled to indicate the zeroing of the y and z components of the displacement data. For example, the normal_component_only_flag is not sent; only the packing type variable is sent. If the packing type variable is set to SAME_PLANE, the x, y, and z components are packed into the 0th (Y) video component. The encoder may choose to zero the y and z components of the displacement. In such an embodiment, on the decoder side, the decoder assumes that all 3 displacement components (x, y, z) are present, regardless of the packing type. The packing type determines whether the displacement data is extracted from the 0th component or all 3 components.
[0071] On the decoder side, in an embodiment, if the displacement data is received in a 4:2:0 video format, the packing type variable may be inferred to be SAME_PLANE. In this case, the x, y, and z components of the displacement data are based on Figure 6B The packing shown in is extracted from the luma (or zeroth) component of the reconstructed displacement video.
[0072] In an embodiment, if the decoder receives displacement video data in a 4:4:4 video format, the x, y, and z components of the displacement data are extracted from the zeroth component, the first component, and the second component of the reconstructed displacement video. In an embodiment, if the decoder receives displacement video data in a 4:0:0 video format, the x (or normal) component of the displacement data is extracted from the zeroth component of the reconstructed displacement video, and the y and z (or tangent and bitangent) components of the displacement data are set to 0.
[0073] Figure 7 An example staggered packing arrangement process 700 according to the present disclosure is shown. Figure 7 The example interleaved packing arrangement process 700 shown in FIG. is for illustration only. For ease of explanation, Figure 7 The process 700 can be described as using Figure 3 However, the process 700 may be used with any other suitable system and any other suitable electronic device (such as the server 200).
[0074] like Figure 7 As shown in , in an embodiment, the electronic device 300 may interleave and store the x, y, and z components of the displacement data in the same plane. As an example, the interleaving may be done at the pixel level or the block level, such as Figure 7 , for a 4:2:0 video frame. A flag (e.g., "interleaving_flag") may be signaled in the bitstream or by an external device to indicate whether the displacement data is interleaved and stored. When the interleaving flag is equal to a predetermined constant (e.g., "INTERLEAVED"), then the x, y, z components of the displacement data in the video frame are pixel or block interleaved. Interleaving may also occur at level of detail (LOD) boundaries. When the interleaving flag is equal to a predetermined value (e.g., "NO_INTERLEAVING"), then the x, y, z data in the frame is not pixel or block interleaved. In this example, all of the x data is stored first, followed by all of the y data, followed by all of the z data, but the data may be interleaved in any other order.
[0075] although Figure 7 An example staggered packing arrangement process 700 is shown, but may be used for Figure 7Various changes may be made. For example, while the x, y, and z components are shown as being stored in certain planes of the video frame, the components may be stored in other ways, such as storing the x, y, and z components in different planes than shown, storing the x, y, z components in a different order than shown, or otherwise storing the components in other ways such as described in the present disclosure without departing from the scope of the present disclosure. In an embodiment, two components may be interleaved in one plane while the other component is stored in a separate plane (e.g., storing the z component values in the Cr and / or Cb planes and interleaving the x and y components in the luma plane). In addition, while Figure 7 The example shows the use of 4:2:0 video frames, but other video formats (such as 4:2:2, 4:0:0, etc.) may be used.
[0076] Figure 8 An example full packaging arrangement process 800 according to the present disclosure is shown. Figure 8 The example full package arrangement process 800 shown in FIG. is for illustration only. For ease of explanation, Figure 8 The process 800 can be described as using Figure 3 However, the process 800 may be used with any other suitable system and any other suitable electronic device (such as the server 200).
[0077] In an embodiment, in order to optimally use the video frames (such as Figure 8 In order to use the space in the 4:2:0 frame shown in , the electronic device 300 can disperse and store the x, y, z components of the displacement data in all luminance, Cr and Cb planes in the video frame. In this example, 50% of the luminance plane is occupied by the x component, the remaining 50% of the luminance plane is occupied by the y component, and 50% of the z component is stored in the Cr plane and the remaining 50% of the z component is stored in the Cb plane. In an embodiment, a packing type variable (e.g., "packing_type") can be used to indicate the use of the packing type (e.g., "packing_type = FULL_420_PACKING").
[0078] although Figure 8 An example full package placement process 800 is shown, but may be used for Figure 8 Various changes may be made. For example, although the x, y, and z components are shown as being stored in certain planes of the video frame, the components may be stored in other ways without departing from the scope of the present disclosure, such as storing the x, y, and z components in different planes than shown, storing the x, y, z components in a different order than shown, or otherwise storing the components in other ways such as described in the present disclosure. In an embodiment, when using the Figure 8When using the fully packed format described above, the x, y, and z components can also be interleaved. Interleaving can be performed at the pixel, block, or LOD level. In addition, although Figure 8 The example shows the use of 4:2:0 video frames, but other video formats (such as 4:2:2, 4:0:0, etc.) may be used. In an embodiment, when the displacement data is encoded as a single component, that is, the y and z components of the displacement data are zeroed, packing_type = FULL_420_PACKING is interpreted to mean packing the 1D displacement data on the luma and Cr and Cb planes. Approximately 2 / 3 of the single component displacement data is stored in the luma component, approximately 1 / 6 is stored in the Cr component, and approximately 1 / 6 is stored in the Cb component.
[0079] In embodiments, components may be packed in different orders (e.g., x followed by y followed by z, or z followed by y followed by x, or y followed by z followed by x, or other orders). In embodiments, components or blocks of components are interleaved. Additionally, in embodiments, interleaving may be done based on LOD. For example, consider there are 3 LODs: LOD0, LOD1, and LOD2. In this case, using the specified reverse packing order, the x, y, and z components of LOD2 are interleaved. This is followed by LOD1, LOD0, and finally padding rows (if any). Thus, the order of data in the shifted video frame would be the x component of LOD2, the y component of LOD2, the z component of LOD2, the x component of LOD1, the y component of LOD1, the z component of LOD1, the x component of LOD0, the y component of LOD0, the z component of LOD0, and any padding blocks / rows at the end.
[0080] Fig. 9 An example split packaging arrangement process 900 according to the present disclosure is shown. Fig. 9 The example split packing arrangement process 900 shown in FIG. is for illustration only. For ease of explanation, Fig. 9 The process 900 can be described as using Figure 3 However, the process 900 may be used with any other suitable system and any other suitable electronic device (such as the server 200).
[0081] like Fig. 9 As shown, in an embodiment, the electronic device 300 may store half of one of the x, y, and z component values in the first plane and store half of these values in the second plane. Fig. 9 , half of the z-component values in the left half (ZL) of the z-component plane are stored in the Cr plane, and half of the z-component values in the right half (ZR) of the z-component plane are stored in the Cb plane.
[0082] Although Fig. 9 An example split packaging arrangement process 900 is shown, but may be used for Fig. 9 Various changes may be made. For example, although the x, y, and z components are shown as being stored in certain planes of the video frame, the components may be stored in other ways without departing from the scope of the present disclosure, such as storing the x, y, and z components in different planes than shown, storing the x, y, z components in a different order than shown, or otherwise storing the components in other ways such as described in the present disclosure. Interleaving may be performed at the pixel, block, or LOD level. In addition, although Fig. 9 The example shows the use of 4:2:0 video frames, but other video formats (such as 4:2:2, 4:0:0, etc.) may be used. In addition, in an embodiment, when using Fig. 9 When splitting the packed format as described, the x, y, and z components may also be interleaved.
[0083] Fig.10 An example split and interleave packing arrangement process 1000 according to the present disclosure is shown. Fig.10 The example split and interleaved packing arrangement process 1000 shown in FIG. is for illustration only. For ease of explanation, Fig.10 The process 1000 can be described as using Figure 3 However, the process 1000 may be used with any other suitable system and any other suitable electronic device (such as the server 200).
[0084] like Fig.10 As shown, in an embodiment, the electronic device 300 may store half of one of the x, y, z component values in a first plane and half of the values in a second plane in an interleaved manner. In this example, odd columns (Z1) of z component values in the z plane are stored in the Cr plane in an interleaved manner, and even columns (Z2) of z component values in the z plane are stored in the Cb plane, or vice versa.
[0085] although Fig.10 An example split and interleave packing arrangement process 1000 is shown, but may be used for Fig.10 Various changes may be made. For example, although the x, y, and z components are shown as being stored in certain planes of the video frame, the components may be stored in other ways without departing from the scope of the present disclosure, such as storing the x, y, and z components in different planes than shown, storing the x, y, z components in a different order than shown, or otherwise storing the components in other ways such as described in the present disclosure. Interleaving may be performed at the pixel, block, or LOD level. In addition, although Fig.10 The example shows the use of 4:2:0 video frames, but other video formats (such as 4:2:2, 4:0:0, etc.) may be used.
[0086] In embodiments of the present disclosure, the Cr and Cb planes may contain a higher level of detail for the x, y, z components, and the luma plane may contain a lower level of detail for the x, y, z components, and vice versa. Furthermore, it will be understood that the x, y, z components of the displacement map may be stored interchangeably. It will also be understood that the x, y, and z components of the displacement data may be scanned in Morton order or other scanning patterns before being stored in a video frame.
[0087] Fig.11A and Fig. 11B An example single component packaging arrangement process 1100 based on a video format according to the present disclosure is shown. Fig.11A and Fig. 11B The single component packing arrangement process 1100 shown in FIG is for illustration only. For ease of explanation, Fig.11A and Fig. 11B The process 1100 can be described as using Figure 3 However, the process 1100 may be used with any other suitable system and any other suitable electronic device (such as the server 200).
[0088] Also about Figure 4 As described, in an embodiment, a flag may be used and signaled to indicate whether only the first component of the displacement is encoded. The first component may be the x component (displacement in the normal direction). For example, the name of the flag may be "normal_component_only_flag". For example, if the value of the flag is 1, and the displacement video is encoded in a 4:2:0 format, the displacement video may be formed by including the first component of the displacement (x or normal direction) in the luma component, and the chroma components may be set to default values, such as Fig.11A and Fig. 11B In an embodiment, for 8-bit video, the Cb and Cr components may be set to a value of 128 or 127. In other embodiments, for 10-bit video, the Cb and Cr components may be set to a value of 512 or 511.
[0089] At the decoder side, when the value of normal_component_only_flag is 1 and a displacement video in 4:2:0 format is received, in an embodiment, only the luma component is decoded to derive the value of the displacement video in the normal direction, such as Fig. 11B In other embodiments, all three components may be decoded, but only the luma component is used to derive the value of the displacement video in the normal direction.
[0090] In an embodiment, if the value of normal_component_only_flag is 0, and the displacement video is encoded in 4:2:0 format, the displacement video may be formed as previously described when packing_type = SAME_PLANE, and also as described above. Figure 6B Here, the chroma components may be assigned default values. On the decoder side, when the value of the flag is 0, and the displaced video is encoded in 4:2:0 format, the displaced x, y, and z components may be retrieved from the luma component of the 4:2:0 video based on the packing arrangement selected during encoding.
[0091] In an embodiment, 4:2:0 video may always be formed according to a packing type variable equal to "SAME_PLANE" without signaling the packing type in the bitstream. For example, if "normal_component_only_flag" is 1 and the video format is 4:2:0 (or 4:0:0 or 4:2:2 or 4:4:4), the 0th (luminance) component of the decoded video corresponds to the x component of the displacement. In this scenario, using the 4:2:0 format may be beneficial because sending data in 4:2:2 or 4:4:4 format wastes bits due to the absence of data in the Cb and Cr components. In an embodiment, depending on whether there is hardware support for the 4:0:0 format, the 4:0:0 format may be used and may provide higher efficiency. If "normal_component_only_flag" is 0 and the video format is 4:2:0 (or 4:0:0 or 4:2:2), the 0th (luminance) component of the decoded video may contain the x, y, and z components of the displacement (x, followed by y, followed by z). In an embodiment, if the video format is 4:4:4 and "normal_component_only_flag" is 0, the x, y, z components of the displacement are extracted from the YcbCr (or RGB) components, respectively.
[0092] In an embodiment, a bitstream consistency constraint may be imposed such that a 4:4:4 format may be used only when all 3 components of the displacement are present. Similarly, in an embodiment, a bitstream consistency constraint may be imposed such that a 4:0:0 format may be used only when only one component of the displacement is present. However, all 3 components may still be packed into 4:0:0 frames. In an embodiment, the condition for bitstream consistency is that the video format is 4:4:4 when the packing type is SEPARATE_PLANES.
[0093] In an embodiment, 4:2:0 video may always be formed according to the packing type variable equal to "FULL_420_PACKING" without signaling the packing type in the bitstream. In an embodiment, when the value of normal_component_only_flag is 0, packing_type may be explicitly signaled in the bitstream.
[0094] In an embodiment, the packing_type is explicitly signaled in the bitstream and also includes a new packing_type represented by, for example, "SUBSAMPLED_PACKING". In an embodiment, when the packing type is set to SUBSAMPLED_PACKING, the scaling factors of the Y, Cr, Cb or respectively x, y and z components are as follows:
[0095] Here, subsampling of the y and z components may be achieved using interpolation, dropping 3 out of 4 samples, or any other method.In an embodiment, 4:2:0 video may always be formed according to a packing type equal to SUBSAMPLED_PACKING without signaling the packing type in the bitstream.
[0096] In an embodiment, if the value of the flag is 0, and the displaced video is encoded in 4:2:0 format, the displaced video may be formed as previously described when packing_type = FULL_420_PACKING, such as Figure 8 However, if the value of the flag is 0, the chroma components may be assigned default values.
[0097] In an embodiment, if the decoder receives a displaced video in 4:4:4 format, the decoder may infer that normal_component_only_flag is 0 (regardless of the actual value signaled in the bitstream) and proceed to extract the displaced x-component, y-component, and z-component from the zeroth, first, and second components of the displaced video in 4:4:4 format. In this example, the signaled value of normal_component_only_flag is ignored.
[0098] In an embodiment, if the decoder receives a displacement video in 4:4:4 format and normal_component_only_flag is 0, the decoder extracts the x, y, and z components of the displacement from the zeroth component, the first component, and the second component of the displacement video in 4:4:4 format. In an embodiment, if the decoder receives a displacement video in 4:4:4 format and normal_component_only_flag is 1, the decoder extracts the x (normal) component of the displacement from the zeroth component of the displacement video in 4:4:4 format, such as Fig.11A shown.
[0099] In an embodiment, if a decoder receives 4:0:0 formatted displaced video, the decoder may infer that normal_component_only_flag is 1 (regardless of the actual value signaled in the bitstream) and proceed to extract the x (normal) component of the displacement from the 4:0:0 formatted displaced video. In an embodiment, a requirement for bitstream conformance may be that when a decoder receives 4:0:0 formatted displaced video, the flag indicating a single displacement component should be equal to 1. In an embodiment, a requirement for bitstream conformance may be that when a decoder receives 4:4:4 formatted displaced video, the flag indicating a single displacement component should be equal to 0.
[0100] although Fig.11A and Fig. 11B An example single component packaging arrangement process 1100 is shown, but may be used for Fig.11A and Fig. 11B Various changes may be made. For example, although the x, y, and z components are shown as being stored in certain planes of the video frame, the components may be stored in other ways without departing from the scope of the present disclosure.
[0101] Figures 12 to 15 An additional example packing arrangement utilizing padding in accordance with the present disclosure is shown. In an embodiment, the size of the displacement frame may vary in time due to a different number of vertices. In this case, the width of the displacement video may be kept fixed, and the number of rows may be padded to make the height of the displacement video constant. For example, in an embodiment, when the packing type is equal to "SAME_PLANE", the padding row may be placed above the frame, followed by the x, y, and z components, such as Fig.12 In an embodiment, when the packing type is equal to "SAME_PLANE", padding lines may be placed below the frame, after the x, y, and z components, such as Fig.13 shown.
[0102] In an embodiment, when normal_component_only_flag is equal to 1, the zeroth component of a 4:2:0 video frame may include a padding line above followed by the x component, such as Fig.14 In an embodiment, a padding row may be placed below the zeroth component after the x component, such as Fig.15 shown.
[0103] although Figures 12 to 15 An example packing arrangement using padding is shown, but Figures 12 to 15 For example, although the x component is shown as being stored in the luma plane of the video frame, the component may be stored in another plane without departing from the scope of the present disclosure. Figures 12 to 15 The example shows the use of 4:2:0 video frames, but other video formats (such as 4:2:2, 4:0:0, etc.) may be used.
[0104] In an embodiment, when the packing arrangement utilizes padding, the sequence parameter set or bitstream may include variables and information identifying the padding region of each displaced video frame. The padding region may be free of displacement data, and the displacement data is retrieved from the region of each displaced video frame that is not padded.
[0105] Various standards for vertex meshes (V-MESH) and dynamic mesh encoding have been proposed. The following documents are incorporated by reference in their entirety as if fully set forth herein: “V-Grid Test Model v1”, ISO / IEC SC29 WG07 N00404, July 2022 “WD 2.0 for V-DMC”, ISO / IEC SC29WG07 N00546, January 2023 “WD 3.0 for V-DMC”, ISO / IEC SC29WG07 N00611, April 2023 “WD 4.0 for V-DMC”, ISO / IEC JTC 1 / SC 29 / WG 07 N00611, August 2023 To provide the packaging arrangement of the disclosed embodiments, WD 2.0 of the V-DMC may be updated to specify the following: 8.3.6.1.3 Atlas sequence parameter set extended RBSP syntax
[0106] 8.4.6.1.3 Atlas sequence parameter set extended RBSP syntax … asps_vmc_ext_packing_method equal to 0 specifies that the displacement component samples are packed in ascending order, and asps_vmc_ext_packing_method equal to 1 specifies that the displacement component samples are packed in descending order.
[0107] Asps_vmc_ext_1D_displacement_flag equal to 1 specifies that only the normal (or x) component of the displacement is present in the compressed geometry video. The remaining two components are inferred to be 0. Asps_vmc_ext_1D_displacement_flag equal to 0 specifies that all 3 components of the displacement are present in the compressed geometry video.
[0108] 11.5 Inverse Image Packing of Wavelet Coefficients The inputs to this process are: width is a variable indicating the width of the displacement video frame, height is a variable indicating the height of the displacement video frame, bitDepth is a variable indicating the bit depth of the displacement video frame, disQuantCoeffFrame is a 3D array of size width × height × 3 indicating the packed quantized shifted wavelet coefficients.
[0109] blockSize is a variable indicating the size of the displacement coefficient block, positionCount is a variable indicating the number of positions in the subdivided sub-grid.
[0110] The output of this process is dispQuantCoeffArray, which is a two-dimensional array of size positionCount × 3 indicating the quantized displacement wavelet coefficients.
[0111] A bitstream conformance requirement is that when DecGeoChromaFormat is equal to 4:0:0, asps_vmc_ext_1D_displacement_flag shall be equal to 1. A bitstream conformance requirement is also that when DecGeoChromaFormat is equal to 4:4:4, asps_vmc_ext_1D_displacement_flag shall be equal to 0.
[0112] The 2D array dispQuantCoeffArray is initialized to 0. The variable DisplacementDim is set as follows: - If asps_vmc_ext_1D_displacement_flag is equal to 1, DisplacementDim is set to 1 - Otherwise (asps_vmc_ext_1D_displacement_flag is equal to 0), DisplacementDim is set to 3, Let the function extracOddBits(x) be defined as follows:
[0113] Let the function computeMorton2D(i) be defined as follows:
[0114] The inverse packing process of wavelet coefficients is performed as follows:
[0115] Figures 16 to 20 An example of a sub-grid packing arrangement according to the present disclosure is shown. In an embodiment, the original grid may be divided into several sub-grids. For each sub-grid, the encoder may derive a corresponding base sub-grid. Multiple base sub-grids may be encoded and decoded in parallel by the base grid codec. The three-dimensional displacement vectors in the displacement field corresponding to different sub-grids occupy non-overlapping rectangular areas in the displacement video frame. Fig.16 An example arrangement 1600 of two sub-grids is shown in FIG. 1 , where the rectangular areas are arranged in a vertical grid.
[0116] In an embodiment, each rectangular area of the displacement frame corresponding to a sub-grid may be regarded as a sub-frame. All packing embodiments described above in the present disclosure may be applied to each sub-frame. In an embodiment, the same filling method may be used for all sub-grids in a grid frame.
[0117] For example, if the flag indicates that only the first component of the displacement is to be encoded, and there are two subgrids, the data may be as follows Fig.17Arranged as shown. In this example arrangement 1700, the luminance (Y) component of the displaced video frame includes the x (or normal) component of subgrid 0, followed by any padding to make the subframes of subgrid 0 have equal sizes. The padding is followed by the x (normal) component of subgrid 1, and finally any padding for subgrid 1. For example, the displaced video may be encoded as monochrome video (4:0:0 format) or 4:2:0 format. In this case, the Cb and Cr components of the displaced video frame take default values. It will be understood that this embodiment can be extended to multiple subgrids. In an embodiment, the subframes of the subgrids may be stacked horizontally or in a rectangular grid structure. The advantage of calculating the padding for each subgrid separately and including it after the corresponding subgrid data is that the position of the subgrid displacement data remains constant for all frames. This makes it easier to tile the displacement data for different subgrids and perform partial decoding.
[0118] If the flag indicates that all three components of the displacement are to be encoded, and there are two subgrids, then Fig.18 Arrange data as shown. When the displacement video is encoded in 4:2:0 format, the Cb and Cr components of the displacement video frame may take default values. In this example arrangement 1800, the luminance (Y) component of the displacement video frame includes the x, y, and z (or normal, tangent, and bitangent) components of subgrid 0, followed by any padding for subgrid 0. The padding is followed by the x, y, and z (or normal, tangent, and bitangent) components of subgrid 1 and any padding. It will be understood that this embodiment can be extended to multiple subgrids. It can also be extended to situations where subframes of a subgrid can be stacked horizontally or in a rectangular grid structure.
[0119] In an embodiment, padding for all sub-grids may be included below the frame, such as Fig.19 1900 is shown in an example arrangement. In such an embodiment, the width of each displacement frame is fixed. Then, for each frame, the height of the subframe corresponding to each subgrid is calculated. The sum of the subframe heights is the height of the displacement frame. Then, the maximum height over all frames is calculated, and (if necessary) padding is added below each frame to make the height of each frame equal to the maximum height. In this example, the position of the displacement data corresponding to each subgrid may change from one frame to another. However, in some cases, less padding is required. It will be understood that this packing arrangement can be extended to multiple subgrids. It can also be extended to situations where the subframes of the subgrids can be stacked horizontally or in a rectangular grid structure.
[0120] In an embodiment, Fig. 20As shown in the example arrangement 2000 of , the x (or normal) components of all sub-meshes may be packed together, followed by the y (or tangent) component, followed by the z (or bitangent) component, and any padding. For such a configuration, the bitstream syntax may be used to signal three rectangles (corresponding to the positions of the x, y, and z components in the video frame) for each sub-mesh. Alternatively, only the rectangle corresponding to the x component of the sub-mesh may be signaled, and the positions of the other two components of the sub-mesh may be inferred.
[0121] although Figures 16 to 20 An example subgrid packing arrangement is shown, but Figures 16 to 20 Various changes may be made. For example, although the subgrids are shown as being stored in the luminance plane of the video frame, the subgrids may be stored in another plane without departing from the scope of the present disclosure. Figures 16 to 20 The example shows the use of 4:2:0 video frames, but other video formats (such as 4:2:2, 4:0:0, etc.) may be used.
[0122] Fig.21 An example encoding method 2100 for improved packing of displacement data in a video frame according to the present disclosure is shown. For ease of explanation, Fig.21 The method 2100 is described as using Figure 3 However, the method 2100 may be used with any other suitable system and any other suitable electronic device.
[0123] like Fig.21 As shown, in step 2102, the electronic device 300 determines or identifies the video format for the video to be encoded, such as whether the video format is 4:4:4, 4:2:0, 4:2:2, 4:0:0, etc. In step 2104, the electronic device 300 sets or determines at least one signaling element for the video. In an embodiment, at least one signaling element is used to identify a displacement data packing arrangement for displacement data. For example, as described in the present disclosure, at least one signaling element may be a format variable indicating a specific video format used, or may be a variable indicating a packing type (e.g., "packing_type") for displacement data, and at least one signaling element may change the displacement data packing arrangement used during encoding in either case. For example, when the video format is 4:4:4, the x, y, and z components may be packed into the zeroth component, the first component, and the second component of the video frame, respectively, or when the video is 4:2:0, such as, in the case where the packing type variable indicates that the displacement data is in the same plane, the x, y, and z components may all be packed into the zeroth component of the video frame. As described in the present disclosure, the zeroth component of the displaced video frame may be a luminance component of the displaced video frame.
[0124] In step 2106, the electronic device 300 encodes the video into a bitstream according to the displacement data packing arrangement. In an embodiment, as described in the present disclosure, encoding the video into a bitstream includes interleaving a normal component, a tangent component, and a bitangent component in a zeroth component of a displacement video frame. In an embodiment, at least one signaling element may include a flag, and the electronic device 300 may set a value of the flag, for example in a sequence parameter set, the flag indicating whether only a normal component of the displacement data is present in the bitstream, or whether a normal component, a tangent component, and a bitangent component of the displacement data are present in the bitstream. For example, when the value of the flag indicates that only a normal component of the displacement data is present in the bitstream, the electronic device 300 may encode the displacement data only in the zeroth component of each displacement video frame. As an example, when the value of the flag indicates that the normal component, the tangent component, and the bitangent component of the displacement data are present in the bitstream, the electronic device 300 may encode the displacement data only in the zeroth component (e.g., storing the x, y, and z components in the zeroth component) or encode the displacement data in the zeroth component, the first component, and the second component (e.g., storing the x component in the zeroth component, the y component in the first component, and the z component in the second component) for each displacement video frame and based on the video format used. In an embodiment, when the video format is the 4:0:0 format, the value of the flag is set to 1, and when the video format is the 4:4:4 format, the value of the flag is set to 0.
[0125] At step 2108, the electronic device 300 outputs a bitstream including the encoded data. The output bitstream may include, for example, Figure 4 , and / or may be a compressed bitstream comprising a compressed displacement bitstream as well as a compressed base grid bitstream and a compressed attribute bitstream, such as Figure 4 The output bitstream may be sent to an external device or a memory on the electronic device 300.
[0126] although Fig.21 An example of an improved packing encoding method 2100 for displacement data in a video frame is shown, but may be used for Fig.21 For example, although shown as a series of steps, Fig.21 The various steps in may overlap, occur in parallel, or occur any number of times.
[0127] Fig. 22 An example decoding method 2200 for improved packing of displacement data in a video frame according to the present disclosure is shown. For ease of explanation, Fig. 22 The method 2200 is described as using Figure 3 However, the method 2200 may be used with any other suitable system and any other suitable electronic device.
[0128] like Fig. 22 As shown, in step 2202, the electronic device 300 identifies the video format for the compressed video. In step 2204, the electronic device 300 determines the displacement data packing arrangement for the compressed video based on at least one signaling element and one or more of the identified video formats. For example, as described in the present disclosure, at least one signaling element may be a format variable indicating a specific video format used, or may be a variable (e.g., "packing_type") indicating a packing type for displacement data, and at least one signaling element may change the displacement data packing arrangement of the displacement data in the compressed video in either case. For example, when the video format is 4:4:4, the x, y, and z components may be packed into the zeroth component, the first component, and the second component of the video frame, respectively, or when the video is 4:2:0, such as, in the case where the packing type variable indicates that the displacement data is in the same plane, the x, y, and z components may all be packed into the zeroth component of the video frame. As described in the present disclosure, the zeroth component of the displacement video frame may be the luminance component of the displacement video frame. In an embodiment, the packing arrangement may also be determined based on the video format used.
[0129] Determining the displacement data packing arrangement informs the electronic device 300 about how to unpack or extract the displacement data from the video frame. In step 2206, the electronic device 300 retrieves the displacement data according to the determined displacement data packing arrangement. In an embodiment, the normal component, the tangent component, and the bitangent component may also be interleaved in the zeroth component of the displacement video frame, and the electronic device 300 may extract the interleaving value. In an embodiment, at least one signaling element may include a flag, and the electronic device 300 may determine whether only the normal component of the displacement data exists in the bitstream of the compressed video, or whether the normal component, the tangent component, and the bitangent component of the displacement data exist in the bitstream of the compressed video based on the value of the flag, for example, in the sequence parameter set. In an embodiment, when only the normal component of the displacement data exists in the bitstream of the compressed video, the electronic device 300 retrieves the displacement data only from the zeroth component of each displacement video frame. In an embodiment, when the normal component, the tangent component, and the bitangent component of the displacement data are present in the bitstream for the compressed video, for each displacement video frame, the electronic device 300 may retrieve the displacement data from the zeroth component (e.g., retrieve the x, y, and z components from the zeroth component), or may retrieve the displacement data from the zeroth component, the first component, and the second component of each displacement video frame (e.g., retrieve the x component from the zeroth component, retrieve the y component from the first component, and retrieve the z component from the second component) based on the video format used. In an embodiment, when the video format is a 4:0:0 format, the value of the flag is inferred to be 1, and when the video format is a 4:4:4 format, the value of the flag is inferred to be 0. In an embodiment, the video frame includes a padding area, wherein the padding area has no displacement data, and the displacement data is retrieved from the unpadded area of each displacement video frame.
[0130] At step 2208, the electronic device 300 outputs a reconstructed mesh based on the displacement data. The output reconstructed mesh frame may be sent to an external device or a memory on the electronic device 300.
[0131] although Fig. 22 An example of a decoding method 2200 for improved packing of displacement data in a video frame is shown, but may be used for Fig. 22 For example, although shown as a series of steps, Fig. 22 The various steps in may overlap, occur in parallel, or occur any number of times.
[0132] Although the present disclosure has been described using exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. It is intended that the present disclosure encompasses these changes and modifications that fall within the scope of the appended claims. None of the descriptions in this application should be interpreted as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the patent subject matter is defined by the claims.
[0133] In an embodiment, a device for decoding a video may include a memory (360) storing one or more instructions; and at least one processor (340) may be configured to execute the one or more instructions stored in the memory (360). The at least one processor (340) may be configured to obtain a bitstream for a compressed video. The at least one processor (340) may be configured to identify a video format for the compressed video. The at least one processor (340) may be configured to determine a displacement data packing arrangement based on one or more of at least one signaling element and the identified video format. The at least one processor (340) may be configured to retrieve displacement data based on the determined displacement data packing arrangement.
[0134] In an embodiment, at least one signaling element may include a packed type variable.
[0135] In an embodiment, when the packing type variable indicates that the displacement data is in the same plane, the displacement data packing arrangement may include the normal component, the tangent component, and the bitangent component of the displacement data being stored in a zeroth component of the displacement video frame.
[0136] In an embodiment, the normal component, the tangent component, and the bitangent component may be interleaved in the zeroth component of the displacement video frame.
[0137] In an embodiment, the zeroth component of the displaced video frame may be a luminance component of the displaced video frame.
[0138] In an embodiment, at least one signaling element may include a flag, and at least one processor (340) may be configured to determine, based on the value of the flag, whether only a normal component of the displacement data is present in the bitstream for compressed video, or whether a normal component, a tangent component, and a bitangent component of the displacement data are present in the bitstream for compressed video. When only a normal component of the displacement data is present in the bitstream for compressed video, the displacement data is retrieved only from the zeroth component of each displacement video frame. When the normal component, the tangent component, and the bitangent component of the displacement data are present in the bitstream for compressed video, for each displacement video frame, at least one processor (340) may be configured to retrieve the displacement data from the zeroth component or one of the zeroth component, the first component, and the second component based on the identified video format.
[0139] In an embodiment, the value of the flag may be inferred to be 1 when the video format is the 4:0:0 format, or the value of the flag may be inferred to be 0 when the video format is the 4:4:4 format.
[0140] In an embodiment, the video format may be a 4:2:0 format.
[0141] In an embodiment, the displacement video frame may include padding areas without displacement data.
[0142] In an embodiment, a method for decoding video is provided. The method may include obtaining a bitstream for compressed video. The method may include identifying a video format for the compressed video. The method may include determining a displacement data packing arrangement based on at least one signaling element and one or more of the identified video formats. The method may include retrieving displacement data based on the determined displacement data packing arrangement.
[0143] In an embodiment, a device for encoding a video may include a memory (360) storing one or more instructions and at least one processor (340) configured to execute the one or more instructions stored in the memory. The at least one processor (340) may be configured to determine a video format for the video. The at least one processor (340) may be configured to set at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for displacement data. The at least one processor (340) may be configured to encode the video into a bitstream according to the displacement data packing arrangement.
[0144] In an embodiment, at least one signaling element may include a packing type variable. When the packing type variable indicates that the displacement data is in the same plane, the displacement data packing arrangement may include that the normal component, the tangent component, and the bitangent component of the displacement data are stored in the zeroth component of the displacement video frame. To encode the video into the bitstream, at least one processor may be configured to interleave the normal component, the tangent component, and the bitangent component in the zeroth component of the displacement video frame.
[0145] In an embodiment, at least one signaling element may include a flag. At least one processor (340) may be configured to execute one or more instructions stored in the memory (360). At least one processor may be configured to set a value of a flag indicating whether only a normal component of the displacement data is present in the bitstream, or whether a normal component, a tangent component, and a bitangent component of the displacement data are present in the bitstream. When the value of the flag indicates that only a normal component of the displacement data is present in the bitstream, at least one processor (340) may be configured to encode the displacement data only in the zeroth component of each displacement video frame. When the value of the flag indicates that a normal component, a tangent component, and a bitangent component of the displacement data are present in the bitstream, for each displacement video frame, at least one processor (340) may be configured to encode the displacement data in the zeroth component or in one of the zeroth component, the first component, and the second component based on the determined video format.
[0146] In an embodiment, a method for encoding a video may include determining a video format for the video. The method may include setting at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for displacement data. The method may include encoding the video into a bitstream according to the displacement data packing arrangement.
[0147] In an embodiment, a computer-readable storage medium storing a bitstream may be provided, the bitstream comprising a video format for a video and at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for displacement data.
[0148] In an embodiment, a device may include a communication interface configured to receive a bitstream of a video for compression and a processor operably coupled to the communication interface. The processor may be configured to identify a video format for the compressed video. The processor may be configured to determine a displacement data packing arrangement based on at least one signaling element and one or more of the identified video formats. The processor may be configured to retrieve the displacement data based on the determined displacement data packing arrangement.
[0149] In an embodiment, at least one signaling element may include a packed type variable.
[0150] In an embodiment, when the packing type variable indicates that the displacement data is in the same plane, the displacement data packing arrangement may include the normal component, the tangent component, and the bitangent component of the displacement data being stored in a zeroth component of the displacement video frame.
[0151] In an embodiment, the normal component, the tangent component, and the bitangent component may be interleaved in the zeroth component of the displacement video frame.
[0152] In an embodiment, the zeroth component of the displaced video frame may be a luminance component of the displaced video frame.
[0153] In an embodiment, at least one signaling element may include a flag. At least one processor may be configured to determine, based on the value of the flag, whether only a normal component of the displacement data exists in the bitstream of the compressed video or whether a normal component, a tangent component, and a bitangent component of the displacement data exist in the bitstream of the compressed video. At least one processor may be configured to retrieve the displacement data only from the zeroth component of each displacement video frame when only a normal component of the displacement data exists in the bitstream of the compressed video. At least one processor may be configured to retrieve the displacement data from the zeroth component or one of the zeroth component, the first component, and the second component for each displacement video frame based on the identified video format when normal components, tangent components, and bitangent components of the displacement data exist in the bitstream of the compressed video.
[0154] In an embodiment, the value of the flag may be inferred to be 1 when the video format is the 4:0:0 format, or the value of the flag may be inferred to be 0 when the video format is the 4:4:4 format.
[0155] In an embodiment, a method may include identifying a video format for compressed video of a received bitstream. The method may include determining a displacement data packing arrangement based on at least one signaling element and one or more of the identified video format. The method may include retrieving displacement data based on the determined displacement data packing arrangement.
[0156] In an embodiment, at least one signaling element may include a packed type variable.
[0157] In an embodiment, when the packing type variable indicates that the displacement data is in the same plane, the displacement data packing arrangement may include the normal component, the tangent component, and the bitangent component of the displacement data being stored in a zeroth component of the displacement video frame.
[0158] In an embodiment, the normal component, the tangent component, and the bitangent component may be interleaved in the zeroth component of the displacement video frame.
[0159] In an embodiment, the zeroth component of the displaced video frame may be a luminance component of the displaced video frame.
[0160] In an embodiment, at least one signaling element may include a flag, and the method further includes determining, based on the value of the flag, whether only a normal component of the displacement data is present in the bitstream for compressed video, or whether a normal component, a tangent component, and a bitangent component of the displacement data are present in the bitstream for compressed video. The method may include: when only a normal component of the displacement data is present in the bitstream for compressed video, retrieving the displacement data only from the zeroth component of each displacement video frame. The method may include: when the normal component, the tangent component, and the bitangent component of the displacement data are present in the bitstream for compressed video, for each displacement video frame, retrieving the displacement data from the zeroth component or one of the zeroth component, the first component, and the second component based on the identified video format.
[0161] In an embodiment, the value of the flag may be inferred to be 1 when the video format is the 4:0:0 format, or the value of the flag may be inferred to be 0 when the video format is the 4:4:4 format.
[0162] In an embodiment, the device may include a communication interface and a processor operably coupled to the communication interface. The processor may be configured to determine a video format for the video. The processor may be configured to set at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for the displacement data. The processor may be configured to encode the video into a compressed video bitstream according to the displacement data packing arrangement.
[0163] In an embodiment, at least one signaling element may include a packed type variable.
[0164] In an embodiment, when the packing type variable indicates that the displacement data is in the same plane, the displacement data packing arrangement may include the normal component, the tangent component, and the bitangent component of the displacement data being stored in a zeroth component of the displacement video frame.
[0165] In an embodiment, the processor may be further configured to interleave the normal component, the tangent component, and the bitangent component in a zeroth component of the displaced video frame to encode the video into a compressed video bitstream.
[0166] In an embodiment, at least one signaling element includes a flag. The processor may be further configured to set a value of the flag, the flag indicating whether only a normal component of the displacement data is present in the compressed video bitstream or whether a normal component, a tangent component, and a bitangent component of the displacement data are present in the compressed video bitstream. The processor may be further configured to encode the displacement data only in the zeroth component of each displacement video frame when the value of the flag indicates that only a normal component of the displacement data is present in the compressed video bitstream. The processor may be further configured to encode the displacement data in the zeroth component or one of the zeroth component, the first component, and the second component for each displacement video frame based on a determined video format when the value of the flag indicates that a normal component, a tangent component, and a bitangent component of the displacement data are present in the compressed video bitstream.
[0167] In an embodiment, when the video format is the 4:0:0 format, the value of the flag may be set to 1, or when the video format is the 4:4:4 format, the value of the flag may be set to 0.
Claims
1. A device for decoding a video, comprising: A memory (360) storing one or more instructions; as well as At least one processor (340) is configured to execute the one or more instructions stored in the memory (360) to perform the following operations: Obtaining a bitstream of a video for compression; Identify the video format used to compress the video; determining a displacement data packing arrangement based on one or more of the at least one signaling element and the identified video format; as well as The displacement data is retrieved according to the determined displacement data packing arrangement.
2. The device according to claim 1, wherein The at least one signaling element comprises a packed type variable.
3. The device according to claim 2, wherein: When the packing type variable indicates that the displacement data is in the same plane, the displacement data packing arrangement includes the normal component, the tangent component, and the bitangent component of the displacement data being stored in the zeroth component of the displacement video frame.
4. The device according to claim 3, wherein: The normal component, the tangent component, and the bitangent component are interleaved in the zeroth component of the displacement video frame.
5. The device according to any one of claims 3 to 4, wherein: The zeroth component of the displaced video frame is the luminance component of the displaced video frame.
6. The device according to any one of claims 1 to 2, wherein: The at least one signaling element comprises a flag, and wherein the at least one processor (340) is configured to execute the one or more instructions stored in the memory to: determining, based on the value of the flag, whether only a normal component of displacement data is present in the bitstream of video for compression, or whether a normal component, a tangent component, and a bitangent component of displacement data are present in the bitstream of video for compression; When only a normal component of displacement data is present in the bitstream for compressed video, retrieving displacement data only from the zeroth component of each displacement video frame; and When normal, tangent, and bitangent components of displacement data are present in the bitstream for compressed video, for each displacement video frame, the displacement data is retrieved from one of the following based on the identified video format: the zeroth component; or The zeroth component, the first component, and the second component.
7. The apparatus of claim 6, wherein: When the video format is 4:0:0 format, the value of the flag is inferred to be 1; or When the video format is 4:4:4 format, the value of the flag is inferred to be 0.
8. The apparatus according to any one of claims 3 to 5, wherein: The video format is 4:2:0 format.
9. The apparatus according to any one of claims 3 to 8, wherein: The displacement video frame includes padding areas without displacement data.
10. A method for decoding a video, comprising: Obtaining a bitstream of a video for compression; Identify the video format used to compress the video; determining a displacement data packing arrangement based on one or more of the at least one signaling element and the identified video format; as well as The displacement data is retrieved according to the determined displacement data packing arrangement.
11. A device for encoding a video, comprising: A memory (360) storing one or more instructions; as well as At least one processor (340) is configured to execute the one or more instructions stored in the memory to perform the following operations: Determine the video format to use for the video; providing at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for the displacement data; as well as The video is encoded into a bitstream according to the displacement data packing arrangement.
12. The device according to claim 11, in, The at least one signalling element comprises a packed type variable; wherein, when the packing type variable indicates that the displacement data is in the same plane, the displacement data packing arrangement includes storing the normal component, the tangent component, and the bitangent component of the displacement data in the zeroth component of the displacement video frame; and Wherein, in order to encode the video into the bitstream, the at least one processor is configured to interleave the normal component, the tangent component and the bitangent component in the zeroth component of the displacement video frame.
13. The device according to claim 11, wherein: The at least one signaling element comprises a flag, and wherein the at least one processor (340) is configured to execute the one or more instructions stored in the memory (360) to: setting a value of the flag, the flag indicating whether only a normal component of displacement data is present in the bitstream, or whether a normal component, a tangent component, and a bitangent component of displacement data are present in the bitstream; When the value of the flag indicates that only a normal component of displacement data is present in the bitstream, encoding only the displacement data in the zeroth component of each displacement video frame; and When the value of the flag indicates that the normal component, the tangent component, and the bitangent component of the displacement data are present in the bitstream, for each displacement video frame, encoding the displacement data in one of the following based on the determined video format: the zeroth component; or The zeroth component, the first component, and the second component.
14. A method for encoding a video, comprising: Determine the video format to use for the video; providing at least one signaling element for the video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for the displacement data; as well as The video is encoded into a bitstream according to the displacement data packing arrangement.
15. A computer-readable storage medium storing a bit stream, the bit stream comprising: Video format to use for videos; as well as At least one signaling element for video, wherein the at least one signaling element is used to identify a displacement data packing arrangement for the displacement data.