Electronic device and method for reconstructing grid frame
By identifying and utilizing the neighbor relationship and motion vector predictors of vertices in the compressed video bitstream of the vertices, the problem of low coding efficiency in the prior art is solved, and a more efficient encoding and decoding process is achieved.
Patent Information
- Application Number
- CN202380073869.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-02
- Filing Date
- 2023-10-19
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has low efficiency when dealing with compression and decoding of vertex grids, especially in the calculation and encoding of vertex motion vector predictors, where complexity and efficiency problems exist.
The grid frame is reconstructed by identifying vertices in the compressed video bitstream based on the limits of the number of vertices or vertices neighbors sent with the signal, and identifying predictors for vertices from the multiple predictors based on the vertex motion vector identifier.
The efficiency of vertex motion vector predictor encoding is improved, the complexity and run time of the encoding process is reduced, specifically manifested as a decrease of about 30% on the motion decoding time, and a decrease of about 50% on the storage use of vertex neighbor tables.
Smart Images

Figure CN120077659A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to multimedia devices and processing. More specifically, the present disclosure relates to improved vertex motion vector predictor coding for vertex meshes (V-MESH). Background Art
[0002] Due to the existing availability of powerful handheld devices such as smart phones, three-hundred-sixty-degree (360°) video and three-dimensional (3D) volumetric video are becoming new ways to experience immersive content. 360° video enables consumers to achieve an immersive "real life", "being on the scene" experience by capturing a 360° external-internal view of the world, while 3D volumetric video can provide a complete six degrees of freedom (DoF) experience of being immersed in and moving within the content. Users can interactively change their viewpoints and dynamically view any part of the captured scene or object they desire. Display and navigation sensors can track the user's head movement in real time to determine the area of the 360° video or volumetric content that the user wants to view or interact with. Multimedia data that is inherently 3D, such as point clouds or 3D polygon meshes, can be used in an immersive environment. This data can be stored in video format and encoded and compressed for transmission as a bitstream to other devices. Summary of the Invention
[0003] Technical Solution The present disclosure provides improved vertex motion vector predictor coding for vertex meshes (V-MESH).
[0004] In an embodiment of the present disclosure, an electronic device may include a memory and at least one processor coupled to the memory. The at least one processor may be configured to identify a compressed video bitstream. The at least one processor may be configured to determine one or more vertex neighbors for a vertex in the compressed video bitstream based on a signal-sent limit on the number of one or more vertex neighbors. The at least one processor may be configured to identify a VMV predictor to be used for the vertex from a plurality of VMV predictors based on a vertex motion vector (VMV) identifier signaled in the compressed video bitstream. The at least one processor may be configured to reconstruct a mesh frame based on the determined one or more vertex neighbors and the identified VMV predictor.
[0005] In an embodiment of the present disclosure, a method may be performed by an electronic device. The method may include identifying a compressed video bitstream. The method may include determining, for vertices in the compressed video bitstream, one or more vertex neighbors based on a signaled limit on the number of one or more vertex neighbors. The method may include identifying, from a plurality of vertex motion vector (VMV) predictors, a VMV predictor to be used for a vertex based on a signaled VMV identifier in the compressed video bitstream. The method may include reconstructing a mesh frame based on the determined one or more vertex neighbors and the identified VMV predictor.
[0006] In an embodiment of the present disclosure, an electronic device may include a memory and at least one processor coupled to the memory. The at least one processor may be configured to identify, for vertices of a mesh frame, one or more vertex neighbors based on a set limit on the number of one or more vertex neighbors. The at least one processor may be configured to determine, based on the identified one or more vertex neighbors, a plurality of vertex motion vector (VMV) predictors for the vertex. The at least one processor may be configured to map each of the plurality of VMV predictors to one of a plurality of VMV identifiers. The at least one processor may be configured to encode a compressed video bitstream that signals the set limit on the number of one or more vertex neighbors and signals one of the plurality of VMV identifiers, the VMV identifier indicating the VMV predictor of the plurality of VMV predictors to be used for the vertex.
[0007] In an embodiment of the present disclosure, a method may be performed by an electronic device. The method may include identifying, for vertices of a mesh frame, one or more vertex neighbors based on a set limit on the number of one or more vertex neighbors. The method may include determining, based on the identified one or more vertex neighbors, a plurality of vertex motion vector (VMV) predictors for the vertex. The method may include mapping each of the plurality of VMV predictors to one of a plurality of VMV identifiers. The method may include encoding a compressed video bitstream that signals the set limit on the number of one or more vertex neighbors and signals one of the plurality of VMV identifiers, the VMV identifier indicating the VMV predictor of the plurality of VMV predictors to be used for the vertex.
[0008] Other technical features may be apparent to those skilled in the art from the following drawings, description, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like reference numerals represent like parts: Figure 1 Illustrates an example communication system in an embodiment of the present disclosure; Figure 2 And Figure 3 Illustrates an example electronic device in an embodiment of the present disclosure; Figure 4 Illustrates an example intra - frame encoding process in an embodiment of the present disclosure; Figure 5A And Figure 5B Illustrates an example inter - frame grid - frame encoding process; Figure 6A And Figure 6B Illustrates an example adjacent - vertex determination process in an embodiment of the present disclosure; Figure 7 Illustrates an example grid - frame decoding process in an embodiment of the present disclosure; Figure 8 Illustrates an example set of vertices in an embodiment of the present disclosure; Figure 9 Illustrates an example encoding method for improved vertex motion vector prediction factor encoding in an embodiment of the present disclosure; and Figure 10 Illustrates an example decoding method for improved vertex motion vector prediction factor encoding in an embodiment of the present disclosure. Detailed Description
[0010] Before proceeding with the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The term "coupled" and its derivatives refer to any direct or indirect communication between two or more elements, whether or not those elements are in physical contact with each other. The terms "send," "receive," and "communicate" and their derivatives cover both direct and indirect communication. The terms "comprise" and "include" and their derivatives mean including but not limited to. The term "or" is inclusive and means and / or. The phrase "associated with" and its derivatives mean including, being included within, interconnecting with, containing, being contained within, connected to or connecting with, coupled to or coupling with, capable of communicating with, cooperating with, interlacing, juxtaposing, proximate to, bound to or binding with, having, having the property of, having a relationship to or a relationship with, and the like. The term "controller" means any device, system, or part thereof that controls at least one operation. Such a controller can be implemented in hardware or in a combination of hardware and software and / or firmware. The functions associated with any particular controller can be centralized or distributed, whether local or remote. When used with a list of items, the phrase "at least one of..." means that different combinations of one or more of the listed items can be used, and it may be necessary to have only one item from the list. For example, "at least one of A, B, or C" includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A and B and C.
[0011] In addition, the various functions described below can be implemented or supported by one or more computer programs, each formed from computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or portions thereof that are adapted to be implemented in suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, compact disc (CD), digital video disc (DVD), or any other type of memory. A "non-transitory" computer-readable medium excludes wired, wireless, optical, or other communication links that transmit transitory electrical or other signals. Non-transitory computer-readable media include media in which data can be permanently stored and media in which data can be stored and later rewritten, such as rewritable compact discs or erasable memory devices.
[0012] Throughout this patent document, definitions of certain words and phrases are provided. One of ordinary skill in the art should understand that in many cases, if not most cases, such definitions apply to both prior and future uses of the words and phrases so defined.
[0013] As described below Figures 1 to 10 and one or more embodiments for describing the principles of the present disclosure are merely exemplary and should not be construed in any way as limiting the scope of the present disclosure. Those skilled in the art will understand that the principles of the present disclosure can be implemented in any type of appropriately arranged device or system.
[0014] As mentioned above, due to the existing availability of powerful handheld devices such as smart phones, three-hundred-sixty-degree (360°) video and three-dimensional (3D) volumetric video are becoming new ways to experience immersive content. 360° video enables consumers to have an immersive "real life", "being there" experience by capturing a 360° external-internal view of the world, while 3D volumetric video can provide a full six degrees of freedom (DoF) experience of being immersed in and moving within the content. Users can interactively change their viewpoints and dynamically view any part of the captured scene or object they desire. Display and navigation sensors can track the user's head movements in real time to determine the regions of the 360° video or volumetric content that the user wants to view or interact with. Multimedia data that is inherently 3D, such as point clouds or 3D polygon meshes, can be used in immersive environments. This data can be stored in video format and encoded and compressed for transmission as a bitstream to other devices.
[0015] A point cloud is a set of 3D points along with properties representing the surface or volume of an object, such as color, normal direction, reflectivity, point size, etc. Point clouds are common in various applications such as gaming, 3D mapping, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view replay, and six degrees of freedom (DoF) immersive media, to name just a few. If uncompressed, point clouds typically require a large amount of bandwidth for transmission. Due to the high bitrate requirements, point clouds are usually compressed before transmission. Compressing 3D objects such as point clouds typically requires dedicated hardware. To avoid dedicated hardware for compressing 3D point clouds, the 3D point clouds can be transformed into traditional two-dimensional (2D) frames, and they can be compressed and later reconstructed and made viewable by the user.
[0016] Polygon 3D meshes, especially triangular meshes, are another common format for representing 3D objects. A mesh typically consists of a set of vertices, edges, and faces that represent the surface of a 3D object. A triangular mesh is a simple polygon mesh where the faces are simple triangles that cover the surface of a 3D object. Typically, one or more properties can be associated with the mesh. In a scene, one or more properties can be associated with each vertex in the mesh. For example, a texture property (RGB) can be associated with each vertex. In the scene, each vertex can be associated with a pair of coordinates (u, v). The (u, v) coordinates can point to a location in a texture map associated with the mesh. For example, the (u, v) coordinates can refer to the row index and column index in the texture map, respectively. A mesh can be considered a point cloud with additional connectivity information.
[0017] A point cloud or mesh can be dynamic, i.e., they can change over time. In these cases, the point cloud or mesh at a specific moment in time can be referred to as a point cloud frame or a mesh frame, respectively. Since point clouds and meshes contain a large amount of data, they need to be compressed for efficient storage and transmission. This is especially true for dynamic point clouds and meshes that may contain 60 frames or more per second.
[0018] As part of the encoding process, an existing mesh codec can be used to encode the base mesh, and a reconstructed base mesh can be built from the encoded original mesh. The reconstructed base mesh can then be subdivided into one or more subdivision meshes, and a displacement field can be created for each subdivision mesh. For example, if the reconstructed base mesh includes triangles that cover the surface of a 3D object, the triangles are subdivided according to the number of subdivision levels so that, depending on how many subdivision levels are applied, a first subdivision mesh is created where each triangle of the reconstructed base mesh is subdivided into four triangles, a second subdivision mesh is created where each triangle of the reconstructed base mesh is subdivided into sixteen triangles, etc. Each displacement field represents the difference between the vertex positions of the original mesh and the subdivision mesh associated with the displacement field. Each displacement field is wavelet-transformed to create a level-of-detail (LOD) signal, which is encoded as part of the compressed bitstream. During decoding, the displacements of each displacement field are added to their associated subdivision meshes to reconstruct the original mesh.
[0019] Previously, for a given vertex, flags were used to indicate whether to directly send the vertex motion vector of that vertex or whether to send the delta difference between the vertex motion vector of that vertex and its predicted value. The predicted value of the vertex motion vector was calculated as the average of the vertex motion vectors of adjacent vertices, but one or more embodiments of the present disclosure may use any combination of the vertex motion vectors of adjacent vertices, such as, an average, a weighted average, a median, a maximum value, a minimum value, etc. The calculation of adjacent motion vectors is a complex process and involves looping through all triangles and their connections.
[0020] The present disclosure provides an improved technique for determining vertex motion vector predictors and generating different mesh data structure tables based on the picture type, thereby providing improved compression efficiency. It has been shown that the techniques described in the present disclosure reduce the runtime of motion encoding by, for example, approximately 30%. These techniques include: for vertices of a mesh frame, identifying one or more vertex neighbors based on a set limit on the number of one or more vertex neighbors; determining a plurality of vertex motion vector (VMV) predictors for a vertex based on the identified one or more vertex neighbors; mapping each of the plurality of VMV predictors to one of a plurality of VMV identifiers; and encoding a compressed video bitstream that signals the set limit on the number of one or more vertex neighbors and signals one of the plurality of VMV identifiers, the VMV identifier indicating the VMV predictor among the plurality of VMV predictors that is used for the vertex, and reusing the identified vertex neighbors associated with an intra mesh frame for an inter mesh frame, as described in detail herein.
[0021] Figure 1 An example communication system 100 in an embodiment of the present disclosure is shown. Figure 1 The embodiment of the communication system 100 shown is for illustrative purposes only. One or more embodiments of the communication system 100 may be used without departing from the scope of the present disclosure.
[0022] As Figure 1 shown, the communication system 100 includes a network 102 that facilitates communication between various components in the communication system 100. For example, the network 102 may transmit IP packets, frame relay frames, asynchronous transfer mode (ATM) cells, or other information between network addresses. The network 102 includes one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), all or a portion of a global network such as the Internet, or any other communication system or systems at one or more locations.
[0023] In this example, network 102 facilitates communication between server 104 and various client devices 106 - 116. Client devices 106 - 116 can be, for example, smart phones, tablet computers, laptop computers, personal computers, TVs, interactive displays, wearable devices, HMDS, etc. Server 104 can represent one or more servers. Each server 104 includes any suitable computing or processing device capable of providing computing services for one or more client devices, such as client devices 106 - 116. Each server 104 can include, for example, one or more processing devices, one or more memories storing instructions and data, and one or more network interfaces facilitating communication over network 102. As described in more detail below, server 104 can send a compressed bitstream representing a point cloud or mesh to one or more display devices, such as client devices 106 - 116. In embodiments of the present disclosure, each server 104 can include an encoder. In embodiments of the present disclosure, server 104 can utilize improved vertex motion predictor coding as described in the present disclosure.
[0024] Each client device 106 - 116 represents any suitable computing or processing device that interacts with at least one server, such as server 104, or other (one or more) computing devices over network 102. Client devices 106 - 116 include desktop computer 106, mobile phone or mobile device 108 (such as a smart phone), PDA 110, laptop computer 112, tablet computer 114, and HMD 116. However, any other or additional client devices can be used in communication system 100. A smart phone represents a class of mobile devices 108 that are handheld devices having a mobile operating system and an integrated mobile broadband cellular network connection for voice, short message service (SMS), and Internet data communication. HMD 116 can display a 360° scene including one or more dynamic or static 3D point clouds. In embodiments of the present disclosure, any one of client devices 106 - 116 can include an encoder, a decoder, or both. For example, mobile device 108 can record 3D volumetric video and then encode the video so that it can be sent to one of client devices 106 - 116. In the example, laptop computer 112 can be used to generate a 3D point cloud or mesh, and then the 3D point cloud or mesh is encoded and sent to one of client devices 106 - 116.
[0025] In this example, some client devices 108 - 116 communicate with the network 102 indirectly. For example, the mobile device 108 and the PDA 110 communicate via one or more base stations 118 (such as cellular base stations or eNodeB (eNB)). Also, the laptop computer 112, the tablet computer 114, and the HMD 116 communicate via one or more wireless access points 120 (such as IEEE 802.11 wireless access points). Note that these are for illustration only, and each client device 106 - 116 can communicate with the network 102 directly or indirectly via any suitable intermediate device(s) or network(s). In embodiments of the present disclosure, the server 104 or any client device 106 - 116 can be used to compress a point cloud or a mesh, generate a bitstream representing the point cloud or the mesh, and send the bitstream to another client device such as any of the client devices 106 - 116.
[0026] In embodiments of the present disclosure, any one of the client devices 106 - 114 securely and efficiently sends information to another device, such as, for example, the server 104. Additionally, any one of the client devices 106 - 116 can trigger the information transfer between itself and the server 104. Any one of the client devices 106 - 114 can be used as a VR display when attached to a headset via a bracket and functions similarly to the HMD 116. For example, when the mobile device 108 is attached to a bracket system and worn over the user's eyes, the mobile device 108 can function similarly to the HMD 116. The mobile device 108 (or any other client device 106 - 116) can trigger the information transfer between itself and the server 104.
[0027] In embodiments of the present disclosure, any one of the client devices 106 - 116 or the server 104 can create a 3D point cloud or a mesh, compress the 3D point cloud or the mesh, send the 3D point cloud or the mesh, receive the 3D point cloud or the mesh, decode the 3D point cloud or the mesh, render the 3D point cloud or the mesh, or a combination thereof. For example, the server 104 can compress a 3D point cloud or a mesh to generate a bitstream and then send the bitstream to one or more of the client devices 106 - 116. As an example, one of the client devices 106 - 116 can compress a 3D point cloud or a mesh to generate a bitstream and then send the bitstream to another of the client devices 106 - 116 or the server 104. In embodiments of the present disclosure, the server 104 and / or the client devices 106 - 116 can utilize improved vertex motion predictor coding as described in the present disclosure.
[0028] Although Figure 1 an example of the communication system 100 is shown, Figure 1Make various changes. For example, communication system 100 can include any number of each component in any suitable arrangement. Generally, computing and communication systems have a wide variety of configurations, and Figure 1 do not limit the scope of the present disclosure to any particular configuration. Although Figure 1 an operating environment is shown in which various features disclosed in this patent document can be used, these features can be used in any other suitable system.
[0029] Figure 2 and Figure 3 illustrate an example electronic device in an embodiment of the present disclosure. In particular, Figure 2 an example server 200 is shown, and server 200 can represent Figure 1 server 104 in Figure 1 . Server 200 can represent one or more encoders, decoders, local servers, remote servers, cluster computers, and components that act as a single pool of seamless resources, cloud-based servers, etc. Server 200 can be accessed by one or more of the client devices 106 - 116 in
[0030] As Figure 2 shown, server 200 can represent one or more local servers, one or more compression servers, or one or more encoding servers, such as an encoder. In an embodiment of the present disclosure, an encoder can perform decoding. As Figure 2 shown, server 200 includes a bus system 205 that supports communication between at least one processing device (such as processor 210), at least one storage device 215, at least one communication interface 220, and at least one input / output (I / O) unit 225.
[0031] Processor 210 executes instructions that can be stored in memory 230. Processor 210 can include any suitable number and (one or more) types of processors or other devices in any suitable arrangement. Example types of processor 210 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application specific integrated circuits, and discrete circuits.
[0032] In an embodiment of the present disclosure, processor 210 can encode a 3D point cloud or mesh stored within storage device 215. In an embodiment of the present disclosure, encoding a 3D point cloud also decodes the 3D point cloud or mesh to ensure that when the point cloud or mesh is reconstructed, the reconstructed 3D point cloud or mesh matches the 3D point cloud or mesh before encoding. In an embodiment of the present disclosure, processor 210 can utilize improved vertex motion predictor coding as described in the present disclosure.
[0033] Memory 230 and persistent memory 235 are examples of storage devices 215, which represent any (one or more) structures capable of storing and facilitating the retrieval of information, such as data, program code, or other suitable information on a temporary or permanent basis. Memory 230 may represent random access memory or any other suitable (one or more) volatile or non-volatile storage device. For example, the instructions stored in memory 230 may include instructions for decomposing a point cloud into patches, instructions for packing the patches onto a 2D frame, instructions for compressing the 2D frame, and instructions for encoding the 2D frames in a specific order to generate a bitstream. The instructions stored in memory 230 may also include instructions for rendering the point cloud on an omnidirectional 360° scene, as viewed through a VR headset (such as Figure 1 the HMD 116). Persistent memory 235 may include one or more components or devices that support long-term storage of data, such as read-only memory, hard disk drives, flash memory, or optical discs.
[0034] Communication interface 220 supports communication with other systems or devices. For example, communication interface 220 may include a network interface card or wireless transceiver that facilitates communication over Figure 1 the network 102. Communication interface 220 may support communication over any suitable (one or more) physical or wireless communication links. For example, communication interface 220 may send a bitstream containing a 3D point cloud to another device, such as one of client devices 106 - 116.
[0035] I / O unit 225 allows for the input and output of data. For example, I / O unit 225 may provide a connection for user input via a keyboard, mouse, keypad, touch screen, or other suitable input device. I / O unit 225 may also send output to a display, printer, or other suitable output device. However, note that I / O unit 225 may be omitted when, for example, I / O interactions with server 200 occur via a network connection.
[0036] Note that although Figure 2 is described as representing Figure 1 the server 104, the same or similar structures may be used in one or more of the various client devices 106 - 116. For example, a desktop computer 106 or a laptop computer 112 may have the same or similar structure as that Figure 2 shown.
[0037] Figure 3 An example electronic device 300 is shown, and electronic device 300 may represent Figure 1one or more of the client devices 106-116 therein. The electronic device 300 can be a mobile communication device, such as, for example, a mobile station, a subscriber station, a wireless terminal, a desktop computer (similar to the Figure 1 desktop computer 106), a portable electronic device (similar to the Figure 1 mobile device 108, PDA 110, laptop computer 112, tablet computer 114, or HMD 116), etc. In an embodiment of the present disclosure, Figure 1 one or more of the client devices 106-116 therein may include the same or similar configuration as the electronic device 300. In an embodiment of the present disclosure, the electronic device 300 is an encoder, a decoder, or both. For example, the electronic device 300 can be used with data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.
[0038] As Figure 3 shown, the electronic device 300 includes an antenna 305, a radio frequency (RF) transceiver 310, a transmit (TX) processing circuit 315, a microphone 320, and a receive (RX) processing circuit 325. The RF transceiver 310 can include, for example, an RF transceiver, a Bluetooth transceiver, a WI-FI transceiver, a ZIGBEE transceiver, an infrared transceiver, and various other wireless communication signals. The electronic device 300 also includes a speaker 330, a processor 340, an input / output (I / O) interface (IF) 345, an input 350, a display 355, a memory 360, and one or more sensors 365. The memory 360 includes an operating system (OS) 361 and one or more applications 362.
[0039] The RF transceiver 310 receives an incoming RF signal sent from an access point (such as a base station, a WI-FI router, or a Bluetooth device) or another device on the network 102 (such as WI-FI, Bluetooth, cellular, 5G, LTE, LTE-A, WiMAX, or any other type of wireless network) from the antenna 305. The RF transceiver 310 down-converts the incoming RF signal to generate an intermediate frequency or baseband signal. The intermediate frequency or baseband signal is sent to the RX processing circuit 325, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or intermediate frequency signal. The RX processing circuit 325 sends the processed baseband signal to the speaker 330 (such as for voice data) or to the processor 340 for further processing (such as for web browsing data).
[0040] The TX processing circuit 315 receives analog or digital voice data from the microphone 320, or other outgoing baseband data from the processor 340. The outgoing baseband data may include web data, email, or interactive video game data. The TX processing circuit 315 encodes, multiplexes, and / or digitizes the outgoing baseband data to generate a processed baseband or intermediate frequency signal. The RF transceiver 310 receives the outgoing processed baseband or intermediate frequency signal from the TX processing circuit 315 and upconverts the baseband or intermediate frequency signal to an RF signal transmitted via the antenna 305.
[0041] The processor 340 may include one or more processors or other processing devices. The processor 340 may execute instructions stored in the memory 360 (such as the OS 361) to control the overall operation of the electronic device 300. For example, the processor 340 may control the reception of forward channel signals and the transmission of reverse channel signals through the RF transceiver 310, the RX processing circuit 325, and the TX processing circuit 315 according to well-known principles. The processor 340 may include any suitable number and type(s) of processors or any other devices arranged in any suitable manner. For example, in an embodiment of the present disclosure, the processor 340 includes at least one microprocessor or microcontroller. Example types of the processor 340 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application specific integrated circuits, and discrete circuits.
[0042] The processor 340 is also capable of executing other processes and programs resident in the memory 360, such as operations for receiving and storing data. The processor 340 may move data into or out of the memory 360 according to the needs of the running processes. In an embodiment of the present disclosure, the processor 340 is configured to execute one or more applications 362 based on the OS 361 or in response to signals received from one or more external sources or an operator. For example, the applications 362 may include encoders, decoders, VR or AR applications, camera applications (for still images and videos), video phone call applications, email clients, social media clients, SMS message clients, virtual assistants, etc. In an embodiment of the present disclosure, the processor 340 is configured to receive and send media content. In an embodiment of the present disclosure, the processor 340 may utilize an improved vertex motion predictor coding as described in the present disclosure.
[0043] The processor 340 is also coupled to the I / O interface 345, which provides the electronic device 300 with the ability to connect to other devices (such as client devices 106 - 114). The I / O interface 345 is a communication path between these accessories and the processor 340.
[0044] The processor 340 is also coupled to an input 350 and a display 355. An operator of the electronic device 300 can use the input 350 to enter data or input into the electronic device 300. The input 350 can be a keyboard, a touch screen, a mouse, a trackball, a voice input, or other devices capable of serving as a user interface to allow a user to interact with the electronic device 300. For example, the input 350 can include speech recognition processing to allow a user to enter voice commands. In an example, the input 350 can include a touch panel, a (digital) pen sensor, keys, or an ultrasonic input device. The touch panel can recognize touch inputs, for example, in at least one of a capacitive scheme, a pressure-sensitive scheme, an infrared scheme, or an ultrasonic scheme. By providing additional input to the processor 340, the input 350 can be associated with one or more sensors 365 and / or a camera. In an embodiment of the present disclosure, the sensors 365 include one or more inertial measurement units (IMUs) (such as accelerometers, gyroscopes, and magnetometers), motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, etc. The input 350 can also include control circuitry. In a capacitive scheme, the input 350 can recognize touch or proximity.
[0045] The display 355 can be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED), an active matrix OLED (AMOLED), or other displays capable of rendering text and / or graphics such as from websites, videos, games, images, etc. The size of the display 355 can be adapted within the HMD. The display 355 can be a single display screen or multiple display screens capable of creating a stereoscopic display. In an embodiment of the present disclosure, the display 355 is a head-up display (HUD). The display 355 can display 3D objects, such as 3D point clouds or meshes.
[0046] The memory 360 is coupled to the processor 340. A portion of the memory 360 can include RAM, and another portion of the memory 360 can include flash memory or other ROM. The memory 360 can include permanent memory (not shown) representing any structure capable of storing and facilitating the retrieval of information such as data, program code, and / or other suitable information. The memory 360 can include one or more components or devices supporting long-term storage of data, such as read-only memory, hard disk drives, flash memory, or optical discs. The memory 360 can also contain media content. The media content can include various types of media, such as images, videos, three-dimensional content, VR content, AR content, 3D point clouds, meshes, etc.
[0047] The electronic device 300 further includes one or more sensors 365. The one or more sensors 365 can measure a physical quantity or detect the activation state of the electronic device 300, and convert the measured or detected information into an electrical signal. For example, the sensor 365 can include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor (such as a gyroscope or a gyroscope sensor and an accelerometer), an eye tracking sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an IR sensor, an ultrasonic sensor, an iris sensor, a fingerprint sensor, a color sensor (such as a red, green, blue (RGB) sensor), etc. The sensor 365 can also include a control circuit for controlling any sensor included therein.
[0048] As discussed in more detail below, one or more of these (one or more) sensors 365 can be used to control a user interface (UI), detect UI input, determine orientation and the user-facing direction for three-dimensional content display recognition, etc. Any one of these (one or more) sensors 365 can be located within the electronic device 300, within a secondary device operably connected to the electronic device 300, within a headset configured to support the electronic device 300, or within a single device in which the electronic device 300 includes a headset.
[0049] The electronic device 300 can create media content, such as generating virtual objects or capturing (or recording) content through a camera. The electronic device 300 can encode the media content to generate a bitstream such that the bitstream can be directly sent to another electronic device, or indirectly sent, such as through Figure 1 the network 102. The electronic device 300 can directly receive a bitstream from another electronic device, or indirectly receive a bitstream, such as through Figure 1 the network 102.
[0050] Although Figure 2 and Figure 3 show examples of electronic devices, various changes can be made to Figure 2 and Figure 3 . For example, Figure 2 and Figure 3 The various components in can be combined, further subdivided, or omitted, and additional components can be added according to specific needs. As a specific example, the processor 340 can be divided into multiple processors, such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs). Additionally, like computing and communication, electronic devices and servers can have a wide variety of configurations, andFigure 2 and Figure 3 does not limit the present disclosure to any particular electronic device or server.
[0051] Figure 4 An example intra - coding process 400 in an embodiment of the present disclosure is shown. Figure 4 The intra - coding process 400 shown is for illustration only. Figure 4 does not limit the scope of the present disclosure to any particular implementation of the intra - coding process.
[0052] As Figure 4 shown, the intra - coding process 400 encodes a mesh frame using an intra - encoder 402. The intra - encoder 402 can be represented or executed by Figure 2 the server 200 shown or Figure 3 the electronic device 300 shown. A base mesh 404, which typically has a smaller number of vertices compared to the original mesh, is created and quantized and compressed in a lossy or lossless manner, and then encoded as a compressed base - mesh bitstream. As Figure 4 shown, a static mesh decoder decodes and reconstructs the base mesh, thereby providing a reconstructed base mesh 406. Then, the reconstructed base mesh 406 undergoes one or more levels of subdivision, and a displacement field is created for each subdivision, which represents the difference between the original mesh and the reconstructed base mesh of the subdivision. In the inter - coding of the mesh frame, the base mesh 404 is encoded by sending vertex motion instead of directly compressing the base mesh. In either case, a displacement field 408 is created. Each displacement in the displacement field 408 has three components represented by x, y, and z. These can be with respect to a canonical coordinate system or a local coordinate system, where x, y, and z represent displacements in the local normal, tangent, and binormal directions. It should be understood that multiple levels of subdivision can be applied, thereby creating multiple subdivided mesh frames, and displacement fields for each subdivided mesh frame are also created.
[0053] Let the number of 3 - D displacement vectors in the displacement 408 of the mesh frame be N. Let the displacement field be represented as d(i)=[dx(i), dy(i), dz(i), 0≤i<N]. The displacement field 408 undergoes one or more levels of wavelet transform 410 to create a level - of - detail (LOD) signal dk(i), i = 0≤i<Nk, 0≤k<numLOD, where k represents the index of the level of detail, Nk represents the number of samples in the level - of - detail signal at level k, and numLOD represents the number of LODs. The LOD signal dk(i) is scalar - quantized.
[0054] As Figure 4As shown, the quantized LOD signal corresponding to the displacement field 408 is encoded into a compressed bitstream. In one or more embodiments of the present disclosure, the quantized LOD signal is packed into a 2D image / video using an image packing operation and losslessly compressed using an image or video encoder. However, another entropy encoder such as an Asymmetric Numeral System (ANS) encoder or a Binary Arithmetic Entropy encoder can be used to encode the quantized LOD signal. There may be other correlations based on previous samples, across components, and across LODs that can be exploited.
[0055] Also as Figure 4 shown, image unpacking of the LOD signal is performed, and an inverse quantization operation and an inverse wavelet transform operation are performed to reconstruct the LOD signal. Another inverse quantization operation is performed on the reconstructed base mesh 406, and the reconstructed base mesh 406 is combined with the reconstructed LOD signal to reconstruct the deformed mesh. An attribute transfer operation is performed using the deformed mesh, the static / dynamic mesh, and the attribute map. A point cloud is a set of 3D points and attributes representing the surface or volume of an object, such as color, normal, reflectivity, point size, etc. These attributes are encoded as a compressed attribute bitstream. As Figure 4 shown, the encoding of the compressed attribute bitstream may also include a padding operation, a color space conversion operation, and a video encoding operation. Figure 4 The various functions or operations shown in Figure 4 can be controlled by the control process 412. The intra-frame encoding process 400 outputs a compressed bitstream that can be sent, for example, to an electronic device such as the server 104 or the client devices 106 - 116 and decoded by the electronic device. As Figure 4 shown, the output compressed bitstream may include a compressed base mesh bitstream, a compressed displacement bitstream, and a compressed attribute bitstream.
[0056] Although Figure 4 a block diagram of an example intra-frame encoding process 400 is shown, various changes can be made to Figure 4 it. For example, the number and placement of the various components of the intra-frame encoding process 400 can vary according to need or desire. Additionally, the intra-frame encoding process 400 can be used in any other suitable process and is not limited to the specific process described above. In an embodiment of the present disclosure, only the first (x) component of the displacement may be created and encoded, and the other two components (y and z) may be assumed to be 0. In this case, a flag can be signaled in the bitstream to indicate that the bitstream contains only data corresponding to the first (x) component, and when decompressing and reconstructing the displacement field 408, the other two components (y and z) should be assumed to be zero. As an example, Figure 4 the intra-frame encoding process 400 of
[0057] Figure 5A andFigure 5B Illustrates an example inter-frame mesh frame encoding process 500. Figure 5A and Figure 5B The inter-frame mesh frame encoding process 500 shown in is for illustrative purposes only. Figure 5A and Figure 5B does not limit the scope of the present disclosure to any particular implementation of the inter-frame mesh frame encoding process. For example, Figure 5A and Figure 5B the process 500 of can be described as being executed using the electronic device 300 of. For example, the process 500 can be used with any suitable system and any suitable electronic device (e.g., Figure 3 the server 200 of). Figure 2
[0058] For an inter-frame mesh frame, the base mesh is encoded by sending vertex 3D motions instead of directly compressing the base mesh. When the mesh connectivity and the number of vertices are the same as those of the previous mesh frame, inter-frame encoding is used. These vertex 3D motions are used to determine the inter-frame mesh frame, as Figure 5A shown, where the motion vector (mv) for each vertex of the intra-frame mesh frame is used to predict each vertex in the inter-frame mesh frame.
[0059] The predicted value of the vertex motion vector is typically calculated as the average of the vertex motion vectors of adjacent vertices, but one or more embodiments of the present disclosure may use any combination of the vertex motion vectors of adjacent vertices, such as an average value, a weighted average value, a median value, a maximum value, a minimum value, etc. For a given vertex, a flag can be used to indicate whether to directly send the 3D vertex motion vector of the vertex or whether to send the difference between the vertex motion vector of the vertex and its predicted value. For example, when using vertex motion vector encoding, for a given vertex A, such as Figure 5A shown, the flag is used to indicate whether to send the vertex motion vector of A or whether to send the difference between the vertex motion vector of A and its predicted value. Therefore, the predicted value of the vertex motion vector can be calculated as the average (or other type of combination) of the vertex motion vectors of adjacent vertices (such as the vertices B, C, and D in Figure 5A ), which can be expressed as follows.
[0060]
[0061] Here, is the difference between the vertex motion vector of A ( ) and its predicted value ( ).
[0062] As explained above, the predictor value can be calculated as the average motion vector of available adjacent vertices. This can be expressed as follows.
[0063]
[0064] To determine available adjacent vertices for a set of vertices 501 in a mesh frame, process 500 includes creating a vertex-to-triangle adjacency list in step 502, as Figure 5B shown. This can include the electronic device 300 iterating through each vertex and calculating, for each vertex, the number of adjacent triangles, which can be a variable number. Then, as Figure 5B shown, in step 502, the electronic device 300 iterates through each vertex and creates a pointer table to reserve a variable amount of storage for each vertex in the vertex-to-triangle adjacency list. The electronic device 300 fills the vertex-to-triangle adjacency list with the variable number of triangle neighbors for all vertices.
[0065] Process 500 also includes determining a list of available vertex neighbors in step 504. As Figure 5B shown, for a given vertex (vertex A in this example), and assuming the vertex transmission order is D, C, A, B, the electronic device 300 calculates a list of adjacent vertices based on the list of adjacent triangles created in step 502. In this example, this provides a list of adjacent vertices that includes vertices C, D, and B. Then, the electronic device 300 determines the adjacent available vertices by pruning the list to include only available vertices (i.e., vertices that have already been received). Thus, in Figure 5B the example shown, since vertex B is received after vertex A, vertex B is pruned from the list because it is not an available vertex. Then, the electronic device 300 uses the determined adjacent available vertices (e.g., the average motion of the available adjacent vertices (vertices C and D)) to determine the motion vector prediction factor for vertex A. It should be understood that this process can be performed for each vertex in the mesh frame to determine the motion vector prediction factor for each vertex, which will be used to create an inter-frame mesh frame from a previous intra-frame mesh frame.
[0066] Although Figure 5A and Figure 5B illustrate an exemplary inter-frame mesh frame encoding process 500, various changes can be made to Figure 5A and Figure 5B . For example, although shown as a series of steps, the individual steps in Figure 5A and Figure 5B can overlap, occur in parallel, or occur any number of times.
[0067] Figure 6A and Figure 6B illustrate an example adjacent vertex determination process 600 in an embodiment of the present disclosure. Figure 6A and Figure 6B The process 600 shown in is for illustrative purposes only. Figure 6A andFigure 6B The scope of the present disclosure is not limited to any particular implementation of the adjacent vertex determination process. For example, the process 600 of FIG. 6 can be described as being executed using Figure 3 the electronic device 300. For example, the process 600 can be used with any suitable system and any suitable electronic device (e.g., Figure 2 the server 200).
[0068] As referred to Figure 5A and Figure 5B The calculation of adjacent motion vectors is a complex process that involves iterating through all triangles and their connectivity. Most of the complexity in vertex neighbor calculations stems from considering a variable number of neighbors for each vertex. The process 600 of the present disclosure places a limit on the maximum number of vertex neighbors used in the calculation of motion vector predictors to improve the efficiency of inter-frame mesh frame prediction.
[0069] As Figure 6A shown, for each vertex in a set of vertices 601 in a mesh frame, a fixed-length table is created in step 602 that stores vertex neighbors according to the imposed maximum number of vertex neighbors (i.e., the limit). The maximum number of vertex neighbors is signaled in the bitstream such that the decoder can determine the maximum number of neighbors for a vertex. For example, as Figure 6A shown, the vertex neighbor limit is set to the value 3 such that at most only two vertex neighbors are stored for each vertex. This avoids the need to prune neighbors to find available neighbors.
[0070] Additionally, as Figure 6B shown, the process 600 can be further optimized by determining the fixed-length vertex neighbor table only for intra-frame mesh frames and reusing the fixed-length vertex neighbor table for inter-frame mesh frames as shown in step 604. In this way, the vertex neighbor table can be reused instead of recalculating neighbors for inter-frame mesh frames because inter-frame mesh frame coding is used only when the mesh connectivity and number of vertices are the same as in the previous mesh frame. This improves efficiency by avoiding having to redetermine adjacent vertices for inter-frame mesh frames.
[0071] Although Figure 6A and Figure 6B show an example adjacent vertex determination process 600, various changes can be made to Figure 6A and 6B . For example, although shown as a series of steps, Figure 6A and Figure 6BThe various steps in [the process] can overlap, occur in parallel, or occur any number of times. Additionally, process 600 can be further optimized with respect to duplicate vertex removal to maintain the temporal consistency of the vertex neighbor table. Generally, duplicate vertex removal changes the transmission order of vertices, which changes the availability of adjacent vertices. Even if the mesh connectivity remains the same, this changes the vertex neighbor table from one frame to the next. However, the present disclosure provides that duplicate vertices can be computed only for intra-mesh frames and that the duplicate vertex table can be reused for inter-mesh frames.
[0072] For example, Figure 7 illustrates an example mesh frame decoding process 700 in an embodiment of the present disclosure. Figure 7 The frame decoding process 700 shown in [the figure] is for illustrative purposes only. Figure 7 It does not limit the scope of the present disclosure to any particular implementation of the mesh frame decoding process.
[0073] The decoding process 700 involves a demultiplexer 702 that receives an incoming bitstream. The demultiplexer separates the individual component bitstreams from the incoming bitstream, including a compressed base mesh bitstream, a compressed displacement bitstream, and a compressed attribute bitstream, such as those described with respect to Figure 4 The compressed attribute bitstream is decoded using a video decoder 704, the decoded attributes are processed using a color space conversion operation 706, and the original attributes of the mesh are restored.
[0074] The decoding process 700 also includes decoding the displacement bitstream using a video decoder 708, which in an embodiment of the present disclosure can be the same video decoder as video decoder 704. As part of restoring the position displacement data 716, the decoded displacement data undergoes an image unpacking operation 710, an inverse quantization operation 712, and an inverse wavelet transform operation 714. Restoring the position displacement data 716 can also include performing one or more subdivision operations 718 on the mesh frame restored using a base mesh decoder 720 and extracting the x, y, z components 722 (normals, tangents, bitangents) from the subdivided mesh frame. The base mesh decoder 720 can perform an inverse quantization operation 721 before the subdivision operations 718 are performed.
[0075] The base mesh decoder 720 obtains the base mesh bitstream provided by the demultiplexer 702 and reconstructs an intra-mesh frame from the base mesh bitstream using a static mesh decoder 724. Data from the intra-mesh frame is used to perform a vertex deduplication operation 726 and build a vertex deduplication table 728. A mesh buffer 730 provides the decoded intra-frame to a motion decoder 732. The motion decoder 732 also receives inter-frame data and uses the intra-frame data, inter-frame data, and associated tables to reconstruct the base mesh at step 734.
[0076] However, prior to this, the reconstruction of the base mesh will be used as part of the vertex deduplication operation 726 and the creation of the vertex deduplication table 728. However, as described above, since the duplicate vertex table can be reused for the inter-frame mesh frames, the process 700 does not need to perform the additional step of using the reconstructed base mesh during vertex deduplication and re-determining the duplicates of the inter-frame mesh frames, thereby improving the overall efficiency of the decoding process 700.
[0077] Although Figure 7 FIG. shows a block diagram of an example frame decoding process 700, various changes can be made to Figure 7 it. For example, the number and arrangement of the various components of the frame decoding process 700 can vary according to need or desire. Additionally, the frame decoding process 700 can be used in any other suitable process and is not limited to the specific process described above. Moreover, although shown as a series of steps, Figure 7 the individual steps in
[0078] can overlap, occur in parallel, or occur any number of times.
[0079] The optimizations described above improve the overall coding efficiency. For example, it has been shown that when the number of adjacent motion vectors is limited to 3, the D1, D2, luminance, Cb, and Cr Bjontegaard Delta (BD) rates are all 0.0%. It has also been shown that the optimizations provide a reduction in motion bits. It has been found that the optimizations described above provide a reduction of approximately 30% in motion decoding time and a reduction of approximately 50% in vertex neighbor table memory usage.
[0079] Additional information, examples, and embodiments regarding these optimizations are described below with respect to Figure 8 it. Figure 8 FIG. shows an example set of vertices 800 in an embodiment of the present disclosure. Figure 8 The example set of vertices 800 shown in Figure 8 is for illustrative purposes only.
[0080] Figure 8 It does not limit the scope of the present disclosure to any particular set, number, or arrangement of vertices. Figure 8 An example set of vertices 800 is provided that includes five vertices: vertex A, vertex B, vertex C, vertex D, and vertex E, which are used hereinafter to emphasize and illustrate various aspects of the present disclosure. When using vertex motion vector coding, for a given vertex, such as vertex A in this example, a flag is used to indicate whether to send the vertex motion vector of A or the difference between the vertex motion vector of A and its predicted value. The predicted value of the vertex motion vector is calculated as a combination (e.g., average, weighted average, median, maximum, minimum, etc.) of the vertex motion vectors of adjacent vertices (e.g., Figure 8 B, C, D, E in
[0081] Encoding of vertex motion information may include indicating which one of multiple vertex motion vectors (VMVs) to use for each vertex at the decoder. For example, for an example set of vertices 800, let mA be the actual VMV of vertex A. Let mB, mC, mD, mE be the actual VMVs of the neighboring vertices of A, as Figure 8 shown. In an embodiment of the present disclosure, multiple VMV predictors are calculated for each vertex A. The syntax element vmv_id is sent to the receiver to indicate which one of the multiple VMV predictors to use at the decoder.
[0082] For example, in an embodiment of the present disclosure, vmv_id takes N + 1 values, where N is the number of neighboring vertices. For Figure 8 the example in, vmv_id can take five values: 0, 1, 2, 3, and 4. In this example, the VMV predictors for A are calculated as shown in Table 1 below.
[0083]
[0084] Table 1: Vertex Motion Vector Predictors In Table 1, vmv_id 0 is associated with mB and indicates that for vertex A, the VMV predictor is the difference between mA and mB. Vmv_id 1 is associated with mC and indicates that for vertex A, the VMV predictor is the difference between mA and mC, and so on. In Table 1, vmv_id with 4 is mapped to the value 0, which indicates that the VMV predictor is mA itself. In Table 1, the mapping between vmv_id and the VMV predictor is just an example mapping. It should be understood that other mappings can be used.
[0085] In an embodiment of the present disclosure, vmv_id takes N + 2 values, where N is the number of neighboring vertices. For Figure 8 the example in, vmv_id can take six values: 0, 1, 2, 3, 4, and 5. In this example, the VMV predictors for A are calculated as shown in Table 2 below.
[0086]
[0087] Table 2: Vertex Motion Vector Predictors The difference between Table 2 and Table 1 is that Table 2 also includes an additional vmv_id with the value 5, which is mapped to the average (or other type of combination) of mB, mC, mD, and mE. In Table 2, the mapping between vmv_id and the VMV predictor is just an example mapping. It should be understood that other mappings can be used.
[0088] In an embodiment of the present disclosure (e.g., as shown in Table 2), the vmv_id signals the use of a particular VMV predictor among a plurality of VMV predictors of a set, the plurality of VMV predictors of the set including the average of a subset of all adjacent VMVs, zero VMV, and / or at least one VMV of adjacent vertices.
[0089] In an embodiment of the present disclosure (e.g., as shown in Table 1), the vmv_id signals the use of a particular VMV predictor among a plurality of VMV predictors of a set, the plurality of VMV predictors of the set including zero VMV and / or at least one VMV of adjacent vertices.
[0090] In an embodiment of the present disclosure (e.g., where other mappings are generalized), the vmv_id signals a particular VMV predictor among a plurality of VMV predictors of a set, the plurality of VMV predictors of the set including the median of a subset or all adjacent VMVs, zero VMV, and / or at least one VMV of adjacent vertices. In an embodiment of the present disclosure, the syntax element vmv_id can be encoded using any entropy coding technique, such as unary code, exponential Golomb code, Huffman code, arithmetic code, etc.
[0091] In an embodiment of the present disclosure, a set of adjacent vertices for each vertex is pre-computed from the vertex connectivity of a reference grid from which the current grid is predicted. The adjacent vertex information can be stored in tabular form or by using other data structures such as linked lists, trees, etc.
[0092] For example, Table 3 below shows Figure 8 an example of an adjacent vertex table for a grid, assuming the vertices are sent in the order of A, B, C, D, E.
[0093]
[0094] Table 3: Adjacent Vertex Table Note that although in Figure 8 vertex A has vertices B, C, D, E as neighbors, since A is the first vertex to be sent, there is no other data available for constructing predictors. Thus, from the perspective of calculating VMV predictors, vertex A has no neighbors. It should be understood that the adjacent vertex table in Table 3 is an unoptimized process for determining all available neighbors from a vertex, such as discussed with respect to Figure 5A and Figure 5B above.
[0095] In an embodiment of the present disclosure, each reference grid in the dynamic grid sequence has an associated adjacent vertex table that is valid as long as the reference grid is valid (i.e., in the reference decoded grid buffer). In this way, when a reference grid is used in temporal prediction, the adjacent vertex table can be easily used to calculate the VMV prediction factor.
[0096] As described in the present disclosure, such as with respect to Figure 6A and Figure 6B , the maximum number of adjacent vectors used in VMV prediction factor calculation can be limited to a fixed value to improve efficiency. For example, if the maximum number of adjacent vertices shown in Table 3 is fixed at 1, this results in Table 4 below.
[0097]
[0098] Table 4: Adjacent Vertex Table with Fixed Maximum Number of Neighbors (max_num_neighbors_vmv = 1) Let N be the total number of available neighbors. Then, in an embodiment of the present disclosure, the max_num_neighbors_vmv - 1 neighbors in order and the last available neighbor are used to calculate the VMC prediction factor. For example, when max_num_neighbors_vmv = 3, instead of using mB, mC, mD in order, mB, mC, mE are used to calculate the VMV prediction factor.
[0099] In an embodiment of the present disclosure, a syntax element (such as "max_num_neighbors_vmv" or some other name) is signaled in the bitstream for the maximum number of adjacent vertices used to calculate the VMV prediction factor. This syntax element can be signaled at the grid level, sub - grid level, sequence level, etc.
[0100] In an embodiment of the present disclosure, a combination of the VMV of adjacent vertices and the VMV of temporally co - located vertices in the reference grid is used to calculate the VMV prediction factor. If the VMV of adjacent vertices is not available, only the VMV of temporally co - located vertices in the reference grid is used.
[0101] In an embodiment of the present disclosure, the predicted VMV is calculated as a weighted sum of the VMV of adjacent vertices and the VMV of temporally co - located vertices in the reference grid. The weights can be uniform or non - uniformly varying. The weights can be signaled in the bitstream or be fixed a priori to known values.
[0102] As in the present disclosure with respect to Figure 5A and Figure 5BAs described, in embodiments of the present disclosure, a table data structure is created to store the adjacent vertices (or vertex indices in embodiments of the present disclosure) of all vertices in the grid. This table data structure can be referred to as a vertex adjacency table. The creation of the vertex adjacency table can be represented as follows: numNeighborsTable[v] = 0 for all vertices v in the grid for all triangles in the grid, Let v1, v2, v3 be the vertices of the triangle AddNeighbor(v1, v2) AddNeighbor(v2, v1) AddNeighbor(v1, v3) AddNeighbor(v3, v1) AddNeighbor(v3, v2) AddNeighbor(v2, v3) AddNeighbor(vertex vA, vertex vB) if (vB is available) if (vB is not yet a neighbor of vA in vertexAdjacencyTable) n = numNeighborsTable[vA] vertexAdjacencyTable[vA][n] = vB numNeighborsTable[vA]++ In one or more embodiments of the present disclosure, the table can be stored as a 1D vector, 2D array, list, etc., or other similar data structures. As shown above, let "numNeighborsTable" be the table that stores the number of available neighbors of each vertex in the grid. In the encoder, a vertex is considered available if it has been processed for transmission. In the decoder, a vertex is considered available if it has been received and decoded. In this example, if vB is available and not yet present in numNeighborsTable, the "AddNeighbor()" procedure adds vertex vB as a neighbor of vA.
[0103] In one or more embodiments of the present disclosure, to calculate the VMV prediction factor for any given vertex, first the adjacent vertex numbers or IDs are read from the adjacent vertex table. Then, the motion vectors corresponding to these adjacent vertices are used to calculate the VMV prediction factor based on vmv_id (such as the average, median, etc. of the VMVs of the adjacent vertices). This can be expressed as follows for calculating the vertex motion vector prediction factor of vertex vA: Read the number of available neighbors N of vA by indexing in numNeighborsTable.
[0104] N = numNeighborsTable[vA] Index in vertexAdjacencyTable to read the list of available adjacent vertices: v0 = vertexAdjacencyTable[vA][0] v1 = vertexAdjacencyTable[vA][1] … vNm1 = vertexAdjacencyTable[vA][N - 1] The vertex motion vector prediction factor is the average of the motion vectors of v1, v2, …, vNm1.
[0105] In an embodiment of the present disclosure, other combination methods such as weighted average, median, maximum, minimum, etc. can be used instead of the average (e.g., based on vmv_id as described above).
[0106] In an embodiment of the present disclosure, the maximum number of available neighbors (the above syntax element: max_num_neighbors_vmv or other names) can be fixed to a predetermined value, which is signaled to the decoder in headers such as sequences, frames, sub - grids, slices, sub - frames, parallel blocks, etc.
[0107] In an embodiment of the present disclosure, vertexAdjacencyTable is only created for intra - grid frames. vertexAdjacencyTable is reused for subsequent inter - grid frames, which can be expressed as follows: If (the current frame is intra - coded) Create vertexAdjacencyTable Use vertexAdjacencyTable to calculate the vertex motion vector prediction factor else Use the vertexAdjacencyTable of the previously intra-coded picture to calculate the vertex motion vector predictor In an embodiment of the present disclosure, a vertex adjacency table is created for an intra-grid frame. In an embodiment of the present disclosure, when hierarchical inter-coding is used, it is inherited from a reference grid frame (which is an intra-grid frame or other inter-grid frame that has been coded) and the vertex adjacency table is not recalculated.
[0108] As described with respect to Figure 6A 、 Figure 6B and Figure 7 A duplicate vertex table can also be used in mesh coding. Duplicate vertices are vertices that have the same geometric position in a reference sub-mesh. For example, when there is a T-junction in a mesh, duplicate vertices may occur. An example duplicate vertex table is shown in Table 5 below.
[0109]
[0110] Table 5: Duplicate Vertex Table Table 5 contains vertex mapping information. In this example, the second vertex (table index 1) and the fourth vertex (table index 3) are duplicate vertices and thus they are mapped to the same value 1. The fifth vertex and the seventh vertex are duplicate vertices and thus they are mapped to the same value 3. Other mechanisms for indicating duplicate vertices can also be used without departing from the scope of the present disclosure.
[0111] In an embodiment of the present disclosure, a table storing information about duplicate vertices is created for an intra-frame. When hierarchical inter-coding is used, it is inherited from a reference grid frame (which is an intra-grid frame or other inter-grid frame that has been coded) and the duplicate vertex table / data structure is not recalculated.
[0112] In an embodiment of the present disclosure, a flag is signaled to indicate that the duplicate vertex table / data structure is not calculated for an inter-frame, but is inherited from a reference grid frame (which is an intra-grid frame or other inter-grid frame that has been coded). In an embodiment of the present disclosure, other mesh-related tables can be created based on the mesh frame type.
[0113] Various standards regarding vertex mesh (V-MESH) and dynamic mesh coding have been proposed. The following documents are incorporated herein by reference in their entirety as if fully set forth herein: “V-Mesh Test Model v1”, ISO / IEC SC29 WG07 N00404, July 2022.
[0114] "WD 2.0 of V-DMC", ISO / IEC SC29 WG07 N00546, January 2023.
[0115] "WD 3.0 of V-DMC", ISO / IEC SC29 WG07 N00611, April 2023.
[0116] "WD 4.0 of V-DMC", ISO / IEC JTC 1 / SC 29 / WG 07 N00611, August 2023.
[0117] To provide improvements to the vertex motion vector prediction factor according to the present disclosure, the standard can be updated to specify the following: H.8.1.3.1.1 General base grid sequence parameter set RBSP syntax
[0118] H.8.3.1.1 General base grid sequence parameter set RBSP semantics … bmsps_inter_mesh_max_num_neighbors_minus1 plus 1 indicates the maximum number of vertex neighbors used in the calculation of the motion vector prediction factor. bmsps_inter_mesh_max_num_neighbors_minus1 shall be in the range of 0 to 255, inclusive of 0 and 255.
[0119] … H.11.1 General principles The reconstruction process is carried out by calling the various procedures described below.
[0120] Reconstruct the intra submesh as defined in H.11.2 and call the post - reconstruction process described in section H.11.4, where the reconstructed submesh is taken as input and the parameters referenceSubmeshIntegratedIndices, referenceSubmeshDupVertCount and the integrated submesh are taken as output. Then export the vertex neighbour table as defined in H.11.6, where the integrated submesh is taken as input and the tables submeshVertexNeighbours and submeshVertexNeighboursCounts are taken as output. Reconstruct the inter submesh as defined in H.11.3 and call the post - reconstruction process described in section H.11.5, where the reconstructed submesh and the parameters referenceSubmeshIntegratedIndices, referenceSubmeshDupVertCount are taken as input and the integrated submesh is taken as output.
[0121] H.11.2 Reconstruction of vertices for the intra submesh H.11.3 Reconstruction of vertices for the inter submesh The inputs to this process are: - motionGroupSize, which is the size of vertex grouping in motion vector coding.
[0122] - submeshFaceCount, which is a variable indicating the number of faces in the current submesh and the reference submesh.
[0123] - submeshFaceIndices, which is a 2D array of size submeshFaceCount times 3, indicating the connectivity indices associated with the current submesh and the reference submesh.
[0124] - referenceSubmeshVertexPositions, which is a 2D array of size submeshVertexCount times 3, indicating the positions of the reference submesh positions.
[0125] - referenceSubmeshDupVertCount, which is a variable indicating the number of pairs of duplicate vertices in the reference submesh.
[0126] - referenceSubmeshVertexCountClean, which is a variable indicating the number of vertices in the reference integrated submesh.
[0127] -referenceSubmeshIntegratedIndices, which is a 2D array of size referenceSub-meshDupVertCount times 2, indicating index pairs of duplicate vertices in the reference submesh. The two indices in each pair indicate that the two vertices are duplicates, i.e., have the same position.
[0128] -referenceSubmeshVertexPositionsClean, which is a 2D array of size referenceSubmesh-VertexCountClean times 3, indicating the positions of the reference integrated submesh positions.
[0129] -submeshVertexNeighboursCounts, which is a 1D array indicating the number of neighbours of each vertex of the submesh.
[0130] -submeshVertexNeighbours, which is a 2D array of size submeshVertexCount times (bmsps_inter_mesh_max_num_neighbors_minus1 + 1), indicating the indices of its neighbours for each vertex v according to the mesh connectivity.
[0131] The output of this process is currentSubmeshVertexPositions, which is a 2D array of size submeshVertexCount times 3, indicating the positions of the current frame submesh.
[0132] The following arrays are exported during the submesh position reconstruction process: -currentSubmeshMotionVectors, which is a 2D array of size submeshVertexCount times 3, indicating the motion vector of each vertex v in the current frame.
[0133] -currentSubmeshPredictedMotionVectors, which is a 2D array of size submeshVertexCount times 3, indicating the predicted motion vector of each vertex v.
[0134] - Since some integrated vertices may have multiple motion vectors signaled by sismu_multi_mv_idx, additional vertices with the number of sismu_multi_mv_num need to be added. This makes submeshMotionCount greater than referenceSubmeshVertexCountClean, and is derived as follows: - submeshMotionCount = referenceSubmeshVertexCountClean + sismu_multi_mv_num - When the vertex index v is greater than referenceSubmeshVertexCountClean, additional vertices va are iteratively added as follows according to sismu_multi_mv_idx.
[0135] idx = sismu_multi_mv_idx[v - referenceSubmeshVertexCountClean]; integrate_to = referenceSubmeshIntegratedIndices[idx][1]; it = lower_bound(referenceSubmeshIntegratedIndices, integrated_to); shift = distance(referenceSubmeshIntegratedIndices, it); va = integrate_to - shift; The function lower_bound(referenceSubmeshIntegratedIndices, value) returns a pointer to the first element in referenceSubmeshIntegratedIndices whose second component is equal to value, or NULL if no such element is found.
[0136] The function distance(vector, pointer) returns the index of the element in the vector pointed to by the pointer.
[0137] Since duplicate vertices are integrated in the reference submesh, submeshVertexCount is greater than referenceSubmeshVertexCountClean, and is derived as follows: submeshVertexCount = referenceSubmeshVertexCountClean + referenceSubmeshDupVertCount The k-th component currentSubmeshVertexPositions[v][k] of the position of the vertex with index v is derived as follows: currentSubmeshVertexPositions[v][k] = referenceSubmeshVertexPositionsClean[vr][k] + currentSubmeshMotionVectors[vm][k] where vr and vm are the corresponding indices assigned as follows:
[0138] The function find_if(sismu_multi_mv_idx, baseIntegrateIndices, v) returns a pointer to the first element i in sismu_multi_mv_idx such that baseIntegrateIndices[i][0] == v, or NULL if no such element is found.
[0139] The k-th component currentSubmeshMotionVectors[v][k] of the motion vector associated with the vertex with index v is derived as follows: The group index g of the vertex with index v is derived as follows: g = v / motionGroupSize The prediction mode sismu_mv_pred_mode_vertex[v] of the vertex with index v is equal to the prediction mode sismu_mv_pred_mode_group[g] of the group with index g: sismu_mv_pred_mode_vertex[v] = sismu_mv_pred_mode_group[g] If the prediction mode sismu_mv_pred_mode[v] is equal to 0, then currentSubmeshMotionVectors[v][k] = VertexMotionVectorResiduals[v][k] Otherwise (when sismu_mv_pred_mode[v] equals 1), currentSubmeshMotionVectors[v][k]=VertexMotionVectorResiduals[v][k]+currentSubmeshPredictedMotionVectors[v][k] The predicted motion vectors currentSubmeshPredictedMotionVectors[v] are derived by applying the following procedure:
[0140] H.11.4 Post - reconstruction procedure for integrating duplicate vertices in intra sub - meshes The inputs to this procedure are: - The sub - mesh reconstructed from the intra sub - mesh The outputs of this procedure are: - referenceSubmeshIntegratedIndices, which is a 2D array of size referenceSubmeshDupVertCount by 2, indicating index pairs of duplicate vertices in the reconstructed sub - mesh.
[0141] - referenceSubmeshDupVertCount.
[0142] - The integrated sub - mesh, in which duplicate vertices are integrated and the connectivity is updated.
[0143] The procedure proceeds as follows: Step 1: Search for duplicate vertices as follows. In the reconstructed sub - mesh, search for all pairs of duplicate vertices in the reconstructed base mesh by iteratively checking if the geometric positions of two vertices are the same. Each pair of duplicate vertices (A(j), B(j)) has exactly the same geometric position, where j = 1, …, referenceSubmeshDupVertCount and A(j)>B(j). The list of index pairs of duplicate vertices forms a 2D array of size referenceSubmeshDupVertCount by 2, i.e., referenceSubmeshIntegratedIndices, which is one of the outputs. If there are no duplicate vertices, then referenceSubmeshIntegratedIndices = NULL. NULL is a special pointer with a zero value indicating that the pointer is not intended to point to an accessible memory location.
[0144] Step 2: Integrate duplicate vertices as follows. After searching for duplicate vertices, we integrate all pairs of duplicate vertices (A(j), B(j)) into a single integrated vertex B(j) in the reconstructed base mesh of the reference frame. The integration is to remove the geometric position of A(j), replace the index of A(j) with B(j), and decrement by 1 all vertex indices greater than A(j) in the connectivity. If referenceSubmeshIntegratedIndices = NULL, this step will be skipped.
[0145] H.11.5 Post - reconstruction process for integrating duplicate vertices of the inter - frame sub - mesh The inputs to this process are: - The reconstructed sub - mesh from the inter - frame sub - mesh - referenceSubmeshIntegratedIndices, which is a 2D array of size referenceSubmeshDupVertCount times 2, indicating index pairs of duplicate vertices in the intra - frame reconstructed sub - mesh.
[0146] - referenceSubmeshDupVertCount.
[0147] The outputs of this process are: - The integrated sub - mesh This process proceeds as follows: Traverse the list of (A(j), B(j)) in referenceSubmeshIntegratedIndices and merge (A(j), B(j)) into a single integrated vertex B(j) in the reconstructed base mesh of the reference frame. The integration is to remove the geometric position of A(j), replace the index of A(j) with B(j), and decrement by 1 all vertex indices greater than A(j) in the connectivity. If referenceSubmeshIntegratedIndices = NULL, this step will be skipped.
[0148] H.11.6 Vertex neighbor table calculation The inputs to this process are: - bmsps_inter_mesh_max_num_neighbors_minus1 - submeshFaceCount, which is a variable indicating the number of faces in the current sub - mesh and the reference sub - mesh.
[0149] - submeshFaceIndices, which is a 2D array of size submeshFaceCount times 3, indicating the connectivity indices associated with the current sub - mesh and the reference sub - mesh.
[0150] The output of the process is: - submeshVertexNeighboursCounts, which is a 1D array indicating the number of neighbours of each vertex of the submesh.
[0151] - submeshVertexNeighbours, which is a 2D array of size submeshVertexCount multiplied by (bmsps_inter_mesh_max_num_neighbors_minus1 + 1), indicating the indices of the neighbours of each vertex v according to the mesh connectivity.
[0152] The maximum number of neighbours maxVertexNeighbourCount is set to be equal to bmsps_inter_mesh_max_num_neighbors_minus1 + 1.
[0153]
[0154]
[0155] The order in which the vertices of the mesh are encoded or decoded can be changed on a per-frame basis (or on a per-sequence basis etc., without loss of generality). Let scan[i][j], i = 0, …, Ns - 1, j = 0, …, maxVertexNeighbourCount - 1 denote the scan order of the vertices for a frame, where Ns is the number of different scan orders. The different scan orders can be based on a depth-first traversal of the mesh or a traversal along the direction of the most available already-encoded vertices of the mesh or any other traversal without loss of generality. In an embodiment of the present disclosure, multiple vertex adjacency lists are computed, one vertex adjacency list for each scan order. If the Ns scan orders to be used in the mesh sequence are known a priori (either via signalling or fixed in a standard), then these vertex adjacency lists can be computed at the end of the intra-grid frame. If the scan order is signalled on a per-frame basis, then they can also be computed in the first inter-grid of a particular scan order. The vertex adjacency lists are then reused for subsequent inter-grid frames of that scan order.
[0156] In an embodiment of the present disclosure, the maximum number of vertex neighbours is signalled in the base mesh sequence parameter set using the syntax element bmsps_inter_mesh_max_num_neighbors_minus1 as shown in Section H.8.1.3.1.1 above.
[0157] In an embodiment of the present disclosure, vertex neighbor tables (submeshVertexNeighbours and submeshVertexNeighboursCounts) are calculated as shown in Section H.11.6 above.
[0158] In an embodiment of the present disclosure, reference vertex tables (referenceSubmeshIntegratedIndices and referenceSubmeshDupVertCount) are created at the end of intra-grid frame processing and are reused for inter-grid frames as shown in Section H.11.5 above.
[0159] In an embodiment of the present disclosure, vertex neighbor tables (submeshVertexNeighbours and submeshVertexNeighboursCounts) are reused for inter-grid frames as shown in Section H.11.3 above.
[0160] Figure 9 An example encoding method 900 for improved vertex motion vector prediction factor encoding in an embodiment of the present disclosure is shown. For example, Figure 9 Method 900 is described as being executed using Figure 3 electronic device 300. For example, method 900 can be used with any suitable system and any suitable electronic device (e.g., Figure 2 server 200).
[0161] As Figure 9 shown, at step 910, electronic device 300 can identify one or more vertex neighbors for vertices of a grid frame based on a set limit on the number of one or more vertex neighbors. This can include the processor of electronic device 300 identifying the last vertex neighbor in a neighbor sequence associated with a vertex, and at least one additional vertex neighbor corresponding to the number of one or more vertex neighbors minus one in the neighbor sequence. In an embodiment of the present disclosure, the identified one or more vertex neighbors are associated with an intra-grid frame. In an embodiment of the present disclosure, electronic device 300 can reuse the identified one or more vertex neighbors for inter-grid frames as described in the present disclosure.
[0162] At step 920, electronic device 300 can determine a plurality of vertex motion vector (VMV) prediction factors for a vertex based on the identified one or more vertex neighbors. At step 930, electronic device 300 maps each of the plurality of VMV prediction factors to one of a plurality of VMV identifiers. For example, a VMV prediction factor can be a combination of one or more vertex neighbors (e.g., average, weighted average, median, maximum, minimum, etc.).
[0163] At step 940, the electronic device 300 may encode a compressed video bitstream that signals a set limit on the number of one or more vertex neighbors and signals one of a plurality of VMV identifiers that indicates the VMV predictor among the plurality of VMV predictors to be used for a vertex, such as described in Table 2 of the present disclosure. In one or more embodiments of the present disclosure, the electronic device 300 encodes the set limit on the number of one or more vertex neighbors in at least one of a sequence header, a frame header, a sub-grid header, a slice header, a sub-frame header, or a parallel block header of the compressed video bitstream. In one or more embodiments of the present disclosure, the electronic device 300 may also generate a duplicate vertex data structure that stores information related to duplicate vertices of one or more mesh frames. The electronic device 300 may also set a flag in the compressed video bitstream that signals that an inter-frame mesh frame inherits the duplicate vertex data structure. In one or more embodiments of the present disclosure, the encoder calculates the VMV predictor and sends the difference between the VMV for the vertex and the associated predicted value of the VMV predictor as part of the bitstream.
[0164] In an embodiment of the present disclosure, the electronic device 300 may output a bitstream. The output bitstream may include, for example, Figure 4 the compressed base mesh bitstream, the displacement bitstream, and the attribute bitstream as shown in, and the above signaling elements. The output bitstream may be sent to an external device or a memory on the electronic device 300.
[0165] Although Figure 9 shows an example of an encoding method 900 for improved vertex motion vector predictor encoding, various changes may be made to Figure 9 it. For example, although shown as a series of steps, the individual steps in Figure 9 may overlap, occur in parallel, or occur any number of times.
[0166] Figure 10 shows an example decoding method 1000 for improved vertex motion vector predictor encoding in an embodiment of the present disclosure. For example, Figure 10 the method 1000 of Figure 3 is described as being performed using the electronic device 300 of Figure 2 . For example, the method 1000 may be used with any suitable system and any suitable electronic device (e.g.,
[0167] As Figure 10As shown, at step 1010, the electronic device 300 may identify a compressed bitstream. At 1020, the electronic device 300 may determine (or identify) one or more vertex neighbors for vertices in the compressed video bitstream based on a signal-sent limit on the number of one or more vertex neighbors. This may include the processor of the electronic device identifying the last vertex neighbor in the received neighbor sequence associated with the vertex, and at least one additional vertex neighbor in the received neighbor sequence corresponding to one less than the number of one or more vertex neighbors. In one or more embodiments of the present disclosure, the determined one or more vertex neighbors are associated with an intra-grid frame, and the processor is further configured to reuse the determined one or more vertex neighbors for an inter-grid frame. In one or more embodiments of the present disclosure, the signal-sent limit on the number of one or more vertex neighbors is included in at least one of a sequence header, a frame header, a sub-grid header, a slice header, a sub-frame header, or a parallel block header of the compressed video bitstream.
[0168] At step 1030, the electronic device 300 may identify a VMV predictor to be used for a vertex based on a signal-sent vertex motion vector (VMV) identifier in the compressed video bitstream. As described in the present disclosure, a VMV predictor may be an associated predicted value. Also as described in the present disclosure, the associated predicted value is a combination of one or more vertex neighbors (e.g., an average value, a weighted average value, a median value, a maximum value, a minimum value, etc.). In one or more embodiments of the present disclosure, the decoder calculates a VMV predictor based on a reconstructed VMV, where the reconstructed VMV is the predicted VMV plus the difference between the received VMV for the vertex and the associated predicted value of the VMV predictor.
[0169] At step 1040, the electronic device 300 may reconstruct a grid frame based on the determined one or more vertex neighbors and the identified VMV predictor. In one or more embodiments of the present disclosure, the electronic device 300 may also obtain a duplicate vertex data structure storing information related to duplicate vertices of one or more grid frames, and determine that an inter-grid frame inherits the duplicate vertex data structure based on a flag signaled in the compressed video bitstream.
[0170] In an embodiment of the present disclosure, the electronic device 300 may output decoded content, such as including the reconstructed grid frame. The reconstructed grid frame corresponds to the original grid frame used during encoding, as described in the present disclosure. The output decoded content may be sent to an external device or a memory on the electronic device 300.
[0171] Although Figure 10 illustrates an example of a decoding method 1000 for improved vertex motion vector predictor coding, it can be applied to Figure 10Make various changes. For example, although shown as a series of steps, each step in Figure 10 can overlap, occur in parallel, or occur any number of times.
[0172] In an embodiment of the present disclosure, an electronic device may include a memory and at least one processor coupled to the memory. In an embodiment of the present disclosure, the electronic device may include a communication interface configured to receive a compressed video bitstream, and the at least one processor may be operably coupled to the communication interface. In an embodiment of the present disclosure, the at least one processor may be configured to identify (or obtain, receive) the compressed video bitstream. In an embodiment of the present disclosure, the at least one processor may be configured to determine one or more vertex neighbors for vertices in the compressed video bitstream based on a signal-sent limit on the number of one or more vertex neighbors. In an embodiment of the present disclosure, the at least one processor may be configured to identify, from a plurality of VMV predictors, a VMV predictor to be used for a vertex based on a signal-sent vertex motion vector (VMV) identifier in the compressed video bitstream. In an embodiment of the present disclosure, the at least one processor may be configured to reconstruct a mesh frame based on the determined one or more vertex neighbors and the identified VMV predictor.
[0173] In an embodiment of the present disclosure, the at least one processor may be configured to identify the last vertex neighbor in a (received) neighbor sequence associated with a vertex, and at least one additional vertex neighbor in the neighbor sequence corresponding to one less than the number of the one or more vertex neighbors.
[0174] In an embodiment of the present disclosure, the determined one or more vertex neighbors may be associated with an intra-frame mesh frame. In an embodiment of the present disclosure, the at least one processor may be configured to reuse the determined one or more vertex neighbors for an inter-frame mesh frame.
[0175] In an embodiment of the present disclosure, the at least one processor may be configured to obtain a duplicate vertex data structure storing information related to duplicate vertices of one or more mesh frames. In an embodiment of the present disclosure, the at least one processor may be configured to determine that an inter-frame mesh frame inherits the duplicate vertex data structure based on a flag signaled in the compressed video bitstream.
[0176] In an embodiment of the present disclosure, the at least one processor may be configured to identify (or receive, obtain) a difference between a VMV for a vertex and an associated predicted value of a VMV predictor.
[0177] In an embodiment of the present disclosure, the associated predicted value may be a combination of one or more vertex neighbors.
[0178] In an embodiment of the present disclosure, a signal-sent limit on the number of the one or more vertex neighbors may be included in at least one of a sequence header, a frame header, a sub-grid header, a slice header, a sub-frame header, or a parallel block header of the compressed video bitstream.
[0179] In an embodiment of the present disclosure, a method may be performed by an electronic device. In an embodiment of the present disclosure, the method may include identifying (or, obtaining, receiving) a compressed video bitstream. In an embodiment of the present disclosure, the method may include determining, for a vertex in the compressed video bitstream, the one or more vertex neighbors based on a signal-sent limit on the number of one or more vertex neighbors. In an embodiment of the present disclosure, the method may include identifying, based on a vertex motion vector (VMV) identifier signaled in the compressed video bitstream, a VMV predictor to be used for the vertex from a plurality of VMV predictors. In an embodiment of the present disclosure, the method may include reconstructing a mesh frame based on the determined one or more vertex neighbors and the identified VMV predictor.
[0180] In an embodiment of the present disclosure, determining the one or more vertex neighbors may include identifying a last vertex neighbor in a (received) neighbor sequence associated with the vertex, and at least one additional vertex neighbor in the neighbor sequence corresponding to a number of the one or more vertex neighbors minus one.
[0181] In an embodiment of the present disclosure, the determined one or more vertex neighbors may be associated with an intra mesh frame. In an embodiment of the present disclosure, the method may include reusing the determined one or more vertex neighbors for an inter mesh frame.
[0182] In an embodiment of the present disclosure, the method may include obtaining a duplicate vertex data structure storing information related to duplicate vertices of one or more mesh frames. In an embodiment of the present disclosure, the method may include determining that an inter mesh frame inherits the duplicate vertex data structure based on a flag signaled in the compressed video bitstream.
[0183] In an embodiment of the present disclosure, the method may include identifying (or receiving, obtaining) a difference between a VMV for a vertex and an associated predicted value of a VMV predictor.
[0184] In an embodiment of the present disclosure, the associated predicted value may be a combination of one or more vertex neighbors.
[0185] In an embodiment of the present disclosure, a signal-sent limit on the number of one or more vertex neighbors may be included in at least one of a sequence header, a frame header, a sub-grid header, a slice header, a sub-frame header, or a parallel block header of the compressed video bitstream.
[0186] In an embodiment of the present disclosure, an electronic device may include a memory and at least one processor coupled to the memory. In an embodiment of the present disclosure, the electronic device may include a communication interface, and the at least one processor may be operably coupled to the communication interface. In an embodiment of the present disclosure, the at least one processor may be configured to identify one or more vertex neighbors for a vertex of a grid frame based on a set limit on the number of one or more vertex neighbors. In an embodiment of the present disclosure, the at least one processor may be configured to determine a plurality of vertex motion vector (VMV) predictors for the vertex based on the identified one or more vertex neighbors. In an embodiment of the present disclosure, the at least one processor may be configured to map each of the plurality of VMV predictors to one of a plurality of VMV identifiers. In an embodiment of the present disclosure, the at least one processor may be configured to encode a compressed video bitstream that signals a set limit on the number of one or more vertex neighbors and signals one of the plurality of VMV identifiers, the VMV identifier indicating the VMV predictor among the plurality of VMV predictors that will be used for the vertex.
[0187] In an embodiment of the present disclosure, the at least one processor may be configured to identify a last vertex neighbor in a neighbor sequence associated with a vertex, and at least one additional vertex neighbor in the neighbor sequence corresponding to a number of one or more vertex neighbors minus one.
[0188] In an embodiment of the present disclosure, the identified one or more vertex neighbors may be associated with an intra-grid frame. In an embodiment of the present disclosure, the at least one processor may be configured to re-use the identified one or more vertex neighbors for an inter-grid frame.
[0189] In an embodiment of the present disclosure, the at least one processor may be configured to generate a duplicate vertex data structure storing information related to duplicate vertices of one or more grid frames. In an embodiment of the present disclosure, the at least one processor may be configured to set a flag in the compressed video bitstream that signals that an inter-grid frame inherits the duplicate vertex data structure.
[0190] In an embodiment of the present disclosure, the at least one processor may be configured to initiate transmission of a difference between a VMV for a vertex and an associated predicted value of a VMV predictor. In an embodiment of the present disclosure, the associated predicted value may be a combination of one or more vertex neighbors.
[0191] In an embodiment of the present disclosure, the at least one processor may be configured to encode a set limit on the number of one or more vertex neighbors in at least one of a sequence header, a frame header, a sub-grid header, a slice header, a sub-frame header, or a parallel block header of the compressed video bitstream.
[0192] In an embodiment of the present disclosure, a method may be performed by an electronic device. In an embodiment of the present disclosure, the method may include: for vertices of a grid frame, identifying one or more vertex neighbors based on a set limit on the number of one or more vertex neighbors. In an embodiment of the present disclosure, the method may include determining a plurality of vertex motion vector (VMV) predictors for a vertex based on the identified one or more vertex neighbors. In an embodiment of the present disclosure, the method may include mapping each of the plurality of VMV predictors to one of a plurality of VMV identifiers. In an embodiment of the present disclosure, the method may include encoding a compressed video bitstream that signals the set limit on the number of one or more vertex neighbors and signals one of the plurality of VMV identifiers, the VMV identifier indicating the VMV predictor of the plurality of VMV predictors that will be used for the vertex.
[0193] In an embodiment of the present disclosure, determining one or more vertex neighbors may include identifying a last vertex neighbor in a neighbor sequence associated with a vertex and at least one additional vertex neighbor in the neighbor sequence corresponding to one less than the number of the one or more vertex neighbors.
[0194] In an embodiment of the present disclosure, the identified one or more vertex neighbors may be associated with an intra-grid frame. In an embodiment of the present disclosure, the method may include reusing the identified one or more vertex neighbors for an inter-grid frame.
[0195] In an embodiment of the present disclosure, the method may include generating a duplicate vertex data structure that stores information related to duplicate vertices of one or more grid frames. In an embodiment of the present disclosure, the method may include setting a flag in the compressed video bitstream that signals that an inter-grid frame inherits the duplicate vertex data structure.
[0196] In an embodiment of the present disclosure, the method may include initiating transmission of a difference between a VMV for a vertex and an associated predicted value of a VMV predictor. In an embodiment of the present disclosure, the associated predicted value may be a combination of one or more vertex neighbors.
[0197] In an embodiment of the present disclosure, the method may include encoding the set limit on the number of one or more vertex neighbors in at least one of a sequence header, a frame header, a sub-grid header, a slice header, a sub-frame header, or a parallel block header of the compressed video bitstream.
[0198] Although the present disclosure has been described using exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. The present disclosure is intended to cover such changes and modifications that fall within the scope of the appended claims. None of the descriptions in this application should be construed as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the patented subject matter is defined by the claims.
Claims
1. An electronic device (300), comprising: a memory (360); and at least one processor (340) coupled to the memory (360), the at least one processor (340) being configured to: identify a compressed video bitstream; determine, for vertices in the compressed video bitstream, one or more vertex neighbors based on a signal-sent limit on the number of one or more vertex neighbors; identify, based on a vertex motion vector VMV identifier signaled in the compressed video bitstream, a VMV predictor factor to be used for the vertex from a plurality of VMV predictor factors; and reconstruct a mesh frame based on the determined one or more vertex neighbors and the identified VMV predictor factor.
2. The electronic device (300) according to claim 1, wherein the at least one processor (340) is further configured to identify a last vertex neighbor in a neighbor sequence associated with the vertex, and at least one additional vertex neighbor corresponding to a decrement by one in the number of the one or more vertex neighbors in the neighbor sequence.
3. The electronic device (300) according to claim 1 or 2, wherein the determined one or more vertex neighbors are associated with an intra mesh frame, and the at least one processor (340) is further configured to reuse the determined one or more vertex neighbors for an inter mesh frame.
4. The electronic device (300) according to any one of claims 1 to 3, wherein the at least one processor (340) is further configured to: obtain a duplicate vertex data structure storing information related to duplicate vertices of one or more mesh frames; and determine, based on a flag signaled in the compressed video bitstream, that an inter mesh frame inherits the duplicate vertex data structure.
5. The electronic device (300) according to any one of claims 1 to 4, wherein the at least one processor (340) is further configured to identify a difference between a VMV for the vertex and an associated predicted value of the VMV predictor factor.
6. The electronic device (300) according to claim 5, wherein the associated predicted value is a combination of the one or more vertex neighbors.
7. The electronic device (300) according to any one of claims 1 to 6, wherein the signal-sent limit on the number of the one or more vertex neighbors is included in at least one of a sequence header, a frame header, a sub-mesh header, a slice header, a sub-frame header, or a parallel block header of the compressed video bitstream.
8. A method (1000) performed by an electronic device (300), the method (1000) comprising: identifying (1010) a compressed video bitstream; determining (1020), for vertices in the compressed video bitstream, one or more vertex neighbors based on a signal-sent limit on the number of one or more vertex neighbors; identifying (1030), based on a vertex motion vector VMV identifier signaled in the compressed video bitstream, a VMV predictor factor to be used for the vertex from a plurality of VMV predictor factors; and Reconstruct (1040) a mesh frame based on the determined one or more vertex neighbors and the identified VMV predictors.
9. An electronic device (300), comprising: a memory (360); and at least one processor (340) coupled to the memory (360), the at least one processor (340) being configured to: Identify, for a vertex of a mesh frame, the one or more vertex neighbors based on a set limit on the number of one or more vertex neighbors; Determine, based on the identified one or more vertex neighbors, a plurality of vertex motion vector VMV predictors for the vertex; Map each VMV predictor of the plurality of VMV predictors to one VMV identifier of a plurality of VMV identifiers; and Encode a compressed video bitstream that signals the set limit on the number of the one or more vertex neighbors and signals one VMV identifier of the plurality of VMV identifiers, the VMV identifier indicating the VMV predictor of the plurality of VMV predictors that will be used for the vertex.
10. The electronic device (300) according to claim 9, wherein the at least one processor (340) is further configured to identify the last vertex neighbor in a neighbor sequence associated with the vertex, and at least one additional vertex neighbor in the neighbor sequence corresponding to the number of the one or more vertex neighbors minus one.
11. The electronic device (300) according to claim 9 or 10, wherein the identified one or more vertex neighbors are associated with an intra mesh frame, and the at least one processor (340) is further configured to reuse the identified one or more vertex neighbors for an inter mesh frame.
12. The electronic device (300) according to any one of claims 9 to 11, wherein the at least one processor (340) is further configured to: Generate a duplicate vertex data structure that stores information related to duplicate vertices of one or more mesh frames; and Set a flag in the compressed video bitstream that signals that an inter mesh frame inherits the duplicate vertex data structure.
13. The electronic device (300) according to any one of claims 9 to 12, wherein the at least one processor (340) is further configured to initiate transmission of a difference between a VMV for the vertex and an associated predicted value of the VMV predictor, and wherein the associated predicted value is a combination of the one or more vertex neighbors.
14. The electronic device (300) according to any one of claims 9 to 13, wherein the at least one processor (340) is further configured to encode the set limit on the number of the one or more vertex neighbors in at least one of a sequence header, a frame header, a sub mesh header, a slice header, a sub frame header, or a parallel block header of the compressed video bitstream.
15. A method (900) performed by an electronic device (300), the method (900) comprising: For vertices of a mesh frame, identify (910) the one or more vertex neighbors based on a set limit on the number of one or more vertex neighbors; Determine (920) a plurality of vertex motion vector (VMV) predictors for the vertex based on the identified one or more vertex neighbors; Map (930) each VMV predictor of the plurality of VMV predictors to one VMV identifier of a plurality of VMV identifiers; and Encode (940) a compressed video bitstream that signals the set limit on the number of the one or more vertex neighbors and signals one VMV identifier of the plurality of VMV identifiers, the VMV identifier indicating the VMV predictor of the plurality of VMV predictors that will be used for the vertex.