Augmenting 3D point clouds with multiple measurements

By refining and compressing point clouds using quantizer shifts and correlated measurements, the method addresses the challenges of processing unorganized point clouds, enhancing accuracy and reducing bitrate consumption for dynamic point clouds.

JP7804580B2Active Publication Date: 2026-01-22INTERDIGITAL VC HOLDINGS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022545937
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-06
Filing Date
2021-02-05
Publication Date
2026-01-22
Estimated Expiration
2041-02-05

AI Technical Summary

Technical Problem

Existing point cloud processing technologies face challenges in efficiently handling unorganized point clouds due to their random distribution in 3D space, making it difficult to apply traditional grid-based algorithms, and there is a need for effective compression methods to deliver dynamic point clouds to end users with reasonable bitrate consumption while maintaining quality.

Method used

A method and apparatus for point cloud decoding and encoding that involve obtaining and refining point clouds using quantizer shifts across multiple observations, shifting quantization bins to enhance accuracy, and encoding refined data in a bitstream, utilizing correlated measurements to improve quality and accuracy.

Benefits of technology

The method enhances point cloud quality by refining and compressing dynamic point clouds, improving accuracy and reducing bitrate consumption, making immersive world delivery networks more practical.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007804580000068
    Figure 0007804580000068
  • Figure 0007804580000069
    Figure 0007804580000069
  • Figure 0007804580000070
    Figure 0007804580000070
Patent Text Reader

Abstract

Systems and methods are described for refining first point cloud data using at least second point cloud data and one or more sets of quantizer shifts. An exemplary point cloud decoding method includes obtaining data representing at least a first point cloud and a second point cloud, obtaining information identifying at least a first set of quantizer shifts associated with the first point cloud, and obtaining refined point cloud data based on at least the first point cloud, the first set of quantizer shifts, and the second point cloud. Obtaining the refined point cloud data may include performing subtraction based on the at least first set of quantizer shifts. Corresponding encoding systems and methods are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application is a nonprovisional application under 35 U.S.C. §119(e) of and claims the benefit of U.S. Provisional Patent Application No. 62 / 970,956, filed February 6, 2020, entitled "3D Point Cloud Enhancement with Multiple Measurements," the contents of which are incorporated herein by reference in their entirety.

[0002] FIELD OF THE INVENTION The present disclosure relates to the field of point cloud processing and compression, which relates to the analysis, interpolation, and representation of point cloud signals. [Background technology]

[0003] As discrete representations of continuous surfaces in 3D space, point clouds can be classified into categories including organized point clouds, such as those collected by 3D sensors like cameras or 3D laser scanners and arranged on a grid, and unorganized point clouds, such as those scanned from multiple viewpoints and then fused together due to their complex structure, causing the loss of index ordering. Organized point clouds are easier to process because the underlying grid implies natural spatial connectivity that can reflect the perceived order. In contrast, processing unorganized point clouds can be more challenging. This is due to the fact that unorganized point clouds differ from 1D audio data or 2D images, which are associated with a regular grid. Conversely, unorganized point clouds are often scattered randomly in 3D space, making it difficult for traditional grid-based algorithms to handle 3D point clouds. For example, the convolution operator, while well-defined on a regular grid, cannot be directly applied to 3D point clouds.

[0004] Furthermore, point clouds can represent continuous representations of the same scene containing multiple moving objects, which are called dynamic point clouds, as opposed to static point clouds captured from a static scene or static objects. Dynamic point clouds can be organized into frames, with different frames captured at different times.

[0005] Point clouds can be used for a variety of purposes, such as representing cultural heritage objects, where an object such as a statue or building is scanned in 3D to share the spatial configuration of the object without sending or visiting the object. Also, if the object is subject to destruction, for example, an earthquake can destroy a temple, a point cloud is a way to ensure knowledge of the object is preserved. Such point clouds are typically static, colored, and often contain a huge number of points.

[0006] Another use case is in topography and cartography, where 3D representations are used so that maps are not limited to flat surfaces but can include relief. Google Maps is currently a good example of a 3D map, but it generally uses meshes instead of point clouds. Nevertheless, point clouds can be the preferred data format for 3D, and such point clouds are usually static, colored, and contain a large number of points.

[0007] The automotive industry and autonomous vehicle technology are also areas where point clouds can be used. Autonomous vehicles need to be able to "explore" their environment and make good driving decisions based on their immediate surroundings. Typical sensors like LIDAR generate dynamic point clouds that are used by decision engines. These point clouds are not intended for human viewing; they are usually small, not necessarily color-coded, and dynamic with a high capture frequency. These point clouds may have other attributes, such as reflectivity provided by LIDAR, that can be a good indicator of the material of the sensed object and help make that decision.

[0008] Virtual reality and immersive worlds are gaining increasing interest. Such technologies attempt to immerse the viewer in the environment all around them, unlike standard TV, where the viewer can only see the virtual world in front of them. There are several levels of immersion depending on the viewer's degrees of freedom within the environment. Point clouds are a good format candidate for delivering virtual reality (VR) worlds. They can be static or dynamic and are typically of average size, e.g., at most a few million points at a time.

[0009] It is desirable to be able to deliver dynamic point clouds to end users with reasonable bitrate consumption while maintaining an acceptable (or preferably very good) quality of experience. Efficient compression of these dynamic point clouds can aid in making immersive world delivery networks practical. Summary of the Invention

[0010] A point cloud decoding method according to some embodiments includes obtaining data representing at least a first point cloud and a second point cloud, obtaining information identifying at least a first set of quantizer shifts associated with the first point cloud, and obtaining refined point cloud data based on at least the first point cloud, the first set of quantizer shifts, and the second point cloud.

[0011] A point cloud decoder apparatus according to some embodiments includes a processor configured to at least obtain data representing at least a first point cloud and a second point cloud; obtain information identifying at least a first set of quantizer shifts associated with the first point cloud; and obtain refined point cloud data based on at least the first point cloud, the first set of quantizer shifts, and the second point cloud.

[0012] In some embodiments, obtaining the refined point cloud data includes performing a subtraction based on at least a first set of quantizer shifts.

[0013] In some embodiments, obtaining refined point cloud data further includes obtaining information identifying a second set of quantizer shifts associated with the second point cloud, the second set of quantizer shifts being different from the first set of quantizer shifts, and obtaining the refined point cloud data is further based on the second set of quantizer shifts. In some such embodiments, obtaining the refined point cloud data includes performing a subtraction based on at least the first set of quantizer shifts and the second set of quantizer shifts.

[0014] In some embodiments, the first point cloud represents a left view of the scene and the second point cloud represents a right view of the scene.

[0015] In some embodiments, the first and second clouds of points are frames associated with different times.

[0016] In some embodiments, the first set of quantizer shifts includes at least a first shift associated with a first point in the first cloud of points and a different second shift associated with a second point in the first cloud of points.

[0017] In some embodiments, the data representing the first point cloud includes a first parameter (y1) having a first quantization range, the data representing the second point cloud includes a second parameter (y2) having a second quantization range, and obtaining the refined point cloud data includes obtaining a third parameter (y2) within the quantization ranges of both the first parameter (y1) and the second parameter (y2).

number

[0018] In some embodiments, the first cloud of points (y l ) and the second point group (y r ) based on the refined point cloud data (x l) is obtained by substituting the refined first point cloud (x ) so as to substantially maximize a product of coefficients including one or more of the following coefficients: l ), the coefficients of which are: The refined first point cloud (x l ) is given by the first point group (y l ) conditional probability Pr(y l |x l )and, The second refined point cloud (x r ) estimate (x l ) is given by the second point group (y r ) conditional probability Pr(y r )|g(x l )) where the estimate (x l ) is the first refined point cloud (x l ) and the conditional probability based on First refined point cloud data (x l ) prior probability Pr(x l )and, The second refined point cloud data (x r ) estimate g(x l ) prior probability Pr(g(x l )) and.

[0019] A point cloud encoding method according to some embodiments includes obtaining data representing at least a first point cloud and a second point cloud; quantizing the first point cloud and the second point cloud, where quantizing the point cloud data includes adding at least a first set of quantizer shifts to the first point cloud; and encoding the quantized first and second point clouds in a bitstream.

[0020] A point cloud encoding apparatus according to some embodiments includes a processor configured to perform at least the following: obtaining data representing at least a first point cloud and a second point cloud; quantizing the first point cloud and the second point cloud, where quantizing the point cloud data includes adding at least a first set of quantizer shifts to the first point cloud; and encoding the quantized first and second point clouds in a bitstream.

[0021] Some embodiments further include encoding, in the bitstream, information indicative of the first set of quantizer shifts.

[0022] In some embodiments, quantizing the first cloud of points and the second cloud of points further includes adding a second set of quantizer shifts to the second cloud of points, and further includes encoding information indicative of the second set of quantizer shifts in the bitstream.

[0023] In some embodiments, the first point cloud represents a left view of the scene and the second point cloud represents a right view of the scene.

[0024] In an exemplary embodiment, the method includes: l ) and at least the initial second point cloud data (y r ), and receiving the refined first point cloud data (x l ), the coefficients of which are calculated based on at least the refined first point cloud data (x l ) is given as the initial first point cloud data (y l ) conditional probability Pr(y l |x l ) and the refined second point cloud data (x r ) estimate g(x l ) is given as the initial second point cloud data (y r ) conditional probability Pr(y r |g(x l )) where the estimate g(xl ) is the refined first point group (x l ) and the refined first point cloud data (x l ) prior probability Pr(x l ) and the refined second point cloud data (x r ) estimate g(x l ) prior probability Pr(g(x l )) and.

[0025] Some embodiments may further include: l ) and then calculate the quantizer shift δ from the refined first point cloud data. l The method further includes subtracting

[0026] In some embodiments, the refined first point cloud data (x l ) is performed repeatedly.

[0027] In some embodiments, the refined first point cloud data (x l ) is selected so as to substantially minimize the sum of the negative logarithms of the coefficients of the refined first point cloud data (x l ) to substantially minimize the sum of the negative logarithms of the coefficients. l ) can be performed using gradient descent.

[0028] In some embodiments, the refined first point cloud data (x l ) is given as the initial first point cloud data (y l ) conditional probability Pr(y l )|(x l )) is the initial first point cloud data (y l ) and the refined first point cloud data (x l ) is expressed as a linear function of the difference between

[0029] In some embodiments, the refined second point cloud data (xr ) estimate g(x l ) is the refined second point cloud data (x r ) is a linear function of

[0030] In some embodiments, the refined second point cloud data (x r ) estimate g(x l ) is given as the initial second point cloud data (y r ) conditional probability Pr(y r |g(x l )) is the initial second point cloud data (y r ) and the refined second point cloud data (x r ) estimate g(x l ) is expressed as a linear function of the difference between

[0031] In some embodiments, the refined first point cloud data (x l ) prior probability Pr(x l )teeth,

number

[0032] In some embodiments, the refined second point cloud data (x r ) estimate g(x l ) prior probability Pr(g(x l )) is exp(-g(x l ) T L r g(x l ) / σ 2 ), where L r is the graph Laplacian matrix.

[0033] Some embodiments further include receiving first point cloud data and at least second point cloud data, and refining the first point cloud data using the at least second point cloud data. Some such embodiments further include, after refining the first point cloud data, subtracting a quantizer shift from the refined first point cloud data. In some embodiments, the quantizer shift is a predetermined quantizer shift.

[0034] In some embodiments, the encoding method includes receiving initial first point cloud data and at least initial second point cloud data, processing the initial first and second point cloud data by a method that includes adding a first set of quantizer shifts δ1 to the first point cloud data, and encoding the processed first and second point cloud data.

[0035] In some embodiments, processing the initial first and second point cloud data further includes adding a second set of quantizer shifts δ to the second point cloud data, where the first set of quantizer shifts δ may be different from the second set of quantizer shifts δ.

[0036] In some embodiments, the encoding method further includes providing information identifying at least one of: (i) a first set of quantizer shifts δ1; or (ii) a second set of quantizer shifts δ2 together with the encoded first and second point cloud data.

[0037] In some embodiments, a decoding method includes receiving encoded first point cloud data and at least encoded second point cloud data, receiving information identifying at least a first set of quantizer shifts δ1, decoding the first point cloud data and the at least second point cloud data, refining the decoded first point cloud data using the decoded second point cloud data, and subtracting the first set of quantizer shifts δ1 from the first point cloud data.

[0038] In some embodiments, any of the systems or methods described herein can be used to refine depth data in formats other than point cloud formats. For example, the systems and methods described herein can be used to refine depth data (e.g., depth data x) in RGB-D or other depth map formats. l , x r ) can be refined.

[0039] Some embodiments of the apparatus include a processor configured to perform at least one or more of the methods described herein.

[0040] Some embodiments of the apparatus include a processor and a computer-readable medium storing instructions operable to perform at least one or more of the methods described herein. The computer-readable medium may be a non-transitory computer-readable medium.

[0041] Some embodiments include a computer-readable medium storing point cloud data encoded using one or more of the methods described herein. The medium may be a non-transitory computer-readable medium. [Brief explanation of the drawings]

[0042] [Figure 1A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed embodiments may be implemented. [Figure 1B] 1B is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used within the communication system illustrated in FIG. 1A, according to one embodiment. [Figure 2A] FIG. 1 is a schematic, illustrative diagram providing an example of refining point cloud data from two independent measurements of the same point. [Figure 2B] FIG. 1 is a schematic illustration showing that when two camera views form a stereo setup, depth refinement cannot be achieved if the quantization bins are well aligned. [Figure 3]FIG. 1 is a schematic, illustrative diagram showing that depth refinement can be achieved using shifting quantization bins when two camera views form a stereo setup, according to some embodiments; [Figure 4] FIG. 10 is a schematic, illustrative diagram showing that refinement can be achieved for highly correlated measurements within the same observation when a shift is introduced into the quantization bins of the individual measurements, according to some embodiments. [Figure 5] 1 illustrates an example of a camera system used in some embodiments. [Figure 6] An example using a linear approximation of a Gaussian distribution is shown. [Figure 7] FIG. 1 is a block diagram of an exemplary encoder for point cloud compression used in some embodiments. [Figure 8] FIG. 1 is a block diagram of a decoder using joint refinement according to some embodiments. [Figure 9] FIG. 10 is a block diagram of a decoder when no refinement is applied, according to some embodiments. [Figure 10] FIG. 10 is a block diagram illustrating components of a refinement module of a decoder, according to some embodiments. [Figure 11] FIG. 2 is a block diagram of a video encoder used in some embodiments. [Figure 12] FIG. 2 is a block diagram of a video decoder used in some embodiments. [Figure 13] FIG. 1 is a block diagram of an example system in which various aspects and embodiments may be implemented. [Figure 14] FIG. 1 is a flow diagram of an exemplary point cloud encoding method, according to some embodiments. [Figure 15] FIG. 1 is a flow diagram of an exemplary point cloud decoding method, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0043] Exemplary Network for Implementing the Embodiments 1A is a diagram illustrating an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.

[0044] 1A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104, CN 106, public switched telephone network (PSTN) 108, Internet 110, and other networks 112, although it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d, any of which may be referred to as a “station” and / or “STA,” may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain contexts), consumer electronics devices, devices operating in commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as a UE.

[0045] The communications system 100 may also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNodeB, a Home Node B, a Home eNodeB, a gNB, an NR Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each shown as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0046] The base station 114a may be part of the RAN 104, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide wireless service coverage for a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one embodiment, the base station 114a may include three transceivers, i.e., one for each sector of the cell. In one embodiment, the base station 114a may employ multiple-input multiple output (MIMO) technology and may utilize multiple transceivers per sector of the cell, for example, using beamforming to transmit and / or receive signals in desired spatial directions.

[0047] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over the air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0048] More specifically, as mentioned above, the communications system 100 may be a multiple access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base station 114 a and the WTRUs 102 a, 102 b, 102 c in the RAN 104 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interfaces 115 / 116 / 117 using wideband CDMA (WCDMA). WCDMA may include communications protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​Uplink Packet Access (HSUPA).

[0049] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-APro).

[0050] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement a radio technology such as New Radio (NR) radio access, which may establish the air interface 116 using NR.

[0051] In one embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0052] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as IEEE 802.11 (Wireless Fidelity, WiFi), IEEE 802.16 (Worldwide Interoperability for Microwave Access, WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), or the like.

[0053] 1A may be, for example, a wireless router, a Home Node B, a Home eNode B, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a local area such as a location such as a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may establish a picocell or a femtocell using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-APro, NR, etc.). As shown in FIG. 1A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 through the CN 106.

[0054] The RAN 104 may communicate with the CN 106, which may be any type of network configured to provide voice, data, application, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1A , it will be understood that the RAN 104 and / or CN 106 may communicate directly or indirectly with other RANs that use the same RAT as the RAN 104 or a different RAT. For example, in addition to being connected to the RAN 104, which may utilize NR radio technology, the CN 106 may also communicate with another RAN (not shown) using GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0055] The CN 106 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a public switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols such as the transmission control protocol (TCP), the user datagram protocol (UDP), and / or the internet protocol (IP) in the TCP / IP Internet protocol suite. The networks 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs that may use the same RAT as the RAN 104 or a different RAT.

[0056] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links.) For example, the WTRU 102c shown in FIG. 1A may be configured to communicate with a base station 114a that can use cellular-based wireless technology and a base station 114b that can use IEEE 802 wireless technology.

[0057] 1B is a system diagram illustrating an example of an exemplary WTRU 102. As shown in FIG. 1B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the foregoing elements while remaining consistent with an embodiment.

[0058] The processor 118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors in conjunction with a DSP core, a controller, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, which may be coupled to the transmit / receive element 122. While FIG. 1B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0059] The transmit / receive element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In one embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0060] 1B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may use MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0061] The transceiver 120 may be configured to modulate signals transmitted by the transmit / receive element 122 and demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.

[0062] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Furthermore, the processor 118 may access information from and store data in any type of suitable memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in memory that is not physically located on the WTRU 102, such as on a server or home computer (not shown).

[0063] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control the power to other components in the WTRU 102. The power source 134 may be any suitable device for providing power to the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0064] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals being received from two or more nearby base stations. It will be understood that the WTRU 102 may obtain location information by way of any suitable location-determination method while remaining consistent with an embodiment.

[0065] The processor 118 may further be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulated (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripheral device 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0066] The WTRU 102 may include a full-duplex radio where transmission and reception of some or all of the signals (e.g., associated with a particular subframe for both the UL (e.g., for transmission) and downlink (e.g., for reception)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and or substantially eliminate self-interference via hardware (e.g., chokes) or processor-based signal processing (e.g., via a separate processor (not shown) or processor 118). In one embodiment, the WTRU 102 may include a half-duplex radio for transmission and reception of either some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or downlink (e.g., for reception)).

[0067] Although the WTRU is illustrated in FIGS. 1A-1B as a wireless terminal device, it is contemplated that in certain representative embodiments, such a terminal device may use a wired communication interface (e.g., temporarily or permanently) with the communication network.

[0068] In a representative embodiment, the other network 112 may be a WLAN.

[0069] 1A-1B and the corresponding description, one or more, or all, of the functions described herein with respect to one or more of the WTRUs 102a-d, base stations 114a-b, and / or any other device(s) described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.

[0070] The emulation devices may be designed to implement one or more tests of other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in the communication network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation devices may be directly coupled to another device for testing purposes and / or may perform testing using terrestrial wireless communication.

[0071] One or more emulation devices may perform one or more functions, inclusive, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices may be utilized in test scenarios in a test lab and / or in an undeployed (e.g., test) wired and / or wireless communication network to implement testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, e.g., one or more antennas) may be used by the emulation devices to transmit and / or receive data. Detailed Description

[0072] 3D point cloud data may represent samples on the surface of an object or scene. When collecting / processing dynamic point clouds, data samples of the same object often appear in different observations, which are highly correlated due to 2D to 3D projection. In other words, the overlapping areas between different observations have more measurements than the same 3D scene, e.g., FIG. 2B . Even in the case of measurements that appear within the same observation, given information that the measurements are highly correlated, e.g., FIG. 4 , the measurements may be considered as different measurements of the same object. Exemplary embodiments operate to exploit this correlation across different observations of the point cloud to improve point cloud quality.

[0073] Some embodiments jointly perform denoising and inverse quantisation on 3D point cloud data via correlated measurements from multiple observations. Not only do we quantise per measurement, but noise-corrupted observations can also be considered as coded quantisation bins for reconstruction performance. Some embodiments intentionally shift the quantisation bins applied to the noise-corrupted point cloud data. By doing so, we can further narrow the range of likely reliable measurements, taking into account the correlation of multiple observations, compared to when the quantisation bins are strictly aligned.

[0074] Exemplary embodiments are not limited to point cloud data acquired from any particular type of depth sensor. A variety of different depth sensors ranging from low-end sensors such as time-of-flight (ToF) sensors and structured light sensors to high-end sensors such as LiDAR, laser scanners, etc., or synthetic data using computer graphics are data sources that may be used in various embodiments.

[0075] In some embodiments, the systems and methods disclosed herein are used to enhance 3D point cloud quality during acquisition using a depth sensor. In other embodiments, the systems and methods herein are applied to point cloud compression (PCC). In some existing PCC approaches, the current frame to be processed may use one or more reference frames for compression. In some embodiments, the reference frames also use information from the current frame to enhance its own reconstruction, so that point cloud frames may reference each other for reconstruction. On the encoder side, there is the ability to intentionally shift the quantization bins applied to the point cloud data, while on the decoder side, multiple correlated observations may be combined to enhance the point cloud data, resulting in a more accurate reconstruction. There are various PCC-based use cases for the systems and methods as described herein. Some embodiments operate to enhance the accuracy of point cloud data encoded using either inter-frame or intra-frame coding. Some embodiments apply when any particular coding structure, e.g., octree, is used in the encoder. When transform coding is adapted to encode point cloud video, some embodiments operate directly in the transform domain to provide greater accuracy.

[0076] Depth sensors are currently lightweight and common, but the measurements they acquire are often coarse in accuracy and can be corrupted by noise. However, when capturing 3D data—especially a series of 3D data of a particular scene—it's possible that the same (or highly correlated) scene (or object) is captured multiple times. In other words, it's highly likely that different observations have been made of the same object. Consequently, exemplary embodiments use different observations of the same object in depth measurements of point cloud data to jointly enhance the accuracy of the observations, while these observations can be corrupted by noise and suffer from quantization errors. In some embodiments, such techniques can be applied for point cloud compression. Overview of Exemplary Embodiments

[0077] Exemplary embodiments jointly enhance the quality of measurements using different observations of point cloud data. FIG. 2A illustrates measuring the depth of point P from two different observations, namely, Observation 1 and Observation 2. The observations here may be views measured by a depth sensor during acquisition, or frames within a point cloud sequence in the case of point cloud compression. In this example, to simplify the explanation, the measurements are assumed to be noise-free. The measurements are quantized by two quantizers, namely, Q1 and Q2, to be represented using a specific number of bits. The quantization produces two different uncertainty ranges. Given that the camera pose is known (or can be estimated by the acquired point cloud data or otherwise), the intersection of both of the two uncertainty ranges can be considered, resulting in a much narrower uncertainty range and even higher accuracy. The exemplary embodiments are not limited to only two observations; similar reasoning can be applied when three or more observations or measurements are made of the same object.

[0078] When two cameras are used to measure a scene, the cameras may be configured as a stereo device, as illustrated in FIG. 2B. In this case, if the quantizer aligns the quantization bins, obtaining a common portion of the uncertainty range does not result in higher accuracy. To address this issue, some embodiments provide the option of shifting the quantization bins, as shown in FIG. 3, which can narrow the uncertainty range.

[0079] Some embodiments do not employ multiple observations. Figure 4 illustrates an embodiment in which only one observation is made. In this case, given that two measured points A and B are highly correlated and have very similar measurements, an embodiment that operates to shift the quantization bins can be applied to these two points individually. As a result, obtaining the intersection of the two uncertainty ranges can increase depth accuracy.

[0080] Exemplary embodiments may be employed with lossy point cloud compression. At the encoder side, point cloud data (either in its original spatial domain or in the transform domain) is quantized, the quantization bins are intentionally shifted, and then the quantized signal is encoded. Information indicating the quantizer shift is available to both the encoder and the decoder. For example, the encoder may signal the shift to the decoder. At the decoder side, point cloud data measuring the same scene may be processed by taking the quantizer shift into account. An example of the application of such techniques for point cloud compression is described below in the section "Point Cloud Compression." Point cloud augmentation via multiple observations

[0081] An embodiment for augmenting 3D point cloud data using imprecise depth measurements from two camera views is described below. The exemplary embodiment uses convex optimization. While the example is described with respect to measurements from two camera views, it should be noted that the described techniques apply to other embodiments using three or more views.

[0082] Because 2D projections of the same 3D object onto two image planes are relative mappings, exemplary embodiments operate to enhance the quality of reconstructed 3D points by simultaneously optimizing the likelihood and presumption of both views. Mathematically, the signal-noise correlation of a pixel row can be modeled using a graph, and the model is learned using data from previous rows via constrained l1-norm minimization. A maximum a posteriori (MAP) problem can be formulated, which, after suitable linear approximation, becomes an unconstrained convex differentiable objective and can be solved using fast gradient methods (FGMs). Stereo Acquisition System

[0083] An exemplary embodiment uses data collected from a camera system, where the same 3D object is observed by two depth cameras from different viewpoints; specifically, there is an overlapping region of the fields of view (FoV) from the two cameras, and thus there are at least two observations of the same 2D object surface. Figure 5 illustrates a camera system as used in some embodiments. Each depth camera returns a depth map of size H x W, and finite precision, where each pixel is a noise-corrupted observation of the physical distance between the camera and the perceived 3D object, which is quantized to a B-bit representation.

[0084] Consider the case where two captured depth maps are rectified, such that a pixel in row i of the left view corresponds to a different pixel in the same row i of the right view. The rectification process may be performed as a pre-processing step before performing point cloud augmentation. Examples of rectification techniques that may be employed in some embodiments include those described in one or more of the following publications: C. Loop and Z. Zhang, "Computing rectifying homographies for stereo vision", in Proceedings,1999 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (Cat.No PR00149), volume 1, pages 125-131, IEEE, 1999; Y.-S. Kang and Y.-S. Ho, "An efficient image rectification method for parallel multi-camera arrangement", IEEE transactions on Consumer Electronics, 57(3):1041-1048, 2011; and Zhengyou Zhang, "A flexible new technique for camera calibration," IIEEE Transactions on pattern analysis and machine intelligence, 22, 2000.

[0085] The example embodiments are not limited to the use of any particular camera pose: for more complex cases other than stereo, the measured depth maps and multi-view geometry can still be used to find the mapping relationship between the views. Exemplary Image Formation Model

[0086] In some embodiments, a 3D object is recorded by two rectified depth maps (left and right) captured by two different cameras. (Other embodiments can be implemented by extending the techniques described in this example to data collected from more than two cameras, for example.)

number

number

number

number

number

number

[0087] Modified left and right views, ix l and x r Pixel row i of x is projected onto two different camera planes from the same 3D object and is therefore related. Consider the use case where there is no occlusion when projecting a 3D object onto two camera views. A 1D twisting procedure is performed to obtain l and x rThe twisting procedure may be a twisting procedure as described in one or more of the following publications: D. Scharstein and R. Szeliski, "High-accuracy stereo depth maps using structured light," in 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2003, Proceedings, volume 1, pages 1-1, IEEE, 2003; and J. Jin, A. Wang, Y. Zhao, C. Lin, and B. Zeng, "Region-aware 3-D warping for DIBR," IEEE Transactions on Multimedia, 18(6):953-966, 2016. The jth pixel x in the left view may be l,j , its horizontal position s(j,x l,j ) is expressed as follows:

number

[0088] Assuming that the object surface is smooth, the left pixel x l Given the right pixel x r is interpolated as follows:

number

number

number

[0089] In this example, the weight ω ij is x l,j The projected position s(j,x l,j ) and x r If the distance between the target pixel position i in is small, then

[0090] Equation (4) can be simplified by using a normalization constant.

number

[0091] Combined with equation (2), equation (4) can be rewritten as:

number

[0092] g(x l ) is differentiable, so the first-order Taylor series expansion is

number

number

number

number

number

number

[0093] The exemplary embodiment provides a zero-mean additive noise

number

number

number

number

number

[0094] The integration of equation (8) over the joint Gaussian PDF is not trivial. However, Pr(n l ) is an affine function of the following equation for the region R l can be approximated over

number

number

number

[0095] With a likelihood term in linear form, let us now turn to the signal before formulating the MAP problem. Depth maps are observed to be piecewise smooth (PWS). We assume that the previous K rows have been reconstructed and the next pixel row i follows a similar image structure. In this example, we choose to use a graph Laplace regularizer as the signal prior, using techniques such as those described in Gene Cheung, Enrico Magli, Yuichi Tanaka, and Michael K Ng, "Graph spectral image processing," Proceedings of the IEEE, 106(5), 907-930, 2018. The use of a Laplace regularizer is achieved by: l But (L l This influences the assumption that the graph should be smoothed for some graph G (with an associated graph Laplace matrix denoted as

number

[0096]

number

[0097] In this example, we do not limit the construction of the graph G. We can use the previously reconstructed results or the current corrupted measurement y l Another embodiment may use the prior probability Pr(x l Other techniques for determining Pr(x) may be used. In some such embodiments, l ) is the value of x l increases as the level of smoothness (or piecewise smoothness) of increases, and that smoothness can be expressed using one or more different techniques. A posteriori maximum (MAP) formulation

[0098] In some embodiments, the MAP problem involves estimating

number

number

[0099] To facilitate optimization, we instead minimize the negative logarithm of (17).

number

[0100] Formula (19) is an unconstrained convex differentiable objective and can be solved using gradient descent. When the above formula (19) is applied, there is no need to restrict the Laplace matrix and the method to estimate the noise parameters. After formula (19) is solved, the quantizer shift δ l is the estimated value

number

number

number

number

[0101] In some embodiments, refinement may be performed using more than two measurements. For example, some embodiments may use a set of K measurements (e.g., from K different camera positions). The set of functions g i (·) is the expression x i =g i (x1) to calculate the measurement x1 from the reference diagram by any other measurement x i (i≧2).

number

number

[0102] In some embodiments, equation (20) can be formulated using a first-order approximation by the procedure shown in equations (17), (18), and (19), which can lead to an unconstrained convex optimization problem that can be solved efficiently. Point Cloud Compression

[0103] Using systems and methods such as those described above, exemplary embodiments are used to enhance the accuracy of 3D point cloud data during acquisition. Further embodiments use such systems and methods in performing lossy point cloud compression (PCC).

[0104] In PCC, quantization is generally performed. The quantization of PCC has the same effect as the quantization performed during acquisition. As mentioned above, the operation of shifting the quantization bins is similar to adding a pre-designed amount to the point cloud signal being coded.

[0105] Figure 7 is a functional block diagram of an exemplary point cloud encoder.

number

number

number

number

number

number

number

[0106] On the decoder side, the encoded point cloud

number

number

number

number

[0107] When the refinement is applied at the decoding stage, as illustrated in FIG.

number

number

number

number

[0000] is also used. The SI here can be provided using any of a variety of different techniques or in a variety of different syntax elements such as the sequence parameter set, picture parameter set, or in a supplemental enhancement information message. The application of the quantization shift and any side information is denoted by module R in Figure 8, whose output is the final composite point cloud

number

[0108] In some embodiments, module R of FIG. 8 may be implemented using components such as those shown in FIG. 10. One module performs correlated signal identification. This module may operate to identify correlated frames or correlated regions within the input point cloud data. One example is view mapping in a stereo camera, as described above. This step may use additional side information such as camera parameters and pose information. Another module illustrated in FIG. 10 may perform joint precision augmentation. Precision augmentation may involve the use of quantization shifts within the point cloud data.

number

[0109] In some cases, the quantization shift δ i may be set to zero. In such cases, applying refinement R (FIG. 8) may still provide more accurate depth measurements as long as multiple observations in different point cloud frames can be distinguished. For example, see the illustrative diagram in FIG. 2A, where the two camera viewpoints are not parallel, the accuracy may be increased by taking the intersection of the uncertainty ranges of the two measurements.

[0110] In some embodiments, encoding and decoding of point cloud data may proceed as follows: A first parameter x1 is obtained for a point in a first point cloud, and a second parameter x2 is obtained for a point in a second point cloud. The parameters x1 and x2 may both be depth values, for example, or the values ​​may be other types of values ​​(e.g., representing horizontal or vertical position). The parameters x1 and x2 may be correlated, for example, they may represent the same or nearby points on the surface of an object captured by the same or different cameras at the same or different times.

[0111] In some embodiments, these parameters x1 and x2 are quantized to values ​​y1 and y2 by a method that includes adding quantizer shifts δ1 and δ2. For example, for a quantization step of 1, the parameters x1 and x2 may be quantized as follows:

number

[0112] Quantization steps of size other than 1 may be used, and the steps need not be the same for the two parameters. In some embodiments, one of the shifts (δ1 or δ2) may be zero. In some embodiments, the same shift δ1 may be used for all points in a first point cloud and the same shift δ2 may be used for all points in a second point cloud, while in other embodiments, different quantizer shifts may be used for different points (or for different parameters) in the point clouds.

[0113] Once the quantized point cloud parameters y1 and y2 are obtained, they may be encoded into a bitstream, for example, for storage and / or transmission to a decoder. In some embodiments, the quantizer shifts δ1 and δ2 may also be encoded into the bitstream. In some embodiments, the quantizer shifts δ1 and δ2 may be provided separately to the decoder, for example, in one or more supplemental enhancement information messages.

[0114] At the decoder, according to some embodiments, the point cloud parameters y1 and y2 and the quantizer shifts δ1 and δ2 are obtained. For example, the parameters y1 and y2 may be received in coded form and decoded by the decoder.

[0115] Given that information is provided to the decoder, the following upper bounds and bounds may be considered and applied to the shifted parameters (x1+δ1) and (x2+δ2):

number

number

[0116] The first quantization range may correspond, for example, to the "y1 range" illustrated in FIG. 3 or FIG. 4, and the second quantization range may correspond, for example, to the "y2 range" in the same figure. In some embodiments, refined point cloud data is obtained by selecting point cloud data (e.g., parameter x) within the overlapping region of these ranges. Such a region may correspond, for example, to the "refined range" illustrated in FIG. 3 or FIG. 4. Depending on which range is higher, the overlapping ranges are:

number

number

[0117] In some embodiments, the refined point cloud parameters

number

number

number

number

[0118] In some embodiments, the decoder may refine the received point cloud data using one or more of a variety of techniques, including, but not limited to, maximum a priori (MAP) refinement as described above.

[0119] An exemplary decoding method is illustrated in Figure 14. In the exemplary decoding method, data representing at least a first cloud of points and a second cloud of points is obtained (1402). Information identifying at least a first set of quantizer shifts associated with the first cloud of points is obtained (1404). In the first set of quantizer shifts, the shifts may be the same for different points or different parameters, or the shifts may be different for different points or different parameters. Refined point cloud data is obtained (1406) based on the first cloud of points and the second cloud of points, where obtaining the refined point cloud data includes performing a subtraction based on the at least first set of quantizer shifts.

[0120] An exemplary encoding method is illustrated in Figure 15. In the exemplary encoding method, data representing at least a first cloud of points and a second cloud of points is obtained (1502). The first cloud of points and the second cloud of points are quantized (1504) by a method that includes adding a first set of quantizer shifts to the first cloud of points. Information indicative of the quantized first and second clouds of points and the first set of quantizer shifts is encoded (1506) within a bitstream. Exemplary Applications

[0121] Various embodiments using PCC are given below.

[0122] Some embodiments are used in conjunction with inter-frame coding. When a frame from a point cloud video is coded using other frames, some embodiments jointly consider quantization shifts and other auxiliary frames before or after the current frame that observe the same scene. Information from other frames can be used to increase the accuracy of reconstruction at the decoder.

[0123] Some embodiments are used in conjunction with intraframe coding: when frames from a point cloud video are coded individually, some embodiments take into account the introduced quantization shift and highly correlated regions within the same point cloud frame to increase the accuracy of the reconstruction.

[0124] Some embodiments are used in conjunction with octree-based point cloud compression. When octrees are employed as the point cloud coding structure, exemplary embodiments operate to actively construct different tree roots with small quantization shifts. The decoder can then perform reconstruction and optimization regardless of whether the frames are coded as inter-frame or intra-frame.

[0125] Some embodiments are used in conjunction with transform coding. When transform coding is used and quantization is applied to transform coefficients, some embodiments operate to perform precision enhancement within the transform domain where the transform coefficients are to be optimized. Additional Embodiments and Information

[0126] In some embodiments, the point cloud decoding method comprises: l ) and at least the initial second point cloud data (y r ), and receiving refined first point cloud data (x) so as to substantially maximize a product of coefficients including one or more of the following: l ), and the coefficients are calculated based on the refined first point cloud data (x l ) is given as the initial first point cloud data (y l ) conditional probability Pr(y l |x l )) and the refined second point cloud data (x r ) estimate g(x l ) is given as the initial second point cloud data (y r ) conditional probability Pr(y r |g(x l )) where the estimate g(x l ) is the refined first point cloud data (x l) and the refined first point cloud data (x l ) prior probability Pr(x l ) and the refined second point cloud data (x r ) estimate g(x l ) prior probability Pr(g(x l )) and.

[0127] Some such embodiments may include: l ) and then calculate the quantizer shift δ from the refined first point cloud data. l This includes subtracting

[0128] In some embodiments, the refined first point cloud data (x l ) is performed repeatedly.

[0129] In some embodiments, the refined first point cloud data (x l ) is selected so as to substantially minimize the sum of the negative logarithms of the coefficients of the refined first point cloud data (x l ) is selected.

[0130] In some embodiments, the refined first point cloud data (x) is calculated so as to substantially minimize the sum of the negative logarithms of the coefficients. l ) is performed using gradient descent.

[0131] In some embodiments, the refined first point cloud data (x l ) is given as the initial first point cloud data (y l ) conditional probability Pr(y l )|(x l )) is the initial first point cloud data (y l ) and the refined first point cloud data (x l ) is expressed as a linear function of the difference between

[0132] In some embodiments, the refined second point cloud data (xr ) estimate g(x l ) is the refined second point cloud data (x r ) is a linear function of

[0133] In some embodiments, the refined second point cloud data (x r ) estimate g(x l ) is given as the initial second point cloud data (y r ) conditional probability Pr(y r |g(x l )) is the initial second point cloud data (y r ) and the refined second point cloud data (x r ) estimate g(x l ) is expressed as a linear function of the difference between

[0134] In some embodiments, the refined first point cloud data (x l ) prior probability Pr(x l )teeth,

number

[0135] In some embodiments, the refined second point cloud data (x r ) estimate g(x l ) prior probability Pr(g(x l )) is exp(-g(x l ) T L r g(x l ) / σ 2 ), where L r is the graph Laplacian matrix.

[0136] A method according to some embodiments includes receiving first point cloud data and at least second point cloud data, and refining the first point cloud data using the at least second point cloud data.

[0137] Some such embodiments further include, after refining the first point cloud data, subtracting the quantizer shift from the refined first point cloud data, hi some embodiments, the quantizer shift is a predetermined quantizer shift.

[0138] An encoding method according to some embodiments includes receiving initial first point cloud data and at least initial second point cloud data, processing the initial first and second point cloud data by a method that includes adding a first set of quantizer shifts δ1 to the first point cloud data, and encoding the processed first and second point cloud data.

[0139] In some embodiments, processing the initial first and second point cloud data further includes adding a second set of quantizer shifts δ2 to the second point cloud data.

[0140] In some embodiments, the first set of quantizer shifts δ1 is different from the second set of quantizer shifts δ2.

[0141] Some embodiments further include providing information identifying at least one of (i) a first set of quantizer shifts δ1, or (ii) a second set of quantizer shifts δ2 that are aligned with the encoded first and second point cloud data.

[0142] A decoding method according to some embodiments includes receiving encoded first point cloud data and at least encoded second point cloud data, receiving information identifying at least a first set of quantizer shifts δ1, decoding the first point cloud data and the at least second point cloud data, refining the decoded first point cloud data using the decoded second point cloud data, and subtracting the first set of quantizer shifts δ1 from the first point cloud data.

[0143] An apparatus according to some embodiments includes a processor configured to perform at least a method according to any of the processes described herein.

[0144] An apparatus according to some embodiments includes a processor and a computer-readable medium storing instructions operable to at least perform a method according to any of the processes described herein. In some such embodiments, the computer-readable medium is a non-transitory computer-readable medium.

[0145] Some embodiments include a computer-readable medium storing point cloud data encoded using any of the methods described herein. In some such embodiments, the computer-readable medium is a non-transitory computer-readable medium.

[0146] This specification describes various aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity, often in a manner that may sound limiting, at least to illustrate their individual characteristics. However, this is for the purpose of clarity of description and does not limit the scope of the disclosure or these aspects. In fact, all of the different aspects can be combined or interchanged to provide further aspects. Furthermore, these aspects can be combined or interchanged with aspects described in previous applications.

[0147] The aspects described and contemplated in this disclosure may be implemented in many different forms. While Figures 11, 12, and 13 below provide some embodiments, other embodiments are contemplated, and discussion of Figures 11, 12, and 13 is not intended to limit the scope of implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects may be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.

[0148] In this disclosure, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side, and the term "decoded" is used on the decoder side.

[0149] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first," "second," and the like may be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." The use of such terms does not imply a modified ordering of operations unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding, but may occur, for example, before, during, or during an overlapping time with the second decoding.

[0150] Various methods and other aspects described in this disclosure can be used to modify modules, such as the inter-prediction module, entropy coding module, and / or decoding module (160, 360, 145, 330) of video encoder 100 and decoder 200, as shown in Figures 11 and 12. Furthermore, aspects of this disclosure are not limited to VVC or HEVC, but can be applied, for example, to other standards and recommendations, whether existing or developed in the future, and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, aspects described in this disclosure may be used individually or in combination.

[0151] Various numerical values ​​are used in this disclosure, and the specific values ​​are for illustrative purposes only and the aspects being described are not limited to these specific values.

[0152] 11 illustrates an encoder 200. Although variations of this encoder 200 are contemplated, the encoder 200 is described below for purposes of clarity without necessarily describing all contemplated variations.

[0153] Before being encoded, the video sequence may undergo pre-encoding processing (204), such as applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components (e.g., using histogram equalization of one of the color components) to obtain a signal delivery that is more resilient to compression. Metadata may be associated with the pre-processing and attached to the bitstream.

[0154] In the encoder 200, a picture is encoded by the encoder elements, as described below. The picture to be encoded is divided (206) into units, e.g., CUs, and processed. Each unit is encoded, e.g., using either intra mode or inter mode. When a unit is encoded in intra mode, the unit performs intra prediction (208). In inter mode, motion estimation and compensation (210) is performed. The encoder determines (214) whether to use either intra mode or inter mode to encode the unit, and indicates the intra / inter decision, e.g., by a prediction mode flag. A prediction residual is calculated (216), e.g., by subtracting the prediction block from the original image block.

[0155] The prediction residual is then transformed (218) and quantized (220). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (232) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can skip both the transform and quantization, in which case the residual is coded directly without applying a transform or quantization process.

[0156] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (222) and inverse transformed (224) to decode the prediction residual. The decoded prediction residual and the predicted block are combined (226) to reconstruct an image block. An in-loop filter (228) is applied to the reconstructed picture, performing, for example, deblocking / SAO (Sample Adaptive Offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (212).

[0157] Figure 12 illustrates a block diagram of video decoder 250. In decoder 250, the bitstream is decoded by decoder elements as described below. Video decoder 250 typically performs a decoding pass that is reverse to the encoding pass, as described in Figure 11. Encoder 200 also typically performs video decoding as part of the encoded video data.

[0158] In particular, the decoder's input includes a video bitstream, which may be generated by the video encoder 200. The bitstream is first entropy decoded (254) to obtain transform coefficients, motion vectors, and other coding information. Picture partition information indicates how the picture is partitioned. Thus, the decoder may partition the picture according to the decoded picture partition information (256). The transform coefficients are inverse quantized (262) and inverse transformed (264) to decode the prediction residual. Combining the decoded prediction residual and the predicted block (266) reconstructs an image block. The predicted block may be obtained from intra prediction (258) or motion-compensated prediction (i.e., inter prediction) (260) (271). An in-loop filter (268) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (270).

[0159] The decoded picture may further undergo post-decoding processing (274), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4 conversion) or an inverse remapping that performs the inverse of the remapping process performed in pre-encoding processing (204). Post-decoding processing may use metadata derived in pre-encoding processing and signaled in the bitstream.

[0160] FIG. 13 is a block diagram of an example system in which various aspects and embodiments can be implemented. System 1000 can be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or separate components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described herein.

[0161] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein, e.g., to implement various aspects described herein. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes storage device(s) 1040, which may include non-volatile and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drives, and / or optical disk drives. Storage device(s) 1040 may include, by way of non-limiting example, internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0162] System 1000 includes, for example, an encoder / decoder module 1030 configured to process data to provide encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 1030 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Furthermore, encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software, as is known to those skilled in the art.

[0163] Program code loaded into the processor 1010 or the encoder / decoder 1030 to perform various aspects described herein may be stored in the storage device 1040 and then loaded into the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items during execution of the processes stored herein. Such stored items may include, but are not limited to, input video, decoded video, or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing of equations, expressions, operations, and operational logic.

[0164] In some embodiments, memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (which may be, for example, either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be the memory 1020 and / or the storage device 1040, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store, for example, the television's operating system. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video coding and decoding operations such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by the Joint Video Experts Team (JVET)).

[0165] Input to the elements of system 1000 can be provided through various input devices, as shown in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives RF signals transmitted, for example, by a broadcaster, (ii) a component (COMP) input terminal (or set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High-Definition Multimedia Interface (HDMI) input terminal. Another example, not shown in FIG. 10, includes composite video.

[0166] In various embodiments, the input devices of block 1130 have associated respective input processing elements, as is known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select a signal frequency band, which in particular embodiments may be referred to as a channel (for example), (iv) demodulating the downconverted, bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0167] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices over USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or within processor 1010, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, for example, in a separate interface IC or within processor 1010, as desired. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and an encoder / decoder 1030, which operates in combination with memory and storage elements to process the data stream as required for presentation on an output device.

[0168] The various elements of system 1000 may be provided within an integrated housing in which the various elements may be interconnected and transmit data between them using suitable connection arrangements 1140, such as internal buses as are known in the art, including Inter-IC (I2C) buses, wiring, and printed circuit boards.

[0169] The system 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented in a wired and / or wireless medium, for example.

[0170] In various embodiments, data is streamed or otherwise provided to system 1000 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received via communication channel 1060 and communication interface 1050 adapted for Wi-Fi communication. Communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, streaming data is provided to system 1000 using a set-top box that delivers data via an HDMI connection in input block 1130. In yet other embodiments, streaming data is provided to system 1000 using an RF connection in input block 1130. As noted above, various embodiments provide data in a manner other than streaming. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0171] The system 1000 can provide output signals to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be for a television, a tablet, a laptop, a mobile phone, or other device. The display 1100 can also be integrated into other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop). The other peripheral devices 1120, in various example embodiments, include one or more of a standalone digital video disc (or digital versatile disc) (DVR, or digital versatile disc, as an abbreviation for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functionality based on the output of the system 1000. For example, a disc player performs the function of playing the output of the system 1000.

[0172] In various embodiments, control signals are communicated between system 1000 and display 1100, speakers 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable control between devices with or without user intervention. Output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 using communication channel 1060 via communication interface 1050. Display 1100 and speakers 1110 can be integrated into a single unit with other components of system 1000 in an electronic device such as a television. In various embodiments, display interface 1070 includes a display driver, such as a timing controller (TCon) chip.

[0173] Alternatively, the display 1100 and speakers 1110 may alternatively be separate from one or more of the other components, for example, if the RF portion of the input 1130 is part of a separate set-top box. In various embodiments in which the display 1100 and speakers 1110 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0174] The embodiments may be performed by computer software implemented by the processor 1010, by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technology environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type suitable for the technology environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0175] Various implementations include decoding. As used in this disclosure, "decoding" may encompass all or some of the processes performed on a received encoded sequence, e.g., to generate a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by decoders of various implementations described in this disclosure, such as refining point cloud data and / or providing or removing quantizer shift from the point cloud data.

[0176] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to refer generally to a broader decoding process will be clear based on the context of the particular description and is believed to be well understood by those skilled in the art.

[0177] Various implementations include encoding. In a manner similar to the above discussion regarding "decoding," "encoding" as used in this disclosure may encompass all or some of the processes performed on an input video sequence, e.g., to generate an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as, for example, partitioning, differential encoding, transforming, quantizing, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by the encoders of various implementations described in this disclosure, e.g., refining point cloud data and / or providing or removing quantizer shift from the point cloud data.

[0178] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to refer generally to a broader encoding process will be clear based on the context of a particular description and is believed to be well understood by those skilled in the art.

[0179] It should be noted that syntax elements, e.g., sequence parameter set, picture parameter set, or supplemental enhancement information message, as used herein are descriptive terms and therefore do not preclude the use of other syntax element names.

[0180] Where a figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.

[0181] Implementations and aspects described herein may be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single type of implementation (e.g., discussed only as a method), the implementation of the discussed feature may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be implemented in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, handheld / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.

[0182] References to "one embodiment" or "embodiment" or "one implementation" or "implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation," as well as any other variations, appear in various places throughout this disclosure, but do not necessarily all refer to the same embodiment.

[0183] Additionally, this disclosure may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0184] Additionally, this disclosure may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or inferring information.

[0185] Additionally, this disclosure may refer to "receiving" various pieces of information. Receiving, like "accessing," is intended to be a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" generally involves in some way, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0186] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass selection of only the first listed alternative (A), or selection of only the second listed alternative (B), or selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass selection of only the first listed alternative (A), or selection of only the second listed alternative (B), or selection of only the third listed alternative (C), or selection of only the first and second listed alternatives (A and B), or selection of only the first and third listed alternatives (A and C), or selection of only the second and third listed alternatives (B and C), or selection of all three alternatives (A, B, and C). This can be expanded as many times as the number of listed items, as would be apparent to one of ordinary skill in this and related arts.

[0187] Also, as used herein, the term "signaling" specifically refers to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a specific quantization bin shift or set of shifts. Thus, in some embodiments, the same parameters are used on both the encoder and decoder sides. Thus, for example, an encoder can transmit specific parameters to a decoder (explicit signaling) so that the decoder can use the same specific parameters. Conversely, if the decoder already has the specific parameters and other parameters, it can use signaling without transmission (implicit signaling) to simply allow the decoder to recognize and select the specific parameters. By avoiding transmitting any actual functions, bit savings are realized in various embodiments. It should be understood that signaling can be achieved in various manners. For example, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder in various embodiments. While the above refers to the verb form of the word "signal," the word "signal" can also be used as a noun herein.

[0188] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be, for example, stored or transmitted. Information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0189] Several embodiments have been described. Features of the embodiments may be provided alone or in any combination across various claim categories and types. Furthermore, embodiments may include one or more of the following features, devices, or aspects, alone or in combination across various claim categories and types: a bitstream or signal including one or more of the described syntax elements or variations thereof; a bitstream or signal including syntax carrying information generated by any of the described embodiments; creating and / or transmitting and / or receiving and / or decoding a bitstream or signal including one or more of the described syntax elements or variations thereof; creating and / or transmitting and / or receiving and / or decoding by any of the described embodiments; and a method, process, apparatus, instruction storage medium, data storage medium, or signal according to any of the described embodiments.

[0190] It should be noted that the various hardware elements of one or more of the described embodiments are referred to as “modules,” which, in combination with the respective modules, perform (e.g., implement, perform, etc.) the various functions described herein. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) deemed suitable by one of ordinary skill in the art for a given implementation. It should be noted that each described module may also include executable instructions for performing one or more functions described as being performed by the respective module, which may take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., and may be stored on any suitable non-transitory computer-readable medium, or medium such as commonly referred to as RAM, ROM, etc.

[0191] Although features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with the other features and elements. Furthermore, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. 1. A point cloud decoding method, comprising: acquiring data representing at least a first point cloud and a second point cloud, the first point cloud representing a first observation of a scene and the second point cloud representing a second observation of the scene; obtaining information identifying at least a first set of quantizer shifts associated with the first cloud of points; obtaining refined point cloud data based on at least the first point cloud, the first set of quantizer shifts, and the second point cloud.

2. The method of claim 1 , wherein obtaining the refined point cloud data comprises performing a subtraction based on at least the first set of quantizer shifts.

3. The method of claim 1 , wherein the first point cloud represents a left view of a scene and the second point cloud represents a right view of the scene.

4. The method of claim 1 , wherein the first and second clouds of points are frames associated with different times.

5. 2. The method of claim 1 , wherein the first set of quantizer shifts includes at least a first shift associated with a first point in the first cloud of points and a different second shift associated with a second point in the first cloud of points.

6. The data representing the first point cloud includes a first parameter (y1) having a first quantization range, the data representing the second point cloud includes a second parameter (y2) having a second quantization range, and obtaining the refined point cloud data includes obtaining a third parameter (y1) within the quantization ranges of both the first parameter (y1) and the second parameter (y2). [Equation 1] The method of claim 1 , comprising obtaining:

7. Obtaining the refined point cloud data (xl) based on the first point cloud (yl) and the second point cloud (yr) includes selecting the refined first point cloud (xl) so as to substantially maximize a product of coefficients including one or more of the following coefficients, wherein the coefficients are: the conditional probability Pr(yl|xl) of the first point set (yl) given the refined first point set (xl); a conditional probability Pr(yr|g(xl)) of the refined second point cloud (xr) given an estimate g(xl) of the second point cloud (yr), where the estimate g(xl) is based on the refined first point cloud (xl); a prior probability Pr(xl) of the refined first point cloud (xl); and a prior probability Pr(g(xl)) of the estimate g(xl) of the refined second point cloud (xr).

8. A point cloud decoder device, comprising at least: acquiring data representing at least a first point cloud and a second point cloud, the first point cloud representing a first observation of a scene and the second point cloud representing a second observation of the scene; obtaining information identifying at least a first set of quantizer shifts associated with the first cloud of points; and obtaining refined point cloud data based on at least the first point cloud, the first set of quantizer shifts, and the second point cloud.

9. The apparatus of claim 8 , wherein obtaining the refined point cloud data comprises performing a subtraction based on at least the first set of quantizer shifts.

10. obtaining information identifying a second set of quantizer shifts associated with the second cloud of points, the second set of quantizer shifts being different from the first set of quantizer shifts; The apparatus of claim 8 , wherein obtaining the refined point cloud data is further based on the second set of quantizer shifts.

11. The apparatus of claim 8 , wherein the first point cloud represents a left view of a scene and the second point cloud represents a right view of the scene.

12. The apparatus of claim 8 , wherein the first and second clouds of points are frames associated with different times.

13. 9. The apparatus of claim 8, wherein the first set of quantizer shifts includes at least a first shift associated with a first point in the first cloud of points and a different second shift associated with a second point in the first cloud of points.

14. A point cloud encoding method, comprising: acquiring data representing at least a first point cloud and a second point cloud, the first point cloud representing a first observation of a scene and the second point cloud representing a second observation of the scene; quantizing the first cloud of points and the second cloud of points, wherein quantizing the first cloud of points and the second cloud of points comprises adding at least a first set of quantizer shifts to the first cloud of points; encoding the quantized first and second point clouds in a bitstream.

15. 15. The method of claim 14, further comprising encoding, within the bitstream, information indicative of the first set of quantizer shifts.

16. The method of claim 14 , wherein the first point cloud represents a left view of a scene and the second point cloud represents a right view of the scene.

17. The method of claim 14 , wherein the first and second clouds of points are frames associated with different times.

18. A point cloud encoding device, comprising at least: acquiring data representing at least a first point cloud and a second point cloud, the first point cloud representing a first observation of a scene and the second point cloud representing a second observation of the scene; quantizing the first cloud of points and the second cloud of points, wherein quantizing the first cloud of points and the second cloud of points comprises adding at least a first set of quantizer shifts to the first cloud of points; encoding the quantized first and second point clouds in a bitstream.

19. 20. The apparatus of claim 18, further comprising encoding, within the bitstream, information indicative of the first set of quantizer shifts.

20. 20. The apparatus of claim 18, wherein quantizing the first cloud of points and the second cloud of points comprises adding a second set of quantizer shifts to the second cloud of points, and further comprising encoding information indicative of the second set of quantizer shifts within the bitstream.

21. The apparatus of claim 18 , wherein the first point cloud represents a left view of a scene and the second point cloud represents a right view of the scene.

22. The apparatus of claim 18 , wherein the first and second clouds of points are frames associated with different times.

Citation Information

Patent Citations

  • Encoding method, decoding method, encoding device, and decoding device

    JP2024061790A

  • Point cloud geometry compression

    WO2019050931A1