Decision rules for attribute smoothing

By converting point clouds into 2D frames and compressing them using existing video codecs, combined with attribute and geometric smoothing techniques, the problems of point cloud transmission bandwidth and visual quality are solved, achieving an efficient compression and decompression process.

CN114467307BActive Publication Date: 2026-04-24SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2020-09-29
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Point cloud data needs to be compressed before transmission to reduce bandwidth requirements, but existing technologies usually require dedicated hardware, and the compression and decompression process can cause visual quality artifacts.

Method used

Point clouds are converted into 2D frames and compressed and reconstructed using existing video codecs. Artifacts are reduced and visual quality is improved through attribute and geometric smoothing techniques.

Benefits of technology

It effectively reduces the bandwidth required for point cloud transmission, avoids the need for dedicated hardware, and improves visual quality through smoothing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114467307B_ABST
    Figure CN114467307B_ABST
Patent Text Reader

Abstract

A method for point cloud decoding includes receiving a bitstream. The method also includes decoding the bitstream into a plurality of frames comprising pixels. Portions of the pixels are organized into patches and correspond to respective point clusters of a 3D point cloud. The method further includes decoding an occupancy map frame from the bitstream. The occupancy map frame indicates portions of the pixels comprising points of the 3D point cloud represented in the plurality of frames. Additionally, the method includes reconstructing the 3D point cloud using the plurality of frames and the occupancy map frame. The method also includes determining whether to perform smoothing on the 3D point cloud based at least in part on a characteristic of the plurality of frames. Based on determining to perform smoothing, the method includes performing smoothing on the 3D point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to multimedia data. More specifically, this disclosure relates to apparatus and methods for compressing and decompressing point clouds. Background Technology

[0002] With the ever-present availability of powerful handheld devices such as smartphones, 360° video is becoming a new way to experience immersive video. 360° video provides consumers with an immersive, “real-life,” “being there” experience by capturing a 360° view of the world. Users can interactively change their viewpoint and dynamically view any part of the captured scene or object they wish to see. Display and navigation sensors track the user’s head movements in real time to determine the area of ​​the 360° video the user wants to view. Multimedia data that is inherently three-dimensional (3D), such as point clouds, can be used in immersive environments.

[0003] A point cloud is a set of points representing an object in 3D space. Point clouds can be used in a wide variety of applications (such as games, 3D maps, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view replay, 6-DOF immersive media, to name a few). If uncompressed, point clouds typically require significant bandwidth for transmission. Due to high bit rate requirements, point clouds are usually compressed before transmission. Compressing 3D objects (such as point clouds) typically requires dedicated hardware. To avoid using dedicated hardware to compress 3D point clouds, they can be manipulated onto traditional 2D frames that can be compressed and reconstructed on different devices for viewing by a user. Compressing and decompressing 2D frames produces artifacts that degrade the visual quality of the point cloud. Summary of the Invention

[0004] Technical issues

[0005] A point cloud is a set of points representing an object in 3D space. Point clouds can be used in a wide variety of applications (such as games, 3D maps, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view replay, 6-DOF immersive media, to name a few). If uncompressed, point clouds typically require significant bandwidth for transmission. Due to high bit rate requirements, point clouds are usually compressed before transmission. Compressing 3D objects (such as point clouds) typically requires dedicated hardware. To avoid using dedicated hardware to compress 3D point clouds, they can be manipulated onto traditional 2D frames that can be compressed and reconstructed on different devices for viewing by a user. Compressing and decompressing 2D frames produces artifacts that degrade the visual quality of the point cloud.

[0006] Solution

[0007] This disclosure provides a decision-making rule for attribute smoothing modifications.

[0008] In one embodiment, a decoding apparatus for point cloud decoding is provided. The decoding apparatus includes a communication interface and a processor. The communication interface is configured to receive a bitstream. The processor is configured to decode the bitstream into a plurality of frames comprising pixels. Portions of the pixels are organized into patches and correspond to corresponding point clusters of a 3D point cloud. The processor is also configured to decode occupancy map frames from the bitstream. Occupancy map frames indicate portions of pixels representing points of the 3D point cloud included in the plurality of frames. The processor is further configured to reconstruct the 3D point cloud using the plurality of frames and the occupancy map frames. Additionally, the processor is configured to determine, at least in part, whether to perform smoothing on the 3D point cloud based on characteristics of the plurality of frames. Based on the determination to perform smoothing, the processor is configured to perform smoothing on the 3D point cloud.

[0009] In another embodiment, a method for point cloud decoding is provided. The method includes receiving a bitstream. The method further includes decoding the bitstream into a plurality of frames comprising pixels. Portions of pixels are organized into patches and correspond to corresponding point clusters of a 3D point cloud. The method further includes decoding occupancy frames from the bitstream. Occupancy frames indicate portions of pixels representing points of the 3D point cloud included in the plurality of frames. Additionally, the method includes reconstructing the 3D point cloud using the plurality of frames and the occupancy frames. The method further includes determining, at least in part, whether to perform smoothing on the 3D point cloud based on characteristics of the plurality of frames. Based on the determination to perform smoothing, the method includes performing the smoothing on the 3D point cloud.

[0010] Other technical features will be obvious to those skilled in the art based on the following figures, description and claims.

[0011] Before proceeding with the detailed description below, it may be advantageous to define certain words and phrases used throughout this patent document. The term “coupled” and its derivatives refer to any direct or indirect communication between two or more elements, whether or not these elements are physically in contact with each other. The terms “transmit,” “receive,” and “communicate,” and their derivatives cover both direct and indirect communication. The terms “include” and “comprise,” and their derivatives, indicate inclusion but are not limiting. The term “or” is inclusive, indicating and / or. The phrase “associated with” and its derivatives indicate inclusion, being included within, interconnected with, containing, being contained within, connected to or connected to, coupled to or coupled to, communicable with, cooperating with, interleaved, juxtaposed, proximate, bound to or bound to, having, possessing the characteristics of, related to, or having a relationship with, etc. The term “controller” means any device, system, or part thereof that controls at least one operation. Such a controller may be implemented in hardware or a combination of hardware and software and / or firmware. The functionality associated with any particular controller can be centralized or distributed, local or remote. When used with a list of items, the phrase “at least one of…” indicates that one or more different combinations of the listed items are available, and that only one item in the list may be required. For example, “at least one of A, B, and C” includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A and B and C.

[0012] Furthermore, the various functions described below can be implemented or supported by one or more computer programs, each computer program being formed by computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium accessible by a computer (such as read-only memory (ROM), random access memory (RAM), hard disk drive, optical disc (CD), digital video disc (DVD), or any other type of storage). "Non-transitory" computer-readable media does not include wired, wireless, optical, or other communication links that transmit transient electrical or other signals. Non-transitory computer-readable media includes media where data can be permanently stored and media where data can be stored and later rewritten (such as rewritable optical discs or erasable memory devices).

[0013] Definitions of certain other words and phrases are provided throughout this patent document. Those skilled in the art will understand that, in many (if not most) cases, such definitions apply to the prior and future use of the words and phrases defined in this way.

[0014] Beneficial effects of the present invention

[0015] According to this disclosure, point clouds can be compressed and decompressed more effectively. Attached Figure Description

[0016] To gain a more complete understanding of this disclosure and its advantages, reference is now made to the following description in conjunction with the accompanying drawings, wherein like reference numerals denote like parts:

[0017] Figure 1 An example communication system according to an embodiment of the present disclosure is shown;

[0018] Figure 2 and Figure 3 An example electronic device is shown according to an embodiment of the present disclosure;

[0019] Figure 4A An example 3D point cloud is shown according to an embodiment of the present disclosure;

[0020] Figure 4B A diagram showing a point cloud surrounded by a plurality of projection planes according to an embodiment of the present disclosure;

[0021] Figure 4C and Figure 4D The illustrations shown include representations according to embodiments of the present disclosure. Figure 4A Example 2D frames of a 3D point cloud;

[0022] Figure 4E Example color artifacts in a reconstructed 3D point cloud according to an embodiment of the present disclosure are shown;

[0023] Figure 5A A block diagram illustrating an example environment architecture according to embodiments of the present disclosure is shown;

[0024] Figure 5B An example block diagram of an encoder according to an embodiment of the present disclosure is shown;

[0025] Figure 5C An example block diagram of a decoder according to an embodiment of the present disclosure is shown;

[0026] Figure 6A and Figure 6B An example method for property smoothing according to an embodiment of the present disclosure is shown;

[0027] Figure 7AAn example method for selecting certain centroids used to perform smoothing is shown according to an embodiment of the present disclosure;

[0028] Figure 7B Example grids and cells are shown according to embodiments of the present disclosure;

[0029] Figure 7C Embodiments according to this disclosure are shown Figure 7B Example 3D cell of a grid;

[0030] Figure 8A A 2D example is shown for identifying and selecting certain centroids used to perform smoothing, according to an embodiment of the present disclosure;

[0031] Figure 8B A 3D example is shown according to an embodiment of the present disclosure for identifying and selecting certain centroids used to perform smoothing; and

[0032] Figure 9 An example method for decoding point clouds according to embodiments of the present disclosure is shown. Detailed Implementation

[0033] The following discussion Figures 1 to 9 The various embodiments used to describe the principles of this disclosure in this patent document are illustrative only and should not be construed as limiting the scope of this disclosure in any way. Those skilled in the art will understand that the principles of this disclosure can be implemented in any suitably arranged system or apparatus.

[0034] Virtual reality (VR) is a rendered version of a visual scene, in which the entire scene is computer-generated. Augmented reality (AR) is an interactive experience of a real-world environment, in which objects residing in the real-world environment are augmented by virtual objects, virtual information, or both. In some embodiments, AR and VR include both visual and audio experiences. Visual rendering is designed to mimic real-world visual stimuli and (if available) auditory sensory stimuli as naturally as possible to the observer or user as they move within constraints defined by the application or the AR or VR scene. For example, VR places the user in an immersive world that responds to the user's head movements. At the video level, VR is achieved by providing a video experience covering as much of the field of view (FOV) as possible, along with synchronizing the viewpoint of the rendered video with head movements.

[0035] Many different types of devices can provide immersive experiences associated with AR or VR. One example device is a head-mounted display (HMD). An HMD represents one of many types of devices that provide AR and VR experiences to users. An HMD is a device that allows users to view VR scenes and adjust the displayed content based on the movement of the user's head. Typically, HMDs rely on a dedicated screen integrated into the device and connected to an external computer (tethered), or on a device plugged into the HMD (untethered) (such as a smartphone). The first approach utilizes one or more lightweight screens and benefits from high computing power. In contrast, smartphone-based systems utilize greater mobility and are cheaper to manufacture. In both cases, the resulting video experience is the same. Note that, as used herein, the term "user" can refer to a person using the electronic device or another device (such as an AI-powered electronic device).

[0036] A point cloud is a virtual representation of a 3D object. For example, a point cloud is a collection of points in 3D space, with each point placed at a specific geometric location within that space and including one or more attributes (such as color). Point clouds can resemble virtual objects in a VR or AR environment. A mesh is another type of virtual representation of an object in a VR or AR environment. A point cloud or mesh can be an object, multiple objects, a virtual scene (which includes multiple objects), etc. Point clouds and meshes are commonly used in a variety of applications, including gaming, 3D mapping, visualization, medicine, AR, VR, autonomous driving, multi-view replay, 6DoF immersive media, to name just a few. As used herein, the terms point cloud and mesh are used interchangeably.

[0037] Point clouds represent volumetric visual data. A point cloud consists of multiple points placed in 3D space, where each point in the 3D point cloud includes a geometric location represented by a tuple of (X, Y, Z) coordinate values. When each point is identified using these three coordinates, its precise location in the 3D environment or space is determined. The position of each point in the 3D environment or space can be relative to the origin, other points in the point cloud, or a combination thereof. The origin is the location where the X, Y, and Z axes intersect. In some embodiments, the points are placed on the outer surface of an object. In other embodiments, the points are placed both within the internal structure and the outer surface of the object.

[0038] In addition to the geometric location of a point (its position in 3D space), each point in a point cloud may also include one or more attributes (such as color, texture, reflectivity, intensity, surface normal, etc.). In some embodiments, a single point in a 3D point cloud may have multiple attributes. In some applications, point clouds can also be used to approximate light field data, in which each point includes multiple viewpoint-dependent color information (R, G, B or Y, U, V triplets).

[0039] A single point cloud can contain hundreds of millions of points, each associated with a geometric location and one or more attributes. The geometric location and each associated attribute occupy a certain number of bits. For example, the geometric location of a single point in a point cloud might consume thirty bits. For instance, if each geometric location of a single point is defined using X, Y, and Z values, then each coordinate (X, Y, and Z) uses ten bits, totaling thirty bits. Similarly, the attribute specifying the color of a single point might consume twenty-four bits. For instance, if the color components of a single point are defined based on red, green, and blue values, then each color component (red, green, and blue) uses eight bits, totaling twenty-four bits. Therefore, a single point with ten bits of geometric attribute data per coordinate and eight bits of color attribute data per color value occupies fifty-four bits. Each additional attribute increases the number of bits required for a single point. If a frame contains one million points, then the number of bits per frame is fifty-four million bits (fifty-four bits per point multiplied by one million points per frame). If the frame rate is 30 frames per second and it is uncompressed, then 1.62 gigabytes per second (54 million bits per frame multiplied by 30 frames per second) will be sent from one electronic device to another so that the second device can display the point cloud. Therefore, sending an uncompressed point cloud from one electronic device to another uses a significant amount of bandwidth due to the size and complexity of the data associated with a single point cloud. Therefore, the point cloud is compressed before transmission.

[0040] Embodiments of this disclosure consider that compressing point clouds is necessary to reduce the amount of data (bandwidth) used when point clouds are transmitted from one device (such as a source device) to another device (such as a display device). Certain dedicated hardware components can be used to meet real-time requirements or reduce latency or lag in transmitting and rendering 3D point clouds; however, such hardware components are typically expensive. Furthermore, many video codecs are not capable of encoding and decoding 3D video content (such as point clouds). Utilizing existing 2D video codecs to compress and decompress point clouds makes point cloud encoding and decoding widely available without requiring new or dedicated hardware. According to embodiments of this disclosure, existing video codecs can be used to compress and reconstruct point clouds when they are converted from a 3D representation to a 2D representation. In some embodiments, the conversion of point clouds from a 3D representation to a 2D representation includes projecting clusters of points from the 3D point cloud onto a 2D frame using a generation slice. Thereafter, video codecs (such as HEVC, AVC, VP9, ​​VP8, VVC, etc.) can be used to compress the 2D frames representing the 3D point cloud, similar to 2D video.

[0041] To transmit a point cloud from one device to another, the geometric locations of the points are separated from their attribute information. The 3D point cloud is projected relative to different projection planes, causing it to be divided into clusters of points represented as patches on 2D frames. The first set of frames may include values ​​representing the geometric locations of the points. Each subsequent set of frames may represent a different attribute of the point cloud. For example, an attribute frame may include values ​​representing color information associated with each point. The decoder uses these frames to reconstruct the 3D point cloud, allowing it to be rendered, displayed, and then viewed by the user.

[0042] When a point cloud is deconstructed to fit multiple 2D frames and compressed, frames can be transmitted using less bandwidth than would be used to transmit the original point cloud. This is described in more detail below. Figures 4A-4D This illustrates the various stages of projecting a point cloud onto different planes and subsequently storing the projection into a 2D frame. For example, Figure 4A The image shows two views of a 3D point cloud, illustrating that the point cloud can be a 360° view of an object. Figure 4B This illustrates the process of projecting a 3D point cloud onto different planes. When projecting a point cloud (such as...) onto different planes... Figure 4A After the point cloud is projected onto different planes, Figure 4C and Figure 4D The geometric frames and attribute frames (which represent the colors of points in a 3D point cloud) are shown separately, including patches corresponding to various projections.

[0043] Embodiments of this disclosure provide systems and methods for converting point clouds into a 2D representation, which can be sent and then reconstructed into a point cloud for rendering. In some embodiments, the point cloud is deconstructed into multiple slices packaged into frames. In some embodiments, a frame includes slices having the same properties. When two slices are placed at the same coordinates, a point of the 3D point cloud represented in one slice of a frame corresponds to the same point represented in another slice of a second frame. For example, a pixel at position (u, v) in a frame representing geometry is the geometric position of a pixel at the same (u, v) position in a frame representing attributes (such as color). In other embodiments, a slice in a frame represents multiple attributes associated with points of the point cloud (such as the geometric position and color of the point in 3D space).

[0044] The encoder separates geometric and attribute information from each point. It groups (or clusters) the points of a 3D point cloud relative to different projection planes and then stores these point groups as slices on 2D frames. Slices representing geometric and attribute information are packaged into geometric video frames and attribute video frames, respectively, where each pixel within any slice corresponds to a point in 3D space. The geometric video frames are used to encode the geometric information, and the corresponding attribute video frames are used to encode the attributes of the point cloud, such as color. The position (U, V) of a pixel in a geometric frame corresponds to the (X, Y, Z) position of the point in 3D space. For example, the two lateral coordinates of a 3D point (relative to the projection plane) correspond to the column and row indices (u, v) in the geometric video frame that determine the position of the entire slice within the video frame, plus a lateral offset. The depth of a 3D point is encoded as the pixel value in the video frame plus the slice's depth offset. The depth of the 3D point cloud depends on whether the projection of the 3D point cloud is obtained from XY, YZ, or XZ coordinates.

[0045] After frames are generated, they can be compressed using various video compression codecs, image compression codecs, or both. For example, the encoder first generates geometry frames, then compresses them using a 2D video codec (such as HEVC). To encode attribute frames (such as the colors of a 3D point cloud), the encoder decodes the geometry frames, which are then used to reconstruct the 3D coordinates of the 3D point cloud. The encoder smoothly reconstructs the point cloud. Afterward, the encoder interpolates the color value of each point based on the color values ​​of the original point cloud. The interpolated color values ​​are then packed into a compressed color frame.

[0046] The encoder can also generate an occupancy map (also called an occupancy frame) that shows the locations of projected points in the 2D video frame. For example, since a piece may not occupy the entire generated frame, the occupancy map indicates which pixels in the geometry and attribute frames correspond to points in the point cloud, and which pixels are empty / invalid and do not correspond to points in the point cloud. In some embodiments, the occupancy frame is compressed. The compressed geometry frame, compressed color frame (and any other attribute frame), and occupancy frame can be multiplexed to generate a bitstream. The encoder or another device then sends the bitstream, including the 2D frames, to a different device.

[0047] The decoder receives the bitstream, decompresses it into frames, and reconstructs the point cloud based on the information within each frame. After reconstructing the point cloud, it can be smoothed to improve its visual quality. The reconstructed 3D points can then be rendered and displayed for user observation.

[0048] Embodiments of this disclosure also provide systems and methods for reducing memory allocations used for mesh-based geometry and color smoothing, wherein mesh-based geometry and color smoothing is used to increase the visual quality of 3D point clouds. For example, smoothing points in a 3D point cloud corresponding to pixels placed at or near sheet boundaries can improve the visual appearance of the point cloud.

[0049] For example, when a 3D point cloud is converted from a 3D representation to a 2D representation, the points of the 3D point cloud are clustered into groups and projected onto a frame, where the clustered points produce patches packed onto the 2D frame. Due to size constraints of certain 2D frames, two patches that are not adjacent to each other on the 3D point cloud may be packed adjacent to each other in a single frame. When two non-adjacent patches of the point cloud are packed adjacent to each other in a 2D frame, pixels from one patch may be unintentionally mixed with pixels from other patches by a block-based video codec. Similarly, even if a patch is not adjacent to any other patch, valid pixels may be unintentionally mixed with invalid pixels (blank spaces in the 2D frame, or spaces including padding to improve coding efficiency but not corresponding to points in the 3D point cloud) by a block-based video codec. When pixels from one patch are unintentionally mixed with other pixels, noticeable artifacts may appear at patch boundaries when the decoder reconstructs the point cloud. Therefore, embodiments of this disclosure provide systems and methods for smoothing the geometry and color of points near patch boundaries to avoid visible artifacts. Removing visible artifacts improves the visual quality of point clouds.

[0050] To perform attribute and geometric smoothing, embodiments of this disclosure provide systems and methods for identifying points in a reconstructed 3D point cloud, wherein points are represented by pixels in frames at or near patch boundaries. The identified points are referred to as boundary points because they are represented by pixels in frames at or near patch boundaries. Boundary points can be smoothed to remove visible artifacts, thereby improving the visual appearance of the point cloud. After identifying boundary points, certain cells including or near boundary points are identified and represented as boundary cells. After identifying boundary cells, a decoder derives the centroid of the boundary cells, which is used to (i) determine whether the boundary points need to be smoothed and (ii) smooth the boundary points. By deriving the centroid only for boundary cells instead of all cells, the decoder can reduce allocated memory by up to 80% for both mesh-based geometric smoothing and mesh-based color smoothing.

[0051] In some embodiments, attribute smoothing is customized as color smoothing in the BT.709 color space. Embodiments of this disclosure also provide systems and methods for improving color smoothing by specifying criteria for when attribute smoothing is performed, such that attribute smoothing is performed on attributes representing color. Additionally, embodiments of this disclosure provide systems and methods for improving color smoothing by implementing color smoothing in both the YUV and RGB color domains.

[0052] Figure 1 An example communication system 100 according to an embodiment of the present disclosure is shown. Figure 1 The embodiment of the communication system 100 shown is for illustrative purposes only. Other embodiments of the communication system 100 may be used without departing from the scope of this disclosure.

[0053] Communication system 100 includes a network 102 that facilitates communication between various components within the communication system 100. For example, network 102 may transmit IP packets, Frame Relay frames, Asynchronous Transfer Mode (ATM) cells, or other information between network addresses. Network 102 may include one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), all or part of a global network such as the Internet, or any other one or more communication systems located in one or more locations.

[0054] In this example, network 102 facilitates communication between server 104 and various client devices 106-116. Client devices 106-116 may be, for example, smartphones, tablets, laptops, personal computers, wearable devices, HMDs, etc. Server 104 may represent one or more servers. Each server 104 includes any suitable computing or processing device capable of providing computing services to one or more client devices (such as client devices 106-116). Each server 104 may, for example, include one or more processing devices, one or more memories storing instructions and data, and one or more network interfaces facilitating communication via network 102. As described in more detail below, server 104 may send compressed bitstreams representing point clouds to one or more display devices (such as client devices 106-116). In some embodiments, each server 104 may include an encoder.

[0055] Each client device 106-116 represents any suitable computing or processing device that interacts with at least one server (such as server 104) or other computing device via network 102. Client devices 106-116 include desktop computer 106, mobile phone or mobile device 108 (such as smartphone), PDA 110, laptop computer 112, tablet computer 114, and HMD 116. However, any other or additional client devices may be used in communication system 100. A smartphone represents a class of mobile devices 108 that are handheld devices with a mobile operating system and integrated mobile broadband cellular network connectivity for voice, short message service (SMS), and Internet data communication. HMD 116 may display a 360° scene including one or more 3D point clouds. In some embodiments, any of client devices 106-116 may include an encoder, a decoder, or both. For example, mobile device 108 may record video and then encode the video so that it can be sent to one of client devices 106-116. In another example, a laptop computer 112 can be used to generate a virtual 3D point cloud, which is then encoded and sent to one of the client devices 106-116.

[0056] In this example, some client devices 108-116 communicate indirectly with network 102. For example, mobile device 108 and PDA 110 communicate via one or more base stations 118 (such as cellular base stations or eNodeBs (eNBs)). Additionally, laptop computer 112, tablet computer 114, and HMD 116 communicate via one or more wireless access points 120 (such as IEEE 802.11 wireless access points). Note that these are for illustrative purposes only, and each client device 106-116 may communicate directly with network 102 or indirectly with network 102 via any suitable intermediary or network. In some embodiments, server 104 or any client devices 106-116 may be used to compress point clouds, generate bitstreams representing the point clouds, and send the bitstreams to another client device (such as any client devices 106-116).

[0057] In some embodiments, any of client devices 106-114 securely and efficiently transmits information to another device (such as, for example, server 104). Furthermore, any of client devices 106-116 can trigger information transmission between itself and server 104. Any of client devices 106-114 can be used as a VR display when attached to a headset via a bracket and functions similarly to HMD 116. For example, mobile device 108 can function similarly to HMD 116 when attached to a bracket system and worn on a user's eyes. Mobile device 108 (or any other client device 106-116) can trigger information transmission between itself and server 104.

[0058] In some embodiments, any of client devices 106-116 or server 104 may generate 3D point clouds, compress 3D point clouds, send 3D point clouds, receive 3D point clouds, render 3D point clouds, or a combination thereof. For example, server 104 receives a 3D point cloud, decomposes the 3D point cloud to fit a 2D frame, and compresses the frame to generate a bitstream. The bitstream may be sent to a storage device (such as a database, or one or more of client devices 106-116). As another example, one of client devices 106-116 may receive a 3D point cloud, decompose the 3D point cloud to fit a 2D frame, and compress the frame to generate a bitstream that may be sent to a storage device (such as a database, another of client devices 106-116) or server 104.

[0059] although Figure 1 An example of a communication system 100 is shown, but it is possible to modify it. Figure 1 Various changes can be made. For example, communication system 100 can include any number of each component in any suitable arrangement. Typically, computing and communication systems have a wide variety of configurations, and Figure 1 This disclosure is not intended to limit the scope to any particular configuration. Although Figure 1 This document illustrates an operating environment in which the various features disclosed in this patent document can be used, but these features can be used in any other suitable system.

[0060] Figure 2 and Figure 3 An example electronic device according to an embodiment of the present disclosure is shown. In particular, Figure 2 Example server 200 is shown, and server 200 can represent Figure 1 Server 104 in the context of server 200. Server 200 can represent one or more encoders, decoders, local servers, remote servers, cluster computers, and components acting as a single seamless resource pool, cloud-based servers, etc. Server 200 can be... Figure 1 One or more of the client devices 106-116 or another server can access the server.

[0061] Server 200 may represent one or more local servers, one or more compression servers, or one or more encoding servers (such as encoders). In some embodiments, the encoder may perform decoding. Figure 2 As shown, server 200 includes a bus system 205 that supports communication between at least one processing device (such as processor 210), at least one storage device 215, at least one communication interface 220 and at least one input / output (I / O) unit 225.

[0062] Processor 210 executes instructions that can be stored in memory 230. Processor 210 may include any suitable number and type of processors or other devices arranged in a reasonable manner. Example types of processor 210 include microprocessors, microcontrollers, digital signal processors, field-programmable gate arrays, application-specific integrated circuits, and discrete circuits. In some embodiments, processor 210 may encode a 3D point cloud stored in storage device 215. In some embodiments, when the encoder encodes the 3D point cloud, the encoder also decodes the encoded 3D point cloud to ensure that when the point cloud is reconstructed, the reconstructed 3D point cloud matches the 3D point cloud before encoding.

[0063] Memory 230 and persistent storage 235, as examples of storage device 215, represent any structure capable of storing and facilitating the retrieval of information (such as data, program code, or other suitable temporary or permanent information). Memory 230 may represent random access memory or any other suitable volatile or non-volatile storage device. For example, instructions stored in memory 230 may include instructions for decomposing a point cloud into pieces, instructions for packing pieces onto 2D frames, instructions for compressing 2D frames, and instructions for encoding 2D frames in a specific order to generate a bitstream. Instructions stored in memory 230 may also include instructions for rendering a 360° scene, such as through a VR headset (e.g.,...). Figure 1 The HMD 116 is for viewing. Persistent storage 235 may contain one or more components or devices (such as read-only memory, hard disk drive, flash memory, or optical disk) that support long-term storage of data.

[0064] Communication interface 220 supports communication with other systems or devices. For example, communication interface 220 may include features that facilitate communication via... Figure 1 The network interface card or wireless transceiver communicates with network 102. Communication interface 220 can support communication over any suitable physical or wireless communication link. For example, communication interface 220 can send a bitstream containing a 3D point cloud to another device (such as one of client devices 106-116).

[0065] I / O unit 225 allows for data input and output. For example, I / O unit 225 can provide connectivity for user input via a keyboard, mouse, keypad, touchscreen, or other suitable input device. I / O unit 225 can also send output to a display, printer, or other suitable output device. However, note that I / O unit 225 can be omitted (e.g., when I / O interaction with server 200 occurs via a network connection).

[0066] Note that, although Figure 2 Described as representing Figure 1 The server 104 may be used, but the same or similar architecture may be used in one or more of the various client devices 106-116. For example, desktop computer 106 or laptop computer 112 may have the same architecture as... Figure 2 The structures shown are the same or similar.

[0067] Figure 3 An example electronic device 300 is shown, and the electronic device 300 can represent Figure 1 One or more of the client devices 106-116. Electronic device 300 may be a mobile communication device, such as, for example, a mobile station, a user station, a wireless terminal, or a desktop computer (similar to...). Figure 1 Desktop computers 106), portable electronic devices (similar to) Figure 1 Mobile devices 108, PDA 110, laptop computer 112, tablet computer 114, or HMD 116, etc. In some embodiments, Figure 1 One or more of the client devices 106-116 may include the same or similar configuration as electronic device 300. In some embodiments, electronic device 300 is an encoder, decoder, or both. For example, electronic device 300 can be used for data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.

[0068] like Figure 3 As shown, electronic device 300 includes an antenna 305, a radio frequency (RF) transceiver 310, a transmit (TX) processing circuitry 315, a microphone 320, and a receive (RX) processing circuitry 325. The RF transceiver 310 may include, for example, an RF transceiver, a Bluetooth transceiver, a Wi-Fi transceiver, a ZigBee transceiver, an infrared transceiver, and various other wireless communication signals. Electronic device 300 also includes a speaker 330, a processor 340, an input / output (I / O) interface (IF) 345, an input 350, a display 355, a memory 360, and a sensor 365. The memory 360 includes an operating system (OS) 361 and one or more applications 362.

[0069] RF transceiver 310 receives incoming RF signals from antenna 305 from an access point (such as a base station, Wi-Fi router, or Bluetooth device) or other devices on network 102 (such as Wi-Fi, Bluetooth, cellular, 5G, LTE, LTE-A, WiMAX, or any other type of wireless network). RF transceiver 310 down-converts the incoming RF signals to generate an intermediate frequency (IF) or baseband signal. The IF or baseband signal is sent to RX processing circuitry 325, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or IF signal. RX processing circuitry 325 sends the processed baseband signal to speaker 330 (e.g., for voice data) or to processor 340 for further processing (e.g., for web browsing data).

[0070] The TX processing circuit 315 receives analog or digital voice data from the microphone 320 or other output baseband data from the processor 340. The output baseband data may include web data, email, or interactive video game data. The TX processing circuit 315 encodes, multiplexes, and / or digitizes the output baseband data to generate a processed baseband or intermediate frequency (IF) signal. The RF transceiver 310 receives the processed baseband or IF signal from the TX processing circuit 315 and up-converts the baseband or IF signal into an RF signal transmitted via the antenna 305.

[0071] Processor 340 may include one or more processors or other processing devices. Processor 340 may execute instructions stored in memory 360 (such as OS 361) to control the overall operation of electronic device 300. For example, processor 340 may control the RF transceiver 310, RX processing circuitry 325, and TX processing circuitry 315 to receive forward channel signals and transmit reverse channel signals according to well-known principles. Processor 340 may include any suitable number and type of processors or other devices in any reasonably arranged manner. For example, in some embodiments, processor 340 includes at least one microprocessor or microcontroller. Example types of processor 340 include microprocessors, microcontrollers, digital signal processors, field-programmable gate arrays, application-specific integrated circuits, and discrete circuits.

[0072] Processor 340 is also capable of executing other processes and programs residing in memory 360 (such as operations for receiving and storing data). Processor 340 may move data into or out of memory 360 as needed by the executing processes. In some embodiments, processor 340 is configured to execute one or more applications 362 based on OS 361 or in response to signals received from an external source or operator. For example, applications 362 may include encoders, decoders, VR or AR applications, camera applications (for still images and video), video call applications, email clients, social media clients, SMS messaging clients, virtual assistants, etc. In some embodiments, processor 340 is configured to receive and send media content.

[0073] The processor 340 is also coupled to an I / O interface 345, which provides the electronic device 300 with the ability to connect to other devices, such as client devices 106-114. The I / O interface 345 is the communication path between these accessories and the processor 340.

[0074] Processor 340 is also coupled to input 350 and display 355. An operator of electronic device 300 can use input 350 to input data or other information into electronic device 300. Input 350 may be a keyboard, touchscreen, mouse, trackball, voice input, or other device capable of acting as a user interface to allow the user to interact with electronic device 300. For example, input 350 may include voice recognition processing, allowing the user to input voice commands. In another example, input 350 may include a touch panel, (digital) pen sensor, key, or ultrasonic input device. Touch panel may recognize touch input, for example, in at least one mode (such as capacitive, pressure-sensitive, infrared, or ultrasonic). Input 350 may be associated with (one or more) sensors 365 and / or a camera by providing additional input to processor 340. In some embodiments, sensors 365 include one or more inertial measurement units (IMUs) (such as accelerometers, gyroscopes, and magnetometers), motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, etc. Input 350 may also include control circuitry. In a capacitive scheme, input 350 can recognize touch or proximity.

[0075] Display 355 may be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED), an active-matrix OLED (AMOLED), or other display capable of rendering text and / or graphics (such as from websites, videos, games, images, etc.). Display 355 may be sized to fit within an HMD. Display 355 may be a single display or multiple displays capable of producing a stereoscopic display. In some embodiments, display 355 is a head-up display (HUD). Display 355 may display 3D objects (such as 3D point clouds).

[0076] Memory 360 is coupled to processor 340. A portion of memory 360 may include RAM, and another portion may include flash memory or other ROM. Memory 360 may include a persistent storage device (not shown) representing any structure capable of storing and facilitating the retrieval of information such as data, program code, and / or other suitable information. Memory 360 may contain one or more components or devices (such as read-only memory, hard disk drive, flash memory, or optical disk) supporting long-term storage of data. Memory 360 may also contain media content. Media content may include various types of media (such as images, videos, 3D content, VR content, AR content, 3D point clouds, etc.).

[0077] The electronic device 300 also includes one or more sensors 365 capable of measuring physical quantities or detecting the activation state of the electronic device 300 and converting the measured or detected information into electrical signals. For example, sensors 365 may include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor (such as a gyroscope or gyroscope sensor and accelerometer), an eye-tracking sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyography (EMG) sensor, an electroencephalography (EEG) sensor, an electrocardiography (ECG) sensor, an IR sensor, an ultrasound sensor, an iris sensor, a fingerprint sensor, and color sensors (such as red-green-blue (RGB) sensors). Sensors 365 may also include control circuitry for controlling any of the included sensors.

[0078] As discussed in more detail below, one or more of these sensors 365 can be used to control the user interface (UI), detect UI input, determine the user's position and orientation for recognition of 3D content display, etc. Any of these sensors 365 may be located within the electronic device 300, within an auxiliary device operatively connected to the electronic device 300, within a headset configured to support the electronic device 300, or within a single device of the electronic device 300 including the headset.

[0079] Electronic device 300 can create media content (such as generating 3D point clouds or capturing (or recording) content via a camera). Electronic device 300 can encode the media content to generate a bitstream (similar to server 200 described above), allowing the bitstream to be sent directly to another electronic device or, for example, via... Figure 1 Network 102 is indirectly transmitted to another electronic device. Electronic device 300 can receive the bit stream directly from the other electronic device, or through, for example, via... Figure 1 Network 102 indirectly receives bit streams.

[0080] When encoding media content (such as point clouds), electronic device 300 or Figure 2 The server 200 can project a point cloud onto multiple slices. For example, clusters of points in the point cloud can be grouped together and represented as slices on a 2D frame. A slice can represent a single attribute of the point cloud (such as geometry, color, etc.). Slices representing the same attribute can be packaged into separate 2D frames. The 2D frames are then encoded to generate a bitstream. During the encoding process, additional content (such as metadata, flags, syntax elements, occupancy graphs, geometric smoothing parameters, one or more attribute smoothing parameters, slice streams, etc.) can be included in the bitstream.

[0081] Similarly, when decoding media content included in a bitstream representing a 3D point cloud, the electronic device 300 decodes the received bitstream into frames. In some embodiments, the decoded bitstream also includes an occupancy map, 2D frames, auxiliary information (such as one or more flags, one or more syntax elements, or quantization parameter sizes), etc. A geometry frame may include pixels indicating the geographic coordinates of points in the point cloud in 3D space. Similarly, an attribute frame may include pixels indicating the RGB (or YUV) color (or any other attribute) of each geometric point in 3D space. Auxiliary information may include one or more flags, one or more syntax elements or quantization parameter sizes, one or more thresholds, geometric smoothing parameters, one or more attribute smoothing parameters, a slice stream, or any combination thereof. After reconstructing the 3D point cloud, the electronic device 300 may render the 3D point cloud in three dimensions via a display 355.

[0082] although Figure 2 and Figure 3 Examples of electronic devices are shown, but are not applicable to... Figure 2 and Figure 3 Make various changes. For example, make them combinable, further subdivided, or omitted. Figure 2 and Figure 3The various components within it, and additional parts can be added as needed. As a specific example, processor 340 can be divided into multiple processors (such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs)). Furthermore, as with computing and communication, electronic devices and servers can have a wide variety of configurations, and Figure 2 and Figure 3 This disclosure is not limited to any particular electronic device or server.

[0083] Figure 4A , Figure 4B , Figure 4C , Figure 4D and Figure 4E The various stages of generating frames representing 3D point clouds are illustrated. Specifically, Figure 4A An example 3D point cloud 400 is shown according to an embodiment of the present disclosure. Figure 4B Figure 405 shows a point cloud surrounded by a plurality of projection planes according to an embodiment of the present disclosure. Figure 4C and Figure 4D The illustrations shown include representations according to embodiments of the present disclosure. Figure 4A A 2D frame of 400 3D point clouds. For example, Figure 4C A 2D frame 430 is shown, representing the geometric positions of points in a 3D point cloud 400, while Figure 4D Frame 440 shows the colors associated with points in the 3D point cloud 400. Figure 4E Example color artifacts are shown in the reconstructed point cloud 450. In some embodiments, the reconstructed point cloud 450 represents... Figure 4A The 3D point cloud is 400, but it is reconstructed for rendering on the user device, and Figure 4A The 3D point cloud 400 can be located on the server.

[0084] Figure 4A A 3D point cloud 400 is a set of data points in 3D space. Each point in the 3D point cloud 400 includes a geometric location that provides the structure of the 3D point cloud and one or more attributes that provide information about each point (such as color, reflectivity, material, etc.). The 3D point cloud 400 represents an entire 360° object. That is, the point cloud can be viewed from various angles (such as front 402, side and back 404, top, bottom).

[0085] Figure 4B Figure 405 includes point cloud 406. Point cloud 406 can be similar to... Figure 4AThe 3D point cloud 400 represents the entire 360° object. Point cloud 406 is surrounded by multiple projection planes (such as projection planes 410, 412, 414, 416, 418, and 420). Projection plane 410 is separated from projection plane 412 by a predefined distance. For example, projection plane 410 corresponds to projection plane XZ0, and projection plane 412 corresponds to projection plane XZ1. Similarly, projection plane 414 is separated from projection plane 416 by a predefined distance. For example, projection plane 414 corresponds to projection plane YZ0, and projection plane 416 corresponds to projection plane YZ1. Additionally, projection plane 418 is separated from projection plane 420 by a predefined distance. For example, projection plane 418 corresponds to projection plane XY0, and projection plane 420 corresponds to projection plane XY1. Note that additional projection planes may be included, and the shapes formed by the projection planes may differ.

[0086] During the partitioning process, each point in the point cloud 406 is assigned to a specific projection plane (such as projection planes 410, 412, 414, 416, 418, and 420). Points that are close to each other and assigned to the same projection plane are grouped together to form patches (such as...). Figure 4C and Figure 4D A cluster of points (any one of the points shown in the image). When assigning points to a particular projection plane, more or fewer projection planes may be used. Furthermore, projection planes may be in various positions and angles. For example, some projection planes may be tilted at 45 degrees relative to other projection planes, and similarly, some projection planes may be at a 90-degree angle relative to other projection planes.

[0087] Figure 4C and Figure 4D 2D frames 430 and 440 are shown respectively. Frame 430 is a geometric frame because it shows... Figure 4A The geometric position of each point in the 3D point cloud 400. Frame 430 includes multiple patches (such as patch 432) representing the depth values ​​of the 3D point cloud 400. The value of each pixel in frame 430 is represented as a lighter or darker color and corresponds to the distance of each pixel from a specific projection plane (such as...). Figure 4B The distance to one of the projection planes 410, 412, 414, 416, 418 and 420.

[0088] Frame 440 is a color frame (a type of attribute), because it provides... Figure 4A The color of each point in the 3D point cloud 400. Frame 440 includes multiple slices (such as slice 442) representing values ​​corresponding to the colors of the points in the 3D point cloud 400.

[0089] Figure 4C and Figure 4DEach slice in a frame can be identified by its index number. Similarly, each pixel within a slice can be identified by its position within the frame and the index number of the slice in which the pixel belongs.

[0090] There is a correspondence (or mapping) between frames 430 and 440. That is, each pixel in frame 430 corresponds to a pixel at the same location in frame 440. Each color pixel in frame 440 corresponds to a specific geometric pixel in frame 430. For example, a mapping is generated between each pixel in frames 430 and 440. For example, each pixel within piece 432 corresponds to a point in 3D space, and each pixel within piece 442 provides color to a point in the 3D point cloud represented at the same location in piece 432. As shown in frames 430 and 440, some pixels correspond to valid pixels representing the 3D point cloud 400, while other pixels (black areas in the background) correspond to invalid pixels that do not represent the 3D point cloud 400.

[0091] Non-nearest points in 3D space can ultimately be represented as pixels that are adjacent to each other in frames 430 and 440. For example, two clusters of points that are not adjacent to each other in 3D space can be represented as patches that are adjacent to each other in frames 430 and 440.

[0092] Frames 430 and 440 can be encoded using video codecs (such as HEVC, AVC, VP9, ​​VP8, VVC, AV1, etc.). The decoder receives the bitstream including frames 430 and 440, reconstructs the geometry of the 3D point cloud from frame 430, and colors the geometry of the point cloud based on frame 440 to generate the reconstructed point cloud 450, such as... Figure 4E As shown.

[0093] Figure 4E The reconstructed point cloud 450 should resemble the 3D point cloud 400. When frames 430 and 440 are encoded and compressed, the values ​​corresponding to pixels can be mixed by a block-based video codec. If pixels within a single slice of frame 430 are mixed, the effect is generally negligible when reconstructing the point cloud, as colors adjacent to each other within a slice are usually similar. However, if pixels at the boundary of a slice (such as slice 432) in frame 430 are mixed with pixels from another slice (or with pixels not belonging to a slice, such as invalid pixels), a result similar to [example image would be inserted here] will be produced when reconstructing the point cloud. Figure 4EThe artifact shown is an artifact of artifact 452. Since the patches can originate from distinctly different parts of the point cloud, the shading of the patches can be different. In a block-based video codec, the encoded block can contain pixels from blocks with drastically different shading. This causes color to leak from one patch to another patch with a drastically different texture. Therefore, visible artifacts that degrade the visual quality of the point cloud are produced. Similarly, if pixels at the boundary of a patch in a patch are mixed with empty pixels (indicated by a black background and not corresponding to points in the point cloud), artifacts will occur when the point cloud is reconstructed because the points corresponding to the pixels swapped with the empty pixels will not be reconstructed.

[0094] The reconstructed point cloud 450 exhibits artifact 452. Artifact 452 can occur when a patch corresponding to the forehead of the model represented in the 3D point cloud 400 is packed into frame 430 (or frame 440), immediately adjacent to a patch corresponding to another portion of the 3D point cloud 400 (such as the dress of the model represented in the 3D point cloud 400 or empty (invalid) pixels in frames 430 and 440). Similarly, color values ​​from a patch representing a portion of the dress may leak into the patch corresponding to the forehead of the model represented in the 3D point cloud 400. In this example, the mixing of color values ​​results in artifacts appearing as cracks or holes in the user's face, degrading the visual quality of the reconstructed point cloud 450. Embodiments of this disclosure provide systems and methods for removing artifacts by smoothing the reconstructed point cloud at regions of artifacts while preserving the quality of the point cloud. For example, identifying and smoothing points near patch boundaries in the reconstructed point cloud.

[0095] although Figure 4A , Figure 4B , Figure 4C , Figure 4D and Figure 4E Example point clouds and 2D frames representing point clouds are shown, but more details are available. Figure 4A , Figure 4B , Figure 4C , Figure 4D and Figure 4E Various modifications can be made. For example, a point cloud or mesh can represent a single object, while in other embodiments, a point cloud or mesh can represent multiple objects, scenery (such as a landscape), virtual objects in AR, etc. In another example, a piece included in a 2D frame can represent other properties (such as brightness, material, etc.). Figure 4A , Figure 4B , Figure 4C , Figure 4D and Figure 4E This disclosure is not limited to any particular 3D object and 2D frame representing a 3D object.

[0096] Figure 5A , Figure 5B and Figure 5C A block diagram is shown according to an embodiment of the present disclosure. Specifically, Figure 5A A block diagram of an example environment architecture 500 according to an embodiment of the present disclosure is shown. Figure 5B Embodiments according to this disclosure are shown Figure 5A An example block diagram of encoder 510, and Figure 5C Embodiments according to this disclosure are shown Figure 5A Example block diagram of decoder 550. Figure 5A , Figure 5B and Figure 5C The embodiments described are for illustrative purposes only. Other embodiments may be used without departing from the scope of this disclosure.

[0097] like Figure 5A As shown, the example environment architecture 500 includes an encoder 510 and a decoder 550 communicating via a network 502. The network 502 can connect to... Figure 1 Network 502 is the same as or similar to network 102. In some embodiments, network 502 represents a "cloud" of computers interconnected via one or more networks, where a network is a computing system that utilizes clustered computers and components that act as a single, seamless pool of resources when accessed. Furthermore, in some embodiments, network 502 is similar to one or more servers (such as...) Figure 1 Server 104, Server 200), one or more electronic devices (such as Figure 1 The network 502 is connected to client devices 106-116, electronic device 300, encoder 510, and decoder 550. Furthermore, in some embodiments, the network 502 may be connected to a database (not shown) containing VR and AR media content, which may be encoded by encoder 510, decoded by decoder 550, or rendered and displayed on an electronic device.

[0098] In some embodiments, encoder 510 and decoder 550 may represent Figure 1 One of server 104 and client devices 106-116 Figure 2 Server 200 Figure 3The encoder 510 and decoder 550 may be an electronic device 300 or another suitable device. In some embodiments, the encoder 510 and decoder 550 may be a “cloud” of computers interconnected via one or more networks, where each network is a computing system utilizing clustered computers and components that act as a single seamless pool of resources when accessed via network 502. In some embodiments, portions of the components included in the encoder 510 or decoder 550 may be included in different devices (such as multiple servers 104 or 200, multiple client devices 106-116, or other combinations of different devices). In some embodiments, the encoder 510 is operatively connected to an electronic device or a server, while the decoder 550 is operatively connected to an electronic device. In some embodiments, the encoder 510 and decoder 550 are the same device or are operatively connected to the same device.

[0099] Below Figure 5B The encoder 510 is described in more detail here. Typically, the encoder 510 is sourced from sources such as servers (similar to...). Figure 1 Server 104 Figure 2 Another device, such as a server 200, a database, or a client device 106-116, receives 3D media content (such as point clouds). In some embodiments, encoder 510 may receive media content from multiple cameras and stitch the content together to generate a 3D scene including one or more point clouds.

[0100] Encoder 510 projects points from the point cloud onto multiple patches representing the projection. Encoder 510 clusters the points of the point cloud into groups that are projected onto different planes (such as the XY plane, YZ plane, and XZ plane). When projected onto a plane, each point cluster is represented by a patch. Encoder 510 packs the information representing the point cloud into a 2D frame. Encoder 510 packs the patches representing the point cloud into a 2D frame. The 2D frame can be a video frame. Note that the points of the 3D point cloud are located in 3D space based on (X,Y,Z) coordinate values, but when the point is projected onto a 2D frame, the pixel representing the projected point is represented by the column and row indices of the frame indicated by coordinates (u,v). Furthermore, 'u' and 'v' can range from zero to the number of rows or columns in the depth image, respectively.

[0101] Each 2D frame represents a specific attribute; for example, one set of frames might represent geometry, and another set might represent attributes such as color. It should be noted that additional frames can be generated based on more layers and each additionally defined attribute.

[0102] The encoder 510 also generates an occupancy map based on geometry and attribute frames to indicate which pixels within a frame are valid. Typically, for each pixel within a frame, the occupancy map indicates whether the pixel is valid or invalid. For example, if a pixel at coordinates (u, v) in the occupancy map is valid, then the corresponding pixel at coordinates (u, v) in both the geometry and attribute frames is also valid. If a pixel at coordinates (u, v) in the occupancy map is invalid, the decoder skips the corresponding pixel at coordinates (u, v) in both the geometry and attribute frames. Invalid pixels may include information (such as padding) that improves encoding efficiency but does not provide any information associated with the point cloud itself. Typically, the occupancy map is binary, such that each pixel has a value of either 1 or 0. For example, when the value of a pixel at position (u, v) in the occupancy map is 1, it indicates that the pixel at (u, v) in both the attribute and geometry frames is valid. Conversely, when the value of a pixel at position (u, v) in the occupancy map is zero, it indicates that the pixel at (u, v) in both the attribute and geometry frames is invalid and therefore does not represent a point in the 3D point cloud.

[0103] Encoder 510 transmits frames representing the point cloud as an encoded bitstream. The bitstream can be transmitted via network 502 to a repository (such as a database) or an electronic device including a decoder (such as decoder 550), or decoder 550 itself. The following... Figure 5B The encoder 510 is described in more detail below.

[0104] Below Figure 5C Decoder 550, described more extensively in the document, receives a bitstream representing media content, such as a point cloud. The bitstream may include data representing a 3D point cloud. In some embodiments, decoder 550 may decode the bitstream and generate multiple frames (such as one or more geometry frames, one or more attribute frames, and one or more occupancy map frames). Decoder 550 uses the multiple frames to reconstruct a point cloud that can be rendered and viewed by a user. Decoder 550 may identify points on or near the boundaries of a slice within a slice of a frame.

[0105] Decoder 550 can also perform smoothing (such as geometric smoothing and attribute smoothing). To perform smoothing, decoder 550 identifies the boundary points of the reconstructed 3D point cloud and then identifies the boundary cells associated with those boundary points. Decoder 550 derives the centroid values ​​of the identified boundary cells. The centroid values ​​are used to determine whether smoothing is necessary. When decoder 550 determines that smoothing is necessary for a particular boundary point, it uses those centroid values ​​to smooth the boundary point based on the centroid values ​​of the cells associated with that particular boundary point.

[0106] For example, the decoder determines whether geometric smoothing is necessary based on the distance between the query point and the output of the trilinear filter at the centroid. Before determining whether color smoothing is necessary, the decoder 550 first determines whether to perform color smoothing on the color frame. When determining that attribute smoothing should be performed based on criteria, the decoder 550 determines whether color smoothing is necessary based on the color and brightness of points near the boundary points.

[0107] Figure 5B An encoder 510 is shown that receives a 3D point cloud 512 and generates a bitstream 534. The bitstream 534 includes data representing the 3D point cloud 512. The bitstream 534 may include multiple bitstreams. The bitstream 534 can be transmitted via... Figure 5A Network 502 is sent to another device (such as decoder 550, electronic device including decoder 550, or information library). Encoder 510 includes slice generator and packer 514, one or more encoding engines (such as encoding engines 522a, 522b, and 522c collectively referred to as encoding engines 522), attribute generator 528, and multiplexer 532.

[0108] The 3D point cloud 512 can be stored in a memory (not shown) or received from another electronic device (not shown). The 3D point cloud 512 can be a single 3D object (similar to...). Figure 4A A 3D point cloud (400) or a grouping of 3D objects. A 3D point cloud (512) can be a static object or a moving object.

[0109] The slice generator and packer 514 generate slices by acquiring projections of the 3D point cloud 512 and pack the slices into frames. In some embodiments, the slice generator and packer 514 divides the geometric and attribute information of each point in the 3D point cloud 512. The slice generator and packer 514 may use two or more projection planes (such as...) Figure 4B Two or more projection planes (410-420) are used to cluster the points of the 3D point cloud 512 to generate patches. The geometric patches are finally packed into a geometric frame 516.

[0110] The slice generator and packer 514 determine the optimal projection plane for each point of the 3D point cloud 512. During projection, each cluster of points in the 3D point cloud 512 is represented as a slice (also called a regular slice). A single point cluster can be represented by multiple slices (located on different frames), where each slice represents a specific aspect of each point within the cluster. For example, a slice representing the geometric location of a point cluster is located on geometry frame 516, and a slice representing the attributes of a point cluster is located on attribute frame 520.

[0111] After determining the optimal projection plane for each point in the 3D point cloud 512, the slice generator and packer 514 divide the points into slice data structures that are packed into frames (such as geometric frames 516). Figure 4C and Figure 4D As shown above, slices are organized by attributes and placed within corresponding frames (e.g., slice 432 is included in geometric frame 430, and slice 442 is included in attribute frame 440). Note that slices representing different attributes of the same cluster of points include correspondences or mappings, where a pixel in one slice corresponds to the same pixel in another slice based on the pixel's position in the corresponding frame.

[0112] The slice generator and packer 514 also generate slice information (providing information about the slices, such as the index number associated with each slice), occupied map frame 518, geometry frame 516, and attribute information (which is used by the attribute generator 528 to generate attribute frame 520).

[0113] Occupancy frame 518 indicates the occupancy map of valid pixels in an indicator frame (such as geometry frame 516). For example, occupancy frame 518 indicates whether each pixel in geometry frame 516 is a valid pixel or an invalid pixel. Each valid pixel in occupancy frame 518 corresponds to a pixel in geometry frame 516 that represents a position point of 3D point cloud 512 in 3D space. Conversely, invalid pixels are pixels within occupancy frame 518 that correspond to points in geometry frame 516 that do not represent points of 3D point cloud 512 (such as...). Figure 4C and 4D The pixels occupying empty / black spaces in frames 430 and 440. In some embodiments, one of the occupies frame 518 may correspond to both geometry frame 516 and attribute frame 520 (discussed below).

[0114] For example, when the slice generator and packer 514 generate occupancy frame 518, occupancy frame 518 includes a predefined value (such as zero or one) for each pixel. For example, when the pixel in the occupancy frame at position (u, v) has a value of zero, it indicates that the pixel at (u, v) in geometry frame 516 is invalid. Similarly, when the pixel in the occupancy frame at position (u, v) has a value of one, it indicates that the pixel at (u, v) in geometry frame 516 is valid, and thus includes information representing points in the 3D point cloud.

[0115] Geometric frame 516 includes pixels representing the geometric values ​​of 3D point cloud 512. Geometric frame 516 includes the geographic location of each point in 3D point cloud 512. Geometric frame 516 is used to encode the geometric information of the point cloud. For example, the two lateral coordinates of a 3D point (relative to the projection plane) correspond to the column and row indices in the geometric video frame (u, v) indicating the position of the entire slice within the video frame, plus a lateral offset. The depth of the 3D point is encoded as the pixel value in the video frame plus the slice's depth offset. The depth of the 3D point cloud depends on whether the projection of the 3D point cloud is obtained from XY, YZ, or XZ coordinates.

[0116] Encoder 510 includes one or more encoding engines 522. In some embodiments, frames (such as geometry frames 516, occupancy map frames 518, and attribute frames 520) are encoded by independent encoding engines 522, as illustrated. In other embodiments, a single encoding engine performs the encoding of frames.

[0117] Encoding engine 522 can be configured to support data with 8-bit, 10-bit, 12-bit, 14-bit, or 16-bit precision. Encoding engine 522 may include video or image codecs (such as HEVC, AVC, VP9, ​​VP8, VVC, EVC, AV1, etc.) to compress 2D frames representing 3D point clouds. One or more of encoding engines 522 can compress information in a lossy or lossless manner.

[0118] As shown in the figure, encoding engine 522a receives geometry frame 516 and performs geometry compression to generate geometry substream 524a. Encoding engine 522b receives occupancy graph frame 518 and performs occupancy graph compression to generate occupancy graph substream 526a. Encoding engine 522c receives attribute frame 520 and performs attribute compression to generate attribute substream 530.

[0119] After the encoding engine 522a generates the geometry substream 524a, the decoding engine (not shown) can decode the geometry substream 524a to generate the reconstructed geometry frame 524b. Similarly, after the encoding engine 522b generates the occupancy map substream 526a, the decoding engine (not shown) can decode the occupancy map substream 526a to generate the reconstructed occupancy map frame 526b.

[0120] The attribute generator 528 generates the attribute frame 520 based on the attribute information from the 3D point cloud 512, the reconstructed geometry frame 524b, the reconstructed occupancy map frame 526b, and the information provided by the slice generator and packer 514.

[0121] For example, to generate one of the attribute frames 520 representing color, geometry frame 516 is compressed by encoding engine 522a using a 2D video codec (such as HEVC). Geometry substream 524a is decoded to generate reconstructed geometry frame 524b. Similarly, occupancy map frame 518 is compressed using encoding engine 522b and then decompressed to generate reconstructed occupancy map frame 526b. Encoder 510 can then reconstruct the geometric locations of points in the 3D point cloud based on the reconstructed geometry frame 524b and the reconstructed occupancy map frame 526b. Attribute generator 528 interpolates the attribute values ​​(such as color) of each point from the color values ​​of the input point cloud to the reconstructed point cloud and the original point cloud 512. The interpolated colors are then partitioned by attribute generator 528 to match patches with the same geometric information. Attribute generator 528 then packs the interpolated attribute values ​​into attribute frame 520 representing color.

[0122] Attribute frames 520 represent different attributes of the point cloud. For example, for one of the geometry frames 516, there may be one or more corresponding attribute frames 520. Attribute frames may include color, texture, normals, material properties, reflection, motion, etc. In some embodiments, one attribute frame 520 may include the color value of each geometric point within a geometry frame of geometry 516, while another attribute frame may include a reflectance value indicating the reflectance level of each corresponding geometric point within the same geometry frame 516. Each additional attribute frame 520 represents other attributes associated with a particular geometry frame 516. In some embodiments, each geometry frame 516 has at least one corresponding attribute frame 520.

[0123] Multiplexer 532 combines chip stream, geometry substream 524a, occupancy graph substream 526a and attribute substream 530 to produce bitstream 534.

[0124] Figure 5C The decoder 550 is shown, which includes a demultiplexer 552, one or more decoding engines, a reconstruction engine 556, a boundary detection engine 558, and a smoothing engine 560.

[0125] Decoder 550 receives bitstream 534 (such as the bitstream generated by encoder 510). Demultiplexer 552 separates bitstream 534 into one or more substreams representing different information. For example, demultiplexer 552 separates various data streams into separate substreams, such as slice substream, geometry substream 524a, occupancy map substream 526a, and attribute substream 530.

[0126] Decoder 550 includes one or more decoding engines. For example, decoder 550 may include decoding engine 554a, decoding engine 554b, decoding engine 554c, and decoding engine 554d (collectively referred to as decoding engine 554). In some embodiments, a single decoding engine performs the operations of all individual decoding engines 554.

[0127] Decoding engine 554a decodes the geometric substream 524a into the reconstructed geometric frame 516a. The reconstructed geometric frame 516a is similar to... Figure 5B The geometric frame 516 differs in that one or more pixels may be shifted due to the encoding and decoding of the frame.

[0128] Decoding engine 554b decodes the occupancy map substream 526a into a reconstructed occupancy map frame 518a. The reconstructed occupancy map frame 518a is similar to... Figure 5B The difference between the 518 occupies a frame is that one or more pixels may be shifted due to the encoding and decoding of the frame.

[0129] Decoding engine 554c decodes attribute substream 530 into reconstructed attribute frame 520a. Reconstructed attribute frame 520a is similar to... Figure 5B The attribute frame 520 differs in that one or more pixels may be shifted due to the encoding and decoding of the frame.

[0130] After decoding the slice information, the reconstructed geometry frame 516a, the reconstructed occupancy frame 518a, and the reconstructed attribute frame 520a, the reconstruction engine 556 generates the reconstructed point cloud. The reconstruction engine 556 reconstructs the point cloud.

[0131] When geometry frame 516, occupancy map frame 518, and attribute frame 520 are encoded by encoding engine 522 and later decoded at decoder 550, pixels from one patch may unintentionally be swapped with pixels or invalid pixels from another patch. Therefore, noticeable artifacts may appear in the reconstructed point cloud, degrading its visual quality. For example, pixels within the reconstructed geometry frame 516a may be slightly shifted due to encoding and decoding processes. Typically, slight shifts may not significantly degrade the visual quality of the point cloud when pixels are located in the middle of a patch. However, slight shifts or transformations of pixels leaving a patch to locations indicated as empty (or invalid) by the occupancy map can cause considerable artifacts because parts of the image will not be rendered. Similarly, slight shifts or transformations of pixels from one patch to another can cause considerable artifacts. For example, if including… Figure 4A A patch representing the face of a 3D point cloud 400 is packaged as a patch adjacent to a patch representing the dress containing the 3D point cloud 400. If the encoding / decoding process causes pixel groups to shift from one patch to another, the reconstructed point cloud's dress will have pixels corresponding to the face, and conversely, the reconstructed and rendered point cloud's face will have pixels corresponding to the dress. This shift can result in noticeable artifacts that degrade the visual quality of the point cloud.

[0132] To reduce artifacts, points in the reconstructed 3D point cloud, represented as pixels near patch boundaries in a 2D frame, can be smoothed. To reduce the occurrence or appearance of visible artifacts and improve compression efficiency, smoothing can be applied to the location of points in the point cloud, each identified attribute of the point cloud (such as color, reflectivity, etc.), or both the geometry and attributes of the point cloud.

[0133] To smooth the geometry, attributes, or both of the 3D point cloud, boundary detection engine 558 identifies certain points in the 3D point cloud represented as boundary points. These boundary points of the 3D point cloud correspond to pixels in the reconstructed geometry frame 516a (or reconstructed attribute frame 520a), and these pixels are placed at or near the boundaries of each patch (based on the corresponding pixels in the reconstructed occupancy map frame 518a). Boundary detection engine 558 identifies the boundary points of the reconstructed point cloud based on the values ​​of pixels within the reconstructed occupancy map frame 518a. For example, boundary detection engine 558 identifies boundary pixels within the reconstructed occupancy map frame 518a.

[0134] To locate boundary pixels, the reconstructed occupancy frame 518a examines the values ​​of the pixels within it to identify valid pixels adjacent to invalid pixels. For example, the boundary detection engine 558 identifies a pixel with a value of 1 that is adjacent to a pixel with a value of 0. An adjacent pixel can be one of eight neighboring pixels (unless that pixel is located on the boundary of the occupancy frame 518a).

[0135] To identify boundary points in the point cloud, boundary detection engine 558 examines pixels within the reconstructed occupancy frame 518a. Boundary detection engine 558 examines the reconstructed occupancy frame 518a to identify a subset of pixels that are valid (value-based) but adjacent (neighboring) to invalid (value-based) pixels. For example, boundary detection engine 558 examines each pixel within occupancy frame 518. The examination includes selecting a query pixel and identifying whether the query pixel is valid based on its value. If the query pixel is invalid, boundary detection engine 558 continues selecting new pixels within occupancy frame 518 until a valid query pixel is identified.

[0136] In some embodiments, the boundary detection engine 558 identifies boundary pixels based on valid pixels being within a threshold distance from invalid pixels. For example, this distance may include pixels within one pixel of the query pixel, pixels within two pixels of the query pixel, pixels within three pixels of the query pixel, etc. As the distance increases, the number of identified boundary points also increases.

[0137] When identifying boundary pixels, the boundary detection engine 558 identifies the corresponding pixel in the reconstructed geometry frame 516a, which is placed at the same location as the boundary pixel in the (occupancy map). Subsequently, the boundary detection engine 558 identifies points in the 3D point cloud corresponding to the corresponding pixel in the reconstructed geometry frame 516a. In other words, based on the correspondence between pixels in the reconstructed occupancy map frame 518a and pixels at the same location in the reconstructed geometry frame 516a, there is a correspondence between the boundary pixels identified in the reconstructed occupancy map frame 518a and the points in 3D space.

[0138] In addition to identifying the boundary points of the reconstructed point cloud, decoder 550 also divides the reconstructed point cloud into a 3D mesh. The 3D mesh consists of multiple non-overlapping cells. The reconstructed point cloud lies within the 3D mesh, such that the points of the 3D point cloud are located within the entire cells of the 3D mesh. For example, decoder 550 generates a 3D mesh around reconstructed point cloud 602. The shape and size of the cells can be uniform throughout the mesh or vary between cells. Figure 7B An example grid is shown.

[0139] Then, decoder 550 identifies specific cells of the mesh that include the identified boundary points. That is, for each boundary point, it identifies the corresponding cell of the 3D mesh. The decoder also identifies certain neighboring cells adjacent to each cell that includes the identified boundary points.

[0140] For example, for a query cell with a single boundary point and multiple other points, the decoder identifies seven other cells adjacent to the query cell. Note that a single cell has twenty-six neighboring (adjacent) cells. Decoder 550 selects a predetermined number of neighboring cells that are geometrically closest to the boundary point within the query cell. For example, if the boundary point is in the lower left portion of the cell, decoder 550 selects neighboring cells located to the left and below the current cell. In some embodiments, the predetermined number of cells adjacent to the cell with the query point is seven neighboring cells. Neighboring cells and the query cell including the boundary point are represented as boundary cells. Similarly, in some embodiments, there are a total of eight boundary cells (seven neighboring cells and the query cell). In other embodiments, a different number of cells may be identified as boundary cells.

[0141] After selecting a predetermined number of boundary cells, the decoder 550 identifies the centroid of each boundary cell (including the query cell containing the boundary point and neighboring cells based on the position of the boundary point within the query cell). The centroid of each boundary cell is based on the points included in each corresponding cell. The centroid of each boundary cell is stored in a lookup table.

[0142] After identifying the boundary points, boundary cells, and centroids of the reconstructed point cloud, the smoothing engine 560 determines whether to perform smoothing relative to each identified boundary point. In some embodiments, the smoothing engine 560 performs attribute smoothing differently from geometric smoothing.

[0143] For example, for geometric smoothing, the smoothing engine 560 performs trilinear filtering of the centroids. The trilinear filter is applied to the centroid of each boundary cell (including the cell containing the boundary point and a predetermined number of neighboring cells) to find smooth geometry for each boundary point. The smoothing engine 560 determines whether the distance between a single boundary point and the filter output is greater than a threshold. When the distance between a single boundary point and the filter output is greater than the threshold, the smoothing engine 560 replaces the boundary point's value with the output of the trilinear filter. Optionally, when the distance between a single boundary point and the filter output is less than the threshold, the smoothing engine 560 determines that geometric smoothing is not needed for the boundary point.

[0144] For attribute smoothing, the smoothing engine 560 applies a set of criteria to determine whether to apply attribute smoothing. First, when the attribute type is texture, attribute smoothing is applied to one of the reconstructed attribute frames 520a. That is, if no reconstructed attribute frame 520a is a texture, the smoothing engine 560 determines that no attribute smoothing is applied. Second, attribute smoothing is applied to the frame when the number of components in the reconstructed attribute frame 520a corresponding to the attribute type of texture is one or three. When the frame is monochrome (such as black and white), the number of components in the frame is one. When the frame is in RGB or YUV or another tri-color stimulus color space, the number of components in the frame is three. Third, attribute smoothing is applied to the reconstructed attribute frame 520a when a flag (or SEI message or syntax element) enabling specified attribute smoothing is identified. That is, if one of the reconstructed attribute frames 520a represents a texture, including a single color component or three color components, and a flag (or syntax element) enabling specified attribute smoothing is identified, the smoothing engine 560 determines that any attribute smoothing is applied. If none of the three criteria are present, the Smoothing Engine 560 determines not to perform any property smoothing.

[0145] When the smoothing engine 560 determines to perform attribute smoothing, it uses the first component (the zeroth (0th) component) of the decoded color video frame (the portion reconstructing attribute frame 520a) to make a color smoothing decision. When the decoded color video frame is in YUV format, the first component (the zeroth (0th) component) will be the Y component. When the decoded color video frame is in RGB format, the first component (the zeroth (0th) component) can be the G component. This is because when directly compressing RGB format video without converting to YUV or YCbCr format, the color components are typically ordered as green-blue-red (GBR), since the green component is very similar to the luminance (Y) component of YUV format video, and most video codecs are optimized for YUV or YCbCr formats.

[0146] For example, when the decoded frame is in RGB 444 (8-bit) format, the smoothing engine 560 performs attribute smoothing using a smoothing decision based on the zeroth (0th) component (which is typically the G component). Alternatively, when the decoded frame is in YUV 420 (8-bit) format, the smoothing engine 560 first performs chroma upsampling to generate a YUV 444 format frame. To maintain precision, the YUV 444 format frame can be stored with 16-bit precision. The smoothing engine 560 then performs attribute smoothing using a smoothing decision based on the zeroth (0th) component (the Y component). After performing smoothing based on the Y component, the smoothing engine 560 then converts the YUV 444 (16-bit) format to RGB 444 (8-bit) format. Depending on the initial format of the associated attribute frame to be smoothed, including information, the smoothing engine 560 performs smoothing slightly differently.

[0147] To perform smoothing, based on either RGB 444 or YUV 444 format, the smoothing engine 560 determines whether to exclude certain identified boundary cells from attribute smoothing based on brightness variations. For example, the smoothing engine 560 determines to exclude boundary cells with internal brightness variations exceeding a threshold. Furthermore, the smoothing engine 560 determines to exclude boundary cells whose neighboring boundary cells, including boundary points, have a brightness difference greater than another threshold. That is, if the brightness variation in a cell is greater than a threshold, that cell is excluded from smoothing. To measure the brightness variation in a cell, the difference between the median and average brightness values ​​is compared to a threshold. If this difference is greater than the threshold, that cell is excluded from smoothing. If the difference between the color centroid of the current cell (containing boundary points) and the color centroids of its neighboring cells is greater than a threshold, that neighboring cell is excluded from smoothing.

[0148] The change in brightness within a cell is represented by the difference between the median and the average brightness. Since the zeroth (0th) component is used to make the decision, the median of the zeroth (0th) component of the points in the cell and the average of the zeroth (0th) component of the points in the cell are used.

[0149] After identifying the cells to be used for color smoothing, the smoothing engine 560 applies a trilinear filter to the color centroids of the cell containing the boundary point and the remaining neighboring cells. If the difference between the boundary point's color and the derived smoothed color is greater than a threshold, the query point's color is replaced with the smoothed color. Optionally, if the difference between the boundary point's color and the derived smoothed color is less than a threshold, color smoothing is not performed on the boundary point.

[0150] After performing geometric smoothing and / or property smoothing (or determining not to perform property smoothing), the decoder 550 can render the reconstructed point cloud 564.

[0151] After the reconstruction engine 556 reconstructs the point cloud and the smoothing engine 560 determines whether to perform geometric and attribute smoothing (and based on the determination to perform smoothing to remove artifacts unintentionally generated during frame encoding and decoding), the decoder 550 renders the reconstructed point cloud 564. The reconstructed point cloud 564 is rendered and displayed on a monitor or head-mounted display, similar to... Figure 1 HMD116. The reconstructed point cloud 564 is similar to the 3D point cloud 512.

[0152] although Figure 5A The environment architecture 500 is shown. Figure 5B Encoder 510 is shown, and Figure 5C The decoder 550 is shown, but it is compatible with... Figure 5A , Figure 5B and Figure 5C Various changes can be made. For example, the Environment-Architecture 500 can include any number of encoders or decoders.

[0153] Figure 6A and Figure 6B An example method for property smoothing according to embodiments of this disclosure is shown. Figure 6A In the middle, the decoder 550 receives attribute video frames in YUV 420 format, while... Figure 6B In the process, the decoder 550 receives attribute video frames in RGB 444 format.

[0154] Figure 6A and Figure 6B The method can be derived from Figure 1 Any one of server 104 or client devices 106-116 Figure 2 Server 200 Figure 3 Electronic devices 300, Figure 5A and Figure 5B encoder 510, Figure 5A and Figure 5C The decoder 550 or any other suitable device or system shall be used to perform this. For ease of illustration, Figure 6A and 6B The method is described as being by Figure 5A and Figure 5C The boundary decoder 550 is executed.

[0155] Different aspects of the encoded and decoded point cloud can be represented on different sets of attribute frames (such as attribute frame 520). For example, one set of attribute frames may represent texture (color), another set may represent reflectivity, an additional set may represent transparency, yet another set may represent normals, and so on. In addition to geometry frames and occupancy map frames, one or more of these sets of attribute frames can be generated by encoder 510 and included in the bitstream. Furthermore, when a set of attribute frames represents texture (color), these frames can be directly input into the video codec in RGB format or encoded in YUV format. Smoothing is based on various standards and the format in which the attribute frames are encoded. As described below, the terms color smoothing and texture smoothing are used interchangeably.

[0156] Decoder 550 determines whether to perform smoothing on a set of attribute frames when the following conditions (criteria) are met. That is, before performing attribute smoothing, decoder 550 determines whether color smoothing is needed based on three criteria. First, decoder 550 determines whether to perform color smoothing based on the type of attribute frames. Since different types of attributes (such as transparency, normals, etc.) can be used to represent point clouds on frames, color smoothing is performed on attribute videos (which consist of a set of attribute frames) when the attribute frame type of the attribute video is texture or color. That is, decoder 550 does not perform color smoothing on any frames of attribute videos that are not identified as texture or color. For example, if the only attribute in the attribute video is normals, color smoothing is not needed. Second, decoder 550 determines whether to perform color smoothing based on the number of color components in the attribute frames. Decoder 550 identifies the number of color components in a set of frames representing texture. The number of color components in an attribute frame representing texture is 1 or 3 for color smoothing to be performed. Finally, decoder 550 determines whether to perform color smoothing based on the presence of flags (syntax elements) in the bitstream. Therefore, after the decoder 550 determines that the attribute type is texture, the number of color components (attribute dimension) is one or three, and enables the flag indicating smoothness (syntax element), the decoder 550 executes method 600 or 650.

[0157] exist Figure 6AIn method 600, decoder 550 receives attribute video encoded in YUV 420 format. In step 602, decoder 550 identifies the frame type as YUV 420 (8 bits). In step 604, the chroma components of the frame are upsampled to make the video in YUV 444 format. The YUV 444 frame may be stored as 16 bits to maintain higher precision before conversion to RGB format. In step 606, decoder 550 performs attribute smoothing to make a smoothing decision by using the Y component (because the Y component is the zeroth component). In step 608, the decoder then converts the smoothed YUV 444 frame to RGB 444 frame (8 bits). The RGB444 frame is used to render the point cloud. In some embodiments, the original texture may be 10 bits or 12 bits instead of 8 bits, in which case the YUV 444 to RGB 444 conversion will produce 10-bit RGB 444 or 12-bit RGB 444 frames, respectively.

[0158] exist Figure 6B In method 650, decoder 550 receives an attribute frame. In step 652, decoder 550 identifies the frame type as RGB 444 (8 bits). In step 654, decoder 550 performs attribute smoothing using the zeroth (0th) component, which can be the green component (G), because smoothing decisions are typically made by directly encoding the RGB format by reordering the color components to GBR (since the G component is closest to the luminance (Y) component of a YUV or YCbCr format frame). If, instead of RGB 444 format, the attribute video is encoded directly using a video codec for another tristimulus color format, then... Figure 6B Method 650 will be applicable.

[0159] although Figure 6A and 6B This shows an example of performing property smoothing, but it is possible to perform property smoothing on other aspects. Figure 6A and Figure 6B Various changes were made. For example, although it is shown as a series of steps, Figure 6A and Figure 6B The steps in the process can overlap, occur in parallel, or occur any number of times.

[0160] Figures 7A-7C The processing for smoothing boundary points is described. Figure 7A An example method 700 for selecting certain centroids used to perform smoothing is shown according to an embodiment of the present disclosure. Figure 7B An example grid 710a and cells are shown according to an embodiment of this disclosure. Figure 7C Example 3D cells 712a and 712b of a grid 710a according to an embodiment of the present disclosure are shown. Method 700 may be derived from... Figure 1Any one of server 104 or client devices 106-116 Figure 2 Server 200 Figure 3 Electronic devices 300, Figure 5A and Figure 5C The method is executed by decoder 550 or any other suitable device or system. For ease of illustration, method 700 is described as being executed by decoder 550.

[0161] Decoder 550 reconstructs point cloud 702 from decoded frames (such as decoded geometry frame 516a, decoded occupancy map frame 518a, and decoded texture frame 520a). The reconstructed point cloud 702 can be... Figure 5C The reconstruction engine 556 generates the data. For example, the reconstructed point cloud 702 is similar to... Figure 5C The reconstructed point cloud is 564, without smoothing.

[0162] At step 710, decoder 550 divides the reconstructed point cloud 702 into a 3D mesh. The 3D mesh consists of multiple 3D cells. The shape and size of the cells can be uniform throughout the mesh or vary between cells. For example, as... Figure 7B As shown, grid 710a consists of multiple cells (such as cell 712). As illustrated, grid 710a consists of 1,000 cells (10 cells in height, 10 cells in width, and 10 cells in length) with uniform size and shape. In other embodiments (not shown), any number of cells can be used, and the cells can be of any shape. Grid 710a is resized such that the reconstructed point cloud 702 is placed within grid 710a. Similarly, the points of the reconstructed point cloud 702 are included in the various cells of grid 710a, while the other cells of the grid remain empty.

[0163] In step 720, decoder 550 identifies the boundary points of the reconstructed point cloud 702. Step 720 can be performed by... Figure 5C The boundary detection engine 558 performs the operation. To identify boundary points, the decoder 550 identifies points within the reconstructed point cloud 702 that are represented as pixels placed on the boundary of a piece within the sheet in geometry frame 516a or texture frame 520a. In some embodiments, the decoder 550 identifies points within the reconstructed point cloud 702 that are represented as pixels placed near the boundary of a piece within the sheet in geometry frame 516a or texture frame 520a based on the values ​​of corresponding pixels at the same location in the occupies frame 518a.

[0164] There is a correspondence between pixels in the reconstructed occupancy frame 518a and points in the reconstructed 3D point cloud 702. That is, for each valid pixel in the reconstructed occupancy frame 518a, there exists a corresponding point in the reconstructed 3D point cloud 702. For example, when a pixel at position (u, v) in the reconstructed occupancy frame 518a is valid, there exists a corresponding pixel at the same position (u, v) in the reconstructed geometry frame 516a with geometric data of the point in the reconstructed 3D point cloud 702. Based on the correspondence between pixels in the reconstructed occupancy frame 518a and points in the reconstructed 3D point cloud 702, the decoder 550 examines the points in the reconstructed 3D point cloud 702 to identify each boundary point. When (i) a point in the reconstructed point cloud 702 corresponds to a valid pixel in the reconstructed occupancy frame 518a and (ii) the valid pixel in the reconstructed occupancy frame 518a is within a predetermined distance from an invalid pixel in the same frame, that point in the reconstructed point cloud 702 is identified as a boundary point. The predetermined distance can be one or more pixels that separate the valid pixels (which correspond to points in the point cloud) of the reconstructed occupied frame 518a from the invalid pixels of the reconstructed occupied frame 518a.

[0165] In step 730, decoder 550 identifies the boundary cells associated with each identified boundary point (from step 720). There are two types of boundary cells. The first type of boundary cell includes the boundary point (as identified in step 720). The second type of boundary cell corresponds to a predefined number of cells adjacent to the first type of boundary cell (the cell containing the boundary point). Similarly, in step 732, decoder 550 identifies the cell containing the boundary point (the first type of boundary cell). Then, in step 734, decoder 550 identifies the location of the boundary point within the cell so that neighboring cells can be identified in step 736.

[0166] For example, Figure 7C Example 3D cells 712a and 712b are shown. 3D cells 712a and 712b can represent data from... Figure 7B Cell 712 of the grid. For example, cell 712a shows Figure 7B The internal view of cell 712, while cell 712b shows Figure 7B The external view of cell 712. Cells 712a and 712b may be referred to as cell 712. A single cell (such as cell 712) is divided into eight parts or sub-cells (top-left back, top-right back, bottom-left back, bottom-right back, top-left front, top-right front, bottom-left front, and bottom-right front) represented as parts 770 to 777. One or more points of the 3D point cloud may be located throughout the entire internal structure of cell 712. Any point located within cell 712 may be a boundary point as identified in step 732.

[0167] When a boundary point is located within cell 712, cell 712 and its seven neighboring cells are represented as boundary cells. For example, when the boundary point is located within portion 776 (the bottom-left front sub-cell) of cell 712, the boundary cells will include cell 712 and the seven cells closest to (nearest) portion 776. That is, for a boundary point located within portion 776, the boundary cells will be cell 712, and the cells placed to the left, below, bottom-left, in front, front-left, bottom-front, and bottom-left front of cell 712. In other words, the seven adjacent cells may include (i) the cells directly below cell 712 (such as the cells that touch portions 774, 775, 776 and 777), (ii) the cells directly in front of cell 712 (such as the cells that touch portions 772, 773, 775 and 776), (iii) the cells directly to the left of cell 712 (such as the cells that touch portions 770, 772, 774 and 776), (iv) the cells in front of and below cell 712 (such as the cells that touch portions 775 and 776), (v) the cells to the left and below cell 712 (such as the cells that touch portions 774 and 776), (vi) the cells in front of and to the left of cell 712 (such as the cells that touch portions 772 and 776), and (vii) the cell at the lower left corner of portion 776 that touches cell 712.

[0168] Based on the position of the boundary point within the boundary cell, in step 736, the decoder 550 identifies a certain number of neighboring cells. A neighboring cell is a cell that touches another cell. Note that there are 26 cells that are neighboring (touching) a single cell. Similarly, the second type of boundary cell is the cell closest to the boundary point within the cell identified in step 732 (the first type of boundary cell). Note that there is a predetermined number of cells that are neighboring cells including the boundary point. In some embodiments, there are a total of 8 boundary cells for a single boundary point. For example, for a single boundary point, there is the boundary cell to which the boundary point belongs and seven neighboring cells. The seven neighboring cells are selected from 26 possible neighboring cells based on the position of the boundary point within the cell. That is, in addition to the cell including the boundary point itself, the seven neighboring cells closest to the boundary point are selected as boundary cells. In other words, depending on the position of the identified boundary point within the first type of boundary cell, only a certain number of second type neighboring cells exist for smoothing.

[0169] The following grammar (1) describes the process of identifying boundary cells in step 730. The following variables are used as inputs to grammar (1) for identifying boundary cells, numCells1D, pointCnt, recPcGeo, isBoundaryPoint, the current point location (pointGeom[k], k = 0 to 2, including the endpoints), and an array containing boolean flags (CellDoSmoothing[x][y][z], for all x, y, and z = 0 to numCells1D - 1, including the endpoints). A boolean value represented as otherClusterPtCnt is one of the outputs of grammar (1). Grammar (1) also outputs an array that contains the upper left corner of a 2×2×2 grid, s[i], where i = 0 to 2, including the endpoints. An array t[i] that contains the 2×2×2 grid location associated with the current position, where i = 0 to 2, including the endpoints.

[0170]

[0171] Grammar (1)

[0172]

[0173]

[0174] After identifying the boundary cells, the decoder 550 generates a lookup table (step 740). The lookup table associates the identified boundary cells with indices representing the cells associated with the reconstructed 3D point cloud. As discussed in more detail below Figure 8A and Figure 8B show an example of generating a lookup table that associates cell indices with boundary cells.

[0175] Grammar (2) describes the generation of the index lookup table. As used in grammar (2), the expressions numCells1D, gridSize, and pointCnt, and recPcGeo[n][k] are inputs. The expression numCells1D is the number of cells in the x, y, or z direction. The expression gridSize is the size of the cells in the x, y, or z direction. The expression pointCnt is the number of reconstructed points. Note that recPcGeo[n][k], n = 0 and PointCnt - 1, k = 0..2.

[0176] Grammar (2) generates an array of cells called cellIdxLut[x][y][z], 0 ≤ x, y, z < numCells1D. Grammar (2) also generates an expression numBoundaryCells for the number of cells related to grid geometry smoothing or attribute smoothing. Note, Figure 7BAll cells in grid 710a are independent of performing geometric or property smoothing. Therefore, decoder 550 generates a lookup table from 3D cell indices to a smaller 1D cell structure storing boundary cells, since boundary cells are the only cells associated with either mesh geometric or property smoothing. An array cellIdxLut of size numCells×numCells×numCells is initialized to -1, and the variable numBoundaryCells is initialized to 0.

[0177]

[0178] Grammar (2)

[0179]

[0180]

[0181] In step 750, the decoder identifies the centroid of each boundary cell. When Figure 5C When the smoothing engine 560 performs geometric smoothing, the centroid can be the geometric centroid. When Figure 5C When the smoothing engine 560 performs texture (color) smoothing, the centroid can be a color centroid. Syntax (3)-Syntax (6), as described below, discusses the identification of the geometric center mesh corresponding to the geometric centroid. Syntax (6)-Syntax (9), as described below, discusses the identification of the color center mesh corresponding to the color centroid.

[0182] To identify the geometric center grid, the expressions gridSize and pointCnt are inputs. gridSize is the size of the geometric grid, and pointCnt is the number of points in the currently reconstructed point cloud frame. For example, the 3D geometric coordinate space is divided into cells of size gridSize × gridSize × gridSize.

[0183] The expression `isBoundaryPoint[n]`, where n = 0..pointCnt-1, is also input and indicates whether the nth reconstructed point is located on or near a patch boundary. Similarly, `recPcGeo[i][k]`, where i = 0..pointCnt-1 and k = 0..2, is another input describing an array containing the locations of reconstructed points. Another array, represented by the expression `pointToPatch[n]`, where n = 0..pointCnt-1, includes the index of the patch to which the nth reconstructed point `recPcGeo[n]` belongs.

[0184] To identify the geometric center grid, decoder 550 identifies the output numCells1D as the number of cells in the x, y, or z direction. The 3D geometric coordinate space is divided into cells of size based on the expression gridSize defined above. The following syntax (3) describes the processing of identifying the expression numCells1D as the number of cells in the x, y, or z direction.

[0185]

[0186] Grammar (3)

[0187]

[0188] Additionally, to identify the geometric center grid, the decoder 550 also identifies the output `numBoundaryCells` as the number of cells containing boundary points or adjacent cells. The expression `cellIdxLut[x][y][z]`, for x, y, and z in the range 0 to `numCells1D-1` (inclusive), is the output of an array mapping 3D cell indices to 1D indices in the range 0..numBoundaryCells-1. The expression `cellCnt[n]` (n = 0..numBoundaryCells-1) is an array containing the number of reconstructed points in the cell. The expression `cellDoSmoothing[n]`, n = 0..numBoundaryCells-1, is a Boolean array indicating whether smoothing should be performed on the cells.

[0189] To generate the index lookup table, the decoder 550 uses the expressions numCells1D, gridSize, pointCnt, isBoundaryPoint[n], and recPcGeo[i][k] as input, and the arrays cellIdxLut[x][y][z] and numBoundaryCells as output. The arrays cellDoSmoothing[n], cellCnt[n], cellPatchIdx[n], and centerGrid[n][k] are initialized to 0 for all x, y, and z in the range of 0 to (numBoundaryCells-1) (inclusive) and k in the range of 0 to 2 (inclusive). Syntax (4) is applied to n = 0 and pointCnt-1.

[0190]

[0191] Grammar (4)

[0192] xIdx=recPcGeo[n][0] / gridSize

[0193] yIdx=recPcGeo[n][1] / gridSize

[0194] zIdx=recPcGeo[n][2] / gridSize

[0195] cellIndex=cellIdxLut[xIdx+yIdx*numCells1D+zIdx*numCells1D*numCells1D]

[0196] When cellIndex is not equal to -1, the following applies. If cellCnt[cellIndex] equals 0, then cellPatchIdx[cellIndex] is set to the slice index of the slice to which the current reconstructed point recPcGeo[n] belongs. Otherwise, if cellDoSmoothing[cellIndex] equals 0 and cellPatchIdx[cellIndex] is not equal to the slice index to which the current point recPcGeom[n] belongs, then cellDoSmoothing[cellIndex] is set to 1. For example, when cellIndex is not equal to -1, syntax (5) is applied to generate the geometric mesh. Syntax (6) is used to identify the centroid of each cell with a non-zero count.

[0197]

[0198] Grammar (5)

[0199] for(k=0;k<3;k++)

[0200] centerGrid[cellIndex][k]+=recPcGeom[n][k]

[0201] cellCnt[cellIndex]++

[0202] Grammar (6)

[0203] if(cellCnt[n]>0)

[0204] for(k=0;k<3;k++)

[0205] centerGrid[n][k]=centerGrid[n][k]÷cellCnt[n]

[0206] To identify the attribute center grid, the inputs are processed as follows. The first input is an expression numComps representing the number of attribute components. The second input is the attribute index, represented as aIdx. The input also includes an array RecPcGeom[i] of reconstructed locations, where i is in the range of 0 to PointCnt-1 (inclusive). Another input is an array containing reconstructed attributes RecPcAttr[aIdx][i], where i is in the range of 0 to PointCnt-1 (inclusive).

[0207] The output of this process includes an array containing the reconstructed center grid attribute values ​​attCenterGrid[i][k], where i is in the range of 0 to numCells1D-1 (inclusive of endpoints) and k is in the range of 0 to numComps-1 (inclusive of endpoints). The output also includes an array containing the reconstructed center grid average brightness value meanLuma[k], where k is in the range of 0 to numComps-1 (inclusive of endpoints), and an array containing the reconstructed center grid brightness median attrMean[i][k], where i is in the range of 0 to numCells1D-1 (inclusive of endpoints) and k is in the range of 0 to numComps-1 (inclusive of endpoints).

[0208] To identify the attribute center grid, the elements of the arrays attrCenterGrid[x][y][z][m] and meanLuma[x][y][z] are initialized to 0 for all x, y, and z in the range 0 to numCells1D-1 (inclusive) and m in the range 0 to numComps-1 (inclusive). For i = 0, PointCnt-1, syntax (7) describes the processing of the identification variables xIdx, yIdx, and zIdx.

[0209]

[0210] Grammar (7)

[0211] xIdx=(RecPcGeom[i][0] / GridSize)

[0212] yIdx=(RecPcGeom[i][1] / GridSize)

[0213] zIdx=(RecPcGeom[i][2] / GridSize)

[0214] If cellDoSmoothing[xIdx][yIdx][zIdx] equals 1, then syntax (8) is applied. If cellCnt[x][y][z] is greater than 0 (inclusive) for x, y, and z in the range 0 to numCells1D-1, then for points belonging to that cell, the mean and median brightness values ​​of the attribute with index aIdx are identified, for example, the expression meanLuma[xIdx][yIdx][zIdx]+=RecPcAttr[aIdx][i][0]. Syntax (9) describes the decoder 550 for identifying the attribute center grid.

[0215]

[0216] Grammar (8)

[0217] for(k=0;k <numComps;k++)

[0218] attrCenterGrid[xIdx][yIdx][zIdx][k]+=RecPcAttr[aIdx][i][k]; syntax (9)

[0219] for(k=0;k <numComps;k++)

[0220] attrCenterGrid[xIdx][yIdx][zIdx][k]=

[0221] attrCenterGrid[xIdx][yIdx][zIdx][k](cellCnt[xIdx][yIdx][zIdx]

[0222] meanLuma[xIdx][yIdx][zIdx]=meanLuma[xIdx][yIdx][zIdx](cellCn t[xIdx][yIdx][zIdx]

[0223] After identifying the centroids of the boundary cells, in step 560a, decoder 550 performs smoothing. In some embodiments, decoder 550 performs geometric smoothing, followed by color smoothing.

[0224] although Figures 7A to 7C This shows an example of performing property smoothing, but it is possible to perform property smoothing on other aspects. Figures 7A to 7C Various changes were made. For example, although it was shown as a series of steps, Figure 7A The steps in method 700 can overlap, occur in parallel, or occur any number of times.

[0225] Figure 8AExample 800 is shown in the embodiment of the present disclosure for identifying and selecting certain centroids in 2D used to perform smoothing. Figure 8B A 3D example 850 is shown, according to an embodiment of the present disclosure, for identifying and selecting certain centroids used to perform smoothing. Examples 800 and 850 may be derived from... Figure 1 Any one of server 104 or client devices 106-116 Figure 2 Server 200 Figure 3 Electronic devices 300, Figure 5A and Figure 5C The decoder 550 or any other suitable device or system shall perform the operation. For ease of illustration, methods 800 and 850 are described as being performed by the decoder 550.

[0226] Figure 8A Example 800 describes identifying boundary cells using only two dimensions of a 3D mesh while generating a centroid table. Typically, the point cloud is first divided into a 3D mesh. Mesh 802 is a 2D mesh, but the concept can be extended for 3D meshes. For a 3D mesh, cells containing boundary points and their seven neighboring cells are identified. The indices of the boundary cells are stored in an index lookup table, while the centroids of the boundary cells are stored in a centroid table. The size of the index lookup table includes all cells, while the size of the centroid lookup table is equal to the number of boundary cells. The index lookup table is used to map the 3D indices of the boundary cells to the 1D indices of the centroid table.

[0227] As shown in Example 800, the 2D grid 802 includes points 804, 805, 806, 807, 808, 809, and 810 placed throughout the 2D grid 802. Each of points 804, 805, 806, 807, 808, 809, and 810 has a (x, y) position. Note that point 810 is a boundary point because it is adjacent to another invalid pixel when corresponding to a pixel in a valid occupied frame. The index lookup table 820 includes three columns: index 822, initial value 824, and modified value 826. Each row of index 822 corresponds to a specific cell. For example, index 0 corresponds to cell I0, which includes point 804. As another example, index 5 corresponds to cell I5, which includes boundary point 810. After identifying the boundary point, adjacent boundary cells I4, I8, and I9 are identified. Since boundary point 810 is located to the left of cell I5, cell I4 is identified as a boundary cell. Because boundary point 810 is located at the bottom of cell I5, cell I9 is ​​identified as a boundary cell. Because boundary point 810 is located at the bottom left corner of cell I5, cell I8 is identified as a boundary cell. Note that even cells I0, I1, I2, I6, and I... 10 Adjacent to and touching cell I5, which contains boundary point 810, cells I0, I1, I2, I6, and I10 It is also not recognized as a boundary cell because, compared to cells I0, I1, I2, I6, and I... 10 The boundary point 810 is closer to cells I4, I8, and I9.

[0228] As shown in the figure, the index lookup table 820 has the same number of entries as the number of cells in grid 802. As shown in the figure, grid 802 contains 16 cells (cells I0 to I...). 15 The index lookup table 820 has 16 rows, with each row corresponding to a specific cell in grid 802. For example, index 0 of index 822 corresponds to cell I0. The initial value of column 824 in index lookup table 820 is set to -1 for each cell. For each boundary cell, the modified value 826 is incremented by 1.

[0229] For example, for cell I4, the value increases from -1 to 0. For cell I5, the value increases from 0 to 1. For cell I8, the value increases from 1 to 2. For cell I9, the value increases from 2 to 3. These modified values ​​correspond to index 832 of the centroid table 830. The centroid table 830 associates the identified boundary cells with their corresponding centroids. For example, the decoder 550 first identifies the boundary cells and then derives the centroid values ​​of the identified boundary cells. Therefore, there exists a correspondence between index 822 and value 834 based on the modified value 826. After identifying the centroids, the decoder 550 can determine whether to perform smoothing on the boundary point 810. For example, the decoder 550 can compare the centroids of cells I4, I8, and I9 with the centroid of cell I5. If smoothing is to be performed, the decoder 550 can perform a trilinear filter on the centroids of cells I4, I8, and I9 and replace the boundary point 810 with the output value of the trilinear filter.

[0230] The following parameters are used Figure 8B Example 850. For Figure 8B Example 850 has a point cloud dimension of 1024*1024*1024. The grid consists of uniformly spaced 8*8*8 cells. Therefore, based on the size of the point cloud and the size of each cell, there can be 2,097,152 cells within the grid (because of 128*128*128). First, lookup table 852 is initialized to -1. That is, the cell index value from index 0 to the maximum number of cells in the grid is initialized to the value of -1.

[0231] Decoder 550 identifies the boundary point located at (66, 87, 46). Since each cell is 8*8*8 in size, the identified boundary point will be located in cell (8, 10, 5) (due to (int[66 / 8], int[87 / 8], int[46 / 8])). The decoder then identifies that this boundary point is located in the lower left front portion of the cell, since 66%8 = 2 < 4, 87%8 = 7 > 4, and 46%8 = 5 > 4. Therefore, the seven neighboring cells will be: (i) the cell to the left (7, 10, 5), (ii) the cell below (8, 11, 5), (iii) the cell to the lower left (7, 11, 5), (iv) the cell in front (8, 10, 6), (v) the cell to the left front (7, 10, 6), (vi) the cell below (8, 11, 6), and (vii) the cell to the lower left (7, 11, 6). Then, decoder 550 converts the 3D index of the cell into a one-dimensional (1D) index. Equation (3) describes the process used to convert the 3D index (represented by the X, Y, Z coordinates) into a 1D index. For a cell located at (ix, iy, iz), the 1D index is represented as CellId.

[0232] CellId=ix+(iy*numCells1D)+(iz*numCells1D*numcells1D) Equation (3)

[0233] The expression `numCells1D` is based on the cell size. In this example, `numCells1D` is 128. Therefore, since (8 + (10 * 128) + (5 * 128 * 128) = 83208, the cell ID of the cell including the boundary point is 83208. Similarly, the 1D indices of the boundary cells will be 83208, 83207, 83336, 83335, 99592, 99591, 99720, and 99719, respectively.

[0234] For each subsequent boundary cell, the expression `numBoundaryCell` is incremented by an integer 1. For example, lookup table 854 shows the lookup table after the first boundary cell, represented by index 83208, is set to 0, after which `numBoundaryCell` is set to 1. Lookup table 856 shows the lookup table after the second boundary cell, represented by index 83207, is set to 1, after which `numBoundaryCell` is set to 2. Lookup table 858 shows the lookup table after the third boundary cell, represented by index 83336, is set to 2, after which `numBoundaryCell` is set to 3. Based on the value of the expression `numBoundaryCell` incremented for each subsequent boundary cell, the process continues to set the value in the lookup table for each boundary cell.

[0235] After modifying all index values ​​of the boundary cells, the decoder 550 identifies the centroid value of each cell whose index is not equal to -1. Then, the centroid value of each cell is stored in a centroid table, similar to... Figure 8A The centroid table 830. In some embodiments, the centroid may be a geometric centroid stored in a geometric centroid table or an attribute centroid (such as color or texture) stored in a corresponding attribute table.

[0236] although Figure 8A and Figure 8B An example of generating a lookup table is shown, but it is possible to modify... Figure 8A and Figure 8B Make various changes. For example, you can include any number of boundary points in a cell.

[0237] Figure 9 An example method for decoding a point cloud is shown according to an embodiment of the present disclosure. Method 900 may be provided by... Figure 1 Any one of server 104 or client devices 106-116 Figure 2 Server 200 Figure 3 Electronic devices 300, Figure 5A and Figure 5C The decoder 550 or any other suitable device or system shall perform this. For ease of illustration, method 900 is described as being performed by... Figure 5A and Figure 5C The decoder 550 is executed.

[0238] Method 900 begins with a decoder (such as decoder 550) receiving a compressed bitstream (step 902). The received bitstream may include an encoded point cloud that has been mapped onto multiple 2-D frames, compressed, then transmitted, and ultimately received by decoder 550.

[0239] In step 904, decoder 550 decodes multiple frames from the bitstream. The multiple frames consist of pixels. In some embodiments, portions of the pixels within a frame are organized into patches corresponding to point clusters in a 3D point cloud. A frame may include at least one geometric frame and at least one attribute frame. For example, a geometric frame includes pixels, and portions of the pixels in the geometric frame represent the geometric positions of points in the 3D point cloud. The pixels of the geometric frame are organized into patches corresponding to the corresponding point clusters in the 3D point cloud. Similarly, an attribute frame includes pixels, and portions of the pixels in the attribute frame represent attribute information of points in the 3D point cloud, and the positions of the pixels in the attribute frame correspond to the corresponding positions of the pixels in the geometric frame.

[0240] Decoder 550 also decodes an occupancy map from the bitstream. An occupancy map indicates the portion of pixels that represent points in a 3D point cloud, encompassing multiple frames. For example, an occupancy map frame includes pixels that represent the geometric locations of points in a geometry frame. In other words, the value of a pixel at position (x, y) in an occupancy map frame indicates whether the pixel at the same (x, y) position in one of the multiple frames represents a point in the point cloud or is empty.

[0241] In step 906, decoder 550 reconstructs the 3D point cloud. In some embodiments, decoder 550 uses multiple frames and occupancy map frames to reconstruct the 3D point cloud.

[0242] In step 908, decoder 550 determines whether to perform smoothing on the 3D point cloud, at least in part, based on characteristics of multiple frames. For example, characteristics may indicate attribute type, number of components, and flags (or syntax elements) indicating whether attribute smoothing is enabled. Flags (or syntax elements) may be included in the bitstream. For example, attribute type may indicate whether the attribute frame is in RGB 444 format. As another example, characteristics may include whether the attribute frame is in YUV 420 format.

[0243] Based on the determined smoothing, decoder 550 performs smoothing on the 3D point cloud (step 910). For example, when the smoothing is attribute smoothing, and one of the multiple frames is an attribute frame with RGB 444 format, decoder 550 determines the smoothing to be performed based on the values ​​associated with the G component of the RGB frame.

[0244] For example, when the smoothing is attribute smoothing, and one of the multiple frames is an attribute frame in YUV 420 format, when decoder 550 determines to perform smoothing, decoder 550 first upscales the frame to YUV 444 format. Once the frame is in YUV 444 format, decoder 550 performs attribute smoothing based on the value associated with the Y of the YUV frame. Once smoothing is performed, decoder 550 converts the YUV 440 format frame to RGB 444 format.

[0245] To perform smoothing, decoder 550 generates a grid comprising multiple cells. The constructed point cloud is placed within the grid such that each point resides in a different cell. Decoder 550 identifies boundary points of the 3D point cloud based on multiple frames and an occupancy map. In some embodiments, to identify boundary points, decoder 550 first identifies a first pixel in an occupancy map frame whose value indicates that the first pixel is valid and adjacent to a second pixel whose value indicates that the second pixel is invalid. The first pixel may be represented as a boundary pixel. After identifying the boundary pixels, decoder 550 identifies the point in the 3D point cloud corresponding to the first pixel.

[0246] After identifying the boundary points of the 3D point cloud, decoder 550 identifies the boundary cells associated with each boundary point. Note that a boundary cell may include the boundary point or a neighboring cell containing the boundary point. A certain number of boundary cells neighboring the cell containing the boundary point are identified based on the position of the boundary point within the cell. For example, decoder 550 may identify a region in a first cell that includes the boundary point. Then, the decoder identifies a predetermined number of neighboring cells adjacent to the region in the first cell. The first cell and the predetermined number of neighboring cells are represented as boundary cells.

[0247] The decoder 550 derives the centroid of each identified boundary cell. Based on the identified centroids of the boundary cells associated with that particular boundary point, smoothing is performed on the boundary points.

[0248] Decoder 550 generates a first lookup table listing the 3D cells. In some embodiments, the first lookup table includes predefined entries for a plurality of 3D cells. After identifying boundary cells, decoder 550 generates a second lookup table including the centroid values ​​of the boundary cells. Decoder 550 modifies predetermined entry values ​​for the boundary cells included in the first lookup table to generate a correspondence between the boundary cells in the first lookup table and the boundary cells in the second lookup table. The size of the first lookup table may be based on the number of 3D cells within the generated mesh, while the size of the second lookup table is based on the number of identified boundary cells associated with boundary points.

[0249] although Figure 9 An example of a method 900 for decoding point clouds is shown, but it is possible to modify... Figure 9 Various changes were made. For example, although it was shown as a series of steps, Figure 9 The steps in the process can overlap, occur in parallel, or occur any number of times.

[0250] Although the accompanying drawings illustrate different examples of user equipment, various changes can be made to the drawings. For example, the user equipment may include any number of each component in any suitable arrangement. Generally, the drawings do not limit the scope of this disclosure to any particular configuration. Furthermore, while the drawings illustrate operating environments in which various user equipment features disclosed in this patent document can be used, these features can be used in any other suitable system. None of the descriptions in this application should be construed as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims.

[0251] Although this disclosure has been described using exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. This disclosure is intended to cover such changes and modifications that fall within the scope of the appended claims.

Claims

1. A decoding device for point cloud decoding, the decoding device comprising: The communication interface is configured to receive bit streams. as well as A processor is operatively coupled to the communication interface, wherein the processor is configured to: The bitstream is decoded into multiple frames comprising pixels, where the pixels are organized into patches and correspond to corresponding point clusters in a 3D point cloud. Occupancy frames are decoded from the bitstream, wherein the occupancy frames indicate portions of pixels representing points of the 3D point cloud within the plurality of frames. The 3D point cloud is reconstructed using the multiple frames and occupancy map frames. Determine the attribute type of the plurality of frames, the number of components associated with the attribute type, or whether at least one of the syntax elements is identified from the bitstream, wherein the syntax element indicates that attribute smoothing is enabled, and Based on the determination, attribute smoothing is performed on the 3D point cloud. Among these, one of the plurality of frames is an attribute frame; and To perform attribute smoothing, the processor is configured as follows: In response to determining that the attribute frame is in YUV 420 format, the attribute frame is enlarged to YUV 444 format. Determine the differences between values ​​associated with the zeroth component in YUV 444 format, where the zeroth component is Y. When the difference exceeds a threshold, attribute smoothing is performed, and After performing attribute smoothing, the YUV 444 format of the attribute frame is converted to RGB 444 format.

2. The decoding device according to claim 1, wherein, To perform attribute smoothing, the processor is configured as follows: In response to determining that the attribute frame is in RGB 444 format, a second difference is determined between the values ​​associated with the zeroth component of the RGB 444 format, wherein the zeroth component is green, and When the second difference is greater than the second threshold, attribute smoothing is performed.

3. The decoding device according to claim 1, wherein, The processor is configured to perform attribute smoothing in the following cases: The attribute type is texture. The quantity is one or three color components, or The syntax elements are identified from the bitstream.

4. The decoding device according to claim 1, wherein, The processor is also configured to: Generate a mesh comprising multiple 3D cells, wherein the 3D point cloud is located within the mesh. Based on the multiple frames and occupancy map frames, multiple boundary points of the 3D point cloud are identified. Identify boundary cells associated with the plurality of boundary points from the plurality of 3D cells. Derive the centroid value of the boundary cell. Determine the differences between the plurality of boundary points and the centroid values, and Based on the differences, smoothing is performed on the multiple boundary points.

5. The decoding device according to claim 4, wherein, In order to identify one of the plurality of boundary points, the processor is further configured to: Identify the first pixel in the occupied frame that is valid and adjacent to the invalid second pixel; as well as When the first point of the 3D point cloud is based on a pixel in one of the multiple frames that shares a position with the first pixel in the occupied frame, the first point is identified as a boundary point.

6. The decoding device according to claim 4, wherein, To identify boundary cells, the processor is also configured to: Identify the first cell among the plurality of 3D cells, which includes the boundary points among the plurality of boundary points; Identify the region within the first cell containing the boundary point; Identify a predetermined number of neighboring cells within the region that is adjacent to the first cell and closest to the boundary point; as well as The first cell and the neighboring cells are identified as boundary cells associated with the boundary point.

7. The decoding device according to claim 4, wherein, The processor is also configured to: Generate a first lookup table that lists the plurality of 3D cells and includes predetermined entry values ​​for the plurality of 3D cells; as well as After identifying the boundary cells, a second lookup table is generated that includes the centroid values ​​of the boundary cells.

8. A method for point cloud decoding, the method comprising: Receive bit stream; The bitstream is decoded into multiple frames that include pixels, where the pixels are organized into patches and correspond to the corresponding point clusters in a 3D point cloud. Occupancy frames are decoded from the bitstream, wherein the occupancy frames indicate portions of pixels representing points of the 3D point cloud that are included in the plurality of frames; The 3D point cloud is reconstructed using the multiple frames and occupancy map frames; Determine the attribute type of the plurality of frames, the number of components associated with the attribute type, or whether at least one of the syntax elements is identified from the bitstream, wherein the syntax element indicates that attribute smoothing is enabled; and Based on the determination, attribute smoothing is performed on the 3D point cloud. Among these, one of the plurality of frames is an attribute frame; and Execution property smoothing, including: In response to determining that the attribute frame is in YUV 420 format, the attribute frame is enlarged to YUV 444 format. Determine the differences between values ​​associated with the zeroth component in YUV 444 format, where the zeroth component is Y. When the difference exceeds a threshold, attribute smoothing is performed, and After performing attribute smoothing, the YUV 444 format of the attribute frame is converted to RGB 444 format.

9. The method according to claim 8, wherein, Execution property smoothing, including: In response to determining that the attribute frame is in RGB 444 format, a second difference is determined between the values ​​associated with the zeroth component of the RGB 444 format, wherein the zeroth component is green, and When the second difference is greater than the second threshold, attribute smoothing is performed.

10. The method according to claim 8, wherein, The method also includes performing attribute smoothing in the following cases: The attribute type is texture. The quantity is one or three color components, or The syntax elements are identified from the bitstream.

11. The method of claim 8, further comprising: Generate a mesh comprising multiple 3D cells, wherein the 3D point cloud is located within the mesh; Identify multiple boundary points of the 3D point cloud based on the multiple frames and occupancy map frames; Identify boundary cells associated with the plurality of boundary points from the plurality of 3D cells; Derive the centroid value of the boundary cell; Determine the differences between the plurality of boundary points and the centroid values, and Based on the differences, smoothing is performed on the multiple boundary points.

12. The method according to claim 11, wherein, Identifying one of the plurality of boundary points includes: Identify the first pixel in the occupied frame that is valid and adjacent to the invalid second pixel; and When the first point of the 3D point cloud is based on a pixel in one of the multiple frames that shares a position with the first pixel in the occupied frame, the first point is identified as a boundary point.

13. The method of claim 11, further comprising: Generate a first lookup table that lists the plurality of 3D cells and includes predetermined entry values ​​for the plurality of 3D cells; as well as After identifying the boundary cells, a second lookup table is generated that includes the centroid values ​​of the boundary cells.

Citation Information

Patent Citations

  • Point cloud compression

    US20190087979A1