Generate new views using point clouds

By interpolating scene geometry using connectivity information and adding geometric primitives to point clouds, the method addresses complexity and latency issues in immersive video systems, enabling efficient rendering of new views with enhanced image quality.

JP2025539556APending Publication Date: 2025-12-05KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025534247
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-14
Filing Date
2023-12-05
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing immersive video systems face challenges in generating new views with low complexity and/or latency due to the limited precision of geometry information and variable number of samples mapped to each scene point, leading to complex encoders that require significant computational and memory resources.

Method used

A method involving receiving an image and a point cloud, obtaining connectivity information, interpolating scene geometry, and rendering a view using the image and interpolated geometry, which includes adding geometric primitives to points in the point cloud to enhance spatial resolution and accuracy.

Benefits of technology

This approach enables efficient rendering of new views with improved image quality and reduced complexity by converting sparse point clouds into dense geometry, suitable for immersive video applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539556000001_ABST
    Figure 2025539556000001_ABST
Patent Text Reader

Abstract

A method for generating a view of a scene at a target viewpoint for immersive video includes receiving an image of the scene and a point cloud of the scene at a point cloud viewpoint, and obtaining connectivity information between points in the point cloud. The point cloud and the connectivity information are used to interpolate a geometry of the scene, and the image and the interpolated geometry are used to render a view of the scene at the target viewpoint.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to generating new views at a target viewpoint using point clouds as depth information. [Background technology]

[0002] Immersive video is video played, for example, in a virtual reality (VR) headset, where the viewer has limited freedom to look around and move around within the scene. This is also known as Free View, 6 degrees of freedom (6DoF), or 3DoF+.

[0003] Immersive video systems typically require a multi-view video camera that captures a 3D scene from various viewpoints. The captured video is provided with a captured or estimated depth map that allows scene elements to be reprojected onto the viewpoints of other (virtual) cameras, independent of the camera's viewpoint. Multi-view imaging therefore involves imaging a scene and retrieving the geometry of the captured scene using, for example, the depth map.

[0004] Rendering from coded multi-view representations is challenging due to the limited precision of the geometry information and the variable number of samples mapped to each scene point. 3D reconstruction, which allows the generation of images from virtual viewpoints, is typically performed using processes such as multi-view calibration, depth estimation, image filtering, and consistency checking, along with coding techniques such as pixel pruning and patch packing. These processing requirements result in complex encoders that require significant computational and memory resources and introduce latency.

[0005] Latency and processing complexity requirements depend on the particular application, so in some applications a low complexity encoder is desirable at low latency (but potentially at the expense of reduced image quality), while in other applications increased complexity and latency may be acceptable in order to improve image quality. Summary of the Invention [Problem to be solved by the invention]

[0006] However, there remains a need for image processing solutions with low complexity and / or latency for encoders and / or decoders. [Means for solving the problem]

[0007] The invention is defined by the claims.

[0008] According to an embodiment of the present invention, there is provided a method for generating a view of a scene at a target viewpoint for immersive video, the method comprising: receiving an image of a scene and a point cloud of the scene at a point cloud viewpoint; obtaining connectivity information between points in the point cloud; interpolating the geometry of the scene using the point cloud and connectivity information; and using the image and the interpolated geometry to render a view of the scene at the target viewpoint.

[0009] The step of obtaining the connectivity information may include receiving the connectivity information and / or determining / estimating the connectivity information.

[0010] Of course, multiple images may be received (e.g., more than two). Typically for immersive / multi-view video, multiple images are received and the colors / textures of the images can be blended (if appropriate) to render the view. Additionally, multiple point clouds can be received.

[0011] Point clouds typically provide sparse depth information of a scene and are therefore generally not particularly suitable for multi-view imaging, as they generally cannot provide both high pixel spatial resolution and acceptable temporal resolution. However, due to the sparse nature of point clouds (e.g., obtained from LiDAR), point clouds offer a low-bitrate solution for providing depth information through the bitstream.

[0012] Therefore, it is proposed to interpolate the scene geometry from the point cloud using connectivity information. The connectivity information provides the geometric relationships between points in the point cloud (e.g., points that correspond to the same surface, etc.). This enables accurate interpolation of the scene geometry. In other words, the method provides relatively accurate sparse-to-dense geometry interpolation at the decoder / client side.

[0013] Preferably, the captured images of the scene provide a different view of the scene than the rendered view of the scene at the target viewpoint. In this manner, rendering may include rendering or compositing images that provide a view of the scene where images were not previously available or generated.

[0014] If multiple images of a scene are acquired (and used in rendering), each image providing an image from a different viewpoint, the target viewpoint will be different from any of the viewpoints provided by the acquired images.

[0015] In general, connectivity information describes the spatial relationships between points in a point cloud.

[0016] The step of interpolating may include determining a geometric primitive for one or more points in the point cloud.

[0017] The interpolating step may include adding a geometric primitive to each of one or more points in the point cloud. The interpolating step may include, for each of one or more points in the point cloud, adding a geometric primitive to the point such that the point touches the added geometric primitive at a non-vertex location of the added geometric primitive. The interpolating step may include, for each of one or more points in the point cloud, adding a geometric primitive to the point such that the added geometric primitive surrounds the point. The interpolating step may include, for each of one or more points in the point cloud, adding a geometric primitive to the point such that the point is at the center of the added geometric primitive.

[0018] In some embodiments, the connectivity information defines one or more properties of each geometric primitive determined for one or more points in the point cloud, and determining the geometric primitives for the one or more points in the point cloud includes defining the one or more properties of each geometric primitive using the connectivity information, in which case primitive sizes can be determined at the encoder and included in the bitstream to improve interpolation (at relatively little data cost).

[0019] The connectivity information may include, for each point in the point cloud, a primitive size that indicates the size of each geometric primitive determined for one or more points in the point cloud.

[0020] Obtaining the connectivity information may include receiving or estimating one or more subsets of related points in the point cloud, each subset of related points representing points corresponding to the same face in the scene, and interpolating the geometry may include determining a geometric primitive for each subset of related points.

[0021] The interpolating step may include adding a geometric primitive to each subset of associated points. In some embodiments, the interpolating step includes, for each subset of associated points, adding a geometric primitive to the subset of associated points such that each of the subset of associated points touches the added geometric primitive at a non-vertex location of the added geometric primitive.

[0022] Interpolating the geometry of the scene may include connecting the relevant points in one or more of the subsets of relevant points.

[0023] The method may further include generating a temporally upsampled point cloud using the point cloud.

[0024] The step of generating the temporally upsampled point cloud may include receiving motion data for points in the point cloud and using the motion data to generate temporally upsampled points for the temporally upsampled point cloud.

[0025] The method may further include receiving sub-frame timing for points in the point cloud, wherein the point cloud is adapted based on the sub-frame timing and / or rendering the view of the scene is based on the sub-frame timing.

[0026] Points in a point cloud may be acquired at slightly different times. This occurs, for example, when acquiring a point cloud using LiDAR. Sub-frame timing provides information about the differences in when the points were acquired, so that they can be taken into account when rendering the view.

[0027] A decoder system may also be provided that includes a processor that performs any of the above steps for generating a view of a scene at a target viewpoint for immersive video.

[0028] The present invention also provides an encoding method for immersive video, the method comprising: acquiring an image of the scene; acquiring a point cloud of a scene at a point cloud viewpoint; determining connectivity information between points in the point cloud; and encoding the image, the point cloud, and the connectivity information.

[0029] In some embodiments, the connectivity information defines one or more properties of one or more geometric primitives that are determined for one or more points in the point cloud.

[0030] Preferably, two or more images of the scene are acquired at different viewpoints.

[0031] The point cloud may be obtained from a LiDAR.

[0032] In this method, acquiring images of the scene may include acquiring two images of the scene, each image representing a different image perspective of the scene, and determining connectivity information may include applying a first geometric primitive at a first primitive size to points in the point cloud, performing a consistency check of the first primitive size using the two images, applying a second geometric primitive at a second primitive size to points in the point cloud, and performing a consistency check of the second primitive size using the two images.

[0033] The connectivity information includes either the first primitive size or the second primitive size based on a consistency check of both the first and second primitive sizes, and preferably the connectivity information includes the largest size of the first and second primitive sizes.

[0034] In one embodiment, the first geometric primitive and / or the second geometric primitive is a region of pixels, and applying the geometric primitive means selecting a region of pixels. In this embodiment, the primitive size is the size of the pixel region (e.g., 2x2, 4x4, 2x4, etc.).

[0035] Consistency checks typically involve comparing pixel values ​​in geometric primitives (e.g., pixel domains). For example, the sum of the color differences between pixels can be calculated.

[0036] Of course, consistency checks can be performed for more than two primitive sizes.

[0037] The step of determining the connectivity information may include comparing pixel colors in the two images at locations corresponding to points in the point cloud, and determining one or more subsets of related points in the point cloud based on the comparison of the pixel colors.

[0038] The method may further include comparing pixel colors in the two images using the point clouds and determining a visibility metric for each point in the point cloud, the visibility metric including view visibility for each of the two image viewpoints.

[0039] In other words, the visibility metric provides information about the image that can be used when blending textures at each point in the point cloud.

[0040] The method may further include determining motion data between the point cloud and a previous point cloud of the scene, and encoding the motion data.

[0041] Therefore, the decoder can use the motion information to temporally upsample the received point cloud over time.

[0042] The method may further include determining tolerance information indicative of a registration error between the points in the point cloud and the two images, and encoding the tolerance information.

[0043] The tolerance information may include an average error and a maximum error. In one embodiment, the average error is less than 2 pixels and the maximum error is less than 10 pixels. In another embodiment, the average error is less than 1 pixel, preferably less than 0.5 pixels, and the maximum error is less than 2 pixels.

[0044] The invention also provides a computer program medium comprising computer program code means which, when executed on a processing system, cause the processing system to perform all the steps of a method for generating and / or encoding a view of a scene.

[0045] The computer program medium may be a relatively long-term data storage solution (e.g., a hard drive, a solid state drive, etc.) Alternatively, the computer program medium may be a data transmission medium such as a bitstream.

[0046] The present invention also provides a system for generating a view of a scene at a target viewpoint in an immersive video, the system comprising: receiving an image of the scene and a point cloud of the scene at a point cloud viewpoint; Obtaining connectivity information between points in the point cloud; Interpolating the geometry of the scene using the point cloud and connectivity information; A processor is provided for rendering a view of the scene at a target viewpoint using the image and the interpolated geometry.

[0047] The processor may perform the interpolation by performing a process that includes determining a geometric primitive for one or more points in the point cloud.

[0048] The present invention also provides a system for encoding immersive video, the system comprising: Acquire an image of the scene, Obtain the point cloud of the scene from the point cloud viewpoint, determining connectivity information between points in the point cloud; A processor is provided for encoding the image, the point cloud, and the connectivity information.

[0049] The connectivity information defines one or more properties of one or more geometric primitives that are determined for one or more points in the point cloud.

[0050] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]

[0051] For a better understanding of the present invention and to show more clearly how it may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings in which:

[0052] [Figure 1] Figure 1 shows an object being imaged with three cameras and a laser system. [Figure 2] Figure 2 shows the positions of four points from the point cloud relative to the positions of two image viewpoints. [Figure 3] FIG. 3 shows an example of temporal upsampling of points in a point cloud. [Figure 4] FIG. 4 shows a method for rendering a view at a target viewpoint. DETAILED DESCRIPTION OF THE INVENTION

[0053] The present invention will now be described with reference to the drawings.

[0054] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the devices, systems, and methods, are for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the devices, systems, and methods of the present invention will become better understood from the following description, the appended claims, and the accompanying drawings. It should be understood that the figures are schematic representations only and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the figures to indicate the same or similar parts.

[0055] The present invention provides a method for generating a view of a scene at a target viewpoint for immersive video, the method including receiving an image of a scene and a point cloud of the scene at a point cloud viewpoint, and obtaining connectivity information between points in the point cloud, interpolating a geometry of the scene using the point cloud and the connectivity information, and rendering a view of the scene at the target viewpoint using the image and the interpolated geometry.

[0056] Traditional multi-view data consists of video data from individual capturing cameras, optionally combined into a single video atlas. Traditional point cloud data consists of points captured by laser systems, e.g., LiDAR.

[0057] LiDAR is a laser scanning instrument that measures the distance and intensity of one or more reflections along the laser path for multiple scan directions of the laser, e.g., 128 longitudes and 64 latitudes. One complete pass of the laser can be called a frame, with a typical frame rate of 20 Hz.

[0058] Assuming synchronization and interleaved calibration of the multi-view video data with the point cloud, the novel view synthesizer can include selecting a target viewpoint for rendering a novel view, projecting the point cloud onto the target viewpoint, and computing a mapping for pixels at the target viewpoint with valid depth values ​​to the two closest views in the set of multi-view camera images. In this way, colors from the closest views can be obtained and blended at the target viewpoint.

[0059] The novel view synthesiser described above is well suited to live streaming contexts, both as an encoder-side and post-decode process, due to the very limited rendering time.

[0060] However, projecting occluded points in the point cloud onto the target view will likely result in erroneous depth values ​​and incorrect color values ​​from the images in the multi-view data, limiting image quality. Additionally, because the point cloud spatial scanning resolution is typically much lower than the image size from the multi-view dataset, the depth map in the target view will have a low resolution, potentially resulting in incorrect color values. Furthermore, the difference in perspective between the laser (used to acquire the point cloud) and the multi-view camera rig will leave holes in the depth map of the target view.

[0061] One solution to solving the hole and occlusion problem is to render a point not just by rasterizing the point at a given output resolution, but by adding a respective geometric primitive to each subset of the point and compositing the added geometric primitives. Each geometric primitive is defined as having at least one dimension (e.g., at least two dimensions, e.g., three dimensions). Thus, each geometric primitive is a line / curve, a plane / surface, and / or a volumetric region.

[0062] 1 shows a use case environment 100 for contextualizing the proposed concepts. In the environment 100, an object 102 is being imaged by three cameras 104, 106, and 108 and a laser system 110. The cameras 104, 106, and 108 each acquire images of the object 102 at different image viewpoints, and the laser system 110 acquires a point cloud of the object 102.

[0063] In the illustrated example, points 112, 114, and 116 correspond to points in a point cloud acquired by the laser system. Points 112, 114, and 116 are directly visible from laser system 110. However, due to differences in perspective between laser system 110 and cameras 104, 106, and 108, some points in the point cloud may not be directly visible from cameras 104, 106, or 108. In this case, points 112 and 114 are visible from all three cameras 104, 106, and 108, but point 116 is not directly visible from any of the three cameras 104, 106, or 108. In other words, point 116 is occluded from the camera's perspective.

[0064] However, when point 116 is used to synthesize a new view, the color fetched from the image (obtained from cameras 104, 106, and / or 108) does not correspond to the actual color of point 116, but instead corresponds to the color of the surface as seen by the camera on which points 112 and 114 reside.

[0065] More specifically, if the images generated by the cameras 104, 106, 108 were used to identify the color of the point 116, the identified color would be the color of the occluded portion of the object 102. This is because the line of sight (shown by the dashed lines) from each of the cameras 104, 106, 108 to the point 116 is blocked by the occluded portion of the object 102. Therefore, the new image that uses the point 116 to provide a new view will not correspond to the actual / true color of the point 116.

[0066] This disclosure proposes applying or attaching geometric primitives 118 and 114 to one or more points in the point cloud (e.g., points 112 and 114). In this manner, the geometry of the scene is interpolated by determining a geometric primitive to be applied to each of one or more points in the point cloud. More specifically, the geometric primitive is attached to each of one or more points of the point cloud.

[0067] Even more specifically, geometric primitives are added to their corresponding points such that the corresponding points are at non-vertex locations (e.g., on the edge of or surrounded by) the (added) geometric primitive, thereby extending the effective size of the points and geometrically interpolating the points.

[0068] In the illustrated example, each geometric primitive 118 and 120 is a mini-mesh that forms a solid black line. Each geometric primitive 118, 120 is applied or attached to a corresponding point such that the point is at a non-vertex location of the applied geometric primitive. In the context of this disclosure, the term "vertex" refers to both the corners of a multidimensional shape and the endpoints of a one-dimensional shape (e.g., a line).

[0069] Geometric primitives 118 and 120, which now cover most of the surfaces visible from the cameras (as applied / attached), effectively cover the occluded point 116, such that no color is applied to point 116 from cameras 104, 106, or 108. Conceptually, if the point cloud (with geometric primitives applied) and the camera representations are formed in the same virtual space, the line of sight between each camera and point 116 will intersect with one of the geometric primitives. Therefore, it is possible to prevent using any image generated by such a camera to determine the color of point 116. In this way, point 116 is correctly occluded by geometric primitives 118 and 120.

[0070] Geometric primitives are predefined and can be constant in shape across a point. They can be a single triangle with the point at its center, or a mini-mesh of triangles. The idea is that a geometric primitive locally models or represents the geometry of the surface around the point captured by the laser. Thus, a geometric primitive is attached to a point, surrounding that point. It is also useful to specify, for each point, how the reference primitive should be scaled, rotated, and / or shifted when generating the scene geometry in the target view.

[0071] One approach to determining whether to associate a geometric primitive with a point and / or defining the size of the geometric primitive is to perform a multi-view consistency check procedure for each point in the point cloud.

[0072] The multiview consistency check procedure for a point can project the point to all image viewpoints. In this way, the relative position of the point in each image generated for each image viewpoint can be identified. For each image viewpoint, multiple pixel blocks of a predetermined size can be defined. Each pixel block has a different size (e.g., 2x2, 4x4, 8x8, 16x16) from the other pixel blocks in the multiple pixel blocks. Each pixel block is positioned (e.g., centered) at the relative position of the point in the image for the image viewpoint. A consistency check between pixel blocks of the same size is performed for each image viewpoint. The consistency check may include determining a measure of correspondence between pixel blocks of the same size, such as the sum of absolute color differences, for one or more pairs of (adjacent) image viewpoints. The smaller the summed absolute difference, the higher the correspondence. The largest block size that remains consistent across a predetermined percentage or number of viewpoints (e.g., the average measure of correspondence for that block size meets some predetermined condition) is selected. This block size can be used to define the size of the geometric primitives applied to the point when synthesizing the new view. In particular, the larger the identified block size, the larger the size of the geometric primitives. The block size and / or the size of the geometric primitives (determined from the block size) may be encoded. Thus, the connectivity information includes primitive sizes that indicate the size of each geometric primitive determined for one or more points in the point cloud.

[0073] In other words, each point in the point cloud can be projected to all image viewpoints before encoding. Multiview consistency checks can then be performed across viewpoints at different spatial scales (e.g., three different block sizes). Rectangular pixel regions (blocks) of pixels (e.g., 2x2, 4x4, 8x8, 16x16) can be used for consistency checks, where the sum of absolute color differences across the block from the source view to the neighboring view is a measure of correspondence. The smaller the summed absolute difference, the higher the correspondence. The largest block size that is still consistent is then selected according to the multiview match error across the smallest percentage of all views (e.g., for at least half of all views). The selected block size can then be used to calculate or define the size of a geometric primitive (e.g., a mesh) in world space using depth in one or more views. Consistency across relatively large blocks of pixel data allows for local representation of large geometric primitives. This situation typically occurs away from disocclusion edges. Note that the blocks do not necessarily need to be aligned with the block grid. That is, multiple shifts can be tested for grid locations (depth transitions) so that the shifted blocks can be selected to describe a geometry closer to the actual object boundary.

[0074] As shown in FIG. 1 , the laser system 110 records the locations of points that are visible in only a subset of all views. Thus, in another further example, a visibility bit can be attached to each point per source view. For example, simply encoding all eight views results in an 8-bit number per point indicating whether the point is visible. This visibility information can be obtained prior to encoding by comparing pixel colors across views via projections of the point onto each view. Outlier color values ​​can be detected, and the corresponding views can be set to invisible. Client-side rendering maintains efficiency by using the visibility bit to ensure that only the visible views of a given point are used in blending operations.

[0075] An alternative or complement to the idea of ​​attaching a geometric primitive to each of multiple single points for rendering is to connect the points or define connection points (i.e., define edges between the points) and then apply a primitive to each of one or more sets of connection points. In particular, for each set of connection points, an applied or attached geometric primitive is attached to each of the connection points such that there are no connection points at the vertices of the attached geometric primitive. In particular, the attached geometric primitive extends outside the area enclosed by the connection points.

[0076] Alternatively, the attached geometric primitive may be a primitive that connects a set of connection points.

[0077] For example, if two (or more) points are found to have similar color statistics or values ​​across views through projection and color comparison across views, then those points can be inferred to be on the same surface. Thus, it is possible to infer that (adjacent or nearby) points that share similar color statistics or values ​​across views are on the same surface, and therefore, it is possible to apply geometric primitives that can operate as the primitives disclosed above to sets of points that share similar color statistics or values.

[0078] For example, for pairs of points that share similar color statistics or values ​​across a view, they can be indicated as points on a connecting line and this information can be sent along with the point cloud. The client-side renderer can then draw a line of a given thickness connecting the points to represent a geometric primitive. More specifically, a geometric primitive having a particular size can be applied to a set of points identified as being on the same line. The geometric primitive applied in this example is a rectangle with a constant thickness and, optionally, a length that is longer than the distance between the two points (e.g., a predetermined percentage of the distance between the two points). In this way, the applied geometric primitive has a size that is larger than the size of the region or area enclosed by the set of points.

[0079] Similarly, for three points that share similar color statistics or values ​​across views, the connection between the points creates a triangle, and this information can be transmitted in the connectivity information (e.g., as metadata). If a triangle (or line, or other shape) is defined for a set of points, the triangle can be rendered instead of rendering all the points separately. More specifically, a geometric primitive can be applied to each of the three points (e.g., a geometric primitive that is larger than the area defined by connecting the points).

[0080] Larger sets or clusters of points can be handled in a similar manner.

[0081] In this way, knowledge of the object surface colors of the points in the point cloud can be used to convert the sparse point cloud of the scene into a dense geometry of the scene.

[0082] FIG. 2 shows the locations of four points from the point cloud relative to the locations of two image viewpoints 202 and 204. Each point has an associated feature vector [f1, f2, f3, f4], which includes, for example, color and / or texture representing the surface being imaged. During rendering of a new view, the points can first be warped to two source views 202 and 204, where pixel-accurate segmentation is performed (lines 208 and 210), and geometric models or geometric primitives (lines 212 and 214) can be fitted, or adapted, to each of the segments 208 and 210. Using the fitted geometry (e.g., a plane), per-pixel depth can be mapped, resulting in dense per-view depth maps. These resulting depth maps can be used to warp and combine source view textures to a target view 206. Using an alternative rendering method, the depth map itself is first warped to the target view, and then texture data is retrieved from the source views and blended based on the warped depth.

[0083] The feature vector for each point can be determined by analyzing color consistency across views before encoding: the most frequent color / texture vector across views can be assigned to that point.

[0084] LiDAR typically involves a trade-off between spatial resolution and temporal capture frequency. In practice, this means that for useful spatial resolution, the temporal resolution is typically around 10 Hz per point. Furthermore, due to the scanning process, each point has its own timestamp. For static geometries, this is not an issue. However, since objects usually move, it is preferable to upsample the point cloud to typical video frame rates of 30 Hz or even 60 Hz. Exploding the point cloud before encoding significantly increases the bitrate, so it is also preferable to do this on the client side.

[0085] The first realization is that providing the point cloud before the video frame time allows time for spatial interpolation during the temporal upsampling of the point cloud. The second realization is that managing a spatial data structure that stores points over time helps with fast lookup of nearby points with different timestamps.

[0086] Figure 3 shows an example of temporal upsampling of points in a point cloud. Using a time sliding window of the two closest LiDAR scans relative to the video frame time ensures that point position interpolation is achievable. The first interpolation method is to determine, for each point in the latest LiDAR scan, the closest point in the previous LiDAR scan. Using the relative times of both points relative to the video frame time and assuming linear motion, linear interpolation is possible.

[0087] The point cloud contains the video frame pixels at times t1 and t2 as a function of spatial coordinates, and the scene points x(t1) and x(t2) for the laser reflectance lines 302 and 304, respectively. The scanning motion of the laser provides a correlation between capture time and spatial location. By transmitting the point cloud data before the video frame image, the positions of the points in the point cloud can be interpolated to the video frame capture time.

[0088] In particular, interpolation of a point x'(t) at time t is achieved by assuming linear motion between the closest point x(t1) at time t1 and the closest point x(t2) at time t2.

[0089] An alternative approach is to determine the motion vectors for each point before encoding and transmit the motion vectors as metadata along with the point cloud. This can be combined with the idea of ​​having connected points, since the motion vector specifies the combined motion of a subset of points. This is advantageous because it requires a lower bitrate to transmit the motion information.

[0090] It will be appreciated that the projection information is used to enable the projection of a point of the point cloud onto the projection plane of the view in order to render the view at the target viewpoint. More generally, the projection information provides information about the relative position / orientation between the viewpoint of the image and the viewpoint of the point cloud.

[0091] Projection information includes view internal parameters (such as projection model parameters, principal point parameters, focal length parameters, and lens distortion parameters), view external parameters (position and orientation = pose), projection surface size (relationship between image coordinates and view coordinates, if not identical), point cloud external parameters (position and orientation = pose), and / or point cloud scale / range parameters (associating code points with distances in scene units).

[0092] It should be appreciated that there are several known representations for view parameters and point cloud parameters, including using matrices and different representations of position and orientation.

[0093] By relating the view and point cloud with external parameters, there exists a scene coordinate space in which both are located. This scene coordinate system has axis semantics such as "front", "left", "top", etc. that are useful for rendering. (In MIV, the definitions can be specified using the coordinate axis system SEI message.)

[0094] Alternatively, the reference frame of the point cloud is the scene reference frame, and the extrinsic parameters are mapped to the point cloud. When there are multiple views and one point cloud, this becomes a logical representation.

[0095] Alternatively, one of the (synthetic) views is a 360-degree or ultra-wide view, an in-paint view, or an environment map, and external parameters map the other views and point clouds onto this view.

[0096] Those skilled in the art will recognize the various types of projection information, how they are used, and what is needed in different situations.

[0097] The way to render a view at a target viewpoint is to warp / project the depth data (in this case a point cloud) to the target viewpoint, fetch pixel texture values ​​(such as color and transparency) from the image for the points in the point cloud, and blend the pixel texture values ​​for each point. The projection information is used to accomplish the warping / projection of the point cloud and fetching of the pixel texture colors.

[0098] In one embodiment of the present invention, the alignment between a point in the point cloud and a corresponding image sample in the coded image is accurate to within a few pixels. For example, the average error is less than 2 pixels and the maximum error is less than 10 pixels. The maximum error is expected to be towards the edge of the projection plane of view.

[0099] The advantage of allowing such a wide tolerance is that it makes it easier to set up the camera system and capture the scene, however, misalignment can lead to ghosting when combining multiple image regions for a region of the viewport, so it is preferable for the decoder or renderer to perform a viewport image alignment step.

[0100] In another embodiment of the invention, the alignment is accurate to an error of less than 0.5 pixels, with a maximum error of less than 2 pixels. The advantage of requiring such a tight tolerance is that the decoder or renderer does not need to perform a viewport image alignment step, thereby reducing processing resource requirements.

[0101] To enable this decision, in a preferred embodiment of the present invention, the bitstream contains tolerance information, and the renderer selects a rendering strategy based on this tolerance information. Examples of tolerance information include one or a combination of a flag for narrow or wide tolerances, an average image sample projection error, a maximum image sample projection error, and / or a map providing different values ​​for more central and more outer positions on the projection plane.

[0102] In larger camera systems, multiple LiDARs can be used, with each LiDAR forming a group with the nearest 2D camera. More generally, multiple point clouds are acquired, each corresponding to a group of images. Each such group can be encoded independently and the bitstreams can be merged. Sub-bitstream access is possible to extract a subset of a group. Group-based coding of multiview + depth is known in MIV.

[0103] To improve the rendering of multiple groups, external parameters within and between groups can be provided. In addition, different tolerances within and between groups can be provided. Also, different temporal coherences within and between groups can be allowed.

[0104] 4 shows a method for rendering a view at a target viewpoint. In step 402, an image and a point cloud of a scene are received. For example, step 402 may include receiving coded representations of attributes of at least first and second image regions (at different image viewpoints), receiving the coded point cloud, receiving projection information, receiving viewport parameters, and decoding all or a portion of the received information.

[0105] In step 404, connectivity information for points in the point cloud is obtained. The connectivity information may be received (e.g., along with the coded image and point cloud) or estimated / determined by the client. For example, the connectivity information may include a subset of points in the point cloud that correspond to faces.

[0106] The connectivity information is used to interpolate the geometry of the scene in step 406. In particular, step 406 includes determining geometric primitives for one or more points of the point cloud. Using the scene geometry, a view is rendered at the target viewpoint in step 408. Step 408 may include projecting points of the decoded point cloud onto the target viewport, projecting points in the point cloud onto positions in the decoded image, and rendering those positions to form the viewport image.

[0107] It will be appreciated that there is a temporal relationship between the image and the point cloud.

[0108] It will also be appreciated that the projection information allows a point in the point cloud to be projected onto a corresponding image sample within a tolerance range.

[0109] Thus, the method can generate a dense viewport image from sparse point cloud data by interpolating geometry information and fetching attribute information.

[0110] Preferably, once the tolerance is received and exceeds a threshold, the method further comprises an alignment refinement step for fine-tuning the alignment between the point cloud and the image.

[0111] Preferably, information about the intensity of light reflections is received and applied in the rendering step (eg, in an image filtering step of the rendering step) to improve the rendering of partially translucent scene elements.

[0112] Preferably, information regarding the number of ray bounces is received and applied in the rendering step to improve the rendering of partially translucent scene elements.

[0113] Preferably, the method further comprises receiving, rendering, etc., a plurality of groups, each group comprising the point cloud and at least one image region of one view.

[0114] Preferably, the method further comprises receiving "deep image" point attributes that enhance the rendering of non-Lambertian scene elements such as color, reflectance, normal, polarization, etc. Deep image points may include additional attributes such as reflectance properties, normal vectors, transparency, etc.

[0115] Preferably, the method further comprises receiving connectivity information regarding relationships (such as labels or edges) between points of the point cloud that further enables the rendering step to filter relevant points differently from non-relevant points and / or enables rendering of relevant points using the first approach and non-relevant points with the second approach, whereby the algorithmic complexity of the first approach is lower than the second approach. In other words, the connectivity information enables different subsets of points to be rendered differently. For example, for a subset of points that define a surface (e.g., via a mini-mesh), the surface can be processed as a whole rather than processing each point separately.

[0116] Optionally, the method further comprises receiving reprojection information that does not comprise view parameters and point cloud parameters but associates relevant points with image regions.

[0117] Preferably, the method further comprises receiving motion information for the cluster of associated points and rendering the cluster of points at a specified time instance (ie, frame rate up-conversion).

[0118] Preferably, the method further includes receiving sub-frame timing information for points in the point cloud or clusters of points in the point cloud, and the rendering step further includes correcting for timing differences to some extent, which can avoid rolling shutter effects or correct for rotating LiDAR.

[0119] Similarly, the sub-frame timing of a viewport display device can be compensated for during rendering. For example, head-mounted displays have high frame rates (e.g., 120 Hz) and often update the viewport during rendering so that line n and line n+1 of the same frame correspond to slightly different viewport positions. Such sub-frame updates reduce the effective update time. These sub-frame timings can be compensated for during rendering.

[0120] By receiving connectivity information, for example by performing mesh-based rendering, the rendering of the view becomes more efficient.

[0121] By estimating / determining connectivity information, for example, mesh-based rendering, the rendering of views is made more consistent.

[0122] Those skilled in the art will be able to readily develop a processor to perform the methods described herein. Accordingly, each step in the flowchart represents a different action performed by a processor, and may be performed by each module of the processing processor.

[0123] As previously mentioned, systems use processors to process data. Processors can be implemented in a variety of ways using software or hardware to perform the various functions required. Typically, a processor uses one or more microprocessors that are programmed using software (e.g., microcode) to perform the required functions. A processor can be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.

[0124] Examples of circuitry that may be used in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs).

[0125] In various implementations, a processor may be associated with one or more storage media, such as volatile and non-volatile computer memory, such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform the necessary functions. The various storage media may be fixed within a processor or controller, or may be transportable so that one or more programs stored thereon can be loaded into a processor.

[0126] Variations of the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps and the word "a" or "an" does not exclude a plurality.

[0127] The functions performed by a processor may be performed by a single processor or by multiple separate processing units which together may be considered to constitute a "processor". Such processing units may be remote from each other and may communicate with each other in wired or wireless manner.

[0128] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0129] The computer program may be stored / distributed on any suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless communication systems.

[0130] It should be noted that when the term "adapted to" is used in the claims or description, it is intended to be equivalent to the term "configured to." When the term "apparatus" is used in the claims or description, it is intended to be equivalent to the term "system," and vice versa.

[0131] Any reference signs in the claims should not be construed as limiting the scope.

[0132] Generally, examples of apparatus and methods for generating a view of a scene at a target viewpoint for immersive video are presented in the following embodiments.

[0133] Embodiment 1. 1. A method for generating a view of a scene at a target viewpoint for immersive video, comprising: receiving an image of a scene and a point cloud of the scene at a point cloud viewpoint; Obtaining connectivity information between points in the point cloud; interpolating the geometry of the scene using the point cloud and connectivity information; and rendering a view of the scene at the target viewpoint using the image and the interpolated geometry. Embodiment 2. 2. The method of embodiment 1, wherein obtaining connectivity information includes receiving or estimating one or more subsets of associated points in the point cloud, each subset of associated points indicating points corresponding to a surface in the scene. Embodiment 3. 3. The method of embodiment 2, wherein interpolating the geometry of the scene includes connecting associated points in one or more of the subsets of associated points. Embodiment 4. 4. A method according to any one of the preceding embodiments, wherein interpolating the geometry of the scene comprises determining geometric primitives for one or more of the points in the point cloud. Embodiment 5. 5. The method of embodiment 4, wherein the connectivity information includes a primitive size indicating the size of the geometric primitive to be determined. Embodiment 6. 6. The method of any preceding embodiment, further comprising using the point cloud to generate a temporally upsampled point cloud. Embodiment 7. receiving subframe timing for points in the point cloud; The point cloud is adapted based on subframe timing, and / or 7. A method according to any one of embodiments 1 to 6, wherein rendering a view of the scene is based on sub-frame timing. Embodiment 8. 1. A method of encoding for immersive video, comprising: acquiring an image of the scene; acquiring a point cloud of a scene at a point cloud viewpoint; determining connectivity information between points in the point cloud; encoding the image, the point cloud, and the connectivity information. Embodiment 9. Determining the connectivity information includes: applying a first geometric primitive at a first primitive size to points in the point cloud; performing a first primitive size consistency check using the two images; applying a second geometric primitive at a second primitive size to points in the point cloud; performing a second primitive size consistency check using the two images; and and encoding either the first primitive size or the second primitive size based on a consistency check of both the first and second primitive sizes, wherein the largest of the first and second primitive sizes is preferably encoded. Embodiment 10. Determining the connectivity information includes: comparing pixel colors in the two images at locations corresponding to points in the point cloud; and determining one or more subsets of relevant points in the point cloud based on a comparison of pixel colors. Embodiment 11. Using point clouds to compare pixel colors in two images; 11. The method of claim 8, further comprising determining a visibility metric for each point in the point cloud, the visibility metric comprising a view visibility for each of the two image viewpoints. Embodiment 12. 12. A method according to any one of embodiments 8 to 11, further comprising determining tolerance information indicative of a registration error between points in the point cloud and the two images, and encoding the tolerance information. Embodiment 13. A computer program medium comprising computer program code means, which when executed on a processing system, causes the processing system to perform all the steps of one or more of the methods described in embodiments 1 to 7 or the methods described in embodiments 8 to 12. Embodiment 14. 1. A system for generating a view of a scene at a target viewpoint in an immersive video, comprising: receiving an image of the scene and a point cloud of the scene at a point cloud viewpoint; Obtaining connectivity information between points in the point cloud; Interpolating the geometry of the scene using the point cloud and connectivity information; A system comprising a processor that uses the image and the interpolated geometry to render a view of the scene at a target viewpoint. Embodiment 15. 1. A system for encoding immersive video, comprising: Acquire an image of the scene, Obtain the point cloud of the scene from the point cloud viewpoint, determining connectivity information between points in the point cloud; A system comprising a processor that encodes an image, a point cloud, and connectivity information.

[0134] More particularly, the present invention is defined by the appended claims.

Claims

1. 1. A method for generating a view of a scene at a target viewpoint for immersive video, comprising: receiving an image of the scene and a point cloud of the scene at a point cloud viewpoint; obtaining connectivity information between points in the point cloud; interpolating the geometry of the scene using the point cloud and the connectivity information; rendering the view of the scene at the target viewpoint using the image and the interpolated geometry; A method comprising:

2. The method of claim 1 , wherein the interpolating step comprises determining a geometric primitive for one or more points in the point cloud.

3. The method of claim 2 , wherein the interpolating step comprises adding a geometric primitive to each of the one or more points in the point cloud.

4. 4. The method of claim 3, wherein the interpolating step comprises, for each of the one or more points in the point cloud, adding a geometric primitive to the point such that the point touches the added geometric primitive at a non-vertex location of the added geometric primitive.

5. 5. The method of claim 4, wherein the interpolating step comprises, for each of the one or more points in the point cloud, adding the geometric primitive to the point such that the added geometric primitive surrounds the point.

6. 6. The method of claim 5, wherein the interpolating step comprises, for each of the one or more in the point cloud, adding the geometric primitive to the point such that the point lies at the center of the added geometric primitive.

7. the connectivity information defines one or more properties of each geometric primitive determined for the one or more points in the point cloud; 7. The method of claim 2, wherein determining geometric primitives for the one or more points in the point cloud comprises using the connectivity information to define the one or more properties of each geometric primitive.

8. The method of claim 7 , wherein the connectivity information includes, for each point in the point cloud, a primitive size indicating a size of each geometric primitive determined for the one or more points in the point cloud.

9. obtaining the connectivity information includes receiving or estimating one or more subsets of related points in a point cloud, each subset of related points representing points corresponding to the same surface in the scene; The method of claim 1 , wherein interpolating the geometry comprises determining a geometric primitive for each subset of associated points.

10. The method of claim 9 , wherein the interpolating step includes adding a geometric primitive to each subset of associated points.

11. 11. The method of claim 10, wherein the interpolating step includes the step of, for each subset of associated points, adding geometric primitives to the subset of associated points such that each of the subset of associated points touches the added geometric primitive at a non-vertex location of the added geometric primitive.

12. acquiring an image of the scene; acquiring a point cloud of the scene at a point cloud viewpoint; determining connectivity information between points in the point cloud; encoding the image, the point cloud, and the connectivity information; A method for encoding for immersive video, comprising:

13. The method of claim 12 , wherein the connectivity information defines one or more properties of one or more geometric primitives determined for one or more points in the point cloud.

14. 14. A computer program medium comprising computer program code means which, when executed on a processing system, cause the processing system to perform all the steps of one or more of the methods of any one of claims 1 to 11, or the methods of claims 12 or 13.

15. 1. A system for generating a view of a scene at a target viewpoint in an immersive video, comprising: receiving an image of the scene and a point cloud of the scene at a point cloud viewpoint; obtaining connectivity information between points in the point cloud; interpolating the geometry of the scene using the point cloud and the connectivity information; a processor that uses the image and the interpolated geometry to render the view of the scene at the target viewpoint.