Techniques for using view-dependent point cloud renditions
By capturing and processing images from multiple camera positions to generate a model with associated attributes, the method addresses the issue of uncertain renderings in point cloud technology, achieving realistic and accurate 3D image rendering.
Patent Information
- Application Number
- JP2023521909
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-12
- Filing Date
- 2021-10-12
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2041-10-12
AI Technical Summary
Current image rendering techniques using point clouds fail to provide realistic views of objects under all angle conditions due to the lack of documented camera settings and angles used for capturing 3D attributes, leading to uncertain and unrealistic renderings.
A method and device for rendering images by receiving images from multiple camera positions, determining camera orientations, and generating a model to provide virtual renderings with appropriate attributes based on viewing orientations, using a decoder and encoder to reconstruct and encode point clouds.
Enables realistic and faithful rendering of 3D objects from various angles by associating camera positions with attribute data, ensuring accurate attribute modulation and improved rendering quality.
Smart Images

Figure 0007791883000003 
Figure 0007791883000004 
Figure 0007791883000005
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to image rendering, and more particularly to image rendering using point cloud techniques. [Background technology]
[0002] Volumetric video capture is a technique that allows for the capture of moving images, often of real-world scenes, so that they can be later observed from any angle. This is very different from regular camera capture, which is limited to capturing images of people and objects from only certain angles. In addition, video capture allows for the capture of scenes in three-dimensional (3D) space. As a result, the acquired data can be used to establish immersive experiences, either real or alternatively computer-generated. As virtual, augmented, and mixed reality environments become more popular, volumetric video capture techniques are also becoming more popular. This is because the technique uses the visual qualities of photography and blends them with the immersion and interactivity of spatialized content. The technique is complex and combines many of the recent advances in computer graphics, optics, and data processing.
[0003] Volumetric visual data is typically captured from real-world objects or provided through the use of computer-generated tools. One common way to provide a common representation of such objects is through the use of point clouds. A point cloud is a set of data points in space that represent a three-dimensional (3D) shape or object. Each point has its own set of X, Y, and Z coordinates. Point cloud compression (PCC) is a method for compressing volumetric visual data. A subgroup of the Motion Picture Expert Group (MPEG) is working on the development of the PCC standard. MPEG PCC requirements for point cloud representation require view-dependent attributes for each 3D location. Patches, or certain points of a point cloud, are observed according to an observer angle. However, observing any 3D object in a scene according to different angles may require modification of different attributes (e.g., color or texture) because some visual aspects may depend on the viewing angle. For example, the properties of light can affect the rendering of an object because the viewing angle can change its color and shading depending on the object's material. This is because texture can depend on the incident light wavelength. Unfortunately, current prior art does not provide realistic views of objects under all angle conditions. Attributes modulated according to the observer angle for captured or scanned images do not always provide a faithful rendition of the original content. Part of the problem is that even when the preferred observer angle is known when rendering an image, the camera settings and angles used to capture the image as they relate to 3D attributes are not always documented in a way that can subsequently provide a realistic rendering, and 3D point cloud attributes can be uncertain at some viewing angles. Therefore, techniques are needed to address these shortcomings of the prior art when rendering realistic views and images. Summary of the Invention
[0004] In one embodiment, a method and device for rendering an image are provided. The method includes receiving images from at least two different camera positions and determining a camera orientation and at least one image attribute associated with each position. A model of the image is then generated based on the attributes associated with the received camera positions of the image and the camera orientation. The model is enabled to provide a virtual rendering of the image at multiple viewing orientations and selectively provide appropriate attributes associated with the viewing orientations.
[0005] In another embodiment, a decoder and encoder are provided. The decoder includes means for decoding from a bitstream having one or more attribute data, the data having at least an associated position corresponding to an attribute capture viewpoint. The decoder also includes a processor configured to reconstruct a point cloud from the bitstream using all received attributes and provide a rendering from the point cloud. The encoder can encode the model and the rendering. [Brief explanation of the drawings]
[0006] The teachings of the present disclosure can be more readily understood by considering the following detailed description in conjunction with the accompanying drawings. [Figure 1] FIG. 1 is an illustrative diagram of an example of a camera rig and a virtual camera rendering an image. [Figure 2] Similar to Figure 1, but with the camera rendering the image at a different angle relative to the system coordinates. [Figure 3] FIG. 1 is an illustrative diagram of an octahedral map of an octant of a sphere projected onto a plane and unfolded into a unit square. [Figure 4] FIG. 10 is an illustrative diagram of dereferencing point values and neighborhoods using octahedral modeling. [Figure 5] FIG. 10 is an illustrative diagram of a table providing capture locations, according to one embodiment. [Figure 6] 6 illustrates an alternative table with similar information as provided in FIG. 5. [Figure 7] FIG. 1 is a flow chart illustration according to one embodiment. [Figure 8] 1 illustrates a general overview of an encoding and decoding system according to one or more embodiments. [Figure 9] FIG. 1 is a flowchart illustration of an encoder, according to one embodiment.
[0007] For ease of understanding, where possible, identical reference numbers have been used to designate identical elements common to the figures. DETAILED DESCRIPTION OF THE INVENTION
[0008] FIG. 1 provides an example of a camera rig and virtual camera that provides an image or video rendering. When providing a rendering, the camera capture parameters must be known to at least the processor providing the rendering in order to select appropriate attribute (e.g., color or texture) point samples using point cloud technology. The captured image in FIG. 1 is designated by the numeral 100. The image can be an object, a scene, or part of a video or live stream. If it is a digital image, such as a video image, a TV image, a still image, or an image generated by a video recorder or computer, or a scanned image, the image traditionally consists of pixels or samples arranged in horizontal and vertical lines. The number of pixels in a single image is typically in the tens of thousands. Each pixel typically contains certain characteristics, such as luminance and chrominance information. Because the sheer amount of information conveyed from an image is difficult, if not impossible, to transmit over conventional broadcast or broadband networks, compression techniques are often used to transmit images, such as from an encoder to an image decoder. Many compression schemes comply with the Motion Picture Expert Group (MPEG) standard, which is provided in different embodiments of the present invention.
[0009] Images are captured and presented in two dimensions, such as provided at 100 in FIG. 1. Providing realistic 3D images or renderings that provide a 3D sense of a two-dimensional (2D) image is challenging. One recently used technique utilizes volumetric video capture, particularly using point cloud technology, as discussed above. A point cloud provides a set of data points. Each point has a set of X, Y, and Z coordinates in space, and together this set of points represents a 3D shape or object. When a compression scheme is used, point cloud compression (PCC) involves a large data set describing three-dimensional points associated with additional information such as distance, color, and other properties and attributes.
[0010] In some embodiments, both the PCC standard and the MPEG standard are used. The MPEG PCC requirements for point cloud representation require view-dependent attributes for each 3D location. For example, patches, or certain points of a point cloud, as specified in V-PCC (FDIS ISO / IEC 23090-5, MPEG-I Part 5), are viewed according to the observer angle. However, viewing a 3D object in a scene represented as a point cloud according to different angles may exhibit different attribute values (e.g., color or texture) as a function of the viewing angle. This is due to the properties of the materials that make up the object. For example, the reflectivity of light on a surface (isotropic, anisotropic, etc.) can change the way an image is rendered. Because the material reflectivity of an object's surface depends on the incident light wavelength, the properties of the light generally affect the rendering.
[0011] The prior art does not provide a solution that allows rendition attributes to be faithfully modulated according to observer angle, either for material captured or scanned under different viewpoints, since the camera settings and angles used to capture each 3D attribute are in most cases not documented and the 3D attributes become unspecific from a particular angle.
[0012] Additionally, when using the PCC and MPEG standards, view-dependent attributes do not address 3D graphics as intended, despite the tiling, volumetric SEI, and viewport SEI messages. Additionally, while some information is carried in the V-PCC stream, the same type of point attribute captured by a multi-angle acquisition system (which may be virtual in the case of CGI) is stored across attribute "counts" (ai_attribute_count in the attribute_information(j) syntax structure) and can be identified by an attribute index (vuh_attribute_index, which indicates the index of the attribute data carried within the attribute video data unit), which causes several problems. For example, there is no information about the acquisition system position or angle used to capture a given attribute according to a given angle. Therefore, because there is no relationship between the capture attributes and their capture positions, such a collection of attributes stored in the attribute dimension can only be arbitrarily modulated according to the observer's viewing angle. This leads to several drawbacks and weaknesses, such as a lack of information about the position of the captured attributes, arbitrary modulation of the content during rendering, and unrealistic rendering that is not faithful to the original content attributes.
[0013] In a point cloud arrangement, the attributes of the points can vary according to the observer's viewpoint. To capture these variations, the following factors need to be considered: 1) the observer's position relative to the observed point cloud; 2) collecting attribute values for several points of the point cloud according to different capture angles; and 3) The position of the capturing camera for a given set of captured attribute values (capture position).
[0014] The Video-Based PCC (or V-PCC) standard and specification addresses some of these issues in providing the observer position through the "viewport SEI message family" (item 1), which allows for rendering of view-dependent attributes. Unfortunately, however, as can be seen, this presents rendering problems. In some of these cases, rendering is affected because there is no indication of the position where the attribute was captured. (Note that in one embodiment, ai_attribute_count only indexes the list of captured attributes, but has no information about where they were captured.) This can be solved by different possibilities in storing the capture position in the description metadata once generated and computed.
[0015] In item 2, it should be noted that a particular capture camera may not capture the attributes (color) of an object (e.g., when considering a head, the front camera captures the cheeks, eyes... but not the back of the head...), and as a result, each point is not provided with actual attributes for each angle.
[0016] The position of the camera used to capture the attributes is provided in an SEI message. This SEI has the same syntax elements and the same semantics as the viewport position SEI message, but differs in the following ways as it defines the capture camera position: - "viewport" is replaced in meaning by "capture". -cp_atlas_id specifies the ID of the atlas corresponding to the associated current V3C unit. The value of cp_atlas_id shall be in the range 0 to 63 inclusive. - cp_attribute_index indicates the index of the attribute data associated with the camera position (i.e., equal to the matching vuh_attribute_index). The value of cp_attribute_index shall be in the inclusive range of 0 to (ai_attribute_count[cp_atlas_id]-1). -cp_attribute_partition_index indicates the index of the attribute dimension group associated with the camera position.
[0017] Further information on this detail is provided in Table 1 as shown in Figure 5. The information can be stored in a common location and retrieved from a repository such as an atlas for later use. For example, as shown, cp_atlas_id is not signaled in the bitstream; its value is inferred from a V3C unit present in the same access unit as the capture position SEI message (i.e., it is equal to vuh_atlas_id), or it takes the value of a preceding or following V3C unit.
[0018] Alternatively, the cp_attribute_index is not signaled and is implicitly derived as being in some order from the attribute data stored in the stream (i.e., the order of the derived cp_attribute_index is the same as the vuh_attribute_index in decoding / stream order).
[0019] In yet another alternative embodiment, the capture position syntax structure loops according to the number of attribute data sets present. The loop size can be explicitly signaled (e.g., cp_attribute_count) or inferred from ai_attribute_count[cp_atlas_id]-1. This is shown in Figure 6 and Table 2.
[0020] Additionally, alternatively or optionally, a flag may be provided to indicate in the capture position SEI message whether the capture position is the same as the viewport position. If this flag is set equal to 1, the cp_rotation-related(quaternion->rotation) and cp_center_view_flag syntax elements are not transmitted.
[0021] Alternatively, at least one indicator may be provided that specifies whether the attribute is view-independent along an axis (x, y, z) or direction. Indeed, view-dependence may only occur for a particular axis or position.
[0022] In another embodiment, again additionally or optionally to one of the previous examples, an indicator associates a sector around the point cloud with the attribute data set identified by cp_attribute_index. Sector parameters such as angle and distance from the center of the reconstructed point cloud may be fixed or signaled.
[0023] In an alternative embodiment, the capture position may be provided through processing of the SEI message. This may be discussed in relation to FIG. 2, which shows the capture camera selection of the same image 100, but in this example with three angles for rendering. In one embodiment, the attributes discussed are used. The angle, in one embodiment, is relative to system coordinates. In this embodiment, the angle (or rotation) is determined with various models known to those skilled in the art, such as the quaternion model. (See cp_attribute_index (and optionally cp_attribute_partition_index, which links the location of the attribute capture system to an index of the attribute information, i.e., the matching vuh_attribute_index, which is the index of the attribute data carried in the attribute video data unit to which it relates). This information makes it possible to match the attribute values seen from the capture system (identified by cp_attribute_index) with the attribute values seen from the observer (possibly identified by a viewport SEI message). Typically, the selected attribute dataset is one for which the viewport position parameters (indicated by the viewport SEI message) are equal to or close (according to some threshold and some metric, such as mean squared error) to the capture position parameters (indicated by the capture position SEI message).
[0024] In one embodiment, at render time, each point in the point cloud is rendered by: - Using the dot product, find the closest capture viewpoint from the Capture Position SEI message (where a is the vector between the rendering camera and the point, and b is the vector between the capture camera and the point) for n points [1, ai_attribute_count[j]] in terms of angular distance (see Figure 2)
[0025]
number
[0026] Alternatively, in a different embodiment, the set of captured viewpoints can be selected by "all" capturing viewpoints within a certain maximum angular distance, and then blended in the same manner as described above.
[0027] Figure 3 provides an octahedral representation that maps the octants of a sphere onto the faces of an octahedron, which is projected onto a plane and unfolded into a unit square. Figure 3 can be used as an alternative method for encoding information for rendering by using an implicit model for encoding the directional sectors per point. In this embodiment and example, the captured data is always encoded in a predefined order in a point multi-value table (attribute data), and the data is dereferenced according to the model used. For example, an octahedral model [2, 3] can be used, which allows for a regular discretization of a sphere (see Figure 2) into eight sections (i.e., eight viewpoints). In this case, the unit square can be discretized along the horizontal and vertical axes of the square unit to include n possible point-wise values (e.g., a maximum of 5 × 5 = 25 camera positions).
[0028] Therefore, it is only necessary to encode the model type (i.e., octahedron, or other let for further use) and the discretization square size (e.g., at most n=11). These two values represent every point and are stored very compactly. As an example, the scanning order of the unit square can be raster scanning or clockwise or counterclockwise. An exemplary syntax can be provided as follows:
[0029] [Table 1] where: - If present, cm_atlas_id specifies the ID of the atlas corresponding to the associated current V3C unit. The value of cp_atlas_id shall be in the range 0 to 63, inclusive. -cm_model_idc indicates the model of representation (or mapping) for the discretization of the capture sphere. cm_model_idc equal to 0 indicates that the discretization model is an octahedral model. Other values are reserved for future use. -cm_square_size_minus1+1 represents the size of the unit square that represents the octahedral model in units of points per attribute value. A default value can be determined (such as 11). Additionally, a syntax element can be provided to allow the camera position to be constrained to a square (e.g., top, right, or upper right).
[0030] Alternatively, only the same representation model can be used, which is not signaled in the bitstream. Filling in the actual regular values from the irregular capture can be done by using the algorithm presented in the previous section in the compression stage, with a user-defined value of n.
[0031] Alternatively, only the same representation model is used, which is not signaled in the bitstream. The actual regular value filing from the irregular capture rig can be done by using the algorithm presented previously, where the user can define the value of n for compression.
[0032] Alternatively, an implicit model SEI message can be used for the process shown in Figure 4, where the dereference point values and neighborhoods are used in the previous octahedral model. In this embodiment, at rendering time, the angular coordinates can be used in global coordinates to retrieve the closest available value. This is done by using the ai_attribute_count value: V = Val[i * n+j], where, for example, n=11 and i and j are indices in the horizontal and vertical system coordinates associated with square units. In one embodiment, more complex filtering may use bilinear, using nearest neighbor in an octahedral map for fast processing.
[0033] FIG. 7 provides a flowchart illustration for processing images according to one embodiment. As shown in step 710 (S710), the images are from at least two different camera positions. In S720, a camera orientation is determined. The camera orientation can include a camera angle, a rotation, a matrix, or other similar orientations that would be understood by one skilled in the art. In one embodiment, the angle can be a compound angle (rotation angles x, y, and z expressed in a quaternion model) determined by several angles according to system coordinates. In another example, the camera orientation can be the position of the camera relative to a 3D rendering of the image to be rendered. Alternatively, it can be expressed as a rotation matrix constructed with respect to the coordinates in which the 3D model is represented. Additionally, this step also determines at least one image attribute associated with each position. In S730, a model is generated. The model can be a 3D or 2D point cloud model. In one embodiment, the model is constructed with all attributes (although some may be selectively provided in the rendering (see S740)). The model is a model of the image to be rendered and is based on the attributes and camera orientation associated with the received camera position of the image. At S740, a virtual rendering of the image is provided. The rendering may be of any arbitrary viewing orientation and may selectively provide appropriate attributes associated with the viewing orientation. In one embodiment, a user may select a preferred viewpoint for the rendering to be provided.
[0034] FIG. 8 generally illustrates a general overview of an encoding and decoding system according to one or more embodiments. The system of FIG. 8 may include a preprocessing module 830 configured to perform one or more functions to prepare received content (including one or more images or video) for encoding by an encoding device 840. The preprocessing module 830 may also perform functions such as multi-image acquisition, merging acquired images into a common space, acquiring omnidirectional video in a specific format, and other functions to prepare a format more suitable for encoding. Another implementation may combine multiple images into a common space with a point cloud representation. The encoding device 840 packages the content in a form suitable for transmission and / or storage for reconstruction by a compatible decoding device 870. Typically, although not strictly required, the encoding device 840 provides some degree of compression, allowing the common space to be represented more efficiently (i.e., using less memory for storage and / or requiring less bandwidth for transmission). In the case of a 3D sphere mapped onto a 2D frame, the 2D frame is effectively an image that can be encoded by any of several image (or video) codecs. In the case of a common space with a point cloud representation, the encoding device may provide well-known point cloud compression, for example, by octree decomposition. After being encoded, the data is sent to a network interface 850, which may be typically implemented in any network interface present in a gateway, for example. The data can then be transmitted over a communication network such as the Internet. Various other network types and components (e.g., wired networks, wireless networks, mobile cellular networks, broadband networks, local area networks, wide area networks, WiFi networks, etc.) may be used for such transmission, and any other communication network may be foreseen.The data may then be received via a network interface 860, which may be implemented in a gateway, an access point, a receiver in an end-user device, or any device with communication reception capabilities. After receipt, the data is transmitted to a decoding device 870. The decoded data is then processed by a device 880, which may also communicate with sensor or user input data. The decoder 870 and device 880 may be integrated into a single device (e.g., a smartphone, game console, STB, tablet, computer, etc.). In another embodiment, a rendering device 890 may also be incorporated. In one embodiment, the decoding device 870 may be used to obtain an image including at least one color component that includes interpolated data and non-interpolated data, and to obtain metadata indicative of one or more locations within the at least one color component that has non-interpolated data.
[0035] 9 is a flowchart illustrating a decoder. In one embodiment, the decoder includes means for decoding at least a position corresponding to an attribute capture viewpoint from a bitstream, as shown in S910. The bitstream may have one or more attributes associated with the position corresponding to the attribute capture viewpoint. The decoder includes at least one processor configured to reconstruct a point cloud from the bitstream using all such attributes received, as shown in S920. The processor can then provide a rendering from the point cloud, as shown in S930.
[0036] Many implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or deleted to produce other implementations. Additionally, those skilled in the art will understand that other structures and processes may be substituted for those disclosed, with the resulting implementation performing at least substantially the same function in at least substantially the same way to achieve at least substantially the same results as the disclosed implementations. Accordingly, these and other implementations are contemplated by this application.
Claims
1. 1. A method for processing an image, comprising: receiving a plurality of images, each of the plurality of images captured from one of a plurality of camera positions; determining a camera position and image attributes associated with each of the plurality of images; generating a bitstream for the plurality of images, the bitstream including metadata providing an indication of the image attributes associated with each of the plurality of camera positions.
2. The method of claim 1 , wherein the metadata provides an index of the image attributes associated with each of the plurality of camera positions.
3. The method of claim 1 , wherein the metadata provides an identification of an atlas corresponding to a volumetric coding unit.
4. The method of claim 1 , wherein the metadata provides an index of an attribute dimension group associated with each of the plurality of camera positions.
5. The method of claim 1 , wherein the image attributes include texture or saturation of each of the plurality of images.
6. The method of claim 1 , wherein the metadata is provided via one or more SEI messages.
7. The method of claim 1 , wherein the metadata includes a flag indicating whether the camera position at which an image of the plurality of images was captured is the same as a viewport position associated with the image.
8. The method of claim 1 , wherein the metadata includes an indicator that specifies whether the image attributes associated with an image of the plurality of images are view-independent along one or more axes or directions.
9. 1. A method comprising: receiving a bitstream including data representing a plurality of images, each of the plurality of images being captured from one of a plurality of camera positions, the bitstream including metadata providing an indication of image attributes associated with each of the plurality of camera positions; selecting metadata based on an orientation for rendering one or more images at a decoder, the rendering orientation being different from each of the plurality of camera positions; and rendering the one or more images using the selected metadata.
10. The method of claim 9 , further comprising rendering the one or more images using the image attributes associated with at least one of the plurality of camera positions.
11. The method of claim 9 , wherein the metadata is selected based on an angular distance between the rendering orientation and the multiple camera positions.
12. The method of claim 11 , wherein the selected metadata provides an indication of image attributes associated with a camera position of the plurality of camera positions that has a minimum angular distance to the rendering orientation.
13. The method of claim 11 , wherein the selected metadata provides indications of image attributes associated with two or more of the camera positions of the plurality of camera positions that are within a threshold angular distance of the rendering orientation.
14. 10. The method of claim 9, wherein the metadata is selected based on one or more of the plurality of camera positions, and wherein rendering the one or more images using the selected metadata comprises blending image attribute values together using blend weights.
15. The method of claim 14 , wherein the blending weights are based on the angular distance between the rendering orientation and the one or more of the plurality of camera positions for which metadata is selected.
Citation Information
Patent Citations
Data generation device, image processing device, method and program
JP2019215622A
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
WO2020116563A1