Method and apparatus for encoding and rendering 3D scenes using patch repair
By generating a central view and patch set, and using patch data items for deprojection and patching, the trade-off between visual quality and data size in 3DoF video is resolved, achieving improved visual quality and enhanced immersion in 3DoF+ and 6DoF rendering.
Patent Information
- Application Number
- CN202080033658.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-14
- Filing Date
- 2020-02-25
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2040-02-25
AI Technical Summary
Existing technologies struggle to achieve a good balance between visual quality and data size in rendering unoccluded portions. 3DoF videos cause dizziness and fail to provide a sense of immersion with 6 degrees of freedom, while traditional patching algorithms produce visual artifacts.
By generating a central view and a patch set, the central view includes color and depth components, while the patches only include color components. The patch data items are used for deprojection and patching, avoiding visual artifacts of traditional patching algorithms.
It achieves improved visual quality in 3DoF+ and 6DoF rendering, reduces the bit rate of the data stream, avoids visual artifacts, and supports free navigation and immersion.
Smart Images

Figure CN113906761B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present principles relate generally to the field of three-dimensional (3D) scene and stereoscopic video content. The present document can also be understood in the context of encoding, formatting and decoding of data representative of textures and geometry of a 3D scene, in order to render stereoscopic content on end-user devices such as mobile devices or head-mounted displays (HMD). BACKGROUND
[0002] This section is intended to introduce the reader to various aspects of art that can be related to various aspects of the present principles that are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.
[0003] Recently, available large field of view content (up to 360°) has increased. Such content can not be fully visible for a user watching the content on an immersive display device such as a head-mounted display, smart glasses, a PC screen, a tablet, a smartphone, etc. This means that at a certain moment, the user can only see a part of the content. However, the user can usually navigate in the content through various means such as head movement, mouse movement, touch screen, voice, etc. It is generally desirable to encode and decode such content.
[0004] Immersive video, also called 360° flat video, allows a user to observe everything around himself by rotating his head around a static viewing point. The rotation only allows a 3 degrees of freedom (3DoF) experience. Even if 3DoF video is enough for a first omnidirectional video experience, for example using a head-mounted display device (HMD), 3DoF video can quickly become frustrating for a viewer expecting more degrees of freedom, for example experiencing parallax. Moreover, 3DoF can also induce dizziness because the user has to rotate his head but also to translate his head in three directions, translations that are not reproduced in a 3DoF video experience.
[0005] Large field of view content can be a three-dimensional computer graphics image scene (3D CGI scene), a point cloud or an immersive video. Many terms can be used to design such immersive video: for example, virtual reality (VR), 360, panoramic, 4π-spherical, immersive, omnidirectional or large field of view.
[0006] Stereoscopic video, also known as 6 Degrees of Freedom (6DoF) video, is an alternative to 3DoF video. When watching a 6DoF video, in addition to rotations, the user can also translate his head and even his body in the content he is watching and experience parallax and even stereoscopy. Such videos greatly increase the sense of immersion and perception of depth in the scene and prevent dizziness by providing consistent visual feedback when the head is translated. These contents are created by dedicated sensors that allow the simultaneous recording of the color and depth of the scene of interest. Even if technically still difficult, using a set of color cameras combined with photogrammetry techniques is still the way to perform such recordings.
[0007] While 3DoF video consists of a sequence of images resulting from the de-mapping of a texture image (for example a spherical image encoded according to a longitude / latitude projection mapping or an equirectangular projection mapping), 6DoF video frames embed information from several viewing points. They can be seen as a temporal sequence of point clouds resulting from a three-dimensional capture. Depending on the viewing conditions, two kinds of stereoscopic videos can be considered. The first one, i.e. full 6DoF, allows a full freedom of navigation in the video content, while the second one, also called 3DoF+, limits the viewing space of the user to a limited stereoscopy called viewing bounding box, allowing a limited translation of the head and a parallax experience. This second case is a valuable trade-off between free navigation and passive viewing conditions for a seated audience.
[0008] The points visible from the central viewing point of the 3D scene prepared for 3DoF rendering are needed to render the 3D scene. The points visible from this central viewing point are projected onto a central view image and transmitted to the renderer. The pixels of the central view are not sufficient to allow a parallax experience in 3DoF+ or 6DoF rendering. At rendering time, as long as the viewing point remains in this central position (i.e. in 3DoF rendering mode), the viewport (i.e. the image prepared for display) is fully filled according to the viewing direction. However, as soon as the viewing point is displaced, some data is lost because points not visible from the central viewing point have to be rendered. If this information is not available, some holes on the viewport image corresponding to unoccluded parts in the 3D scene appear, creating annoying visual artifacts. To fill these unoccluded parts, two methods are considered.
[0009] A technical approach to stereoscopic video coding is based on the projection of a 3D scene on multiple 2D images, called patches, which can be further compressed using a traditional video coding standard (e.g. HEVC), packed for instance into atlases. At decoding time, the depth and color information comprised by the pixels of a patch is used to de-project the points of the 3D scene according to the metadata representing their acquisition, i.e. the projection and mapping operations. Each point visible from any viewing point of a predetermined viewing zone and in any viewing direction is encoded as a patch with color components and depth components, the patch being associated with metadata comprising its de-projection parameters. This approach is efficient but requires a large amount of data.
[0010] Another approach consists in filling the holes in the viewport with inpainting techniques, for instance using a patch-based inpainting algorithm as described in C. Barnes et al., "PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing", ACM Trans. on Graphics (Proc. SIGGRAPH), Vol. 28, No. 3, August 2009. The drawback of this approach is that any unoccluded part to be rendered in the viewport has to be inferred from the transmitted visible parts, that is by means of inpainting techniques. This solution produces annoying artifacts at the boundaries of foreground objects.
[0011] There is currently a lack of solution that allows a good trade-off between the visual quality of the rendering of the unoccluded parts and the size of the data required to achieve this goal. The present principles propose such a solution. SUMMARY
[0012] The following presents a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or critical elements of the present principles. The sole purpose of the following summary is to present some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below.
[0013] The present principles relate to a method of encoding a 3D scene according to a viewing zone. The method comprises obtaining a central view of the 3D scene. The view is image data representing the projection on an image plane of points of the 3D scene visible from a viewing point, for example at the center of the viewing zone. The pixels of the view comprise a color component and a depth component encoding the texture and the geometry of the 3D scene as seen from the viewing point. The method further comprises obtaining a set of patches. A patch is associated with data and is image data representing the projection of a portion of the 3D scene visible from a viewing point of the viewing zone, the pixels of the patch having at least a color component encoding the texture of the patch. The data associated with a patch, also called patch data item, comprises a pose description of a virtual camera associated with the projection viewing point, and a distance between the projection viewing point and the portion of the 3D scene. The distance can be determined in different ways. The method comprises encoding in a data stream the view, the set of patches associated with their patch data.
[0014] The present principles also relate to a device implementing the method, and to a stream encoding a 3D scene generated by the method.
[0015] The present principles also relate to a method of rendering a 3D scene for a current viewing point located in a viewing zone. The method comprises decoding from a data stream a view and a set of patches, a patch being associated with data. The view is image data representing the projection on an image plane of points of the 3D scene visible from a viewing point within the viewing zone, for example the center of the viewing zone, the pixels of the view having a color component and a depth component. The patches are image data representing the projection of portions of the 3D scene visible from different viewing points within the viewing zone, the pixels of the patches having at least a color component. The data comprises information representing the pose of a camera corresponding to the projection viewing point used to generate the patch, and the distance from this projection viewing point to this portion of the 3D scene. The method comprises rendering the 3D scene as seen from said current viewing point onto a viewport image by de-projecting the pixels of the view. Since the current viewing point can be different from the viewing point used to project the view, the viewport image comprises at least one scene region filled with information from the view, and at least one unoccluded region for which information is missing because this portion of the 3D space is not visible from the viewing point. The method then comprises patching the unoccluded regions of the viewport image with the patches decoded from the data stream. In an embodiment, the patches are warped according to the distance in the patch data item, the pose of the virtual camera with respect to the current viewing point, and the pose of the virtual camera with respect to the projection viewing point. BRIEF DESCRIPTION OF DRAWINGS
[0016] The present disclosure will be better understood, and further specific features and advantages will become apparent, when the following description reads in conjunction with the annexed drawings, on which:
[0017] Figure 1Points of a point cloud corresponding to a 3D model of an object according to non-limiting embodiments of the present principles are shown;
[0018] Figure 2 Non-limiting examples of encoding, transmitting and decoding data representing a sequence of 3D scenes according to non-limiting embodiments of the present principles are shown;
[0019] Figure 3 Examples of a graph of a stream when data is transmitted through a packet-based transmission protocol according to non-limiting embodiments of the present principles are shown; Figure 10 and Figure 11 Example architecture of a device of the described method;
[0020] Figure 4 Examples of a graph of a stream when data is transmitted through a packet-based transmission protocol according to non-limiting embodiments of the present principles are shown;
[0021] Figure 5 Spherical projection from a central viewing point according to non-limiting embodiments of the present principles is illustrated;
[0022] Figure 6 Examples of atlases including texture information of points of a 3D scene according to non-limiting embodiments of the present principles are shown;
[0023] Figure 7 Examples of atlases including depth information of points of a 3D scene according to non-limiting embodiments of the present principles are shown; Figure 6
[0024] Figure 8 Viewport image generated from a central view for a current viewing point to the right of the central viewing point according to non-limiting embodiments of the present principles is shown;
[0025] Figure 9 Patch being warped for use as input to a patch-based inpainting algorithm according to non-limiting embodiments of the present principles is illustrated;
[0026] Figure 10 Method of encoding a 3D scene for viewing from a viewing point within a viewing zone according to non-limiting embodiments of the present principles is illustrated;
[0027] Figure 11 Method 110 of rendering a 3D scene for a viewing point and a viewing direction within a viewing zone according to non-limiting embodiments of the present principles is illustrated. DETAILED DESCRIPTION
[0028] The present principles will be described more fully hereinafter with reference to the accompanying drawings, in which example embodiments of the present principles are shown. The present principles may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present principles to those skilled in the art. Like numbers refer to like elements throughout.
[0029] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the present principles. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including" when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Additionally, when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or indirectly responsive or connected to the other element through one or more other elements. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to another element, there are no intervening elements.
[0030] It will be understood that, although the terms first, second, etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the teachings of the present principles.
[0031] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it is to be understood that communication can occur in the opposite direction to the depicted arrows.
[0032] Some examples are described with respect to block diagrams and operational flowcharts, in which each block represents circuitry, modules, or portions of code which include one or more executable instructions for implementing the specified (one or more) logical functions. It should also be noted that in other embodiments, the functions noted in the blocks can occur out of the order noted in the flowcharts. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in reverse order, depending on the functionality involved.
[0033] Reference to“according to an example” or“in an example” in this text means that a particular feature, structure, or characteristic described in connection with the example can be included in at least one implementation of the present principles. The appearances of the phrase“according to an example” or“in an example” in various places in the specification are not necessarily all referring to the same example, nor are separate or alternative examples necessarily mutually exclusive of one another.
[0034] Reference signs appearing in the claims are merely illustrative, and do not limit the scope of the claims. Although examples and variants are presented, they can be combined or sub-combined in any way, even if not explicitly described.
[0035] In the following, a viewpoint is a point from which a camera or a virtual camera can work when performing a view from it. A viewpoint is the association of a 3D coordinate of a point with a viewing direction. Thus, a viewpoint refers to the data that a camera captures of a scene, including the position and viewing direction of the camera and the projection parameters of the camera.
[0036] According to the present principles, a central view of a 3D scene is captured from a central viewpoint. The central view is image data representing the 3D scene as viewed from the central viewpoint. The pixels of the central view include a depth component to be de-projected and a color component. In a variant, two central views are generated, one including color information and one including depth information. According to the present principles, patches are generated and encoded in association with the central views with their metadata, a patch being image data representing a portion of the 3D scene, in particular a portion of the 3D scene that is not visible from the central viewpoint but is visible from another viewpoint located in the viewing zone. A patch has only a color component. The decoded central views are rendered and, depending on the position and direction of the current view, a viewport image is generated using the points de-projected from the pixels of the central view. According to the present principles, holes corresponding to unoccluded portions of the 3D scene are filled by using a patch-based inpainting method that uses color patches decoded from the stream to progressively fill the unoccluded portions. The advantage of the present principles is that, since the patches are not extrapolated from the central view but are taken from the original 3D scene, visual artifacts due to the inpainting algorithm can be avoided, while, since the depth information relative to the patches is not encoded, the bit rate of the data stream encoding the 3D scene can be limited.
[0037] Figure 1A three-dimensional (3D) model 10 of an object and points of a point cloud 11 corresponding to the 3D model 10 are shown. The 3D model 10 and the point cloud 11 can for example correspond to a possible 3D representation of an object in a 3D scene comprising other objects. The model 10 can be a 3D mesh representation, the points of the point cloud 11 can be the vertices of the mesh. The points of the point cloud 11 can also be points distributed on the surface of the faces of the mesh. The model 10 can also be represented as a stitched version of the point cloud 11, the surface of the model 10 being created by stitching the points of the point cloud 11. The model 10 can be represented in many different representations, like voxels or splines. Figure 1 The figure illustrates the fact that a point cloud can be defined with a surface representation of a 3D object and that a surface representation of a 3D object can be generated from the points of the cloud. As used herein, projecting points of a 3D object (by extension points of a 3D scene) onto an image is equivalent to projecting any representation of this 3D object, for example a point cloud, a mesh, a spline model or a voxel model.
[0038] A point cloud can be represented in memory, for example as a vector-based structure where each point has its own coordinates in the frame of reference of the point of view (for example three-dimensional coordinates XYZ, or entity angle and distance from / to the point of view (also called depth) and one or more attributes, also called components. An example of a component is a color component, which can be expressed in various color spaces, for example RGB (red, green, blue) or YUV (Y is the luminance component, UV are two chrominance components). The point cloud is a representation of a 3D scene comprising an object. The 3D scene can be seen from a given point of view or a series of points of view. The point cloud can be obtained in many ways from, for example:
[0039] - capture of a real object by a set of cameras, optionally supplemented by depth active sensing devices;
[0040] - capture of a virtual / synthetic object by a set of virtual cameras in a modeling tool;
[0041] - a mix of both real and virtual objects.
[0042] Figure 2 Non-limiting examples of encoding, transmission and decoding of data representing a sequence of 3D scenes are shown. The encoding format can for example and simultaneously be compatible with 3DoF, 3DoF+ and 6DoF decoding.
[0043] A sequence of 3D scenes 20 is obtained. Since a sequence of pictures is a 2D video, the sequence of 3D scenes is a 3D (also called stereoscopic) video. The sequence of 3D scenes can be provided to a stereoscopic video rendering device for 3DoF, 3Dof+ or 6DoF rendering and display.
[0044] 3D scene sequence 20 is provided to an encoder 21. Encoder 21 takes as input a 3D scene or a sequence of 3D scenes and provides a bitstream representing the input. The bitstream can be stored in a memory 22 and / or on an electronic data medium and can be transmitted through a network 22. The bitstream representing the sequence of 3D scenes can be read from memory 22 and / or received from network 22 by a decoder 23. Decoder 23 takes as input the bitstream and provides a sequence of 3D scenes, for example in point cloud format.
[0045] Encoder 21 can comprise several circuits implementing several steps. In a first step, encoder 21 projects each 3D scene onto at least one 2D picture. A 3D projection is any method of mapping three-dimensional points into a two-dimensional plane. The use of this type of projection is widespread, especially in computer graphics, engineering and cartography, since most methods of displaying graphical data are currently based on two-dimensional media (pixel information from several bitplanes). Projection circuit 211 provides at least one two-dimensional frame 2111 for the 3D scenes of sequence 20. Frame 2111 comprises color information and depth information representing the 3D scene projected onto frame 2111. In variants, color information and depth information are encoded in two separate frames 2111 and 2112.
[0046] Metadata 212 are used and updated by projection circuit 211. Metadata 212 comprise information about the projection operation (e.g. projection parameters) and information about the way color and depth information are organized in frames 2111 and 2112, as described in the detailed description of the application. Figures 5 to 7
[0047] Video encoding circuit 213 encodes the sequence of frames 2111 and 2112 as a video. The pictures of 3D scenes 2111 and 2112 (or the sequence of pictures of a 3D scene) are encoded as a stream by video encoder 213. Video data and metadata 212 are then encapsulated in a data stream by data encapsulation circuit 214.
[0048] Encoder 213 is for example compliant with an encoder such as:
[0049] - JPEG, specification ISO / CEI 10918-1 UIT-T proposal T.81, https: / / www.itu.int / rec / T-REC-T.81 / en;
[0050] - AVC, also called MPEG-4 AVC or h264. Specified in both UIT-T H.264 and ISO / CEI MPEG-4 Part 10 (ISO / CEI 14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (whose specification can be found on the ITU website, T Recommendation, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en);
[0051] - 3D-HEVC (extension of HEVC, whose specification can be found on the ITU website, T Recommendation, H series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annexes G and I);
[0052] - VP9 developed by Google; or
[0053] - AV1 (AOMedia Video 1) developed by the Alliance for Open Media.
[0054] The data stream is stored in a memory accessible by the decoder 23, for example through the network 22. The decoder 23 comprises different circuits implementing the different steps of the decoding. The decoder 23 takes as input the data stream generated by the encoder 21 and provides a sequence of 3D scenes 24, rendered and displayed by a stereoscopic video display device, such as a head-mounted device (HMD). The decoder 23 obtains the stream from a source 22. For example, the source 22 belongs to the set comprising:
[0055] - a local memory, for example a video memory or a RAM (or Random Access Memory), a flash memory, a ROM (or Read Only Memory), a hard disk;
[0056] - a storage interface, for example an interface with a mass storage, a RAM, a flash memory, a ROM, an optical disk or a magnetic support;
[0057] - a communication interface, for example a wired interface (for example a bus interface, a wide area network interface, a local area network interface) or a wireless interface (such as an IEEE 802.11 interface or a Bluetooth® interface); and - a user interface, such as a graphical user interface, enabling a user to input data.
[0058] - a user interface, such as a graphical user interface, enabling a user to input data.
[0059] The decoder 23 comprises a circuit 234 for extracting the data encoded in the data stream. The circuit 234 takes as input the data stream and provides metadata 232 corresponding to the metadata 212 encoded in the data stream and a two-dimensional video. The video is decoded by a video decoder 233 which provides a sequence of frames. The decoded frames comprise color and depth information. In a variant, the video decoder 233 provides two sequences of frames, one comprising color information and the other comprising depth information. The circuit 231 uses the metadata 232 to de-project the color and depth information from the decoded frames to provide a sequence of 3D scenes 24. The sequence of 3D scenes 24 corresponds to the sequence of 3D scenes 20 with possible loss of precision related to the encoding as a 2D video and the video compression.
[0060] According to the present principles, the circuit 211 generates a frame called central view by projecting the points of the 3D scene visible from the central viewing point onto the image plane. The pixels of the central view have a color component and a depth component representing the color of the projected point and the distance between the projected point and the central viewing point in the 3D scene space. In a variant, two central views are generated, one encoding the color component and the other encoding the depth component of the central projection. Other parts of the 3D scene, visible from other points in the 3D scene and according to several viewing directions, are encoded as images called patches. Only the color component is encoded in the patches. The projection can be the same as for the central view or different. The metadata describing the projection parameters for the patches are associated with the patches. In addition, the patch data item comprises the distance between the viewing point from which the patch has been taken and the part of the 3D scene projected onto this patch. This distance is, for example, the average distance or the median distance, or the shortest distance or the longest distance, between the point and the points of the part of the 3D scene.
[0061] According to the present principles, the central view and the patches are decoded by the circuits 234 and 233 and the metadata 232 is retrieved from the data stream. Knowing the coordinates of the central viewing point (for example, the origin of the 3D space reference frame of the 3D scene), and according to the projection parameters of the central view and the depth and color information encoded in the central view, the pixels of the central view are de-projected. According to the present principles, the viewport image is generated from these de-projected points according to the current position and viewing direction. In 3DoF+ or 6DoF rendering, the current viewing point can be different from the central viewing point and there are holes in the viewport image corresponding to the unoccluded parts. According to the present principles, these holes are filed by using a patch-based algorithm which does not use the patches created from the current viewport image but uses the color patches decoded from the stream according to the associated patch data.
[0062] Figure 3 An example architecture of a device implementing the method described with respect to Figure 10 and Figure 11 An example architecture of a device implementing the method described with respect to Figure 2The encoder 21 and / or the decoder 23 can implement such an architecture. In addition, each circuit of the encoder 21 and / or the decoder 23 can be a device according to the architecture of Figure 3 linked together, for example, via its bus 31 and / or via the I / O interface 36.
[0063] The device 30 comprises the following elements, linked together by a data and address bus 31 :
[0064] - a microprocessor 32 (or CPU), for example, it is a DSP (or Digital Signal Processor);
[0065] - a ROM (or Read Only Memory) 33;
[0066] - a RAM (or Random Access Memory) 34;
[0067] - a storage interface 35;
[0068] - an I / O interface 36 for receiving data to be transmitted from an application; and
[0069] - a power supply, for example a battery.
[0070] According to an example, the power supply is external to the device. In each of the memories mentioned, the word "register" used in the description can correspond to an area of small capacity (a few bits) or very large (for example, the entire program or a large amount of received or decoded data). The ROM 33 comprises at least the program and the parameters. The ROM 33 can store algorithms and instructions to perform the techniques according to the present principles. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.
[0071] The RAM 34 comprises the program executed in register by the CPU 32 and uploaded in the RAM after switching on of the device 30, the input data in register, the intermediate data of the different states of the method in register, and other variables used to execute the method in register.
[0072] The embodiments described herein can be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or as an apparatus), the implementation discussed can also be implemented in other forms (for example, a program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, an apparatus such as, for example, a processor, which is to be understood in a broad sense to encompass any processing device, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0073] According to the example, device 30 is configured to implement about Figure 10 and Figure 11 The methods described belong to the following set:
[0074] -mobile device;
[0075] - Communication equipment;
[0076] -Gaming devices;
[0077] - Tablet PC (or tablet computer);
[0078] - Laptop;
[0079] - Still image camera;
[0080] -Camera;
[0081] - Encoding chip;
[0082] - Servers (e.g., broadcast servers, video-on-demand servers, or web servers).
[0083] Figure 4 An example of an embodiment of the syntax of a stream is shown when data is transmitted via a packet-based transport protocol. Figure 4 Example structure 4 for a stereoscopic video stream is shown. This structure exists within a container, organizing the stream into individual syntax elements. This structure may include a header section 41, which is a dataset common to each syntax element of the stream. For example, the header section includes some metadata about the syntax elements, describing the properties and roles of each of them. The header section may also include... Figure 2 The metadata 212 in the image contains, for example, coordinates used to project points of the 3D scene onto the central viewing point on frames 2111 and 2112. The structure includes a payload comprising syntax element 42 and at least one syntax element 43. Syntax element 42 includes data representing color and depth frames. The image may have been compressed according to a video compression method.
[0084] Syntax element 43 is part of the payload of the data stream and may include metadata about how the frames of syntax element 42 are encoded, such as parameters for projecting points of the 3D scene and packing them onto the frames. This metadata may be associated with each frame or group of frames of the video (also known as a group of pictures (GoP) in video compression standards).
[0085] Figure 5The illustration shows a patch atlas method with an example of four projection centers. A 3D scene 50 includes characters. For example, projection center 51 is a perspective camera, and camera 53 is an orthographic camera. The camera could also be an omnidirectional camera, such as a spherical map (e.g., an isometric map) or a cube map. Based on the projection operations described in the projection data of the metadata, 3D points of the 3D scene are projected onto a 2D plane associated with a virtual camera located at the projection center. Figure 5 In the example, the projection of the point captured by camera 51 is mapped onto patch 52 according to perspective mapping, while the projection of the point captured by camera 53 is mapped onto patch 54 according to orthographic mapping.
[0086] Clustering of projected pixels generates multiple 2D patches, which are packaged in a rectangular atlas 55. The organization of patches within the atlas defines the layout of the atlas. In an embodiment, two atlases have the same layout: one for texture (i.e., color) information and one for depth information. Two patches captured by the same camera or two different cameras may include information representing the same part of a 3D scene, for example, patches 54 and 56.
[0087] The packing operation generates patch data for each generated patch. Patch data includes a reference to the projection data (e.g., an index in a projection data table or a pointer to the projection data (i.e., an address in memory or a data stream)) and information describing the patch's position and size within the atlas (e.g., top-left corner coordinates, pixel dimensions, and width). Patch data items are added to metadata and encapsulated in the data stream along with compressed data from one or both atlases.
[0088] Figure 6 An example of an atlas 60 comprising texture information (also known as color components, such as RGB or YUV data) of points in a 3D scene, according to a non-limiting embodiment of this principle, is shown. (See also: Regarding...) Figure 5 The atlas, as explained, is a collection of images that are patched together; a patch is an image obtained by projecting a portion of a point from a 3D scene.
[0089] exist Figure 6 In the example, atlas 60 includes a central view 61 and one or more second parts 62. The central view 61 includes texture information of points in the 3D scene visible from the viewing point. The texture information of the central view 61 can be obtained, for example, from an isometric projection map, which is an example of a spherical projection map. Figure 6In the example of Fig. 6, the second portion 62 is arranged at the left and right boundaries of the central view 61, but the second portion can be arranged differently. The second portion 62 comprises the texture information of the portion of the 3D scene complementary to the portion visible from the viewing point. The second portion can be obtained by removing from the 3D scene the points visible from the first viewing point, whose texture is stored in the first portion, and by projecting the remaining points from the same viewing point. This latter process can be iteratively repeated so as to obtain each time the hidden portion of the 3D scene. According to a variant, the second portion can be obtained by removing from the 3D scene the points visible from a viewing point, e.g. the central viewing point, whose texture is stored in the first portion, and by projecting the remaining points from a viewing point different from the first viewing point, e.g. from one or more second viewing points of a viewing space centered on the central viewing point, e.g. the viewing space of the 3DoF rendering.
[0090] The central view 61 can be seen as a first large texture patch (corresponding to the first portion of the 3D scene), the second portion 62 comprising smaller texture patches (corresponding to the second portion of the 3D scene complementary to the first portion). This atlas has the advantage of being compatible with both 3DoF rendering (when only the central view 61 is rendered) and 3DoF+ / 6DoF rendering.
[0091] Figure 7 Fig. 7 shows an example of a central view 71 comprising depth information of the points of a 3D scene visible from a central viewing point, according to a non-limiting embodiment of the present principles. The central view 71 comprises depth information of the points of the 3D scene visible from the central viewing point. According to the present principles, Figure 6 The depth information of the patches 62 is not encoded in the stream. So, there is no depth patch corresponding to the color patch 62. The central view 71 can be obtained in the same way as the central view 61, but comprises depth information associated with the points of the 3D scene instead of texture information. Figure 6
[0092] For 3DoF rendering of a 3D scene, only one viewing point is considered, typically the central viewing point. The user can rotate his head around the first viewing point with three degrees of freedom to observe various portions of the 3D scene, but the user cannot move this unique viewing point. The points of the scene to be encoded are the points visible from this unique viewing point, for which only the texture information needs to be encoded / decoded for 3DoF rendering. For 3DoF rendering, there is no need to encode the points of the scene that are not visible from this unique viewing point, as they are not accessible to the user.
[0093] For 3DoF+ or 6DoF rendering, the depth information of the points not visible from the central views 61 and 71 is missing. Figure 2 The circuit 231 of the renderer 23 produces a viewport image from the position and orientation of the current viewpoint, but holes appear in the viewport image, the holes corresponding to parts of the 3D scene that are not occluded by the displacement of the viewpoint from the central pose to the current pose. According to the present principles, these holes are filled by a patch-based inpainting algorithm that uses color patches 62, instead of patches created from the generated viewport image.
[0094] Figure 8 A viewport image 80 generated from the central view for a current viewpoint to the right of the central viewpoint is shown. The pixels of the viewport 80 are set by the capture of the de-projected points of the 3D scene by a virtual camera located at the current viewpoint and oriented in the viewing direction. In Figure 8 In the example, the current viewpoint is displaced 10 centimeters to the right. Therefore, a part 81 of the viewport cannot be filled because there are no points to capture in this place according to the current position and viewing direction. The area 81 corresponds to a part of the 3D scene that is unoccluded when the viewpoint is displaced from the central position to the right of the current position. According to the present principles, the viewport is rendered by first de-projecting the point cloud reconstructed from the central texture plus depth view. As the viewpoint is displaced from the central viewpoint, unoccluded holes appear in the rendered viewport, as Figure 8 depicted. Then, the front edge of each unoccluded hole to fill is segmented between the foreground object boundary 82 (no continuity) and the background boundary 85 (called "filling front", continuity must be enforced) because the depth information is available on this boundary. The clustering of the two classes of depth values produces this segmentation. State-of-the-art exemplar-based or patch-based inpainting algorithms are adapted to fill the unoccluded holes, taking candidate patches from the occluded texture patches instead of in the already projected visible background. Starting from the filling front, the best matching patch 83 is iteratively stitched to the filling front until the hole is completely inpainted. The patch 83 is chosen as the best fitting patch according to its position in the 3D space; this position is retrieved by using the information of the patch data item associated with this patch.
[0095] Figure 9 A patch distorted for use as input of a patch-based inpainting algorithm is illustrated. The patch 91 is image data representing the projection 92 of a part 90 of the 3D scene on an image plane. The projection 92 is equivalent to the acquisition of the part 90 of the 3D scene by a camera 92. The position, orientation and parameters of the projection 92 are formatted in a segment of metadata, called patch data item, and encoded in association with the patch image data 91. The part 90 can not be flat. According to the present principles, only the color components of the points of the part 90 are projected on the patch 91. Therefore, the depth information is lost. Only the distance 93 between the projection center 92 and the part 90 is stored in the patch data item.
[0096] At the decoding side, the texture patch (i.e. color patch) is post-processed to dynamically adapt the appearance of the patch to the current viewpoint. To be used as input for the patch-based inpainting algorithm, the patch is warped to take into account the change of viewpoint between the virtual camera 92 that captured the patch 91 and the current virtual camera 94 used to generate the viewport image. Assuming a planar patch surface model (and a perspective projection model), such warping can be reduced to a homography whose coefficients only depend on the translation and rotation between the two cameras and the distance 93. The following equations provide the coefficients of such parametric transformation, from the translation T and rotation Ω of the cameras, the distance 93 to the patch Z (and the camera focal length f and pixel aspect ratio a).
[0097]
[0098] where:
[0099]
[0100] In the context of the simplified 3DoF+ transmission format, this principle guarantees the reconstruction of the unoccluded areas even in cases where the propagation of a simple inpainting from the background would provide annoying visual artifacts. This enables valuable trade-offs between the complexity of the decoder, the transmission bitrate and the visual quality.
[0101] Figure 10A method of encoding a 3D scene intended to be viewed from a viewing point within a viewing zone is illustrated according to a non-limiting embodiment of the present principles. In step 101, image data representative of a central view of the 3D scene is obtained. The central view is obtained by projecting points of the 3D scene visible from a viewing point determined to be central to the viewing zone onto an image plane. Different projection modes are suitable for this operation, such as sphere projection and mapping, cube projection and mapping, perspective projection or orthogonal projection. The projection mode is chosen according to the viewing zone, for example whether the user has the possibility to look behind him with respect to the initial viewing direction. The pixels of the central view comprise at least a color component and a depth component. The color component of the image pixels forms a texture of the image. In a variant, two central views are obtained, one carrying texture information and one carrying depth information. In step 102, a set of patches is obtained, each patch being associated with a patch data item. The set of patch data items forms a segment of metadata called patch data. A patch is image data representative of a portion of the 3D scene projected onto the image plane. In this sense, the central view is a large patch. The pixels of a patch comprise a color component. In a preferred embodiment, the patches carry only texture information, no depth information. This has the advantage that the bitrate of the stream can be reduced. The patch data item comprises data representative of the projection used to generate the patch and the distance from the center of the projection, i.e. the point in the viewing zone, to the projected portion of the 3D scene. This distance can be for example the average distance between the points of the projected portion and the center of the projection. In a preferred embodiment, the clustering algorithm used to generate the patches is parameterized to enforce patches small enough so that they can be approximated with planar surfaces. The patches are generated in order that they encode points of the 3D scene that are not visible from the central viewing point but that are visible from another viewing point of the viewing zone. In other words, the patches encode texture information of potentially unoccluded portions of the possible viewports on the rendering side. In step 103, the central view(s), the patches and the metadata are encoded in a data stream.
[0102] Figure 11A method 110 of rendering a 3D scene for a viewing point and a viewing direction within a viewing zone according to a non-limiting embodiment of the present principles is illustrated. At step 111, a data stream encoding a 3D scene is obtained according to the present principles. From this stream, a central view and a set of patches are decoded. The central view is image data representing the projection onto an image plane of points of the 3D scene visible from a 3D center point at the center of the viewing zone. The pixels of the central view comprise a color component and a depth component. A patch is image data representing the projection of a portion of the 3D scene, the pixels of the patch having a color component. The patch is associated with a segment of metadata, called a patch data item, comprising parameters representing the projection and a distance from the center of the projection to the portion of the 3D scene. At step 112, the 3D scene is rendered onto a viewport image viewed from a current viewing point. The viewport image is generated from points of the 3D scene retrieved from the central view according to the current viewing point. Since the current viewing point can be offset with respect to the central viewing point in the viewing zone, the viewport image comprises a scene area (i.e. pixels valorized with the color component of the projected point) and an unoccluded area (i.e. pixels not valorized because they correspond to a region of the 3D scene not visible from the central viewing point, so not encoded in the central view), as illustrated in Figure 8 step 113, the patch is warped according to the relative pose of the virtual camera associated with the projection viewing point from which the patch has been captured (and described in the associated patch data item), and the current viewing point and the distance are stored in the patch data item. Then, the warped patch is input into an exemplar-based or patch-based inpainting algorithm adapted to fill the unoccluded hole, taking candidate patches from the warped patch instead of in the already projected visible background.
[0103] Embodiments described herein can be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method or as a device), the implementation of features discussed can also be implemented in other forms (for example, a program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, an apparatus such as, for example, a processor, which is an example of a processing device. A processor includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, a smartphone, a tablet, a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate communication of information among end-users.
[0104] Implementations of the various processes and features described herein can be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications related to data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and / or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video coder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a phone, a PDA, and other communication devices. As should be apparent, these devices can be mobile, even if not always in motion.
[0105] Moreover, these methods can be implemented by way of machine, programmable, processor, or computer readable instructions, and / or other computer program product(s). Some embodiments can therefore comprise a computer program product for use by or to control the operation of a processor. By way of example, a computer program product can comprise a computer readable medium having instructions stored for execution by a processor. The computer readable medium can comprise a hard disk, a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc, a flash memory, a ROM, a RAM, a PROM, an EPROM, a EEPROM, a solid state drive, a cache, a register, any combination thereof, or the like. The computer program product can comprise instructions for implementing a process, or a portion thereof, or the like. The computer program product can comprise a computer readable medium having instructions stored for execution by a processor, the instructions comprising instructions for implementing a process, or a portion thereof, or the like.
[0106] As will be evident to one of skill in the art, implementations can produce signals formatted to carry information that can be, for example, stored or transmitted. For example, a signal can include instructions for carrying out a method, or data produced by one of the described embodiments. For example, a signal can be formatted to carry data, which, when loaded into a machine (for example, a processor) can cause the machine to facilitate carrying out a process described herein. A signal can also carry another signal. For example, a modem can receive a signal carrying the instructions described herein, and can format the instructions as a telephone signal and transmit them to another modem. A signal can be analog or digital. The various implementations can modulate a carrier with a signal, produce the signal on a conductive medium, or produce the signal on a carrier. As noted above, a signal can be formatted to carry data. The data can comprise, for example, instructions to implement a process, or data produced by one of the described embodiments. For example, a signal can be formatted to carry data, which, when loaded into a machine (for example, a processor) can cause the machine to facilitate carrying out a process described herein. A signal can also carry another signal. For example, a modem can receive a signal carrying the instructions described herein, and can format the instructions as a telephone signal and transmit them to another modem. A signal can be analog or digital. The various implementations can modulate a carrier with a signal, produce the signal on a conductive medium, or produce the signal on a carrier.
[0107] A number of implementations have been described. Nevertheless, it will be understood that various modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill will understand that other structures and processes can be substituted for those disclosed and the resulting implementations will perform at least substantially the same function(s), in at least substantially the same way(s), to achieve at least substantially the same result(s) of the implementations disclosed. Accordingly, these and other implementations are contemplated by this application.
Claims
1. A rendering method, comprising: Decode the view and at least one patch from the data stream. The view is image data representing the points of a 3D scene visible from a first viewing point at the center of the viewing area, projected onto an image plane according to an omnidirectional mapping, wherein the pixels of the view have color and depth components. A patch is image data representing a portion of the 3D scene projected from a second viewing point within the viewing area, wherein the pixels of the patch have color components; and The patch is associated with data, which includes the position and orientation of the second viewing point and the distance between the second viewing point and the portion of the 3D scene; The pixels of the view are projected onto an image based on the rendering viewpoint, the image including at least one scene area and at least one unobstructed area; Based on the data, a patch among the at least one patch is iteratively selected by segmenting the foreground and background boundaries of the unoccluded area; as well as At least one patch is used to repair at least one unoccluded area of the image.
2. The method according to claim 1, wherein, The patching includes distortion of the patch based on the distance, the second viewing point, and the rendered viewing point.
3. The method according to claim 2, wherein, The distortion of the patch is due to homography parameterized by the following: The distance from the point to the portion of the 3D scene, and Translation and rotation between the rendering viewing point and the second viewing point.
4. The method according to any one of claims 1 to 3, wherein, The repair includes: Segmenting the unoccluded region in the viewport image between the foreground and background boundaries; and The twisted patch in the unoccluded region is iteratively stitched from the background boundary to the foreground boundary.
5. A rendering device, including a processor, the processor being configured to: Decode the view and at least one patch from the data stream. The view is image data representing the points of a 3D scene visible from a first viewing point at the center of the viewing area, projected onto an image plane according to an omnidirectional mapping, wherein the pixels of the view have color and depth components. A patch is image data representing a portion of the 3D scene projected from a second viewing point within the viewing area, wherein the pixels of the patch have color components; and The patch is associated with data, which includes the position and orientation of the second viewing point and the distance between the second viewing point and the portion of the 3D scene; The pixels of the view are projected onto an image based on the rendering viewpoint, the image including at least one scene area and at least one unobstructed area; Based on the data, a patch among the at least one patch is iteratively selected by segmenting the foreground and background boundaries of the unoccluded area; as well as At least one patch is used to repair at least one unoccluded area of the image.
6. The device according to claim 5, wherein, The processor is configured to distort the patch based on the distance, the second viewing point, and the rendered viewing point to perform the patching.
7. The device according to claim 6, wherein, The processor is configured as a twisted patch with homography parameterized by the following: The distance from the point to the portion of the 3D scene, and Translation and rotation between the rendering viewing point and the second viewing point.
8. The device according to any one of claims 5 to 7, wherein, The processor is configured to perform the patch by: Segmenting the unoccluded region in the viewport image between the foreground and background boundaries; and The twisted patch in the unoccluded region is iteratively stitched from the background boundary to the foreground boundary.
9. An encoding method, comprising: Obtain a view of the 3D scene. The view is image data representing the points of the 3D scene visible from a first viewing point at the center of the viewing area, projected onto an image plane according to an omnidirectional mapping, wherein the pixels of the view have color and depth components. Obtain the supplementary film set. A patch is image data representing a portion of the 3D scene visible from a second viewing point within the viewing area, and the pixels of the patch have color components; The patch is associated with data, which includes the position and orientation of the second viewing point and the distance between the second viewing point and the portion of the 3D scene; as well as The view, the patch set, and the data are encoded in the data stream.
10. The method according to claim 9, wherein, The distance between the second viewing point and the portion of the 3D scene is the average distance, median distance, shortest distance, or longest distance between the points of the second viewing point and the portion of the 3D scene.
11. An encoding device, comprising a processor, the processor being configured to: Obtain a view of the 3D scene. The view is image data representing the points of the 3D scene visible from a first viewing point at the center of the viewing area, projected onto an image plane according to an omnidirectional mapping, wherein the pixels of the view have color and depth components. Obtain the supplementary film set. A patch is image data representing a portion of the 3D scene visible from a second viewing point within the viewing area, and the pixels of the patch have color components; The patch is associated with data, which includes the position and orientation of the second viewing point and the distance between the second viewing point and the portion of the 3D scene; as well as The view, the patch set, and the data are encoded in the data stream.
12. The device according to claim 11, wherein, The processor determines that the distance between the second viewing point and the portion of the 3D scene is the average distance, median distance, shortest distance, or longest distance between points of the second viewing point and the portion of the 3D scene.
13. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method according to claim 1 or 9.
Citation Information
Patent Citations
Methods, devices and stream for encoding and decoding volumetric video
EP3432581A1
Target Region Fill Utilizing Transformations
US20150097827A1