METHOD AND APPARATUS FOR ENCODING, TRANSMITTING AND DECODING VOLUMETRIC VIDEO - Patent application
The method of encoding and decoding pruned multi-view frames using an acyclic graph addresses the limitations of 3DoF and 6DoF video technologies, providing seamless navigation and reducing visual artifacts for enhanced immersive experiences.
Patent Information
- Application Number
- JP2022518235
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-07
- Filing Date
- 2020-09-22
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2040-09-22
AI Technical Summary
Existing 3DoF video technologies fail to provide a seamless immersive experience due to limited navigation freedom and potential dizziness, while 6DoF video technologies face challenges in efficiently encoding and decoding non-redundant multi-view images for optimal bitstream and rendering quality.
A method for encoding and decoding pruned multi-view frames using an acyclic graph to prioritize view pruning, ensuring efficient data transmission and accurate view synthesis by maintaining non-redundant and complementary pixel information.
Enhances immersive experiences by allowing seamless navigation and reducing visual artifacts, achieving efficient data compression and high-quality rendering for 3D scenes on devices like head-mounted displays.
Smart Images

Figure 0007744334000006 
Figure 0007744334000007 
Figure 0007744334000008
Abstract
Description
[Technical Field]
[0001] The present principles generally relate to three-dimensional (3D) scenes and Volumetric This document also relates to the domain of video content on end-user devices such as mobile devices or head-mounted displays (HMDs). Volumetric For content rendering, 3D scene textures and Encoding, formatting and processing of data representing geometric shapes Decryption Among other topics, the present principles relate to pruning pixels of multi-view images to ensure optimal bitstream and rendering quality. [Background technology]
[0002] This section is intended to introduce the reader to various aspects of art that may be related to various aspects of the present principles, which are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.
[0003] In recent years, there has been a growth in available large field-of-view content (up to 360°). Such content may not be fully visible to users viewing the content on immersive display devices such as head-mounted displays, smart glasses, PC screens, tablets, or smartphones. This means that at any given moment, only a portion of the content is visible to the user. However, users can typically navigate within the content by various means, such as head movements, mouse movements, touchscreens, and voice. It is typically desirable to encode and decode this content.
[0004] Immersive video, also known as 360° flat video, allows users to view everything around them through head rotation around a stationary point. Rotation only allows for a three-degrees-of-freedom (3DoF) experience. For example, even if 3DoF video is sufficient for a first-order omnidirectional video experience using a head-mounted display device (HMD), it can quickly become frustrating for viewers who expect more degrees of freedom, such as by experiencing parallax. Furthermore, 3DoF can also induce dizziness because users not only rotate their head but also translate it in three directions, a translation that is not reproduced in a 3DoF video experience.
[0005] The large field-of-view content can be, among others, a three-dimensional computer graphic image scene (3D CGI scene), a point cloud, or an immersive video. Many terms can be used to design such immersive video, such as virtual reality (VR), 360, panoramic, 4π steradian, immersive, omnidirectional, or large field-of-view.
[0006] Volumetric Six-degrees-of-freedom (6DoF) video is an alternative to 3DoF video. When watching 6DoF video, in addition to rotation, users can also translate their head and even their body within the viewed content, experiencing parallax and even volume. Such video significantly increases the sense of immersion and the perception of scene depth, and prevents dizziness by providing consistent visual feedback during head translation. Content is created by means of dedicated sensors that allow simultaneous recording of the color and depth of the desired scene. The use of color camera rigs combined with photogrammetry techniques is a method for performing such recording, even though technical difficulties remain.
[0007] While 3DoF video contains a sequence of images resulting from the unmapping of texture images (for example spherical images encoded according to a latitude / longitude projection mapping or equirectangular mapping), 6DoF video frames embed information from several viewpoints. They can be viewed as a temporal sequence of points resulting from three-dimensional capture. Depending on the viewing conditions, there are two types of 6DoF video: Volumetric Consider video. The first (i.e., full 6DoF) allows for complete freedom of navigation within the video content, while the second (also known as 3DoF+) restricts the user's visual space to a limited volume called the visual bounding box, allowing for a limited volume of head and parallax experiences. This second context is a valuable trade-off between freedom of navigation and passive viewing conditions for seated audience members.
[0008] 3DoF+ content can be provided as a set of Multi-View+Depth (MVD) frames. Such content may be captured by a dedicated camera, or it can be generated from existing computer graphics (CG) content by dedicated (potentially photorealistic) rendering. Volumetric The information is conveyed as a combination of color and depth patches stored in corresponding color and depth atlases, which are then video encoded using a codec (e.g., HEVC). Each combination of color and depth patches represents a portion of the MVD input view, and the set of all patches is designed in the encoding stage to cover the entire scene with as little redundancy as possible. Decryption Atlas first introduced the video Decryption The patches are then rendered in a view synthesis process to recover the viewport associated with the desired viewing position. A problem with such a solution relates to how the patches can be created so that they are sufficiently non-redundant and complementary. Summary of the Invention
[0009] The following presents a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or critical elements of the present principles. The following summary merely presents some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below.
[0010] The present principles relate to a method for encoding pruned multiview frames in a data stream, the method comprising: - obtaining an acyclic graph connecting the views of the unpruned multiview frame, the links of the graph representing view pruning priorities; pruning pixels of the views of the multi-view image in a determined order such that a first view is pruned after views connected to it by pruning priority links; - encoding the graph and the pruned view in the data stream.
[0011] The present principles also relate to a device comprising a processor configured to carry out this method.
[0012] The present principles also relate to a method for decoding multi-view frames pruned from a data stream, the method comprising: obtaining pruned multiview frames from the data stream; - obtaining an acyclic graph from the data stream, the graph connecting views of the multi-view image, and links in the graph representing view pruning priorities; generating a viewport frame according to the viewing pose by determining the contribution of each view of the pruned multiview frame as a function of the pruning priority of the graph.
[0013] The present principles also relate to a device comprising a processor configured to carry out this method.
[0014] The present principles also provide a data stream comprising: data representing pruned multiview frames; and - data representing an acyclic graph, the graph connecting views of a multi-view image, and the links of the graph representing view pruning priorities. [Brief explanation of the drawings]
[0015] The present disclosure will be better understood, and other particular features and advantages will become apparent, on reading the following description, which makes reference to the accompanying drawings, in which: [Figure 1] 1 illustrates a three-dimensional (3D) model of an object and points of a point cloud corresponding to the 3D model, in accordance with a non-limiting embodiment of the present principles. [Figure 2] 1 shows a non-limiting example of encoding, transmission and decoding of data representing a sequence of 3D scenes, in accordance with a non-limiting embodiment of the present principles; [Figure 3] 13 shows an exemplary architecture of a device that may be configured to implement the method described in connection with FIGS. 11 and 12, in accordance with a non-limiting embodiment of the present principles. [Figure 4] 1 illustrates an example of one embodiment of the syntax of a stream when data is transmitted via a packet-based transmission protocol, in accordance with a non-limiting embodiment of the present principles. [Figure 5] 10 illustrates a patch atlas approach with four example centers of projection, in accordance with a non-limiting embodiment of the present principles. [Figure 6] 1 shows an example of an atlas containing texture information for points of a 3D scene, in accordance with a non-limiting embodiment of the present principles; [Figure 7] 7 shows an example of an atlas containing depth information for points in the 3D scene of FIG. 6, in accordance with a non-limiting embodiment of the present principles. [Figure 8] 10 illustrates the process used by a view synthesis device when generating an image for a given viewport from unpruned MVD frames, in accordance with a non-limiting embodiment of the present principles. [Figure 9] 9 shows the same view synthesis as FIG. 8 from pruned MVD frames, in accordance with a non-limiting embodiment of the present principles. [Figure 10] 10 shows an exemplary pruning graph for a 4x4 multiview frame and such an MVD frame, in accordance with a non-limiting embodiment of the present principles; [Figure 11] 1 illustrates a method for encoding multiview frames in a data stream, in accordance with a non-limiting embodiment of the present principles; [Figure 12] 5 illustrates a method for decoding multi-view frames pruned from a data stream, in accordance with a non-limiting embodiment of the present principles.
[0016] The present principles are more fully described below with reference to the accompanying drawings, in which examples of the present principles are shown. However, the present principles may be embodied in many alternative forms and should not be construed as limited to the embodiments set forth herein. Accordingly, while the present principles are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and are described in detail herein. It is to be understood, however, that there is no intention to limit the present principles to the particular forms disclosed, but on the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present principles as defined by the appended claims.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the present principles. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that as used herein, the terms "comprises," "comprising," "includes," and / or "including" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, when an element is referred to as "responsive to" or "connected to" another element, it may be directly responsive to or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as "directly responsive to" or "directly connected to" another element, there are no intervening elements present. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".
[0018] In this specification, terms such as "first," "second," etc. may be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the teachings of the present principles.
[0019] Some of the figures include arrows on communication paths to indicate the primary direction of communication, however, it should be understood that communication may occur in the opposite direction to the depicted arrow.
[0020] Some examples are described with reference to block diagrams and operational flowcharts, in which each block represents circuit elements, modules, or portions of code, with each block including one or more executable instructions for implementing a specified logical function. It should also be noted that in other implementations, the functions noted in the blocks may occur out of the order noted. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved.
[0021] As used herein, "by one example" or "in one example" means that a particular feature, structure, or characteristic described in connection with this embodiment may be included in at least one implementation of the present principles. The appearances of the phrase "by one example" or "in one example" in various places in this specification do not necessarily all refer to the same embodiment, and in separate or alternative embodiments are not necessarily mutually exclusive of other embodiments.
[0022] Reference numerals appearing in the claims are by way of example only and shall have no limiting effect on the scope of the claims. Although not expressly stated, the present embodiments and variations may be used in any combination or subcombination.
[0023] FIG. 1 illustrates a three-dimensional (3D) model 10 of points of an object and a point cloud 11 corresponding to the 3D model 10. The 3D model 10 and point cloud 11 may correspond to, for example, a potential 3D representation of an object in a 3D scene containing other objects. The model 10 may be a 3D mesh representation, and the points of the point cloud 11 may be vertices of the mesh. The points of the point cloud 11 may also be points spread on the surface of a face of the mesh. The model 10 may also be represented as a splatted version of the point cloud 11, where the surface of the model 10 is created by splatting the points of the point cloud 11. The model 10 may be represented by many different representations, such as voxels or splines. FIG. 1 illustrates the fact that a point cloud may be defined as a surface representation of a 3D object, and that a surface representation of a 3D object may be generated from a cloud of points. As used herein, projecting a point of a 3D object (by an elongation point of the 3D scene) onto an image is equivalent to projecting any representation of this 3D object, for example a point cloud, a mesh, a spline model or a voxel model.
[0024] A point cloud can be represented in memory, for example, as a vector-based structure, with each point having its own coordinates in the reference frame of the viewpoint (e.g., three-dimensional coordinates XYZ, or solid angle and distance (also called depth) from / to the viewpoint) and one or more attributes, also called components. Examples of components are color components, which can be expressed in various color spaces, for example, RGB (red, green, and blue) or YUV (Y is the luminance component and UV are the two color difference components). A point cloud is a representation of a 3D scene containing objects. A 3D scene can be viewed from a given viewpoint or range of viewpoints. Point clouds can be represented in many ways, for example, From capturing real objects photographed by a camera rig, optionally complemented by an active depth sensing device; From capturing virtual / synthetic objects photographed by virtual camera rigs in modeling tools, · It can be obtained from a mixture of both real and virtual objects.
[0025] A 3D scene, especially when prepared for 3DoF rendering, can be represented by Multi-View+Depth (MVD) frames. Volumetric A video is a sequence of MVD frames. In this approach, Volumetric The information is conveyed as a combination of color and depth patches stored in corresponding color and depth atlases, which are then video encoded using a codec (typically HEVC). Each combination of color and depth patches typically represents a portion of the MVD input view, and the set of all patches is designed in the encoding stage to cover the entire scene with as little redundancy as possible. Decryption Atlas first introduced the video Decryption The patches are then rendered in a view synthesis process to recover the viewport associated with the desired viewing position.
[0026] 2 shows a non-limiting example of encoding, transmission and decoding of data representing a sequence of 3D scenes, e.g., a coding format that can accommodate 3DoF, 3DoF+ and 6DoF decoding simultaneously.
[0027] A sequence of a 3D scene 20 is acquired. Picture When the sequence is a 2D video, the sequence of 3D scenes is a 3D ( Volumetric A sequence of 3D scenes is a video for 3DoF, 3Dof+ or 6DoF rendering and display. Volumetric It may be provided to a video rendering device.
[0028] A sequence of 3D scenes 20 is provided to an encoder 21. The encoder 21 takes as input a 3D scene or a sequence of 3D scenes and provides a bitstream representing the input. The bitstream may be stored in a memory 22 and / or on an electronic data medium and may be transmitted over a network 22. The bitstream representing the sequence of 3D scenes may be read from the memory 22 and / or received from the network 22 by a decoder 23. The decoder 23 is input with the bitstream and provides the sequence of 3D scenes, for example in point cloud format.
[0029] The encoder 21 may comprise several circuits that implement several steps. In a first step, the encoder 21 encodes each 3D scene into at least one 2D Picture 3D projection is any method of mapping three-dimensional points onto a two-dimensional plane. Since most modern methods for displaying graphic data are based on flat (pixel information from several bit planes) two-dimensional media, the use of this type of projection is widespread, especially in computer graphics, manipulation, and drafting. The projection circuit 211 provides at least one two-dimensional frame 2111 for a 3D scene of the sequence 20. The frame 2111 includes color and depth information representing the 3D scene projected onto the frame 2111. In a variant, the color and depth information are coded in two separate frames 2111 and 2112.
[0030] The metadata 212 is used and updated by the projection circuitry 211. The metadata 212 includes information about the projection operation (e.g., projection parameters) and how color and depth information is organized within frames 2111 and 2112, as described in connection with Figures 5-7.
[0031] The video encoding circuit 213 encodes the sequence of frames 2111 and 2112 as a video. Picture (or 3D scene PictureThe sequence of video data and metadata 212 is then encoded in a stream by a video encoder 213. The video data and metadata 212 is then encapsulated in a data stream by a data encapsulation circuit 214.
[0032] The encoder 213 may, for example, -JPEG, specification ISO / CEI10918-1UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-REC-T.81 / en; -Compliant with encoders such as AVC, also known as MPEG-4 AVC or h264. UIT-TH.264 and ISO / CEI MPEG-4-Part 10 (ISO / CEI14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (whose specifications can be found on the ITU website, T Recommendation, H Series, h265, http: / / www.tigh.int / rec / T-REC-H.265-201612-I / en), -3D-HEVC (an extension of HEVC whose specification can be found on the ITU website, T Recommendation, H Series, h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annex G and I), -VP9, developed by Google, -AV1 (AO Media Video 1) developed by the Alliance for Open Media; or Adaptable to encoders of future standards such as the Versatile Video Coder or future versions of MPEG-I or MPEG-V.
[0033] The data stream is stored by the decoder 23 in a memory accessible, for example, via the network 22. The decoder 23 Decryption The decoder 23 takes as input the data stream generated by the encoder 21 and outputs it to a device such as a head-mounted device (HMD). VolumetricThe decoder 23 obtains a stream from a source 22, for example, a video display device that provides a sequence of 3D scenes 24 to be rendered and displayed by the video display device. - local memory, such as video memory or RAM (or Random Access Memory), flash memory, ROM (or Read Only Memory), hard disk, etc. -for example, trout a storage interface, such as an interface to storage, RAM, flash memory, ROM, optical disk, or magnetic support; a communication interface, such as a wired interface (e.g., a bus interface, a wide area network interface, a local area network interface) or a wireless interface (e.g., an IEEE 802.11 interface or a Bluetooth interface), a user interface, such as a graphical user interface, that allows a user to input data; belongs to the set containing
[0034] The decoder 23 includes a circuit 234 for extracting data encoded within the data stream. The circuit 234 takes the data stream as input and provides metadata 232 corresponding to the stream and the metadata 212 encoded in the two-dimensional video. The video is decoded by a video decoder 233, which provides a sequence of frames. The decoded frames include color and depth information. In a variant, the video decoder 233 provides two sequences of frames, one containing color information and the other containing depth information. The circuit 231 does not project the color and depth information from the decoded frames, but instead provides a sequence of 3D scene 24. The sequence of 3D scene 24 corresponds to the potentially reduced accuracy associated with encoding the sequence of 3D scene 20 as 2D video and video compression.
[0035] Figure 3 shows an example architecture of a device 30 that may be configured to perform the methods described in connection with Figures 11 and 12. The encoder 21 and / or decoder 23 of Figure 2 may implement this architecture. Alternatively, the circuits of the encoder 21 and / or decoder 23 may be devices according to the architecture of Figure 3, coupled together, for example, via their bus 31 and / or via an I / O interface 36.
[0036] The device 30 includes the following elements coupled together by a data and address bus 31: a microprocessor 32 (or CPU), for example a DSP (or Digital Signal Processor), -ROM (or read-only memory) 33; a RAM (or Random Access Memory) 34; a storage interface 35; an I / O interface 36 for receiving data to be transmitted from an application; - a power source, for example a battery;
[0037] According to one example, the power source is external to the device. In each of the mentioned memories, the word "register" as used herein may correspond to a small area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). The ROM 33 contains at least programs and parameters. The ROM 33 can store algorithms and instructions for performing techniques according to the present principles. When switched on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.
[0038] The RAM 34 contains in registers the programs executed by the CPU 32 and uploaded after switching on the device 30, input data in registers, intermediate data for different states of the methods in registers, and other variables used for the execution of the methods in registers.
[0039] The implementations described herein may be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when discussed only in the context of a single form of implementation (e.g., discussed only as a method or device), the implementation of the discussed features may also be implemented in other forms (e.g., a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be implemented in an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, handheld / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0040] According to an embodiment, the device 30 is configured to implement the method described in relation to Figures 11 and 12, -Mobile devices and a communication device; -A gaming device, -a tablet (or tablet computer); -Laptop and -A still camera, -Video camera and - an encoding chip; - a server (for example a broadcast server, a video-on-demand server or a web server), and belongs to the set including:
[0041] Figure 4 shows an example embodiment of the syntax of a stream when data is transmitted via a packet-based transmission protocol. Volumetric 4 shows an exemplary structure of a video stream. The structure organizes the stream in syntactically independent elements. containerThe structure may include a header section 41, which is a set of data common to all syntax elements of the stream. For example, the header section may include some of the metadata about the syntax elements, describing the nature and role of each of them. The header section may also include some of the metadata 212 of FIG. 2, for example, the coordinates of the central viewpoint used to project points of the 3D scene onto frames 2111 and 2112. The structure includes a payload that includes an element of syntax 42 and at least one element of syntax 43. Syntax element 42 includes data representing color and depth frames. The images may be compressed according to a video compression method.
[0042] Syntax 43 elements are part of the payload of the data stream and may contain metadata about how the frames of syntax 42 elements are encoded, for example, parameters used to project or pack points of a 3D scene onto the frame. Such metadata may be stored with each frame of video or (as in video compression standards) Group of Pictures It may be associated with a Group of Frames (also known as GoP).
[0043] FIG. 5 illustrates a patch atlas approach with four example centers of projection. 3D scene 50 includes features. For example, center of projection 51 is a perspective projection camera, and camera 53 is an orthographic projection camera. The cameras may also be omnidirectional cameras with spherical mapping (e.g., equirectangular mapping) or cubic mapping. 3D points of the 3D scene are projected onto a 2D plane associated with a virtual camera located at the center of projection according to a projection operation described in the projection data of the metadata. In the example of FIG. 5, the projection of points captured by camera 51 is mapped onto patch 52 according to a perspective mapping, and the projection of points captured by camera 53 is mapped onto patch 54 according to an orthogonal mapping.
[0044] Clustering the projected pixels results in a number of 2D patches, which are packed into a rectangular atlas 55. The organization of the patches within the atlas defines the atlas layout. In one embodiment, two atlases have the same layout: one for texture (i.e., color) information and one for depth information. Two patches captured by the same camera or two separate cameras may contain information representing the same portion of a 3D scene, such as patches 54 and 56.
[0045] The packing operation generates patch data for each generated patch. The patch data includes a reference to the projection data (e.g., an index into a table of projection data or a pointer to the projection data (an address in memory or the data stream)) and information describing the location and size of the patch within the atlas (e.g., a pixel's top-left corner coordinate, size, and width). The patch data items are associated with the compressed data of one or two atlases and are added to the metadata encapsulated in the data stream.
[0046] 6 shows an example of an atlas 60 containing texture information (e.g., RGB or YUV data) for points of a 3D scene, in accordance with a non-limiting embodiment of the present principles. As explained in connection with FIG. 5, the atlas is an image packing patch, where the patch is obtained by projecting a portion of the points of the 3D scene. Picture is.
[0047] In the example of FIG. 6 , the atlas 60 includes a first portion 61 containing texture information for points of the 3D scene visible from the viewpoint and one or more second portions 62. The texture information for the first portion 61 may be obtained, for example, according to equirectangular projection mapping, which is an example of spherical projection mapping. In the example of FIG. 6 , the second portion 62 is located at the left and right boundaries of the first portion 61, but the second portion may be located differently. The second portion 62 contains texture information for portions of the 3D scene that are complementary to the portions visible from the viewpoint. The second portion can be obtained by removing from the 3D scene the points visible from the first viewpoint (the texture stored in the first portion) and projecting the remaining points according to the same viewpoint. The latter process can be repeated iteratively so that hidden portions of the 3D scene are obtained each time. According to a variant, the second part can be obtained by removing from the 3D scene the points that are visible from a viewpoint, for example a central viewpoint (the texture stored in the first part), and by projecting the remaining points from one or more second viewpoints according to a viewpoint different from the first viewpoint, for example a space of views centered on the central viewpoint (for example a viewing space of a 3DoF rendering).
[0048] The first portion 61 can be seen as a first large texture patch (corresponding to a first portion of the 3D scene), and the second portion 62 comprises smaller texture patches (corresponding to a second portion of the 3D scene that is complementary to the first portion). Such an atlas has the advantage of being simultaneously compatible with 3DoF rendering and 3DoF+ / 6DoF rendering (when rendering only the first portion 61).
[0049] Figure 7 shows an example of an atlas 70 containing depth information for points in the 3D scene of Figure 6, in accordance with a non-limiting embodiment of the present principles. Atlas 70 can be seen as a depth image corresponding to texture image 60 of Figure 6.
[0050] Atlas 70 includes a first portion 71 containing depth information for points in the 3D scene as seen from a central viewpoint and one or more second portions 72. Atlas 70 may be obtained in the same manner as atlas 60, but includes depth information associated with points in the 3D scene instead of texture information.
[0051] For 3DoF rendering of a 3D scene, only one viewpoint is considered, typically a central viewpoint. The user can rotate their head with three degrees of freedom around the first viewpoint to view different parts of the 3D scene, but the user cannot move this unique viewpoint. The scene points that are coded are those that are visible from this unique view, and only texture information needs to be coded / decoded for 3DoF rendering. There is no need to code scene points that are not visible from this unique viewpoint for 3DoF rendering, as the user does not have access to them.
[0052] For 6DoF rendering, the user can move the viewpoint all over the scene. In this case, it is necessary to encode all points of the scene (depth and texture) in the bitstream, since all points are potentially accessible by a user who can move his / her viewpoint. At the encoding stage, there is no way to know a priori from which viewpoint the user will observe the 3D scene.
[0053] For 3DoF+ rendering, the user can move their viewpoint within a limited space around a central viewpoint. This allows them to experience parallax. Data representing a portion of a scene visible from any point in the view space should be encoded into a stream containing data representing the 3D scene as seen from the central viewpoint (i.e., the first portions 61 and 71). The size and shape of the view space can be determined, for example, in the encoding step and encoded in the bitstream. The decoder can obtain this information from the bitstream, and the renderer limits the view space to the space determined by the obtained information. According to another example, the renderer determines the view space according to hardware constraints, for example, related to the capabilities of a sensor that detects the user's movements. In such a case, if a point visible from a point in the renderer's view space is not encoded in the bitstream during the encoding stage, this point will not be rendered. According to a further example, data representing all points of the 3D scene (e.g., textures and / or geometric shapes) is encoded in the stream without taking into account the view rendering space. To optimize the size of the stream, only a subset of the points of the scene can be coded, for example the subset of points that are visible according to the rendering space of the view.
[0054] The patches are created to be fully non-redundant and complementary. The process of generating patches from a Multi-View+Depth (MVD) representation of a 3D scene consists of "pruning" the input source views to remove any redundant information. To do so, each input view (color + depth) is iteratively pruned from each other. A set of unpruned views, called base views, is first selected among the source views and transmitted in full. The remaining set of views, called additional views, is then iteratively processed to remove information that is redundant (in terms of color and depth similarity) with respect to the base view and the already pruned additional views. The color or depth values of the pruned pixels are replaced with a predetermined value, for example, 0 or 255.
[0055] FIG. 8 shows the process used by the view synthesis unit 231 of FIG. 2 when generating an image for a given viewport from unpruned MVD frames. Volumetric To convey video, a key step consists of removing redundant information between the base view and the additional views. However, simply removing redundant information without any other signaling can significantly reduce the amount of information to be transmitted. Decryption This could significantly alter the view synthesis process at this stage and severely degrade the end-user experience. When attempting to synthesize pixel 81 for viewport 80, the synthesizer (e.g., circuit 231 in FIG. 2) does not cast a ray (e.g., rays 82 and 83) that passes through this given pixel, but instead checks the contribution of each source camera 84-87 along this ray. As shown in FIG. 8, consensus among all source cameras 84-87 regarding the pixel's characteristics for synthesis may not be found when some objects in the scene create occlusions from one camera to another or cannot be ensured to be visible due to camera settings. In the example of FIG. 8, the first group of three cameras 84-86 "votes" to synthesize pixel 81 using the color of foreground object 88 when they all "see" this object along the ray for synthesis. The second group of one single camera 87 cannot see this object because it is outside its viewport. Therefore, camera 87 "votes" for background object 89 to synthesize pixel 81. A strategy to disambiguate such situations is to blend and / or merge the contributions of each camera by weight depending on their distance to the viewport for compositing. In the example of Figure 8, the first group of cameras 84-86 provide the largest contributions when they are more numerous and closer to the viewport for compositing. Finally, pixel 81 is composited by using the characteristics of foreground object 88, as expected.
[0056] Figure 9 shows the same view synthesis as Figure 8 from a pruned MVD frame. In a pruned MVD frame, pixels from cameras that share the same information are cleared and are no longer transmitted or considered. In the example of Figure 9, the previous group of three cameras is now reduced to one single camera 96, which carries information for foreground object 88. Corresponding pixel information 92 in the views from cameras 84 and 85 has been pruned. The second group of cameras related to background object 89 remains unchanged and contains only the view of camera 87. In this case, the contribution of the background to synthesize pixel 91 is no longer negligible relative to the contribution of the foreground, as the "opposites" become one-to-one. Even if the weight of object 88 is slightly higher than that of background 89, the blending of the two contributions includes a significant amount from the background, which does not correspond to what the user expects and leads to visual artifacts. Therefore, the loss of some camera contribution information after the pruning stage can be significant in the decoding stage when attempting to synthesize new views from the atlas.
[0057] According to the present principles, a method is disclosed to overcome these drawbacks. During the encoding phase, a pruning graph is obtained. The pruning graph constrains each camera to pruning a given subgroup of other cameras. Data representing the pruning graph is encoded in the data stream and provided to the decoder in a compact manner. During the decoding phase, the pruning graph can be recovered by using these metadata and used to restore the contribution information of all pruned cameras.
[0058] FIG. 10 shows an example pruning graph for a 4x4 multiview frame and such an MVD frame. According to the present principles, for each camera (i.e., views 111-144), a set of other cameras is determined. Each camera is aperiodically related to zero, one, or several other cameras by a pruning priority relation (i.e., the pruned graph obtained from the pruning priority relation does not contain any cycles). To have an efficient pruning relation, the priority relation is selected so that two connected views have a high potential amount of redundancy. This potential can be determined, for example, based on the distance between the optical centers of the two cameras of interest, their overlap ratio, or the angle / distance between their optical axes. To obtain an acyclic graph, a two-step strategy can be envisioned: first, densely connect all cameras according to the criterion selected for priority, and second, greedily prune the obtained graph to retain the minimum amount of connections that guarantees acyclicity. The base view (133 in the example of Figure 10) does not point towards any other camera because the base view has not been pruned. Some views (111, 114, 141 and 144 in the example of Figure 10) have no predecessors in the graph.
[0059] During the pruning procedure, a pruning order is determined so that a camera is always pruned after all its parents in terms of pruning priority. In the example of FIG. 10, the pruning order may be (133, 123, 132, 134, 143, 113, 122, 124, 131, 142, 144, 112, 114, 121, 141). The pruning procedure for all cameras is performed in this order: A pixel of a camera to be pruned is pruned for its associated camera if and only if it can be pruned for all cameras in its reference set (i.e., the same information is carried by all reference cameras). If one part of the parent camera set has already been pruned during the process, pruning is attempted recursively for its unique or multiple parents until an unpruned region is found to avoid any drift effects. If no consensus is found, the pixel considered for pruning is not pruned, and its value remains unchanged. Otherwise, the pixel (and its value) is discarded. With each pairwise comparison that occurs along the path of the pruning tree, there is a small registration error in depth. The error is lower than the threshold for comparisons between two close cameras (i.e., topologically adjacent views), but not for two remote cameras that are compared indirectly through a path of the pruning tree. The drift effect is the accumulation of small registration errors in depth between cameras along the path of the pruning tree.
[0060] For use in the decoding stage, the pruned graph is encoded within the data stream in accordance with a non-limiting embodiment of the present principles.
[0061] In a first embodiment, data representing all priority relationships in the pruning graph is encoded as a list containing, for each camera, a list of the cameras it is associated with, according to the syntax format as shown in Table 2, with each camera identified by its position in the camera parameters list, according to the syntax format as proposed in Table 1. If the number of cameras is small (e.g., lower than 64), a mask / bit array may alternatively be used to describe the pruning priority, with each ith bit being set to 1 if pruning is to be performed on the ith camera, for example, according to the syntax format described in Table 3. [Table 1] [Table 2] [Table 3]
[0062] In another embodiment, the pruning relations are integrated into the camera parameter list as new parameters for each camera (as an array or as a mask), for example according to the syntax format suggested in Tables 4 and 5. [Table 4] [Table 5]
[0063] During the decoding stage, the pruning graph is recovered from the metadata and used to correctly process the renderer's weighting strategy. In one embodiment, for each pixel to be composited, the contributions of all cameras are iteratively considered. For each camera that provides a valid contribution, all cameras pruned to this camera are iteratively considered by browsing the pruning graph in pruning order (from parent to its children). If the browsed camera is pruned to the camera of interest for the pixel being considered, its weight is combined (e.g., added) to the weight of the current camera, and then its children are processed similarly. If the browsed camera is not pruned to this camera because it holds different valid information, browsing is stopped along the associated branch of the graph, and the weight of the camera of interest remains unchanged.
[0064] According to the present principles, the contributions of the pruned cameras are correctly restored at the decoder stage after pruning, preventing visual artifacts such as those described in connection with FIG.
[0065] FIG. 11 shows a method 110 for encoding multiview frames in a data stream according to a non-limiting embodiment of the present principles. In step 111, an MVD frame is obtained from a source. In this step, the MVD frame requires a large amount of data to be encoded. In step 112, a graph is constructed to determine the concatenated views of the MVD according to a precedence relationship. The graph is constructed to be aperiodic, meaning that no view can be preceded in the pruning process by a view that precedes it. Some views have no predecessors and views that are not meant to be pruned (also called base views) have no successors in the graph. In step 113, views are pruned according to the precedence relationship of the graph, as described in connection with FIG. 10. In this stage, redundant information (color and depth) in the initial MVD obtained in step 111 is removed, resulting in less data being required to be encoded. The remaining useful information can be organized into unique frames called atlases, as described in connection with FIGS. 5-7. In step 114, the pruned MVD or the corresponding atlas is associated with dedicated metadata and encoded in the stream. According to the present principles, the pruning precedence relations of the pruned graph are also encoded in the stream, for example following one of the proposed syntax forms. In a further step, the data stream can be stored in a memory or a non-transitory storage medium, or transmitted to a remote or local device via a network or data bus.
[0066] FIG. 12 shows a method 120 for decoding pruned multi-view frames from a data stream, according to a non-limiting embodiment of the present principles. In step 121, a data stream is obtained, and data representing a pruned MVD, e.g., in the format of an atlas, is obtained from the data stream. For example, the pruned MVD is decoded from the data by using a video codec. In step 122, a pruning graph connecting views of the MVD is obtained from the data stream. Steps 121 and 122 can be performed in any order or in parallel. The pruning graph is an acyclic structure of pruning priority relationships between views of the MVD, as described in detail herein. In step 123, a viewport frame is generated for a viewing pose (i.e., the location and orientation in the 3D space of the renderer). For a pixel of the viewport frame, the weight of the contribution of each view (also referred to as a "camera" in this application) is determined according to the pruning priority relationships between views in the obtained pruning graph. For each camera that provides a valid contribution, all cameras that are pruned to this camera are iteratively considered by browsing the pruning graph in pruning order (from parent to its children). If the browsed camera is pruned to the camera of interest for the pixel being considered, its weight is combined (e.g., added) to the weight of the current camera, and then its children are processed similarly. If the browsed camera is not pruned to this camera because it holds different valid information, browsing is stopped along the associated branch of the graph, and the weight of the camera of interest remains unchanged.
[0067] In one embodiment, the decoding stage can use the pruned graph to unplane the pruned input views. According to the present principles, all source views of the received pruned MVD are reconstructed by recovering the missing redundancies suppressed by the pruning process. To do so, a backward procedure is applied. Starting from the root node to the leaves, a valid (unpruned) pixel p of the view associated with node N is considered. Then, 1) Pixel p is unprojected onto the (not yet "pruned") views associated with its view's children, and if it contributes to those viewports, then the associated unprojected pixel status is captured. 2) If an unprojected pixel is identified as pruned (and remains without a valid value), its color and depth values are set to the value of pixel p (color and / or depth), and the process is repeated iteratively for the children of the latter view. 3) If the deprojected pixel is identified as unpruned (and has valid values), its color and depth values remain unchanged and no further graph inspection is done towards the children of this latter view. 4) If pixel p does not fall within the viewport of one of its children, the process is repeated recursively for its grandchildren.
[0068] Doing so makes it possible to provide a multi-view display, which requires displaying all views of the MVD content at all times (not just the synthesized virtual view at the HMD, but also the synthesized virtual view at the HMD) while transmitting pruned content at a reduced bitrate.
[0069] The implementations described herein may be implemented in, for example, a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when discussed in the context of only a single form of implementation (e.g., discussed only as a method or device), the implementation of the discussed features may also be implemented in other forms (e.g., a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The method may be implemented in an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices such as, for example, smartphones, tablets, computers, mobile phones, handheld / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0070] Implementations of the various processes and features described herein may be embodied in a variety of different devices or applications, particularly, for example, devices or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and associated texture and / or depth information. Examples of such devices include encoders, decoders, post-processors that process output from decoders, pre-processors that provide input to encoders, video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communication devices. As should be clear, the devices may be mobile and installed in mobile vehicles.
[0071] Furthermore, a method may be implemented by instructions executed by a processor, and such instructions (and / or data values produced by an implementation) may be stored on a processor-readable medium, such as, for example, an integrated circuit, software carrier, or other storage device, e.g., a hard disk, a compact diskette ("CD"), an optical disk (e.g., a DVD, often referred to as a digital versatile disk or digital video disk), a random access memory ("RAM"), or a read-only memory ("ROM"). The instructions may form an application program tangibly embodied on the processor-readable medium. The instructions may be, for example, hardware, firmware, software, or a combination. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. A processor may therefore be characterized as both, for example, a device configured to execute a process and a device that includes a processor-readable medium (e.g., a storage device) having instructions for executing a process. Furthermore, a processor-readable medium may store data values produced by an implementation in addition to, or in place of, instructions.
[0072] As will be apparent to those skilled in the art, implementations may generate various signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry, as data, rules for writing or reading syntax of a described embodiment, or to carry, as data, the actual syntax value written by a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
[0073] Many implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or deleted to produce other implementations. Moreover, those skilled in the art will understand that other structures and processes may be substituted for those disclosed, with the resulting implementation performing at least substantially the same function in at least substantially the same way to achieve at least substantially the same results as the disclosed implementations. Accordingly, these and other implementations are contemplated by this application.
Claims
1. 1. A method for encoding views of a multiview frame into a data stream, comprising: obtaining an acyclic graph connecting views of the multiview frame, wherein links of the acyclic graph represent pruning priority relationships, and at least one base view of the multiview frame has no pruning priority links; pruning pixels of views of the multiview frame in a determined order such that a given view is pruned after views connected to the given view by the pruning priority link, wherein pixels of the given view are pruned when they correspond to information coded in pixels of a base view or in a pruned view; encoding the acyclic graph, the at least one base view, and the pruned views into the data stream; A method comprising:
2. The method of claim 1 , wherein pruning pixels of a view comprises replacing values of the pixels with determined values.
3. The method of claim 1 , wherein the acyclic graph is signaled in the data stream as a list that includes, for each view of the multiview frame, a list of views that are connected to the view.
4. 1. A device for encoding views of a multiview frame into a data stream, comprising: obtaining an acyclic graph connecting views of the multiview frame, wherein links of the acyclic graph represent pruning priority relationships, and at least one base view of the multiview frame has no pruning priority links; pruning pixels of views of the multiview frame in a determined order such that a given view is pruned after views connected to the given view by the pruning priority link, wherein pixels of the given view are pruned when they correspond to information coded in pixels of a base view or in a pruned view; encoding the acyclic graph, the at least one base view, and the pruned views into the data stream; 10. A device comprising: a processor configured to:
5. The device of claim 4 , wherein pruning pixels of a view comprises replacing values of the pixels with determined values.
6. The device of claim 4 , wherein the acyclic graph is signaled in the data stream as a list that includes, for each view of the multiview frame, a list of views that are concatenated with the view.
7. 1. A method for decoding a view of a multiview frame from a data stream, comprising: obtaining the views of the multiview frame from the data stream, wherein at least one base view is unpruned and other views are pruned; obtaining an acyclic graph from the data stream, the acyclic graph connecting views of the multiview frame, links of the acyclic graph representing pruning priority relationships, and the at least one base view of the multiview frame having no pruning priority links; generating a viewport frame according to a viewing pose by determining the contribution of each view of the multiview frame as a function of the pruning priority relationship of the acyclic graph; A method comprising:
8. The method of claim 7 , wherein pruned pixels of the pruned view have determined values.
9. The method of claim 7 , wherein the acyclic graph is signaled in the data stream as a list that includes, for each view of the multiview frame, a list of views that are connected to that view.
10. 1. A device for decoding a view of a multiview frame from a data stream, comprising: obtaining the views of the multiview frame from the data stream, wherein at least one base view is unpruned and other views are pruned; obtaining an acyclic graph from the data stream, the acyclic graph connecting views of the multiview frame, links of the acyclic graph representing view pruning priority relationships, and the at least one base view of the multiview frame having no pruning priority links; generating a viewport frame according to a viewing pose by determining the contribution of each view of the multiview frame as a function of the pruning priority relationship of the acyclic graph; 20. A device comprising: a processor configured to:
11. The device of claim 10 , wherein pruned pixels of the pruned view have determined values.
12. The device of claim 10 , wherein the acyclic graph is signaled in the data stream as a list that includes, for each view of the multiview frame, a list of views that are concatenated with the view.
Citation Information
Patent Citations
Multi-view media data
JP2012505569A
Method and system for encoding multi-view video content
JP2014520409A
Method and system for encoding multi-view video content
US20140111611A1
Multi-view media data
WO2010041998A1