A method and apparatus for encoding and decoding stereoscopic video

By selecting an appropriate atlas layout and utilizing rate-distortion optimization standards to encode stereoscopic video, the problem of layout data size in stereoscopic video encoding is solved, thereby improving encoding efficiency and quality.

CN114270863BActive Publication Date: 2025-12-23INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080037690.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-29
Filing Date
2020-03-19
Publication Date
2025-12-23
Estimated Expiration
2040-03-19

AI Technical Summary

Technical Problem

Existing stereoscopic video coding technologies struggle to effectively reduce the size of layout data in atlas-based coding without compromising coding quality, resulting in low coding efficiency.

Method used

By acquiring multiple atlas layouts of a 3D scene, selecting a suitable layout based on rate-distortion optimization criteria, generating an atlas and encoding it into a video data stream, and using the same layout to encode consecutive 3D scene sequences, the size of the layout metadata is reduced.

Benefits of technology

It improves the efficiency of stereoscopic video encoding, reduces the size of layout data, and maintains encoding quality while optimizing encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114270863B_ABST
    Figure CN114270863B_ABST
Patent Text Reader

Abstract

A method and apparatus for encoding stereoscopic video in patch-based atlas format in different length intra-periods is disclosed. A first atlas layout is constructed for a first sequence of 3D scenes. The number of 3D scenes in the sequence is chosen to fit the size of the GoP of the codec. A second sequence is iteratively built by appending the next 3D scene of the sequence when the number of patches of the layout constructed for the second sequence is less than or equal to the number of patches of the first layout. At the end of the iteration, one of the layouts is chosen to generate each atlas of the group. In this way, the size of the metadata is reduced and the compression is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to the field of three-dimensional (3D) scenes and stereoscopic video content. It is also understood herein in the context of encoding, formatting and decoding data representing textures and geometry of 3D scenes for rendering stereoscopic content on end-user devices such as mobile devices or head-mounted displays (HMDs). BACKGROUND

[0002] This section is intended to introduce the reader to various aspects of art that can be related to various aspects of the present disclosure that are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present disclosure. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.

[0003] Recently, there has been a growth in available large field of view content (up to 360°). Such content can not be fully visible to a user watching it on an immersive display device such as a head-mounted display, smart glasses, a PC screen, a tablet, a smartphone, etc. This means that at a given moment, the user can only be watching a part of the content. However, the user can typically roam within the content through various means such as moving his head, moving a mouse, touching a screen, voice, etc. It is typically required to encode and decode the content.

[0004] Immersive video (also called 360° planar video) allows a user to watch everything around him by rotating his head around a static viewpoint. The rotation only allows a 3 degrees of freedom (3DoF) experience. Even if 3DoF video is enough for a first omnidirectional video experience, for example using a head-mounted display device (HMD), 3DoF video can quickly become frustrating for a viewer expecting more freedom, for example by experiencing parallax. Moreover, 3DoF can also cause dizziness because a user never only rotates his head, but also translates his head in three directions, which is not reproduced in a 3DoF video experience.

[0005] The large field of view content can be in particular a three-dimensional computer graphics image scene (3D CGI scene), a point cloud or an immersive video. Many terms can be used to refer to such immersive video: for example, virtual reality (VR), 360, panorama, 4π steradians, immersive, omnidirectional or large field of view.

[0006] Stereoscopic video, also called 6 Degrees of Freedom (6DoF) video, is an alternative to 3DoF video. When watching a 6DoF video, in addition to rotations, the user can also translate his head, and even his body, within the content he is watching, and experience parallax and even stereoscopy. This video significantly increases the sense of immersion and perception of the depth of the scene and prevents dizziness by providing consistent visual feedback during head translation. This content is created from dedicated sensors that allow simultaneous recording of the color and depth of the scene of interest. The use of a platform of color cameras combined with photogrammetry techniques is one way to perform this recording, even though technical difficulties remain.

[0007] While 3DoF video consists of a sequence of images resulting from the demapping of texture images (for example, spherical images encoded according to a latitude / longitude projection mapping or an equirectangular projection mapping), 6DoF video frames embed information from several viewpoints. They can be considered as a temporal sequence of point clouds resulting from a three-dimensional capture. Two kinds of stereoscopic video can be considered depending on the viewing conditions. The first one, i.e. full 6DoF, allows a complete freedom to roam within the video content, while the second one, also known as 3DoF+, restricts the user view space to a limited stereoscopy called viewing bounding box, allowing a limited head translation and parallax experience. The second scenario is a valuable trade-off between free roaming and passive viewing conditions for a seated audience.

[0008] The technical approach to encode stereoscopic video is based on the projection of a 3D scene onto a plurality of 2D pictures called patches, which are packed into atlases that can be further compressed using a traditional video encoding standard, for example HEVC. The patches are packed in an atlas following an organization of a given layout. The atlas is encoded in the data stream in association with metadata describing its layout; this is a description of the position, shape and size of each patch in the atlas. These metadata have a non-negligible size because an atlas can include hundreds of patches. To limit the size of the layout metadata, one approach consists in using the same layout for a given number of successive atlases corresponding to the projection of a successive stereoscopic sequence of the same number of consecutive 3D scenes. This number is chosen to fit the number of frames in a group of pictures (GoP) of the chosen codec, for example 8 or 12. Even dividing the number of layout metadata by this number, the size of these data remains important. There is a lack of a technique to reduce the size of the layout data in an atlas-based encoding of a stereoscopic video without degrading the quality of the encoded sequence. SUMMARY

[0009] The following presents a simplified summary of the disclosure to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the disclosure. It is not intended to identify key or critical elements of the disclosure. The following summary merely presents some aspects of the disclosure in a simplified form as a prelude to the more detailed description provided below.

[0010] The disclosure relates to a method comprising acquiring a first atlas layout of a first sequence of 3D scenes. An atlas layout defines the organization of at least one patch within an atlas, an atlas being a packed image of at least one patch of the same 3D scene. A patch is an image representing the projection of a part of a 3D scene on an image plane, so different projections of a part of a 3D scene provide a patch set. These patch sets are packed in a larger image called an atlas. The organization of the patches within their atlas is called an atlas layout. The method further comprises acquiring a second atlas layout of a second sequence of 3D scenes. The second sequence is the first sequence to which a next 3D scene of the sequence of 3D scenes to be encoded is appended. If the number of patches of the second atlas layout is greater than the number of patches of the first atlas layout, then the atlases are generated for the first sequence of 3D scenes according to the first atlas layout. Otherwise, the method selects a layout among the first and second layouts, and generates the atlases for the second sequence of 3D scenes according to the selected layout.

[0011] In a particular embodiment, when the number of patches of the second atlas layout is less than or equal to the number of patches of the first atlas layout, the step of acquiring an atlas layout is iterated, the first sequence of 3D scenes becoming the second sequence of 3D scenes of the previous iteration. At each iteration, the method stores the acquired first atlas layout. When the iteration ends, a layout is selected among the stored layouts, and the atlases for the last first sequence of 3D scenes are generated according to the selected atlas layout. In a variant, a given number is set as the maximum number of atlases generated according to the same atlas layout. In this case, the iteration of the method ends when the second sequence comprises more than the given number of 3D scenes.

[0012] According to a particular embodiment, the selection of the layout is performed based on a rate-distortion optimization criterion. According to a particular embodiment, the generated sequence of atlases is encoded as one intra period in a video data stream.

[0013] The disclosure also relates to a device comprising a memory storing instructions to cause a processor to implement the steps of the method. The disclosure also relates to a video data stream generated by the device. BRIEF DESCRIPTION OF DRAWINGS

[0014] The disclosure will be better understood by reading the following description, given with reference to the attached drawings, in which:

[0015] Figure 1 A three-dimensional (3D) model of an object and points of a point cloud corresponding to the 3D model are shown, according to a non-limiting embodiment of the disclosure;

[0016] Figure 2Non-limiting examples of encoding, transmitting and decoding data representative of a sequence of 3D scenes according to non-limiting embodiments of the present disclosure are shown;

[0017] Figure 3 Examples of architectures of devices that can be configured to implement the methods described in connection with Figure 8 described methods are shown;

[0018] Figure 4 Examples of syntax of a stream when data is sent through a packet-based transmission protocol according to non-limiting embodiments of the present disclosure are shown;

[0019] Figure 5 Patch atlas methods with 4 projection centers examples according to non-limiting embodiments of the present disclosure are shown;

[0020] Figure 6 Examples of atlases including texture information of points of a 3D scene according to non-limiting embodiments of the present disclosure are shown;

[0021] Figure 7 Examples of atlases including depth information of points of a 3D scene in Figure 6 according to non-limiting embodiments of the present disclosure are shown;

[0022] Figure 8 A method 80 of encoding a sequence of 3D scenes in a data stream according to non-limiting embodiments of the present disclosure is shown;

[0023] Figure 9 Example structures of intra-periods of atlas layouts of different lifetimes according to non-limiting embodiments of the present disclosure are shown. DETAILED DESCRIPTION

[0024] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which non-limiting examples of the present disclosure are shown. The present disclosure may, however, be embodied in many alternate forms and should not be construed as limited to the examples set forth herein. Accordingly, while the present disclosure is susceptible to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and will be described herein in detail. It should be understood that the present disclosure is not to be limited to the particular examples disclosed, but it is intended to cover modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure as defined by the claims.

[0025] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises," "comprising," "includes" and / or "including," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Additionally, it will be understood that when an element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or indirectly responsive or connected to the other element via one or more other elements. In contrast, when an element is referred to as being "directly responsive" or "directly connected" to another element, there are no intervening elements. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and can be abbreviated as " / ".

[0026] It should be understood that, although the terms first, second, etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the teachings of the present disclosure.

[0027] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it is to be understood that communication can occur in the opposite direction to the depicted arrows.

[0028] Some examples are described with respect to block and operational flow diagrams in which each block represents a portion of code comprising one or more executable instructions for implementing the specified logical function(s). It should also be noted that in other implementations, the function(s) noted in the blocks can occur out of the order noted in the figure. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in reverse order, depending on the functionality involved.

[0029] Reference herein to "one example" or "an example" means that a particular feature, structure, or characteristic described in connection with the example can be included in at least one implementation of the disclosure. The appearances of the phrase "in one example" or "an example" in various places in the specification are not necessarily all referring to the same example, nor are they necessarily referring to some common but alternative example, nor are they necessarily referring to some operating example, but for illustrating the

[0030] The reference signs appearing in the claims are only for explanatory purposes and do not have any limiting effect on the scope of the claims. The examples and variants can be implemented in any combination or sub-combination, even if not explicitly described.

[0031] According to the present disclosure, methods and devices for encoding and decoding a stereoscopic video in a tile-based format are disclosed herein. A 3D scene is projected onto a plurality of 2D pictures called patches, which are packed into tiles according to a layout. According to the present disclosure, the same layout of patches in a tile is used for different numbers of consecutive tiles encoding a sequence of consecutive 3D scenes.

[0032] Figure 1 A three-dimensional (3D) model 10 of an object is shown as well as points of a point cloud 11 corresponding to the 3D model 10. The 3D model 10 and the point cloud 11 can for example correspond to a possible 3D representation of an object in a 3D scene comprising other objects. The model 10 can be a 3D mesh representation, while the points of the point cloud 11 can be the vertices of the mesh. The points of the point cloud 11 can also be points scattered on the surface of the faces of the mesh. The model 10 can also be represented as an unfolded version of the point cloud 11, the surface of the model 10 being created by unfolding the points of the point cloud 11. The model 10 can be represented by many different representations, such as voxels or splines. Figure 1 It is shown the fact that a point cloud can be defined with a surface representation of a 3D object and that a surface representation of a 3D object can be generated from a point cloud. As used herein, projecting points of a 3D object (through an extended point of a 3D scene) onto an image is equivalent to projecting any representation of that 3D object, such as a point cloud, a mesh, a spline model or a voxel model.

[0033] For example, a point cloud can be represented in memory as a vector-based structure, where each point has its own coordinates in the reference frame of a viewpoint (e.g. three-dimensional coordinates XYZ, or stereographic angle and distance from / to the viewpoint, also known as depth) and one or more attributes, also called components. One example of a component is a color component, which can be represented in various color spaces, such as RGB (red, green and blue) or YUV (Y is the luminance component, UV are two chrominance components). The point cloud is a representation of a 3D scene comprising an object. The 3D scene can be viewed from a given viewpoint or range of viewpoints. The point cloud can be acquired in many ways, for example:

[0034] • from a capture of a real object by a camera platform, optionally complemented by depth active sensing devices;

[0035] • from a capture of a virtual / synthetic object by a virtual camera platform in a modeling tool;

[0036] • from a mix of real and virtual objects.

[0037] Figure 2 Non-limiting examples of encoding, transmitting and decoding data representative of a sequence of 3D scenes are illustrated. For example, the encoding format can be compatible with 3DoF, 3DoF+ and 6DoF decoding at the same time.

[0038] A sequence of 3D scenes 20 is acquired. As a sequence of pictures is a 2D video, a sequence of 3D scenes is a 3D (also called stereoscopic) video. The sequence of 3D scenes can be provided to a stereoscopic video rendering device for 3DoF, 3DoF+ and 6DoF rendering and display.

[0039] The sequence of 3D scenes 20 is provided to an encoder 21. The encoder 21 takes as input one 3D scene or a sequence of 3D scenes and provides a bitstream representative of the input. The bitstream can be stored in a memory 22 and / or on an electronic data medium and can be transmitted on a network 22. A decoder 23 can read from the memory 22 and / or receive from the network 22 a bitstream representative of a sequence of 3D scenes. The bitstream is input to the decoder 23 and the decoder 23 provides a sequence of 3D scenes in a point cloud format for example.

[0040] The encoder 21 can comprise some circuits implementing some steps. In a first step, the encoder 21 projects each 3D scene onto at least one 2D picture. A 3D projection is any method of mapping three-dimensional points into a two-dimensional plane. As most current methods of displaying graphical data are based on planar (pixel information from some bitplanes) two-dimensional media, the use of this type of projection is widespread, especially in computer graphics, engineering, and cartography. The projection circuit 211 provides at least one two-dimensional frame 2111 for the sequence of 3D scenes 20. The frame 2111 comprises color information and depth information representative of a 3D scene projected onto the frame 2111. In a variant, the color information and the depth information are encoded in two separate frames 2111 and 2112.

[0041] The metadata 212 is used and updated by the projection circuit 211. The metadata 212 comprises information about the projection operation (e.g. projection parameters) and about the way the color and depth information are organized within the frames 2111 and 2112 as described with reference to Figures 5 to 7 According to the disclosure, the 3D scenes are projected onto a plurality of 2D pictures called patches which are packed into atlases. The same layout of patches in an atlas is used for different numbers of consecutive atlases encoding a sequence of consecutive 3D scenes.

[0042] The video encoding circuit 213 encodes the sequence of frames 2111 and 2112 into a video. The video encoder 213 encodes the pictures of the 3D scenes 2111 and 2112 (or the sequence of pictures of a 3D scene) in a stream. Then, the video data and the metadata 212 are encapsulated in a data stream by the data encapsulation circuit 214.

[0043] For example, the encoder 213 is an encoder complying with:

[0044] - JPEG, norm ISO / CEI 10918-1 UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-REC-T.81 / en;

[0045] - AVC, also called MPEG-4 AVC or h264. Both specified in UIT-T H.264 and ISO / CEI MPEG-4 Part 10 (ISO / CEI 14496-10), http: / / www.itu.int / rec / T-REC-H.264 / en, HEVC (whose norm is found on the ITU website - T Recommendation - H series - h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en ) ;

[0046] - 3D-HEVC (extension of HEVC, whose norm is found on the ITU website - T Recommendation - H series - h265, http: / / www.itu.int / rec / T-REC-H.265-201612-I / en annex G and I) ;

[0047] - VP9 developed by Google; or

[0048] - AV1 (Alliance for Open Media Video 1) developed by the Alliance for Open Media.

[0049] The data stream is stored in a memory accessible by the decoder 23 through the network 22, for example. The decoder 23 comprises different circuits implementing different decoding steps. The decoder 23 takes as input the data stream generated by the encoder 21 and provides a sequence of 3D scenes 24 to be rendered and displayed by a stereoscopic video display device, like a head-mounted device (HMD). The decoder 23 gets this stream from the source 22. For example, the source 22 belongs to the set comprising:

[0050] - a local memory, for example a video memory or a RAM (or Random Access Memory), a flash memory, a ROM (or Read Only Memory), a hard disk;

[0051] - a storage interface, for example an interface of a mass memory, a RAM, a flash memory, a ROM, an optical or magnetic support;

[0052] - a communication interface, for example a wired interface (for example a bus interface, a wide area network interface, a local area network interface) or a wireless interface (for example an IEEE 802.11 interface or a Bluetooth interface); and

[0053] - a user interface such as a graphical user interface enabling the user to input data.

[0054] The decoder 23 comprises a circuit 234 for extracting the data encoded in the data stream. The circuit 234 takes as input the data stream and provides metadata 232 corresponding to the metadata 212 encoded in the stream and a two-dimensional video. The video is decoded by a video decoder 233 providing a sequence of frames. The decoded frames comprise color and depth information. In one variant, the video decoder 233 provides two sequences of frames, one comprising color information and the other comprising depth information. The circuit 231 uses the metadata 232 to de-project the color and depth information from the decoded frames, in turn providing a sequence of 3D scenes 24. The sequence of 3D scenes 24 corresponds to the sequence of 3D scenes 20, possibly with a loss of precision related to the encoding as a 2D video and the video compression.

[0055] Figure 3 An example architecture of a device 30 that can be configured to implement the encoder 21 and / or the decoder 23 of the method described in connection with Figure 8 An example architecture of a device 30 that can be configured to implement the encoder 21 and / or the decoder 23 of the method described in connection with Figure 2 The encoder 21 and / or the decoder 23 can implement such an architecture. Optionally, each circuit of the encoder 21 and / or the decoder 23 can be a device according to the architecture of Figure 3 The encoder 21 and / or the decoder 23 can implement such an architecture. Optionally, each circuit of the encoder 21 and / or the decoder 23 can be a device according to the architecture of

[0056] The device 30 comprises the following elements linked together by a data and address bus 31 :

[0057] - a microprocessor 32 (or CPU), for example a DSP (or Digital Signal Processor);

[0058] - a ROM (Read Only Memory) 33;

[0059] - a RAM (Random Access Memory);

[0060] - a storage interface 35;

[0061] - an I / O interface 36 for receiving data to be transmitted from an application; and

[0062] - a power supply, for example a battery.

[0063] According to one example, the power supply is external to the device. In each of the mentioned memories, the word "register" used in the present description can correspond to an area of small capacity (a few bits) or to a very large area (for example, the entire program or a large amount of received or decoded data). The ROM 33 comprises at least the program and the parameters. The ROM 33 can store algorithms and instructions to perform the techniques according to the present disclosure. When switched on, the CPU 32 loads the program into the RAM and executes the corresponding instructions.

[0064] RAM 34 includes the program in registers executed by CPU 32 and uploaded after the device 30 is switched on, the input data in registers, the intermediate data of the different states of the method in registers, and other variables used to perform the method in registers.

[0065] Implementations described herein can be implemented in, for example, a method or a flow, an apparatus, a computer program product, data stream, or a signal. Even if only discussed in the context of a single implementation form (for example, as a method or device), the implementation of the discussed features can also be implemented in other forms (for example, a program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, an apparatus such as a processor which refers to processing device in general, including for example a computer, a

[0066] According to an example, the device 30 is configured to implement the method described in connection with Figure 8 The method described is, according to an example, implemented by a device 30 belonging to the set comprising:

[0067] - a mobile device;

[0068] - a communication device;

[0069] - a gaming device;

[0070] - a tablet (or tablet computer);

[0071] - a laptop computer;

[0072] - a still picture camera;

[0073] - a video camera;

[0074] - an encoding chip;

[0075] - a server (for example a broadcast server, a video on demand server or a web server).

[0076] Figure 4 An example of an embodiment showing the syntax of a stream when data is sent through a packet-based transmission protocol is shown. Figure 4 An example structure 4 of a stereoscopic video stream is shown. The structure is included in a container that organizes the stream in independent syntax elements. The structure can include a header 41 that is a set of data common to each syntax element of the stream. For example, the header includes some metadata about the syntax elements, describing the nature and role of each syntax element. The header can also include Figure 2part of the metadata 212, e.g. coordinates of the center viewpoint used to project points of the 3D scene onto the frames 2111 and 2112. This structure comprises a payload comprising a syntax element 42 and at least one syntax element 43. The syntax element 42 comprises data representing color and depth frames. The images can have been compressed according to a video compression method.

[0077] The syntax element 43 is part of the payload of the data stream and can comprise metadata about how the frames of the syntax element 42 have been encoded, e.g. parameters used to project and pack points of the 3D scene onto the frames. This metadata can be associated with each frame or group of frames (also called Group of Pictures (GoP) in video compression standards) of the video.

[0078] Figure 5 A patch atlas approach with 4 projection centers of the example is illustrated. The 3D scene 50 comprises one person. For example, the projection center 51 is a perspective camera and the camera 53 is an orthographic camera. The cameras can also be omnidirectional cameras with e.g. spherical mapping (e.g. equirectangular mapping) or cubic mapping. The 3D points of the 3D scene are projected onto 2D planes associated with virtual cameras located at the projection centers according to the projection operations described in the projection data of the metadata. In the example, the projection of the points captured by the camera 51 is mapped onto the patch 52 according to a perspective mapping and the projection of the points captured by the camera 53 is mapped onto the patch 54 according to an orthographic mapping. Figure 5

[0079] The clustering of the projected pixels yields a plurality of 2D patches which are packed in a rectangular atlas 55. The organization of the patches within the atlas defines an atlas layout. In one embodiment, two atlases have the same layout: one for texture (i.e. color) information and one for depth information. Two patches captured by the same camera or two different cameras can comprise information representing the same part of the 3D scene, e.g. the patches 54 and 56.

[0080] The packing operation generates one patch data for each generated patch. The patch data comprises a reference to the projection data (e.g. an index in a table of projection data or a pointer (i.e. an address in memory or data stream) pointing to the projection data) and information describing the position and size of the patch within the atlas (e.g. top-left coordinates, size and pixel width). The patch data items are added to the metadata to be encapsulated in the data stream in association with the compressed data of one or two atlases. The set of patch data items is also referred to as layout metadata in this specification.

[0081] Figure 6 An example of an atlas 60 comprising texture information (e.g. RGB data or YUV data) of points of a 3D scene according to a non-limiting embodiment of the disclosure is illustrated. As for the atlas 55, the atlas 60 comprises a plurality of patches 61, 62, 63 and 64. The atlas 60 is packed in a rectangular shape. The atlas 60 comprises two patches 61 and 62 captured by the same camera 51. The atlas 60 also comprises two patches 63 and 64 captured by the same camera 53. Figure 5 ​As explained, the atlas is a packed patch of images, the patches being pictures obtained by projecting a portion of points of a 3D scene. The layout of the atlas is metadata describing the position, shape and size of the patches within the atlas. In one embodiment, the shape of the patches is rectangular by default, so that the shape is not described, this information can be omitted in the layout metadata. In this embodiment, the position can be the top-left coordinates of the patch rectangle in pixels in the atlas, and its size can be described by a width and a height in pixels. In a variant, the size of a patch is described by an arc, pointing to the center of projection, comprising the points projected onto this patch, in a solid angle corresponding to a solid angle of the three-dimensional space. In other embodiments, the patches are ellipses and / or polygons described for example with Scalable Vector Graphics (SVG) instructions.

[0082] In Figure 6 In the example of FIG. 1, the atlas 60 comprises a first portion 61 comprising texture information of points of the 3D scene visible from the viewpoint, and one or more second portions 62. The texture information of the first portion 61 can be obtained for example according to an equirectangular projection mapping, which is one example of a spherical projection mapping. In Figure 6 In the example of FIG. 1, the second portions 62 are arranged at the left and right borders of the first portion 61, but the second portions can be arranged differently. The second portions 62 comprise texture information of portions of the 3D scene complementary to the portion visible from the viewpoint. The second portions can be obtained by removing from the 3D scene the points visible from the first viewpoint, whose texture is stored in the first portion, and by projecting the remaining points according to the same viewpoint. This latter process can be repeated iteratively to obtain each time a hidden portion of the 3D scene. According to a variant, the second portions can be obtained by removing from the 3D scene the points visible from a viewpoint (for example the central viewpoint), whose texture is stored in the first portion, and by projecting the remaining points according to a viewpoint different from the first viewpoint (for example one or more second viewpoints from the view space (for example the view space of the three-dimensional rendering) centered on the central viewpoint).

[0083] The first portion 61 can be seen as a first large texture patch (corresponding to a first portion of the 3D scene), and the second portions 62 comprise smaller texture patches (corresponding to second portions of the 3D scene complementary to the first portion). This atlas has the advantage of being compatible with both 3DoF rendering (when only the first portion 61 is rendered) and 3DoF+ / 6DoF rendering.

[0084] Figure 7 An example of an atlas 70 comprising depth information of points of a 3D scene according to non-limiting embodiments of the disclosure is shown. Figure 6 The atlas 70 can be seen as a depth image corresponding to the texture image 60 of Figure 6

[0085] ​The atlas 70 comprises a first part 71 comprising depth information of points of the 3D scene visible from the center viewpoint and one or more second parts 72. The atlas 70 can be acquired in the same way as the atlas 60, but contains depth information associated with points of the 3D scene instead of texture information.

[0086] For 3DoF rendering of the 3D scene, only one viewpoint is considered (typically the center viewpoint). The user can rotate his head around the first viewpoint with three degrees of freedom to watch various parts of the 3D scene, but the user cannot move this unique viewpoint. The scene points to be encoded are the points visible from this unique viewpoint and only the texture information needs to be encoded / decoded for 3DoF rendering. For 3DoF rendering, there is no need to encode scene points that are not visible from this unique viewpoint as they are not accessible by the user.

[0087] For 6DoF rendering, the user can move the viewpoint everywhere in the scene. In this case, every point of the scene (depth and texture) needs to be encoded in the bitstream as each point can be accessed by a user able to move his / her viewpoint. At the encoding stage, there is no way to know a priori from which viewpoint the 3D scene will be observed by the user.

[0088] For 3DoF+ rendering, the user can move the viewpoint within a limited space around the center viewpoint. This enables the experience of parallax. Data representative of the parts of the scene visible from any point of the view space will be encoded in the stream, including data representative of the 3D scene according to the center viewpoint (i.e. the first parts 61 and 71). The size and shape of the view space can be decided and determined, for example, in the encoding step and encoded in the bitstream. The decoder can get this information from the bitstream and the renderer limits the view space to the space determined by the information obtained. According to another example, the renderer determines the view space according to hardware constraints (for example related to the capabilities of the sensors detecting the user movements). In this case, if at the encoding stage, a point visible from a point within the view space of the renderer is not encoded in the bitstream, this point will not be rendered. According to another example, data representative of each point of the 3D scene (e.g. texture and / or geometry) are encoded in the stream without considering the rendering space of the view. To optimize the size of the stream, only a subset of points of the scene (e.g. a subset of points visible according to the rendering space of the view) can be encoded.

[0089] Figure 5The parameters of the projection surface can vary frequently over time in order to adapt to the pose and geometry variations between the 3D scene sequence. These parameters are chosen by the projection algorithm as a function of criteria to be respected, for example the number of patches in the projected information or the redundancy rate. From one 3D scene to its successor in the sequence, these parameters can vary leading to a modification of the number and / or size of the patches and a temporary modification of the atlas structure and related layout metadata. In order to limit these variations, the projection parameters are evaluated in small fixed length segments of N consecutive frames, typically N equals 8. Therefore, the transmitted de-projection parameters, including the layout metadata (i.e. patch data items) are regularly updated every N frames. In addition, the encoding structure of the video stream, consisting of a sequence of depth and texture patch atlases, is accordingly adapted to have a closed GOP of fixed size, N frames length, aligned. In this way, the encoding efficiency is optimized by resetting the temporal prediction at each video content change (i.e. patch atlas structure update, which creates a "scene cut" in the atlas video). This method of updating the projection surface parameters at fixed time instants every N frames is sub-optimal in terms of transmission bitrate, as a given projection surface (and thus atlas layout) can be efficient over a longer duration. If the 3D geometry of the scene does not change too fast:

[0090] • By avoiding unnecessary updates of the de-projection parameters, the metadata bitrate can be reduced;

[0091] • By avoiding too frequent unnecessary scene changes (corresponding to atlas structure updates) and adapting accordingly the encoding parameters, the video compression efficiency of the projected depth and texture atlases can be significantly improved.

[0092] According to the present disclosure, instead of evaluating the parameters of the patch-based projection surface adapted to the fixed length group of N consecutive point clouds (and thus determining the atlas layout of N consecutive atlases), the number N of consecutive frames varies over time according to the temporal evolution of the scene geometry.

[0093] Figure 8 A method 80 of encoding a sequence of 3D scenes in a data stream according to a non-limiting embodiment of the present disclosure is illustrated.

[0094] At step 81, data is acquired from a source. For example, a series of 3D scenes is acquired. Variables are initialized. In particular, a first sequence of consecutive 3D scenes is selected. The size N of the first sequence can be set to the size of a GoP for the codec selected to encode the atlases representing the 3D scenes, for example N = 8 for HEVC. The maximum size can also be initialized. Therefore, the 3D scenes of the first sequence are from index i to index i+N-1, where i is the index of the first 3D scene of the first sequence. For the sake of clarity of the present description, the index n is initialized to 0.

[0095] At step 82, a atlas layout is constructed for the first sequence according to known methods. As shown in Figure 5 The set of patches is generated by projecting the points of the 3D scene onto a projection surface. The patches are packed in N atlases according to the same layout. The number of patches packed in each atlas is called the size of the layout. The obtained layout is stored in a table S indexed by n in memory. At step 83, n is incremented. A second sequence of 3D scenes is constructed by appending i+N scenes of the sequence to the first sequence. That is, the upcoming 3D scenes are added to the first sequence to form a second sequence. n is also incremented. At step 84, a new atlas layout is constructed for the second sequence. The second layout is stored in the layout table indexed by n which has been incremented.

[0096] At step 85, the size of the second layout S[n] is compared to the size of the first layout S[n-1]. If the number of patches of the second layout is less than or equal to the number of patches of the first layout, the method iterates at step 83. N and n are incremented, the second sequence becomes the first sequence and a new second sequence is established by appending the next 3D scene to the new first sequence. An atlas layout is constructed for this new second sequence and stored in the table S. The method iterates when the size of S[n] is less than or equal to S[n-1]. In a variant, the iteration ends when the second sequence comprises a number of 3D scenes that exceeds a given number, for example 9, 10, 128, 256 or 512. This given number can be defined according to the maximum intra period size of the codec chosen to encode the atlas.

[0097] At step 86, the table S comprises n+1 layouts. The last one, i.e. the layout built for the last second sequence, is removed because its patch number exceeds the patch number of the first first layout (or because the number of stored layouts exceeds the maximum intra-period size in a variant). One of these layouts is selected for generating the n atlases for the last first sequence. The selected layout can be the last one or a layout stored in the table. In another embodiment, when n consecutive layouts {S[k]}, k e [1, K] have been evaluated, the method selects the atlas layout that generates the patch of the texture and depth atlas video with the best compression properties. First, for each computed atlas layout, it is verified whether all points from the final point cloud segmentation [1, N] can be paired with one of its patches, thus yielding a set of valid candidate patch sets {S[k]}, k e [1, K'], K' < K. Then, for each candidate patch set, the depth and texture atlas videos of each incremental sequence of the 3D scene are encoded with the same encoding parameters (i.e. in a single GOP) and the patch set with the best compression properties is selected. Only the bitrate is considered and the patch set yielding the smallest compressed atlas file size R is selected. In a variant, the rate-distortion optimization (RDO; as described in G. Sullivan and T. Wiegand "Rate-Distortion Optimization for Video Compression") method described by equation 1 is used with the constraint R c .

[0098] Equation 1 with the constraint R < R c

[0099] This optimization task can be solved using Lagrangian optimization, where the distortion term is weighted against the rate term as in equation 2.

[0100] Equation 2

[0101] For a given value of the Lagrangian coefficient λ, each solution of equation 2 is a solution of equation 1 for a given constraint Rc.

[0102] The distortion D is thus needed. Image-based criteria based on the rendered images at the user's viewpoint are superior to point-to-point distortion of compressed point clouds because they are closer to the user experience. More precisely, a pixel-to-pixel distance metric is evaluated between the rendered frames before and after the stereo video compression and averaged over those frames belonging to the predefined viewpoint path. In other variants, other criteria can be considered depending on the video stream properties that must be guaranteed or optimized.

[0103] The pseudo-code of the method 80 can be:

[0104]

[0105]

[0106] The depth and texture atlas videos generated by the projection method described above can be encoded in any conventional video coding standardization method, such as HEVC. However, these atlas videos have specific properties inherent to their generation process: between two atlas structure updates, the video content is highly temporally correlated, as frames consist of the same patch layout (e.g. rectangular) including partial projections of depth or texture. The atlas layout update to a new set of 3D scene input breaks the temporal consistency and can be described as a "scene cut". This property can be exploited to optimize the compression efficiency by setting appropriate encoding parameters accordingly.

[0107] The usual temporal organization of coded pictures is based on Groups of Pictures (GOPs). Typically, a GOP is 8 pictures long. An intra period is usually composed of several GOPs.

[0108] To benefit from the temporal predictability of frames in an atlas video, the method according to the disclosure aligns the variable GOP and intra period structure with the atlas updates, a new intra period starting at each atlas update.

[0109] Figure 9 An example structure of an intra period of atlas layout of different longevity N is shown. In this example, the intra period starts with a first GOP of fixed size k pictures long. For example, k is equal to 8 or 10. (|N / k|-1) fixed size GOPs follow. The last GOP of the intra period has a shortened length N modulo the fixed size. In this example, the intra period ends with a GOP of shortened length N. Figure 9 In the example of Fig. 9, a GOP contains 8 pictures. The first intra period 93 contains the first GOP 91 of 8 atlas and 3 shortened GOPs 92 of 3 atlas. GOP 92 has been shortened because appending the next atlas to this sequence would increase the number of patches and therefore modify the layout. Figure 9 Another intra period 96 is shown, which includes two GOPs 94 and 95 of eight pictures. At this stage, the intra period 96 is not over yet and will receive at least another GOP which can be shortened.

[0110] Such a structure can be embedded into the general elements of the ISOBMMF syntax, e.g. to signal de-projection parameter metadata of different durations as described in accordance with the present disclosure. The de-projection parameters of a given patch-based projection surface, including the patch list with its characteristics (i.e. patch data item) and the associated atlas of packed patches (layout metadata), are defined as metadata samples with structured sample format. The de-projection metadata samples are placed in a timed metadata track, with samples of different durations. The two video tracks of the depth and texture atlas-based projection are combined in one track group. The de-projection metadata track is linked to the projected video track group by a track parameter (i.e. with the 'cdtg' track parameter).

[0111] The synchronization of the duration of the timed metadata samples, which do not match the duration of the associated video samples, the projected depth / texture video frames and the de-projection metadata at the rendering side is solved by parsing the sample decoding time.

[0112] The implementations described herein can be implemented in, for example, a method or a process, an apparatus, a computer program product, a data stream, or a signal. Even if only discussed in the context of a single implementation form (for example, as only a method or as only an apparatus), the implementation discussed can be implemented in other forms (for example, a program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, an apparatus such as a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, smart phones, tablets, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.

[0113] Implementations of the various processes and features described herein can be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and associated texture and / or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video coder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a cell phone, a PDA, and other communication devices. As should be apparent, the equipment can be mobile, even if the equipment is installed in a moving vehicle.

[0114] Moreover, these methods can be implemented by way of machine, hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks can be stored in a processor readable medium such as a storage medium or other storage(s). A processor readable medium can include any medium that can store or transfer information in a form readable by a machine such as a computer or other electronic device. Examples of a non-exhaustive list of

[0115] It will be apparent to those skilled in the art that implementations can involve various signals. Naturally, these signals can be formatted in a variety of different ways and can be transmitted on various types of media. Briefly, non-limiting examples of media that can be used to transmit or store data include physical media such as volatile and non-volatile storage media, physical transmission media such as physical wires, physical taps, physical interfaces, physical busses, physical channels and physical connections. Non-limiting examples of physical storage media include compact discs (CDs), digital versatile discs (DVDs), floppy discs, memory cards, hard drives, optical discs and recordable media. Non-limiting examples of physical transmission media include telephone lines, wireless transmission media such as microwave inter-core access, free-space optics, global areas of refraction, and global areas of reflection. It will be appreciated that other types of media can be used to transmit or store data usable by implementations. In general, a computer-readable medium can include virtually any medium that can store or transfer data for use by or in connection with an implementation. Non-limiting examples of computer-readable media include physical computer storage media and physical transmission media.

[0116] Many implementations have been described. However, various modifications can be made without departing from the scope of the application. For example, the elements of one implementation can be combined with those of another implementation. Similarly, some elements of each implementation can be left out, omitted, or never used. Further, substitutions of equivalent elements can be made. Therefore, it is contemplated that these and other implementations cover any and all modifications that fall within the scope of the application.

Claims

1. A method comprising: a) determining a first atlas layout for a first sequence of 3D scenes, wherein the atlas layout defines an organization of at least one patch within an atlas; an atlas being a picture packing at least one patch of the same 3D scene; and a patch being a picture representing a projection of a portion of a 3D scene on an image plane; b) determining a second atlas layout for a second sequence of 3D scenes, the second sequence being the first sequence to which one 3D scene is appended; in case the number of patches of the second atlas layout is less than or equal to the number of patches of the first atlas layout, iterating steps a) and b) using the second sequence as first sequence and storing the first atlas layout; and otherwise, selecting an atlas layout among the stored atlas layouts and generating an atlas sequence for the first sequence of 3D scenes using the selected atlas layout.

2. The method of claim 1, wherein, The iteration ends when the second sequence of 3D scenes comprises more than a given number of 3D scenes.

3. The method of claim 1 or 2, wherein, The selection of the layout is performed based on a rate-distortion optimization criterion.

4. The method of one of claims 1 to 3, further comprising encoding the generated atlas sequence as an intra-frame period in a video data stream.

5. The method of claim 4, wherein, The intra-frame period comprises at least one group of pictures, the group of pictures comprising a number of atlases equal to the number of scenes of the initial first sequence of 3D scenes.

6. A device comprising a memory storing instructions causing a processor to perform: a) determining a first atlas layout for a first sequence of 3D scenes, wherein wherein the atlas layout defines an organization of at least one patch within an atlas; an atlas being a picture packing at least one patch of the same 3D scene; and a patch being a picture representing a projection of a portion of a 3D scene on an image plane; b) determining a second atlas layout for a second sequence of 3D scenes, the second sequence being the first sequence to which one 3D scene is appended; in case the number of patches of the second atlas layout is less than or equal to the number of patches of the first atlas layout, iterating steps a) and b) using the second sequence as first sequence and storing the first atlas layout; and otherwise, selecting an atlas layout among the stored atlas layouts and generating an atlas sequence for the first sequence of 3D scenes using the selected atlas layout.

7. The apparatus of claim 6, wherein, The iteration ends when the second sequence of 3D scenes comprises more than a given number of 3D scenes.

8. The apparatus of claim 6 or 7, wherein, The selection of the layout is performed based on a rate-distortion optimization criterion.

9. The apparatus of one of claims 6 to 8, wherein, The processor is further configured to implement encoding the generated atlas sequence as an intra-frame period in a video data stream.

10. The apparatus of claim 9, wherein, The intra-frame period comprises at least one group of pictures, the group of pictures comprising a number of atlases equal to the number of scenes of the initial first sequence of 3D scenes.

Citation Information

Patent Citations

  • Method, apparatus and stream for immersive video format

    EP3349182A1