Method for rendering a 3D scene on a terminal, method for generating a 3D scene at a server, and terminal

By segmenting the 3D scene into angle sectors and depth ranges, and only the patches and drawings of useful parallax information are transmitted, the problems of inefficiency and delay in 3D scene transmission and rendering in the prior art are solved, and more efficient rendering performance is achieved.

CN114208201BActive Publication Date: 2025-06-17INTERDIGITAL VC HOLDINGS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080055215.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-15
Filing Date
2020-07-15
Publication Date
2025-06-17
Estimated Expiration
2040-07-15

AI Technical Summary

Technical Problem

The prior art has inefficiency and latency problems when transmitting and rendering 3D scenes, especially in complex scenarios and high bandwidth requirements.

Method used

By dividing the space into angular sectors and depth ranges, patches and drawings containing only useful parallax information are generated and transmitted, and the transmission content is dynamically adjusted according to the terminal's delivery standards.

Benefits of technology

It realizes the efficiency of 3D scene transmission and rendering, reduce latency, and provide faster rendering performance under complex scenes and high bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114208201B_ABST
    Figure CN114208201B_ABST
Patent Text Reader

Abstract

The present disclosure discloses methods and devices for transmitting and rendering a 3D scene. The method for transmission includes: dividing a space into m angular sectors, each of the m angular sectors corresponding to an angular distance from a viewport, and dividing the space into n depth ranges; obtaining (11) at least one first patch generated from a first view of the 3D scene, the at least one first patch including a texture component and a depth component; obtaining (12) at least one atlas generated from at least one second view of the 3D scene, the at least one atlas being constructed by packing together at least one second patch generated from at least one point that is invisible in another view of the 3D scene and belongs to the same angular sector among the m angular sectors and the same depth range among the n depth ranges for one of the second views, at least one of m or n being greater than or equal to 2, the at least one second patch including a texture component and a depth component, wherein each of the at least one first patch and the at least one second patch is based on at least one of a sector and a depth; generating (13) a first subset of streams including m' pairs of streams and a second subset of streams including m'×n' pairs of streams; and transmitting (14) the first subset of streams and the second subset of streams to the terminal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of European Patent Application No. 19305939.1, filed on July 15, 2019, the content of which is incorporated herein by reference. Technical field

[0003] The present disclosure relates to the field of video processing and, more particularly, to the field of volumetric video content. The present disclosure provides a technique for adaptively transmitting a representation of a 3D scene to a terminal by considering at least one terminal - based delivery standard. Such adaptive transmission can be used to enhance the rendering of the 3D scene, for example, for immersive rendering on terminals such as mobile or head - mounted display devices (HMDs).

[0004] The present disclosure can be adapted to any application that must deliver volumetric content, particularly 3DoF + video content. Background art

[0005] This section is intended to introduce aspects of the art that may be related to aspects of the present disclosure described and / or claimed below. This discussion helps to provide background information to facilitate a better understanding of the aspects of the present disclosure. Accordingly, it should be understood that these statements should be read from this perspective and not as an admission of prior art.

[0006] Immersive video (also known as 360° planar video) allows a user to view everything around themselves by rotating their head around a stationary viewing angle. The rotation allows only a 3 - degree - of - freedom (3DoF) experience. Even though 3DoF video is sufficient to meet the requirements of a first omnidirectional video experience (e.g., using an HMD device), 3DoF video may quickly become frustrating for viewers who expect more freedom (e.g., by experiencing parallax). In addition, 3DoF may also cause dizziness because the user not only rotates their head but also translates their head in three directions, and these translations are not reproduced in the 3DoF video experience.

[0007] Volumetric video (also known as 6 - degree - of - freedom (6DoF) video) is an alternative to 3DoF video. When viewing 6DoF video, in addition to rotation, the user can also translate their head and even their body within the viewed content and experience parallax and even volume. Such video significantly increases the immersion and the perception of scene depth and prevents dizziness by providing consistent visual feedback during head translation.

[0008] An intermediate approach between 3DoF and 6DoF is also proposed, called 3DoF+. This video-based method (e.g., disclosed in WO2019 / 055389) includes transmitting volumetric input information as a combination of color and depth patches. Each patch is generated by a sequential spherical 2D projection / mapping of a sub-part of the original 3D scene.

[0009] Basically, this decomposition strips / breaks down the scene into: (1) a central patch that contains the part of the scene visible from the main central viewpoint; and (2) peripheral patches that are embedded with supplementary information not visible from this central viewpoint.

[0010] To transmit 3DoF+ video content, the following two video frames are defined: (1) a color frame that carries the texture of both the central patch and the peripheral patches to carry parallax information; and (2) a depth frame that carries the depth of both the central patch and the peripheral patches to carry parallax information.

[0011] To limit the number of decoder contexts, the color frame and the depth frame have a fixed size, which corresponds to the size of the central patch (e.g., 4K pixels × 2K pixels) plus an additional room size to carry parallax information from the source viewpoint in all 360° directions.

[0012] However, packing parallax information into a fixed-size frame is sufficient for simple scenes with not many hidden objects, but may be inefficient for the transmission of complex scenes, where many hidden objects require a large amount of data for the peripheral video patches and parallax information. In addition, the prior art 3DoF+ techniques have latency in rendering 3D scenes. For example, this may occur when an HMD user quickly turns their head in one direction. According to the prior art, the rendering terminal has to wait for the color frame to be received and wait for the depth frame to be received for volumetric rendering before displaying any content. Summary of the Invention

[0013] Therefore, a new technique is needed to transmit 3D scenes that overcomes at least one of the disadvantages of the known techniques.

[0014] According to one aspect of the present disclosure, a method for transmitting a representation of a 3D scene to a terminal is disclosed. The method includes: dividing a space into m angular sectors, each of the m angular sectors corresponding to an angular distance from a viewport, and dividing the space into n depth ranges; obtaining at least one first patch generated from a first view of the 3D scene, the at least one first patch including a texture component and a depth component; obtaining at least one atlas generated from at least one second view of the 3D scene, the at least one atlas being constructed by packing together at least one second patch generated for one of the second views for at least one point that is invisible in another view of the 3D scene and belongs to the same angular sector among the m angular sectors and the same depth range among the n depth ranges, at least one of m or n being greater than or equal to 2, the at least one second patch including a texture component and a depth component, wherein each of the at least one first patch and the at least one second patch is based on at least one of a sector and a depth; generating the following items according to at least one terminal-based delivery criterion: a first stream subset that includes m' streams from the one or more first patches, where m' is the whole or a subset of the m angular sectors; and a second stream subset that includes m'×n' streams from the at least one atlas, where m'≤m and n'≤n, each stream including a stream for transmitting the texture component and a stream for transmitting the depth component, and transmitting the first stream subset and the second stream subset to the terminal.

[0015] According to the present disclosure, therefore, it is possible to transmit only a subset of the streams for transmitting the depth component and the texture component to the terminal, taking into account at least one terminal-based delivery criterion.

[0016] More specifically, for at least one second view, points (or voxels) that are invisible in another view (the first view or another second view) of the second view can be identified, and the depth range and / or angular sector to which these points belong can be determined. Therefore, the second patches that can be used to transmit disparity information obtained from these points can be grouped in an atlas, with at least one atlas for each depth range and / or each angular sector.

[0017] In this way, it is possible to transmit only the disparity information that is "useful" to the terminal (user), rather than transmitting all the disparity information. For example, it is possible to transmit only the disparity information corresponding to the viewpoint of the terminal user, or only the disparity information corresponding to the minimum depth range from the user's viewpoint, especially when the available bandwidth of the communication channel with the terminal is limited.

[0018] Accordingly, at least one embodiment of the present disclosure aims to solve the problem of the fixed-size framework according to the prior art. In fact, only useful parallax information can be transmitted, thus solving the problem of complex scenes or heterogeneous scenes, where the parallax information in some sectors of the 360° space is poor, while the amount of parallax information in other sectors is very large, which may not be suitable for the additional room size.

[0019] At least one embodiment of the present disclosure also aims to solve the latency problem in rendering. In fact, only useful parallax information can be transmitted, thus enabling fast rendering.

[0020] According to another embodiment, a corresponding device for transmitting a representation of a 3D scene to a terminal is disclosed. Such a device may be particularly suitable for implementing the method for transmitting the representation of the 3D scene described above. For example, such a device is a server.

[0021] The present disclosure also discloses a method for rendering a 3D scene on a terminal. This method includes: dividing the space into m angular sectors, each of the m angular sectors corresponding to an angular distance from the viewport, and dividing the space into n depth ranges; receiving a first stream subset and a second stream subset generated according to at least one terminal-based delivery criterion, the first subset including m′ pairs of streams generated from at least one first patch and the second subset including m1×n′ pairs of streams generated from at least one atlas, each pair of streams including a stream for transmitting a texture component and a stream for transmitting a depth component, m1 being the whole or a subset of the m angular sectors and n1 being the whole or a subset of the n depth ranges, the at least one first patch being generated from a first view of the 3D scene and including a texture component and a depth component, the at least one atlas being generated from at least one second view of the 3D scene and constructed by packing together at least one second patch generated for at least one point in the 3D scene that is invisible in another view of the 3D scene and belongs to the same angular sector among the m angular sectors and the same depth range among the n depth ranges, at least one of m or n being greater than or equal to 2, the at least one second patch including a texture component and a depth component, where m′≤m and n′≤n, where each of the at least one first patch and the at least one second patch is based on at least one of a sector and a depth; and constructing a representation of the 3D scene from the first stream subset and the second stream subset.

[0022] Specifically, such a method can be implemented to render the 3D scene transmitted by the method for transmitting the representation of the 3D scene as described above.

[0023] As already mentioned, since the terminal can only receive "useful" parallax information, the method according to at least one embodiment allows for fast rendering of the 3D scene.

[0024] According to another embodiment, a corresponding terminal for rendering a 3D scene is disclosed. Such a terminal (also referred to as a device for rendering) may be particularly adapted to implement the method for rendering the above-mentioned 3D scene. For example, such a device is an HMD, a mobile phone, a tablet computer, etc.

[0025] The present disclosure also discloses a method for generating patches representing a 3D scene. This method includes: obtaining a first view of the 3D scene from a first viewpoint; generating at least one first patch from the first view, the at least one first patch including a texture component and a depth component; obtaining at least one second view of the 3D scene from at least one second viewpoint; and dividing the 3D scene space into m angular sectors, each of the m angular sectors corresponding to a distance from a given viewport and into n depth ranges, wherein for at least one of the at least one second views in the second views, the method further includes: identifying at least one point in the second view that is not visible in another view of the 3D scene; determining the depth range to which the at least one point belongs; for at least one of the m angular sectors and for at least one of the n depth ranges, at least one of m or n is greater than or equal to 2, generating at least one second patch from the second view for points belonging to the angular sector and the depth range, the at least one second patch including a texture component and a depth component, wherein each of the at least one first patch and the at least one second patch is based on at least one of the sector and the depth; and constructing at least one atlas by packing together at least one of the second patches generated for points belonging to the same angular sector and the same depth range.

[0026] Specifically, this method can be implemented to generate patches and atlases obtained by the method for transmitting a representation of a 3D scene as described above.

[0027] According to a first embodiment, the method for generating patches and the method for transmitting a representation of a 3D scene can be implemented by the same device (e.g., a server).

[0028] According to a second embodiment, the method for generating patches and the method for transmitting a representation of a 3D scene can be implemented by two different devices, and the two devices can communicate via wire or wireless according to any communication protocol.

[0029] Therefore, a corresponding device for generating patches representing a 3D scene according to the second embodiment is disclosed. Such a device may be particularly adapted to implement the method for generating patches representing the above-mentioned 3D scene.

[0030] Another aspect of the present disclosure relates to at least one computer program product downloadable from a communication network and / or recordable on a computer-readable and / or processor-executable medium, the at least one computer program product including software code adapted to execute a method for transmitting a representation of a 3D scene, a method for rendering a 3D scene, or a method for generating a patch representing a 3D scene, wherein the software code is adapted to execute at least one step of the above methods.

[0031] Furthermore, another aspect of the present disclosure relates to a non-transitory computer-readable medium including a computer program product recorded thereon and executable by a processor, the computer program product including program code instructions for implementing a method for transmitting a representation of a 3D scene, a method for rendering a 3D scene, or a method for generating a patch representing the previously described 3D scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] With reference to the accompanying drawings, the present disclosure will be better understood and illustrated by the following embodiments and implementation examples, which are in no way limiting, in which:

[0033] Figure 1 is a flowchart showing a method for transmitting a representation of a 3D scene according to an embodiment of the present disclosure;

[0034] Figure 2 is a flowchart showing a method for generating a patch representing a 3D scene according to an embodiment of the present disclosure;

[0035] Figure 3 is a flowchart showing the main steps of a method for processing a 3D scene according to an embodiment of the present disclosure;

[0036] Figure 4 shows the position of a camera for generating a patch according to the prior art;

[0037] Figure 5 gives an example of a patch generated by a peeling technique according to the prior art;

[0038] Figure 6A and Figure 6B gives an example of a patch generated according to the present disclosure;

[0039] Figure 7 shows an example of a depth-first representation;

[0040] Figure 8 shows an example of a sector and a depth-first representation; and

[0041] Figure 9It is a block diagram of a device implementing at least one of a method for generating patches representing a 3D scene, a method for transmitting a representation of a 3D scene, or a method for rendering a 3D scene according to at least one embodiment of the present disclosure.

[0042] In the drawings, the represented blocks are pure functional entities, which do not necessarily correspond to physically separated entities. That is, they can be developed in the form of software, hardware, or implemented in one or more integrated circuits, including one or more processors. Detailed implementation

[0043] It should be understood that the drawings and descriptions of the present disclosure have been simplified to illustrate elements relevant to a clear understanding of the present disclosure, while eliminating many other elements found in typical transmission or rendering devices for clarity.

[0044] The general principles of the present disclosure will be discussed below.

[0045] The present disclosure proposes a technology for volumetric data organization and associated terminal - related delivery modes (e.g., viewport - related).

[0046] According to at least one embodiment, this technology provides progressive rendering on a terminal, thereby reducing latency by delivering the first basic elements for immediate volumetric rendering.

[0047] This technology relies on a new method to construct patches containing parallax information (volumetric data), thereby allowing patches to be constructed according to the viewing point (e.g., a real camera or a virtual camera) and / or the point position within the space (i.e., the position of the point / voxel in the 3D scene relative to the viewing point): the farther, the less important. The criteria for determining the priority of volumetric data elements (point positions) can be depth (distance from the viewing point), angular sector (distance from the center of the delivered viewport), or a combination of both. For example, the client can first download the video information required for basic planar 360° rendering and further download improved data for the parallax experience according to the available throughput.

[0048] According to at least one embodiment, the volumetric data is thus organized in a list of video frames, which may have the same size (e.g., 4K), but with different patch arrangements, thereby allowing each sector of the 360° space and each distance to the source viewing point (e.g., near to far) to be rendered.

[0049] The volumetric data can be included in a variable patch list, and for a given spatial sector, the content of the patches is distributed over the transmission of consecutive video frames.

[0050] In order to be able to switch from one viewpoint to another while optimizing the amount of received data, volumetric content can be segmented into chunks with a fixed duration. The chunks stored on the server side show a three-level organization: each time interval, each sector, and each depth (i.e., level of detail) to the source viewpoint. Thanks to this method, the terminal (or client application) can retrieve data in order of priority: first, the necessary video information for planar 360° rendering, and then, depending on the available throughput, improved data for the parallax experience. This priority of data retrieval may be proportional to the proximity of the user's position in the scene. This means that video patches and associated metadata corresponding to more objects can only be used when network resources are sufficient.

[0051] Now in conjunction with Figures 1 to 3 present at least one embodiment of the present disclosure.

[0052] Figure 1 FIG. schematically shows the main steps implemented by a device (e.g., a server) for transmitting a representation of a 3D scene. According to this embodiment, the server (10) obtains (11) at least one first patch, which includes a texture component and a depth component. Such one or more first patches (also referred to as main patches or central patches) can be generated from a first view (also referred to as the main view or source view) of a 3D scene captured from a first viewpoint (by a real camera or a virtual camera). This first view can be a projective representation of the 3D scene.

[0053] The server also obtains (12) at least one atlas. Such one or more atlases can be generated from at least one second view of the 3D scene obtained from at least one second viewpoint (by a real camera or a virtual camera). More specifically, for one second view in the second views (and advantageously for each second view in the second views), at least one second patch (also referred to as a peripheral patch) can be generated. In order to reduce the amount of data that has to be transmitted, such second patches can be generated only for points in the second view that are not visible in the first view or in a second view captured from another viewpoint. Taking into account the angular sectors and / or depth ranges to which the corresponding points belong, such second patches can be packed or grouped together in the atlas. In this way, several angular sectors centered on one of the viewpoints and / or several depth ranges originating from one of the viewpoints can be considered, and at least one atlas for each angular sector and / or each depth can be constructed and obtained by the server. For example, a first depth range corresponds to a distance between 0 and 50 cm from one of the viewpoints, a second depth range corresponds to a distance between 50 cm and 1 m from one of the viewpoints, a third depth range corresponds to a distance between 1 m and 2 m from one of the viewpoints, and a fourth depth range corresponds to a distance greater than 2 m.

[0054] It can be noted that the steps for obtaining the first patch and obtaining the atlas can be implemented simultaneously or consecutively in any order.

[0055] After obtaining the first patch and the atlas, the server can generate (13) the following streams based on at least one terminal-based delivery criterion: (1) a first stream subset that includes m' pairs of streams from the one or more first patches; and (2) a second stream subset that includes m'×n' pairs of streams from the one or more atlases, where m'≤m and n'≤n, and each pair of streams includes a stream for transmitting a texture component and a stream for transmitting a depth component.

[0056] For example, if the bandwidth of the communication channel between the server and the terminal is very large, there is no need to transmit only a subset of the streams: m' can be equal to m and n' can be equal to n. On the contrary, if the bandwidth is limited, m' can be equal to 1 and n' can be equal to n, or m' can be equal to m and n' can be equal to 1, or other combinations.

[0057] Then, the server can transmit (14) or deliver the first stream subset and the second stream subset to the terminal. Thus, the first patch and the second patch are transmitted in different frames.

[0058] For example, the terminal-based delivery criterion can be selected from the group including: the bandwidth available on the communication channel between the terminal and the server, at least one angular sector observed by the terminal user, the capabilities of the terminal, and the request received from the terminal.

[0059] The generation of the streams and the transmission of the streams can be implemented periodically and / or after the at least one terminal-based delivery criterion changes.

[0060] In this way, the generation of the streams to be transmitted to the terminal can be adapted to the terminal. Specifically, the generation of the streams can change over time to adapt the content carried by the streams to the terminal and, for example, to the viewpoint of the terminal user. The generation of the streams can be determined by the server, for example, after analyzing the available bandwidth or according to a request from the terminal.

[0061] According to at least one embodiment, the server obtains all the first patches generated from the first view, and all the atlases for each angular sector and each depth range generated from all the second views of the 3D scene.

[0062] In this way, the server can have complete knowledge of the 3D scene and can generate only the streams useful for the terminal based on at least one terminal delivery criterion. Specifically, the server can generate a first stream set including m pairs of streams from all the first patches, and a second stream set including m×n pairs of streams from all the atlases, where each pair of streams includes a stream for transmitting a texture component and a stream for transmitting a depth component.

[0063] According to the first embodiment, the first patch, the second patch, and the corresponding atlas can be generated by such a server. In this first embodiment, the steps for obtaining (11) the first patch and obtaining (12) the atlas can correspond to the steps for generating the first patch and generating the atlas.

[0064] According to the second embodiment, the first patch, the second patch, and the corresponding atlas can be generated by another device for generating patches and then transmitted to the server. In this second embodiment, the steps for obtaining (11) the first patch and obtaining (12) the atlas can correspond to the steps for receiving the first patch and receiving the atlas.

[0065] Figure 2 Illustrated are the main steps for generating the first patch and the atlas implemented by a device for generating patches according to such a second embodiment. According to this embodiment, such a device (20) for generating patches can include a memory (not shown) associated with at least one processor, the processor being configured to: obtain (21) a first view of a scene from a first viewpoint; generate (22) at least one first patch from the first view, the at least one first patch including a texture component and a depth component; and obtain (23) at least one second view of the scene from at least one second viewpoint. For at least one second view of the scene (and advantageously for each second view), the at least one processor is further configured to: identify (24) at least one point in the second view that is not visible in another view (the first view or another second view) of the 3D scene; determine (25) the depth range to which the at least one point belongs, where for at least one of m angular sectors centered on one of the viewpoints and for at least one of n depth ranges originating from one of the viewpoints, at least one of m or n is greater than or equal to 2; generate (26) at least one second patch for points belonging to the angular sector and the depth range, the at least one second patch including a texture component and a depth component; and construct (27) at least one atlas by packing together at least one (preferably all) second patch generated for points belonging to the same angular sector and the same depth range.

[0066] According to the first embodiment or the second embodiment, the first patch can be generated by projecting a first view of the 3D scene onto a 2D representation. For example, such a 2D projection can be an equirectangular projection (ERP) or a cube projection, such as proposed in the omnidirectional media format (OMAF) standard currently being developed by the Moving Picture Experts Group (MPEG). Other 3D-to-2D projection representations can also be used. For more complex projections, rectangles in the projected picture may map to more complex 3D regions than angular sectors, but it can be advantageously ensured that there is a one-to-one correspondence between the tiles and the sub-parts of the point cloud.

[0067] According to at least one embodiment, the description data that describes the organization of the first stream subset and the second stream subset can also be transmitted from the server to the terminal. Before transmitting the first stream subset and the second stream subset, such description data can be transmitted in a manifest file. The description data can be transmitted offline on a dedicated channel in response to a request from the terminal, or previously stored in the terminal and downloaded from the server when the present disclosure is first used, and so on.

[0068] For example, the description data can include: (1) the number of available depth ranges and their values; (2) the number of available angular sectors and their positions; (3) the resolution of one or more atlases for each stream of the second subset, and whether the atlases are packed together in a GOP; (4) the average bitrate of each GOP and each stream of the second stream subset. The description data can also include the positions of patches within the 3D scene, for example, represented in spherical coordinates. The terminal can use the description data to select and decode the streams and render the 3D scene.

[0069] Figure 3 The main steps for a terminal to implement for rendering a 3D scene are schematically shown in. According to this embodiment, the terminal (30) receives (31) a first stream subset and a second stream subset generated according to at least one terminal-based delivery criterion.

[0070] For example, the terminal-based delivery criterion can be selected from the group consisting of: the bandwidth available on the communication channel of the device used to transmit the representation of the 3D scene, at least one angular sector observed by the end user, the capabilities of the terminal, and the request sent by the terminal.

[0071] Such stream subsets can be generated by Figure 1 the server 10 shown in. For example, the terminal can send a request to the server to receive only the disparity information useful for the terminal. In a variant form, the server can analyze the terminal-based delivery criterion (e.g., the communication channel between the server and the terminal or the position / viewpoint of the end user) and select the streams that must be transmitted. For example, the server may provide only the patches corresponding to the user's field of view.

[0072] The first stream subset may include m′ pairs of streams generated from at least one first patch, and the second stream subset may include m′×n′ pairs of streams generated from at least one atlas, each pair of streams including a stream for transmitting texture components and a stream for transmitting depth components. The at least one first patch may be generated from a first view of the 3D scene. The at least one atlas may be generated from at least one second view of the 3D scene and may be constructed by packing together at least one second patch generated for at least one point that is invisible in another view of the 3D scene and belongs to the same corner sector among m corner sectors and the same depth range among n depth ranges for one of the second views, where m′≤m and n′≤n, and at least one of m or n is greater than or equal to 2. The at least one first patch and the at least one second patch may each include texture components and depth components.

[0073] Then, the terminal may construct (32) and render a representation of the 3D scene from the first and second stream subsets.

[0074] According to at least one embodiment, the second stream subset may include at least an atlas constructed for a minimum depth range originating from the viewpoint of the end user.

[0075] According to at least one embodiment, the second stream subset may include at least an atlas constructed for a corner sector centered on the viewpoint of the end user.

[0076] According to these embodiments, the viewpoint of the terminal (or end user) may first be determined by the terminal, a server, or another device. If determined by the terminal, the terminal may send a request to the server in order to obtain the second stream subset taking into account the viewpoint.

[0077] Thus, the way of generating disparity patches according to at least one embodiment of the present disclosure may allow for scalable delivery of volumetric video and thus allow for reaching a larger audience or improving the experience for the same transmission cost.

[0078] According to at least one embodiment, this way may allow for transmitting the same 3DoF+ content over heterogeneous networks with different bandwidth capabilities, since each terminal may adjust the amount of disparity information retrieved from the server according to its network characteristics.

[0079] According to at least one embodiment, this way may also aim to provide fast first rendering (low latency) on the device by implementing progressive rendering and displaying disparity information in order of importance and reception order.

[0080] A detailed description of embodiments of the present disclosure will be described below.

[0081] First, several embodiments of the present disclosure will be discussed in the context of 3DoF+.

[0082] Now briefly discuss generating patches according to the prior art. To explain the differences between the prior art and the present disclosure, the following gives tips on 3DoF+ technology for generating patches according to the prior art.

[0083] As mentioned in the prior art section, 3DoF+ has been developed to enrich the immersive video experience with parallax. The volumetric input information (e.g., volumetric video) can be decomposed into several components: the color / depth of the projection representation of the 360° scene observed from the central point (also called the first patch or the central patch); the color / depth patches of the scene parts displayed by the natural displacement of the head, also called the second patch or the peripheral patch; the metadata containing information for exploiting the patches.

[0084] Basically, the components of the volumetric video can be generated in the following ways: capturing a 360° scene, for example, by the equipment of N 360° cameras; generating a point cloud from the N-camera capture; introducing four virtual cameras, with three cameras placed at the three vertices of a tetrahedron concentric with the central observation point where the camera C0 is located, as Figure 4 shown; generating a projection representation with texture and depth, as seen from the central camera ( Figure 4 the camera C0 above), where the scene forms two video streams C0 (color) and D0 (depth), and this projection representation can be obtained by any 3D-to-2D projection (e.g., equirectangular projection (ERP) or cube map projection (CMP)); a peeling process for generating color / depth patches for points not visible to the previous cameras, where for each camera ( Figure 4 the cameras C1, C2, and C3 in it), this process can be completed iteratively; packing the central color / depth patches and the peripheral color / depth patches generated in the previous steps in a rectangular patch atlas, where the packing algorithm provides the patch positions on the GOP and the metadata is generated accordingly; encoding the atlas using a traditional HEVC video codec, where the depth atlas and the color atlas can be feathered and quantized separately in a dedicated manner first to encode the artifacts robustly enough and optimize the overall bitrate.

[0085] As Figure 5 shown, the points captured by the camera C0 can be placed in the first patches C0.I0, C0.I1, and C0.I2, where they are collected by adjacent points. Thus, the patches can be defined by adjacent point sets. The patches can be segmented according to size criteria.

[0086] Then, the peeling process can deliver the peripheral patches. As Figure 5 shown, the points captured by the camera Cl and not visible to the camera C0 are placed in the second patches C1.I0, C1.I1, where they are collected by adjacent points. This process can be implemented iteratively for each of the cameras C1, C2, and C3.

[0087] Then, a dedicated packing algorithm can place the patches in the color and depth atlases in a GOP-consistent manner (within a GOP / IntraPeriod, the patch positions remain unchanged). Then, the atlas can be encoded using a traditional HEVC video codec. For each patch, an additional metadata set can be provided that specifies the information required to reconstruct the volumetric scene (patch position / size, projection parameters). Thus, the entire stream is completely video-based and compatible with existing video streaming pipelines.

[0088] Generating patches according to the present disclosure will be described below.

[0089] According to the present disclosure, a new algorithm is proposed for generating patches per viewpoint based on the angular sectors of the central viewpoint and / or the distance from the central viewpoint to the angular sectors. The technique aims to distinguish points based on their positions. Globally, the farthest points may require less texture or depth precision.

[0090] More specifically, the components of the volumetric video can be generated as disclosed in the previous section, but can also be generated by considering the depth range and / or angular sectors to which the points of the point cloud belong.

[0091] According to a first example, the capture from the central camera (C0) delivering the first view (reference Figure 2 in 21) is not modified according to the prior art. It still provides a projective representation of the 3D scene, such as an equirectangular representation with color and depth (reference Figure 2 in 22).

[0092] According to a second example, the capture from the central camera (C0) can be modified according to the prior art. For example, the first patches are defined according to the depth range or angular sectors to which they belong.

[0093] The second patches are constructed by capturing points from various cameras (e.g., C1, C2, and / or C3 can deliver the second view) (reference Figure 2 in 23) in order to show points that were masked by the previous capture (reference Figure 2 in 24). It should be noted that cameras C0 to C3 according to the present disclosure can be real cameras or virtual cameras or a combination thereof. Additionally, the number of cameras is not limited to the four cameras disclosed in the prior art.

[0094] According to the present disclosure, the second patches can be defined by the depth range or angular sector to which they belong, rather than (or in addition to) by adjacent points. The depth here can be the distance from the central viewport (i.e., the position of C0), or the distance from the capture point (i.e., the position of C1, C2, or C3). The second method regarding the capture point is more relevant because the depth determined from the capture point may be equivalent to the depth seen by the user visualizing volumetric content. In the same way, the angular sector can be centered on the central viewport or any capture point.

[0095] Figure 6A and Figure 6B Two examples of generating second patches considering the depth of points are shown, thus allowing the patches to be adaptively used according to the depth of the points represented by the patches. This distance is considered when constructing the patches. All points of a patch must belong to the same depth range. This allows patches to be constructed according to the depth of the viewing point, enabling them to be selected accordingly for optimal delivery.

[0096] According Figure 6A to the first example shown, the space is divided into three regions D0, D1, D2, which correspond to three different depth ranges (also called distance ranges) from the central camera C0.

[0097] Thus, two patches C1.D0.I0, C1.D0.I1 and one patch C1.D1.I0 and one patch C1.D2.I0 are generated according to the present disclosure (C i represents the corresponding camera, D j represents the corresponding depth range, and I k represents the patch index within the considered depth range), while only one patch C1.I0 and one patch C1.I1 are generated according to the Figure 5 prior art shown.

[0098] According Figure 6B to the second example shown, the space is divided into three regions D0, D1, D2, which correspond to three different depth ranges (also called distance ranges) from the capture camera C1.

[0099] In this case, five patches C1.D0.I0, C1.D0.I1 and C1.D1.I0, C1.D1.I1 and C1.D2.I0 are generated according to the present disclosure (C i represents the corresponding camera, D j represents the corresponding depth range, and I k represents the patch index within the considered depth range), while only one patch C1.I0 and one patch C1.I1 are generated according to the Figure 5 prior art shown.

[0100] Thus, if the second patches are defined by grouping adjacent points according to the depth range to which they belong, five patches C1.D0.I0, C1.D0.I1 and C1.D1.I0, C1.D1.I1 and C1.D2.I0 can be generated according to the present disclosure. In a variant, if the second patches are not defined by adjacent points but according to the depth range or angular sector to which they belong, three patches C1.D0, C1.D1 and C1.D2 can be generated.

[0101] Of course, the number and size of the depth ranges are not limited to Figure 6A and Figure 6B those shown.

[0102] Once the patches are built, they may be packed into atlases together with other patches having the same depth range (even if the depth is from another viewing point).

[0103] Once all the patches / atlases for each depth and / or each sector are generated, they are stored in the memory of the device for patch generation for later use.

[0104] When the available throughput is not sufficient to deliver everything, this patching for each depth and / or each angular sector according to the present disclosure may allow giving privileges to the nearest volume data or viewport-based volume data.

[0105] For example, when the farthest patches are not delivered, the repair technique may limit the impact of the missing parts in the scene. The available throughput is dedicated to the nearest objects, which optimizes the rendering.

[0106] Before playing the content, the player / device for rendering may instantiate and configure a fixed number of video decoders without reconfiguring them during use, even though the amount of data in the atlases may vary over time.

[0107] The following description will discuss the delivery of patches.

[0108] According to the present disclosure, a new algorithm for delivering patches, i.e., transmitting the representation of a 3D scene, is also proposed.

[0109] This transmission is adaptive and depends on at least one terminal-based delivery criterion. According to at least one embodiment, this patch delivery algorithm aims to optimize the user experience according to the available network and terminal resources.

[0110] In other words, the device for transmitting the representation of a 3D scene may select some patches / atlases to be transmitted from all the patches / atlases previously generated and stored by the device for generating patches. As already mentioned, the device for generating patches and the device for transmitting the representation of a 3D scene may be the same device, such as a server.

[0111] The following discloses different methods for adaptive volumetric content delivery aimed at optimizing bitrate and player resources.

[0112] The following description will first discuss depth-based patch delivery.

[0113] According to a first example, the texture component and the depth component of a projected representation (first patch) of a 3D scene can be fully delivered from a device (e.g., server 10) for transmitting the representation of the 3D scene to a device (e.g., terminal 30) for rendering the 3D scene.

[0114] If the first patches are generated by sectors and / or by depth in the device for generating patches, they can all be transmitted to server 10, and server 10 can connect or merge the first patches to cover a 360° angular sector.

[0115] In the same way, if the second patches are generated by sectors and / or by depth in the device for generating patches, they can all be transmitted to server 10, and server 10 can connect or merge the second patches to cover a 360° angular sector.

[0116] In this depth-based method, the content can be organized in 2+(n×2) streams as follows: a first set of streams containing a pair of streams to respectively transmit the texture component and the depth component of the first patch; a second set of n pairs of streams to respectively transmit the texture component and the depth component of n atlases generated for the second patch associated with n depth ranges horizontally and the metadata associated with the atlas.

[0117] For example, the first set of streams can carry a central patch of size W×H, where W and H can depend on the visual quality defined by pixels per degree (PPD). For example, a 4K×2K frame provides a quality of 4K / 360° = 11 pixels per degree.

[0118] These atlases can be put together in the form of a group of pictures (GOP). For all streams, the duration of the GOP may be the same and does not always contain frame numbers.

[0119] A manifest can describe the organization of the different streams.

[0120] For example, the manifest indicates: for each stream associated with depth ranges d = 1..n, the number n of available depth ranges and their values, the resolution Wd×Hd of the atlas carried by the stream; and for each stream associated with depth ranges d = 1..n, for each GOP index t, the average bitrate Ratet,d.

[0121] The value of the resolution Wd×Hd of the atlas can be defined, for example, as: at least equal to the average number of points (i.e., pixels) per second of patches within the depth range d; or at least equal to the maximum number of points (i.e., pixels) per second of patches within the depth range d.

[0122] In the latter case, each rendered video frame may have exactly one atlas frame.

[0123] As already mentioned, the manifest can be transferred offline at the start of content distribution (in the same or a dedicated channel) or via any suitable means, such as an explicit request from the client (terminal) to the server.

[0124] If there is no bandwidth limitation, the server can transfer the first set of streams (containing a pair of streams) and the second set of streams (containing n pairs of streams) to the terminal.

[0125] In a variant form, for each depth range d, knowing the necessary bandwidth Ratet,d, the terminal can select a first subset of streams and a second subset of streams. For example, as discussed above, the projected representation (first patch) of the 3D scene can be fully delivered to the terminal, and the first subset of streams can be the same as the first set of streams. The second subset of streams includes n' pairs of streams, where n' ≤ n, and the number of atlas streams to be downloaded is selected according to at least one terminal-based criterion, such as available bandwidth or terminal capabilities.

[0126] Streams corresponding to the nearest depth can be preferentially downloaded.

[0127] Rendering can be decoupled from the full reception of all streams and can start immediately after the first atlas stream is completed. This can allow for dynamic progressive rendering. The first level of detail brought by the first atlas stream (for d = 1) is first rendered with the lowest latency and is gradually completed by receiving the next streams to be processed (for d = 2...n').

[0128] In the absence of sectorization (i.e., with one 360° angular sector), the priority of the patches retrieved by the rendering device may be the depth index, where the smallest index corresponds to the shortest distance from the viewpoint, as Figure 7 shown.

[0129] According to at least one embodiment, for the same content, the number n of available depth ranges can vary over time. For example, this number can be reduced to one for most of the time (e.g., by merging atlases generated for different depth ranges on the server side) and can increase during periods when the scene becomes more complex. In this case, the adaptive behavior of the player may allow it to select only the most essential depth atlases based on its available bandwidth.

[0130] The following description will first discuss viewport-based patch delivery.

[0131] According to a second example, texture components and depth components of a projection representation (first patch) of a 3D scene can be delivered partially from a device (e.g., server 10) for transmitting a representation of the 3D scene to a device (e.g., terminal 30) for rendering the 3D scene in a viewport-based manner.

[0132] In fact, volumetric content may require delivering a large amount of data, and thus, this does not always fit existing networks where the bandwidth may be limited. Therefore, such content is typically delivered partially in a viewport-based manner.

[0133] For example, high-quality content (e.g., for a full-scene representation, 8K 3DoF+ content or more) can be tiled in m corner sectors (for longitude, [Θi1, Θi2], for latitude, centered on the central viewport (i.e., the position of C0) or any capture point (i.e., positions C1, C2, and C3). The second method regarding the capture point is more relevant because the corner sectors observed from the capture point may be equivalent to the corner sectors seen by the user visualizing the volumetric content.

[0134] For each sector, a set of streams carrying volumetric data corresponding to the subpart of the scene is disclosed.

[0135] In this viewport-based method, the content can be organized in 2 + (n×2) streams as follows: a first set containing m pairs of streams to transmit the texture components and depth components of the first patch for the m sectors respectively; a second set containing m×n pairs of streams to transmit the texture components and depth components of n atlases generated for the second patch and metadata associated with the atlas for the m sectors respectively, which are horizontally associated with n depth ranges.

[0136] If there is no bandwidth limitation, the server can transmit the first set of streams (containing m pairs of streams) and the second set of streams (containing m×n pairs of streams) to the terminal. In this case, if the available bandwidth is sufficient to deliver all the scene to the player, the depth-based delivery is actually a viewport-based delivery where m = 1 (e.g., only one sector).

[0137] In a variant form, the client can select a first subset of m' pairs of streams and a second subset of m'×n' pairs of streams, where m' ≤ m and n' ≤ n, and the number of streams to be downloaded is selected according to at least one terminal-based criterion such as available bandwidth or terminal capabilities.

[0138] On the rendering device, for each time interval (GOP), the next viewport and the sectors covering this next viewport can be predicted. Therefore, the terminal can download only the streams related to this part from the server during the next GOP duration. This can be repeated for each GOP.

[0139] In another embodiment, for over - supply purposes, in addition to the stream related to the next predicted viewport, the terminal may also download supplementary streams to cover adjacent portions of the predicted viewport.

[0140] According to at least one embodiment, an atlas may be defined by depth and angular sectors. In this case, the priority of the atlas may be defined according to two parameters: the depth from the user's location; and the angle with the user's gaze direction. Figure 8 Shows the priority of the atlas that the player is to retrieve according to the position of the points represented by the atlas. As Figure 8 shown, the atlas obtained for the minimum depth index (depth 1) and the angular sector corresponding to the user's viewpoint (S0) may be retrieved first. Then, the atlas obtained for the directly higher depth index (depth 2) and the angular sector corresponding to the user's viewpoint (S0), as well as the atlas obtained for the minimum depth index (depth 1) and the angular sectors adjacent to the angular sector corresponding to the user's viewpoints (S1, S - 1), and so on, may be retrieved.

[0141] Of course, the number and size of the depth ranges and angular sectors are not limited to Figure 8 those shown. Specifically, the size of the depth ranges of the respective angular sectors may be different for each depth range and each sector.

[0142] Similar to the depth - based delivery method, a manifest may describe the organization of different streams. To benefit from sectorization and provide patch positions to the client, the manifest may also include patch positions within the 3D scene, for example, represented in spherical coordinates. Since a patch may represent a volume rather than a point (which has a texture component and a depth component), its position may be represented as a single coordinate indicating the center point of the patch (e.g., the centroid), or as a set of spherical coordinates of the volume elements that contain the patch (r, θ, ) / size (dr, r dθ, ).

[0143] Finally, it should be noted that both the depth - based patch delivery method and the viewport - based patch delivery method can be used in combination.

[0144] The following description will discuss the device.

[0145] Figure 9 Schematically shows an example of a device for generating patches representing a 3D scene, a device for transmitting a representation of a 3D scene, or a device for rendering a 3D scene according to at least one embodiment of the present disclosure.

[0146] The apparatus for generating patches representing a 3D scene may include, for example, a non-volatile memory 93G (e.g., read-only memory (ROM) or hard disk), a volatile memory 91G (e.g., random access memory or RAM), and at least one processor 92G. The non-volatile memory 93G may be a non-transitory computer-readable carrier medium. It may store executable program code instructions that are executed by the processor 92G to enable the methods described above in their various embodiments.

[0147] Specifically, the processor 92G is configured to perform the following processes: obtain a first view of the 3D scene from a first viewpoint; generate at least one first patch from the first view, the at least one first patch including a texture component and a depth component; and obtain at least one second view of the 3D scene from at least one second viewpoint. For at least one of the second views among the second views, the processor 92G is further configured to perform the following processes: identify at least one point in the second view that is not visible in another view of the 3D scene; determine a depth range to which the at least one point belongs; for at least one of m angular sectors and for at least one of n depth ranges, where at least one of m or n is greater than or equal to 2, generate at least one second patch from the second view for points belonging to the angular sector and the depth range, the at least one second patch including a texture component and a depth component; and construct at least one atlas by packing together at least one of the second patches generated for points belonging to the same angular sector and the same depth range.

[0148] Upon initialization, the foregoing program code instructions may be transferred from the non-volatile memory 93G to the volatile memory 91G for execution by the processor 92G. The volatile memory 91G may also include registers for storing variables and parameters required for this execution.

[0149] The apparatus for transmitting a representation of a 3D scene may include, for example, a non-volatile memory 93T (e.g., read-only memory (ROM) or hard disk), a volatile memory 91T (e.g., random access memory or RAM), and at least one processor 92T. The non-volatile memory 93T may be a non-transitory computer-readable carrier medium. It may store executable program code instructions that are executed by the processor 92T to enable the methods described above in their various embodiments.

[0150] Specifically, the processor 92T can be configured to perform the following process: obtain at least one first patch generated from a first view of a 3D scene, the at least one first patch including a texture component and a depth component; obtain at least one atlas generated from at least one second view of the 3D scene, the at least one atlas being constructed by packing together at least one second patch generated for one of the second views for at least one point that is invisible in another view of the 3D scene and belongs to the same corner sector among m corner sectors and the same depth range among n depth ranges, at least one of m or n being greater than or equal to 2, the at least one second patch including a texture component and a depth component; generate a first subset of m' streams from the one or more first patches and a second subset of m'×n' streams from the one or more atlases according to at least one terminal-based delivery criterion, where m'≤m and n'≤n, each stream including a stream for transmitting the texture component and a stream for transmitting the depth component, and transmit the first stream subset and the second stream subset to the terminal.

[0151] Upon initialization, the foregoing program code instructions can be transferred from the non-volatile memory 93T to the volatile memory 91T for execution by the processor 92T. The volatile memory 91T can also include registers for storing variables and parameters required for this execution.

[0152] A device for rendering a 3D scene can include, for example, a non-volatile memory 93R (e.g., read-only memory (ROM) or hard disk), a volatile memory 91R (e.g., random access memory or RAM), and at least one processor 92R. The non-volatile memory 93R can be a non-transitory computer-readable carrier medium. It can store executable program code instructions that are executed by the processor 92R to enable the methods described above in their various embodiments.

[0153] Specifically, the processor 92R may be configured to receive a first subset of streams and a second subset of streams generated according to at least one terminal-based delivery criterion, the first subset including m' pairs of streams generated from at least one first patch and the second subset including m'×n' pairs of streams generated from at least one atlas, each pair of streams including a stream for transmitting a texture component and a stream for transmitting a depth component, the at least one first patch being generated from a first view of the 3D scene and including a texture component and a depth component, the at least one atlas being generated from at least one second view of the 3D scene and constructed by packing together at least one second patch generated for one of the second views for at least one point that is invisible in another view of the 3D scene and belongs to the same corner sector among m corner sectors and the same depth range among n depth ranges, at least one of m or n being greater than or equal to 2, the at least one second patch including a texture component and a depth component, where m'≤m and n'≤n. The processor 92R may be further configured to construct a representation of the 3D scene from the first subset of streams and the second subset of streams.

[0154] Upon initialization, the foregoing program code instructions may be transferred from the non-volatile memory 93R to the volatile memory 91R for execution by the processor 92R. The volatile memory 91R may also include registers for storing variables and parameters required for such execution.

[0155] The method according to at least one embodiment of the present disclosure may be equally well implemented in one of the following ways: (1) executing a set of program code instructions executed by a reprogrammable computing machine such as a PC-type device, a DSP (Digital Signal Processor), or a microcontroller. Such program code instructions may be stored in a separable (e.g., floppy disk, CD-ROM, or DVD-ROM) or inseparable non-transitory computer-readable carrier medium; or (2) a dedicated machine or component such as an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or any dedicated hardware component.

[0156] In other words, the present disclosure is not limited to a purely software-based implementation in the form of computer program instructions, but the present disclosure may also be implemented in hardware form or in any form combining a hardware part and a software part.

[0157] The flowcharts and / or block diagrams in the figures illustrate the possible configurations, operations, and functions of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function.

[0158] It should also be noted that in some alternative specific implementations, the functions labeled in the blocks may not occur in the order labeled in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order, or the blocks may be executed in alternative orders depending on the functions involved. It should also be noted that each block of the block diagram and / or flowchart illustration, and combinations of blocks in the block diagram and / or flowchart illustration, can be implemented by a dedicated system based on hardware that performs the specified functions or actions, or a combination of dedicated hardware and computer instructions. Although not explicitly described, the embodiments of the present invention can be employed in any combination or sub-combination.

Claims

1. A method for rendering a 3D scene on a terminal, the method comprising: Receive a manifest at the terminal, where the manifest includes: References to at least one first data stream available on the server, the at least one first data stream including central patch content of the 3D scene, the central patch content including the part visible from a central viewpoint in the scene; References to a plurality of second data streams available on the server, the plurality of second data streams including parallax patch content of the 3D scene, where the parallax patch content includes the part visible from a second non-central viewpoint in the 3D scene; And The association of each second data stream in the second data streams with a depth range; Request at least one available first data stream from the server; Request a subset of available second data streams selected at least based on the depth range associated with at least one available second data stream from the server; And Render the 3D scene using the central patch content from the requested first data stream and the parallax patch content of the selected subset from the requested available second data streams.

2. The method according to claim 1, the method further comprising: Receive the requested first data stream and the subset of the requested available second data streams at the terminal.

3. The method according to claim 1, the method further comprising: Specify the association of each second data stream in the second data streams with an angular sector.

4. The method according to claim 3, wherein the associated angular sector includes an angular range within the view of the user of the terminal.

5. The method according to claim 3, wherein the subset of the available second data streams is selected based on the depth range and the angular sector associated with the available second data stream.

6. The method according to claim 1, wherein the subset of the available second data streams is selected based on at least one of the bandwidth available on the communication channel between the terminal and the server and the capabilities of the terminal.

7. The method according to claim 1, wherein the subset of the available second data streams is selected based on the priority of the depth range, wherein the priority of the second data stream associated with a closer depth range is higher than the priority of the second data stream associated with a farther depth range.

8. The method according to claim 1, wherein the manifest further specifies the association of each second data stream in the second data streams with a time interval.

9. The method according to claim 1, wherein the terminal is a head-mounted display.

10. The method according to claim 1, wherein the parallax patch content includes a depth patch and a texture patch.

11. The method according to claim 1, wherein each second data stream carries the parallax patch content in the form of a patch atlas.

12. A method for generating a 3D scene at a server, the method comprising: Transmit a manifest, where the manifest includes; References to at least one first data stream, the at least one first data stream including central patch content of the 3D scene, the central patch content including the part visible from a central viewpoint in the scene; References to a plurality of second data streams, the plurality of second data streams including parallax patch content of the 3D scene, where the parallax patch content includes the part visible from a second non-central viewpoint in the 3D scene; and The association of each second data stream in the second data streams with a depth range; Receive a request for at least one available first data stream; Receive a request for a subset of available second data streams selected at least based on the depth range associated with at least one available second data stream; And Transmit the at least one available first data stream and the subset of available second data streams for rendering the 3D scene.

13. The method according to claim 12, wherein the method further comprises: Specify the association of each second data stream in the second data streams with an angular sector in the manifest.

14. The method according to claim 13, wherein the associated angular sector includes an angular range within the view of the user of the terminal.

15. The method according to claim 12, wherein the list further specifies the association of each second data stream in the second data stream with a time interval.

16. The method according to claim 12, wherein the parallax patch content includes a depth patch and a texture patch.

17. The method according to claim 12, wherein each second data stream carries the parallax patch content in the form of a patch atlas.

18. A terminal, the terminal comprising: Receiver; Transmitter; And Processor; Where the receiver is configured to receive a manifest, where the manifest includes: References to at least one first data stream available on the server, the at least one first data stream including central patch content of the 3D scene, the central patch content including the part visible from a central viewpoint in the 3D scene; References to a plurality of second data streams available on the server, the plurality of second data streams including parallax patch content of the 3D scene, where the parallax patch content includes the part visible from a second non-central viewpoint in the 3D scene; and The association of each second data stream in the second data streams with a depth range; Where the transmitter is configured to request at least one available first data stream from the server; wherein the transmitter is further configured to request from the server a subset of the available second data streams selected at least based on a depth range associated with at least one available second data stream; and wherein the processor is configured to render the 3D scene using the central patch content from the requested first data stream and the disparity patch content from the selected subset of the requested available second data streams.

Citation Information

Patent Citations

  • Methods and devices for encoding and decoding three degrees of freedom and volumetric compatible video stream

    WO2019055389A1

  • Intermediate view synthesis and multi-view data signal extraction

    CN102239506A

  • Method and device for generating disparity vector

    CN103916652A