Supporting Multiview Video Motion with Disocclusion Atlases

By employing a disocclusion atlas to transmit disocclusion data in multi-view video operations, the inefficiencies in existing technologies are addressed, resulting in reduced data requirements and improved visual quality in volumetric video streaming.

JP7679014B2Active Publication Date: 2025-05-19DOLBY LABORATORIES LICENSING CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023119205
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-16
Filing Date
2023-07-21
Publication Date
2025-05-19
Estimated Expiration
2041-06-16

AI Technical Summary

Technical Problem

Existing multi-view video technologies face challenges in efficiently encoding and transmitting volumetric video due to large gaps or holes in synthesized views, leading to perceivable visual artifacts.

Method used

The use of a disocclusion atlas to support multi-view video operations by transmitting a small amount of disocclusion data, which includes texture and depth information for occluded regions, thereby reducing redundancy and improving video streaming efficiency.

Benefits of technology

This approach effectively reduces the amount of data required for video streaming, improves compression efficiency, and minimizes visual artifacts by efficiently filling in disoccluded regions in synthesized views.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679014000001
    Figure 0007679014000001
  • Figure 0007679014000002
    Figure 0007679014000002
  • Figure 0007679014000003
    Figure 0007679014000003
Patent Text Reader

Abstract

To provide support of multi-view video operations with a disocclusion atlas.SOLUTION: Occluded image fragments are sorted in size. The largest image fragment is used to size a quadtree node in a layout mask for a disocclusion atlas used to store the image fragments. The sorted image fragments are stored into the disocclusion atlas using the layout mask such as each image fragment is hosted with a best fit quadtree node in the disocclusion atlas. A video signal may be generated by encoding one or more reference images and the disocclusion atlas storing the image fragments. The image fragments can be used by a recipient device to fill disoccluded image data in disoccluded spatial regions in a display image synthesized from the reference images.SELECTED DRAWING: Figure 4B
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Provisional Application No. 63 / 039,595, filed on June 16, 2020, and European Patent Application No. 20180179.2, each of which is hereby incorporated by reference in its entirety.

[0002] [Technical Field] The present invention generally relates to image coding and rendering, and more particularly to using a disocclusion atlas to support multi - view video operations.

Background Art

[0003] View synthesis is used in applications such as 3 - dimensional (3D) TV, 360 - degree video, volumetric video, virtual reality (VR), and augmented reality (AR). Virtual views are synthesized from existing views using associated depth information. The existing views can be warped or mapped to the depicted 3D world and then back - projected to the target view position.

[0004] As a result, background regions occluded by foreground objects in the existing views can be disoccluded from the target view position in the target view (without available image data from the existing views), thereby creating gaps or holes in the target view. In addition, discontinuities in the depth image(s) can also cause gaps or holes in the synthesized view. As the total number of views to be encoded or transmitted in a video signal is reduced or minimized in an actual video display application, the area of holes in the synthesized view generated from the reduced or minimized number of views becomes relatively large and increases, resulting in easily perceivable visual artifacts.

[0005] The techniques described in this section are techniques that could be pursued, but are not necessarily techniques that have been previously devised or pursued. Accordingly, unless otherwise indicated, no assumption should be made that any of the techniques described in this section are to be regarded as prior art solely for the reason that they are included in this section. Similarly, unless otherwise indicated, no assumption should be made that any problem identified with respect to one or more of the techniques is recognized in any prior art based on this section.

Brief Description of the Drawings

[0006] The present invention is shown by way of example, and not by way of limitation, in the figures of the accompanying drawings, and like reference numerals refer to like elements.

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 2C

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 4C

Figure 5

[0007] Exemplary embodiments are described herein with respect to using a disocclusion atlas to support multi-view video operations. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are not described in detail to avoid obscuring, obfuscating, or rendering the present invention unintelligible.

[0008] Exemplary embodiments are described herein according to the following overview: 1. Overview 2. Volumetric Video 3. Exemplary Video Streaming Server and Client 4. Image Fragments in Disocclusion Data 5. Image Mask for Disocclusion Data 6. Generation of Disocclusion Atlas 7. Temporally Stable Group-Level Layout Mask 8. Exemplary Process Flow 9. Implementation Mechanism - Hardware Overview 10. Equivalents, Extensions, Alternatives, and Others

[0009] 1. Overview This summary presents a basic description of some aspects of exemplary embodiments of the present invention. Note that this summary is not a comprehensive or exhaustive summary of the aspects of the exemplary embodiments. Further note that this summary is not intended to identify particularly important aspects or elements of the exemplary embodiments, nor is it intended to define the scope of the exemplary embodiments or generally the present invention. This summary merely presents some concepts related to the exemplary embodiments in a simplified form and should be understood as a mere conceptual prelude to the more detailed description of the exemplary embodiments that follows. Although separate embodiments are described herein, note that additional embodiments can be formed by combining the embodiments and / or any combination of partial embodiments described herein.

[0010] A common approach for transmitting volumetric video is to accompany a wide - field - of - view (often 360 - degree) captured or rendered image with depth from a finite set of view positions (also called "recorded view", "reference view", or "represented view"). Due to the depth values at each pixel, these pixels can typically be re - projected (and z - buffered) to an estimated view between the recorded view positions (or reference views). A single re - projected view image, such as a warped image synthesized from the recorded images at the recorded view positions, will have holes and gaps corresponding to the occluded regions that are not visible from the original viewpoints represented by the recorded images. By adding more surrounding source viewpoints or more recorded view positions, the holes remaining in the re - projected view image can be reduced, but at the expense of a large amount of redundant data (e.g., pixels visible within each of the added multiple recorded views).

[0011] For comparison, the techniques described herein can be used to transmit a relatively small amount of disocclusion data in an atlas representation. The disocclusion data includes only texture and depth information for fragments that are not visible from a single (nearest) reference view position, thereby avoiding redundancy in additional recorded views and significantly reducing the amount of data during video streaming and decoding. Using these techniques, image fragments (e.g., in a combined image such as a rectangle, square, etc.) can be laid out within the combined image so as to leave as little empty space as possible within the combined image.

[0012] Furthermore, the problem of video compression efficiency arising from temporal variations between frames in different successive atlas layouts is effectively addressed by the techniques described herein, and motion prediction (e.g., inter prediction, etc.) can be improved. For example, the layout of successive disocclusion atlases between atlas "I frames" (frames that can be encoded or decoded without motion prediction) can be temporally stabilized to achieve a relatively high compression efficiency.

[0013] In some operating scenarios, one or more video streams corresponding to one or more representative views of a multi-view video (or including image data therefrom) can be transmitted to a recipient video decoder together with or separately from disocclusion data for one or more representative views. The disocclusion data includes texture and / or depth image data for image details that may be hidden or occluded in the representative views within the video stream. Some of the occluded image details depicted by the disocclusion data may be visible in the current view of a viewer (also referred to as a "virtual view" or "target view") adjacent to one or more of the representative views within the video stream.

[0014] As described above, the occlusion data can be packaged or encoded into an occlusion atlas. The occlusion atlas can be used by a video encoder to support encoding a volumetric video signal, including a video stream of a view representation, with multi-depth information such as visible image details at one or more depths and occluded image details at other depths (in some cases for multiple view representations). The occlusion atlas can be used by a recipient video decoder of the video signal to render view-dependent effects such as image details specific to the current view of a viewer adjacent to one or more of the view representations.

[0015] The volumetric video signal can include the occlusion atlas as part of the image metadata to assist a recipient video decoder in using the image data of the view representation in the video stream to render an image specific to the current view of the viewer. The video stream and the image metadata can be encoded with a coding syntax based on a video coding standard or a proprietary standard, including but not limited to the MPEG (Moving Picture Experts Group) video standard, H.264 / AVC (H.264 / Advanced Video Coding), HEVC (High-Efficiency Video Coding), MPEG-I, Dolby's ViX file format, etc. Additionally, optionally or alternatively, the occlusion atlas can be encoded in a sub-stream associated with the video stream including the image data of the view representation and decoded from the sub-stream.

[0016] The recipient video decoder can decode the occlusion data packed in the occlusion atlas within the image metadata (or substream) carried by the volumetric video signal and the image data of the representative view in the video stream encoded in the volumetric video signal. The occlusion data and the image data can be used by the video decoder to fill in holes or gaps when generating or constructing an image of the current view of a viewer adjacent to one or more of the representative views. The current view of the viewer may not coincide with any of the representative views in the video stream, and the image of the current view (or view position) of the viewer can be obtained from the received images of the representative views through an image warping operation. Exemplary image warping and / or compositing operations are described in U.S. Provisional Patent Application No. 62 / 518,187, filed on June 12, 2017, the entire content of which is incorporated herein by reference as if fully set forth herein.

[0017] To fill in holes or gaps in the deformed image, for example, by an efficient lookup operation or an index-based search operation, access and retrieve some or all of the occlusion data in the occlusion atlas, which can provide image details that are occluded in the representative view but are disoccluded in the current view of the viewer. As a result, the viewer can view view-specific image details that are not provided in the image of the representative view encoded in the video stream of the volumetric video signal, according to the current view of the viewer.

[0018] The exemplary embodiments described herein relate to streaming volumetric video. Image fragments that are occluded in one or more reference images depicting a visual scene from one or more reference views and are at least partially disoccluded in non-reference views adjacent to the one or more reference views are sorted by size. The image fragments include a first image fragment that is larger in size than any other image fragment within the image fragment. A layout mask for a disocclusion atlas used to store the image fragments is generated. The layout mask is covered by a quadtree including a first best fit node sized specifically for the first image fragment. The sorted image fragments are stored in descending order in the best fit nodes identified within the layout mask. Each image fragment within the sorted image fragments is stored in respective best fit nodes within the best fit nodes. The best fit nodes include at least one best fit node obtained by repeatedly dividing at least one node within the quadtree covering the layout mask. A volumetric video signal encoded with the one or more reference images is generated. The volumetric video signal is further encoded using the image fragments within the disocclusion atlas. The one or more reference images are used by a recipient device of the volumetric video signal to synthesize a display image in a non-representative view for rendering on an image display. The image fragments within the disocclusion atlas are used by the recipient device to fill in disoccluded image data within disoccluded spatial regions in the display image.

[0019] The exemplary embodiments described herein relate to the rendering of volumetric video. One or more reference images are decoded from a volumetric video signal. Image fragments within a disocclusion atlas are decoded from the volumetric video signal. Based on the one or more reference images, a display image in a non-representative view is synthesized. The image fragments within the disocclusion atlas are used to fill in the disoccluded image data within the disoccluded spatial regions in the display image. The display image is rendered on an image display.

[0020] In some exemplary embodiments, the mechanisms described herein form part of a media processing system including, but not limited to, a cloud-based server, a mobile device, a virtual reality system, an augmented reality system, a head-up display device, a helmet-mounted display device, a CAVE-type system, a wall-sized display, a video game device, a display device, a media player, a media server, a media generation system, a camera system, a home system, a communication device, a video processing system, a video codec system, a studio system, a streaming server, a cloud-based content service system, a handheld device, a gaming console, a television, a cinema display, a laptop computer, a netbook computer, a tablet computer, a cellular radiotelephone, an electronic book reader, a POS (point of sale) terminal, a desktop computer, a computer workstation, a computer server, a computer kiosk, or any of various other types of terminals and media processing units.

[0021] Various modifications to the preferred embodiments and the general principles and features described herein will be readily apparent to those skilled in the art. Accordingly, the present disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.

[0022] 2. Volumetric Video The techniques described herein can be used to provide view-specific video with full parallax for a viewer corresponding to the movement of the viewer's body or head in all six degrees of freedom at most. As used herein, the term "view-specific" video (image) can mean a position-specific and / or orientation-specific video (image) that is generated and / or rendered at least in part based on (or in response to a determination of) the position and / or orientation of the viewer.

[0023] To achieve this, video at a set or subset of different points in space (corresponding to a set or subset of different positions and / or different orientations across a viewing volume in which the viewer can move freely) can be used to generate view-specific images to be rendered to the viewer. The video at these different points in space includes texture video and depth video and can form the reference view (or reference viewpoint) of the volumetric video.

[0024] A virtual view, such as the current view of the viewer for a given position and / or orientation of the viewer (which may not coincide with any of these reference views), can be synthesized from these reference views represented by the volumetric video using image-based rendering techniques.

[0025] As used herein, texture video refers to a sequence of texture images over multiple time points and includes the spatial distribution of pixels each specified by individual color or luminance information such as RGB pixel values, YCbCr pixel values, luma and / or chroma pixel values. Depth video corresponding to the texture video refers to a sequence of depth images over multiple time points and includes the spatial distribution of pixels each specified by the corresponding spatial depth information of the corresponding pixels of the corresponding texture images such as z-axis values, depth values, spatial parallax values, parallax information.

[0026] A occlusion atlas containing occlusion data for one or more reference views represented by one or more video streams within a volumetric video can be used to support encoding multi-depth information for view-dependent effects. For example, image details such as highlight speckles can appear in some but not all views and, when visible, appear differently in different views (e.g., different virtual views such as different reference views, the viewer's current view at different points in time, etc.). The multi-depth information of view-dependent image details that are hidden or occluded in the reference view is included in the occlusion data and is delivered to the recipient video decoder as part of the image metadata, thereby enabling the correct rendering or presentation of view-dependent image details (or effects) to the viewer in response to a detected change in the viewer's position or orientation.

[0027] Additionally, optionally or alternatively, the image metadata can include descriptions of fragments, portions, patches, etc. within the occlusion atlas described herein. The image metadata is delivered from an upstream device to a recipient device as part of a volumetric video and can be used to assist the recipient device in rendering the decoded image data from the video stream and the occlusion atlas.

[0028] 3. Exemplary Video Streaming Server and Client FIG. 1A shows an exemplary upstream device such as a video streaming server 100 comprising a multi-view stream receiver 132, a viewpoint processor 134, a stream composer 136, etc. Some or all of the components of the video streaming server (100) can be implemented by one or more devices, modules, units, etc. in software, hardware, a combination of software and hardware, etc.

[0029] The multi-view stream receiver (132) includes software, hardware, a combination of software and hardware, etc., configured to receive reference textures and / or depth videos (106) of a plurality of reference views directly or indirectly from an external video source.

[0030] The viewpoint processor (134) receives the viewpoint data of the viewer from a video client device operated by the viewer in real time or near real time, and includes software, hardware, a combination of software and hardware, etc., configured to establish / determine the position or orientation of the viewer at a plurality of time points over the time interval / duration of an AR, VR, or volumetric video application. In a video application, the display image derived from the reference texture and / or depth video (106) should be rendered at a plurality of time points within the viewer's viewport so as to be provided on an image display that operates in conjunction with a video client device or the like. The viewer's viewport refers to the size of a window or visible area on the image display.

[0031] The stream composer (136) generates a volumetric video signal 112 (including one or more video streams representing one or more reference views and a disocclusion atlas including a disocclusion atlas of views adjacent to the representation view, but not limited thereto), for example, in real time, from the reference texture and / or depth video (106), based at least in part on the viewpoint data 114 indicating the position or orientation of the viewer received as part of the input from a receiver device or the like, and includes software, hardware, a combination of software and hardware, etc., configured to do so.

[0032] The video streaming server (100) can be used to support AR applications, VR applications, 360-degree video applications, volumetric video applications, real-time video applications, near-real-time video applications, non-real-time omnidirectional video applications, automotive entertainment, helmet-mounted display applications, head-up display applications, games, 2D display applications, 3D display applications, multi-view display applications, etc.

[0033] FIG. 1B shows an exemplary receiver device such as a video client device 150 including a real-time stream receiver 142, a viewpoint tracker 144, a volumetric video renderer 146, an image display 148, etc. Some or all of the components of the video client device (150) can be implemented by one or more devices, modules, units, etc. in software, hardware, a combination of software and hardware, etc.

[0034] The viewpoint tracker (144) operates with one or more viewer position / orientation tracking sensors (e.g., motion sensors, position sensors, eye trackers, etc.) to collect real-time or near-real-time viewpoint data 114 related to the viewer, and is configured to perform operations such as transmitting the viewpoint data (114) or the position / orientation of the viewer determined from the viewpoint data to the video streaming server (100). The viewpoint data (114) can be sampled or measured at a relatively fine time scale (e.g., every millisecond, every 5 milliseconds, etc.). The viewpoint data can be used to establish / determine the position or orientation of the viewer at a given time resolution (e.g., every millisecond, every 5 milliseconds, etc.).

[0035] The real-time stream receiver (142) includes software, hardware, a combination of software and hardware, etc., configured to receive and decode a volumetric video signal (112) (such as in real time).

[0036] The volumetric video renderer (146) includes software, hardware, a combination of software and hardware, etc., configured to perform image warping, image warping, blending (such as blending of multiple deformed images from multiple camera sources), image composition, hole filling, etc. on the image data decoded from the volumetric video (112) to generate a view-specific image corresponding to the predicted or measured position or orientation of the viewer, and output the view-specific image to an image display (148) for rendering.

[0037] As used herein, the video content within the video streams described herein can include, but is not necessarily limited to, any of a visual and auditory program, a movie, a video program, a TV broadcast, a computer game, augmented reality (AR) content, virtual reality (VR) content, in-vehicle entertainment content, etc. Exemplary video decoders can include, but are not necessarily limited to, any of a display device, a computing device having a near-eye display, a head-mounted display (HMD), a mobile device, a wearable display device, a set-top box having a display such as a TV, a video monitor, etc.

[0038] As used herein, a "video streaming server" may refer to one or more upstream devices that prepare video content and stream it to one or more video streaming clients, such as a video decoder, for rendering at least a portion of the video content on one or more displays. The display on which the video content is rendered may be part of one or more video streaming clients or may operate in conjunction with one or more video streaming clients.

[0039] Exemplary video streaming servers may include, but are not necessarily limited to, cloud-based video streaming servers located remotely from the video streaming client(s), local video streaming servers connected to the video streaming client(s) on a local wired or wireless network, VR devices, AR devices, automotive entertainment devices, digital media devices, digital media receivers, set-top boxes, gaming machines (e.g., Xbox), general-purpose personal computers, tablets, dedicated digital media receivers such as Apple TV or Roku boxes, etc.

[0040] 4. Image Fragments in Occlusion Data The occlusion data within the occlusion atlas may include image fragments that are occluded in the representation (reference) view within the volumetric video signal. The image fragments described herein refer to continuous non-convex (or occluded) regions of pixels having per-pixel image texture information (e.g., color, luminance / chrominance values, RGB values, YCbCr values, etc.) and per-pixel depth information. The per-pixel image texture and depth information specified for the image fragments within the occlusion atlas can visually depict hidden or occluded image features / objects / structures in the representation view of the volumetric video signal, but in views adjacent to the representation view, they may become at least partially de-occluded or visible.

[0041] For a given reference view that does not contain holes lacking image texture and depth information, a synthesized image for adjacent views around the representation view can be generated using depth image based rendering (DIBR) and the image texture / depth information available for the reference view. The synthesized image may have holes where image texture information and depth information cannot be obtained from the image texture / depth information available for the reference view. Using the synthesized image, an image mask can be generated to identify the holes within the synthesized image for adjacent views.

[0042] In some operation scenarios, the image mask can be at least partially generated for a given reference view by identifying an image region (or area) within the given reference view that includes a large depth gap between adjacent pixels, as compared to other image regions within the given reference view having a relatively smooth transition of depth between adjacent pixels.

[0043] Image texture information and depth information about image fragments within holes (or non-convex regions of pixels) identified in an image mask can be obtained from spatially different reference views or from temporally different reference views. For example, spatially different reference views that are at the same time as a given reference view but are spatially different from the given reference view include and can provide image texture and depth information about holes in the composite image in adjacent views. These spatially different reference views including the given reference view can collectively form a multi-view image for the same time point.

[0044] Additionally, optionally or alternatively, temporally different reference views that are at a time different from that of a given reference view include and can provide image texture and depth information about holes in the composite image in adjacent views. These temporally different reference views including the given reference view can belong to the same visual scene, the same picture group (GOP), etc.

[0045] Additionally, optionally or alternatively, artificial intelligence (AI) or machine learning (ML) can be trained by training on images and then applied to generate or predict some or all of the image texture and depth information about holes in the composite image in adjacent views.

[0046] Image fragments included in a disocclusion atlas for a given time point can be partitioned into different subsets of image fragments for different reference views. Each subset of image fragments within different subsets can include the (occluded) image fragments within their respective reference views of different reference views.

[0047] The occlusion atlas technique described in this specification can be used to pack these image fragments into a combined image (or "atlas") that covers a minimal total area and has no overlapping fragments (e.g., adaptively, optimally, etc.). Each fragment in the combined image representing the occlusion atlas has a dedicated area (or region) that is not overlapped by other fragments included in the occlusion atlas.

[0048] A volumetric video signal can be generated from a sequence of consecutive multi-view images. The sequence of consecutive multi-view images includes a plurality of multi-view images for a plurality of time points that form a sequence of consecutive time points. Each multi-view image within the plurality of multi-view images includes a plurality of single-view images with respect to a plurality of reference views for each of the plurality of time points within the plurality of time points.

[0049] A sequence of consecutive occlusion atlases can be generated for a sequence of consecutive time points. The sequence of consecutive occlusion atlases includes a plurality of occlusion atlases for a plurality of time points within the sequence of consecutive time points. Each occlusion atlas within the plurality of occlusion atlases includes an image fragment that includes one or more subsets of image fragments for one or more of the reference views within a plurality of reference views, represented by a volumetric video signal, for each of the plurality of time points within the plurality of time points.

[0050] For a partial interval (e.g., one fraction of a second, one second or more) within a time interval (e.g., 30 minutes, one hour or more) covered by a sequence of consecutive time points, the volumetric video signal may be encoded with one or more sub - sequences of a picture group (GOP) for one or more reference views represented by this signal. Each sub - sequence of the GOP within one or more sub - sequences of the GOP includes a sub - sequence of texture images and a sub - sequence of depth images for each reference view within one or more reference views represented by the volumetric video signal.

[0051] Each sub - sequence of the GOP includes one or more GOPs. Each GOP is delimited by an I - frame or starts with a starting I - frame and ends with the frame immediately preceding the next starting I - frame. In some embodiments, the starting I - frame and the next starting I - frame may be the two closest I - frames with no other I - frames between them. In some embodiments, the starting I - frame and the next starting I - frame may be nearby I - frames, but not necessarily the two closest I - frames. The I - frames within a GOP can be decoded without depending on the image data from other frames, while non - I - frames such as B - frames or P - frames within a GOP can be at least partially predicted from other frames within the GOP. The I - frame(s) and / or non - I - frame(s) within a GOP can be generated from temporally stable or temporally similar source / input images. These temporally stable source / input images can facilitate relatively efficient inter - prediction or intra - prediction, and data compression or encoding when generating the I - frame(s) and / or non - I - frame(s) within the GOP.

[0052] For the same sub-interval within the interval covered by a sequence of consecutive time points, the volumetric video signal may be encoded with one or more sub-sequences of a group of disocclusion atlases for one or more reference views represented by the signal. Each sub-sequence of the group of disocclusion atlases within one or more sub-sequences of the group of disocclusion atlases includes a sub-sequence of texture images and a sub-sequence of depth images for holes in views adjacent to each reference view within one or more reference views represented by the volumetric video signal.

[0053] Each subsequence of a group of occlusion atlases includes one or more groups of occlusion atlases. Each group of occlusion atlases is delimited by an atlas I frame, or begins at a start atlas I frame and ends at an atlas frame immediately preceding the next start atlas I frame. In some embodiments, the start atlas I frame and the next start atlas I frame may be the two closest atlas I frames with no other atlas I frames between them. In some embodiments, the start atlas I frame and the next start atlas I frame may be nearby atlas I frames, but need not necessarily be the two closest atlas I frames. Atlas I frames within a group of occlusion atlases can be decoded without depending on occlusion data from other atlas frames, but atlas non-I frames, such as atlas B frames or atlas P frames within a group of occlusion atlases, can be at least partially predicted from other atlas frames within the group of occlusion atlases. The atlas I frame(s) and / or atlas non-I frame(s) within a group of occlusion atlases can be generated from temporally stable or temporally similar occlusion atlases. These temporally stable occlusion atlases can facilitate relatively efficient inter-prediction or intra-prediction, and data compression or encoding, when generating the atlas I frame(s) and / or atlas non-I frame(s) within a group of occlusion atlases.

[0054] 5. Image Mask for Occlusion Data FIG. 2A shows an exemplary texture image in a reference view (e.g., a 360-degree "baseball cover" view, etc.). The texture image includes texture information such as color, luminance / chrominance values, RGB values, YCbCr values, etc. for an array of pixels within the image frame. The texture image may correspond to or be indexed by a point in time within a time interval covered by a sequence of consecutive time points, e.g., as a texture image I-frame or a texture image non-I-frame within a group of pictures (GOP) of pictures or images within a video stream, and may be encoded in a video stream for the reference view.

[0055] FIG. 2B shows an exemplary depth image in the same reference view as the texture image of FIG. 2A (e.g., a 360-degree "baseball cover" view, etc.). The depth image of FIG. 2B includes depth information such as depth values, z-values, spatial disparity values, disparity values, etc. for some or all of the pixels within the array of pixels within the texture image of FIG. 2A. The depth image may correspond to or be indexed by the same point in time within a time interval covered by a sequence of consecutive time points, e.g., as a depth image I-frame or a depth image non-I-frame within a group of pictures (GOP) of depth images of pictures or images within a video stream, and may be encoded in a video stream for the reference view.

[0056] FIG. 2C shows an exemplary image mask that can be a bit mask having an array of bits. Indicators or bits (e.g., 1-1, etc.) in the array of bits within the image mask can correspond to respective pixels within the array of pixels represented by the texture image of FIG. 2A and / or the depth image of FIG. 2B. Each indicator or bit within the image mask indicates or can specify whether a disocclusion data portion such as a disoccluded pixel texture value (e.g., color, luminance / chrominance value, RGB value, YCbCr value, etc.) and / or a disoccluded pixel depth value (e.g., depth value, z value, spatial disparity value, disparity value, etc.) should be provided in a disocclusion atlas to be used with the texture image of FIG. 2A and / or the depth image of FIG. 2B in image warping and hole filling operations.

[0057] An exemplary hole filling operation is described in U.S. Provisional Patent Application No. 62 / 811,956, filed Apr. 1, 2019, entitled “HOLE FILLING FOR DEPTH IMAGE BASED RENDERING” by Wenhui Jia et al., the entire content of which is incorporated herein by reference as if fully set forth herein.

[0058] Image warping and hole filling operations can be used to generate a composite image for the viewer's current view, which can be a view adjacent to the reference view. The occluded pixel texture values and / or occluded pixel depth values provided in the occlusion atlas depict image details that are occluded in the texture image of FIG. 2A and / or the depth image of FIG. 2B but can become partially visible in views adjacent to the reference view. The occlusion atlas can correspond to or be indexed by the same point in time within the time interval covered by a sequence of consecutive time points, for example, as an atlas I-frame or atlas non-I-frame within a group of occlusion atlases in a video stream or a separate accompanying video stream, and can be encoded in a video stream for the reference view or a separate accompanying video stream.

[0059] An image mask as shown in FIG. 2C does not appear to align with the corresponding texture image of FIG. 2A or the corresponding depth image of FIG. 2B because the mask covers portions of the texture image and / or depth image that are not visible from one or more adjacent views adjacent to the reference view. The purpose of the occlusion atlas generated using the image mask is to provide texture and depth image data to fill holes in a composite view such as the viewer's current view, where the holes are caused by occlusion in the reprojection of the composite view (or a selected "reference" view). In various operating scenarios, the texture and depth data in the occlusion atlas can cover more, less, or the same spatial area as the holes in the composite view.

[0060] In some operation scenarios, the spatial region covered by the disocclusion atlas may include a safety margin, whereby the disocclusion atlas can ensure that the disoccluded texture and depth data in the disocclusion atlas are available to fully fill the holes in the views adjacent to the reference view.

[0061] In some operation scenarios, the spatial region covered by the disocclusion may not include a safety margin, and as a result, the disocclusion atlas may not guarantee that the disoccluded texture and depth data in the disocclusion atlas are available to fully fill the holes in the views adjacent to the reference view. In these operation scenarios, the receiver video decoder may apply a hole-filling algorithm to generate at least a portion of the texture and depth information for at least a portion of the holes in the synthesized view adjacent to or close to the reference view represented in the video stream.

[0062] Additionally, optionally, or alternatively, the masked spatial region covered in the disocclusion atlas may be used to select prominent visual objects from the visual scene depicted in the reference view. For example, the disocclusion atlas may not convey or provide to the receiver video decoder the texture or depth information for the spatial region away from the prominent visual object. The spatial region for which the disocclusion atlas conveys or provides the receiver video decoder texture or depth information may indicate to the receiver video decoder that the spatial region contains a prominent visual object.

[0063] 6. Generation of the Disocclusion Atlas FIG. 3A shows an exemplary (output) disocclusion atlas that includes or is packaged with image fragments representing occluded regions for one or more reference views. Image metadata may be generated to indicate which reference view each of these image fragments in the disocclusion atlas corresponds to.

[0064] As an example, a volumetric video signal is generated from a sequence of multi-view images. Each multi-view image in the sequence of multi-view images may include a set of N single-view (input / source) texture images for N reference views and a set of N single-view (input / source) depth images for N reference views at a point in time within the sequence of successive time points.

[0065] View parameters may be received and used to specify or define an injective function that maps image (pixel) coordinates (e.g., pixel position, pixel row and column, etc.) and depth to a coordinate system such as a world (3D) coordinate system. The view parameters may be used to synthesize images in adjacent views, identify holes or regions that may be occluded in a reference view but at least partially disoccluded in an adjacent view, and determine, estimate, or predict disocclusion texture data and disocclusion depth data for these holes or regions for each reference view for a part or all of the reference views.

[0066] For each reference view and single-view texture image and single-view depth image at a given point in time, an image mask such as a bitmask may be generated for the reference view to identify the spatial regions in the disocclusion atlas at the given point in time where disocclusion texture and depth data should be provided, as shown in FIG. 3A.

[0067] FIG. 3B shows an exemplary sequence of successive disocclusion atlases that can be created for a sequence of multi-view images within a received or input multi-view video. The sequence of disocclusion atlases can be encoded into a group of disocclusion atlases. Each such group of disocclusion atlases includes a temporally stable disocclusion atlas and can be encoded into the video stream relatively efficiently.

[0068] FIG. 4A shows an exemplary processing flow for generating a disocclusion atlas as shown in FIG. 3A for multi-view images within a sequence of multi-view images covering a time interval. In some exemplary embodiments, one or more computing devices or components can execute this process flow.

[0069] The multi-view images correspond to or are indexed to points in time within the time interval and include N (source / input) single-view texture images for N reference views and N (source / input) single-view depth images for N reference views. Each single-view texture image within the N single-view texture images corresponds to a respective single-view depth image within the N single-view depth images.

[0070] In block 402, the system described herein (e.g., 100 of FIG. 1A) performs an initialization operation on the disocclusion atlas before the disocclusion atlas is used to store (e.g., copy, stamp, place, etc.) image fragments for spatial regions or holes that may be present in the synthesized / deformed images within views adjacent to the N reference views.

[0071] The initialization operation of block 402 may include the following: (a) receiving or loading N image masks that identify spatial regions or holes within N reference views adjacent to N reference views where texture or depth data may be missing in the composite / deformed image; (b) receiving or loading texture and depth information for the image fragments identified in the N image masks; (c) sorting the image fragments by size into a list of image fragments; and so on.

[0072] Here, "size" refers to a metric for measuring the spatial dimensions of an image fragment. Various metrics can be used to measure the spatial dimensions of an image fragment. For example, the smallest rectangle that completely encloses the image fragment can be determined. The horizontal size (denoted as "xsize"), the vertical size (denoted as "ysize"), a combination of the horizontal size and the vertical size, etc. can be used individually or collectively as the metric(s) for measuring the size of the image fragment.

[0073] In some operation scenarios, the size of the image fragment can be calculated as 64 * max(xsize, ysize) + min(xsize, ysize), where each of xsize and ysize is represented in units of pixels, or in units of the horizontal or vertical dimension of a pixel block of a specific size, such as 2 pixels for a 2×2 pixel block, 4 pixels for a 4×4 pixel block, etc., which can be a non - negative integer power of 2.

[0074] Each of the N loaded image masks corresponds to each of the N reference views within the N reference views. The image mask includes an image mask portion for an image fragment that is occluded in the reference view but becomes at least partially visible in a view adjacent to the reference view. Each image mask portion within the image mask portion of the image mask spatially defines or demarcates each image fragment within the image fragment that is occluded in the reference view to which the image mask corresponds but becomes at least partially visible in a view adjacent to the reference view. For each pixel represented by the image mask, if the pixel belongs to one of the image fragments, the bit indicator is set to true or 1, and if the pixel does not belong to any of the image fragments, it is set to false or 0.

[0075] In some operation scenarios, the disocclusion atlas indicates the spatial arrangement of the image fragments (such as all, for example) of the multi-view image and includes a layout mask used to identify or track the image fragments in which the disocclusion data is stored or maintained in the disocclusion atlas. The layout mask can include an array of pixels arranged within a spatial shape such as a rectangular shape. The image fragments spatially defined or demarcated in the layout mask of the disocclusion atlas are mutually exclusive and do not overlap with each other (such as completely, for example) in the layout mask.

[0076] The initialization operation of block 402 may further include the following: (d) creating a single quadtree root node, which is initialized to the best size that just covers the size of the largest image fragment. The quadtree can be incrementally grown by a factor of two in each dimension as needed to keep the corresponding layout mask as small as possible; (e) linking the largest image fragment to the first node of the quadtree by stamping the image fragment (e.g., the image mask part of the image fragment) into the designated area for the first node in the layout mask of the occlusion atlas; among others. The first node of the quadtree here refers to the first quadtree node among the quadtree nodes at the first level under the root node representing the entire layout mask. Here, "stamp" means copying, transferring, or fitting the image fragment or its image mask part into the layout mask of the occlusion atlas. Here, "quadtree" refers to a tree data structure where each internal node has four child quadtree nodes.

[0077] The quadtree initially includes four nodes of the same spatial shape, such as rectangles of the same size. The spatial shape of the quadtree nodes described herein may have special dimensions with a count of pixels that is a non - negative integer power of two.

[0078] Following stamping the largest image fragment into the layout mask of the occlusion atlas, the largest image fragment is removed from the list of image fragments (sorted by size), and the next quadtree node after the first quadtree node is set as the current quadtree node. The current quadtree node represents an empty or candidate quadtree node (not yet populated by the image fragment or its respective image mask part) to be used next to host the image fragment.

[0079] In block 404, the system determines whether the list of size-sorted image fragments includes any image fragments that still need to be stamped or spatially placed on the layout mask of the disocclusion atlas. In some embodiments, any image fragment that is below a minimum fragment size threshold may be removed from the list or ignored within the list. Exemplary minimum fragment size thresholds can be one of 4 pixels in one or both of the horizontal and vertical dimensions, 6 pixels in one or both of the horizontal and vertical dimensions, and the like.

[0080] In response to determining that the list of (size-sorted) image fragments does not include any image fragment(s) that still need to be stamped or spatially placed on the layout mask of the disocclusion atlas, the processing flow ends.

[0081] Otherwise, in response to determining that the list of (size-sorted) image fragments includes any image fragment(s) that still need to be stamped or spatially placed on the layout mask of the disocclusion atlas, the system selects the next larger image fragment as the current image fragment from the list of (size-sorted) image fragments.

[0082] In block 406, the system determines whether the current quadtree node within the quadtree is large enough to host the current image fragment or the corresponding image mask portion for the current image fragment.

[0083] In response to determining that the current quadtree node within the quadtree is not large enough to host the current image fragment, the processing flow proceeds to block 410.

[0084] Otherwise, in response to determining that the current quadtree node within the quadtree is large enough to host the current image fragment, the processing flow proceeds to block 408.

[0085] At block 408, the system determines whether the current quadtree node is the "best" fitting quadtree node for the current image fragment. The "best" fitting quadtree node refers to a quadtree node that is just large enough to host the image fragment or its image mask portion. In other words, the best "fitting" quadtree node represents the smallest sized quadtree node that completely encloses or hosts the image fragment within the layout mask of the occlusion atlas.

[0086] In response to determining that the current quadtree node is not the "best" fitting quadtree node for the current image fragment, the system subdivides the current quadtree node (e.g., repeatedly, iteratively, recursively, etc.) until the "best" fitting quadtree node is found. The "best" fitting quadtree node is set to be the current quadtree node.

[0087] Once it is determined that the current quadtree node is the "best" fitting quadtree node for the current image fragment, the system stamps or spatially demarcates the current image fragment at the "best" fitting quadtree node.

[0088] Subsequent to stamping the current image fragment onto the occlusion atlas or the layout mask of the current quadtree node, the current image fragment is removed from the list of image fragments (sorted by size), and the next quadtree node after the (removed) current quadtree node is set as the (new or current (present)) current quadtree node.

[0089] In block 410, the system determines whether there is an empty quadtree node or candidate quadtree node available to host the current image fragment somewhere under the root node representing the entire layout mask of the occlusion atlas. If so, the empty quadtree node or candidate quadtree nodes are used (collectively, if more than one node is used, etc.) to host the current image fragment. Thereafter, the process flow proceeds to block 404. Thus, if the current image fragment does not (entirely) fit into any existing (child) quadtree node under the current quadtree node, an attempt can be made to fit the fragment anywhere within the layout mask. Note that in many operating scenarios, the quadtree(s) are merely acceleration data structures designed to make atlas construction faster. The quadtree described herein may not be saved or needed once the layout (or layout mask) is determined. Further, there are no (e.g., absolute, unique, etc.) restrictions imposed by the quadtree described herein on where any image fragment can be placed. Image fragments can overlap multiple quadtree nodes and often do overlap in some operating scenarios. Thus, if the "best fit" method (e.g., to find a single best fit node for a fragment such as the current image fragment) fails, a more exhaustive (and expensive) search can be performed across the entire layout mask to fit the nearby fragment. If successful, all quadtree nodes where the fragments so placed overlap are marked as "occupied" and the process continues. If it fails, the process flow proceeds to block 412 to grow the quadtree. The overall algorithm as shown in FIG. 4A remains efficient and effective because in most cases in many operating scenarios, the best fit quadtree search is successful.Only if no best-fit node for the current image fragment is found, perform a more expensive or comprehensive fallback search, or call to find quadtree nodes that might overlap to host the current image fragment. This can involve searching (e.g., in a search loop) through all empty quadtree nodes or candidate quadtree nodes (not yet occupied by any image fragment) in the entire layout mask of the occlusion atlas.

[0090] In response to determining that none of the remaining empty quadtree nodes or candidate quadtree nodes in the layout mask are large enough to host the current image fragment, the processing flow proceeds to block 412.

[0091] Otherwise, in response to determining that an empty quadtree node or candidate quadtree node in the layout mask is large enough to host the current image fragment, the empty quadtree node or candidate quadtree node is set as the (new) current quadtree node, and the processing flow proceeds to block 408.

[0092] In block 412, the system expands or increases the size of the occlusion atlas or the layout mask of the occlusion atlas by a factor of two (2x) in each of the horizontal and vertical dimensions. The existing quadtree (or old quadtree) before this expansion can be linked or placed in the first quadtree node (e.g., the upper left quadrant of the newly expanded quadtree). The second quadtree node (e.g., the upper right quadrant of the newly expanded quadtree) is set to be the (new) current quadtree node. The processing flow proceeds to block 408.

[0093] The texture values and depth values of each pixel identified in the layout mask of the occlusion atlas as belonging to the image fragment can be stored, cached, or buffered as part of the occlusion atlas together with the layout mask of the occlusion atlas.

[0094] 7. Temporally Stable Group-Level Layout Mask To stabilize consecutive occlusion atlases within a video sequence, the layout masks of the occlusion atlases within a group of consecutive occlusion atlases (which may correspond to texture image GOPs, depth image GOPs, etc. within the video sequence) can be separately combined by an "or" operation to form a group-level layout mask for the group of consecutive occlusion atlases.

[0095] Each layout mask within the layout mask can be of equal size, and the same pixel array is compressed using respective indicators or bits to indicate whether any pixel within the layout mask belongs to the image fragment hosted in each occlusion atlas.

[0096] The group-level layout mask may be of the same size as the (individual) layout masks for the group of consecutive occlusion atlases and includes the same array of pixels as in the case of the (individual) layout masks. To generate the group-level layout mask through a union operation or a separate "OR" operation, the indicator or bit for the pixel at a specific pixel location or index can be set to true or 1 if any of the indicators or bits for the corresponding pixel at the same specific pixel location or index in the (individual) layout masks for the group of consecutive occlusion atlases is true or 1.

[0097] A group-level layout mask or an instance thereof may be repeatedly used for each of a plurality of consecutive time points covered in a group of consecutive disocclusion atlases to host or layout an image fragment represented by a disocclusion atlas for each of the consecutive time points. Pixels having no disocclusion texture and depth information for a given time point may be excluded (e.g., undefined, unoccupied, etc.) in the corresponding instance of the group-level layout mask for that time point (or timestamp). FIG. 3C shows an exemplary group of consecutive disocclusion atlases generated using a common group-level layout mask described herein.

[0098] In some operation scenarios, multiple instances of the same group-level layout mask may be used to generate separate disocclusion atlases. The first disocclusion atlas in a group of consecutive disocclusion atlases may be used with the first instance of the combined group-level layout mask to generate a start atlas I frame, followed by other atlas frames generated from other disocclusion atlases in the group of consecutive disocclusion atlases. The start atlas I frame and the other atlas frames may form a group of consecutive atlas frames delimited by the start atlas I frame and the next start atlas I frame before the end of the group. Using a temporally stable group-level layout mask can facilitate data compression operations, such as applying inter-prediction and / or intra-prediction to find data similarity, and reduce the overall data within a group of consecutive disocclusion atlases to be transmitted to a recipient video decoder. In some implementations, video compression can be improved by more than two times by using the union of layout masks (or bitmasks) over (I-frame) time intervals.

[0099] In some operation scenarios, for each of the individual layout masks for a sequence of occlusion atlases, the texture and depth data for all the pixels identified in each may be included or transmitted without generating a combined group-level layout mask for the sequence of occlusion atlases. Data compression operations that use individual layout masks that are temporally and spatially different from each other may not be as efficient in reducing the amount of data as the data compression operations that use the group-level layout masks described herein.

[0100] In some operation scenarios, in order to layout image fragments on the layout mask of an occlusion atlas, these image fragments may first be adapted to available spatial regions such as empty quadtree nodes or candidate quadtree nodes without being rotated first. In some operation scenarios, in order to increase packing efficiency, the image fragments may first be rotated and then placed in the "best" fitting quadtree node. As a result, a quadtree node that may not have been able to host an image fragment before rotation may become able to host the image fragment after rotation.

[0101] A multi-view image or any single-view image therein may be a 360-degree image. The image data of a 360-degree image (including occlusion data) may be represented by an image frame such as a rectangular frame (e.g., a "baseball cover" view, etc.). As shown in FIGS. 2A and 2B, such an image includes a plurality of image segments that are combined into a rectangular image frame, for example, in a "baseball cover" view. However, the plurality of image segments may be combined into different view shapes, such as a square view shape.

[0102] The occlusion atlas described herein may or may show mask striping within a layout mask to indicate that an image fragment includes texture and depth information that touches the boundaries of image segments. The reason for placing mask striping within the layout mask is to avoid cases where the atlas-hosted image fragment crosses a C 0 (or zero-th order) discontinuity. For example, in a baseball cover representation of a 360-degree image, there is one long horizontal seam in the center, where adjacent pixels on opposite sides of the seam do not correspond to adjacent parts of the (e.g., actual) view of the visual scene. Mask striping can be implemented in an occlusion atlas (es) by zeroing out lines in the input mask along this seam to ensure that the image fragment does not cross this boundary. Thus, an image fragment with mask striping is constrained to be applied to fill holes or gaps on the same side of the line as the image fragment and can be correctly interpreted.

[0103] 8. Exemplary Process Flow Figure 4B shows an exemplary process flow according to an exemplary embodiment of the present invention. In some exemplary embodiments, one or more computing devices or components may execute this process flow. In block 422, the upstream device sizes and sorts image fragments that will be occluded in one or more reference images depicting a visual scene from one or more reference views and at least partially de-occluded in non-reference views adjacent to the one or more reference views. The image fragments include a first image fragment that is larger than any other image fragment within the image fragment.

[0104] In block 424, the upstream device generates a layout mask for a disocclusion atlas that is used to store image fragments. The layout mask is covered by a quadtree that includes a first best-fit node that is specifically sized for a first image fragment. The first fit node is sized to cover (e.g., completely) the first image segment.

[0105] In block 426, the upstream device stores the image fragments sorted in descending order in the best-fit nodes identified within the layout mask. Each image fragment within the sorted image fragments is stored in each respective best-fit node within the best-fit nodes. The best-fit nodes include at least one best-fit node obtained by repeatedly splitting at least one node within the quadtree that covers the layout mask. Each of the best-fit nodes can be identified as the smallest-sized quadtree node for completely covering each of the respective image fragments.

[0106] In block 428, the upstream device generates a volumetric video signal encoded with one or more reference images. The volumetric video signal is further encoded using the image fragments within the disocclusion atlas. The one or more reference images are used by a recipient device of the volumetric video signal to synthesize a display image in a non-representational view for rendering on an image display. The image fragments within the disocclusion atlas are used by the recipient device to fill in the disoccluded image data within the disoccluded spatial regions in the display image.

[0107] In one embodiment, each of the one or more reference images represents one of a 360-degree image, a 180-degree image, a viewport image, an image within a regular spatial shape image frame, or an image within an irregular spatial shape image frame.

[0108] In one embodiment, each of the image fragments includes texture image values and depth image values for a spatial region formed by consecutive pixels occluded in one or more reference views.

[0109] In one embodiment, a set of one or more salient video streams includes a first salient video stream assigned a first saliency rank and a second salient video stream assigned a second saliency rank lower than the first saliency rank, and the second video stream is removed from the set of one or more salient video streams to be transmitted to a video streaming client at a later time in response to determining that the available data rate has been reduced.

[0110] In one embodiment, one or more reference images are included in the multi-view images within a multi-view image group that includes a plurality of consecutive multi-view images for a plurality of consecutive time points, and the disocclusion atlas is included in a disocclusion atlas group that includes a plurality of disocclusion atlases for a plurality of consecutive time points.

[0111] In one embodiment, the layout mask is included in a plurality of individual layout masks generated for a plurality of disocclusion atlases, the group-level layout mask is generated from the plurality of individual layout masks by a union operation, and the disocclusion atlas encoded in the volumetric video signal is represented by the group-level layout mask.

[0112] In one embodiment, the disocclusion atlas group is encoded in the volumetric video signal as a group of atlas frames, and the group of atlas frames starts from an atlas I-frame and ends before different atlas I-frames.

[0113] In one embodiment, the occlusion atlas includes mask striping, and the mask striping indicates that image fragments stored in the occlusion atlas are in contact at one or more boundaries of the image segments.

[0114] In one embodiment, the layout mask is expanded in response to determining that no best-fit node is found within the layout mask of the pre-expanded size.

[0115] In one embodiment, the spatial region for an image fragment is identified by a bitmask, and the image fragments are sorted using the size of the spatial region for the image fragment.

[0116] In one embodiment, the image fragments stored in the occlusion atlas are located in one or more salient regions identified from the visual scene. The salient regions can be the more interesting or more important parts or regions of interest of the visual scene.

[0117] FIG. 4C shows an exemplary process flow according to an exemplary embodiment of the present invention. In some exemplary embodiments, one or more computing devices or components may execute this process flow.

[0118] In block 460, a downstream decoder (e.g., a receiver device or a decoder) receives a volumetric video signal. The volumetric video signal can be encoded / generated in any of the embodiments described above with reference to, for example, FIG. 4B. The volumetric video signal is encoded with one or more reference images and image fragments within a disocclusion atlas. The disocclusion atlas is used to store the image fragments. The image fragments that are occluded in one or more reference images depicting a visual scene from one or more reference views and are at least partially disoccluded in non-reference views adjacent to the one or more reference views are sorted by size (block 422) as described with reference to FIG. 4B.

[0119] In block 462, the downstream device decodes one or more reference images from the volumetric video signal.

[0120] In block 464, the downstream device decodes the image fragments within the disocclusion atlas from the volumetric video signal.

[0121] In block 466, the downstream device synthesizes a display image in a non-representative view based on the one or more reference images.

[0122] In block 468, the downstream device uses the image fragments within the disocclusion atlas to fill in the disoccluded image data within the disoccluded spatial regions in the display image.

[0123] In block 470, the upstream device renders the display image on an image display.

[0124] In one embodiment, each of the image fragments includes texture image values and depth image values for a spatial region formed by contiguous pixels that are occluded in one or more reference views.

[0125] In one embodiment, synthesizing the display image at block 466 includes using texture and depth image values ​​available for the one or more reference views.

[0126] In one embodiment, the disoccluded regions of space in the synthesized display image are identified by determining that texture and depth image values ​​available for the one or more reference views are not available for non-reference views adjacent to the one or more reference views.

[0127] In one embodiment, the image fragments stored in the disocclusion atlas are located in one or more salient regions identified from the visual scene, and the disocclusion atlas does not include texture or depth image values ​​to cover spatial regions away from the one or more salient regions such that the one or more salient regions are identified in the synthesized display image.

[0128] In one embodiment, the volumetric video signal includes image metadata that specifies an injective function that maps each pixel in the image fragment from the pixel location of the pixel in the image frame to a corresponding location in a three-dimensional coordinate system in which the visual scene is represented.

[0129] In various exemplary embodiments, an apparatus, a system, an apparatus, or one or more other computing devices performs any or part of the above-described methods described. In one embodiment, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause the method described herein to be performed.

[0130] It should be noted that although separate embodiments are described herein, any combination of the embodiments and / or partial embodiments described herein may be combined to form further embodiments.

[0131] 9. Implementation Mechanism - Hardware Overview According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing device can be hardwired to execute the techniques, or can include one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) that are permanently programmed to execute the techniques, or can include one or more general-purpose hardware processors programmed to execute the techniques according to program instructions in firmware, memory, other storage devices, or a combination. Such special-purpose computing devices can also combine custom hardwired logic, ASICs, or FPGAs with custom programming. The special-purpose computing device can be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that incorporates hardwired and / or programmed logic to implement the techniques.

[0132] For example, FIG. 5 is a block diagram showing a computer system 500 in which an exemplary embodiment of the present invention can be implemented. The computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled to the bus 502 for processing information. The hardware processor 504 can be, for example, a general-purpose microprocessor.

[0133] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 504. When such instructions are stored on a non-transitory storage medium accessible to processor 504, computer system 500 is customized into a dedicated machine that executes the operations specified by the instructions.

[0134] Computer system 500 further includes read only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504.

[0135] A storage device 510, such as a magnetic disk or optical disk, solid state RAM, etc., is provided and coupled to bus 502 for storing information and instructions.

[0136] Computer system 500 can be coupled via bus 502 to a display 512, such as a liquid crystal display, for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to processor 504 and controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, e.g., a first axis (e.g., x) and a second axis (e.g., y), whereby the device can specify a position within a plane.

[0137] Computer system 500 may implement the techniques described herein using customized hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that programs computer system 500 in combination with a computer system to be or become a special-purpose machine. According to one embodiment, the techniques herein are performed by computer system 500 in response to a processor 504 executing one or more sequences of one or more instructions included in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequence of instructions included in main memory 506 causes processor 504 to perform the process steps described herein. In an alternative embodiment, hardwired circuitry may be used in place of, or in combination with, software instructions.

[0138] As used herein, the term "storage media" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular fashion. Such storage media may include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 510. Volatile media includes dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy (registered trademark) disk, flexible disk, hard disk, solid state drive, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, RAM, PROM, and EPROM, FLASH (registered trademark)-EPROM, NVRAM, any other memory chip or cartridge.

[0139] The memory medium is separate from the transmission medium but can be used together with the transmission medium. The transmission medium is involved in the transfer of information between memory media. For example, the transmission medium includes coaxial cables, copper wires, and optical fibers, including the wires that make up bus 502. The transmission medium can also take the form of acoustic or light waves such as those generated during radio waves and infrared data communication.

[0140] Various forms of media can be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions can first be carried on the magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and transmit the instructions over a telephone line using a modem. A modem local to computer system 500 can receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. The infrared detector can receive the data carried by the infrared signal, and an appropriate circuit can place the data on bus 502. Bus 502 carries the data to main memory 506, and processor 504 fetches and executes the instructions therefrom. The instructions received by main memory 506 can be stored in storage device 510 either before or after execution by processor 504.

[0141] Computer system 500 also includes a communication interface 518 coupled to bus 502. The communication interface 518 provides a two-way data communication coupling to a network link 520 connected to a local network 522. For example, the communication interface 518 can be an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, the communication interface 518 can be a LAN card that provides a data communication connection to a compatible local area network (LAN). A wireless link can also be implemented. In any such implementation, the communication interface 518 transmits and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.

[0142] Network link 520 typically provides data communication to other data devices through one or more networks. For example, network link 520 can provide a connection through local network 522 to a host computer 524 or to a data device operated by an Internet service provider (ISP) 526. The ISP 526 then provides data communication services through a worldwide packet data communication network currently commonly referred to as the "Internet" 528. Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks that carry digital data between computer system 500 and network link 520 and through communication interface 518 are exemplary forms of transmission media.

[0143] The computer system 500 can send messages and receive data including program code through a network (s), network link 520, and communication interface 518. In an example of the Internet, the server 530 can send the requested code for an application program through the Internet 528, ISP 526, local network 522, and communication interface 518.

[0144] The received code can be executed by the processor 504 when it is received and / or stored in the memory device 510 or other non-volatile storage device for later execution.

[0145] 10. Equivalents, Extensions, Substitutions, and Others In the foregoing specification, exemplary embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the only and exclusive indicator of what is the invention and what is intended by the applicant to be the invention is the set of claims issued from this application, including any subsequent amendments, in the specific form in which such claims are issued. Any definitions expressly set forth herein for terms contained in such claims shall apply to the meaning of such terms as used in the claims. Accordingly, no limitations, elements, characteristics, features, advantages or attributes not expressly recited in the claims should ever limit the scope of such claims. Accordingly, the specification and drawings are to be considered in an illustrative rather than a limiting sense.

[0146] Aspects of some embodiments include the following listed Exemplary Embodiments (EEE).

[0147] EEE1. A method, comprising: Sorting, by size, image fragments that will be occluded in one or more reference images depicting a visual scene from one or more reference views and at least partially de-occluded in non-reference views adjacent to the one or more reference views, the image fragments including a first image fragment that is larger than any other image fragment within the image fragment, Generating a layout mask for a disocclusion atlas used to store the image fragments, the layout mask being covered by a quadtree including a first best-fit node sized to cover the first image fragment, the disocclusion atlas being a combined image of the smallest total area including a plurality of non-overlapping image fragments, Storing, in descending order, the image fragments sorted to the best-fit nodes identified within the layout mask, each image fragment within the sorted image fragments being stored to respective best-fit nodes within the best-fit nodes, the best-fit nodes including at least one best-fit node obtained by repeatedly splitting at least one node within the quadtree covering the layout mask, Generating a volumetric video signal encoded with one or more reference images, the volumetric video signal being further encoded using the image fragments within the disocclusion atlas, the one or more reference images being used by a recipient device to synthesize a display image in a non-representative view for rendering on an image display, the image fragments within the disocclusion atlas being used by the recipient device to fill in disoccluded image data within disoccluded spatial regions in the display image, A method comprising.

[0148] For each of the one or more reference images, the method according to EEE1, which represents one of a 360-degree image, a 180-degree image, a viewport image, an image within a regular spatial shape image frame, or an image within an irregular spatial shape image frame.

[0149] For each of the image fragments, the method according to EEE1 or EEE2, which includes texture image values and depth image values for a spatial region formed by consecutive pixels occluded in one or more reference views.

[0150] The one or more reference images are included in multi-view images within a multi-view image group that includes multiple consecutive multi-view images for multiple consecutive time points, and the disocclusion atlas is included in a disocclusion atlas group that includes multiple disocclusion atlases for multiple consecutive time points, according to any one of EEE1 to 3.

[0151] The layout mask is included in a plurality of individual layout masks generated for a plurality of disocclusion atlases, the group-level layout mask is generated from the plurality of individual layout masks by a union operation, and the disocclusion atlas encoded in the volumetric video signal is represented by the group-level layout mask, according to EEE4.

[0152] The disocclusion atlas group is encoded in the volumetric video signal as a group of atlas frames, and the group of atlas frames starts from an atlas I frame and ends before different atlas I frames, according to EEE4 or EEE5.

[0153] EEE7. The occlusion atlas includes mask stripping, and the mask stripping is the method according to any one of EEE1 to 6, indicating that the image fragments stored in the occlusion atlas are in contact with one or more boundaries of the image segment.

[0154] EEE8. The layout mask is extended in response to determining that the best-fit node is not found within the layout mask of the pre-expansion size, according to the method according to any one of EEE1 to 7.

[0155] EEE9. The spatial region for the image fragment is identified by a bitmask, and the image fragments are sorted using the size of the spatial region for the image fragment, according to the method according to any one of EEE1 to 8.

[0156] EEE10. The image fragments stored in the occlusion atlas are located in one or more salient regions identified from the visual scene, according to the method according to any one of EEE1 to 9.

[0157] EEE11. One or more salient video streams include a first salient video stream assigned a first saliency rank and a second salient video stream assigned a second saliency rank lower than the first saliency rank, according to the method according to EEE10.

[0158] EEE12. A method comprising: decoding one or more reference images from the volumetric video signal; decoding image fragments in the occlusion atlas from the volumetric video signal; synthesizing a display image in a non-representative view from the one or more reference images; filling the occluded image data in the occluded spatial region in the display image using the image fragments in the occlusion atlas; A step of rendering a display image on an image display, and A method including the same.

[0159] EEE13. The volumetric video signal includes image metadata specifying an injective function that maps each pixel in an image fragment from its pixel location in the image frame to a corresponding location in a three-dimensional coordinate system in which the visual scene is represented, according to the method described in any one of EEE1 to 12.

[0160] EEE14. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause the method described in any one of EEE1 to 13 to be executed.

[0161] EEE15. A computing device comprising one or more processors and one or more storage media storing a set of instructions that, when executed by one or more processors, cause the method described in any one of EEE1 to 13 to be executed.

Claims

1. 1. A method comprising: - sorting by size image fragments that become occluded in one or more reference images depicting a visual scene from one or more reference views and at least partially disoccluded in non-reference views adjacent to said one or more reference views; generating a layout mask for a disocclusion atlas used to store the image fragments, the disocclusion atlas being a combined image of smallest total area containing multiple non-overlapping image fragments; storing the sorted image fragments in descending order in nodes within the layout mask; generating a volumetric video signal encoded with the one or more reference images, the volumetric video signal being further encoded using the image fragments in the disocclusion atlas; The method includes:

2. 2. The method of claim 1 , wherein each of the one or more reference images represents one of a 360 degree image, a 180 degree image, a viewport image, an image in a regular spatially shaped image frame, or an image in an irregular spatially shaped image frame.

3. The method of claim 1 or 2, wherein each of the image fragments comprises texture and depth image values ​​for a spatial region formed by contiguous occluded pixels in the one or more reference views.

4. 4. The method of claim 1 , wherein the one or more reference images are included in a multiview image in a multiview image group comprising a plurality of consecutive multiview images for a plurality of consecutive time points, and the disocclusion atlas is included in a disocclusion atlas group comprising a plurality of disocclusion atlases for the plurality of consecutive time points.

5. 5. The method of claim 4, wherein the layout mask is included in a plurality of individual layout masks generated for the plurality of disocclusion atlases, a group-level layout mask is generated from the plurality of individual layout masks by a union operation, and the disocclusion atlas encoded into the volumetric video signal is represented in the group-level layout mask.

6. The method of claim 4 or 5, wherein the disocclusion atlas group is encoded into the volumetric video signal as a group of atlas frames, the group of atlas frames starting with an atlas I-frame and ending before a different atlas I-frame.

7. 7. The method of claim 1, wherein the disocclusion atlas includes mask striping, the mask striping indicating that image fragments stored in the disocclusion atlas meet at one or more boundaries of image segments.

8. The method of claim 1 , wherein the layout mask is expanded in response to determining that a best-fit node is not found in the layout mask in a pre-expanded size.

9. The method of claim 1 , wherein spatial regions for the image fragments are identified in a bitmask and the image fragments are sorted using sizes of the spatial regions for the image fragments.

10. The method of claim 1 , wherein the image fragments stored in the disocclusion atlas are located in one or more salient regions identified from the visual scene.

11. 1. A method comprising:

11. A method according to claim 1, further comprising the steps of: receiving a volumetric video signal encoded with one or more reference images and image fragments in a disocclusion atlas; decoding the one or more reference images from the volumetric video signal; decoding the image fragments in the disocclusion atlas from the volumetric video signal; synthesizing a display image at a non-representation view from the one or more reference images; filling disoccluded image data in disoccluded spatial regions in the display image using the image fragments in the disocclusion atlas; rendering the display image on an image display; The method includes:

12. The method of claim 11 , wherein each of the image fragments includes texture and depth image values ​​for a spatial region formed by contiguous occluded pixels in the one or more reference views.

13. The method of claim 11 or 12, wherein synthesizing the display image comprises using texture and depth image values ​​available for the one or more reference views.

14. 14. The method of claim 13, wherein the disoccluded spatial regions within the synthesized display image are identified by determining that the texture image values ​​and the depth image values ​​available for the one or more reference views are not available for the non-reference views adjacent to the one or more reference views.

15. 14. The method of claim 12 or 13, wherein the image fragments stored in the disocclusion atlas are located in one or more salient regions identified from the visual scene, and the disocclusion atlas does not include texture or depth image values ​​to cover spatial regions away from the one or more salient regions such that one or more salient regions are identified in the synthesized display image.

16. 16. A method according to claim 1, wherein the volumetric video signal includes image metadata specifying an injective function that maps each pixel in the image fragment from its pixel location in an image frame to a corresponding location in a three-dimensional coordinate system in which the visual scene is represented.

17. A computer program product causing one or more processors to carry out the method of any one of claims 1 to 16.

18. 17. A computing device comprising one or more processors and one or more storage media storing a set of instructions, the instructions, when executed by the one or more processors, cause the computing device to perform a method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Method and device for processing a depth-map

    US20100215251A1

  • Method and system for encoding a 3D video signal, enclosed 3D video signal, method and system for decoder for a 3D video signal

    WO2009001255A1

  • Method for generating and reconstructing a three-dimensional video stream, based on the use of the occlusion map, and corresponding generating and reconstructing device

    WO2013168091A1

  • Auxiliary data for artifacts –aware view synthesis

    WO2017080420A1