Supporting multi-view video operations with disoccluded atlas

By generating and encoding de-occlusion atlases, the problem of holes and gaps in multi-view videos is solved, achieving efficient data transmission and visual effects, and ensuring that observers see the complete image from different perspectives.

CN115769582BActive Publication Date: 2026-03-03DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180042986.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-16
Filing Date
2021-06-16
Publication Date
2026-03-03
Estimated Expiration
2041-06-16

AI Technical Summary

Technical Problem

Existing technologies suffer from holes and gaps in multi-view videos, leading to visual artifacts and significant data redundancy, making efficient encoding and transmission difficult.

Method used

By employing de-occlusion atlas technology, the hole areas in multi-view videos are filled by generating and encoding de-occlusion atlases. Virtual views are synthesized using image masks and depth information, reducing redundant data transmission.

Benefits of technology

It effectively reduces the amount of data in the video stream, improves encoding efficiency, ensures that observers can see complete image details from different perspectives, and reduces the occurrence of visual artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115769582B_ABST
    Figure CN115769582B_ABST
Patent Text Reader

Abstract

Occluded image fragments are sorted by size. The largest image fragment is used to determine the size of a quadtree node in the layout mask of the de-occlusion atlas used to store the image fragments. These sorted image fragments are stored in the de-occlusion atlas using this layout mask, for example, each image fragment is carried on a best-fit quadtree node in the de-occlusion atlas. A video signal can be generated by encoding one or more reference images and a de-occlusion atlas storing these image fragments. These image fragments can be used by a receiving device to fill de-occluded image data in de-occluded spatial regions of a display image synthesized from these reference images.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 039,595 and European Patent Application No. 20180179.2, both filed on June 16, 2020, each of which is incorporated herein by reference in its entirety. Technical Field

[0003] This invention relates generally to image encoding and rendering, and more specifically to using deocclusion atlases to support multi-view video operations. Background Technology

[0004] View compositing is used in applications such as 3D television, 360-degree video, volumetric video, virtual reality (VR), and augmented reality (AR). A virtual view is synthesized from an existing view and its associated depth information. An existing view can be distorted or mapped to the depicted 3D world and then back-projected onto the target view's location.

[0005] Therefore, background areas occluded by foreground objects in an existing view may be deoccluded in a target view from the target view location (where no image data from the existing view is available), thus creating gaps or holes in the target view. Additionally, discontinuities in (multiple) depth images can also lead to gaps or holes in the composite view. As the total number of views to be encoded or transmitted in a video signal is reduced or minimized in real-world video display applications, the areas of holes in the composite view generated from the reduced or minimized number of views become relatively large and numerous, resulting in easily noticeable visual artifacts.

[0006] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise indicated, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, any issues identified with respect to one or more methods should not be considered to have been identified in any prior art based on this section. Attached Figure Description

[0007] The invention is illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings:

[0008] Figure 1A and Figure 1B The illustration shows an example volumetric video stream;

[0009] Figure 2A and Figure 2B The illustration shows example texture and depth images;

[0010] Figure 2C The illustration shows an example image mask used to identify spatial regions that are occluded in a reference view and become at least partially unoccluded in adjacent views;

[0011] Figure 3A The illustration shows an example of an unmasked image set; Figure 3B The illustration shows an example of a continuous deocclusion atlas sequence; Figure 3C The illustration shows an example of a continuous demasking atlas group generated using a common group-level layout mask;

[0012] Figures 4A to 4C The example processing flow is illustrated; and

[0013] Figure 5 An example hardware platform on which a computer or computing device as described herein can be implemented is illustrated. Detailed Implementation

[0014] This document describes example embodiments involving the use of deocclusion atlases to support multi-view video operations. In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent that the invention can be practiced without these specific details. In other instances, well-known structures and devices have not been described in detail to avoid unnecessarily obscuring, obscuring, or obscuring the invention.

[0015] This document describes example embodiments based on the following summary:

[0016] 1. General Overview

[0017] 2. Video size

[0018] 3. Example video streaming server and client

[0019] 4. Remove occlusion from image fragments in the data

[0020] 5. Image masks used for removing occluded data

[0021] 6. Generation of unoccluded image atlases

[0022] 7. Time-stable group-level layout mask

[0023] 8. Example Processing Flow

[0024] 9. Implementation Mechanism – Hardware Overview

[0025] 10. Equivalents, extensions, substitutes, and others

[0026] 1. General Overview

[0027] This overview provides a basic description of some aspects of exemplary embodiments of the invention. It should be noted that this overview is not a broad or exhaustive summary of all aspects of the exemplary embodiments. Furthermore, it should be understood that this overview is not intended to identify any particularly important aspects or elements of the exemplary embodiments, nor is it intended to be construed as particularly depicting any scope of the exemplary embodiments, nor is it a general description of the invention. This overview only introduces some concepts related to the exemplary embodiments in a condensed and simplified format and should be understood as merely a conceptual prelude to a more detailed description of the following exemplary embodiments. Note that although individual embodiments are discussed herein, any combination of the embodiments and / or portions of the embodiments discussed herein can be combined to form further embodiments.

[0028] A common method for transmitting volumetric video is to accompany it with a wide field-of-view (typically 360-degree) image with depth, captured or rendered from a finite set of view locations (also known as “recorded views,” “reference views,” or “represented views”). The depth value at each pixel allows those pixels to be reprojected (and z-buffered) into a hypothetical view, typically located between the recorded view locations (or reference views). A single reprojected view image (such as a distorted image synthesized from a recorded image at a recorded view location) will have holes and gaps corresponding to de-occluded areas not visible from the original viewpoint represented in the recorded image. By adding more surrounding source views or more recorded view locations, fewer holes may remain in the reprojected view image, but at the cost of a large amount of redundant data (e.g., pixels visible in each of the multiple added recorded views).

[0029] By comparison, the techniques described in this paper can be used to transmit a relatively small amount of deocclusion data in atlas representations. The deocclusion data includes only the texture and depth information of those fragments not visible from a single (nearest) reference view location, thus avoiding redundancy in additional recorded views and significantly reducing the amount of data in video streaming and decoding. These techniques can be used to layout image fragments in combined images (e.g., rectangular, square, etc.) to leave as little blank space as possible in the combined image.

[0030] Furthermore, the video compression efficiency problem caused by frame-to-frame temporal variations in different successive atlas layouts can be effectively addressed by techniques such as those described herein for enhancing motion prediction (e.g., inter-frame prediction, etc.). For example, the layout of successive deocclusion atlases between atlas “I-frames” (frames that can be encoded or decoded without motion prediction) can be temporally stabilized to achieve a relatively efficient compression ratio.

[0031] In some operational scenarios, one or more video streams corresponding to one or more represented views of a multi-view video (or including image data from those represented views) may be sent together or separately to a receiving video decoder along with deocclusion data for one or more represented views. The deocclusion data includes texture and / or depth image data of image details that may be hidden or occluded in the represented views of the video stream. Some of the occluded image details depicted by the deocclusion data may become visible in the current view (also referred to as a "virtual view" or "target view") of one or more adjacent observers in the represented views of the video stream.

[0032] As described, deocclusion data can be encapsulated or encoded in a deocclusion atlas. The deocclusion atlas can be used by a video encoder to support the encoding of multi-depth information (for potentially multiple represented views) such as visible image details at one or more depths and occluded image details at other depths into a volumetric video signal that includes the video stream of the represented views. The deocclusion atlas can be used by a video decoder at the receiver of the video signal to render view-dependent effects, such as image details specific to the current view of one or more neighboring observers in the represented views.

[0033] Volumetric video signals may include a deocclusion atlas as part of image metadata to assist the receiver's video decoder in rendering an image specific to the observer's current view using image data representing the view in the video stream. The video stream and image metadata may be encoded using coding syntax based on video coding standards or proprietary specifications, including but not limited to the Moving Picture Experts Group (MPEG) video standard, H.264 / Advanced Video Coding (H.264 / AVC), High Efficiency Video Coding (HEVC), MPEG-I, Dolby ViX file format, etc. Alternatively, the deocclusion atlas may be encoded in and decoded from a substream of the video stream accompanying the video stream, which includes image data representing the view.

[0034] The receiving video decoder can decode de-occlusion data in a de-occlusion atlas packaged in image metadata (or substream) carried by a volumetric video signal, and can decode image data of the represented view encoded in the video stream of the volumetric video signal. The de-occlusion data and image data can be used by the video decoder to fill holes or gaps when generating or constructing images for the current view of one or more adjacent observers in the represented view. An observer's current view may not coincide with any represented view in the video stream, and an image of the observer's current view (or view position) can be obtained from the received image of the represented view through an image warping operation. Example image warping and / or compositing operations are described in U.S. Provisional Patent Application No. 62 / 518,187, filed June 12, 2017, the entire contents of which are fully set forth herein and are incorporated herein by reference.

[0035] To fill holes or gaps in a distorted image, some or all of the de-occlusion data in the de-occlusion atlas can be accessed and retrieved, for example, through efficient lookup operations or indexed search operations, to provide image details that are occluded in the represented view but de-occluded in the observer's current view. Therefore, depending on the observer's current view, the observer can see view-specific image details not provided in the image of the represented view encoded in the video stream of the volumetric video signal.

[0036] The example embodiments described herein relate to streaming volumetric video. Image segments occluded in one or more reference images depicting a visual scene from one or more reference views and at least partially de-occluded in non-reference views adjacent to the one or more reference views are sorted by size. Each image segment includes a first image segment whose size is not smaller than any other image segment in the image segment set. A layout mask is generated for a de-occlusion atlas used to store the image segments. The layout mask is overlaid with a quadtree including a first best-fit node specifically sized for the first image segment. The sorted image segments are stored in descending order into the best-fit nodes identified in the layout mask. Each image segment in the sorted image segments is stored in a corresponding best-fit node within the best-fit nodes. The best-fit nodes include at least one best-fit node obtained by iteratively partitioning at least one node in the quadtree overlaying the layout mask. A volumetric video signal encoded using the one or more reference images is generated. This volumetric video signal is further encoded using the image segments in the de-occlusion atlas. The one or more reference images are used by a receiving device of the volumetric video signal to synthesize a display image in a non-representation view for rendering on an image display. Image fragments in the demasking atlas are used by the receiving device to fill the demasking image data in the demasking space area of ​​the displayed image.

[0037] The example embodiments described herein relate to rendering volumetric video. One or more reference images are decoded from the volumetric video signal. Image segments from a deocclusion atlas are decoded from the volumetric video signal. A display image in a non-representational view is synthesized based on the one or more reference images. The image segments from the deocclusion atlas are used to fill deocclusion image data in deocclusion spatial regions of the display image. The display image is then rendered on an image display.

[0038] In some example embodiments, mechanisms such as those described herein form part of a media processing system, including but not limited to any of the following: cloud-based servers, mobile devices, virtual reality systems, augmented reality systems, head-up displays, head-mounted displays, cAVE systems, wall-mounted displays, video game devices, display devices, media players, media servers, media production systems, camera systems, home-based systems, communication devices, video processing systems, video codec systems, studio systems, streaming media servers, cloud-based content service systems, handheld devices, game consoles, televisions, cinema displays, laptops, notebook computers, tablets, cellular phones, e-book readers, point-of-sale terminals, desktop computers, computer workstations, computer servers, computer kiosks, or various other types of terminals and media processing units.

[0039] Various modifications to the preferred embodiments, general principles, and features described herein will be readily apparent to those skilled in the art. Therefore, this disclosure is not intended to be limited to the illustrated embodiments, but rather to be accorded the maximum scope consistent with the principles and features described herein.

[0040] 2. Video size

[0041] The techniques described herein can be used to provide an observer with view-specific video with full parallax in response to movement of the observer's body or head (up to all six degrees of freedom). As used herein, the term "view-specific" video (image) may mean a position-specific and / or orientation-specific video (image) generated and / or rendered at least in part based on the observer's position and / or orientation (or in response to determining the observer's position and / or orientation).

[0042] To achieve this, a view-specific image rendered to the observer can be generated using video from sets or subsets of different points in space (corresponding to sets or subsets of different locations and / or orientations across the observation volume in which the observer moves freely). The video from these different points in space can include texture video as well as depth video, and together form a reference view (or reference viewpoint) of the volumetric video.

[0043] Virtual views can be synthesized from these reference views represented in volumetric video using image-based rendering techniques. These virtual views could be, for example, the observer's current view for a given position and / or orientation (which may not coincide with any of these reference views).

[0044] As used in this article, texture video refers to a sequence of texture images at multiple time points, including spatial pixel distribution. Each pixel is assigned individual color or luminance information, such as RGB pixel values, YCbCr pixel values, and luminance and / or chrominance pixel values. In contrast, depth video refers to a sequence of depth images at multiple time points, including spatial pixel distribution. Each pixel is assigned spatial depth information of the corresponding pixel in the texture image, such as z-axis values, depth values, spatial difference values, and parallax information.

[0045] Deocclusion atlases, including deocclusion data representing one or more reference views in one or more video streams within a volumetric video, can be used to support the encoding of multi-depth information about view-dependent effects. For example, image details such as speckle may appear in some, but not all, views, and may appear differently in different views when visible (e.g., different reference views, different virtual views (e.g., the observer's current view at different points in time)). Multi-depth information about view-dependent image details that are hidden or occluded in the reference views can be included in the deocclusion data and passed as part of the image metadata to the receiving video decoder, enabling the view-dependent image details (or effects) to be correctly rendered or presented to the observer in response to detected changes in the observer's position or orientation.

[0046] Alternatively, image metadata may include descriptions of segments, sections, patches, etc., in the deocclusion atlas as described herein. Image metadata may be passed from the upstream device to the receiving device as part of the volumetric video and may be used to assist the receiving device in rendering image data decoded from the video stream and the deocclusion atlas.

[0047] 3. Example video streaming server and client

[0048] Figure 1A The illustration shows an example upstream device such as a video streaming server 100, which includes a multi-view stream receiver 132, a viewpoint processor 134, a stream synthesizer 136, etc. Some or all of the components of the video streaming server (100) may be implemented by one or more devices, modules, units, etc., in a software, hardware, or a combination of software and hardware manner.

[0049] The multi-view stream receiver (132) includes software, hardware, or a combination of software and hardware configured to receive reference textures and / or depth video (106) of multiple reference views directly or indirectly from an external video source.

[0050] The viewpoint processor (134) includes software, hardware, or a combination of software and hardware configured to: receive viewpoint data from a video client device operated by the observer in real-time or near real-time; establish / determine the position or orientation of the observer at multiple points in time within a time interval / duration in an AR, VR, or volumetric video application. In the video application, a display image obtained from a reference texture and / or depth video (106) is rendered at multiple points in time within the observer's viewport, which is provided, for example, in conjunction with an image display operating in conjunction with the video client device; and so on. The observer's viewport refers to the size of a window or visible area on the image display.

[0051] The stream synthesizer (136) includes software, hardware, a combination of software and hardware, etc., configured to perform the following: generating (e.g., in real time, etc.) volumetric video signals 112 (including, but not limited to, one or more video streams representing one or more reference views and a deocclusion atlas containing deocclusion atlases of views adjacent to the represented views) from reference textures and / or depth video (106) based at least in part on viewpoint data 114 received as part of input from a receiving device that indicates the position or orientation of an observer.

[0052] The video streaming server (100) can be used to support AR applications, VR applications, 360-degree video applications, volumetric video applications, real-time video applications, near real-time video applications, non-real-time omnidirectional video applications, automotive entertainment, helmet-mounted display applications, head-up display applications, games, 2D display applications, 3D display applications, multi-view display applications, etc.

[0053] Figure 1B The illustration shows an example receiver device such as video client device 150, which includes a real-time stream receiver 142, a viewpoint tracker 144, a volumetric video renderer 146, an image display 148, etc. Some or all of the components of the video client device (150) may be implemented by one or more devices, modules, units, etc., in a software, hardware, or a combination of software and hardware manner.

[0054] The viewpoint tracker (144) includes software, hardware, a combination of software and hardware, etc., configured to: operate in conjunction with one or more observer position / orientation tracking sensors (e.g., motion sensors, position sensors, eye trackers, etc.) to collect real-time or near-real-time viewpoint data 114 related to the observer; transmit the viewpoint data (114) or the observer's position / orientation determined from the viewpoint data to a video streaming server (100); etc. The viewpoint data (114) can be sampled or measured at relatively fine time scales (e.g., every millisecond, every five milliseconds, etc.). The viewpoint data can be used to establish / determine the observer's position or orientation at a given time resolution (e.g., every millisecond, every five milliseconds, etc.).

[0055] The real-time streaming receiver (142) includes software, hardware, or a combination of software and hardware configured to receive and decode (e.g., in real time) volumetric video signals (112).

[0056] The volumetric video renderer (146) includes software, hardware, a combination of software and hardware, etc., configured to perform image warping, image distortion, blending (e.g., blending multiple warped images from multiple camera sources, etc.), image compositing, hole filling, etc. on image data decoded from the volumetric video (112) to generate a view-specific image corresponding to the observer's predicted or measured position or orientation; output the view-specific image to an image display (148) for rendering; etc.

[0057] As used herein, the video content in the video stream described herein may include, but is not limited to, any of the following: audiovisual programs, movies, video programs, television broadcasts, computer games, augmented reality (AR) content, virtual reality (VR) content, in-vehicle entertainment content, etc. Example video decoders may include, but are not limited to, any of the following: display devices, computing devices with near-eye displays, head-mounted displays (HMDs), mobile devices, wearable display devices, set-top boxes with displays such as televisions, video monitors, etc.

[0058] As used herein, a “video streaming server” can refer to one or more upstream devices that prepare video content and stream that video content to one or more video streaming clients (such as video decoders) for rendering at least a portion of the video content on one or more displays. The display on which the video content is rendered can be part of one or more video streaming clients, or can operate in conjunction with one or more video streaming clients.

[0059] Example video streaming servers may include, but are not limited to, any of the following: cloud-based video streaming servers located remotely from the video streaming clients(s), local video streaming servers connected to the video streaming clients(s) via local wired or wireless networks, VR devices, AR devices, car entertainment devices, digital media devices, digital media receivers, set-top boxes, game consoles (e.g., Xbox), general-purpose personal computers, tablet computers, dedicated digital media receivers such as Apple TV or Rokubox, etc.

[0060] 4. Remove occlusion from image fragments in the data

[0061] De-occlusion data in a de-occlusion atlas can include occluded image segments in a represented (reference) view of a volumetric video signal. As described herein, an image segment refers to a contiguous, non-convex (or occluded) pixel region having per-pixel image texture information (e.g., color, luminance / chrominance values, RGB values, YCbCr values, etc.) and per-pixel depth information. The per-pixel image texture and depth information specified for an image segment in the de-occlusion atlas can visually depict image features / objects / structures that are hidden or occluded in the represented view of the volumetric video signal but may become at least partially de-occluded or visible in views adjacent to the represented view.

[0062] For a given reference view that does not contain holes with missing image texture and depth information, composite images can be generated for neighboring views surrounding the represented view using depth-based image rendering (DIBR) and image texture / depth information available to the reference view. The composite images may contain holes whose image texture and depth information cannot be obtained from the image texture / depth information available to the reference view. Using the composite images, image masks can be generated to identify holes in the composite images of neighboring views.

[0063] In some operational scenarios, an image mask can be generated at least partially for a given reference view by identifying image regions (or areas) in the given reference view that contain large depth gaps between or within adjacent pixels from other image regions in the given reference view that have relatively smooth depth transitions between or within adjacent pixels.

[0064] Image texture and depth information (such as that identified in an image mask) of image fragments within holes (or non-convex regions of pixels) can be obtained from spatially or temporally different reference views. For example, spatially different reference views, used at the same time point as a given reference view but spatially different, can contain and provide image texture and depth information of holes in a composite image of adjacent views. These spatially different reference views, including a given reference view, can collectively form a multi-view image at the same time point.

[0065] Alternatively, or alternatively, the temporally distinct reference views used at different times than the given reference view may include and provide image texture and depth information of holes in the composite image of adjacent views. These temporally distinct reference views, including the given reference view, may belong to the same visual scene, the same group of pictures (GOP), etc.

[0066] Alternatively, or alternatively, artificial intelligence (AI) or machine learning (ML) may be trained from training images and then applied to generate or predict some or all of the image texture and depth information of holes in synthetic images in neighboring views.

[0067] Image segments included in a de-occluded atlas at a given time point can be segmented into different subsets of image segments from different reference views. Each subset of image segments may include (occluded) image segments from the corresponding reference view in the different reference views.

[0068] The de-occlusion atlas technique described herein can be used (e.g., adaptively, optimally, etc.) to package these image fragments into a combined image (or "atlas") that covers the minimum total area and has no overlapping fragments. Each fragment in the combined image of the de-occlusion atlas has its own region (or area) that does not overlap with other fragments included in the de-occlusion atlas.

[0069] A volumetric video signal can be generated from a continuous sequence of multiview images. This continuous sequence of multiview images comprises multiple multiview images at multiple time points forming a continuous time point sequence. Each of these multiple multiview images comprises multiple single-view images of multiple reference views at a corresponding time point within those multiple time points.

[0070] A sequence of deocclusion atlases can be generated for a continuous time-point sequence. This sequence of deocclusion atlases includes multiple deocclusion atlases for multiple time points within the continuous time-point sequence. Each deocclusion atlas in the multiple deocclusion atlases includes an image segment comprising one or more subsets of image segments from one or more reference views represented by the corresponding time point in the volumetric video signal.

[0071] For sub-intervals (e.g., fractions of a second, a second, or more seconds) within a time interval (e.g., 30 minutes, an hour, or multiple hours) covered by a continuous time point sequence, the volumetric video signal can be encoded using one or more group of picture (GOP) subsequences of one or more reference views represented in the signal. Each GOP subsequence in the one or more GOP subsequences includes a texture image subsequence and a depth image subsequence of the corresponding reference view in the one or more reference views represented in the volumetric video signal.

[0072] Each GOP subsequence comprises one or more GOPs. Each GOP is defined by an I-frame, or begins with a starting I-frame and ends with a frame exactly before the next starting I-frame. In some embodiments, the starting I-frame and the next starting I-frame can be two most recent I-frames with no other I-frames in between. In some embodiments, the starting I-frame and the next starting I-frame can be nearby I-frames, but not necessarily two most recent I-frames. I-frames in a GOP can be decoded independently of image data from other frames; however, non-I-frames (such as B-frames or P-frames) in the GOP can be predicted at least partially from other frames in the GOP. I-frames and / or non-I-frames in a GOP can be generated from temporally stable or temporally similar source / input images. These temporally stable source / input images can facilitate relatively efficient inter-frame or intra-frame prediction and data compression or coding when generating I-frames and / or non-I-frames in a GOP.

[0073] For the same sub-interval within an interval covered by a sequence of consecutive time points, the volumetric video signal can be encoded using one or more deoccluded atlas group subsequences of one or more reference views represented in the signal. Each of the one or more deoccluded atlas group subsequences includes a texture image subsequence and a depth image subsequence of holes in a view adjacent to the corresponding reference view in one or more reference views represented in the volumetric video signal.

[0074] Each deocclusion atlas group subsequence comprises one or more deocclusion atlas groups. Each deocclusion atlas group is defined by an atlas I-frame, or begins with a starting atlas I-frame and ends with an atlas frame that is exactly before the next starting atlas I-frame. In some embodiments, the starting atlas I-frame and the next starting atlas I-frame can be two nearest atlas I-frames with no other atlas I-frames in between. In some embodiments, the starting atlas I-frame and the next starting atlas I-frame can be nearby atlas I-frames, but not necessarily two nearest atlas I-frames. Atlas I-frames in a deocclusion atlas group can be decoded independently of deocclusion data from other atlas frames; however, atlas non-I-frames (such as atlas B-frames or atlas P-frames) in a deocclusion atlas group can be predicted at least partially from other atlas frames in the deocclusion atlas group. Atlas I-frames and / or non-I-frames in a deocclusion atlas group can be generated from temporally stable or temporally similar deocclusion atlases. These temporally stable de-occlusion atlases can facilitate relatively efficient inter-frame or intra-frame prediction and data compression or coding when generating (multiple) atlas I-frames and / or (multiple) atlas non-I-frames within a de-occlusion atlas set. 5. Image masks for de-occlusion data

[0075] Figure 2A The illustration shows an example texture image in a reference view (e.g., a 360-degree "baseball overlay" view, etc.). A texture image includes texture information of an array of pixels in an image frame, such as color, luminance / chrominance values, RGB values, YCbCr values, etc. The texture image may correspond to or be indexed by a time point in a time interval covered by a continuous sequence of time points, and may be encoded into the video stream of the reference view, such as as a texture image I-frame or a texture image non-I-frame in a group of text images (GOPs) of pictures or images in the video stream.

[0076] Figure 2B The diagram illustrates the relationship with Figure 2A The texture image is an example depth image in the same reference view (e.g., a 360-degree "baseball overlay" view, etc.). Figure 2B The depth images include Figure 2A Depth images contain depth information of some or all pixels in a pixel array within a texture image, such as depth values, z-values, spatial difference values, and disparity values. A depth image can correspond to or be indexed by the same point in time within a time interval covered by a sequence of consecutive time points, and can be encoded into a video stream of a reference view, such as a depth image I-frame or a depth image non-I-frame in a group of depth images (GOP) within a video stream.

[0077] Figure 2C The illustration shows an example image mask, which can be a bit mask with a bit array. Indicators or bits in the bit array of the image mask can (e.g., 1-1, etc.) be... Figure 2ATexture images and / or Figure 2B The corresponding pixel in the pixel array represented in the depth image corresponds to the image in the image mask. Each indicator or bit in the image mask can indicate or specify the pixel to be used in image warping and hole filling operations. Figure 2A Texture images and / or Figure 2B Does the demasking atlas used with the depth image provide demasking data portions such as demasking pixel texture values ​​(e.g., color, brightness / chromaticity values, RGB values, YCbCr values, etc.) and / or demasking pixel depth values ​​(e.g., depth values, z values, spatial difference values, parallax values, etc.)?

[0078] An example hole-filling operation is described in U.S. Provisional Patent Application No. 62 / 811,956, filed April 1, 2019, by Wenhui Jia et al., the entire contents of which are fully set forth herein and are incorporated herein by reference.

[0079] Image warping and hole-filling operations can be used to generate a composite image of an observer's current view, which can be a view adjacent to a reference view. De-occluded pixel texture values ​​and / or de-occluded pixel depth values, as provided in the de-occlusion atlas, depict the image in... Figure 2A Texture images and / or Figure 2B Image details that are occluded in the depth image but may become partially visible in a view adjacent to the reference view. The deocclusion atlas may correspond to or be indexed by the same time point in a time interval covered by a sequence of consecutive time points, and may be encoded into a video stream of the reference view or a separate accompanying video stream, such as as an atlas I-frame or atlas non-I-frame in a deocclusion atlas group in a video stream or a separate accompanying video stream.

[0080] like Figure 2C The image mask shown in the middle diagram does not appear to be consistent with... Figure 2A The corresponding texture image or Figure 2B The corresponding depth image is aligned because the mask covers portions of the texture and / or depth images that are not visible from one or more adjacent views near the reference view. The purpose of the deocclusion atlas generated using the image mask is to provide texture and depth image data to fill holes in the composite view (such as the observer's current view), where the holes are caused by deocclusion during reprojection of the composite view (or the selected "reference" view). In various operational scenarios, the texture and depth data in the deocclusion atlas may cover more, less, or the same spatial region as the holes in the composite view.

[0081] In some operational scenarios, the spatial area covered by the deocclusion atlas can include a safety margin, ensuring that the deocclusion texture and depth data in the deocclusion atlas can be used to completely fill holes in views adjacent to the reference view.

[0082] In some operational scenarios, the spatial region covered by the deocclusion atlas may not include a safety margin, making it impossible for the deocclusion atlas to guarantee that the deoccluded texture and depth data in the atlas can be used to completely fill holes in views adjacent to the reference view. In these operational scenarios, the receiver video decoder may apply a hole-filling algorithm to generate at least a portion of the texture and depth information of a portion of the holes in a composite view adjacent to or near the reference view represented in the video stream.

[0083] Alternatively, or alternatively, the occluded spatial regions covered in the deocclusion atlas can be used to select salient visual objects from the visual scene depicted in the reference view. For example, the deocclusion atlas may not carry or provide any texture or depth information to the receiving video decoder to cover spatial regions far from salient visual objects. Spatial regions that the deocclusion atlas carries or provides texture or depth information to the receiving video decoder can indicate to the receiving video decoder that these spatial regions contain salient visual objects.

[0084] 6. Generation of unoccluded image atlases

[0085] Figure 3A The illustration shows an example (output) deocclusion atlas that includes (or encapsulates) image fragments representing occluded regions of one or more reference views. Image metadata can be generated to indicate which reference views each image fragment in the deocclusion atlas corresponds to.

[0086] This example demonstrates the generation of a volumetric video signal from a multiview image sequence. Each multiview image in the sequence can include a set of N single-view (input / source) texture images of N reference views and a set of N single-view (input / source) depth images of N reference views at consecutive time points in a time-point sequence.

[0087] It can receive view parameters, which specify or qualify an injective function that maps image (pixel) coordinates (e.g., pixel position, pixel row and column, etc.) and depth to a coordinate system such as a world (3-D) coordinate system. View parameters can be used to synthesize images in adjacent views, identify holes or regions that may be occluded in a reference view but could become at least partially undisturbed in adjacent views, and determine, estimate, or predict undisturbed texture data and undisturbed depth data for some or all of these holes or regions on a per-reference-view basis.

[0088] For a reference view and each single-view texture image and single-view depth image at a given time point, image masks such as bitmasks can be generated for the reference view to identify the spatial regions that need to provide their de-occluded texture and depth data in the de-occlusion atlas at a given time point. Figure 3A As shown in the diagram.

[0089] Figure 3B The illustration shows an example sequence of consecutive deocclusion atlases that can be created for a sequence of multiview images in a received or input multiview video. The deocclusion atlas sequence can be encoded into a deocclusion atlas group. Each such deocclusion atlas group comprises a temporally stable deocclusion atlas and can be encoded into a video stream relatively efficiently.

[0090] Figure 4A The illustration shows the deocclusion atlas of multi-view images used to generate a multi-view image sequence with overlapping time intervals (e.g., ...). Figure 3A The example processing flow is illustrated in the diagram. In some example embodiments, one or more computing devices or components may perform this processing flow.

[0091] The multi-view images correspond to or are indexed to a time point in the time interval, and include N (source / input) single-view texture images of N reference views and N (source / input) single-view depth images of N reference views. Each of the N single-view texture images corresponds to a corresponding single-view depth image of the N single-view depth images.

[0092] In box 402, before the deocclusion atlas is used to store (e.g., copy, stamp, place, etc.) image fragments of spatial regions or holes that may exist in a composite / distorted image in views adjacent to N reference views, the system as described herein (e.g., Figure 1A (e.g., 100) performs initialization operations on the de-occlusion atlas.

[0093] The initialization operation of box 402 may include: (a) receiving or loading N image masks that identify spatial regions or holes in the N reference views that may have missing texture or depth data in the composite / distorted images of the N reference views; (b) receiving or loading texture and depth information of image fragments identified in the N image masks; (c) sorting the image fragments by size into a list of image fragments; and so on.

[0094] Here, "size" refers to a metric used to measure the spatial dimensions of an image segment. Various metrics can be used to measure the spatial dimensions of an image segment. For example, the smallest rectangle that completely encloses an image segment can be determined. Horizontal size (denoted as "xsize"), vertical size (denoted as "ysize"), and combinations of horizontal and vertical sizes can be used individually or collectively as multiple metrics for measuring the size of an image segment.

[0095] In some operational scenarios, the size of an image fragment can be calculated as: 64*max(xsize, ysize)+min(xsize, ysize), where each of xsize and ysize can be represented in units of pixels or in the horizontal or vertical dimensions of a pixel block of a specific size (such as 2 pixels in a 2×2 pixel block, 4 pixels in a 4×4 pixel block, etc.). The horizontal or vertical dimensions can be non-negative integer powers of 2.

[0096] Each of the N loaded image masks corresponds to a corresponding reference view in the N reference views. An image mask comprises image mask portions of image segments that are occluded in the reference view but become at least partially visible in a view adjacent to the reference view. Each image mask portion of the image mask spatially delineates or defines a corresponding image segment in the image segments that are occluded in the reference view corresponding to the image mask but become at least partially visible in a view adjacent to the reference view. For each pixel represented in the image mask, a bit indicator is set to true or 1 if the pixel belongs to one of the image segments; otherwise, the bit indicator is set to false or 0 if the pixel does not belong to any of the image segments.

[0097] In some operational scenarios, de-occlusion atlases include layout masks used to describe the spatial arrangement of image segments (e.g., all, etc.) of a multi-view image, and to identify or track image segments whose de-occlusion data is stored or maintained in the de-occlusion atlas. A layout mask may include an array of pixels arranged within a spatial shape (such as a rectangle). Image segments spatially defined or confined within the layout mask of the de-occlusion atlas are mutually exclusive and do not overlap each other within the layout mask (e.g., completely, etc.).

[0098] The initialization operation of box 402 may further include: (d) creating a single-number quadtree root node. This root node is initialized to an optimal size to just cover the size of the largest image fragment. The quadtree grows at a doubling rate in each size as needed so that the corresponding layout mask remains as small as possible; (e) linking the largest image fragment to the first node of the quadtree by stamping the image fragment (e.g., a portion of the image mask) in a specified region of the layout mask of the demasking atlas for the first node; and so on. Here, the first node of the quadtree refers to the first quadtree node among the first-level quadtree nodes below the root node representing the entire layout mask. Here, "stamping" means copying, transferring, or fitting an image fragment or a portion of its image mask in the layout mask of the demasking atlas. Here, "quadtree" refers to a tree data structure in which each internal node has four child quadtree nodes.

[0099] A quadtree initially consists of four nodes with equal spatial shapes (such as rectangles of equal size). As described in this paper, the spatial shape of a node in a quadtree can have a special size with a pixel count that is a non-negative integer power of 2.

[0100] After stamping the largest image fragment into the layout mask of the demasking atlas, the largest image fragment is removed from the list of (size-sorted) image fragments, and the next quadtree node after the first quadtree node is set as the current quadtree node. The current quadtree node represents an empty or candidate quadtree node (not yet occupied by any image fragment or corresponding image mask portion) that will be used to carry the next image fragment.

[0101] In box 404, the system determines whether the list of image segments sorted by size contains any image segments that still need to be stamped or are spatially arranged into the layout mask of the demasking atlas. In some embodiments, any image segment below a minimum segment size threshold may be removed from the list or may be ignored in the list. Example minimum segment size thresholds may be one of the following: four (4) pixels in one or both of the horizontal and vertical dimensions, six (6) pixels in one or both of the horizontal and vertical dimensions, etc.

[0102] The processing flow ends when the list of image segments (sorted by size) does not contain (multiple) image segments that still need to be stamped or are spatially arranged in the layout mask of the demasking atlas.

[0103] Otherwise, in response to the determination that the list of (size-sorted) image segments contains (multiple) image segments that still need to be stamped or are spatially arranged in the layout mask of the demasking atlas, the system selects the next largest image segment from the list of (size-sorted) image segments as the current image segment.

[0104] In box 406, the system determines whether the current quadtree node in the quadtree is large enough to hold the current image fragment or the corresponding image mask portion of the current image fragment.

[0105] In response to the determination that the current quadtree node in the quadtree is not large enough to hold the current image segment, the processing flow is transferred to box 410.

[0106] Otherwise, in response to determining that the current quadtree node in the quadtree is large enough to hold the current image fragment, the processing flow goes to box 408.

[0107] In box 408, the system determines whether the current quadtree node is the "best" fit quadtree node for the current image segment. A "best" fit quadtree node is a quadtree node that is just large enough to hold the image segment or a portion of its image mask. In other words, a "best" fit quadtree node represents the smallest size quadtree node that can be used to completely enclose or hold the image segment in the layout mask of the de-occluded atlas.

[0108] In response to determining that the current quadtree node is not the “best” fitting quadtree node for the current image segment, the system subdivides (e.g., repeatedly, iteratively, recursively, etc.) the current quadtree node until a “best” fitting quadtree node is found. The “best” fitting quadtree node is then set as the current quadtree node.

[0109] Once it is determined that the current quadtree node is the "best" fitted quadtree node for the current image segment, the system adds a stamp to the "best" fitted quadtree node or defines the current image segment in space.

[0110] After stamping the current image fragment into the layout mask of the demasking atlas or the current quadtree node, the current image fragment is removed from the list of (size-sorted) image fragments, and the next quadtree node after the (removed) current quadtree node is set as the (new or current) current quadtree node.

[0111] In box 410, the system determines whether an empty or candidate quadtree node exists anywhere below the root node of the entire layout mask representing the de-occluded atlas that can be used to carry the current image fragment. If so, the empty or candidate quadtree node is used (e.g., if more than one node is used, it is used collectively, etc.) to carry the current image fragment. The processing flow then proceeds to box 404. Thus, if the current image fragment does not fit (overall) into any existing (sub)quadtree node below the current quadtree node, an attempt can be made to fit the fragment at any location in the layout mask. It should be noted that in many operational scenarios, (multiple) quadtrees are merely designed as an acceleration data structure to make atlas construction faster. Once the layout (or layout mask) is determined, quadtrees as described herein may not be saved or needed. Furthermore, there are no (e.g., absolute, inherent, etc.) restrictions imposed by quadtrees as described herein on the location where any image fragment can be placed. Image fragments can (and often in some operational scenarios) overlap multiple quadtree nodes. Therefore, if the "best fit" method (e.g., to find a single best-fit node for a segment such as the current image segment) fails, a more exhaustive (and more expensive) search can be performed on the entire layout mask to fit nearby segments. When successful, all quadtree nodes overlapping with such a placed segment are marked as "occupied," and processing continues. When it fails, the processing flow jumps to box 412 to allow the quadtree to grow. Figure 4A The reason the overall algorithm illustrated in the diagram remains efficient and effective is that, most of the time, the best-fit quadtree search is successful in many operational scenarios. Only when no best-fit node is found for the current image fragment will a more expensive or exhaustive backtracking search be performed or invoked to find potentially overlapping quadtree nodes to accommodate the current image fragment. This can involve searching (e.g., in a search loop, etc.) all empty or candidate quadtree nodes (not yet occupied by any image fragment) in the entire layout mask of the de-occluded atlas.

[0112] In response to the determination that none of the remaining empty or candidate quadtree nodes in the layout mask are large enough to hold the current image fragment, the processing flow proceeds to box 412.

[0113] Otherwise, in response to determining that an empty or candidate quadtree node in the layout mask is large enough to hold the current image fragment, the empty or candidate quadtree node is set as the (new) current quadtree node, and the processing flow goes to box 408.

[0114] In box 412, the system expands or doubles (2x) the size of the deocclusion atlas or its layout mask in both the horizontal and vertical dimensions. The existing quadtree (or old quadtree) prior to this expansion can be linked to or placed into a first quadtree node (e.g., the top-left quadrant of the newly expanded quadtree, etc.). A second quadtree node (e.g., the top-right quadrant of the newly expanded quadtree, etc.) is set as the (new) current quadtree node. The process then proceeds to box 408.

[0115] Texture and depth values ​​of each pixel identified as belonging to an image fragment in the layout mask of the demasking atlas can be stored, cached, or buffered as part of the demasking atlas along with the layout mask of the demasking atlas.

[0116] 7. Time-stable group-level layout mask

[0117] To stabilize successive demasking atlases in a video sequence, the layout masks of demasking atlases in multiple successive time points (which can correspond to texture image GOPs, depth image GOPs, etc. in the video sequence) can be separately combined using an "OR" operation to form a group-level layout mask for the successive demasking atlas group.

[0118] Each layout mask in the layout mask can be of equal size and the same pixel array can be compressed using a corresponding indicator or bit that indicates whether any pixel in the layout mask belongs to an image segment carried in the corresponding demasking atlas.

[0119] A group-level layout mask can be the same size as the (individual) layout mask of a successive demasking atlas group and includes the same pixel array as the pixel array in the (individual) layout mask. To generate a group-level layout mask via a union or split OR operation, the indicator or bit of the pixel at a specific pixel position or index in the (individual) layout mask of the successive demasking atlas group can be set to true or one (1) if either the indicator or bit of the corresponding pixel at the same specific pixel position or index is true or one (1).

[0120] Group-level layout masks or instances thereof can be repeatedly used for each demasking atlas in a successive demasking atlas group to carry or lay out the image fragment to be represented in the demasking atlas for a corresponding time point among multiple consecutive time points covered in the successive demasking atlas group. Pixels that do not have demasking texture and depth information for a given time point can be omitted (e.g., unqualified, unoccupied, etc.) in the corresponding instance of the group-level layout mask for that time point (or timestamp). Figure 3C The illustration shows an example of a continuous demasking atlas group generated using a public group-level layout mask as described in this article.

[0121] In some operational scenarios, multiple instances of the same group-level layout mask can be used to generate individual demasking atlases. The initial demasking atlas in a successive demasking atlas group (along with the initial instance of the combined group-level layout mask) can be used to generate a starting atlas I-frame, followed by other atlas frames generated from other demasking atlases in the successive demasking atlas group. The starting atlas I-frame and other atlas frames can form a successive atlas frame group defined by the starting atlas I-frame and the next starting atlas I-frame before the end of the group. Temporally stable group-level layout masks can be used to facilitate data compression operations, such as applying inter-frame prediction and / or intra-frame prediction to identify data similarities and reduce the overall data in the successive demasking atlas groups to be transmitted to the receiving video decoder. In some implementation examples, using the union of layout masks (or bit masks) within (I-frame) time intervals can improve video compression by 2x or better.

[0122] In some operational scenarios, texture and depth data of all pixels identified in each individual layout mask of a successive demasking atlas can be included or transmitted without generating a combined group-level layout mask of the successive demasking atlas. Data compression operations (using individual layout masks that are different from each other in time and space) may not be as efficient as data compression operations using group-level layout masks as described in this paper in terms of reducing data volume.

[0123] In some operational scenarios, to lay out image fragments onto a layout mask of an occluded atlas, these fragments can be fitted into available spatial regions (such as empty or candidate quadtree nodes) without prior rotation. In other scenarios, to improve wrapping efficiency, image fragments can be rotated before being placed into the "best" fitted quadtree node. Therefore, a quadtree node that might not have been able to hold an image fragment before rotation can now hold one after rotation.

[0124] A multi-view image, or any single-view image therein, can be a 360-degree image. The image data (including de-occlusion data) of a 360-degree image can be represented in an image frame such as a rectangular frame (e.g., in a "baseball coverage" view). Figure 2A and Figure 2B As illustrated, such an image can include multiple image segments that are combined together to form, for example, a rectangular image frame in a "baseball overlay" view. However, multiple image segments can be combined together to form the shape of a different view, such as the shape of a square view.

[0125] As described in this paper, demasking atlases can include or define mask stripes within a layout mask to indicate that image fragments include texture and depth information adjoining the boundaries of the image fragments. The reason for placing mask stripes in the layout mask is to prevent image fragments carried by the atlas from crossing C-shaped seams corresponding to seams in a 360-degree image.0 In the case of (or zero-order) discontinuities, the 360-degree image can include multiple image segments bound to one or more seams. For example, in a baseball overlay representation of a 360-degree image, there is a long horizontal seam in the middle, where neighboring pixels on different sides of the seam do not correspond to neighboring portions of the visual scene (e.g., the actual view, etc.). By zeroing out the lines in the input mask along this seam, mask stripes can be implemented in the demasking atlas(s) to ensure that image segments do not cross this boundary. Thus, image segments with mask stripes can be constrained and correctly interpreted to fill holes or gaps on the same side of the line as the image segments.

[0126] 8. Example Processing Flow

[0127] Figure 4B An example processing flow according to an exemplary embodiment of the present invention is illustrated. In some example embodiments, one or more computing devices or components may perform this processing flow. In block 422, the upstream device sorts image fragments by size, which are occluded in one or more reference images depicting a visual scene from one or more reference views and become at least partially deoccluded in non-reference views adjacent to one or more reference views. The image fragments include a first image fragment, which is not smaller in size than any other image fragment in the image fragments.

[0128] In box 424, the upstream device generates a layout mask for a de-occlusion atlas used to store image segments. The layout mask is covered by a quadtree that includes a first best-fit node specifically sized for the first image segment. The size of the first best-fit node is used to (e.g., completely) cover the first image segment.

[0129] In box 426, the upstream device stores the sorted image segments in descending order into the best-fit nodes identified in the layout mask. Each image segment in the sorted image segments is stored in the corresponding best-fit node. The best-fit nodes include at least one best-fit node obtained by iteratively partitioning at least one node in the quadtree covering the layout mask. Each of the best-fit nodes can be identified as the smallest-sized quadtree node used to completely cover each corresponding image segment.

[0130] In box 428, the upstream device generates a volumetric video signal encoded using one or more reference images. The volumetric video signal is further encoded using image segments from a demasking atlas. The one or more reference images are used by the receiving device of the volumetric video signal to synthesize a display image for rendering on an image display in a non-representational view. The image segments from the demasking atlas are used by the receiving device to fill demasking image data in demasking spatial regions within the display image.

[0131] In an embodiment, each of one or more reference images represents one of the following: a 360-degree image, a 180-degree image, a viewport image, an image in a regular spatial shape image frame, or an image in an irregular spatial shape image frame.

[0132] In an embodiment, for a spatial region formed by contiguous pixels that are occluded in one or more reference views, each of the image fragments includes a texture image value and a depth image value.

[0133] In one embodiment, the set of one or more prominent video streams includes a first prominent video stream assigned a first prominent level and a second prominent video stream assigned a second prominent level lower than the first prominent level; the second video stream is removed from the set of one or more prominent video streams to be transmitted to the video streaming client at a later time in response to determining that the available data ratio has decreased.

[0134] In an embodiment, one or more reference images are included in a multiview image group, which includes multiple consecutive multiview images at multiple consecutive time points; and a de-occlusion atlas is included in a de-occlusion atlas group, which includes multiple de-occlusion atlases at multiple consecutive time points.

[0135] In the embodiments, the layout mask is included in multiple individual layout masks generated for multiple deocclusion atlases; a group-level layout mask is generated from the multiple individual layout masks by a union operation; and the deocclusion atlas encoded in the volumetric video signal is represented in the group-level layout mask.

[0136] In an embodiment, the demasking atlas group is encoded as an atlas frame group in the volumetric video signal; wherein the atlas frame group begins with an atlas I-frame and ends before a different atlas I-frame.

[0137] In an embodiment, the demasking atlas includes mask stripes; the mask stripes indicate that image segments stored in the demasking atlas border each other at one or more boundaries of the image segments.

[0138] In one embodiment, the layout mask is expanded in response to determining that no best-fit node is found within the pre-expanded size of the layout mask.

[0139] In one embodiment, spatial regions for image segments are identified in a bitmask; the image segments are then sorted using the size of the spatial regions for the image segments.

[0140] In this embodiment, the image fragments stored in the deocclusion atlas are located in one or more prominent regions identified from the visual scene. A prominent region can be a more interesting or important part or area of ​​interest to the visual scene.

[0141] Figure 4C An example processing flow according to an exemplary embodiment of the present invention is illustrated. In some example embodiments, one or more computing devices or components may perform this processing flow.

[0142] In box 460, the downstream decoder (e.g., a receiving device or decoder) receives the volumetric video signal. The volumetric video signal can be encoded / generated using any of the embodiments described above, for example, referring to... Figure 4B Volumetric video signals are encoded using one or more reference images and image segments from a deocclusion atlas. The deocclusion atlas is used to store image segments. Image segments that are occluded in one or more reference images depicting a visual scene from one or more reference views and become at least partially deoccluded in non-reference views adjacent to one or more reference views are ordered by size, such as reference... Figure 4B As described in (box 422).

[0143] In box 462, the downstream device decodes one or more reference images from the volumetric video signal.

[0144] In box 464, the downstream device decodes image segments from the volumetric video signal in the demasking atlas.

[0145] In box 466, the downstream device synthesizes a display image in a non-representational view based on one or more reference images.

[0146] In box 468, the downstream device uses image fragments from the demasking atlas to fill the demasking space region in the displayed image with demasking image data.

[0147] In box 470, the upstream device renders and displays an image on an image display.

[0148] In an embodiment, for a spatial region formed by contiguous pixels that are occluded in one or more reference views, each of the image fragments includes a texture image value and a depth image value.

[0149] In an embodiment, in box 466, compositing the display image includes using texture image values ​​and depth image values ​​that can be used in one or more reference views.

[0150] In an embodiment, de-occlusion spatial regions in a synthetic display image are identified by determining that texture image values ​​and depth image values ​​available for one or more reference views are unavailable for non-reference views adjacent to one or more reference views.

[0151] In an embodiment, the image fragments stored in the deocclusion atlas are located in one or more prominent regions identified from the visual scene, and wherein the deocclusion atlas does not include any texture image values ​​or depth image values ​​to cover spatial regions far from the one or more prominent regions, so that the one or more prominent regions are identified in the composite display image.

[0152] In an embodiment, the volumetric video signal includes image metadata specifying an injective function; the injective function maps each pixel in the image segment from its pixel position in the image frame to its corresponding position in a three-dimensional coordinate system representing the visual scene.

[0153] In various example embodiments, an apparatus, system, device, or one or more other computing devices performs any or a portion of the methods described herein. In embodiments, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause the methods described herein to be performed.

[0154] Note that although individual embodiments are discussed herein, any combination of the embodiments and / or some of the embodiments discussed herein can be combined to form further embodiments.

[0155] 9. Implementation Mechanism – Hardware Overview

[0156] According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to perform these techniques, or may include digital electronic devices persistently programmed to perform these techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hardwired and / or program logic to implement the techniques.

[0157] For example, Figure 5This is a block diagram illustrating a computer system 500 on which an example embodiment of the invention may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 to process information. The hardware processor 504 may be, for example, a general-purpose microprocessor.

[0158] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When stored in non-transitory storage media accessible to processor 504, such instructions enable computer system 500 to become a dedicated machine defined to perform the operations specified in the instructions.

[0159] The computer system 500 further includes a read-only memory (ROM) 508 or other static storage device coupled to the bus 502 for storing static information and instructions of the processor 504.

[0160] Storage devices 510, such as disks or optical discs, solid-state RAM, etc., are provided and coupled to bus 502 for storing information and instructions.

[0161] Computer system 500 can be coupled to display 512, such as an LCD, via bus 502 for displaying information to a computer user. Input device 514, including alphanumeric keys and other keys, is coupled to bus 502 for transmitting information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, trackball, or arrow keys, for transmitting directional information and command selections to processor 504 and for controlling cursor movement on display 512. Typically, this input device has two degrees of freedom on two axes (a first axis (e.g., x-axis) and a second axis (e.g., y-axis)), allowing the device to specify a position in a plane.

[0162] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic. These custom hardwired logics, one or more ASICs or FPGAs, firmware, and / or program logic, combined with the computer system, enable computer system 500 to be a dedicated machine or programmed to be a special-purpose machine. According to one embodiment, the techniques described herein are executed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium (such as storage device 510). Execution of the sequence of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0163] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, floppy hard disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, flash EPROMs, NVRAMs, any other memory chips or memory cartridges.

[0164] Storage media differ from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including conductors containing bus 502. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.

[0165] Various forms of media can involve loading one or more sequences of one or more instructions to processor 504 for execution. For example, instructions may initially be loaded onto a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 500 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 502. Bus 502 loads the data into main memory 506, from which processor 504 fetches and executes the instructions. Instructions received in main memory 506 may optionally be stored on storage device 510 before or after execution by processor 504.

[0166] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides bidirectional data communication coupled to network link 520, which connects to local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity with a corresponding type of telephone line. As another example, communication interface 518 may be a Local Area Network (LAN) card for providing data communication connectivity with a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying streams of digital data representing various types of information.

[0167] Network link 520 typically provides data communication to other data devices via one or more networks. For example, network link 520 may provide a connection via local network 522 to host computer 524 or to data devices operated by Internet Service Provider (ISP) 526. ISP 526, in turn, provides data communication services via a global packet data communication network now commonly referred to as the “Internet” 528. Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals that carry streams of digital data. Signals through various networks, as well as signals on network link 520 and through communication interface 518 (which carries digital data to and from computer system 500), are example forms of transmission media.

[0168] Computer system 500 can send messages and receive data, including program code, through multiple networks, network links 520, and communication interfaces 518. In the Internet example, server 530 can transmit application request codes through the Internet 528, ISP 526, local network 522, and communication interface 518.

[0169] The received code can be executed by processor 504 upon receipt and / or stored in storage device 510 or other non-volatile memory for later execution.

[0170] 10. Equivalents, extensions, substitutes, and others

[0171] In the foregoing description, exemplary embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the sole and exclusive indication of the invention and the applicant's inventive intent is the set of claims issued in specific form according to this application, wherein such claims include any subsequent amendments. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, characteristics, advantages, or attributes not expressly referenced in the claims should not in any way limit the scope of such claims. Therefore, this specification and drawings should be viewed in an illustrative rather than restrictive sense.

[0172] Some aspects of the embodiments include the following enumerated example embodiments (EEE):

[0173] EEE1. A method comprising:

[0174] Image segments are sorted by size, the image segments being occluded in one or more reference images depicting a visual scene from one or more reference views and becoming at least partially unoccluded in a non-reference view adjacent to the one or more reference views, the image segments including a first image segment whose size is not smaller than any other image segment in the image segments;

[0175] A layout mask is generated for a demasking atlas used to store the image fragments. The layout mask is covered with a quadtree, which includes a first best-fit node of a specific size for the first image fragment. The demasking atlas is a combined image containing the minimum total area of ​​multiple non-overlapping image fragments.

[0176] The sorted image fragments are stored in descending order into the best-fit nodes identified in the layout mask. Each image fragment in the sorted image fragments is stored in the corresponding best-fit node in the best-fit nodes, which includes at least one best-fit node obtained by iteratively partitioning at least one node in the quadtree covering the layout mask.

[0177] A volumetric video signal encoded using the one or more reference images is generated, the volumetric video signal being further encoded using the image segments in the deocclusion atlas, the one or more reference images being used by a receiving device of the volumetric video signal to synthesize a display image in a non-representational view for rendering on an image display, and the image segments in the deocclusion atlas being used by the receiving device to fill deocclusion image data in deocclusion spatial regions of the display image.

[0178] EEE2. The method as described in EEE1, wherein each of the one or more reference images represents one of the following: a 360-degree image, a 180-degree image, a viewport image, an image in a regular spatial shape image frame, or an image in an irregular spatial shape image frame.

[0179] EEE3. The method as described in EEE1 or EEE2, wherein, for a spatial region formed by contiguous pixels occluded in the one or more reference views, each of the image segments includes a texture image value and a depth image value.

[0180] EEE4. The method of any one of EEE1 to EEE3, wherein the one or more reference images are included in a multiview image group in a multiview image group, the multiview image group including a plurality of consecutive multiview images at a plurality of consecutive time points; wherein the deocclusion atlas is included in a deocclusion atlas group, the deocclusion atlas group including a plurality of deocclusion atlases at the plurality of consecutive time points.

[0181] EEE5. The method as described in EEE4, wherein the layout mask is included in a plurality of individual layout masks generated for the plurality of deocclusion atlases; wherein a group-level layout mask is generated from the plurality of individual layout masks by a union operation; wherein the deocclusion atlas encoded in the volumetric video signal is represented in the group-level layout mask.

[0182] EEE6. The method as described in EEE4 or EEE5, wherein the de-occlusion atlas group is encoded as an atlas frame group in the volumetric video signal; wherein the atlas frame group begins with an atlas I-frame and ends before a different atlas I-frame.

[0183] EEE7. The method of any one of EEE1 to EEE6, wherein the demasking atlas includes mask stripes; wherein the mask stripes indicate that image segments stored in the demasking atlas are adjacent at one or more boundaries of the image segments.

[0184] EEE8. The method of any one of EEE1 to EEE7, wherein the layout mask is expanded in response to determining that no best-fit node is found within the pre-expansion size of the layout mask.

[0185] EEE9. The method of any one of EEE1 to EEE8, wherein a spatial region of the image fragment is identified in a bitmask; wherein the image fragment is sorted using the size of the spatial region of the image fragment.

[0186] EEE10. The method of any one of EEE1 to EEE9, wherein the image fragment stored in the demasking atlas is located in one or more prominent regions identified from the visual scene.

[0187] EEE11. The method as described in EEE10, wherein the one or more prominent video streams include a first prominent video stream assigned a first prominent level and a second prominent video stream assigned a second prominent level lower than the first prominent level.

[0188] EEE12. A method comprising:

[0189] Decode one or more reference images from a volumetric video signal;

[0190] Decode image segments from the demasking atlas from the volumetric video signal;

[0191] Synthesize the display image in the non-representation view from the one or more reference images;

[0192] The image fragments in the demasking atlas are used to fill the demasking space region in the displayed image with demasking image data;

[0193] The displayed image is rendered on the image display.

[0194] EEE13. The method of any one of EEE1 to EEE12, wherein the volumetric video signal includes image metadata with a specified injective function; wherein the injective function maps each pixel in the image segment from the pixel position of the pixel in the image frame to a corresponding position in a three-dimensional coordinate system, wherein the visual scene is represented in the three-dimensional coordinate system.

[0195] EEE14. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause to perform the method as described in any one of EEE1 to EEE13.

[0196] EEE15. A computing device comprising one or more processors and one or more storage media storing an instruction set, which, when executed by the one or more processors, causes to perform a method as described in any one of EEE1 to EEE13.

Claims

1. A method comprising: Image segments are sorted by size, the image segments being occluded in one or more reference images depicting a visual scene from one or more reference views and becoming at least partially unoccluded in a non-reference view adjacent to the one or more reference views, the image segments including a first image segment whose size is not smaller than any other image segment in the image segments; A layout mask is generated for a demasking atlas used to store the image fragments. The layout mask is covered by a quadtree, which includes a first best-fit node of size for covering the first image fragment. The demasking atlas is a combined image containing the minimum total area of ​​multiple non-overlapping image fragments. The sorted image segments are stored in descending order into the best-fit nodes identified in the layout mask, wherein each of the best-fit nodes is identified as a minimum-sized quadtree node for completely covering each of the corresponding image segments, and each of the sorted image segments is stored in the corresponding best-fit node, the best-fit node comprising at least one best-fit node obtained by iteratively partitioning at least one node in the quadtree covering the layout mask. A volumetric video signal encoded using the one or more reference images is generated, the volumetric video signal being further encoded using the image segments in the deocclusion atlas, the one or more reference images being used by a receiving device of the volumetric video signal to synthesize a display image in a non-representational view for rendering on an image display, and the image segments in the deocclusion atlas being used by the receiving device to fill deocclusion image data in deocclusion spatial regions of the display image.

2. The method as described in claim 1, wherein, Each of the one or more reference images represents one of the following: a 360-degree image, a 180-degree image, a viewport image, an image in a regular spatial shape image frame, or an image in an irregular spatial shape image frame.

3. The method as described in claim 1 or 2, wherein, For a spatial region formed by connected pixels that are occluded in one or more reference views, each of the image segments includes a texture image value and a depth image value.

4. The method as described in claim 1 or 2, wherein, The one or more reference images are included in a multiview image group, which includes multiple consecutive multiview images at multiple consecutive time points; wherein, the de-occlusion atlas is included in a de-occlusion atlas group, which includes multiple de-occlusion atlases at the multiple consecutive time points.

5. The method of claim 4, wherein, The layout mask is included in a plurality of individual layout masks generated for the plurality of deocclusion atlases; wherein a group-level layout mask is generated from the plurality of individual layout masks by a union operation; wherein the deocclusion atlas encoded in the volumetric video signal is represented in the group-level layout mask.

6. The method of claim 4, wherein, The de-occlusion atlas group is encoded as an atlas frame group in the volumetric video signal; wherein the atlas frame group begins with an atlas I-frame and ends before a different atlas I-frame.

7. The method as described in claim 1 or 2, wherein, The demasking atlas includes mask stripes; wherein the mask stripes indicate that image segments stored in the demasking atlas are adjacent at one or more boundaries of the image segments.

8. The method as claimed in claim 1 or 2, wherein, The layout mask is expanded in response to the determination that no best-fit node is found within the pre-expanded size of the layout mask.

9. The method as claimed in claim 1 or 2, wherein, The spatial regions of the image fragments are identified in a bitmask; wherein the image fragments are sorted using the size of the spatial regions of the image fragments.

10. The method as claimed in claim 1 or 2, wherein, The image fragments stored in the de-occlusion atlas are located in one or more prominent regions identified from the visual scene.

11. The method as claimed in claim 1 or 2, wherein, The volumetric video signal includes image metadata with a specified injective function; wherein the injective function maps each pixel in the image segment from its pixel position in the image frame to its corresponding position in a three-dimensional coordinate system, wherein the visual scene is represented in the three-dimensional coordinate system.

12. A method comprising: Receive a volumetric video signal encoded using one or more reference images and image segments from a demasking atlas, wherein the volumetric video signal is encoded by the method according to any one of claims 1 to 11; Decode the one or more reference images from the volumetric video signal; Decode the image segments from the demasking atlas from the volumetric video signal; Synthesize the display image in the non-representation view from the one or more reference images; The image fragments in the demasking atlas are used to fill the demasking space region in the displayed image with demasking image data; The displayed image is rendered on the image display.

13. The method of claim 12, wherein, For a spatial region formed by connected pixels that are occluded in one or more reference views, each of the image segments includes a texture image value and a depth image value.

14. The method according to any one of claims 12 to 13, wherein, Synthesizing the display image includes using texture image values ​​and depth image values ​​that can be used in the one or more reference views.

15. The method of claim 14, wherein, The de-occlusion spatial region in the synthetic display image is identified by determining that the texture image values ​​and the depth image values ​​available for the one or more reference views are unavailable for the non-reference views adjacent to the one or more reference views.

16. The method of claim 13, wherein, The image fragments stored in the deocclusion atlas are located in one or more prominent regions identified from the visual scene, and wherein the deocclusion atlas does not include any texture image values ​​or depth image values ​​to cover spatial regions far from the one or more prominent regions, so that one or more prominent regions are identified in the composite display image.

17. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause the method of any one of claims 1 to 16 to be performed.

18. A computing device comprising one or more processors and one or more storage media storing an instruction set, the instruction set causing, when executed by the one or more processors, to perform the method as claimed in any one of claims 1 to 16.

Citation Information

Patent Citations

  • Devices and methods for warping and hole filling during view synthesis

    CN103518222A

  • Multi-view scene segmentation and propagation

    US20170358092A1