Image synthesis system and method therefor

The image synthesis device improves immersive video by adapting transparency based on distance and depth, addressing quality degradation and user freedom issues in AR, VR, and MR applications.

JP7787187B2Active Publication Date: 2025-12-16KONINKLIJKE PHILIPS NV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023544041
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-21
Filing Date
2022-01-13
Publication Date
2025-12-16
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

Existing immersive video technologies suffer from limited view space and quality degradation when viewers move outside the capture area due to insufficient 3D data for de-occluded areas and incomplete depth maps, leading to artifacts and distortions.

Method used

An image synthesis device that adapts transparency of image regions based on distance and depth, making foreground objects invisible when the viewer moves too far, ensuring high-quality rendering by using multi-view with depth data.

Benefits of technology

Enhances user experience by maintaining image quality and reducing artifacts, allowing more freedom of movement without distortion, suitable for AR, VR, and MR applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007787187000003
    Figure 0007787187000003
  • Figure 0007787187000004
    Figure 0007787187000004
  • Figure 0007787187000001
    Figure 0007787187000001
Patent Text Reader

Abstract

The image synthesis device comprises a first receiver 201 for receiving three-dimensional image data describing at least a portion of a three-dimensional scene, and a second receiver 203 for receiving a view pose for a viewer. An image region circuit 207 identifies at least a first image region in the three-dimensional image data, and a depth circuit 209 identifies a depth indication for the first image region from depth data of the three-dimensional image data. A region circuit 211 identifies a first region for the first image region. A view synthesis circuit 205 generates a view image from the three-dimensional image data, the view image representing a view of the three-dimensional scene from the view pose. The view synthesis circuit 205 is configured to adapt a transparency for the first image region in the view image depending on the depth indication and a distance between the view pose and the first region.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to image synthesis systems, and in particular, but not exclusively, to image synthesis devices that support view synthesis for immersive video applications. [Background technology]

[0002] The variety and range of image and video applications has increased significantly in recent years, as new services and methods for using and consuming video continue to be developed and introduced.

[0003] For example, one service that is gaining popularity is the presentation of image sequences in a manner that allows the viewer to actively and dynamically interact with the system to change the parameters of the rendering. A very attractive feature in many applications is the ability to change the viewer's effective viewing position and direction, e.g., allowing the viewer to move around and look around in the presented scene.

[0004] Such features may allow a virtual reality experience in particular to be provided to the user, allowing the user, for example, to move around (relatively) freely in the virtual environment and to dynamically change their position and where they are looking. Typically, such virtual reality (VR) applications are based on a three-dimensional model of the scene, with the model being dynamically evaluated to provide a specific requested view. This approach is well known, for example, in gaming applications for computers and consoles, for example in the first-person shooter category; other examples include augmented reality (AR) or mixed reality (MR) applications.

[0005] An example of a proposed video service or application is immersive video, where video is played, for example, in a VR headset, to provide a three-dimensional experience. With immersive video, a viewer can freely look around and move around in a presented scene, which is then perceived as being seen from different viewpoints. However, in many typical approaches, the amount of movement is limited to a relatively small area around a nominal viewpoint, for example, which typically corresponds to the viewpoint from which video capture of the scene was performed. In such applications, three-dimensional scene information is often provided that allows high-quality view image synthesis for viewpoints relatively close to the reference viewpoint, although this deteriorates when the viewpoint deviates excessively from the reference viewpoint.

[0006] Immersive video is often also called six degrees of freedom (6DoF) or 3DoF+ video. MPEG Immersive Video (MIV) [1] is an emerging standard in which metadata is used in addition to existing video codecs to enable and standardize immersive video.

[0007] A problem with immersive video is the limited view space, the 3D space in which the viewer has a sufficient quality 6DoF experience. As the viewer moves outside the view space, the degradation and errors introduced by synthesizing the view images become increasingly large, resulting in an unacceptable user experience. Errors, artifacts, and inaccuracies in the generated view images occur particularly because the provided 3D video data does not provide sufficient information (e.g., de-occlusion data) for view synthesis.

[0008] For example, immersive video data is provided in the form of a multi-view-with-depth (MVD) representation of a scene. The scene is captured by many spatially distinct cameras, and the captured images are provided along with a depth map. However, as the viewpoints become increasingly different from the reference viewpoint from which the MVD data is captured, the likelihood that such a representation will not contain sufficient image data for de-occluded areas increases significantly. Thus, as a viewer moves away from the nominal position, image portions that must be de-occluded for the new viewpoint but are missing from the source view cannot be directly synthesized from image data describing such image portions. Furthermore, incomplete depth maps introduce distortions when performing view synthesis, particularly as part of view warping, which is an essential part of the synthesis process. The farther the synthesized viewpoint is from the original camera viewpoint, the more severe the distortion in the synthesized view. Thus, as a user moves out of the viewing space, the quality of the rendered view image deteriorates, and the quality typically becomes unacceptable even for relatively small movements outside the viewing space.

[0009] To solve this fundamental problem, the 5th Working Draft of the MPEG Immersive Video (MIV) standard ISO / IEC JTC1 SC29 WG11 (MPEG) N19212 includes a proposal for handling such movements outside the viewing space. The standard proposes different actions and modes to be implemented when the viewer moves outside the viewing space.

[0010] [Table 1]

[0011] However, while these approaches provide desirable performance in some scenarios, they tend not to be ideal for all applications and services. These approaches are particularly relatively complex or impractical, often resulting in a less than optimal user experience. In particular, the VHM_RENDER and VHM_EXTRAP modes result in a distorted view but keep the viewer oriented, whereas the VHM_FADE, VHM_RESET, VHM_STRETCH, and VHM_ROTATE modes prevent distortion but at best disrupt immersion or even cause the viewer to feel disoriented.

[0012] Therefore, improved approaches would be beneficial, particularly approaches that allow for improved operation, improved flexibility, an improved immersive user experience, reduced complexity, smoothed implementation, improved synthetic image quality, improved rendering, increased (possibly virtual) freedom of movement for the user, improved user experience, and / or improved performance and / or operation. Summary of the Invention [Problem to be solved by the invention]

[0013] SUMMARY OF THE INVENTION Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]

[0014] According to one aspect of the present invention, there is provided an image synthesis device comprising: a first receiver configured to receive three-dimensional image data describing at least a portion of a three-dimensional scene; an image region circuit configured to identify at least a first image region in the three-dimensional image data; a depth circuit configured to identify a depth indication for the first image region from depth data in the three-dimensional image data for the first image region; a region circuit configured to identify the first region for the first image region; a second receiver configured to receive a view pose for a viewer; and a view synthesis circuit configured to generate a view image from the three-dimensional image data, the view image representing a view of the three-dimensional scene from the view pose, the view synthesis circuit configured to adapt a transparency of the first image region in the view image depending on the depth indication and a distance between the view pose and the first region, the view synthesis circuit configured to increase transparency as the distance between the view pose and the first region increases and as the depth indication indicates less depth for the first image region.

[0015] The present invention provides an improved user experience in many embodiments and scenarios. The present invention realizes an improved trade-off between image quality and freedom of movement, for example, for AR, VR, and / or MR applications. The present approach often provides a more immersive user experience and is well suited for immersive video applications. The present approach reduces the perception of quality degradation, for example, reducing the risk that significant artifacts or errors in the view image will cause the perception of the experience to be artificial or incomplete. The present approach provides the user with an improved experience, for example, with consistent and stable movement in the scene.

[0016] This approach enables improved AR / VR / MR applications, for example, based on limited scene capture.

[0017] The transparency is semi-transparent. The first region is a set of view poses for which the 3D image data is indicated to be sufficient for image synthesis. Such indication is in response to a synthesis quality criterion being met, the synthesis quality criterion including a requirement that a quality metric for an image representation of the first image region be higher than a threshold, and the image representation is generated (by a view synthesis circuit) from the received 3D data. The view synthesis circuit is configured to identify quality metrics for image representations generated from the 3D image data for different view poses. The first region is generated to include view poses for which the quality metric is higher than a threshold.

[0018] 3D image data is a complete or partial description of a 3D scene. Pose is a position and / or orientation.

[0019] The 3D image data includes a multi-view image set. The 3D image data includes depth information, such as a depth map, for one or more images. The 3D image data includes multiple images of a scene for different view poses. The 3D image data includes a multi-view-plus-depth (MVD) representation of a scene.

[0020] The image region corresponds to an image object. The term first image region is in some embodiments replaced by the term first image object or first scene object. The first image region is in some embodiments a pixel. The term first image region is in some embodiments replaced by the term first pixel.

[0021] Any suitable distance or difference measure may be used to determine the distance, ie, any suitable distance measure may be used for the distance between the view pose and the visibility region.

[0022] The image regions are specifically generated to correspond to regions of the input image that do not correspond to background regions. A first image region is one that does not contain background pixels. A first image region is an image region that represents a foreground object of a scene. A foreground object is an object that is not a background object.

[0023] According to an optional feature of the invention, the view synthesis circuitry is configured to generate a view image that includes a fully transparent image region when a distance between the view pose and the first region is greater than a threshold.

[0024] This provides beneficial, and typically highly efficient, operation and results in an improved user experience in many scenarios. In particular, it typically results in foreground objects becoming invisible if the view pose deviates too far from the viewing area. In particular, having foreground objects disappear, rather than being presented at a significantly lower quality, provides a more intuitive experience for many users in many scenarios.

[0025] The threshold depends on the depth indication, and in some embodiments is zero.

[0026] According to an optional feature of the invention, the view synthesis circuitry is configured to generate a view image that includes image regions that are not fully transparent if the distance is not greater than a threshold.

[0027] This provides beneficial, and typically more efficient, operation and / or an improved user experience in many scenarios.

[0028] According to an optional feature of the invention, the view synthesis circuitry is configured to generate a view image that includes an opaque image region if the distance is not greater than a threshold.

[0029] This provides beneficial, and typically more efficient, operation and / or an improved user experience in many scenarios, for example, it is beneficial in many embodiments for foreground objects to be perceived as either completely present (fully opaque) or completely invisible / absent (fully transparent).

[0030] According to an optional feature of the invention, the image synthesis further comprises an image region circuit that identifies a second region relative to the first image region, and the view synthesis circuit is configured to generate a view image including an image region that is opaque when the view pose is within the second region, partially transparent when the view pose is outside the second region and within the first region, and fully transparent when the view pose is outside the first region.

[0031] This provides an improved user experience in many embodiments. For example, the present approach presents foreground objects so that they are perceived as fully present / opaque when the view pose is close enough to the capture pose, fully absent / transparent when the view pose is too far from the capture pose, and with gradually increasing transparency between these regions.

[0032] The second viewing area is inside / surrounded by the first area.

[0033] According to an optional feature of the invention, the first region depends on the depth indication.

[0034] This provides beneficial operation and / or an improved user experience in many embodiments.

[0035] In some embodiments, the image region circuitry is configured to adapt at least one of a shape and a size of the first viewing region in response to the depth indication.

[0036] According to an optional feature of the invention, the first region depends on the complexity of the shape of the image region.

[0037] This provides beneficial operation and / or an improved user experience in many embodiments.

[0038] In some embodiments, the image region circuitry is configured to adapt at least one of a shape and a size of the first viewing region in response to a measure of complexity of the shape.

[0039] In some embodiments, the image region circuitry is configured to adapt at least one of a shape and a size of the first viewing region in response to a parallax variation measure for the image region.

[0040] The disparity variation measure indicates the variation in disparity for pixels in an image region for a given viewpoint shift.

[0041] According to an optional feature of the invention, the first region depends on the view shift / pose change sensitivity for the image region.

[0042] This provides beneficial operation and / or an improved user experience in many embodiments.

[0043] According to an optional feature of the invention, the first region depends on an amount of de-occlusion data for the first image region included in the three-dimensional image data.

[0044] This provides beneficial operation and / or an improved user experience in many embodiments.

[0045] According to an optional feature of the invention, the function for determining transparency as a function of distance includes a hysteresis related to changes in viewing pose.

[0046] This provides beneficial operation and / or an improved user experience in many embodiments.

[0047] According to an optional feature of the invention, the three-dimensional image data further includes an indication of an image region for at least one of the input images of the three-dimensional image, and the image region circuit is configured to identify the first image region in response to the indication of the image region.

[0048] This provides beneficial operation and / or an improved user experience in many embodiments. The approach reduces complexity and / or computational burden in many embodiments.

[0049] According to an optional feature of the invention, the three-dimensional image data further includes an indication of a given region for at least one input image of the three-dimensional image, and the region circuit is configured to identify the first region according to the indication of the given region.

[0050] This provides beneficial operation and / or an improved user experience in many embodiments. The approach reduces complexity and / or computational burden in many embodiments.

[0051] According to an optional feature of the invention, the view synthesis circuit is configured to select from a plurality of candidate pixel values ​​derived from different images of the multi-view image for at least a first pixel of the view image, the view synthesis circuit being configured to select the rearmost pixel if the distance is greater than a threshold, and to select the frontmost pixel if the distance is less than the threshold, the rearmost pixel being associated with a depth value indicating a depth furthest from the view pose, and the frontmost pixel being associated with a depth value indicating a depth closest to the view pose.

[0052] This provides beneficial operation and / or an improved user experience in many embodiments.

[0053] It allows for particularly efficient and less complicated operation.

[0054] According to one aspect of the invention, there is an image signal comprising three-dimensional image data describing at least a portion of a three-dimensional scene and a data field indicating whether rendering of the three-dimensional image data should include depth indication for the image region and adapting transparency for an image region of the three-dimensional image data in the rendered image depending on the distance between the view pose for the rendered image and a reference region for the image region.

[0055] According to an optional feature of the invention, the image signal comprises at least one of an indication of an image region and a reference region.

[0056] According to one aspect of the present invention, there is provided an image signal device configured to generate an image signal as described above.

[0057] According to one aspect of the present invention, there is provided a method of image synthesis, the method comprising: receiving three-dimensional image data describing at least a portion of a three-dimensional scene; identifying at least a first image region in the three-dimensional image data; identifying a depth indication for the first image region from depth data in the three-dimensional image data for the first image region; identifying a first region for the first image region; receiving a view pose for a viewer; and generating a view image from the three-dimensional image data, the view image representing a view of the three-dimensional scene from the view pose, wherein generating the view image comprises adapting a transparency for the first image region in the view image in response to the depth indication and a distance between the view pose and the first region, the transparency increasing as the distance between the view pose and the first region increases and as the depth indication indicates a decreasing depth for the first image region.

[0058] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0059] Embodiments of the present invention will now be described, by way of example only, with reference to the drawings, in which: [Brief explanation of the drawings]

[0060] [Figure 1] FIG. 1 illustrates an example of image and depth capture of a 3D object. [Figure 2] FIG. 2 illustrates example elements of an image synthesis device according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0061] Three-dimensional video capture, distribution, and presentation are becoming increasingly common and desirable for several applications and services. A particular approach, known as immersive video, involves providing views of real-world scenes and often real-time events that typically allow for small viewer movements, e.g., relatively small head movements and rotations. For example, real-time video broadcasts of sporting events, for example, that allow for local, client-based generation of views that follow the viewer's small head movements, provide the impression that a user is sitting in the stands while watching the sporting event. The user can, for example, look around and have a natural experience similar to that of a spectator present at that location in the stands. Recently, display devices that include position tracking and 3D interaction support applications based on 3D capture of real-world scenes have become popular. Such display devices are well suited for immersive video applications that provide an enhanced three-dimensional user experience.

[0062] To provide such services for real-world scenes, the scene is typically captured from different positions, with different camera capture poses used. As a result, multi-camera capture and, for example, 6DoF (six degrees of freedom) processing are rapidly gaining relevance and importance. Applications include live concerts, live sports, and telepresence. The freedom to choose one's own viewpoint improves these applications by enhancing the sense of realism beyond regular video. Furthermore, immersive scenarios can be envisioned, where observers navigate and interact with the live-captured scene. For broadcast applications, this requires real-time depth estimation at the production side and real-time view synthesis at the client device. Both depth estimation and view synthesis introduce errors, which depend on the implementation details of the algorithms.

[0063] In the art, the terms configuration and pose are used as general terms for position and / or orientation. For example, the combination of position and orientation / orientation of an object, camera, head, or view is called a pose or configuration. Thus, a configuration or pose indication includes six values / components / degrees of freedom, and each value / component typically describes an individual characteristic of the configuration / position or orientation / orientation of the corresponding object. Of course, in many situations, a configuration or pose can be considered or represented using fewer components, for example, when one or more components are considered fixed or irrelevant (e.g., if all objects are considered to be at the same height and horizontally oriented, four components provide a complete representation of the object's pose). In the following, the term pose will be used to refer to a position and / or orientation represented by one to six values ​​(corresponding to the maximum possible degrees of freedom). The term pose may be replaced by the term configuration. The term pose may be replaced by the terms position and / or orientation. The term pose may be replaced by the term position and orientation (if the pose provides both position and orientation information), by the term position (if the pose provides position (or possibly only position) information), or by orientation (if the pose provides orientation (or possibly only orientation) information).

[0064] A frequently used approach to representing a scene is known as multi-view with depth (MVD) representation and capture. In such an approach, a scene is represented by multiple images with associated depth data, where the images represent different view poses from a typically limited capture area. The images are in practice captured by using a camera rig with multiple cameras and depth sensors.

[0065] An example of such a capture system is shown in Figure 1. The figure shows a scene to be captured, including a scene object 101 in front of a background 103. Multiple capture cameras 105 are located in a capture area 105. The result of the capture is a representation of the 3D scene by means of multi-view images and depth representations, i.e., images and depths provided for multiple capture poses. The multi-view images and depth representations thus provide a description of the 3D scene from the capture zone. The data representing the 3D scene thus provides a representation of the 3D scene from the capture zone from which the visual data provides a description of the 3D scene.

[0066] The MVD representation is used to perform view synthesis, in which a view image of a scene from a given view pose can be generated. The view pose requires a view shift of the images of the MVD representation to the view pose so that the image of the view of the scene from the view pose can be generated and presented to a user. The view shift and synthesis is based on depth data, for example, along with a disparity shift between positions in the MVD image and the view pose image according to the depth of the corresponding object in the scene.

[0067] The quality of the generated view images depends on the image and depth information available for the view synthesis operation, which in turn depends on the amount of view shift required.

[0068] For example, view shifting typically results in the deocclusion of portions of an image that are not visible in, e.g., the primary image used for the view shift. If another image captures the deoccluded elements, such holes are filled with data from the other image; however, it is also typical for the deoccluded image portion for the new viewpoint to be missing from other source views. In this case, view synthesis requires estimating the data based on, e.g., peripheral data. The deocclusion process inherently tends to introduce inaccuracies, artifacts, and errors. Furthermore, this tends to increase with the amount of view shift; in particular, the likelihood of missing data (holes) during view synthesis increases as the image's distance from the capture pose increases.

[0069] Another source of possible distortion is incomplete depth information. Often, depth information is provided by a depth map, and the depth values ​​are generated by depth estimation (e.g., by disparity estimation between source images) or imperfect measurement (e.g., ranging), and therefore the depth values ​​contain errors and inaccuracies. View shift is based on depth information, and incomplete depth information will result in errors or inaccuracies in the synthesized image. The farther the synthesized viewpoint is from the original camera viewpoint, the more severe the distortion in the synthesized target-view image.

[0070] Thus, as the view pose moves further and further from the capture pose, the quality of the synthesized image tends to degrade. If the view pose is far enough away from the capture pose, the image quality will degrade to an unacceptable level, and a poor user experience will be experienced.

[0071] Many different approaches have been proposed to address these issues, but these tend to be suboptimal, and in particular tend to undesirably restrict user movement or introduce undesirable user effects. Figure 2 illustrates a view synthesis device / system that offers performance and approaches that may achieve a more desirable user experience in many scenarios.

[0072] 2 shows an example of an image synthesizer that may be used to generate view images, for example for an immersive video experience. The image synthesizer comprises a first receiver 201 configured to receive 3D image data that describes at least a portion of a 3D scene. The 3D image data may in particular describe a real-world scene captured by cameras at different positions.

[0073] In many embodiments, the 3D image data comprises a multi-view image, and thus comprises multiple (simultaneous) images of a scene from different viewpoints. In many embodiments, the 3D image data is in the form of an image and depth map representation, where one image and an associated depth map are provided. The following description focuses on embodiments where the 3D image data is a multi-view plus depth representation comprising at least two images from different viewpoints, at least one of the images comprising an associated depth map. It will be appreciated that if the received data is a multi-view data representation that does not, for example, include an explicit depth map, the depth map may be generated using a suitable depth estimation algorithm, in particular a disparity estimation approach using different images of the multi-view representation.

[0074] Thus, in a particular example, the first receiver 201 receives MVD image data that describes a 3D scene using a plurality of images and depth maps, also referred to below as source images and source depth maps, and it will be appreciated that a temporal sequence of such 3D images is provided for a video experience.

[0075] The image synthesis system further comprises a second receiver 203 configured to receive a viewpose for a viewer (and in particular for a three-dimensional scene). The viewpose represents a position and / or orientation from which the viewer views the scene, and in particular provides a pose of a subject from which a view of the scene must be generated. It will be appreciated that many different approaches for identifying and providing a viewpose are known, and that any suitable approach may be used. For example, the second receiver 203 may be configured to receive pose data from an eye tracker, or the like, from a VR headset worn by the user.

[0076] The first receiver and the second receiver may be implemented in any suitable manner and may receive data from any suitable source, including local memory, a network connection, a wireless connection, a data medium, or the like.

[0077] The receiver may be implemented as one or more integrated circuits, such as, for example, an application specific integrated circuit (ASIC). In some embodiments, the receiver may be implemented as one or more programmed processing units, such as, for example, firmware or software running on a suitable processor, such as a central processing unit, digital signal processing unit, or microcontroller. In such embodiments, it is understood that the processing unit includes on-board or external memory, clock driver circuits, interface circuits, user interface circuits, etc. Such circuitry may further be implemented as part of the processing unit, as an integrated circuit, and / or as discrete electronic circuitry.

[0078] The first receiver 201 and the second receiver 203 are coupled to a view synthesis circuit 205 configured to generate at least one view image from the received three-dimensional image data, the view image being generated to represent a view of the three-dimensional scene from a view pose. Thus, the view synthesis circuit 205 generates the view image of the 3D scene from the received image data and the view pose.

[0079] It will be appreciated that in many cases, a stereo image / image object is generated that includes a view image / object for the right eye and a view image / object for the left eye, such that when the view images are presented to a user, for example via an AR / VR headset, the 3D scene appears as it is observed from the view pose.

[0080] Thus, the view synthesis circuitry 205 is typically configured to perform depth-based view shifting of the multi-view image, which typically includes techniques such as pixel shifting (changing pixel positions to reflect appropriate disparity corresponding to disparity changes), de-occlusion (typically based on infilling from other images), combining pixels from different images, etc., as will be understood by those skilled in the art.

[0081] It will be appreciated that many algorithms and approaches are known for synthesizing images, and that any suitable approach may be used by view synthesis circuitry 205.

[0082] Thus, the image synthesizer generates view images for a 3D scene. Furthermore, as the view pose dynamically changes in response to a user moving around the scene, the view of the 3D scene is continuously updated to reflect the change in view pose. For static scenes, the same source-view images are used to generate the output-view images, whereas for video applications, different source images are used to generate different view images, e.g., a new set of source images and depths is received for each output image. Thus, the processing is frame-based. In the following, static scenes are considered for clarity and brevity of explanation. However, it will be understood that the approach applies equally to dynamic scenes by simply generating output-view images over a given period / frame based on the source images and depths received over that period / frame.

[0083] The view synthesis circuit 205 is configured to generate views of the scene and scene objects to be from different angles relative to lateral movement of the view pose. When the view pose is changed so that the view pose is in a different direction / orientation, the view synthesis circuit 205 is configured to generate views of the three-dimensional scene objects to be from different angles. Thus, as the view pose changes, the scene objects may be perceived as static and with a fixed orientation in the scene. The viewer effectively moves and views the objects from different directions.

[0084] The view synthesis circuitry 205 may be implemented in any suitable manner, including as one or more integrated circuits, such as, for example, an application specific integrated circuit (ASIC). In some embodiments, the receiver is implemented as one or more programmed processing units, such as, for example, firmware or software running on a suitable processor, such as a central processing unit, digital signal processing unit, or microcontroller. In such embodiments, it will be understood that the processing unit includes on-board or external memory, clock driving circuits, interface circuits, user interface circuits, etc. Such circuits may further be implemented as part of the processing unit, as an integrated circuit, and / or as discrete electronic circuits.

[0085] As mentioned above, a problem with view synthesis is that the quality degrades the more the view pose of the object from which the views are synthesized differs from the capture pose of the provided scene image data. In fact, if the view pose moves too far from the capture pose, the generated images will contain significant artifacts and errors and will become unacceptable.

[0086] The apparatus of Figure 2 provides functionality for solving and mitigating such problems, and implements an approach for solving and mitigating such problems. In particular, the view synthesis circuit 205 is configured to identify a first region for an image region in a 3D image and to adapt the transparency of that image region depending on the distance between the view pose and the visibility region. The first region is hereinafter referred to as a / first region, or more frequently as a / (first) visibility region.

[0087] The view synthesis circuit 205 adapts, for example, the transparency of an object depending on how close the view pose is to the viewing region; in particular, the view synthesis circuit 205 increases transparency as the distance of the view pose relative to the viewing region increases. As a particular example, if the viewer moves such that the view pose is too far from the viewing region / capture pose, one or more of the foreground objects will be rendered to be completely transparent. In such an example, if the view pose moves too far from the capture pose, the foreground objects will, for example, become invisible and disappear from the scene rather than being rendered / presented with significant errors and artifacts.

[0088] The adaptation of transparency to an image region further depends on the depth marking for the image region, with greater transparency being preferred for shallower depths. Thus, transparency is adapted based on several considerations, and depends in particular on both the depth of the image region and the distance between the view pose and the viewing region.

[0089] This provides an improved user experience in many scenarios and applications compared to presenting significantly degraded foreground objects. This approach reflects the inventors' realization that improved performance can be achieved by processing regions / objects at different depths with different morphologies, and in particular that more frontal regions / objects tend to exhibit substantially greater degradation in quality than more rearward regions / objects (and especially than the background).

[0090] The view synthesis circuit 205 further comprises an image region circuit 207 configured to identify one or more image regions in the 3D image, and in particular an image region of one of the images in the multi-view image representation. The image region is identified, for example, to correspond to a scene object or a portion of a scene object. In some embodiments, the image region is identified as a relatively small region, for example, an area of ​​less than, for example, 10,000, 1000, 100, or even 10 pixels. Indeed, in some embodiments, the image region is only one pixel.

[0091] Image regions are objects (especially scene objects).

[0092] Different approaches may be used to identify one or more image regions. For example, in some embodiments, each pixel is considered to be a separate image region. In other embodiments, for example, the input image is tiled into different tiles, with each tile being an image region. For example, a predetermined tiling arrangement is implemented, such that each image region corresponds to a predetermined image region.

[0093] However, in many embodiments, dynamic identification of image regions is performed. For example, the image is segmented into a number of image segments that are believed to correspond to scene objects or parts thereof. For example, the segmentation results are a function of image properties, such as pixel color and brightness. Thus, image regions that have similar visual properties and are therefore likely to be part of the same object are generated. Segmentation may alternatively or additionally be based on detecting transitions in the image and using such transitions as indicators of boundaries between image regions.

[0094] In many embodiments, the identification of image regions is alternatively or additionally based on consideration of depth maps / depth information. For example, image regions may additionally or alternatively be formed to further consider depth uniformity to account for visual uniformity, such that image regions have similar depths, thereby increasing the likelihood that they belong to the same scene object. Similarly, depth transitions may be identified and used to find the edges of image regions.

[0095] In some embodiments, scene objects are detected and image regions corresponding to the objects are identified.

[0096] It will be appreciated that many different approaches and algorithms are known for identifying image regions, and in particular for object detection / estimation and / or image segmentation, and therefore any suitable approach may be used.

[0097] In the above examples, image regions are generated based on a 3D image. In some embodiments, the image regions are identified based on received metadata describing the image regions. For example, the 3D image is received in a bitstream that further includes metadata identifying one or more image regions. For example, for each pixel or block of pixels (e.g., for each macroblock), metadata is received that identifies whether the pixel is a background pixel or a foreground pixel. Image regions are then identified as regions proximate to the foreground pixels.

[0098] The image region circuit 207 is coupled to a depth indication circuit 209 that is configured to identify a depth indication for each image region. The depth indication indicates the depth of the image region.

[0099] A depth indication for an image region is any indication or value that reflects a depth characteristic for the image region, and in particular any indication that reflects the depth of the image region.

[0100] It will be understood that any suitable function or algorithm for identifying such depth indicators from the depth data of the three-dimensional image data may be used, and that the particular identification and depth indicators used will depend on the particular characteristics and requirements of each individual embodiment.

[0101] In many embodiments, the depth indication for an image region is specified as a depth value that is given as a function of the depth values ​​for the image region, the depth values ​​being included in the depth data of the three-dimensional image data, the function and relationship between the depth indication and the depth values ​​of the depth data of the three-dimensional image data depending on the particular embodiment.

[0102] The depth indication may be determined, for example, by considering all depth values ​​for pixels in an image region and determining the depth indication as, for example, the average depth, median depth, maximum depth, or minimum depth for the pixels in the image region. In some embodiments, the depth indication is simply a binary value or an indication of the depth interval to which the image region belongs. For example, the depth indication may simply be an indication of whether the corresponding image region is background or foreground. Of course, many other options are possible and useful and may be used to provide the desired effect and performance for a particular embodiment. In fact, the above is merely exemplary, and many other options for generating a depth indication for an image region are possible and may be used without detracting from the invention.

[0103] The input source receiver 401 and the depth indication circuit 209 are coupled to a view synthesis circuit 205 that is provided with data describing the identified visibility region and associated depth indication.

[0104] The view synthesis device further comprises a view region circuit 211 configured to identify a viewing region for an image region. In some embodiments where multiple image regions are identified / generated, the view region circuit 211 is configured to generate a viewing region that is common to all or some of the image regions. In other embodiments, a separate viewing region is generated for each separate image region. Thus, different image regions may be linked to the same viewing region or to different viewing regions.

[0105] The visible / first region is a nominal or reference region relative to the image region. The reference or nominal region is identified relative to the image region as one for which a criterion is met. The exact criterion depends on the particular embodiment. In many embodiments, the criterion is, for example, a geometric criterion, and the visible / nominal / reference region is identified as a region for which a geometric criterion associated with the image region and / or capture region for the first image region is met. For example, 3D image data provides image data representing a view of a three-dimensional scene from one or more capture regions and / or points. The visible region is identified as a region for which a geometric criterion associated with the capture region / point is met. The visible region is particularly identified as a region for which a proximity criterion associated with the capture region / point is met.

[0106] The visibility region for an image region is the region of poses for which it is believed that the image region can be synthesized / view-shifted with a given minimum quality, and in particular, the set of poses for which the representation provides data that allows a view image to be generated that includes image regions of sufficiently high quality. Thus, for view poses that fall within the visibility region for the image region, it is believed possible to generate a view image of sufficient quality for the image region. For view poses outside the visibility region, it is believed that it is not guaranteed that a view image of sufficient quality for the image region can be generated.

[0107] The exact selection / identification / characterization of the visibility region (typically represented by its boundary, contour, or edge) will, of course, depend on the particular preferences and requirements of each individual embodiment. For example, in some embodiments, the visibility region is identified to correspond directly to the capture zone, i.e., it is the zone spanned by the capture pose. In many embodiments, the visibility region is identified to include poses where a distance measure between the pose and the nearest capture pose meets a criterion.

[0108] A visibility region is identified in some embodiments as being a region where proximity criteria related to the capture region for three-dimensional image data are satisfied, the exact proximity requirements depending on the requirements and preferences of a particular embodiment.

[0109] In some embodiments, the visibility region is identified as a region where an image quality measure for compositing image regions is higher than a threshold. The image quality measure used depends on the particular desirability of the implementation. For example, in some embodiments, the quality measure is identified as a function of the magnitude of the view shift required to perform the compositing from the received 3D image data and / or as an estimate of the extent to which deocclusion must be compensated for, for example, by interpolation. In some embodiments, the visibility region is static, specifically the same for all image regions. In other embodiments, the visibility region is dynamically identified depending on the characteristics of the image region. In this case, different image regions contain different visibility regions, and a visibility region is identified specifically for each image region.

[0110] In many embodiments, the viewing area is R N is defined as a subset of poses in space, where N is the number of dimensions considered. In many embodiments, for example, particularly many 6DoF applications, N is equal to 6, typically corresponding to three coordinates / dimensions indicating position and three coordinates indicating orientation ( / direction / rotation). In some embodiments, N is less than 6, corresponding to some dimensions not being considered (and specifically ignored or considered fixed).

[0111] In some embodiments, only the position dimension or coordinate is considered, and in some embodiments, only the orientation dimension is considered, but in many embodiments, at least one position dimension and one orientation dimension are considered.

[0112] The visibility region is typically at least two-dimensional and includes poses with different values ​​in at least two coordinates / dimensions. In many embodiments, the visibility region is at least three-dimensional and includes poses with different values ​​in at least three coordinates / dimensions. The visibility region is typically a zone in at least two or three dimensions. The visibility region typically includes poses that vary in at least two dimensions.

[0113] In many embodiments, the visibility region includes poses with different orientations, and therefore often has a non-zero extent in at least one orientation coordinate / dimension.

[0114] In most embodiments, the visibility region spans at least one orientation dimension and at least one position dimension, and thus in most embodiments both position and orientation are taken into account by the system.

[0115] In many embodiments, the visibility region is identified simply as a region of a pose where a predetermined distance to a reference or preferred viewing pose is less than a given threshold. In other embodiments, the distance is measured relative to a given capture region. As described below, in some embodiments, more complex considerations are applied, with the visibility region depending on many different parameters, etc. In general, however, it will be understood that any suitable approach for identifying the visibility region for an image region may be used, and that the approach is not limited to any particular manner of identifying the visibility region.

[0116] In many embodiments, a visibility region for a given image region is identified as a region where high-quality compositing of the given image region is likely to be achieved from the received 3D image data, but this is not required, and other approaches may be used. For example, a visibility region may be identified as a region where it is desirable to bias the user toward that region. For example, in a game or virtual reality application, it may be desirable to bias the user toward a particular location or region. Such an approach may be used, for example, to bias the user toward a position directly in front of a virtual object, even though this object in the 3D image data is represented by image data captured substantially from one or both sides of the object. Thus, adaptive transparency may be used to bias the user toward a position that does not provide optimal compositing quality, but is preferred for other purposes, including purposes unrelated to compositing quality / process.

[0117] In many embodiments, the view region circuit 211 is configured to identify the first region in response to at least one capture pose for the three-dimensional image data. In particular, the view region circuit 211 is configured to identify the first region in response to a proximity criterion being satisfied for one or more capture poses for the three-dimensional image data. For example, the first region is identified as a region for which a proximity criterion is satisfied for at least one capture pose for the three-dimensional image data.

[0118] In many embodiments, the first region is a capture region with reference to which three-dimensional image data is provided.

[0119] The three-dimensional image data provides three-dimensional image data linked to a capture / reference pose. The capture / reference pose is a pose from which at least some of the three-dimensional image data is perceived / captured. The first region is identified as a location near the capture / reference pose (e.g., when a proximity criterion is met, e.g., the distance from a pose in the first region to the reference / capture pose is less than a given distance).

[0120] In some examples, more than one reference / capture pose is provided for the three-dimensional image data, in such cases, identifying the visibility region comprises selecting one, more, or all of the reference / capture poses and identifying the visibility region as a region of poses for which a closeness criterion to at least one of the selected capture / reference poses is satisfied.

[0121] The first (visible) region and the first image region do not overlap in many embodiments, and in many embodiments, no pose (and possibly no position) falls within both the first (visible) zone and the first image zone.

[0122] The viewing area may be specified as any reference or nominal area that provides a basis for adaptive transparency according to a particular preference desired. The first viewing area is a preferred viewing area that indicates a preferred area for the view pose.

[0123] In some embodiments, the received data includes an indication of a visibility area or parameters used to identify the visibility area. Thus, the received data includes data describing or allowing a nominal and / or reference area to be identified. This nominal / reference area is then used as a reference for the described adaptive transparency to provide a desired effect.

[0124] For example, the 3D image data may be generated by MVD capture as described above, and together with the image and depth map, an indication of the capture area, or more directly the view area, may be included in the 3D image data.

[0125] The view synthesis circuit 205 is configured to generate an image of the scene for a view pose (particularly a stereo image set for a VR headset) based on the received 3D images, and thus in a particular example based on the MVD images and depth.

[0126] However, the view synthesis circuitry 205 is further configured to perform adaptive rendering of image regions based on the association between the view pose and the visibility region for the image region, and in particular to adapt the transparency of the image region based on the association between the view pose and the visibility region for the image region.

[0127] In particular, the view synthesis circuitry 205 is configured to adapt the transparency / translucency of image regions in the view image depending on the depth indication for the image region and the distance between the view pose and the visibility region. The view synthesis circuitry 205 is configured to adapt the transparency such that the transparency increases as the distance between the view pose and the first region increases and as the depth indication indicates a decrease in depth relative to the first image region.

[0128] For example, the transparency is specified as a function of the depth indicator and the distance between the view pose and the viewing region. The function is monotonic with respect to the depth indicator, and in particular, monotonically increasing as the depth indicator indicates decreasing depth. Depth is considered to increase toward the background. The function is also a monotonically increasing function of the distance between the view pose and the first region.

[0129] In the following, the approach is described with reference to one image region, referred to as the first image region, but it is understood that the approach is repeated for many more, and typically all, identified image regions. It is further understood that in some embodiments, transparency is specified to be common to multiple image objects.

[0130] As a particular example, in some embodiments, view synthesis circuitry 205 is configured to increase transparency the greater the distance between the view pose and the visibility region. For example, when the view pose is within the visibility region, the image region is rendered fully opaque, but becomes more transparent as the view pose moves further outside the visibility region, until at a given distance the image region is rendered fully transparent; i.e., for view poses farther from the visibility region, the image objects represented by the image region become invisible and the image background is revealed rather than the image region being shown.

[0131] Thus, in such an example, when applied to an immersive video application, a view pose beyond the visible region results in all image regions becoming invisible and fully transparent, so that only the background of the immersive video scene is presented. In such an example, the foreground objects may be replaced by a background, which may be generated, for example, from a different image of the MVD representation if available, or by inpainting if the data is not available. Such an approach may result in an extreme form of deocclusion by making the foreground objects fully transparent. This requires or is based on the expectation that the occluded data is available (made available) from the 3D image data, or is generated on the fly (e.g., inpainted from the surrounding area).

[0132] Such an approach effectively expands the primary viewing space, in which the scene is fully presented / rendered with high quality, along with a secondary viewing space in which only the background is shown. Indeed, the inventors have noticed that image quality degradation for the background and more distant objects tends to be less than that for closer foreground objects, so the perceived quality of the secondary viewing space remains high. Thus, instead of rendering increasingly lower-quality foreground objects as the viewer moves further and further away from the viewing area, these become invisible, while the background and, by extension, the overall scene remains visible. The user is provided with an experience in which the lower-quality rendered images closer to the viewer disappear, but the scene as a whole is still maintained and is still of sufficient quality. While such an experience may feel unnatural to the user for some applications and scenarios, it provides a substantially more beneficial and often intuitive user experience in many embodiments and applications. For example, when a user notices that foreground objects begin to disappear, they intuitively realize that they have moved farther away and begin to move back toward the viewing area. Furthermore, in some situations, a user moves away from the viewing area precisely to be able to look around a foreground object, i.e., to be able to see the object or the background behind the object. In such cases, it is a highly desirable experience for the foreground object to become transparent, allowing the user to see through the foreground object. Furthermore, in contrast to other proposed approaches for solving the problem of degraded quality when a user moves too far away from the desired viewing area, the present approach allows the user to still experience consistency in their perception of their position in the scene, and can, for example, navigate to a more desirable location. The approach provides a more immersive experience in many scenarios.

[0133] The view synthesis circuit 205 is configured to specify a transparency for a first image region and to generate a view image including the first image region with the specified transparency. Thus, the view synthesis circuit 205 is configured to generate a view image including the first image region generated with a specified transparency depending on the depth indication and the distance between the viewing pose and the first region. The view synthesis circuit 205 adapts the transparency of the first image region in the view image by specifying and applying the (adapted) transparency to the first image region. The (adapted) transparency is specified depending on the depth indication and the distance between the viewing pose and the visibility region. The adapted transparency is specified in particular as an alpha value for objects / components / portions in the first image region, and the view synthesis circuit 205 is configured to generate the view image using the alpha values ​​for these objects / components / portions. It is understood that algorithms for generating a view image based on such transparency / alpha values ​​are known to those skilled in the art and therefore will not be described in further detail herein.

[0134] Thus, for example, the transparency of an object depends on different distances. In many embodiments, relying on depth indication results in relying on the distance from the view pose to the object, i.e., the object distance (to the view pose) is used in specifying the transparency. In addition, the distance from the view pose to the visibility region is used to adapt the transparency of the object. Therefore, the view pose change distance, which reflects a change in the view pose (relative to the visibility region), is also taken into account. For example, for a visibility region corresponding to a capture region, the transparency is adapted according to both the object distance and the view pose change distance. Such consideration provides a significantly improved effect.

[0135] In particular, the adaptation is performed such that the transparency / translucency increases as the object depth / object distance decreases and as the distance between the view pose and the visibility area increases, i.e., as the view pose change distance increases.

[0136] Different considerations include non-separable and / or non-linear and / or multiplicative effects. For example, adaptations along the following lines may be applied:

[0137] [Table 2]

[0138] The dependence of transparency on depth marking and viewpose distance (to the viewing area) is inseparable in many embodiments and is non-linear and / or multiplicative.

[0139] Adaptation is a constructive combination of depth indication and view pose distance. In particular, let A denote the distance between the view pose and the view region, and let B denote the depth indication, and let f(A,B) for the transparency of the first image region have the following property: ·For any B, there is a non-negative correlation between f(A,B) and A. For any A, there is a non-negative correlation between f(A,B) and B. ·There is a positive correlation between f(A,B) and A for some B. ·For some A, there is a positive correlation between f(A,B) and B.

[0140] In the aforementioned example, the image region is identified based on the received 3D image or metadata. In this example, the image region is referenced to the 3D input image, and in particular, a region in one of the images used for view synthesis, such as the nearest MVD image. In other embodiments, the image region is referenced, for example, to the output image. For example, for an image object or area in an input 3D input image, a corresponding area in the output image is identified taking into account the depth of the object or area. For example, a required parallax shift for depth is identified, and an image region in the output image corresponding to the image object in one or more input images is identified. Typically, an image region is occupied by the foremost pixel calculated by translation from different images (as this is what a person sees along their line of sight), but in current approaches, the transparency of one or more of the pixels in this image region is adapted based on the distance between the view pose and the viewing region. In particular, for a given pixel in an image region of the output-view image, the transparency of the foremost pixel value (or just one pixel, if only one image provides the pixel after parallax shift) is pixel-dependent.

[0141] The image region is an image region in an input image of a 3D input image. The image region is an image region in an input image of a 3D multi-view input image. The image region is an image region in an input image of a 3D input image that provides the front-most pixel for an image region in the synthesized output-view image. The image region is an image region in an input image that corresponds to a given pixel region in the synthesized output-view image.

[0142] The transparency of a pixel is specifically its alpha value, and thus the alpha value for at least one pixel depends on the distance between the view pose and the visibility region. The transparency for a pixel value reflects the degree to which further back scene objects (including background) are visible. In particular, for a pixel in the output view image, the pixel value is generated as a combination of the front-most pixel value and the further back pixel value generated from the 3D input image (typically by parallax shifting). The further back pixel value is generated from the 3D input image (typically by parallax shifting or by infilling). The further back pixel value is a background pixel.

[0143] As a specific example, the output-view image is generated by finding, for each pixel in the output-view image, a corresponding source pixel in each of the view input images. A source pixel in a given input image is identified as the pixel that results in the source pixel being at the location of the output pixel after a parallax shift caused by a viewpoint shift from the source image to the output-view image. For some source images, no such pixel exists (e.g., such a pixel is a deoccluded pixel), and thus the view synthesis circuit 205 identifies a number of source pixels that is no greater than the number of source images. Each source pixel is further associated with a depth. Traditionally, the source pixel with the lowest depth, i.e., closest to the source camera, is selected because this source pixel corresponds to the foremost object and is therefore the one seen by the viewer from the view pose in the view direction represented by the pixel. However, in an example of the current approach, the view synthesis circuit 205 proceeds to use this approach if the view pose falls within (or within a threshold distance of) the visibility region; otherwise, the view synthesis circuit 205 proceeds to select the source pixel that is furthest toward the rear, i.e., farthest from the view pose. Typically, this pixel is a background pixel. Thus, this effectively renders the object represented by the foreground pixel completely transparent or invisible, and instead of presenting this object, the background is presented. In such an approach, the image region in the source image is identified as the pixel that is at the location of a given output pixel after view shifting / warping.

[0144] It is understood that further considerations are included, for example, if all of a set of source pixels correspond to non-background objects (e.g., their distance is less than a threshold), then none of the source pixels are determined to be suitable for the output image, and instead, suitable values ​​are found, for example, by infilling from adjacent background pixels.

[0145] Thus, in some examples, the view synthesis circuit 205 is configured to select from multiple candidate pixel values ​​from image regions of the input multi-view images for at least a first pixel of the output image. In this example, the selection is based on depth relative to the pixel, which depends on whether the distance between the view pose and the first visibility region is less than a threshold. If the distance is less than the threshold, the view synthesis circuit 205 selects the front-most pixel; otherwise, the view synthesis circuit 205 selects the rear-most pixel.

[0146] The rearmost pixel is the pixel associated with a depth value that indicates a depth closest to the rear and / or furthest from the view pose. The frontmost pixel is the pixel associated with a depth value that indicates a depth closest to the front / closest to the view pose. The rearmost pixel is the pixel furthest from the view pose. The rearmost pixel is the pixel closest to the view pose.

[0147] Such an approach provides a highly effective implementation where modifications to existing techniques result in a less complex algorithm where foreground objects become invisible / disappear when the view pose moves too far from the visibility region.

[0148] In some embodiments, view synthesis circuitry 205 is configured to generate a view image that includes fully transparent image regions when the distance between the view pose and the visibility region (hereinafter referred to simply as the viewing distance) is greater than a given threshold, which may be zero. Thus, in such cases, view synthesis circuitry 205 renders the view image such that foreground objects are completely invisible / non-existent when the viewing distance is sufficiently long.

[0149] Similarly, in some embodiments, the view synthesis circuitry (205) is additionally or alternatively configured to generate view images that include opaque image regions when the viewing distance is not greater than a threshold. Thus, in such cases, the view synthesis circuitry 205 renders view images with foreground objects fully visible / present when the viewing distance is sufficiently short.

[0150] This approach is combined so that the foreground object is either completely present or completely absent (completely opaque or completely transparent) depending on whether the viewing distance is greater than a threshold.

[0151] This can be a highly desirable effect in some embodiments, providing a clear indication to the user that they have moved too far from the preferred pose and must move back towards the viewing area, for example.

[0152] In many embodiments, the view synthesis circuit 205 is configured to generate a view image to render a first image region with transparency applied, the transparency being determined as a function of both a depth indication for the first image region and the distance between the view pose and the first region.

[0153] When the transparency is not completely opaque, the view image is generated to include contributions from other visual elements for the first image region. A pixel light intensity value for a pixel of the view image representing the first image region is generated as a weighted combination of a contribution from at least one visual element of the first image region and a contribution from at least one other visual element. The other visual element is particularly an object (including the scene background) behind the first image region from the view pose. However, in some scenarios, the at least one other visual element is possibly an element that does not directly correspond to the scene, e.g., a particular visual characteristic (e.g., a black or gray background). The weighting of the contributions from visual elements of the first image region increases with decreasing transparency. The weighting of the contributions from visual elements not belonging to the first image region increases with increasing transparency.

[0154] Thus, the view image is generated such that the greater the transparency, the more the first image region is visible through it. Typically, increasing the transparency provides the effect of the first image region in the view image being more "see-through" so that the scene behind the first image region is partially visible. Thus, increasing the transparency typically allows scene objects behind the first image region to become increasingly visible through the first image region. In particular, in many embodiments, the background of the scene becomes increasingly visible through the first image region.

[0155] In some embodiments, transparency is generated by including visual contributions from elements that are not part of the scene, but instead are elements that have fixed or constant visual characteristics, such as, for example, a uniform color or a predetermined texture.

[0156] Thus, in many embodiments, the greater the transparency, the more the first image region becomes visible, and therefore the objects in the first image region fade. In most embodiments, therefore, visibility of the scene behind the first image region becomes visible, and therefore the objects in the first image region fade to reveal the scene behind them.

[0157] The view synthesis circuit 205 is configured to apply the (adapted) transparency by increasing the (relative) contribution from the first image region to the (light intensity pixel values ​​for) pixels in the view image corresponding to the first image region the lower the transparency.

[0158] Alternatively, or in addition, the view synthesis circuit 205 is configured to apply the (adapted) transparency by increasing the (relative) contribution from elements that are not of the first image region to pixels (for light intensity pixel values) in the view image corresponding to the first image region the higher the transparency.

[0159] In many embodiments, the view synthesis circuit 205 is configured to adapt the transparency of the first image region in the view image so that the greater the transparency, the more visible part of the three-dimensional scene behind the first image region becomes.

[0160] In many embodiments, the view synthesis circuit 205 is configured to adapt the transparency of a first image region in the view image such that a higher transparency allows a portion of the three-dimensional scene behind the first image region to contribute more to the view image. In some embodiments, hysteresis is included so that the threshold is adapted depending on whether the view distance is currently greater than or less than a threshold. Thus, to change an opaque object to transparent, the user is required to move to increase the view distance from less than a given first value to greater than the given first value, and to change a transparent object to opaque, the user is required to move to decrease the view distance from greater than a given second value to less than a given second value, the first value being greater than the second value. Such an approach prevents a ping-pong effect, in which foreground objects flicker between being perceived as present and not present.

[0161] Thus, in some embodiments, the transparency as a function of distance has a hysteresis related to changes in viewing pose.

[0162] The hysteresis is velocity-independent. The transparency as a function of distance is a hysteresis loop. The transparency value depends on the direction of change in distance. In some embodiments, the view synthesis circuit 205 is configured to generate view images with more gradual changes in transparency of image regions, and in particular of foreground objects. For example, in some embodiments, the transparency (often represented by an alpha value in the field) is gradually increased with increasing view distance. The transparency is a monotonically increasing function of view distance.

[0163] Such an approach to generating partially transparent objects may be combined with a binary approach. For example, instead of simply identifying one visibility region, two visibility regions may be identified, one within the other. In such an embodiment, a view image is generated that includes an image region that is opaque when the view pose is within the inner visibility region and completely transparent when the view pose is outside the outer visibility region. For viewer poses between the two regions, transparency is gradually increased as a monotonically increasing function of the distance to the inner visibility region (or similarly gradually decreased as a monotonically increasing function of the distance to the outer visibility region). Such an approach provides a gradual user experience, where objects do not immediately appear or disappear as the viewer moves, but rather transition gradually through an intermediate region. In such an approach, objects appear / disappear gradually, which, for example, can alleviate any viewing discomfort experienced due to the effect.

[0164] In some embodiments, a view pose that exceeds the viewing area by more than a given amount results in only the background of the immersive video scene being visible, thus maintaining the viewer's immersion. To do so, foreground objects are replaced by the background if available, or inpainted if not.

[0165] Our approach extends the primary viewing region, where the scene is fully rendered / synthesized, with a secondary viewing region, where only the background is rendered / synthesized. The secondary viewing region is larger than the primary viewing region, but is more limited because depth maps are involved in view synthesis.

[0166] Thus, in some embodiments, if the viewing distance is greater than a threshold, the scene is, for example, no longer presented, e.g., one of the prior art modes mentioned above is applied in this situation.

[0167] In the above examples, the approach has been described primarily with reference to one visibility region, but as explained, the approach may be applied separately to different image regions. Different visibility regions may be identified for different image regions. For example, depending on depth, image regions may be categorized into a set of predetermined categories, each of which is associated with image regions of different sizes.

[0168] This approach is particularly implemented so that when the view pose exceeds the primary viewing space, foreground objects are replaced by the background when available and inpainted when unavailable. The size of the inpainting region of missing data can be large, depending on the size of the foreground object and the availability of background information in other views. In some embodiments, only foreground objects that include substantial background available from other views are removed; i.e., the transparency depends on whether data is available for deocclusion. Such foreground objects are typically the smallest objects and are most forward / foreground. The inpainted region results in the perception of some blurring of the background. However, this blurring is insignificant or tolerable, and is typically constant over time. It has been found that even if some blurring of the background occurs, any visual distortion is perceived as relatively unobtrusive compared to existing approaches.

[0169] In many embodiments, the viewing region dynamically depends on different parameters, particularly parameters that affect the quality of the synthesis process. For example, the more data provided in the 3D input images, the better the quality of the view image that can be synthesized for a given view pose, and therefore the less quality degradation there will be. In some embodiments, the viewing region circuit 211 is configured to adapt the viewing region, and in particular to adapt at least one of the size and shape of the viewing region, in response to the quality affecting parameters.

[0170] In many embodiments, the visibility area for an image region depends on the view-shift / view-pose change sensitivity for the image region. The view-shift / view-pose change sensitivity for an image region reflects how sensitive the image region is to distortions introduced by performing view-shift / view-pose change synthesis. The viewpoint change sensitivity for an image region indicates the sensitivity of the image region to changes in the view-pose, which can be used to improve performance. For example, a relatively complex object that is relatively close to the camera will have a smaller visibility area than a relatively distant, flat object.

[0171] In some embodiments, the view region circuitry 211 is configured to adapt a viewing region for an image region / object in response to a depth indication for the image region / object. In particular, the view region circuitry 211 is configured to adapt at least one of a shape and a size of the viewing region for an image region in response to a depth indication for the viewing region.

[0172] In many embodiments, the size of the image region is increased as the depth indication indicates that the image region is further back. For example, the visibility region for an object that is relatively closer to the view pose is smaller than the visibility region for an object that is relatively farther from the view pose. Thus, the closer an object is to the foreground, the smaller the visibility region and therefore the less movement of the view pose is required before the foreground object becomes invisible.

[0173] Typically, the closer an object is to the viewer, the greater the quality degradation; therefore, by adapting the viewing area to the depth indication, a more gradual user experience can be achieved, where the transparency of the object is more flexibly adapted to reflect the quality degradation.

[0174] In some embodiments, the size and / or shape of the view region, and in particular the visibility region, for a given image region depends on the complexity of the shape of the image region, hi some embodiments, the view region circuitry 211 is configured to adapt at least one of the shape and size of the first visibility region depending on a measure of the complexity of the shape.

[0175] The view-shift portion of view synthesis tends to introduce less distortion for image regions / objects with simple shapes compared to more complex ones. For example, simple shapes have higher consistency between neighboring pixels and tend to be less de-occluded than complex shapes. Therefore, the more complex the shape, the larger the size of the visibility region for that image region.

[0176] Shape complexity is determined according to different measures in different embodiments, for example, algebraic complexity, such as how many sides a viewable area is represented by, the angles between such sides, etc.

[0177] In some embodiments, the visibility region for an image region depends on a disparity variation measure for the image region. The view synthesis circuit 205 is configured to adapt at least one of a shape and a size of the visibility region for the image region in response to the disparity variation measure for the image region. The disparity variation measure indicates the variation of disparity for pixels of the image region for a given viewpoint shift. The disparity variation measure is in particular a depth variation measure for the image region.

[0178] The view region circuit 211 is configured to determine the view region to be smaller, for example, for large amounts of disparity or depth variation in an image region. If there is a large variation in depth, and therefore parallax is required when implementing viewpoint shifting, there is a greater likelihood of introducing distortion or inaccuracy, which may result in, for example, greater deocclusion. Thus, the greater the amount of disparity or depth variation for a given image region, the smaller the view region will be, and therefore the smaller the view pose will be before the image region begins to become transparent.

[0179] The view region circuit 211 is configured to determine a view region for an image region based on the depth quality of the depth information provided for the view region. The depth quality of an object may be an indication of how well the object can be reprojected from a first (actual) camera view to a second (actual) camera view. For example, a floor surface (with a less complex shape) is likely to have high depth quality. Depth quality is relatively easy to determine in many embodiments. For example, a view shift of an input image of the MVD representation to the position of another input of the MVD representation is performed based on the depth data for the input image. The result is compared with the corresponding data in the input image of the MVD representation for the image region, and the depth quality is based on this comparison. The closer the synthesized image is to the input image, the higher the depth quality.

[0180] The view region for an image region depends on the amount of de-occlusion data for the image region that is included in the three-dimensional image data. In many embodiments, the view region circuit 211 is configured to adapt at least one of the shape and size of the view region for an image region depending on how much de-occlusion data is available for the image region in the received 3D data.

[0181] For example, if the received 3D image data includes another image that views the image area from a capture pose at a significantly different angle, this provides substantially additional data that enables improved de-occlusion. The more de-occlusion data available, the larger the viewing area. This reflects the fact that the more de-occlusion data there is available, the less degradation is likely to occur from view shift.

[0182] The amount of de-occlusion data for an image region in the input images of the multi-view representation is determined, for example, by performing a view shift of all different view images of the representation to the capture pose for the input image. The data, and in particular the depth, for the image region determined by such view synthesis is then compared with the original image region. The larger the difference, the more de-occlusion data is considered to be present, since the difference reflects that different images have captured different objects in line of sight from the current input image pose.

[0183] In some embodiments, the more de-occlusion data available, the smaller the visibility region is created, reflecting that the more de-occlusion data available, the easier it is to generate an accurate view of the background, and therefore the higher the quality of the scene presented after removal of foreground image regions / objects.

[0184] Indeed, if deocclusion data is not available, deocclusion requires inpainting, which introduces further degradation to the rendering and makes foreground objects invisible. This depends, for example, on the size of the foreground object and the availability of background information in other views. In some embodiments, only foreground objects with substantial background available from other views are removed. Thus, if deocclusion data is available to composite background in the absence of foreground objects, small visibility regions are identified, whereas if deocclusion data is available, very large (possibly infinite) visibility regions are generated. These are typically the smallest objects and are furthest to the front.

[0185] The above example of how the visibility area is adapted also applies to the dependence of transparency on view distance, i.e. the function therefore also depends on any of the parameters mentioned above to affect visibility area specification.

[0186] In some embodiments, the described view synthesis device performs operations to identify image regions and, for example, to divide the received data into foreground and background image regions. Similarly, in the description so far, operations are performed to identify visibility regions for different image regions. However, in some embodiments, the received input data includes data describing image regions and / or visibility regions.

[0187] For example, the view synthesis apparatus of Figure 2 is a decoder-based implementation, where the input data is received from an encoder. In addition to providing 3D image data, the image data stream includes further data describing image regions for at least one of the input images.

[0188] For example, the received 3D image data is for a given input image (e.g., of a multi-view representation) that includes an image region map that indicates, for each pixel, whether the pixel is a foreground pixel or a background pixel. In other embodiments, the 3D image data indicates, for example, for each non-background pixel, the identity of the image region to which the pixel belongs.

[0189] In such an embodiment, the image region circuit 207 is configured to identify an image region according to the received data indication. For example, the image region circuit 207 considers each foreground pixel to be an image region. As another example, the image region circuit 207 groups a set of contiguous foreground pixels into an image region. If the received data includes an image region identification, the image region circuit 207 groups pixels provided with the same identification into an image region.

[0190] In some embodiments, the received 3D data includes an indication of the visibility region that must be applied: the visibility region may be a fixed visibility region that must be applied to all image regions / objects, or for example different visibility regions may be defined for different image regions or for different properties associated with the image regions.

[0191] In such a case, the view region circuit 211 determines the view region according to the received indication of the view region, e.g., the view region circuit 211 simply uses the view region defined in the received data.

[0192] The advantage of using a data stream containing such information is that it significantly reduces the complexity and resource demands on the decoder side. This is important, for example, for embodiments where data is distributed to many decoders, thus centralized operation reduces overall resource demands and provides a consistent experience for different users. Typically, more information and / or options for control are also available on the encoder side; for example, manual identification of visibility or image regions is practical.

[0193] In many embodiments, an image signal device, e.g., an encoder, is configured to generate an image signal that includes 3D image data and further includes a data field / flag indicating whether the described approach for rendering should be applied or not.

[0194] Thus, the image signal device generates an image signal including three-dimensional image data describing at least a portion of a three-dimensional scene, a depth indication for an image region, and a data field indicating whether rendering of the three-dimensional image data should include adapting transparency for an image region of the three-dimensional image data in the rendered image depending on the distance between the view pose for the rendered image and the viewing area for the image region.

[0195] As a specific example, the described approach is added as an additional operating mode to the list of possible modes provided in the 5th Working Draft of the MPEG Immersive Video (MIV) standard ISO / IEC JTC1 SC29 WG11 (MPEG) N19212, which includes proposals for handling such motion outside the viewing space. For example, a mode ID bit for an unassigned value (e.g., between 7 and 63) may be used to indicate that the described approach of making one or more foreground objects transparent is used.

[0196] As will be described, in some embodiments, image region identification is based on processing in the synthesizer (decoder) or in the image signal device (encoder). Doing it in the decoder makes decoding more (computationally) expensive. Doing it in the encoder is more feasible but requires information about background regions to be transmitted to the decoder. The preferred tradeoff depends on the embodiment.

[0197] In the following, a particular approach is described that is based on a binary separation of the image into foreground (FG) and background (BG) and using image regions corresponding to FG regions / pixels. In a particular example, the segmentation into FG and BG is performed on the encoder side.

[0198] The approach follows the following steps:

[0199] 1. Compute a dense FG / BG segmentation for each source view. As a result, next to the color and depth attributes, each pixel has an FG or BG label.

[0200] 2.a) The MIV "entity" extension is used to transmit the FG / BG segmentation map to the decoder. To do so, the MIV encoder receives as a further input a binary entity map containing the obtained per-pixel FG / BG segmentation. The resulting bitstream therefore contains metadata identifying the entity ID (e.g., "background" label) for each rectangular texture atlas patch, and the modification of that label to the pixel level via "occupancy". This allows the decoder to reconstruct the segmentation map. b) Alternatively, a new "background" flag is added to the standard specifically for this purpose.

[0201] 3. A second visual space for background visibility is put into the bitstream metadata. The MIV standard currently does not support multiple visual spaces. However, it allows a "guard_band_size" to be specified for a (primary) visual space. This effectively obtains a secondary visual space of the same shape as the primary visual space (visibility region), but larger. Alternatively, a modification to the MIV standard must be used to allow for multiple visual spaces, or a non-standardized approach must be chosen.

[0202] At the decoder, dense FG / BG labels are reconstructed from the decoded bitstream and attached to the vertices of rendering primitives (e.g., triangles). When using, for example, OpenGL for view synthesis, the labels are attached to a "texture" and can be sampled by the vertex shader. Optionally, the vertex shader can attach segmentation labels to vertices as attributes. When the viewer moves beyond the view space boundary, all vertices with FG labels can be directly discarded by setting their output values ​​outside the valid clip space, or the attached segmentation labels can be used later to discard them there (geometry and fragment shaders have direct means to discard primitives). Discarding foreground objects in the view synthesis process is likely to increase the size of missing data. The process of inpainting missing data is already available in the normal decoding process and will not be described in detail.

[0203] Different approaches are used to segment an image into foreground and background (FG / BG segmentation). At the core of the process is a connectivity measure, i.e. a measure that reflects how connected adjacent pixels are. In one example, the world space distance (meters) is used for this purpose. Each pixel has a world space (x,y,z) coordinate through the use of a depth map. When two adjacent pixels have a distance below a certain threshold (e.g. 2 cm depending on the depth map quality), they are considered to be connected. Distinct objects are defined, which are clusters (regions) of pixels that are connected to themselves or to the floor surface only.

[0204] The following steps are used to perform FG / BG segmentation.

[0205] 1. Find the floor surface. In this embodiment, it is assumed that the z-component (height) of the selected world coordinate system is orthogonal to the floor surface. If this is not the case, further steps are performed to make it so. For all pixels in the image, the minimum value of the z-component (height) is found (the "z-floor"). To do so robustly, for example, the average of the 1st percentile of the minimum z-values ​​is taken. All pixels in the image with z-values ​​around the "z-floor" are labeled. A threshold (possibly the same as the connectivity threshold) is used for this purpose.

[0206] 2. A connected component analysis is performed on the unlabeled pixels in the image.

[0207] 3. Find regions of "available hidden layers," meaning regions of foreground pixels for which background data is available from other source views. To find these for a particular source view, the depth map of that source view is composited from all other available source views. Only here, using inverse z-buffering (OpenGL: glDepthFunc(GL_TRUE)), significant priority is given to the background in the compositing process. In normal view compositing, priority is given to the foreground. By using inverse z-buffering, the composited result contains all available background. It is a distorted image in which many foreground objects have disappeared or are eroded by the background. By slicing (thresholding) the difference between the original depth map and the inversely composited one, foreground regions containing available hidden layers are identified through a binary pixel map, which is only used for analysis.

[0208] 4. Components that contain significant hidden layer portions are classified as "foreground." Significance is then determined by dividing the area of ​​the connected components that contain "available hidden layers" by the total area of ​​that component. The larger the percentage, the more background is occluded by that component indicating it is a foreground object. Optionally, "foreground" classifications can be appended with their component number to distinguish them from other foreground objects.

[0209] 5. Unclassified pixels are classified as "background."

[0210] The invention may be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the invention may be physically, functionally, and logically implemented in any suitable way. Indeed, functionality may be implemented in one unit, in several units or as part of other functional units. Thus, the invention may be implemented in a single unit, or may be physically and functionally distributed between different units, circuits, and processors.

[0211] In accordance with standard terminology in the art, the term pixel is used to refer to pixel-related properties, such as light intensity, depth, and position of the portion / element of the scene represented by the pixel. For example, the depth of a pixel, or pixel depth, is understood to refer to the depth of the object represented by that pixel. Similarly, the brightness of a pixel, or pixel brightness, is understood to refer to the brightness of the object represented by that pixel.

[0212] While the present invention has been described in connection with several embodiments, it is not intended that the present invention be limited to the specific form set forth herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, although functions may appear to be described in connection with particular embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprise" does not exclude the presence of other elements or steps.

[0213] Furthermore, even if listed independently, a plurality of means, elements, circuits, or method steps may be implemented by, for example, a single circuit, unit, or processor. Furthermore, although individual functions may be included in different claims, they may also be beneficially combined in some cases, and their inclusion in different claims does not imply that the combination of features is not feasible and / or beneficial. Furthermore, the inclusion of functions in a certain category of claims does not imply limitation to this category, but rather indicates that the functions are equally applicable to other claim categories, as appropriate. Furthermore, the order of functions in the claims does not imply any particular order in which the functions must be performed, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in that order. Rather, steps may be performed in any suitable order. In addition, the singular does not exclude a plurality; thus, terms such as "first," "second," etc. do not exclude a plurality. Reference signs in the claims are provided merely as examples for clarity and are not to be construed as limiting the scope of the claims in any manner.

[0214] Generally, examples of an image synthesis device, an image signal, a method of image synthesis, and a computer program for implementing the method are illustrated by the following embodiments.

[0215] Embodiment Embodiment 1. A first receiver (201) configured to receive three-dimensional image data describing at least a portion of a three-dimensional scene; an image region circuit (207) configured to identify at least a first image region in the three-dimensional data; a depth circuit (209) configured to identify a depth indication for the first image region from depth data of the three-dimensional image data; a visibility circuit (211) configured to identify a first visibility region for a first image region; a second receiver (203) configured to receive a viewpose for a viewer; a view synthesis circuit (205) configured to generate a view image from the three-dimensional image data, the view image representing a view of the three-dimensional scene from a view pose, the view synthesis circuit (205) configured to adapt a transparency for a first image region in the view image according to a depth indication and a distance between the view pose and the first viewing region; An image synthesis device comprising:

[0216] Embodiment 2. The view synthesis circuit (205) is configured to generate a view image including a fully transparent image region when the distance is greater than a threshold value. An image synthesis device according to embodiment 1.

[0217] Embodiment 3. The view synthesis circuit (205) is configured to generate a view image including image regions that are not fully transparent if the distance is not greater than a threshold. An image synthesis device according to embodiment 2.

[0218] Embodiment 4. The view synthesis circuit (205) is configured to generate a view image including an opaque image region if the distance is not greater than a threshold. An image synthesis device according to embodiment 2.

[0219] Embodiment 5. The system further comprises an image region circuit (207) that identifies a second viewing region relative to the first image region, and wherein the view synthesis circuit (205) is configured to generate a view image including image regions that are opaque when the view pose is within the second viewing region, partially transparent when the view pose is outside the second viewing region but within the first viewing region, and fully transparent when the view pose is outside the first viewing region. 10. The image synthesis device according to any one of the preceding embodiments.

[0220] Embodiment 6. The first viewing area depends on the depth marking. 10. The image synthesis device according to any one of the preceding embodiments.

[0221] Embodiment 7. The first visibility region depends on the complexity of the shape of the image region. 10. The image synthesis device according to any one of the preceding embodiments.

[0222] Embodiment 8. The first visibility region depends on the view-shift sensitivity to the image region; 10. The image synthesis device according to any one of the preceding embodiments.

[0223] Embodiment 9. The first visibility region depends on the amount of de-occlusion data for the first image region included in the three-dimensional image data. 10. The image synthesis device according to any one of the preceding embodiments.

[0224] Embodiment 10. The function for determining transparency as a function of distance includes a hysteresis related to changes in viewing pose. 10. The image synthesis device according to any one of the preceding embodiments.

[0225] Embodiment 11. The three-dimensional image data further includes an indication of an image region for at least one of the input images of the three-dimensional image, and the image region circuit (207) is configured to identify the first image region according to the indication of the image region; 10. The image synthesis device according to any one of the preceding embodiments.

[0226] Embodiment 12. The three-dimensional image data further includes an indication of a viewing area for at least one of the input images of the three-dimensional image, and the viewing area circuit (211) is configured to identify the first viewing area according to the indication of the viewing area; 10. The image synthesis device according to any one of the preceding embodiments.

[0227] Embodiment 13. A view synthesis circuit (205) is configured to select from a plurality of candidate pixel values ​​derived from different images of the multi-view image for at least a first pixel of the view image, wherein the view synthesis circuit (205) is configured to select the rearmost pixel if the distance is greater than a threshold, and to select the frontmost pixel if the distance is less than the threshold; 10. An image synthesis system according to any one of the preceding embodiments.

[0228] Embodiment 14. Three-dimensional image data describing at least a portion of a three-dimensional scene; a data field indicating whether the rendering of the three-dimensional image data must include adapting the transparency for an image region of the image of the three-dimensional image data in the rendered image depending on the depth marking for the image region and the distance between the view pose for the rendered image and the visibility region for the image region; an image signal, including:

[0229] Embodiment 15. Further comprising at least one of an image area and a viewing area indication. An image signal as described in embodiment 14.

[0230] Embodiment 16: An image signal device configured to generate the image signal according to embodiment 14 or embodiment 15.

[0231] Embodiment 17. A method of image synthesis, comprising: receiving three-dimensional image data describing at least a portion of a three-dimensional scene; identifying at least a first image region in the three-dimensional data; identifying a depth indication for the first image region from depth data of the three-dimensional image data; Identifying a first viewing area for a first image area; receiving a viewpose for a viewer; generating a view image from the three-dimensional image data, the view image representing a view of the three-dimensional scene from a view pose, said generating the view image comprising adapting a transparency for a first image region in the view image according to a depth indication and a distance between the view pose and the first viewing region; A method comprising:

[0232] 18. A computer program product comprising computer program code means adapted to perform all the steps described in embodiment 17 when said program is run on a computer.

[0233] The invention is defined in more detail in the accompanying claims.

Claims

1. a first receiver for receiving three-dimensional image data describing at least a portion of a three-dimensional scene; an image region circuit for identifying at least a first image region in the three-dimensional image data; a depth circuit for determining a depth indication for the first image region from depth data of the three-dimensional image data for the first image region; a region circuit for identifying a first region for the first image region; a second receiver for receiving a viewpose for a viewer; a view synthesis circuit that generates a view image from the three-dimensional image data, the view image representing a view of the three-dimensional scene from the view pose, the view synthesis circuit adapting a transparency of the first image region in the view image in response to the depth indication and a distance between the view pose and the first region, the view synthesis circuit increasing transparency as the distance between the view pose and the first region increases and as the depth indication indicates a decreasing depth relative to the first image region; An image synthesis device comprising:

2. The view synthesis circuit generates the view image including a completely transparent image region when the distance between the view pose and the first region is longer than a threshold. The image synthesis device according to claim 1 .

3. the view synthesis circuit generates the view image including the image region that is not fully transparent if the distance between the view pose and the first region is not greater than the threshold. The image synthesis device according to claim 2 .

4. The view synthesis circuit generates the view image including the image region that is opaque when the distance between the view pose and the first region is not longer than the threshold. The image synthesis device according to claim 2 .

5. the image synthesis device further comprising an image region circuit that identifies a second region relative to the first image region, the view synthesis circuit generating the view image including image regions that are opaque when the view pose is within the second region, partially transparent when the view pose is outside the second region but within the first region, and fully transparent when the view pose is outside the first region. The image synthesis device according to any one of claims 1 to 4.

6. the first region being dependent on the depth marking; The image synthesis device according to any one of claims 1 to 5.

7. The first region depends on the complexity of the shape of the image region. The image synthesis device according to any one of claims 1 to 6.

8. the first region is dependent on view pose change sensitivity for the image region; The image synthesis device according to any one of claims 1 to 7.

9. The first region depends on the amount of de-occlusion data for the first image region included in the three-dimensional image data. The image synthesis device according to any one of claims 1 to 8.

10. The method of claim 1, wherein the function for determining the transparency as a function of the distance between the view pose and the first region includes hysteresis associated with changes in viewing pose. The image synthesis device according to any one of claims 1 to 9.

11. the three-dimensional image data further includes an indication of an image region for at least one three-dimensional image of the three-dimensional image data, and the image region circuit identifies the first image region according to the indication of the image region; The image synthesis device according to any one of claims 1 to 10.

12. the three-dimensional image data further includes an indication of a given region for at least one three-dimensional image of the three-dimensional image data, and the region circuit identifies the first region according to the indication of the given region; The image synthesis device according to any one of claims 1 to 11.

13. the view synthesis circuit selects from a plurality of candidate pixel values ​​derived from different images of the multi-view image of the three-dimensional image data for at least a first pixel of the view image, the view synthesis circuit selecting a rearmost pixel when the distance between the view pose and the first region is greater than a threshold, and selecting a frontmost pixel when the distance between the view pose and the first region is less than the threshold, the rearmost pixel being associated with a depth value indicating a depth furthest from the view pose and the frontmost pixel being associated with a depth value indicating a depth closest to the view pose; The image synthesis device according to any one of claims 1 to 12.

14. three-dimensional image data describing at least a portion of a three-dimensional scene; a data field indicating whether the rendering of the three-dimensional image data must include adapting the transparency for an image region of the image of the three-dimensional image data in the rendered image depending on a depth indication for the image region and a distance between a view pose for the rendered image and a reference region for the image region; An image signal device for generating an image signal, comprising:

15. The image signal further includes at least one of an indication of the image region and the reference region.

15. An image signal device according to claim 14.

16. 1. A method of image synthesis, said method comprising: receiving three-dimensional image data describing at least a portion of a three-dimensional scene; identifying at least a first image region in the three-dimensional image data; determining a depth indication for the first image region from depth data of the three-dimensional image data for the first image region; identifying a first region for the first image region; receiving a view pose for a viewer; generating a view image from the three-dimensional image data, the view image representing a view of the three-dimensional scene from the view pose, the generating the view image comprising adapting a transparency for the first image region in the view image in response to the depth indication and a distance between the view pose and the first region, the transparency increasing as the distance between the view pose and the first region increases and as the depth indication indicates a decreasing depth relative to the first image region; A method comprising:

17. The program comprises computer program code means for performing all the steps of the method according to claim 16 when the program is run on a computer. Computer program.

Citation Information

Patent Citations

  • Display control program, display control device, display control method, and display control system

    JP2012123337A

  • Image processing device, image processing method, and program

    JP2019144638A

  • Image generation device, image generation method, image generation system, and program

    JP2020135290A