Image Generation
Patent Information
- Application Number
- JP2023580771
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-29
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-24
AI Technical Summary
Existing immersive video systems face limitations in viewing space with noticeable degradations and errors in image quality when viewers move outside the sweet spot, due to insufficient 3D data for field of view synthesis, leading to a suboptimal user experience.
A video rendering device that combines captured video data with a three-dimensional mesh model to generate output images, using different data types for regions based on viewing pose deviations, ensuring high-quality image generation across a wider range of movements.
This approach enhances user experience by reducing image quality degradation and increasing freedom of movement, while requiring fewer cameras and lowering data communication needs, thus providing a more immersive and efficient rendering process.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to image generation approaches, particularly, but not exclusively, to the generation of images for a 3D video signal for different viewpoints. [Background technology]
[0002] Recently, the variety and range of image and video applications has increased significantly, with the continuous development and introduction of new services and methods for using and consuming images and videos.
[0003] For example, one increasingly popular service is one that provides image sequences in such a way that the viewer can actively and dynamically interact with the view, such that the viewer can change their viewing position or direction in the scene so that the presented animation adapts to present the view from the changed position or direction.
[0004] The capture, distribution, and presentation of three-dimensional video has become increasingly popular and attractive in some applications and services. A particular approach is known as immersive video, and typically involves providing a real-world scene that allows for small viewer movements, such as relatively small head movements and rotations, and a view of the event, often in real time. For example, a real-time video broadcast, such as a sporting event, that allows for local client-based view generation following small viewer head movements, provides the user with the impression of sitting in the stands and watching the sporting event. The user will have a natural experience similar to that of a spectator present at that location in the stands, including being able to look around. Recently, display devices with applications that support position tracking and 3D interaction based on 3D capturing of real-world scenes have become more widespread. Such display devices are particularly suitable for immersive video applications that provide an enhanced three-dimensional user experience.
[0005] To provide such services for real-world scenes, the scene is typically captured from different positions and with different camera capture poses. As a result, the relevance and importance of multi-camera capturing and 6DoF (six degrees of freedom) processing, etc., is rapidly increasing. Applications include live concerts, live sports, and telepresence. The freedom to choose one's own viewpoint enriches these applications by increasing the sense of presence over regular video. Furthermore, immersive scenarios can be imagined, where the observer can navigate and interact with the captured live scenes. For broadcast applications, this requires real-time depth estimation at the production side and real-time view synthesis at the client device. Both depth estimation and view synthesis introduce errors, and these errors depend on the implementation details of the algorithms used. In many such applications, 3D scene information is often provided that allows the synthesis of high-quality view images for viewpoints relatively close to a reference viewpoint(s), but degrades the synthesis of high-quality view images if the viewpoint deviates too much from the reference viewpoint.
[0006] For example, a set of mutually offset video cameras capture the scene to provide 3D image data in the form of multiple 2D images from offset positions and / or as image data plus depth data. A rendering device dynamically processes the 3D data to generate images for different changing viewing positions / directions. The rendering device can dynamically shift viewpoints or project, etc. to dynamically follow the user's movements.
[0007] A problem with immersive video, etc., is that the viewing space, the space in which the viewer has a sufficient image quality experience, is limited. As the viewer moves outside the viewing space, degradations and errors resulting from synthesizing the viewing image become more and more noticeable, resulting in an unacceptable user experience. Errors, artifacts, and inaccuracies in the generated viewing image arise in particular due to the fact that 3D video data is provided that does not provide enough information (e.g., non-occlusion data) for viewing synthesis.
[0008] For example, typically when multiple cameras are used to capture a 3D representation of a scene, playback in a virtual reality headset tends to be spatially constrained to a virtual viewpoint located near the original camera positions, ensuring that the render quality of the virtual viewpoint does not exhibit artifacts, typically the result of missing information (occluded data) or 3D estimation errors.
[0009] Within the so-called sweet spot viewing region, rendering can be done directly from one or more reference camera images with associated depth maps or meshes using standard texture mapping combined with view blending.
[0010] Outside the sweet spot viewing area, image quality degrades, often to an unacceptable extent. In current applications, this is addressed by providing the viewer with a blurred or even black picture for parts of the scene that cannot be rendered accurately enough. However, such approaches tend to be suboptimal and tend to provide a suboptimal user experience. EP 3422711 A1 discloses an example of a rendering system in which blur is introduced to bias the user away from parts of the scene that are not represented by an incomplete representation of the scene. Summary of the Invention [Problem to be solved by the invention]
[0011] Thus, improved approaches would be advantageous, particularly approaches that allow for improved operation, increased flexibility, an improved immersive user experience, reduced complexity, facilitated implementation, improved perceived and synthesized image quality, improved rendering, increased (possibly virtual) freedom of movement for the user, improved user experience, and / or improved performance and / or operation.
[0012] SUMMARY OF THE DISCLOSURE Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination. [Means for solving the problem]
[0013] According to one aspect of the present invention, an apparatus is provided comprising: a first receiver configured to receive video data for a real world scene, the video data being captured and linked to a capture pose area; a store configured to store a three-dimensional mesh model of at least a portion of the real world scene; a second receiver configured to receive a viewing pose; and a renderer configured to generate an output image for a viewport for the viewing pose, the renderer comprising: a first circuit configured to generate first image data for the viewport for at least a portion of the output image by projecting the captured video data onto the viewing pose; a second circuit configured to generate second image data for the output viewport for at least a first region of the output image from the three-dimensional mesh model; a third circuit configured to generate the output image including at least a portion of the first image data and the second image data for the first region; and a fourth circuit configured to determine the first region in response to a deviation of the viewing pose from the capture pose area.
[0014] The present invention provides an improved user experience in many embodiments and scenarios. It allows for an improved tradeoff between image quality and freedom of movement for many applications. This approach often provides a more immersive user experience and is particularly suitable for immersive video applications. This approach reduces the degradation of perceived image quality for different viewing poses. This approach provides, for example, the user with an improved experience with respect to a larger range of position and / or orientation changes. In many embodiments, this approach relaxes the requirements for the capture of real-world scenes. For example, fewer cameras are used. The requirements on how much of the scene is captured are relaxed. This approach in many embodiments relaxes the requirements for data communication, allowing, for example, lower latency interactive services.
[0015] This approach allows, for example, for an improved immersive video experience.
[0016] A pose is a position and / or orientation. A pose region is a consecutive set of poses. A capture pose region is a region where the captured video data provides data that allows image data to be generated having a quality that meets a quality criterion. An output image is an image of an image sequence, in particular a frame / image of a video sequence.
[0017] The 3D mesh model further includes at least one pixel map having pixel values linked to vertices of the 3D mesh of the 3D mesh model.
[0018] According to an optional feature of the invention, the renderer is configured to determine the first region as a region in which a quality of the first image data generated by the first circuit does not meet a quality standard.
[0019] In some embodiments, the renderer is configured to determine an intermediate image comprising the first image data and to determine the first region as a region where the quality of the image data of the intermediate image does not meet a quality standard.
[0020] This provides improved and / or facilitated operation in many embodiments, and it provides a particularly efficient approach for determining first regions that are particularly suitable for providing an engaging user experience.
[0021] According to an optional feature of the invention, the third circuit is configured to determine the first region in response to a difference between the viewing pause and the capture pause regions.
[0022] This provides improved and / or facilitated operation in many embodiments, and it provides a particularly efficient approach for determining first regions that are particularly suitable for providing an engaging user experience.
[0023] In many embodiments, the third circuit is configured to determine the first region as a function of a distance between the viewing pose and the capture pose region, the distance being determined according to a suitable distance measure, the distance measure reflecting a position and / or orientation of the viewing pose relative to the capture pose region.
[0024] According to an optional feature of the invention, the difference is an angular difference.
[0025] This provides improved and / or facilitated operation in many embodiments.
[0026] According to an optional feature of the invention, the renderer is configured to adapt the second image data in response to the captured video data.
[0027] This provides an improved user experience in many embodiments, as it provides a more consistent and consistently generated output image in many scenarios and reduces the perceived visibility of differences between portions of the output image generated from the video data and portions of the output image generated from the 3D mesh model.
[0028] According to an optional feature of the invention, the renderer is configured to adapt the first image data in response to the three-dimensional mesh data.
[0029] This provides an improved user experience in many embodiments, as it provides a more consistent and consistently generated output image in many scenarios and reduces the perceived visibility of differences between portions of the output image generated from the video data and portions of the output image generated from the 3D mesh model.
[0030] According to an optional feature of the invention, the renderer is configured to adapt the second image data in dependence on the first image data.
[0031] This provides an improved user experience in many embodiments, as it provides a more consistent and consistently generated output image in many scenarios and reduces the perceived visibility of differences between portions of the output image generated from the video data and portions of the output image generated from the 3D mesh model.
[0032] According to an optional feature of the invention, the renderer is configured to adapt the first image data in response to the second image data.
[0033] This provides an improved user experience in many embodiments, as it provides a more consistent and consistently generated output image in many scenarios and reduces the perceived visibility of differences between portions of the output image generated from the video data and portions of the output image generated from the 3D mesh model.
[0034] According to an optional feature of the invention, the renderer is configured to adapt the three-dimensional mesh model in response to the first image data.
[0035] This provides an improved user experience in many embodiments, as it provides a more consistent and consistently generated output image in many scenarios and reduces the perceived visibility of differences between portions of the output image generated from the video data and portions of the output image generated from the 3D mesh model.
[0036] According to an optional feature of the invention, the apparatus further comprises a model generator for generating a three-dimensional mesh model in response to the captured video data.
[0037] This provides for improved operation and / or easier implementation in many embodiments.
[0038] According to an optional feature of the invention, the first receiver is configured to receive video data from a remote source and further receive a three-dimensional mesh model from the remote source.
[0039] This provides for improved operation and / or easier implementation in many embodiments.
[0040] According to an optional feature of the invention, the second circuitry is configured to vary a level of detail for the first region in response to a deviation of the viewing pose relative to the capture pose region.
[0041] This, in many embodiments, provides a further improved user experience and provides improved perceptual adaptation to viewer pose changes.
[0042] According to an optional feature of the invention, the first receiver is further configured to receive second captured video data for the real world scene, the second video data being linked to a second capture pose area, and the first circuit is further configured to determine third image data for at least a portion of the output image by projecting the second captured video data onto the viewing pose, and the third circuit is configured to determine the first area in response to a deviation of the viewing pose from the second capture pose area.
[0043] This provides an enhanced user experience in many scenarios and embodiments.
[0044] According to one aspect of the present invention, there is provided a method comprising the steps of receiving captured video data for a real-world scene, the video data being linked to a capture pose area; storing a 3D mesh model of at least a portion of the real-world scene; receiving a viewing pose; and generating an output image for a viewport for the viewing pose, the generating the output image comprising the steps of generating first image data for the viewport for at least a portion of the output image by projecting the captured video data onto the viewing pose; generating second image data for the output viewport for at least a first region of the output image from the 3D mesh model; generating the output image to include at least a portion of the first image data and the second image data for the first region; and determining the first region in response to a deviation of the viewing pose from the capture pose area.
[0045] These and other aspects, features and advantages of the present invention will be apparent from and elucidated with reference to the embodiments described hereinafter.
[0046] Implementations of the present invention will now be described, by way of example only, with reference to the following drawings in which: [Brief description of the drawings]
[0047] [Figure 1] 1 is an illustration of an example of elements of a video distribution system according to some embodiments of the present invention. [Diagram 2] 1 illustrates an example of capturing a 3D scene. [Diagram 3] 1 illustrates an example of a field of view generated for a particular viewing pose. [Figure 4] 1 illustrates an example of a field of view generated for a particular viewing pose. [Diagram 5] 1 illustrates an example of a field of view generated for a particular viewing pose. [Figure 6] 1 is an illustration of an example of elements of a video rendering device according to some embodiments of the present invention. [Figure 7] 1 illustrates an example of a field of view generated for a particular viewing pose. [Figure 8] 1 illustrates an example of a field of view generated for a particular viewing pose. [Figure 9] 1 illustrates an example of capturing a 3D scene using two sets of capture cameras. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0048] The following description focuses on immersive video applications, but it will be understood that the principles and concepts described can be used in many other applications and embodiments.
[0049] In many approaches, immersive video is provided locally to the viewer, for example by a stand-alone device without using or even having any access to any remote video server. In other applications, however, the immersive application may be based on data received from a remote or central server. For example, video data is provided from a remote central server to a video rendering device and processed locally to generate the desired immersive video experience.
[0050] 1 illustrates such an example of an immersive video system in which a video rendering device 101 communicates with a remote immersive video server 103 via a network 105, e.g., the Internet. The server 103 is configured to support a potentially large number of client video rendering devices 101 simultaneously.
[0051] Immersive Video Server 103 supports immersive video experiences, for example, by transmitting three-dimensional video data describing a real-world scene, specifically describing the visual features and geometric properties of the scene as generated from real-time capture of the real world by a set of (possibly 3D) cameras.
[0052] 2, a set of cameras may be arranged with appropriate capture settings (e.g., linear, etc.) and offset individually, each capturing an image of the scene 203. The captured data may be used to generate a 3D video data stream that is transmitted from the immersive video server 103 to a remote video rendering device.
[0053] The 3D video data may be, for example, a video stream, and may include, for example, directly captured images from multiple cameras, and / or may include processed data such as images and depth data generated from the captured images. Many techniques and approaches for generating 3D video data are known, and it will be appreciated that any suitable approach and 3D video data format / representation may be used without detracting from the present invention.
[0054] The immersive video rendering device 101 is configured to receive 3D video data and process the received 3D video data to generate an output video stream, where the generated output video stream dynamically reflects changes in a user's pose, thereby providing an immersive video experience in which the presented views adapt to changes in field of view / user pose / position.
[0055] In this field, the terms configuration and pose are used as general terms for position and / or orientation. For example, a combination of position and orientation of an object, camera, head, or view is referred to as a pose or configuration. Thus, a configuration or pose indication contains six values / components / degrees of freedom, each value / component typically describing an individual property of the position / location or orientation / direction of the corresponding object. Of course, in many situations, a configuration or pose will be considered or represented using fewer components, for example when one or more components are considered to be fixed or meaningless (for example, four components provide a complete representation of the pose of an object when all objects are considered to be at the same height and have a horizontal orientation). In the following, the term pose is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom). The term pose can be interchanged with the term configuration. The term pose can be interchanged with the term position and / or orientation. The term pose can be interchanged with the terms position and orientation (when the pose provides both position and orientation information), position (when the pose provides information (perhaps only) about position), and orientation (when the pose provides information (perhaps only) about orientation).
[0056] The quality of the generated view images depends on the image and depth information available to the view synthesis operation, which in turn depends on the amount of reprojection and view shifting required.
[0057] For example, view shifting typically results in the de-occlusion of parts of an image that cannot be seen, for example, in the main image used for view shifting. Such holes are filled by data from other images if they capture the de-occluded object, but it is also typically possible that parts of the image that are de-occluded for the new viewpoint are also missing from the other source views. In such cases, view synthesis needs to estimate data, for example, based on surrounding data. The process of de-occlusion is inherently prone to being a process that introduces inaccuracies, artifacts, and errors. Moreover, this tends to increase with the amount of view shift, and in particular, during view synthesis, the probability of missing data (holes) increases with increasing distance from the image capture pose.
[0058] Another source of possible distortion is imperfect depth information. Often, depth information is provided by a depth map where depth values are generated by depth estimation (e.g., disparity estimation between source images) or measurement (e.g., ranging), but this is not perfect, so the depth values contain errors and inaccuracies. View shifting is based on depth information, and imperfect depth information causes errors or inaccuracies in the synthesized image. The further the synthesized viewpoint is from the viewpoint of the original camera, the more severe the distortion in the synthesized target view image becomes.
[0059] Thus, as the viewing pose moves further and further from the capture pose, the quality of the synthesized image tends to deteriorate: if the viewing pose is far enough away from the capture pose, the image quality will degrade to an unacceptable degree, resulting in a poor user experience.
[0060] Figures 3 to 5 illustrate the problems associated with moving away from the capture pose. Figure 3 illustrates an example where the synthesized viewport is closely aligned with the capture camera's viewport so that a specific image for the viewing pose viewport can be predicted from the capture camera using depth image-based rendering, resulting in a high quality image. In contrast, in the examples of Figures 4 and 5, the viewing pose and the capture pose differ only by the angular orientation of the viewport, which differs from the capture viewport. As illustrated, as a result of the angular change in the viewing direction, a large portion of the image (the right or left side of the image in this example) is not provided with adequate image data. Furthermore, while information for extrapolation from the image data to unknown ranges may provide some improved perception, as illustrated, it results in very significant degradation and distortion, leading to an unrealistic representation of the scene.
[0061] The viewing pose and the capture pose differ only by deviations in the position and / or angle of the field of view, and their effects are different. Position changes, such as translation, tend to increase the disocclusion range behind foreground objects, increasing the unreliability of view synthesis due to uncertainties in 3D (depth / geometry) estimation. An angular change in viewpoint, rotating away from the capture camera angle, can result in a situation where image data is not available for a large area of the new viewport, for example (as illustrated in Figures 4 and 5).
[0062] The above problems result in a poor immersive effect because the entire field of view of the display (e.g., typically 110 degrees) is filled and head rotation does not introduce new content. Also, spatial content is often lost and it can be more difficult to navigate when the image is blurred or of otherwise poor quality. Several different approaches have been proposed to address these problems, but they tend to be suboptimal and, in particular, undesirably restrict user movement or result in undesirable user effects.
[0063] 6 illustrates a video rendering apparatus / system / device that provides performance and approaches that can achieve a more desirable user experience in many scenarios. This apparatus may specifically be the video rendering device 101 of FIG.
[0064] The video rendering device comprises a first receiver 601 configured to receive captured video data for a real-world scene. In this example, the video data is provided by a video server 103.
[0065] The video data is captured video data of a real-world scene, typically three-dimensional video data generated from capture of the scene by multiple mutually offset cameras. The video data may be, for example, multiple video streams from different cameras, or video data for one or multiple capture positions with depth information. It will be appreciated that many different approaches are known for capturing video data of a real-world scene, generating (three-dimensional) video data representative of the capture, and communicating / distributing the video data, and that any suitable approach may be used without detracting from the value of the present invention.
[0066] In many embodiments, the 3D video data includes images of multiple views, and thus multiple (simultaneous) images of a scene from different viewpoints. In many embodiments, the 3D video data has the form of an image and depth map representation, where an image / frame is provided with an associated depth map. The 3D image data is in particular a multiple view plus depth representation, and includes for each frame at least two images from different viewpoints, at least one of these images having an associated depth map. If the received data is, for example, a multiple view data representation without an explicit depth map, a depth map can be generated using a suitable depth estimation algorithm, in particular a disparity estimation based approach using different images of the multiple view representation.
[0067] In this particular example, a first receiver 601 receives MVD (Multiple Views and Depth) video data describing a 3D scene using a sequence of multiple simultaneous images and a depth map, also referred to below as source images and source depth map, it will be understood that for a video stream a sequence of such 3D images is provided in time.
[0068] The received video data is linked to a capture pose region, which is typically a region of the scene proximate to the capture pose, typically including the capture pose. The capture pose region is a range of intervals for one, several or all parameters representing the capture pose and / or the view pose. For example, if the pose is represented by a two-dimensional position, the capture pose region is represented by a corresponding range of two positions, i.e. as a two-dimensional range. In other embodiments, the pose is represented by six parameters, typically three position parameters and three orientation parameters, in which case the capture pose region is given by the limits on these six parameters, i.e. by a full 6DoF representation of the pose.
[0069] In some examples, the capture pose region is a single capture pose that corresponds to a single pose that corresponds to the viewport (position and orientation of the field of view) for which the captured video data is being provided. The capture pose region may be a set of poses that indicate / include one or more poses in which the scene was captured.
[0070] In some embodiments, the capture pause area is provided directly from the source of the video data, specifically, it is included in the received video data stream. In some embodiments, it is provided specifically as metadata in the video data stream. In the example of Figure 2, the video data communicated to the video rendering device 101 is provided based on a camera array 205 that is positioned within the capture pause area 205.
[0071] The video rendering device, in some embodiments, is configured to use the capture pose region directly as received, while in other embodiments the video rendering device may be configured to modify the capture pose region or may itself generate the capture pose region.
[0072] For example, in some embodiments, the received data only includes video data corresponding to a given capture pose, without any indication of the capture pose itself, of what magnified areas, or of how the image data is suitable for viewing composites for poses other than the given capture pose. In such a case, the receiver 601 will, for example, generate a capture pose area based on the received capture pose. For example, it considers that the provided video data is linked to a reference pose, so that the video data is directly rendered to this reference pose without any view shifting or projection. Then, all poses are measured relative to this reference pose, and the capture pose area is determined as the reference pose or as a pre-determined area, for example centered on the reference pose. As the user moves, the viewing pose is then represented / measured relative to this reference pose.
[0073] In some embodiments, the capture pose region is considered to simply correspond to a single pose, e.g., that of the received video data. In other embodiments, the receiver 401 will generate an extended capture pose region, e.g., by performing an assessment of the quality degradation as a function of the difference from or distance to the capture pose. For example, for various test poses that deviate from the capture pose region by different amounts, the first receiver 601 will evaluate how large a proportion of the corresponding viewport is covered by image data and how large a proportion corresponds, e.g., to de-occluded areas / objects or to areas / objects for which no data is provided, due to the viewport extending into parts of the scene not covered by the capture camera. The capture pose region is determined, e.g., as a six-dimensional region for which the proportion of the corresponding viewport not covered by image data is below a given threshold. It will be appreciated that many other approaches to assess the quality level or degradation as a function of the deviation between the capture pose and the viewing pose are possible and any suitable operation may be used.
[0074] As another example, the first receiver 601 may, for example, modify the capture pose region to a region including all poses whose distance to the nearest capture pose is less than a given threshold, for example to the nearest camera pose if multiple camera poses are provided, or to the nearest pose of the received capture pose region for which video images are provided. The distance may be determined according to any suitable distance measure, possibly including consideration of both positional distance and angular (orientation) distance.
[0075] It will be understood that other approaches to determining the capture pose region are used in other embodiments, and the particular approach to determining the capture pose region reflecting a set of poses for which an image can be generated with suitable quality will depend on the requirements and preferences of the particular embodiment.
[0076] The video rendering apparatus of Figure 6 further comprises a second receiver 603 configured to receive a viewing pose for a viewer (and in particular in a three-dimensional scene). The viewing pose represents the position and / or orientation from which the viewer views the scene, and specifically provides the pose in which a view of the scene should be generated.
[0077] It will be appreciated that many different approaches for determining and providing a viewing pose are known and any suitable approach may be used. For example, the second receiver 603 may be configured to receive pose data from a VR headset worn by a user, an eye tracker, etc. In some embodiments, a relative viewing pose is determined (e.g., a change from an initial pose is determined), which may be relative to a reference pose, such as the camera pose or the center of a capture pose area.
[0078] The first and second receivers 601, 603 may be implemented in any suitable manner and may receive data from any suitable source including local memory, a network connection, a wireless connection, a data medium, etc.
[0079] These receivers are implemented as one or more integrated circuits, such as application specific integrated circuits (ASICs). In some embodiments, these receivers are implemented as one or more programmed processing devices, such as firmware or software running on a suitable processor, such as a central processing unit, digital signal processor, or microcontroller. In such embodiments, it will be understood that the processing device includes on-board or external memory, clock driving circuits, interface circuits, user interface circuits, etc. These circuits may further be implemented as part of the processing device, as integrated circuits, and / or as discrete electronic circuits.
[0080] The first and second receivers 601, 603 are coupled to a view synthesis or projection circuit or renderer 605 configured to generate view frames / images from the received three-dimensional video data, where view images are generated to represent a view of the three-dimensional scene from a viewing pose. The renderer 605 thus generates a video stream of view images / frames for the 3D scene from the received video data and the viewing pose. In the following, the operation of the renderer 605 is described with reference to the generation of a single image. However, it will be understood that in many embodiments the image is part of a sequence of images, in particular a frame of a video sequence. In practice, the described approach applies to multiple frames / images, often to the entire frame / image of an output video sequence.
[0081] It will be appreciated that in many cases a stereo video sequence will be generated that includes a video sequence for the right eye and a video sequence for the left eye, so that when the images are presented to a user, for example via an AR / VR headset, it appears as if the 3D scene is being viewed from a viewing pose.
[0082] The renderer 605 is typically configured to perform field shifting or projection of the received video images based on the depth information, which typically includes techniques such as pixel shifting (changing pixel positions to reflect appropriate disparities corresponding to parallax changes), occlusion removal (typically based on filling from other images), combining pixels from different images, etc., as known to those skilled in the art.
[0083] It will be appreciated that many algorithms and approaches are known for image compositing, and any suitable approach may be used by the renderer 605.
[0084] The image synthesizer thus generates a view image / video for the scene. Moreover, as the viewing pose changes dynamically in response to the user moving around in the scene, the view of the scene is continuously updated to reflect the change in viewing pose. For static scenes, the same source view image is used to generate the output view image, but for video applications, different source images are used to generate different view images, e.g., a new set of source images and depths is received for each output image. Thus, the processing is frame-based.
[0085] The renderer 605 is configured to generate views of the scene from different angles for lateral movement of the viewing pose. When the viewing pose changes to assume a different direction / orientation, the renderer 605 is configured to generate views of the three-dimensional scene objects from different angles. Thus, as the viewing pose changes, the objects in the scene may be perceived as static and having a fixed orientation in the scene. The viewer can effectively move and view the objects from different directions.
[0086] The view synthesis circuitry 205 may be implemented in any suitable manner, including as one or more integrated circuits, such as an application specific integrated circuit (ASIC). In some embodiments, these receivers are implemented as one or more programmed processing devices, such as firmware or software running on a suitable processor, such as a central processing unit, digital signal processor, or microcontroller. In such embodiments, it will be appreciated that the processing device includes on-board or external memory, clock driving circuits, interface circuits, user interface circuits, etc. These circuits may further be implemented as part of the processing device, as integrated circuits, and / or as discrete electronic circuits.
[0087] As mentioned above, a problem with view synthesis is that the quality decreases as the viewing pose from which the view is synthesized becomes increasingly different from the capture pose of the video data of the scene being provided. In fact, if the viewing pose strays too far from the capture pose region, the resulting image will be unacceptable with significant artifacts and errors.
[0088] The video rendering device further comprises a store 615 for storing a three-dimensional mesh model of at least a portion of the real-world scene.
[0089] A mesh model provides a three-dimensional description of at least a portion of a scene. A mesh model is composed of a set of vertices interconnected by edges that create faces. A mesh model provides a number of faces, e.g., triangles or rectangles, that provide a three-dimensional representation of the elements of the scene. Typically, a mesh is described by the three-dimensional positions of the vertices, for example.
[0090] In many embodiments, the mesh model also includes texture data and texture information because the mesh is provided to indicate a texture for the faces of the mesh. In many embodiments, the 3D mesh model includes at least one pixel map having pixel values that are linked to the vertices of the 3D mesh of the 3D mesh model.
[0091] A mesh model of a real-world scene provides an accurate yet realistic representation of the three-dimensional information of the scene that is used in a video rendering device to provide improved image data for viewing poses that differ by large angles from the capture pose region.
[0092] The mesh model, in many embodiments, provides a static representation of the scene, and in many embodiments the video signal provides a dynamic (typically real-time) representation of the scene.
[0093] For example, the scene may be a football pitch or stadium, and a model is generated to represent the permanent parts of the scene, such as the pitch, the goals, the lines, the stands, etc. The video data provided is a capture of a particular match, and includes the dynamic elements, such as players, coaches, spectators, etc.
[0094] The renderer 605 comprises a first circuit 607 configured to determine image data for at least a portion of an output image by projection of the received video data onto a viewing pose. The first circuit 607 is thus configured to generate image data for a viewpoint of the current viewing pose from the received video data. To generate image data for a viewport of the viewing pose, the first circuit 607 applies any suitable view shifting and reprojection process, in particular generating a complete or partial intermediate image corresponding to a current viewport (which is a viewport for the current viewing pose). The projection / view shifting is from a capture pose of the video data, in particular a projection from a capture pose of one or more capture cameras onto the current viewing pose. As described above, any suitable approach may be used, including techniques such as parallax shifting, occlusion removal, etc.
[0095] The renderer 605 further comprises a second circuit 609 configured to determine second image data for an output viewport for at least the first region in response to the three-dimensional mesh model. The second circuit 609 is thus configured to generate image data for a viewport for a current viewing pose from the stored mesh model, typically taking into account texture information. The second circuit 609 applies any suitable approach for generating image data from the mesh model for a given viewing pose, including using a technique for mapping vertices to image positions in the output image depending on the viewer's pose, filling in ranges based on vertex positions and textures, etc. The second circuit 609 specifically generates a second intermediate image corresponding to the viewport for the current viewing pose. This second intermediate image is a partial image and includes image data for only one or more regions of the viewport.
[0096] It will be appreciated that many different approaches, algorithms, and techniques are known for synthesizing image data from captured image data and 3D data, including from 3D mesh models, and any suitable approach and algorithm may be used without detracting from the present invention.
[0097] Examples of suitable view synthesis algorithms include, for example: “A review on image-based rendering” Yuan HANG,Guo-Ping ANG Virtual Reality & Intelligent Hardware,Volume 1,Issue 1,February 2019,Pages 39-54 https: / / doi.org / 10.3724 / SP.J.2096-5796.2018.0004 or “A Review of Image-Based Rendering Techniques” Shum; Kang Proceedings of SPIE - The International Society for Optical Engineering 4067:2-13, May 2000 DOI:10.1117 / 12.386541 Or, for example, the Wikipedia article on 3D rendering: https: / / en.wikipedia.org / wiki / 3D_rendering can be found in.
[0098] The renderer 605 thus generates image data for the current viewpoint in two separate ways: based on the received video data and based on a stored mesh model.
[0099] The renderer 605 further comprises a third circuit 611 configured to generate an output image to include the first image data and the second image data. In particular, for at least the first region, the output image is generated to include the second image data generated from the mesh model, and for at least a portion of the output image outside the first region, the output image is generated to include the first image data generated from the video signal.
[0100] In many scenarios, an output image is generated that includes first image data for all ranges where the resulting image quality is considered to be sufficiently high, and second image data is included for ranges where the image quality is not considered to be sufficiently high.
[0101] The renderer 605 comprises a fourth circuit 613 configured to determine one or more regions of the output image where second image data should be used, i.e. image data generated from the mesh model rather than from the video data should be included in the output image. The fourth circuit 613 is configured to determine a first such region in response to a deviation of the viewing pose from the capture pose region. Thus, the renderer 605 is configured to determine a region of the output image where the video-based image data is replaced by the model-based image data, which region depends on the viewing pose and how different it is from the capture pose region.
[0102] In some embodiments, the fourth circuit 613 is configured to determine the first region depending on the difference between the viewing pause and the capture pause regions. For example, if the distance between them is less than a given threshold (according to a suitable distance measure), no region is defined, i.e. the entire output image is generated from the received video data. However, if this distance is greater than the threshold, the fourth circuit 613 determines a region that is likely to be considered of insufficient quality and controls the second circuit 609 to use the second image data for this region. This region is determined, for example, based on the direction of change (typically in the 6 DoF space).
[0103] For example, a video rendering device may be configured to model the scene using a graphics package, and the graphics model may be rendered onto a viewport after the capture-guided composite image such that this data is replaced by a generated model in one or more regions when the difference between the viewing pose and the capture pose regions is too great.
[0104] As a specific example, the fourth circuit 613 is configured to take into account the horizontal angular orientation of the viewing pose (reflecting the viewer rotating his / her head). As long as the viewing pose reflects a horizontal angular rotation that is less than a given threshold angle, the output image corresponding to the viewport of the viewing pose is generated based exclusively on the video data. However, if the viewing pose indicates an angular rotation above this threshold, the fourth circuit 613 determines that there is a left or right region of the image that will be occupied by the second image data. Whether this region is on the left or right side of the output image depends on the direction of rotation indicated by the viewing pose (i.e., whether the viewer rotates his / her head to the left or right) and the size of the region, which depends on how large the angular rotation is. Figures 7 and 8 show examples of how this approach improves the images of Figures 4 and 5.
[0105] If the viewing pose moves too far from the capture pose area, the image quality of the synthesized view will degrade. In this case, the user experience is significantly improved, instead of providing low quality or blurry data, which is typically generated, for example, by evaluating a static graphics model of the scene. This, among other things, provides the viewer with improved spatial content about his / her presence in the scene.
[0106] It should be noted that in a typical practical system, it is desirable to be able to use a capture camera with a limited field of view, since this allows more distant objects to be captured with higher resolution for a given sensor resolution. To obtain the same resolution using a wide-angle lens, for example 180 degrees, would require a sensor with a very high resolution, which is not always practical, since such a sensor would be more costly in terms of camera and processing hardware, and would be more demanding on resources in terms of processing and communication.
[0107] As discussed above, in some embodiments, the video rendering device determines regions for which model-based image data is used, specifically whether such regions should be included based on the distance between the viewing pose and the capture pose region. In some embodiments, the determination of the region based on the deviation from the viewing pose to the capture pose region is based on considering the effect of the deviation on the quality of image data that can be synthesized for the viewing pose using the video data.
[0108] In some embodiments, the first circuit 607 generates the intermediate image based on projecting the received video data from an appropriate capture pose onto the viewing pose.
[0109] The fourth circuit 613 then proceeds to evaluate the resulting intermediate images and in particular to determine quality measures for different parts / areas / regions of the image. The quality measures are determined, for example, based on the algorithm or process used to generate the image data. For example, image data that can be generated by disparity shifting is assigned a high quality value, which is further graded depending on how large the shift is (for example, in case of a remote background, the disparity shift is zero, so that, for example, in the disparity estimation, it is not sensitive to errors and noise). Image data that is generated by extrapolation from other image data to a de-occluded region is assigned a lower quality value, which is further graded depending on how much extrapolation of data is required, the degree of texture variation in neighboring regions, etc.
[0110] A fourth circuit 613 then evaluates the determined quality measure to determine one or more regions whose quality does not meet the quality criterion. A simple criterion is simply to determine the regions as areas where the quality criterion is lower than a threshold. More complex criteria include, for example, requirements on a minimum size or shape of the regions.
[0111] The second circuit 609 then proceeds to generate an output image as a combination of the video-based (synthesized) image data from the intermediate image and the model-based image data. For example, the output image is generated by overwriting the image data of the intermediate video-based image with the model-based image data in areas determined by the fourth circuit 613 to not have sufficient image quality.
[0112] It will be appreciated that in general, a number of different approaches to assessing quality are used.
[0113] For example, depth quality may be determined for different reasons and regions using model data may be determined based on depth quality, specifically image regions generated using depth data deemed to have a quality below a threshold.
[0114] To explicitly determine the depth data, a reprojection error can be calculated (either at the encoder or decoder side). This means that a view from the image data, in particular a multi-view dataset, is reprojected (with depth) from the set of multi-views to another known view. A color difference measure (per pixel or averaged over a region) can then be used as an indication of quality. Occlusion / de-occlusion, although undesirable, affects this error calculation. This is avoided by only accumulating the error in the metric when the absolute difference between the pixel's depth and the warped depth is below a threshold. Such a process is used, for example, to identify depth data that is not considered sufficiently reliable. When generating a new image for any desired viewpoint, the regions that would result from the use of such unreliable depth data are identified and overwritten by image data generated from the model.
[0115] In some cases, a small overall warp error is not a sufficient indication of the rendering quality for any new viewpoint. For example, when any new viewpoint is close to the original capture viewpoint, such as close to the center of the viewing area, the rendering quality typically results in a relatively high quality even if the depth quality of the depth data used is relatively low. Thus, the region is determined by considering the depth quality and identifying regions that result from low quality depth data, but also depends on other parameters, such as how large a shift is made (and in particular on the distance between the viewpoint where the image is generated and the capture pose region defined for that image data).
[0116] Another way to determine the rendering quality to a given viewpoint is to compare image feature statistics of the synthesized image for that viewpoint with image feature statistics of one or more reference images. A relevant statistic is, for example, curvature. Curvature can be calculated directly for one of the color channels or during summation with a local filter window. Alternatively, edge / contour detection can be used first, after which curvature statistics can be calculated. Statistics can be calculated over a given region in the synthesized view. This region can then be warped to one or more reference views and compared with the statistics found in the region there. Because a (larger) region is used, the evaluation becomes less dependent on strict pixel correspondence. Instead of physically meaningful features such as curvature, deep neural nets can be used to calculate view-invariant quality features based on multiple reference views. Such an approach can be applied and evaluated in regions, allowing low-quality regions to be determined.
[0117] In some cases, so-called "reference-free" metrics are used to evaluate the quality of the synthesized view without any reference. A neural network is typically trained to predict image quality.
[0118] Such quality rate is determined without explicitly determining the deviation or difference between the viewing pause and the capture pause region (i.e., such determination is indirect in the quality measurement reflecting the deviation of the viewing pause from the capture pause region).
[0119] As mentioned above, a video rendering device stores a mesh model of a scene, and typically also stores a pixel map with pixel values linked to the vertices of a 3D mesh of the 3D mesh model. A pixel map is specifically a map that indicates visual properties (intensity, color, texture) with a mapping that links a mesh to a part of the pixel map that reflects the local visual properties. The pixel map may specifically be a texture map, and the model of the scene may be a mesh plus a texture model and a representation.
[0120] In some embodiments, the server 103 is configured to transmit the model information to the video rendering device, and thus the first receiver 601 is configured to receive the model data from the server 103. In some embodiments, the model data is combined with the video data into a single data stream, and the first receiver 601 is configured to store the data locally once it is received. In some embodiments, the model data is received independently of the video data, for example at a different time and / or from a different source.
[0121] In some embodiments, the video rendering device is configured to locally generate the model, in particular configured to generate the model from the received video data. The video rendering device specifically comprises a model generator 617 configured to generate a three-dimensional mesh model in response to the captured video data.
[0122] The model generator 617 is equipped with some predefined information (e.g., goals), e.g., an expectation that the scene is a room with some predefined objects therein, etc., and is configured to generate a model by combining and adapting these parameters, e.g., the texture and dimensions of the room are determined based on the received video data, and the positions of predefined objects within the room are determined based on the video data.
[0123] In some embodiments, a (simple) graphics model is inferred from the received multi-view video. For example, flat surfaces like floors, ceilings, walls can be detected and converted into graphics. Ancillary textures can optionally be extracted from the video data. Such inferences do not have to be derived on a frame-by-frame basis, but can be accumulated and improved over time. When presented / rendered to the viewer, such relatively simple visual elements provide a better experience, lacking details but less interesting, as they are not compared to any image or to images with distortions. They keep the viewer immersive and navigable (VR) without feeling disorienting.
[0124] In some embodiments, the model generator is configured to use object detection techniques to recognize objects or people present in the scene, and such objects are represented by pre-existing graphical models or avatars. A pose of the object or body can optionally be determined and applied to the graphical representation.
[0125] It will be appreciated that a variety of techniques and approaches for detecting objects and scene characteristics are known and any suitable approach may be used without detracting from the invention.
[0126] In some embodiments, the mesh model is provided from a remote source, specifically the server 103. In such a case, the server 103 may use, for example, some of the approaches described above.
[0127] In some embodiments, the mesh model is pre-generated and represents the static parts of the scene, as described above. For example, a dedicated capture of the static parts of the second common network element 707 is performed prior to the capture of the event (such as a football match). For example, a camera is moved around the scene to provide images for developing a more accurate mesh model. The development of the model is further based on input from, for example, a dedicated 3D scanner and / or manual adaptation of the model. Such an approach is more laborious but provides a more accurate model. It is particularly useful in the case of events where the same model can be reused for many users and / or events. For example, a lot of effort goes into developing an accurate model of a football stadium, but this can be reused for millions of viewers and for many matches / events.
[0128] In some embodiments, the renderer 605 is configured to adapt the video database processing and / or data in response to the model processing and / or data. Alternatively or additionally, the renderer 605 is configured to adapt the model processing and / or data in response to the video database processing and / or data.
[0129] For example, the mesh model defines the components of the goal, such as the goal posts and the crossbar. The video data includes data for the part of the goal that is visible from the current viewing pose, which is complemented by the mesh model that provides data for the rest of the goal. However, the generated image data is adapted so that the different data match more closely. For example, part of the crossbar is generated from the video data and part of the crossbar is generated from the mesh model. In such an example, the data is adapted to provide a better interface between these parts. For example, the data is adapted so that in the generated output image, the crossbar forms a straight object. This is done, for example, by shifting image data for the crossbar generated from one source so that it matches and has the same orientation as image data from another source for the crossbar. The renderer 605 is configured to adapt the model-based image data to match the received video-based image data, and is configured to match the received video-based image data to match the model-based image data, or to adapt them to match each other.
[0130] In some embodiments, the adaptation is based directly on the generated image data, while in other embodiments, the adaptation is based directly on the mesh model data using a suitable approach. Similarly, in some embodiments, the video rendering device is configured to adapt the mesh model in response to the generated video-based image data. For example, rather than adapting the model-based image data to match the video-based image data, the video rendering device may modify the model, e.g., by moving some vertices until the resulting model-based image data matches the video-based image data.
[0131] Specifically, in some embodiments, the renderer 605 is configured to adapt the generated model-based image data in response to the captured video data. For example, colors from the model-based image may deviate from the actual captured colors. This may be due to (dynamic) circumstances such as lighting or shading conditions or limitations in the accuracy of the model. Thus, the renderer 605 modifies the colors to (closer) match the colors of the captured data.
[0132] As an example of adapting the model-based image, the color distribution may be sampled over the entire image range for both intermediate images, i.e., the video-based intermediate image and the model-based intermediate image. As a result, a single color offset that minimizes the difference in color distribution is adapted to the model-based image. An improvement is to adapt multiple color offsets that are linked to components or clusters in the color distribution. Another improvement is to both sample the distribution and adapt the offsets to specific spatial visual elements (e.g., surfaces).
[0133] In some embodiments, the renderer 605 is configured to adapt the generated video-based image data in response to a three-dimensional mesh model.
[0134] For example, the colors of the generated video-based image may be modified to more closely match those recorded by the mesh model, or the video-based image may be rotated to more closely match the resulting straight lines of the mesh model.
[0135] In some embodiments, the renderer 605 is configured to adapt the generated video-based image data in response to the generated model-based image data.
[0136] For example, the orientation of linear image structures in the model-based image data can be used to correct distortions of the same types of structures in the video-based image data. Specifically, this can be done using filtering operations that use knowledge about the orientation and position of straight lines detected in the model-based image.
[0137] In some embodiments, the renderer 605 is configured to adapt the generated model-based image data in response to the generated video-based image data.
[0138] For example, the example provided above regarding adapting color of a model-based image can also be used to directly modify stored colors (e.g., texture maps) for a model, thereby allowing corrections to be adapted for future images / frames.
[0139] In some embodiments, the renderer 605 is configured to adapt the three-dimensional mesh model in response to the generated video-based image data.
[0140] For example, the positions of the light sources used to illuminate the model can be modified to match the lighting conditions at the stadium (but perhaps without knowledge of the light source positions because they are not available). As another example, the positions of the vertices are adapted to result in a generated model-based intermediate image that matches the video-based image data. For example, different model-based images are generated for slightly perturbed positions of vertices near the transition, and the image that results in the closest match to the video-based image (e.g., most closely aligns the straight lines across the edges) is selected. The positions of the vertices in the mesh model are then modified to the positions for the selected image.
[0141] In some embodiments, the second circuit 609 is configured to vary a level of detail for the first region in response to a deviation of the viewing pose to the capture pose region. In particular, the level of detail is decreased as the difference between the viewing pose and the capture pose region increases. The level of detail is reflected, for example, by the number of object and model features included in the generated image data.
[0142] In some embodiments, the intermediate images are gradually blended into one another.
[0143] In some embodiments, the first receiver 601 is configured to receive further captured video data of the scene for a second capture pose region, for example as illustrated in Figure 9, the scene is captured by two different camera rigs 901, 903 at different positions.
[0144] In such an embodiment, the video rendering device applies a similar approach to both capture pose regions, in particular the first circuit 607 is configured to determine third image data for an output image of the viewport of the current viewing pose based on the video data for the second capture pose. An output image is then generated taking into account the first image data and the second image data. For example, image data is selected based on which of the image data derived from the first capture pose and the second capture pose allows the best combination to be performed.
[0145] In some embodiments, the second circuit 609 simply selects one of the sources on an image-by-image basis (or for groups of images), but in other embodiments, the selection is made separately for different regions, or even for each individual pixel.
[0146] For example, the output image is generated from video data from the closest capture pose region unless this would result in occlusion removal. For these ranges, image data is instead generated from video data from the farthest capture pose region unless this would result in occlusion removal for the pixels in that range.
[0147] In such an approach, the fourth circuit 613 is further configured to generate a first region of the output image, i.e., a region where the output image is dense based on a mesh model, in response to consideration of the viewing pose for both the first and second capture pose regions.
[0148] As a low-complexity example, mesh-model-based data can be used for all regions where the current viewing pose is de-occlusion for both capture pose regions.
[0149] In some embodiments, the capture of a scene may be from two or more separate regions and video data may be provided that is linked to two different capture pose regions, and for a given viewing pose, the video rendering device considers deviations or differences to the different capture pose regions to determine a range of images that may or should be generated based on the mesh model data.
[0150] The following may be provided: a first receiver (601) configured to receive captured video data for a real-world scene, the video data being linked to a capture pause area; a store (615) configured to store a three-dimensional mesh model of at least a portion of a real-world scene; a second receiver (603) configured to receive a viewing pause; and a renderer (605) configured to generate an output image for a viewport for a viewing pose, the renderer (605) comprising: a first circuit (607) configured to generate first image data for a viewport for at least a portion of an output image by projecting the captured video data onto a viewing pose; a second circuit (609) configured to determine second image data for an output viewport for at least a first region of the output image in response to the three-dimensional mesh model; a third circuit (611) configured to generate an output image including at least a portion of the first image data and the second image data for the first region; It is equipped with:
[0151] This device is a fourth circuit (613) configured to determine the first region in response to an image quality measure for the first image data for the first region; a fourth circuit (613) configured to determine an intermediate image comprising the first image data and to determine the first region as a region in which the quality of the image data of the intermediate image does not meet a quality standard; and / or The system may include a fourth circuit (613) configured to determine the first region in response to a quality measure for the first image data.
[0152] The apparatus and / or the fourth circuit may not determine the deviation and / or difference of the viewing pause relative to the capture pause area.
[0153] This approach provides a particularly compelling user experience in many embodiments. As an example, consider a football match captured by a camera rig at the center line and a second camera rig closer to the goal. The viewer takes a viewing pose close to the center line and is presented with a high quality image of the match. The user then decides to virtually travel closer to the goal and, upon arriving at this destination, is provided with a high quality video of the match based on the camera rig positioned closer to the goal. However, in contrast to the traditional approach of teleporting between multiple positions, the user is provided with an experience of a continuous change of position from the center line to the goal (e.g., by emulating a user physically walking between these positions). However, since there may not be enough video data to accurately render the view from a position between the center line and the goal, video data is rendered from model data for at least a portion of the image. This provides an improved and more immersive experience in many scenarios compared to the traditional experience of a user simply teleporting from one position to another.
[0154] The described approach thus generates images for a viewing pose / viewport, the images being generated from two fundamentally different types of data and specifically adaptively generated to include regions generated from these different types of data, namely some regions generated from captured video data of the real-world scene and other regions generated from data of a 3D mesh model for the real-world scene.
[0155] This approach specifically addresses the problem that in many scenarios, capture of real-world scenes is often imperfect, and allows an improved output image / view of the scene to be generated and / or allows for reduced video capture of the real-world scene.
[0156] In contrast to conventional approaches, where images for scene regions for which no captured video data are available are generated by extrapolation of available data, the described approach uses two fundamentally different representations of the scene and combines them in generating a single image: the first type is the captured video data and the second type is a three-dimensional mesh model. In this way, both the captured video data and the data of the 3D mesh model are used. In particular, the data of the mesh model are used to complement the captured video data so that parts of the generated image for which the captured video data does not provide any information can still be presented.
[0157] This approach adaptively combines two fundamentally different types of scene representations to provide improved image quality, and in particular allows image data to be generated for views of the scene about which the captured video data does not contain any information.
[0158] As an example, the described approach allows, for example, an image to be generated for a given viewpoint that includes a portion of a scene for which no captured video data exists, where the generated image even includes features and objects in the scene for which no captured data exists.
[0159] The described approach offers many advantages.
[0160] In particular, images can be generated that provide improved views of real-world scene features for many more viewing poses and scenarios can be achieved for a given capture, such as allowing parts of a scene to be displayed that would not otherwise be possible for a given viewing pose, including the presentation of objects for which the captured video does not contain any data. This approach actually facilitates capture, including allowing fewer cameras to be used for capture while still allowing a large portion of the scene (potentially the entirety) to be seen in some form.
[0161] This approach also reduces the data rate required to communicate video data for a scene: the capture is scaled down to a smaller portion of the scene because it is deemed acceptable to replace parts of the scene by model data (e.g. the playing area of a football pitch is captured in real time by a video camera, whereas the top of the stadium is represented by static 3D mesh model data). Because video data is typically dynamic and real time, it tends to require much higher data rates in practice. The data rate required to represent, for example, the top of a stadium by 3D mesh model data is in practice much lower than if it had to be captured by a video camera and represented by video data.
[0162] This approach typically allows for a significantly improved user experience, including increased freedom: the technical effect is that the restrictions on movement caused by the imperfect capture of video data are reduced (compare, for example, with D1).
[0163] This approach also often results in easier implementation and / or lower complexity and / or reduced computational burden, e.g., reduced encoding / decoding of video capture is achieved and easier rendering is achieved (rendering based on 3D mesh models is typically less complex and more computationally intensive than rendering of captured video).
[0164] The invention can be implemented in any suitable form including hardware, software, firmware or any combination of these. The invention may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of the embodiments of the invention may be physically, functionally and logically implemented in any suitable way. Indeed, the functionality may be implemented in a single unit, in multiple units or as part of other functional units. Thus, the invention may be implemented in a single unit or may be physically and functionally distributed between different units, circuits and processors.
[0165] In this application, any reference to one of the terms "in response," "based on," "according to," and "functionally" should be considered as a reference to the term "in response / based on / in response / functionally." Any of these terms should be considered a disclosure of any of the other terms, and the use of only a single term should be considered as a shorthand concept including the other options / terms.
[0166] Although the present invention has been described in relation to certain embodiments, it is not intended that the present invention be limited to the specific form described herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while certain features may appear to be described in relation to specific embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. The terms "having" and "comprising" in the claims do not exclude the presence of other elements or steps.
[0167] Moreover, even if individually recited, a plurality of means, elements, circuits, or method steps may be implemented by a single circuit, unit, processor, or the like. Additionally, although individual features may be included in different claims, they may be advantageously combined, and the inclusion of the features in different claims does not imply that the combination of the features is infeasible and / or disadvantageous, etc. Furthermore, the inclusion of a feature in one category of claims does not imply a limitation to this category, but rather indicates that the feature is equally applicable to other claim categories, as appropriate. Furthermore, the order of features in the claims does not imply a particular order in which the features must be performed, and in particular the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps may be performed in any suitable order. Additionally, a reference to the singular does not exclude a plural. Thus, a reference to "a", "first", "second", etc. does not exclude a plural. Reference signs in the claims are provided merely as a clarifying example and shall not be construed in any way as limiting the scope of the claims.
[0168] Generally, examples of the apparatus and methods are illustrated by the following embodiments.
[0169] Embodiments: Claim 1. A method for detecting a captured video image of a real-world scene, comprising: a first receiver (601) configured to receive captured video data of a real-world scene, the video data being linked to a capture pause area; and a store (615) configured to store a three-dimensional mesh model of at least a portion of a real-world scene; a second receiver (603) configured to receive a viewing pause; a renderer (605) configured to generate an output image for a viewport for a viewing pose; The apparatus includes: a first circuit (607) configured to generate first image data for a viewport for at least a portion of an output image by projecting the captured video data onto a viewing pose; a second circuit (609) configured to determine second image data for an output viewport for at least a first region of the output image in response to the three-dimensional mesh model; a third circuit (611) configured to generate an output image to include at least a portion of the first image data and the second image data for the first region; a fourth circuit (613) configured to determine the first region in response to a deviation of the viewing pause from the capture pause region; The device comprises:
[0170] Claim 2. The renderer (605) determining an intermediate image including the first image data; The apparatus of claim 1 , configured to determine the first region as a region in which the quality of the image data of the intermediate image does not meet a quality criterion.
[0171] Claim 3. The apparatus of claim 1 or 2, wherein the third circuit (609) is configured to determine the first region in response to a difference between the view pause and capture pause regions.
[0172] Claim 4. The apparatus of claim 3, wherein the difference is an angular difference.
[0173] Claim 5. The apparatus of any one of claims 1 to 4, wherein the renderer (605) is configured to adapt the second image data in response to the captured video data.
[0174] Claim 6. The apparatus of any one of claims 1 to 5, wherein the renderer (605) is configured to adapt the first image data in response to the three-dimensional mesh data.
[0175] Claim 7. The apparatus of any one of claims 1 to 6, wherein the renderer (605) is configured to adapt the second image data in response to the first image data.
[0176] Claim 8. The apparatus of any one of claims 1 to 7, wherein the renderer (605) is configured to adapt the first image data in response to the second image data.
[0177] Claim 9. The apparatus of any one of claims 1 to 8, wherein the renderer (605) is configured to adapt the three-dimensional mesh model in response to the first image data.
[0178] Claim 10. Apparatus according to any one of claims 1 to 9, further comprising a model generator (617) for generating a three-dimensional mesh model in response to the captured video data.
[0179] Claim 11. The apparatus according to any one of claims 1 to 10, wherein the first receiver (601) is configured to receive video data from the remote source (103) and further to receive a 3D mesh model from the remote source (103).
[0180] Claim 12. An apparatus as claimed in any one of claims 1 to 11, wherein the second circuit (609) is configured to vary the level of detail for the first region in response to a deviation of the viewing pause relative to the capture pause region.
[0181] Claim 13. The first receiver (601) is further configured to receive a second captured video data of the real-world scene, the second video data being linked with a second capture pause area; The first circuit (607) is further configured to determine third image data for at least a portion of the output image by projecting the captured second video data onto the viewing pose; 13. An apparatus according to claim 1, wherein the third circuit is configured to determine the first region in response to a deviation of the viewing pause relative to the second capture pause region.
[0182] 14. The method of claim 1, further comprising the steps of: receiving captured video data for a real-world scene, the video data being linked with a capture pause area; storing a 3D mesh model of at least a portion of a real-world scene; receiving a viewing pause; generating an output image for a viewport for a viewing pose; generating an output image comprising: generating first image data for a viewport for at least a portion of an output image by projecting the captured video data onto a viewing pose; determining second image data for an output viewport for at least a first region of the output image in response to the three-dimensional mesh model; generating an output image to include at least a portion of the first image data and the second image data for the first region; determining a first region in response to a deviation of the viewing pose from the capture pose region; The method comprising:
Claims
1. A first receiver that receives captured video data providing a dynamic representation of a real-world scene, wherein the video data is linked to a capture pose region; a first receiver, A store that stores a 3D mesh model providing a static representation of at least a part of the real-world scene; A second receiver that receives a viewing pose; A renderer that generates an output image for a viewport for the viewing pose; An apparatus comprising: wherein the renderer, A first circuit that generates first image data for a viewport for at least a part of the output image by a field-of-view shift of the captured video data from the capture pose of the captured video data to the viewing pose; A second circuit that generates second image data for the viewport for at least a first region of the output image from the 3D mesh model; A third circuit that generates the output image so as to include at least a part of the first image data and the second image data for the first region; A fourth circuit that determines the first region according to a deviation of the viewing pose with respect to the capture pose region An apparatus comprising.
2. The apparatus according to claim 1, wherein the renderer determines the first region as a region where the quality of the first image data generated by the first circuit does not meet a quality standard.
3. The apparatus according to claim 1 or 2, wherein the third circuit determines the first region according to a difference between the viewing pose and the capture pose region.
4. The apparatus according to claim 3, wherein the difference is an angular difference.
5. The apparatus according to any one of claims 1 to 4, wherein the renderer adapts the second image data according to the captured video data.
6. The apparatus according to any one of claims 1 to 5, wherein the renderer adapts the first image data according to the 3D mesh model.
7. The apparatus according to any one of claims 1 to 6, wherein the renderer adapts the second image data according to the first image data.
8. The apparatus according to any one of claims 1 to 7, wherein the renderer adapts the first image data according to the second image data.
9. The renderer adapts the three-dimensional mesh model according to the first image data, the apparatus according to any one of claims 1 to 8.
10. The apparatus according to any one of claims 1 to 9, further comprising a model generator for generating the three-dimensional mesh model according to the captured video data.
11. The first receiver receives the video data from a remote source and further receives the three-dimensional mesh model from the remote source, the apparatus according to any one of claims 1 to 10.
12. The second circuit varies a detail level for the first region according to the deviation of the viewing pose with respect to the capture pose region, the apparatus according to any one of claims 1 to 11.
13. The first receiver further receives captured second video data for the real-world scene, the second video data being linked to a second capture pose region, The first circuit further determines third image data for at least a part of the output image by projecting the captured second video data onto the viewing pose, The third circuit determines the first region according to the deviation of the viewing pose with respect to the second capture pose region, the apparatus according to any one of claims 1 to 12.
14. Receiving captured video data that provides a dynamic representation of a real-world scene, the video data being linked to a capture pose region, a step; Storing a three-dimensional mesh model that provides a static representation of at least a part of the real-world scene; Receiving a viewing pose; Generating an output image for a viewport with respect to the viewing pose; A method having, wherein the step of generating the output image includes: Generating first image data for the viewport for at least a part of the output image by a field-of-view shift of the captured video data from the capture pose of the captured video data to the viewing pose; Generating second image data for the viewport for at least a first region of the output image from the three-dimensional mesh model; generating the output image so as to include at least a part of the first image data and the second image data for the first region; determining the first region according to a deviation of the viewing pose with respect to the capture pose region; A method comprising: **Claim 15** A computer program including computer program code, wherein the computer program code, when the computer program is run on a computer, executes all steps of the method according to Claim 14.