Apparatus and method for generating an image signal
By generating image signals that combine images and fragments, the problems of high computing resources and high data rates in virtual reality technology are solved, enabling more efficient image generation and rendering and improving the user experience of virtual reality applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-02-14
- Publication Date
- 2026-04-14
AI Technical Summary
Existing virtual reality technologies suffer from high computational resource consumption, high data rates, excessive redundant information, and limited user experience when generating and rendering images, especially in virtual reality applications based on real-world scenarios.
By receiving multiple source images, a composite image is generated, with each pixel representing a scene for the ray pose. Segments are determined by evaluating prediction quality metrics, and an image signal is generated, including image data of the composite image and segments. The image representation is optimized to reduce data rate and improve quality.
It enables more efficient image generation and rendering in virtual reality applications, reduces data rates, improves image quality and user experience, adapts to the flexibility of viewing posture, and is particularly suitable for broadcast video services.
Smart Images

Figure CN113614776B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image signals representing a scene, and particularly, but not exclusively, to the generation of image signals representing a scene and the rendering of images based on such image signals as part of a virtual reality application. Background Technology
[0002] In recent years, the variety and scope of image and video applications have greatly increased with the continuous development and introduction of new services and the ways in which video is utilized and used.
[0003] For example, an increasingly popular service provides image sequences in a way that allows viewers to actively and dynamically interact with the system to change rendering parameters. In many applications, a particularly appealing feature is the ability to change the viewer's effective viewing position and orientation, such as allowing the viewer to move around and "look around" within the presented scene.
[0004] Such features specifically allow for the provision of virtual reality experiences to users. This allows users to move (relatively) freely within a virtual environment and dynamically change their position and the position they are looking at. Typically, such virtual reality applications are based on a 3D model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well-known for computer and console applications, such as in first-person shooter games.
[0005] Especially for virtual reality applications, it is desirable to present three-dimensional images. In fact, to optimize the viewer's immersion, it is generally preferred that the user experience the presented scene as a three-dimensional environment. Ideally, the virtual reality experience should allow the user to choose their position, camera viewpoint, and time relative to the virtual world.
[0006] Virtual reality applications are typically limited by their pre-defined scene-based models and often by artificial models of the virtual world. There is a frequent desire to deliver virtual reality experiences based on real-world captures. However, in many cases, such an approach is restrictive or tends to require constructing a virtual model of the real world based on real-world captures. The virtual reality experience is then generated by evaluating that model.
[0007] However, many current methods tend to be suboptimal and often have high computational or communication resource requirements and / or provide a suboptimal user experience with, for example, reduced quality or limited degrees of freedom.
[0008] In many systems, particularly when based on real-world scenes, an image representation of the scene is provided, comprising an image and depth for one or more capture points / viewpoints in the scene. Image-plus-depth representation provides a highly efficient representation, especially of real-world scenes, where the representation is not only relatively easy to generate by capturing the real-world scene, but is also well-suited for renderers to synthesize views for other viewpoints rather than those captured. For example, a renderer can be configured to dynamically generate views that match the current local viewer's pose. For instance, the viewer's pose can be dynamically determined, and a view can be dynamically generated to match that viewer pose based on a provided image and, for example, a depth map.
[0009] However, for a given image quality, this image representation often results in very high data rates. To provide good scene capture and, in particular, to address occlusion, it is desirable to capture the scene from capture positions that are close to each other and cover a wide area. Therefore, a relatively large number of images are expected. Furthermore, camera capture viewports often overlap, and thus the image set tends to include a significant amount of redundant information. Such problems are often independent of the specific capture configuration and are not specifically related to whether a linear or, for example, circular capture configuration is used.
[0010] Therefore, although many traditional image representations and formats can provide good performance in many applications and services, they tend to be suboptimal, at least in some cases.
[0011] Therefore, improved methods for processing and generating image signals, including image representations of scenes, would be advantageous. In particular, systems and / or methods that allow for improved operation, increased flexibility, improved virtual reality experiences, reduced data rates, increased efficiency, easier distribution, reduced complexity, easier implementation, reduced storage requirements, improved image quality, improved rendering, improved user experience, improved trade-offs between image quality and data rates, and / or improved performance and / or operation would be advantageous. Summary of the Invention
[0012] Therefore, the present invention seeks to reduce, mitigate or eliminate one or more of the above-mentioned disadvantages, either alone or in any combination.
[0013] According to one aspect of the present invention, an apparatus for generating an image signal is provided, the apparatus comprising: a receiver for receiving a plurality of source images representing a scene from different viewing poses; a composite image generator for generating a plurality of composite images based on the source images, each composite image being derived from a set of at least two source images from the plurality of source images, each pixel of the composite image representing a scene for a ray pose, and the ray pose for each composite image including at least two distinct positions, the ray pose for a pixel representing the pose of a ray from the viewing direction for the pixel and from the viewing position for the pixel; and an evaluator for determining elements of the plurality of source images. The system includes: a prediction quality metric for elements of a first source image, indicating the difference between pixel values for pixels in the first source image and predicted pixel values for pixels in the element, the predicted pixel values being pixel values obtained by predicting pixels in the element based on the multiple combined images; a determiner for determining a segment of the source image, the segment including elements whose difference is indicated by the prediction quality metric above a threshold; and an image signal generator for generating an image signal including image data representing the combined image and image data representing the segment of the source image.
[0014] This invention can provide an improved representation of a scene and can provide improved image quality of the rendered image relative to the data rate of the image signal in many embodiments and scenarios. In many embodiments, a more efficient scene representation can be provided, for example, allowing a given quality to be achieved by reducing the data rate. This method can provide a more flexible and efficient approach for rendering images of a scene and can allow for improved adaptation to, for example, scene properties.
[0015] In many embodiments, the method can employ image representations suitable for flexible, efficient, and high-performance virtual reality (VR) applications. In many embodiments, it can allow or enable VR applications with a trade-off between significant improvements in image quality and data rate. In many embodiments, it can allow improved perceived image quality and / or reduced data rate.
[0016] This method can be applied to, for example, broadcast video services that support adaptation to movement and head rotation at the receiving end.
[0017] The source image can specifically be a light intensity image with relevant depth information, such as a depth map.
[0018] This method can particularly allow for the optimization of combined images for separate foreground and background information, wherein the segments provide additional data where particularly appropriate.
[0019] Image signal generators can be configured to use encoding of combined images that is more efficient than encoding of segments. However, segments can typically represent a relatively small proportion of the data in a combined image.
[0020] According to an optional feature of the invention, the combined image generator is arranged to generate the at least first combined image in the plurality of combined images by view synthesis of pixels from a first combined image of the plurality of source images, wherein each pixel of the first combined image represents a scene for a ray pose, and the ray pose for the first image includes at least two distinct positions.
[0021] This can provide particularly advantageous operation in many embodiments and can, for example, allow the generation of combined images for the viewing pose, in which they can (often in combination) provide a particularly advantageous representation of the scene.
[0022] According to an optional feature of the invention, for at least 90% of the pixels of the first combined image, the dot product between the vertical vector and the pixel cross product vector is non-negative, and the pixel cross product vector for a pixel is the cross product between the ray direction for the pixel and the vector from the center point for different viewing poses to the ray position for the pixel.
[0023] In many embodiments, this can provide particularly efficient and advantageous generation of combined images. It can particularly provide a low-complexity method for determining combined images, which provides a favorable representation of the background data by tending to provide an offset view oriented towards the side view.
[0024] According to an optional feature of the invention, the combined image generator is arranged to generate a second combined image from the plurality of source images by view composition of pixels of the second combined image, wherein each pixel of the second combined image represents a scene for a ray pose, and the ray pose for the second image includes at least two distinct positions. Furthermore, for at least 90% of the pixels of the second combined image, the dot product between the vertical vector and the pixel cross product vector is non-positive.
[0025] In many embodiments, this can provide particularly efficient and advantageous generation of combined images. It can particularly provide a low-complexity method for determining combined images, which provides a favorable representation of the background data by tending to provide biased views oriented towards different side views.
[0026] According to an optional feature of the invention, the ray pose of the first combined image is selected to be close to the boundary of a region comprising different viewing poses of multiple source images.
[0027] This can provide advantageous operation in many embodiments and can provide improved background information, for example, through image signals, thereby facilitating and / or improving image signal-based view synthesis.
[0028] According to an optional feature of the invention, each of the ray poses of the first combined image is determined to be less than a first distance from the boundary of a region comprising different observation poses of multiple source images, the first distance not exceeding 50% of the maximum internal distance between points on the boundary.
[0029] This can provide advantageous operation in many embodiments and can provide improved background information, for example, via image signals, thereby facilitating and / or improving image signal-based view synthesis. In some embodiments, the first distance does not exceed 25% or 10% of the maximum interior distance.
[0030] In some embodiments, at least one viewing pose of the combined image is determined to be less than a first distance from the boundary of the region comprising the different viewing poses of the multiple source images, the first distance being no more than 20%, 10%, or even 5% of the maximum distance between two viewing poses in the different viewing poses.
[0031] In some embodiments, at least one viewing pose of the combined image is determined to be at least a minimum distance from the center point of different viewing poses, the minimum distance being at least 50%, 75%, or even 90% of the distance from the center point to the boundary of the region comprising the different viewing poses of multiple source images along a line passing through the center point and at least one viewing pose.
[0032] According to an optional feature of the invention, the combined image generator is arranged for each pixel of a first combined image in the plurality of combined images to: determine a corresponding pixel in each view source image where a corresponding pixel exists in the view source image, the corresponding pixel being a pixel representing the same ray direction as the pixel in the first combined image; select the pixel value of the pixel in the first combined image as the pixel value of the corresponding pixel in the view source image, wherein, for the view source image, the corresponding pixel represents the ray with the maximum distance from the center point for different viewing postures, the maximum distance being located along a first direction along a first axis perpendicular to the ray direction for the corresponding pixel.
[0033] In many embodiments, this can provide particularly efficient and advantageous generation of combined images. It can particularly provide a low-complexity method for determining combined images, which provides a favorable representation of the background data by tending to provide an offset view oriented towards the side view.
[0034] According to an optional feature of the invention, the corresponding pixel includes resampling each source image to represent at least a portion of the spectral surface surrounding the viewing posture, and determining the corresponding pixel to have the same position in the image representation.
[0035] This can provide a particularly effective and accurate determination of the corresponding pixel.
[0036] The surface of the sphere can be represented, for example, by an isometric histogram or a cubic diagram. Each pixel of the sphere can have a ray direction, and resampling the source image can include setting the pixel values of the sphere to the pixel values of the source image with the same ray direction.
[0037] According to an optional feature of the invention, the combined image generator is arranged for each pixel of the second combined image to select a pixel value for a pixel in the second combined image as the pixel value of the corresponding pixel in the view source image, wherein, for the view source image, the corresponding pixel represents a ray having the maximum distance from the center point in a direction opposite to the first direction.
[0038] In many embodiments, this can provide particularly efficient and advantageous generation of combined images. It can particularly provide a low-complexity method for determining combined images, which provides a favorable representation of background data by tending to provide an offset view oriented towards the side view. Furthermore, the second combined image can supplement the first combined image by providing a side view from the opposite direction, thereby combining with the first combined image to provide a particularly advantageous representation of the scene and, in particular, a representation of background information.
[0039] According to an optional feature of the invention, the combined image generator is arranged such that, for each pixel in the third combined image, a pixel value for the pixel in the third combined image is selected as the pixel value of the corresponding pixel in the view source image, wherein the corresponding pixel in the view source image represents the ray with the minimum distance from the center point in the first direction.
[0040] In many embodiments, this can provide particularly effective and advantageous generation of combined images. The third combined image can supplement (one or more) the first (and second) combined images by providing a more frontal view of the scene, which can provide an improved representation of foreground objects in the scene.
[0041] According to an optional feature of the invention, the combined image generator is arranged such that, for each pixel in the fourth combined image, a pixel value for the pixel in the fourth combined image is selected as the pixel value of the corresponding pixel in the view source image, wherein, for the view source image, the corresponding pixel represents a ray with the maximum distance from the center point along a second direction on a second axis perpendicular to the ray direction of the corresponding pixel, the first axis and the second axis having different directions.
[0042] In many embodiments, this can provide particularly efficient and advantageous generation of combined images and can provide an improved representation of the scene.
[0043] According to an optional feature of the invention, the combined image generator is arranged to generate origin data for the first combined image, the origin data indicating which source image is the origin for each pixel of the first combined image; and the image signal generator is arranged to include the origin data in the image signal.
[0044] In many embodiments, this can provide particularly advantageous operation.
[0045] According to an optional feature of the invention, the image signal generator is arranged to include source observation pose data in the image signal, the source observation pose data indicating different observation poses for the source image.
[0046] In many embodiments, this can provide particularly advantageous operation.
[0047] According to one aspect of the present invention, an apparatus for receiving an image signal is provided, the apparatus comprising: a receiver for receiving the image signal, the image signal comprising: multiple composite images, each composite image representing image data derived from a set of at least two source images of multiple source images representing scenes from different viewing poses, each pixel of the composite image representing a scene for a ray pose, and the ray pose for each composite image including at least two distinct positions, the ray pose for a pixel representing the pose of a ray emanating from the viewing position of the pixel in a viewing direction for the pixel; image data for a set of segments of the multiple source images, the segment for a first source image including at least one pixel of the first source image, and for the at least one pixel, a prediction quality metric for the prediction of the segment from the multiple composite images being below a threshold; and a processor for processing the image signal.
[0048] According to one aspect of the present invention, a method for generating an image signal is provided, the method comprising: receiving multiple source images representing a scene received from different viewing postures; generating multiple composite images based on the source images, each composite image being derived from a set of at least two source images from the multiple source images, each pixel of the composite image representing a scene for a ray posture, and the ray posture for each composite image including at least two distinct positions, the ray posture for a pixel representing the posture of a ray in the viewing direction for the pixel and the viewing position for the pixel; determining a prediction quality metric for elements of the multiple source images, the prediction quality metric for elements of a first source image indicating the difference between a pixel value for a pixel in the first source image and a predicted pixel value for a pixel in the element, the predicted pixel value being a pixel value obtained by predicting the pixel in the element based on the multiple composite images; determining a segment of the source images including elements for which the prediction quality metric indicates a difference above a threshold; and generating an image signal including image data representing the composite images and image data representing segments of the source images.
[0049] According to one aspect of the present invention, a method for processing an image signal is provided, the method comprising: receiving an image signal comprising: multiple composite images, each composite image representing image data derived from a set of at least two source images of multiple source images representing scenes from different viewing poses, each pixel of the composite image representing a scene for a ray pose, and the ray pose for each composite image including at least two distinct positions, the ray pose for a pixel representing the pose of a ray emanating from the viewing position of the pixel in a viewing direction for the pixel; image data for a set of segments of the multiple source images, the segment for a first source image including at least one pixel of the first source image, wherein for the at least one pixel, a prediction quality metric for the segment from the multiple composite images is below a threshold; and processing the image signal.
[0050] According to one aspect of the invention, an image signal is provided, comprising: multiple composite images, each composite image representing image data derived from a set of at least two source images of multiple source images representing scenes from different viewing poses, each pixel of the composite image representing a scene for a ray pose, and the ray pose for each composite image including at least two distinct locations, the ray pose for a pixel representing the pose of a ray emanating from the viewing location of the pixel in the viewing direction for the pixel; image data for a set of segments of the multiple source images, the segment for a first source image including at least one pixel of the first source image, wherein for the at least one pixel, a prediction quality metric for the segment from the multiple composite images is below a threshold.
[0051] These and other aspects, features, and advantages of the invention will become apparent and will be explained with reference to one or more embodiments described below. Attached Figure Description
[0052] Embodiments of the invention are described by way of example only with reference to the accompanying drawings, wherein,
[0053] Figure 1 The illustration shows an example of a setup used to provide a virtual reality experience;
[0054] Figure 2 The illustration shows an example of scene capture setup;
[0055] Figure 3 The illustration shows an example of scene capture setup;
[0056] Figure 4 Examples of elements of a device according to some embodiments of the present invention are illustrated;
[0057] Figure 5 Examples of elements of a device according to some embodiments of the present invention are illustrated;
[0058] Figure 6 Examples of pixel selection according to some embodiments of the present invention are illustrated; and
[0059] Figure 7 Examples of pixel selection according to some embodiments of the present invention are illustrated;
[0060] Figure 8 The illustration shows an example of the elements of a ray pose arrangement for a composite image generated according to some embodiments of the present invention;
[0061] Figure 9 The illustration shows an example of the elements of a ray pose arrangement for a composite image generated according to some embodiments of the present invention;
[0062] Figure 10 The illustration shows an example of the elements of a ray pose arrangement for a composite image generated according to some embodiments of the present invention;
[0063] Figure 11 The illustration shows an example of the elements of a ray pose arrangement for a composite image generated according to some embodiments of the present invention;
[0064] Figure 12 Examples of elements of a ray pose arrangement for a composite image generated according to some embodiments of the present invention are illustrated; and
[0065] Figure 13Examples of elements of a ray pose arrangement for a composite image generated according to some embodiments of the present invention are illustrated. Detailed Implementation
[0066] Virtual experiences that allow users to move freely within virtual worlds are becoming increasingly popular, and services are being developed to meet these needs. However, providing effective virtual reality services is very challenging, especially if the experience is based on the capture of real-world environments rather than on entirely virtual, artificial worlds.
[0067] In many virtual reality applications, the viewer's pose input is determined to reflect the pose of the virtual viewer in the scene. The virtual reality device / system / application then generates one or more images corresponding to the viewer's pose and the viewport of the scene.
[0068] Typically, virtual reality applications generate 3D output as separate view images for the left and right eyes. These can then be presented to the user in a suitable manner, such as the left and right eye displays of a typical VR headset. In other embodiments, the images can be presented, for example, on an autostereoscopic display (in which case a large number of viewing images can be generated based on the viewer's posture), or in some embodiments, only a single 2D image can be generated (e.g., using a conventional 2D display).
[0069] Viewer or gesture input can be determined in different ways across different applications. In many embodiments, the user's body movements can be tracked directly. For example, a camera surveying the user's area can detect and track the user's head (or even the eyes). In many embodiments, the user can wear a VR headset that can be tracked by external and / or internal devices. For example, the headset may include accelerometers and gyroscopes that provide information about the headset and therefore the movement and rotation of the head. In some examples, the VR headset may send signals or include (e.g., visual) identifiers that enable external sensors to determine the movement of the VR headset.
[0070] In some systems, viewer posture can be provided manually, such as by the user manually controlling a joystick or similar manual input. For example, a user can manually move the virtual viewer around the scene by using one hand to control a first analog joystick and manually control the virtual viewer's viewing direction by using the other hand to manually move a second analog joystick.
[0071] In some applications, a combination of manual and automatic methods can be used to generate input viewer poses. For example, a head-mounted device can track head orientation, and the viewer's movement / position within the scene can be controlled by the user using a joystick.
[0072] Image generation is based on an appropriate representation of the virtual world / environment / scene. In some applications, a complete 3D model of the scene can be provided, and the view of the scene from a specific viewer's pose can be determined by evaluating that model.
[0073] In many practical systems, a scene can be represented by an image representation that includes image data. Image data typically includes images associated with one or more capture or anchoring poses, and specifically, images for one or more viewports, each corresponding to a particular pose. An image representation comprising one or more images can be used, where each image represents a view of a given viewport for a given viewpoint pose. Such an observation pose or position for which image data is provided is often referred to as an anchoring pose or position or a capture pose or position (because the image data can typically correspond to, or will be, an image captured by a camera positioned in the scene with a position and orientation corresponding to the capture pose).
[0074] Many typical VR applications can build upon this image representation to provide view images corresponding to the viewport of the scene in the current viewer's pose. These images are dynamically updated to reflect changes in the viewer's pose, and are based on image data representing a (possible) virtual scene / environment / world. Applications can perform this by executing view composition and view shifting algorithms known to those skilled in the art.
[0075] In this art, the terms placement and pose are used as general terms for position and / or orientation / orientation. For example, a combination of position and orientation / orientation of an object, camera, head, or view can be referred to as a pose or placement. Therefore, a placement or pose indication may include six values / components / degrees of freedom, where each value / component typically describes a separate attribute of the position / location or orientation / orientation of the corresponding object. Of course, in many cases, placement or pose may be represented using fewer components, for example, if one or more components are considered fixed or unrelated (e.g., if all objects are considered to be at the same height and have a horizontal orientation, four components can provide a complete representation of the object's pose). Hereinafter, the term "pose" is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom).
[0076] Many VR applications are based on pose, which has the maximum number of degrees of freedom; that is, three degrees of freedom for each position and orientation result in a total of six degrees of freedom. Therefore, pose can be represented by a set or vector of six values representing the six degrees of freedom, thus the pose vector can provide a three-dimensional position and / or three-dimensional orientation indication. However, it will be appreciated that in other embodiments, pose can be represented by fewer values.
[0077] Attitude can be at least one of orientation and position. Attitude values can indicate at least one of orientation and position values.
[0078] A system or entity that provides the maximum degrees of freedom for the viewer is typically defined as having 6 degrees of freedom (6DoF). Many systems and entities provide only orientation or position, and are often referred to as having 3 degrees of freedom (3DoF).
[0079] In some systems, VR applications can be provided locally to viewers by a standalone device that does not use or even have access to any remote VR data or processing. For example, a device such as a game console may include: a memory for storing scene data, an input unit for receiving / generating viewer poses, and a processor for generating corresponding images from the scene data.
[0080] In other systems, VR applications can be implemented and executed remotely from the viewer. For example, a user's local device can detect / receive motion / pose data (which is then sent to a remote device that processes the data) to generate the viewer's pose. The remote device can then generate a suitable viewing image for the viewer's pose based on scene data describing the scene. This viewing image is then transmitted to the viewer's local device. For example, the remote device can directly generate a video stream (typically a stereo / 3D video stream) that is directly rendered by the local device. Therefore, in such an example, the local device may not perform any VR processing other than sending motion data and rendering the received video data.
[0081] In many systems, functionality can be distributed across a local device and a remote device. For example, a local device might process received input and sensor data to generate a viewer pose, which is then continuously sent to a remote VR device. The remote VR device can then generate a corresponding view image and send it to the local device for rendering. In other systems, the remote VR device may not directly generate a view image, but instead may select relevant scene data and transmit it to the local device, which can then generate the rendered view image. For instance, the remote VR device might identify the nearest capture point and extract the corresponding scene data (e.g., a spherical image and depth data from the capture point) and send it to the local device. The local device can then process the received scene data to generate an image specific to the current viewer pose. A viewer pose typically corresponds to a head pose, and it can generally be considered equivalent to a reference to a head pose.
[0082] In many applications, especially for broadcast services, the source can transmit scene data as an image (including video) representation of the scene, independent of the viewer's pose. For example, an image representation of a single view sphere for a single capture location can be sent to multiple clients. Each client can then locally synthesize a view image corresponding to its current viewer pose.
[0083] One application that has garnered particular attention is the support for limited amounts of movement, allowing the presented view to be updated to follow small movements and rotations corresponding to a largely static observer who makes only minor head movements and rotations. For example, a seated viewer might turn their head and move it slightly, and the presented view / image would adjust to follow these changes in posture. This approach can provide highly immersive experiences, such as video experiences. For instance, a viewer watching a sporting event might feel as if they are located in a specific part of the stadium.
[0084] Such applications with limited degrees of freedom offer the advantage of providing an improved experience while significantly reducing capture requirements by not needing to accurately represent the scene from many different locations. Similarly, the amount of data that needs to be provided to the renderer can be greatly reduced. In fact, in many scenes, only an image and general depth data for a single viewpoint are required, from which the local renderer is able to generate the desired view.
[0085] This method may be particularly suitable for applications that need to transmit data from a source to a destination through a bandwidth-limited communication channel, such as broadcast or client-server applications.
[0086] Figure 1 The illustration shows an example of a VR system in which a remote VR client device 101 communicates with a VR server 103, for example, via a network 105 such as the Internet. The server 103 can be configured to simultaneously support a potentially large number of client devices 101.
[0087] VR server 103 can support broadcast experiences, for example, by transmitting image signals that include image representations in the form of image data, which can be used by client devices to locally synthesize view images corresponding to appropriate poses.
[0088] In many applications, for example Figure 1 Applications of this technology may therefore require capturing a scene and generating an effective image representation that can be efficiently contained within an image signal. The image signal can then be transmitted to various devices that can locally synthesize views of observation poses other than the capture pose. To do this, the image representation can typically include depth information and, for example, provide images with associated depth. For instance, depth maps can be obtained using stereo capture combined with parallax estimation or using a distance sensor, and these depth maps can be provided along with light intensity images.
[0089] However, a particular problem with this approach is that changing the viewing pose can alter the occlusion properties, causing background fragments that are invisible in a given captured image to become visible under different viewing poses.
[0090] To solve this problem, a relatively large number of cameras are typically used to capture the scene. Figure 2 An example captured by a circular 8-view camera setup is shown. In this example, the camera is facing outwards. It can be seen that different cameras and different captured / source images may have different visibility of different parts of the scene. For example, background area 1 is only visible from camera 2. However, it can also be seen that many scenes can be seen from multiple cameras, thus generating a large amount of redundant information.
[0091] Figure 3 An example of a linear set of cameras is shown. Similarly, the cameras provide information about different parts of the scene; for example, c1 is the only camera capturing area 2, c3 is the only camera capturing area 4, and c4 is the only camera capturing area 3. Meanwhile, some parts of the scene are captured by more than one camera. For example, all cameras capture the frontal views of foreground objects fg1 and fg2, with some cameras providing better captures than others. Figure 3 Example A with four cameras and example B with two cameras are shown. It can be seen that the four-camera setup provides better capture, including capturing a portion of the scene (region 4 of the background bg), but of course it also generates a larger amount of data, including more redundant data.
[0092] The obvious disadvantage of multi-view capture compared to a single central view is the increased amount of image data. Another disadvantage is the large number of pixels generated, requiring processing and a higher pixel rate for the decoder. This also increases the complexity and resource usage of view compositing during playback.
[0093] The following section describes a specific method for using a more efficient and less redundant image representation of the captured view. It seeks to preserve some spatial and temporal coherence of the image data, thereby making the video encoder more efficient. It reduces the bitrate, pixel rate, and view composition complexity at playback sites.
[0094] This representation includes multiple composite images, each generated from two or more source images (which may specifically be captured 3D images, e.g., represented as image depth maps), typically considering only portions of each source image. The composite images can provide a reference for view composition and offer a wealth of scene information. Composite images can be generated to favor more external views of the scene, and specifically to favor the boundaries of the captured area. In some embodiments, one or more central composite images may also be provided.
[0095] In many embodiments, each composite image represents a view from a different observation position, meaning each image may include at least pixels corresponding to different observation / capture / anchor poses. Specifically, each pixel of the composite image may represent a ray pose corresponding to the origin / position and direction / orientation of a ray, which originates from the origin / position, points to the direction / orientation, and terminates at a scene point / object represented by the pixel value of the pixel. At least two pixels of the composite image may have different ray origins / positions. For example, in some embodiments, the pixels of the composite image may be divided into N groups, where all pixels in a group have the same ray origin / position, but are not the same for each group. N may be two or greater. In some embodiments, N may be equal to the maximum number of horizontal pixels in a row (and / or the number of columns in the composite image), and in fact, in some embodiments, N may be equal to the number of pixels, meaning all pixels may have a unique ray origin / pose.
[0096] Therefore, the ray pose of a pixel can represent the origin / position, and / or the orientation / direction of the ray between the origin / position and the scene point represented by the pixel. Specifically, the origin / position can be the view position of the pixel, and the orientation / direction can be the view direction of the pixel. It can effectively represent the light ray captured at the ray position from the ray direction of the pixel, and thus reflect the light ray represented by the pixel value.
[0097] Therefore, each pixel can represent the scene as seen from the viewing position in the viewing direction. Accordingly, the viewing position and viewing direction define a ray. Each pixel can have an associated viewing ray from the pixel's viewing position and viewing direction. Each pixel represents the scene of the (viewing) ray pose, which is the pose of the ray from the pixel's viewpoint / position along the viewing direction. A pixel can specifically represent a scene point (a point in the scene) where the viewing ray intersects with scene objects (including the background). A pixel can represent a light ray from a scene point to the viewing position along the viewing direction. The viewing ray can be a ray emanating from the viewing position along the direction intersecting with the scene point.
[0098] Furthermore, the composite image is supplemented by segments or fragments of captured views that have been identified as not being adequately predicted from the composite image. Therefore, numerous, typically relatively numerous, and usually small segments are defined and included to specifically represent individual portions of the captured image, which can provide information about elements of the scene not adequately represented by the composite image.
[0099] The advantage of this representation is that different codes can be applied to different parts of the image data to be transmitted. For example, efficient and complex coding and compression can be applied to composite images, as this will tend to apply to the largest part constituting the image signal, while less efficient coding can generally be applied to segments. Furthermore, composite images well-suited for efficient coding can be generated, for example, by being generated as similar to regular images, thus allowing the use of efficient image coding methods. In contrast, the properties of segments may vary more greatly depending on the specific characteristics of the image, and therefore may be more difficult to encode efficiently. However, this is not a problem, as these segments tend to provide less image data.
[0100] Figure 4 An example of a device for generating an image signal is illustrated, the image signal comprising a representation of multiple source images of a scene from different source viewpoints (anchor poses) as described above. This device will also be referred to as an image signal transmitter 400. The image signal transmitter 400 may, for example, be included in… Figure 1 VR server 103.
[0101] Figure 5 An example of an apparatus for presenting a view image based on a received image signal comprising a representation of multiple images including a scene is illustrated. Specifically, the apparatus can receive an image signal from... Figure 4 The device generates image data signals and continues to process them in order to render images for a specific viewing posture. Figure 5 The device will also be referred to as image signal receiver 500. Image signal receiver 500 may, for example, be included in... Figure 1 In client device 101.
[0102] Image signal transmitter 400 includes image source receiver 401, which is arranged to receive multiple source images of a scene. The source images can represent views of the scene from different viewing postures. The source images can typically be captured images, such as those captured by cameras in a camera setup. The source images can, for example, include images from a row of equidistant capture cameras or from a ring of cameras.
[0103] In many embodiments, the source image may be a 3D image that includes a 2D image with associated depth information. Specifically, the 2D image may be a view image from a scene viewport of a corresponding capture pose, and the 2D image may be accompanied by a depth image or a map including depth values for each pixel of the 2D image. The 2D image may be a texture map. The 2D image may be a light intensity image.
[0104] Depth values can be, for example, parallax or distance values, indicated by, for instance, the z-coordinate. In some embodiments, the source image can be a 3D image in the form of a texture map with an associated 3D mesh. In some embodiments, such a texture map and mesh representation can be converted into an image-plus-depth representation by the image source receiver before further processing by the image signal transmitter 400.
[0105] Image source receiver 401 accordingly receives multiple source images characterizing and representing a scene from different source view poses. Such a set of source images would allow the generation of view images for other poses using algorithms such as view shifting known to those skilled in the art. Therefore, image signal transmitter 400 is arranged to generate an image signal including image data from the source images and transmit this data to a remote device for local rendering. However, directly transmitting all source images would require an impractical high data rate and would contain a large amount of redundant information. Image signal transmitter 400 is arranged to reduce the data rate by using image representation as described above.
[0106] Specifically, the input source receiver 401 is coupled to a combined image generator 403, which is arranged to generate multiple combined images. The combined image includes information derived from the multiple source images. In different embodiments, the exact method used to derive the combined image may differ, and specific examples will be described in more detail later. In some embodiments, the combined image can be generated by selecting pixels from different source images. In other embodiments, the combined image may alternatively or additionally be generated by view composition from the source images, producing one or more of the combined images.
[0107] However, although each composite image includes contributions from at least two, and often more, source images, typically only portions of individual source images are considered for each composite image. Therefore, for each source image used to generate a given composite image, some pixels are excluded / discarded. Consequently, the pixel values generated for a particular composite image do not depend on the pixel values of these excluded pixels.
[0108] Composite images can be generated such that each image represents not just one view / capture / anchor position, but two or more view / capture / anchor positions. Specifically, the ray origin / position for at least some pixels in a composite image will be different, and therefore a composite image can represent scene views from different directions.
[0109] The composite image generator 403 can therefore be arranged to generate multiple composite images from source images, wherein each composite image is derived from a set of at least two source images, and wherein the derivation of the first composite image typically includes only a portion of each of these at least two source images. Furthermore, each pixel of a given composite image represents a scene with a ray pose, and the ray pose for each composite image can include at least two distinct locations.
[0110] A combined image generator 403 is coupled to an evaluator 405, which is fed a combined image and source images. The evaluator 405 is configured to determine a predicted quality metric for elements of the source images. Elements can be individual pixels, and the evaluator 405 can be configured to determine a predicted quality metric for each pixel of each source image. In other embodiments, elements can include multiple pixels, and each element can be a group of pixels. For example, the predicted quality metric can be determined for blocks of, for example, 4x4 or 16x16 pixels. This can reduce the granularity of the determined segments or fragments, but can significantly reduce processing complexity and resource usage.
[0111] A prediction quality metric is generated for a given element to indicate the difference between the pixel values in the first source image for the pixels in the element and the predicted pixel values for the pixels in the element. Thus, an element can consist of one or more pixels, and the prediction quality metric for that element can indicate the difference between the pixel values for those pixels in the original source image and the pixel values for the pixels to be predicted based on the combined image.
[0112] It will be understood that different methods for determining prediction quality metrics may be used in different embodiments. Specifically, in many embodiments, evaluator 405 may continue to actually perform predictions for each source image from the combined image. It can then determine the difference between the original pixel value and the predicted pixel value for each individual image and each individual pixel. It should be understood that any suitable measure of difference may be used, such as a simple absolute difference, the summation of the square root difference applied to, for example, the pixel value components of multiple color channels, etc.
[0113] Such predictions can therefore simulate prediction / view synthesis that can be performed by the image signal receiver 500 to generate a view for the observation pose of the source image. Therefore, the prediction quality metric reflects how well the receiver of the combined image can generate the original source image based solely on the combined image.
[0114] The predicted image for a source image from a combined image can be an image of the view pose of the source image generated by view synthesis based on the combined image. View synthesis typically involves view pose shifting and often includes view position shifting. View synthesis can be view-shifted image synthesis.
[0115] Predicting a first image from a second image can specifically be a view synthesis of an image based on the second image (and its viewing pose) at the viewing pose of the first image. Therefore, the prediction operation of predicting a first image from a second image can be a viewing pose shift of the second image from its associated viewing pose to the viewing pose of the first image.
[0116] It should be understood that different methods and algorithms for view synthesis and prediction can be used in different embodiments. In many embodiments, a view synthesis / prediction algorithm can be used, which takes as input a synthetic view pose for which a synthetic image is to be generated and multiple input images (each of which is associated with a different view pose). The view synthesis algorithm can then generate a synthetic image for that view pose based on the input image, which typically includes both texture maps and depth maps.
[0117] Many such algorithms are known, and any suitable algorithm can be used without departing from the invention. As an example of this approach, intermediate synthetic / predicted images can be generated first for each input image. This can be achieved, for example, by first generating a mesh for the input image based on the image's depth map. The mesh can then be warped / moved from the view pose of the input image to the synthetic view pose based on geometric calculations. The vertices of the resulting mesh can then be projected onto the intermediate synthetic / predicted image, and a texture map can be overlaid on this image. Such a process can be implemented, for example, using vertex processing and fragment shaders known according to, for example, standard graphics pipelines.
[0118] In this way, an intermediate synthesized / predicted image (hereinafter referred to as the intermediate predicted image) can be generated for each input image to synthesize the observation pose.
[0119] The intermediate prediction images can then be combined, for example, by weighted combination / summing or by selection combination. For example, in some embodiments, each pixel of the synthetic / predicted image for the synthetic view pose can be generated by selecting the foremost pixel from the intermediate prediction images, or by generating the pixel by a weighted sum of the corresponding pixel values over all intermediate prediction images, where the weights for a given intermediate prediction image depend on the depth determined for that pixel. The combination operation is also called a blending operation.
[0120] In some embodiments, prediction quality metrics can be performed without performing full prediction, or instead an indirect metric of prediction quality can be used.
[0121] Prediction quality metrics can be determined indirectly, for example, by evaluating parameters of the processes involved in view offset. For instance, the amount of geometric distortion (stretching) that produces primitives (typically triangles) during view pose transformation. The greater the geometric distortion, the lower the prediction quality metric for any pixel represented by that primitive.
[0122] The evaluator 405 can therefore determine a prediction quality metric for elements of the multiple source images, wherein the prediction quality metric for elements of the first source image indicates the difference between the predicted pixel values of pixels in the elements and the pixel values in the first source image, based on the predicted pixel values of pixels in the elements predicted according to the multiple combined images.
[0123] The evaluator 405 is coupled to the determiner 407, which is arranged to determine segments of the source image that include elements for which a predicted quality metric indicates a difference above a threshold or a predicted quality metric indicates a predicted quality below a threshold.
[0124] A fragment may correspond to an independent element determined by evaluator 405 whose predicted quality metric is below a quality threshold. However, in many embodiments, determiner 407 may be arranged to generate fragments by grouping such elements, and in practice, the grouping may also include some elements whose predicted quality metric is above the threshold.
[0125] For example, in some embodiments, the determiner 407 may be arranged to generate fragments by grouping all adjacent elements (hereinafter referred to as low predicted quality metric and low quality element, respectively) that have a predicted quality metric below a quality threshold.
[0126] In other embodiments, the determiner 407 may be arranged, for example, to fit fragments of a given size and shape to an image such that they include as many low-quality elements as possible.
[0127] The determiner 407 accordingly generates a set of segments that include low-quality elements and therefore cannot be predicted accurately enough from the combined image. Typically, these segments will correspond to a low proportion of the source image and thus to a relatively small amount of image data and pixels.
[0128] Determiner 407 and combined image generator 403 are coupled to image signal generator 409, which receives the combined image and segments. Image signal generator 409 is arranged to generate an image signal including image data representing the combined image and image data representing the segments.
[0129] The image signal generator 409 can specifically encode combined images and segments, and can specifically encode them in different ways, using different algorithms and encoding standards for combined images and segments.
[0130] Typically, if the image is a frame of a video signal, the combined image is encoded using efficient image coding algorithms and standards or efficient video coding algorithms and standards.
[0131] Encoding fragments is generally inefficient. For example, fragments can be combined into fragment images, where each image typically includes segments from multiple source images. Such combined fragment images can then be encoded using standard image or video coding algorithms. However, due to the mixed and partial nature of such combined fragment images, the encoding efficiency is generally lower than that of normal, complete images.
[0132] As another example, due to the sparsity of fragments, they may not be stored in a complete frame / image. In some embodiments, for example, VRML (Virtual Reality Modeling Language) can be used to represent fragments as a mesh in 3D space.
[0133] The image data of a segment is often accompanied by metadata indicating the origin of the segment, such as the original image coordinates and the camera / source image origin.
[0134] In the example, image signals are sent to an image signal receiver 500, which is part of the VR client device 101. The image signal receiver 500 includes an image signal receiver 501 that receives image signals from an image signal transmitter 400. The image signal receiver 501 is arranged to decode the received image signals to recover combined images and segments.
[0135] Image signal receiver 501 is coupled to image processor 503, which is arranged to process image signals, particularly combining images and segments.
[0136] In many embodiments, the image processor 503 may be arranged to synthesize view images for different viewing poses based on the combination of images and fragments.
[0137] In some embodiments, the image processor 503 may continue to first synthesize a source image. Fragments of this source image, included in the image signal as part of the synthesized source message, can then be replaced by image data of the provided fragments. The resulting source image can then be used for conventional image synthesis.
[0138] In other embodiments, the combined images and fragments can be used directly without first restoring the source images.
[0139] It should be understood that the image signal transmitter 400 and the image signal receiver 500 include the functions required for transmitting image signals, including functions for encoding, modulating, transmitting, and receiving image signals. It should be understood that such functions will depend on the preferences and requirements of each embodiment, and such techniques are known to those skilled in the art; therefore, for clarity and brevity, they will not be discussed further here.
[0140] In different embodiments, different methods can be used to generate combined images.
[0141] In some embodiments, the combined image generator 403 may be arranged to generate a combined image by selecting pixels from source images. For example, for each pixel in the combined image, the combined image generator 403 may select a pixel from one of the source images.
[0142] An image and / or depth map comprises pixels having values that can be considered to represent corresponding image attributes (light intensity / strength or depth) of a scene, with rays following ray directions (orientations) originating from a ray origin (position). The ray origin is typically relative to the image's viewing pose, but in some representations may vary on a pixel-by-pixel basis (e.g., for an omnidirectional stereo, where such an image can be considered to have a viewing pose corresponding to the center of the omnidirectional stereo circle, but each pixel has a separate viewing pose corresponding to its position on the omnidirectional stereo circle). Ray directions can typically vary on a pixel-by-pixel basis, particularly for images where all pixels have the same ray origin (i.e., a single common image viewing pose exists). The ray origin and / or direction are also frequently referred to as ray pose or ray projection pose.
[0143] Therefore, each pixel is associated with a position that serves as the origin of a ray / line. Each pixel is also linked to a direction, which is the direction of the ray / line originating from the origin. Thus, each pixel is linked to a ray / line defined by the position / origin and the direction from that position / origin. The pixel value is given by the appropriate properties of the scene at the first intersection of the pixel's ray and the scene object (including the background). Therefore, the pixel value represents the properties of the scene along a ray / line originating from the ray origin position and having a ray direction associated with the pixel. The pixel value also represents the properties of the scene along a ray having the pixel's ray pose.
[0144] The combined image generator 403 can therefore determine, for a given first pixel in the combined image being generated, the corresponding pixel in the source image as a pixel representing the same ray direction. The corresponding pixel can therefore be a pixel representing the same ray direction but possibly at a different location, since the source image may correspond to different locations.
[0145] Therefore, in principle, the combined image generator 403 can determine the ray direction for a given pixel of the combined image, and then identify all pixels in the source image that have the same ray direction (within a given similarity requirement) and consider these as corresponding pixels. Thus, corresponding pixels will typically have the same ray direction but different ray positions / origins.
[0146] Views from different source observation pose images can, for example, be resampled so that the corresponding image coordinates have corresponding ray directions. For instance, when source views are represented in a partially isometric cylindrical projection format, they are resampled to a full 360° / 180° version. For example, a view sphere can be defined around the entire view source configuration. This view sphere can be divided into pixels, where each pixel has a ray direction. For a given source image, each pixel can be resampled into a view sphere representation by setting the view sphere pixel value for a given ray direction to the pixel value of a pixel in the source view with the same ray direction.
[0147] Resampling a source image onto a full sphere surface representation typically produces an image with N partially filled parts, since a single image usually has a finite viewport, and N is the number of source images. However, viewports tend to overlap, and therefore the set of sphere surface representations tends to provide multiple pixel values for any given direction.
[0148] The combined image generator 403 can now continue to generate at least one, but usually multiple, combined images by selecting between corresponding pixels.
[0149] Specifically, a first composite image can be generated to cover a portion of the scene. For example, a composite image with a predetermined size can be generated to cover a specific pixel region in the spectral representation, thereby describing that portion of the scene. In some embodiments, each of the composite images can cover the entire scene and include the entire spectral surface.
[0150] For each pixel in the first combined image, the combined image generator 403 can now consider the corresponding pixel in the view sphere representation and continue selecting one of the pixels. Specifically, the combined image generator 403 can generate the first combined image by selecting a pixel value for the combined image as the pixel value of the corresponding pixel in the view source image, wherein the corresponding pixel represents the ray with the largest distance from the center point along a first direction perpendicular to the ray direction for the corresponding pixel.
[0151] The distance from the center point to the ray direction can be determined as the distance between the ray from the center point and the corresponding pixel of that pixel in the combined image.
[0152] This option can be selected by Figure 6For example, the diagram is based on an example of a circular source observation posture configuration with a center point C.
[0153] In this example, consider determining the pixels of a combined image with a ray direction rc. Camera / Source View Figure 1-4 This direction is captured, and therefore there are four corresponding pixels. Each of these corresponding pixels represents a different pose, and thus represents a ray originating from a different location, as shown in the figure. Therefore, there is an offset distance p1-p4 between the ray and the ray in the combined image rc, corresponding to the distance between the center point C and the ray as it extends backward (crosses the axis).
[0154] Figure 6 The direction / axis perpendicular to ray rc is also shown. For the first combined image, the combined image generator 403 can now select the corresponding pixel for which the ray distance is greatest in this direction. Therefore, in this case, the combined image pixel value will be selected as the camera / view value. Figure 1 The pixel value, because p1 is the maximum distance in that direction.
[0155] The combined image generator 403 can also further determine a second combined image by performing the same operation but selecting the corresponding pixels with the maximum distance in opposite directions (it can be considered that generating the first and second combined images can be achieved by selecting the maximum positive and negative distances relative to the first direction, respectively, if the distance is measured positive in the same direction as the axis and negative in the other direction). Therefore, in this case, the combined image generator 403 will select the combined image pixel values as the camera / view... Figure 4 The pixel value, because p4 is the maximum distance in that direction.
[0156] In many embodiments, the combined image generator 403 can also continue generating a third combined image by performing the same operation but selecting the corresponding pixels that have the minimum distance (minimum absolute distance) in any direction. Therefore, in this case, the combined image generator 403 will select the combined image pixel values as the camera / view... Figure 3 The pixel value, because p3 is the minimum distance.
[0157] In this way, the combined image generator 403 can generate three combined images for the same part of the scene (and possibly the entire scene). One image corresponds to a pixel selection providing a view of the scene from one direction, one representing a view of the scene from the opposite direction from the sidemost side, and one representing a view of the scene from the center. This can be achieved by... Figure 7 illustrate, Figure 7 The images show the viewing directions selected from each view / camera, respectively, for the central composite image and the two side composite images.
[0158] The resulting images provide a very effective representation of the scene, with one combined image typically providing the best representation of the foreground objects, while the other two are combined to provide background focus data.
[0159] In some embodiments, the combined image generator 403 may be arranged to also generate one or more combined images by selecting corresponding pixels according to an axial direction perpendicular to the ray direction but different from the previously used axial direction. This approach may be applicable to non-planar source viewing pose configurations (i.e., three-dimensional configurations). For example, for a spherical source viewing pose configuration, more than two planes may be considered. For example, planes of 0 degrees, 60 degrees, and 120 degrees may be considered, or two orthogonal planes (e.g., left-right planes and top-bottom planes) may be considered.
[0160] In some embodiments, the composite image can be generated through view synthesis / prediction from a source image. The image generator 103 can specifically generate composite images representing scene views based on different viewing positions, and specifically based on viewing positions different from those of the source image. Furthermore, unlike conventional image synthesis, the composite image is not generated to represent a view of the scene from a single view / capture position, but rather can represent the scene from different viewing positions or even within the same composite image. Therefore, the composite image can be generated by generating pixel values for the pixels used in the composite image from the source image through view synthesis / prediction, but the pixel values represent different viewing positions.
[0161] Specifically, for a given pixel in the composite image, view synthesis / prediction can be performed to determine the pixel value corresponding to a specific ray pose for that pixel. This can be repeated for all pixels in the composite image, but at least some pixels will have ray poses at different locations.
[0162] For example, a single composite image can provide a 360° representation of a scene corresponding to, for example, a sphere configured around the entire source's viewing pose. However, views of different parts of the scene can be represented from different locations within the same composite image. Figure 8 An example is shown where the composite image includes pixels representing two distinct ray positions (and therefore pixel view positions): a first ray origin 801 for a pixel representing one hemisphere and a second ray origin 803 for a pixel representing the other hemisphere. For each of these ray positions / origins, a different ray direction is provided for the pixel, as indicated by the arrows. In a particular example, the source view pose configuration includes eight source views (1-8) arranged in a circle. Each camera view provides only a partial view, such as a 90° view, but there is overlap between the views. For a given pixel in the composite image, there may be an associated ray pose, and the pixel value of that ray pose can be determined through view synthesis / prediction from the source views.
[0163] In principle, each pixel of the composite image can be synthesized individually; however, in many embodiments, composite synthesis is performed for multiple pixels. For example, a single 180° image can be synthesized from a source image (e.g., using positions 2, 1, 8, 7, 6, 5, 4) for a first position 801, and a single 180° image can be synthesized from a source image for a second position 803 (e.g., using positions 6, 5, 4, 3, 2, 1, 8). These can then be combined to generate a composite image. If the individually synthesized images overlap, they can be combined or blended to generate the composite image. Alternatively, the overlapping portions of the composite image can be weakened, for example, by assigning preserved color or depth values, thereby improving video coding efficiency.
[0164] In many embodiments, one or more of the combined images can be generated to represent the scene from a more lateral viewpoint providing the scene. For example, in Figure 8 In this example, the center of the view circle corresponds to the center point of the source view pose and the center of the origin position of the rays for the combined image. However, the ray directions for the given ray origins 801 and 803 are not in the primary radial direction, but rather provide a side view of the scene. Specifically, in this example, both the first ray origin 801 and the second origin 803 provide a view to the left; that is, when facing the ray origins 801 and 803 from the center point, the ray directions of both are to the left.
[0165] Image generator 103 can continue to generate second combined images representing different views of the scene, and specifically, it can advantageously generate second views of the scene that are complementary to the first view but viewed in the opposite direction. For example, image generator 103 can generate second combined images that use the same ray origin but with ray directions in opposite directions. For example, image generator 103 can generate images corresponding to... Figure 9 The second combination image of the configuration.
[0166] These two images can provide a very advantageous and complementary representation of the scene, and can often provide an improved representation of the background portion of the scene.
[0167] In many embodiments, the combined image may also include one or more generated images to provide a more frontal view, such as, for example, corresponding to Figure 10 The image is configured as such. In many embodiments, such an example can provide an improved representation of the foreground object's frontal view.
[0168] It should be understood that different ray origin configurations can be used in different embodiments, and in particular, more origins can be used. For example, Figure 11 and 12Examples of two complementary configurations are shown for generating a side-view composite image, where the ray origins are distributed along curves (particularly circles), in this case around the view source configuration (such curves are typically chosen to closely fit the view source's pose configuration). The accompanying figures only show the origins and poses of the circular / curved portions, and it should be understood that in many embodiments, a global plane or 360° view will be generated.
[0169] Figure 7 This can actually be considered as another exemplary configuration, where three combined images are generated based on the positions of eight rays on a circle around a center point. For the first combined image, a radial direction that is circular is chosen; for the second image, a ray direction that rotates 90° to the right is chosen; and for the third image, a ray direction that rotates 90° to the left is chosen. This combination of combined images can provide an efficient composite representation of the scene.
[0170] In some embodiments, the image generator 103 can therefore be arranged to generate pixel values of a composite image for a specific ray pose by view synthesis from a source image. Different ray poses can be selected for different composite images.
[0171] Specifically, in many embodiments, the ray pose of one image can be selected to provide a side view of the scene from the ray origin, and the ray pose of another image can be selected to provide a complementary side view.
[0172] Specifically, the ray pose of the first combined image is such that the dot product between the vertical vector and the pixel cross product vector is non-negative for at least 90% (sometimes 95% or even all) of the pixels in the first combined image. The pixel cross product vector for a pixel is determined as the cross product between the ray direction of the pixel and the vector from the center point of the different source observation poses to the ray position of the pixel.
[0173] The center point of the source observation pose can be generated as an average or average position over the source observation pose. For example, each coordinate (e.g., x, y, z) can be averaged individually, and the resulting average coordinates can be the center point. It should be noted that the center point for the configuration is not (necessarily) located at the center of the smallest circle / sphere containing the source observation pose.
[0174] Therefore, for a given pixel, the vector from the center point to the origin of the ray is a vector in scene space that defines the distance and direction from the center point to the viewpoint of that pixel. The ray direction can be represented by a (ny) vector with the same direction, that is, it can be a vector from the origin of the ray to the scene point represented by the pixel (and therefore can also be a vector in scene space).
[0175] The cross product of these two vectors will be perpendicular to both. For the horizontal plane (in the scene coordinate system), the ray direction to the left (viewed from the center point) will produce a cross product vector with an upward component, that is, a positive z-component in the x, y, z scene coordinate system, where z indicates height. Regardless of the ray origin, for any left-facing view, the cross product vector will be upward, for example, for... Figure 8 All pixel / ray poses will be upward.
[0176] Conversely, for a right-facing view, for all ray poses, the cross product vector will be downwards, for example, for... Figure 9 The pose of all pixels / rays will result in a negative z-coordinate.
[0177] In scene space, the dot product of a vertical vector and all vectors with positive z-coordinates will have the same sign; specifically, upward-pointing vertical vectors are positive, and downward-pointing vertical vectors are negative. Conversely, for negative z-coordinates, the dot product of upward-pointing vertical vectors will be negative, while the dot product of downward-pointing vertical vectors will be positive. Therefore, the dot product has the same sign for right-hand ray poses and the opposite sign for all left-hand ray poses.
[0178] In some cases, zero vectors or dot products may be generated (e.g., for poles on the view circle), and for such ray poses, the sign will not differ from the left or right view.
[0179] It should be understood that the above considerations, after necessary modifications, also apply to three-dimensional representations, such as when the origin of the ray is located on a sphere.
[0180] Therefore, in some embodiments, at least 90% of the combined image, and in some embodiments at least 95% or even all pixels, result in dot products without different signs, i.e., at least many pixels will have side views pointing to the same side.
[0181] In some embodiments, the combined image can be generated with a protective band, or for example, some specific edge pixels may have a dot product that might not meet the requirements. However, for the vast majority of pixels, the requirements are met, and the pixels provide the corresponding side view.
[0182] Furthermore, in many embodiments, at least two combined images satisfy these requirements, but the sign of the dot product is reversed. Thus, for one combined image, at least 90% of the pixels can represent a right-facing view, while for another combined image, at least 90% of the pixels can represent a left-facing view.
[0183] It is possible to generate composite images for poses that provide particularly advantageous views of a scene. The inventors have realized that in many scenarios, it may be particularly advantageous to generate composite images for observation poses that result in more lateral views of the main part of the scene, and further, for a given configuration of the source view, it may be advantageous to generate at least some views that are closer to the extreme positions of the configuration rather than closer to the center of the configuration.
[0184] Therefore, in many embodiments, at least one, and typically at least two, composite images are generated for the ray pose near the boundary of the region corresponding to the source observation pose configuration.
[0185] The region can specifically be a spatial region (a collection or set of points in space) defined by the largest polygon that can be formed using the vertices of lines that can be used as the polygon's lines at least some of the view locations. The polygon can be a planar figure bounded by a finite chain of line segments that close in a loop to form a closed chain or circuit, and this can include a one-dimensional configuration (also known as a degenerate polygon). For a three-dimensional configuration, the region can correspond to the largest possible polyhedron formed by at least some of the source view locations. Therefore, the region can be the largest polygon or polyhedron that can be formed using at least some of the source view locations as the vertices of lines that are the polygon or polyhedron.
[0186] Alternatively, the region encompassing different viewing postures of multiple source images can be a minimal line, circle, or sphere that includes all viewpoint positions. Specifically, this region can be a minimal sphere that includes all source viewpoint positions.
[0187] Therefore, in many embodiments, the ray pose of at least one of the combined images is selected to be close to the boundary of the region including the source view pose configuration.
[0188] In many embodiments, at least one ray position of the combined image is determined to be less than a first distance from the region boundary, wherein the first distance does not exceed 50% of the maximum (internal) distance between points on the region boundary, or in many cases, does not exceed 25% or 10%. Therefore, from the position of the observation pose, the minimum distance to the boundary may not exceed 50%, 25%, or 10% of the maximum distance to the boundary.
[0189] This can be achieved through Figure 13 To explain, Figure 13 An example of a source viewpoint indicated by a black dot is shown. Figure 13 The diagram also illustrates the region corresponding to the smallest sphere, including the viewing posture. In this example, the view configuration is a two-dimensional planar configuration, and the sphere is reduced to a circle 1301. Figure 13Ray pose 1303 for a combined image near the boundary of a sphere / circle / region is also shown. Specifically, the minimum distance dmin to the region boundary / edge is much smaller (approximately 10%) than the maximum distance dmax to the region boundary / edge.
[0190] In some embodiments, the ray pose of the combined image can be determined to be less than a first distance from the region boundary, wherein the first distance does not exceed 20% of the maximum distance between the two source view poses, or typically even 10% or 5%. In an example where the region is determined to be the smallest sphere / circle encompassing all source view poses, the maximum distance between the two view poses is equal to the diameter of the sphere / circle, and therefore the combined image view pose can be selected such that the minimum distance dmin satisfies this requirement.
[0191] In some embodiments, the ray pose of the combined image can be determined as at least a minimum distance from the center point of different observation poses, wherein the minimum distance is at least 50% of the distance from the center point to the boundary along a line passing through the center point and the ray pose, and often even 75% or 90%.
[0192] In some embodiments, the two viewing poses of the combined image are selected such that the distance between them is at least 80%, and sometimes even 90% or 95%, of the maximum distance between two points on the boundary through which the line intersects the viewing pose. For example, if a line is drawn through the two poses, the distance between the two poses is at least 80%, 90%, or 95% of the distance between the points where the line intersects the circle.
[0193] In some embodiments, the maximum distance between two ray poses of the first combined image is at least 80% of the maximum distance between points on the boundary of regions including different viewing poses of multiple source images.
[0194] The inventors have recognized that methods for generating composite images at locations near the boundaries / edges of a region including the source viewpoint can be particularly advantageous, as they tend to provide additional information about background objects in the scene. Most background data is typically captured by a camera or image region with the maximum lateral distance relative to the central viewpoint. This can be advantageously combined with a more central composite image, as this tends to provide improved image information for foreground objects.
[0195] In many embodiments, the image signal generator 409 may be arranged to also include metadata for the generated image data. Specifically, the combined image generator 403 may generate origin data for the combined image, wherein the origin data indicates which source image is the origin of individual pixels in the combined image. The image signal generator 409 may then include this data in the generated image signal.
[0196] In many embodiments, the image signal generator 409 may include source view pose data indicating the view pose of the source images. Specifically, the data may include data defining the position and orientation of each source image / view.
[0197] The image signal may accordingly include metadata, which may individually indicate the position and orientation of the pixel value for each pixel, i.e., ray pose indication. Therefore, the image signal receiver 500 may be arranged to process this data to perform, for example, view composition.
[0198] For example, for each pixel in one of the three views generated by selecting the corresponding pixel, metadata indicating the identity of the source view can be included. This might result in three labeled maps, one for the center view and two for the side views. The labels can then be further linked to specific observation pose data, including, for example, camera optics and device geometry.
[0199] It should be understood that, for clarity, the above description has referenced various functional circuits, units, and processors in describing embodiments of the invention. However, it will be apparent that any suitable functional distribution among the different functional circuits, units, or processors can be used without departing from the invention. For example, functions shown to be performed by separate processors or controllers may be performed by the same processor. Therefore, references to specific functional units or circuits are to be considered merely as references to suitable means of providing the described functions, and not as indications of a strict logical or physical structure or organization.
[0200] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Optionally, the invention can be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the invention can be implemented in a single unit or physically and functionally distributed among different units, circuits, and processors.
[0201] Although the invention has been described in conjunction with some embodiments, it is not intended to limit the invention to the specific forms set forth herein. Rather, the scope of the invention is limited only by the appended claims. Furthermore, while features may appear to have been described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.
[0202] Furthermore, although listed separately, multiple devices, elements, circuits, or method steps can be implemented, for example, by a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features can be advantageously combined, and inclusion in different claims does not imply that the combination of features is infeasible and / or disadvantageous. Including a feature in one class of claims does not imply a limitation on that class, but rather indicates that the feature is equally applicable to other classes of claims where appropriate. Furthermore, the order of features in a claim does not imply any particular order in which the features must operate, and in particular, the order of steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Additionally, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. Reference numerals in the claims are provided only for clarity of example and should not be construed as limiting the scope of the claims in any way.
Claims
1. An apparatus for generating an image signal, the apparatus comprising: Receiver (401) is used to receive multiple source images representing a scene from different viewing postures; A composite image generator (403) is used to generate multiple composite images based on the source images, each composite image being derived from a set of at least two source images from the multiple source images, each pixel of the composite image representing the scene of the ray pose, and the ray pose of each composite image including at least two distinct positions, the ray pose of the pixel representing the pose of a ray emitted from the viewing position of the pixel in the viewing direction of the pixel. An evaluator (405) is used to determine a prediction quality metric for an element of the multiple source images. The prediction quality metric for an element of the first source image indicates the difference between a pixel value for a pixel in the element and a predicted pixel value for a pixel in the element in the first source image. The predicted pixel value is a pixel value obtained by predicting the pixel in the element based on the multiple combined images. Determiner (407), which determines a segment of the source image, the segment including elements that predict quality metrics indicating differences above a threshold; and An image signal generator (409) is used to generate an image signal, the image signal including image data representing the combined image and image data representing the segment of the source image.
2. The apparatus according to claim 1, wherein, The combined image generator (403) is arranged to generate at least the first combined image from the multiple source images by view synthesis of pixels of the first combined image from the multiple source images, wherein each pixel of the first combined image represents the scene for a ray pose, and the ray pose for the first combined image includes at least two distinct positions.
3. The apparatus according to claim 2, wherein, For at least 90% of the pixels in the first combined image, the dot product between the vertical vector and the pixel cross product vector is non-negative, and the pixel cross product vector for a pixel is the cross product between the ray direction for the pixel and the vector from the center point for different viewing poses to the ray position for the pixel.
4. The apparatus according to claim 3, wherein, The combined image generator (403) is arranged to generate the second combined image from the multiple source images by view composition of pixels of the second combined image, wherein each pixel of the second combined image represents the scene for a ray pose, and the ray pose for the second combined image includes at least two distinct positions; and Wherein, for at least 90% of the pixels in the second combined image, the dot product between the vertical vector and the pixel cross product vector is non-positive.
5. The apparatus according to claim 2, wherein, The ray pose of the first combined image is selected to be close to the boundary of the region including the different viewing poses of the multiple source images.
6. The apparatus according to claim 2 or 3, wherein, Each ray pose in the first combined image is determined to be less than a first distance from the boundary of the region including the different observation poses of the multiple source images, the first distance not exceeding 50% of the maximum internal distance between points on the boundary.
7. The apparatus according to claim 2 or 3, wherein, The combined image generator (403) is arranged for each pixel of the first combined image in the plurality of combined images: In each of the multiple source images, a corresponding pixel is identified, where the corresponding pixel represents a pixel with the same ray direction as the pixel in the first combined image. The pixel value of the pixel in the first combined image is selected as the pixel value of the corresponding pixel in the source image, wherein the corresponding pixel represents a ray having a maximum distance from the center point for different viewing postures, the maximum distance being located along a first direction along a first axis perpendicular to the ray direction for the corresponding pixel.
8. The apparatus according to claim 7, wherein, Determining the corresponding pixels includes: resampling each source image into an image representation representing at least a portion of the spectral surface surrounding the viewing posture, and determining the corresponding pixels as pixels having the same position in the image representation.
9. The apparatus according to claim 7, wherein, The combined image generator (403) is arranged for each pixel of the second combined image: The pixel value of the pixel in the second combined image is selected as the pixel value of the corresponding pixel in the source image, wherein the corresponding pixel represents a ray having the maximum distance from the center point in a direction opposite to the first direction.
10. The apparatus according to claim 7, wherein, The combined image generator (403) is arranged as follows: For each pixel in the third combined image: The pixel value of the pixel in the third combined image is selected as the pixel value of the corresponding pixel in the source image, wherein the corresponding pixel represents the ray having the minimum distance from the center point.
11. The apparatus according to claim 7, wherein, The combined image generator (403) is arranged as follows: For each pixel in the fourth combined image: The pixel value of a pixel in the fourth combined image is selected as the pixel value of the corresponding pixel in the source image, wherein the corresponding pixel represents a ray with the maximum distance from the center point along a second direction on a second axis perpendicular to the ray direction of the corresponding pixel, and the first axis and the second axis have different directions.
12. The apparatus according to claim 7, wherein, The combined image generator (403) is arranged to generate origin data for a first combined image, the origin data indicating which of the source images is the origin for each pixel of the first combined image; and the image signal generator (409) is arranged to include the origin data in the image signal.
13. The apparatus according to any one of claims 1-5, wherein, The image signal generator (409) is arranged to include source observation pose data in the image signal, the source observation pose data indicating different observation poses for the source image.
14. An apparatus for receiving image signals, the apparatus comprising: A receiver for receiving image signals, the image signals including: Multiple composite images, each composite image representing image data derived from a set of at least two source images representing a scene from different viewing poses, each pixel of the composite image representing the scene with a ray pose, and the ray pose of each composite image including at least two distinct positions, the ray pose of a pixel representing the pose of a ray emanating from the viewing position of the pixel in the viewing direction of the pixel. Image data for a set of fragments from the multiple source images, wherein a fragment from a first source image includes at least one pixel of the first source image, and for the at least one pixel, the prediction quality metric for the fragment from the multiple combined images is below a threshold; and Processor (503) for processing the image signal.
15. A method for generating an image signal, the method comprising: Receive multiple source images representing a scene from different viewing postures; Multiple composite images are generated based on the source images. Each composite image is derived from a set of at least two source images. Each pixel of the composite image represents the scene of the ray pose, and the ray pose of each composite image includes at least two distinct positions. The ray pose of a pixel represents the pose of a ray emitted from the viewing position of the pixel in the viewing direction of the pixel. Determine a prediction quality metric for elements of the multiple source images. The prediction quality metric for elements of the first source image indicates the difference between the pixel value for a pixel in the first source image and the predicted pixel value for a pixel in the element, wherein the predicted pixel value is a pixel value obtained by predicting the pixel in the element based on the multiple combined images. Identify segments in the source image that include elements whose difference in the predicted quality metric is above a threshold; and An image signal is generated, the image signal including image data representing the combined image and image data representing the segment of the source image.
16. A method for processing an image signal, the method comprising: Receive image signals, the image signals including: Multiple composite images, each composite image representing image data derived from a set of at least two source images representing a scene from different viewing poses, each pixel of the composite image representing a scene with ray poses, and the ray pose of each composite image including at least two distinct locations, the ray pose of a pixel representing the pose of a ray emanating from the viewing location of the pixel in the viewing direction of the pixel; image data for a set of segments of the multiple source images, the segment for a first source image including at least one pixel of the first source image, for the at least one pixel, the prediction quality metric for the segment from the multiple composite images being below a threshold; and The image signal is processed.
17. A computer-readable medium storing an image signal, the image signal comprising: Multiple composite images, each composite image representing image data derived from a set of at least two source images representing a scene from different viewing poses, each pixel of the composite image representing the scene with a ray pose, and the ray pose of each composite image including at least two distinct locations, the ray pose of a pixel representing the pose of a ray emanating from the viewing location of the pixel in the viewing direction of the pixel; image data for a set of segments of the multiple source images, the segment of a first source image including at least one pixel of the first source image, and for the at least one pixel, the prediction quality metric for the segment from the multiple composite images being below a threshold.
18. A computer program product comprising a computer program code module, wherein when the program is run on a computer, the computer program code module adapts the computer to perform all the steps of the method according to claim 15 or 16.
Citation Information
Patent Citations
Apparatus and method for generating a representation of a scene
EP3441788A1
Method for vision field computing
US20110158507A1