Apparatus and method for evaluating quality of image capture of a scene
By generating model-based depth data and estimated depth data, virtual capture images are synthesized, solving the problem of difficulty in evaluating the quality of multi-camera capture configurations. This provides an efficient and accurate quality evaluation method, reduces experimental costs and complexity, and improves the quality of virtual reality and real-time video broadcasting.
Patent Information
- Application Number
- CN202080063964.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-12
- Filing Date
- 2020-09-08
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-09-08
Smart Images

Figure CN114364962B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to apparatus and methods for evaluating the quality of image captures of a scene by multiple cameras, for example, particularly for evaluating the quality of video captures of real-life events for virtual display rendering. Background Technology
[0002] In recent years, with the continuous development and introduction of new services and ways to utilize and consume images and videos, the variety and scope of image and video applications have greatly increased.
[0003] For example, an increasingly popular service provides image sequences in such a way that viewers can actively and dynamically interact with the system to change the rendering parameters. In many applications, a particularly attractive feature is the ability to change the viewer's effective viewing position and orientation, such as allowing the viewer to move around and "look around" within the presented scene.
[0004] Such features particularly allow for the provision of virtual reality experiences to users. This allows users to move (relatively) freely within a virtual environment and dynamically change their position and the place they are viewing. Typically, such virtual reality applications are based on a 3D model of the scene, which is dynamically evaluated to provide a specific requested view. This approach is well-known from applications used, for example, in computer and console games (such as in first-person shooter genres).
[0005] Another application that is attracting significant interest is providing views of real-world scenes (and often real-time events) that allow for small viewer movements, such as small head movements and rotations. For example, live video broadcasting of sporting events can allow for the generation of views based on local clients that follow small head movements of the viewer, thus providing the user with the impression of sitting in the stands watching the sporting event. The user can, for example, look around and will have a natural experience similar to that of a spectator in that position in the stands.
[0006] To provide such services for real-world scenes, it is necessary to capture the scene from different locations, thus using different cameras to capture poses. However, this often requires a complex and expensive capture process, frequently involving the simultaneous use of multiple cameras. Therefore, it is desirable to minimize the complexity and resource requirements of the capture process. However, it is often difficult to determine the minimum capture settings that can achieve the desired performance, and different capture configurations typically need to be practically implemented and tested in real-world environments.
[0007] Recently, display devices supporting positional tracking and 3D interaction based on 3D capture of real-world scenes have become increasingly popular. As a result, the relevance and importance of multi-camera capture and processing such as 6DoF (6 degrees of freedom) are rapidly increasing. Applications include live concerts, live sporting events, and telepresence. The ability to freely choose one's own viewpoint enhances the sense of presence compared to regular video, thus enriching these applications. Furthermore, it becomes possible to envision immersive scenes where observers can navigate and interact with the captured scene. For broadcast applications, this requires real-time depth estimation on the production side and real-time view compositing on the client device. Both depth estimation and view compositing introduce errors, and these errors depend on the implementation details of the algorithms. Moreover, the optimal camera configuration depends on the intended application and the 3D structure of the scene being captured.
[0008] Competing methods for 6DoF video capture / creation and compression are often compared visually, and quantitatively in the case of compression. However, quality is generally determined more by the type of camera sensor, the spatial configuration of the camera sensors (e.g., spacing), and camera parameters. Comparing such capture configurations is often expensive, as it involves costly instrumentation and the creation of labor-intensive setups.
[0009] Creating, for example, live 6DoF video requires video capture using multiple cameras, real-time depth estimation, compression, streaming, and playback. In order to make the right choices during development, it is desirable to be able to predict in advance the impact of system parameters (e.g., the number of cameras and the distance between them) and depth estimation algorithms or other processing on image quality.
[0010] As a result, there is a growing desire to evaluate various capture configurations and processes, but this is inherently a difficult process, which typically involves creating experimental settings and evaluating these settings by capturing experimental events and scenarios using these settings.
[0011] Therefore, there will be a need for improved methods for evaluating the quality of capture / camera configuration and / or associated processing. In particular, methods that allow for improved operation, increased flexibility, ease of implementation, convenient operation, convenient evaluation, reduced costs, reduced complexity, and / or improved performance will be advantageous. Summary of the Invention
[0012] Therefore, the present invention seeks to mitigate, reduce or eliminate one or more of the disadvantages mentioned above, either alone or in any combination.
[0013] According to one aspect of the present invention, an apparatus for evaluating the quality of image capture is provided.
[0014] This invention provides an advantageous method for evaluating the quality of camera configurations and / or associated processing. For example, it can be particularly advantageous for evaluating and / or comparing different camera configurations and / or associated processing without requiring the implementation and testing of the capture system. This method allows for the evaluation of different capture methods for a specific application prior to its implementation. Therefore, design decisions and capture parameters can be evaluated and selected based on the performed analysis.
[0015] Consider the following types of images, which can provide particularly useful information: images generated directly for the test pose, without considering the captured image and based on view images generated from depth data generated according to the model; and view images generated based on estimated depth data. For example, this allows for the differentiation of errors and artifacts that can be mitigated by improving depth estimation (whether by adding more camera poses to the camera configuration or by improving the depth mode) versus errors and artifacts that cannot be mitigated by improving depth estimation.
[0016] This method can provide an accurate evaluation of the entire processing path from capturing a view image for the test pose to synthesizing a view image for the test pose, thereby providing a more accurate evaluation of the quality of the achievable results.
[0017] The processing of the first and second synthesizing circuits may in particular include some or all of the processing blocks of the allocated path, including encoding and decoding.
[0018] In some embodiments, the apparatus may be arranged to generate quality metrics for a plurality of different camera configurations and to select a camera configuration from the plurality of different camera configurations in response to the quality metrics.
[0019] Posture can be position and / or orientation.
[0020] According to an optional feature of the invention, at least one of the processes performed by the first synthesis circuit and the second synthesis circuit includes: generating a depth map model for a first virtual capture image in the virtual capture image, and using the depth map model to shift a view of the first virtual capture image to one of the plurality of test poses.
[0021] This method can provide a particularly advantageous evaluation of capture and rendering systems that use depth maps for view composition.
[0022] According to an optional feature of the invention, at least one of the processes performed by the first synthesis circuit and the second synthesis circuit includes: determining a set of 3D points using at least one depth model determined based on the virtual capture image, determining a color for each 3D point using at least one virtual capture image from the virtual capture images, and synthesizing a new image for one of the plurality of test poses based on the projection of the 3D points.
[0023] This method can provide a particularly advantageous evaluation of capture and rendering systems that use 3D point depth representation for view composition.
[0024] According to an optional feature of the invention, the quality circuit is arranged to determine the quality metric, including a first quality metric for the first view image and a second quality metric for the second view image.
[0025] This can provide particularly advantageous evaluations in many embodiments and can especially allow for the distinction between the effects resulting from depth estimation and those resulting from suboptimal depth estimation.
[0026] According to an optional feature of the invention, the quality circuit is arranged to: determine a quality metric for a plurality of camera configurations, and select among the plurality of camera configurations in response to both the first quality metric and the second quality metric.
[0027] This method can provide a particularly advantageous approach for evaluating and selecting between different camera configurations.
[0028] According to an optional feature of the invention, the quality circuit is arranged to select a camera configuration among the plurality of camera configurations in response to at least the following: the first quality metric satisfies a first criterion; the second quality metric satisfies a second criterion; and a measure of the difference between the first quality metric and the second quality metric satisfies a third criterion.
[0029] In many embodiments, this can provide particularly advantageous performance.
[0030] According to an optional feature of the invention, the quality circuit is arranged to generate a signal-to-noise ratio metric for each second view image, and to generate the quality metric in response to the signal-to-noise ratio metric for the second view image.
[0031] This can provide a particularly advantageous method for determining quality metrics. In particular, it has been recognized that signal-to-noise ratio (SNR) metrics can be especially useful for evaluating the effects of camera configuration and associated processing.
[0032] Signal-to-noise ratio (SNR) can specifically be the peak signal-to-noise ratio (PSNR).
[0033] In some embodiments, the quality circuitry may be arranged to generate a signal-to-noise ratio metric for each first view image and to generate a quality metric in response to the signal-to-noise ratio metric for the first view image.
[0034] In other embodiments, other metrics besides signal-to-noise ratio or peak signal-to-noise ratio metrics may be used, such as video multi-method evaluation fusion metrics.
[0035] According to an optional feature of the invention, the processing of at least one of the first synthesis circuit and the second synthesis circuit includes encoding and decoding the virtual capture image based on the encoded virtual capture image and the decoded virtual capture image before image synthesis.
[0036] This method can provide particularly advantageous evaluations, including considering the effects of both camera configuration and encoding / decoding algorithms.
[0037] Encoding and decoding can include, for example, frame / video encoding / decoding, and can include a range of operations such as scaling down an image or depth, packing an image and depth together into a single texture (image), bitstream formatting, etc.
[0038] According to an optional feature of the invention, the processing of at least one of the first synthesis circuit and the second synthesis circuit includes encoding and decoding the depth data associated with the virtual capture image and the estimated depth data based on at least one of the model depth data and the estimated depth data prior to image synthesis.
[0039] This method can provide particularly advantageous evaluations, including considering the effects of both camera configuration and encoding / decoding algorithms.
[0040] According to an optional feature of the invention, the encoding includes lossy encoding.
[0041] According to optional features of the invention, at least some camera poses are the same as at least some test poses.
[0042] According to an optional feature of the invention, the test pose is not less than 10 times the camera pose.
[0043] According to optional features of the invention, the camera positions are arranged in a one-dimensional configuration, while the test positions are arranged in a two-dimensional or three-dimensional configuration.
[0044] According to an optional feature of the invention, the maximum distance between two test locations does not exceed 1 meter.
[0045] In some embodiments, the maximum distance between two test locations does not exceed 10 meters.
[0046] In some embodiments, the maximum distance between two test locations is not less than 10 meters.
[0047] According to one aspect of the present invention, a method for evaluating the quality of image capture is provided.
[0048] These and other aspects, features, and advantages of the invention will be apparent and explained with reference to one or more embodiments described below. Attached Figure Description
[0049] Embodiments of the invention will be described with reference to the accompanying drawings, and by way of example only, in which:
[0050] Figure 1 The illustration shows an example of the components of a device used to evaluate the quality of image captures of a scene by multiple cameras;
[0051] Figure 2 The diagram illustrates the target Figure 1 An example of the viewing area for the test posture of the device;
[0052] Figure 3 The diagram shows that it can be made by Figure 1 Examples of processing functions simulated by the second synthesis circuit and / or the first synthesis circuit of the device;
[0053] Figure 4 The diagram shows that it can be made by Figure 1 Examples of processing functions simulated by the second synthesis circuit and / or the first synthesis circuit of the device;
[0054] Figure 5 An example of an experimental setup for capturing and drawing a scene is illustrated;
[0055] Figure 6 The diagram illustrates the process of... Figure 1 The device is used to select an example of the captured image;
[0056] Figure 7 The diagram illustrates the process of... Figure 1 An example of a quality measure determined by a device;
[0057] Figure 8 The diagram illustrates the process of... Figure 1 An example of a quality measure determined by a device;
[0058] Figure 9 The diagram illustrates the process of... Figure 1 Examples of devices that determine depth maps based on different camera configurations;
[0059] Figure 10 The diagram illustrates the process of... Figure 1 An example of the details of the view image generated by the device. Detailed Implementation
[0060] Figure 1 The illustration shows an example of elements of a device for evaluating the quality of image captures of a scene by multiple cameras. Specifically, the device can determine quality metrics for image capture and the processing of these captured images in order to synthesize images seen from other viewpoints. The device is based on an evaluation of a model of the scene based on the processing of captured images with respect to a particular camera configuration and / or scene.
[0061] The apparatus includes a model storage unit 101 that stores models of scenes. The scene can be a virtual scene representing a real scene or a completely artificially created scene. However, the advantage of this method is that scenes can be selected or created to closely correspond to the scenes in which capture camera configurations and processing are performed. For example, if the system for which evaluation is performed is intended to capture a football match within that system, the virtual scene can be selected to correspond to a football field. As another example, if the quality assessment is for an application of capturing a concert in a concert hall, a virtual scene of the concert hall can be used. In some scenarios, more general scenes can be considered. For example, if the system under investigation is intended to capture scenery, a general, typical virtual landscape scene can be captured. In some cases, the model can be generated based on a real scene, so the scene represented by the model can be either a virtual scene or a real scene.
[0062] The scene model can be any 3D model that allows for the determination of a view image and depth for positions within the scene / model. Typically, the model can be represented by 3D objects, object properties (e.g., optical properties), and light sources. As another example, the model can include multiple meshes with associated textures. Properties such as albedo can be attached to object surfaces. Advanced ray tracing methods can be used to form view images based on the model, taking into account physical properties such as object transparency and multiple scattering.
[0063] Based on this model, the device can synthesize images of multiple test poses in a region of the scene using various methods described below. The results of different methods can then be compared, and a quality metric can be determined based on this comparison.
[0064] In this field, the terms placement and attitude are used as common terms for position and / or orientation. For example, a combination of position and orientation of an object, camera, head, or view can be referred to as attitude or placement. Therefore, a placement or attitude indication may include six values / components / degrees of freedom, where each value / component typically describes a particular attribute of the corresponding object's position / location or orientation / direction. Of course, in many cases, fewer components can be used to consider or represent placement or attitude, for example, if one or more components are considered fixed or unrelated (e.g., if all objects are considered to be at the same height and have a horizontal orientation, then four components can provide a complete representation of the object's attitude). In the following text, the term attitude is used to refer to a position and / or orientation that can be represented by one to six values (corresponding to the maximum possible degrees of freedom).
[0065] The test poses and the areas covered by these poses can be selected based on the specific application / system being evaluated. In many embodiments, the test poses can be selected to cover a relatively small area. In particular, in many embodiments, the test poses can be selected to have a maximum distance of no more than 1 meter between any two test poses. For example, as... Figure 2 As shown, a relatively large number of test poses can be selected as a regular horizontal grid within an area of (approximately) 0.5m × 0.5m. In the example shown, the number of test poses is 15 × 15 (i.e., 225 poses), with a grid spacing of 3cm. It will be appreciated that, depending on the preferred trade-off between desired accuracy and computational complexity, more or fewer test poses may be used in different embodiments. However, in many embodiments, having no fewer than 50, 100, 200, or 5000 test poses is advantageous in order to provide a high level of accuracy for a suitable computational complexity.
[0066] Using numerous examples of test poses within a small area can provide highly accurate results, particularly for applications where captured images are used to provide the viewer with some restricted movement; for example, in this case, the user cannot move freely around the scene but can slightly move or turn their head from a nominal position. Such applications are becoming increasingly popular and offer many desirable uses, such as watching sporting events from a designated location.
[0067] In other embodiments, it may be desirable to view the scene from more varied locations; for example, it may be desirable for the user to be able to move further around the scene or view events from different positions. In such embodiments, test poses covering a larger area / region can be selected.
[0068] The determination of quality metrics is based on capture / camera configuration; that is, quality metrics can be determined for a specific camera configuration. A camera configuration includes one or (typically) multiple camera poses from which the camera can capture images of the scene. Therefore, the camera pose of a camera configuration represents the pose used to capture the scene, and evaluation and quality metrics can be used to determine how well a particular camera configuration is suited for capturing the scene. A camera configuration can also be referred to as a capture configuration.
[0069] Therefore, the model and camera configuration can accordingly represent the actual scene and camera pose that can be used in the settings for capturing the scene.
[0070] In many applications, camera configurations include a relatively small number of cameras, and in practice, the number of cameras in a camera pose is usually no more than 15, 10, or 5.
[0071] Therefore, the number of test poses is typically much greater than the number of camera poses, usually no less than 10 times. This usually provides an accurate, comprehensive, and advantageous quality metric determination for the system.
[0072] In some embodiments, a large number of capture cameras can be considered. For example, for a football field, the number of cameras can easily reach hundreds, depending on the type of flight motion desired. However, even in such embodiments, there may be (potentially) a large number of test postures for evaluation.
[0073] In addition, such as Figure 2 For example, the camera pose / position of the capturing camera may often be consistent with the test pose / position (at least for some cameras). This can provide a practical approach and, for example, reduce some computational complexity. Additionally, having consistent capturing and test poses provides a basic test that the algorithm is working correctly, since MSE = 0, therefore PSNR is not defined (it includes division by ).
[0074] In many embodiments, the camera configuration includes camera positions forming a one-dimensional arrangement, which typically corresponds to a linear arrangement of capturing cameras. This is often very practical, and many practical camera setups are arranged in a linear configuration. In such embodiments, the positions of the test pose are typically arranged in a two-dimensional or three-dimensional configuration. Therefore, the test pose can reflect not only the effects of lateral view displacement but also the effects of displacement in other directions, thus reflecting more typical user behavior.
[0075] Figure 2A specific example is shown where six online camera poses are aligned with six of the 225 test poses (indicated by a ring around the test poses). The test poses are arranged around the camera poses, allowing determination of the extent to which movement from the nominal center position can affect quality.
[0076] The model storage unit 101 is coupled to the reference circuit 103, which is arranged to generate a reference image for multiple test poses by drawing images for multiple test poses based on the model.
[0077] Reference circuit 103 is arranged to generate a reference image by directly evaluating the model and drawing the image. Therefore, the drawing of the reference image is independent of the captured image or camera configuration. The drawing depends directly on the model and the specific test pose. It will be appreciated that different drawing algorithms can be used in different embodiments. However, in many embodiments, the drawing of the reference image is performed based on the stored model using ray tracing techniques.
[0078] As a specific example, rendering can utilize commercially available software packages such as Unity, Unreal Engine, and Blender (open source), which have been developed to create photorealistic game and film content. Such high-level packages typically not only provide photorealistic images but also allow the output of additional data (e.g., depth).
[0079] Therefore, reference images are based solely on the model and test pose, and can typically be generated with very high accuracy because the rendering does not require any assumptions or potentially noisy or distorted processes. Thus, reference images can be considered to provide an accurate representation of the view seen from a specific test pose.
[0080] The model is also coupled to a capture circuit 105, which is arranged to generate virtual capture images of the camera pose for the camera configuration. The capture circuit 105 thus renders virtual capture images reflecting the view seen from the camera pose, and thus renders images to be captured by cameras located in those poses.
[0081] It should be noted that the capturing camera may, in some cases, include, a wide-angle fisheye lens. When ray-tracking such a camera, wide-angle images and depth will be produced with visual distortion. This makes these images different from test images, which can predict a more limited viewport received by a given human eye.
[0082] The rendering algorithm used to render the virtual capture image is based on this model, and in particular, it can be the same algorithm used by reference circuit 103 to render an image for the test pose. In fact, in examples where the camera pose and some test poses are consistent, the same rendering can be used to generate both the reference image for those poses and the virtual camera image for the camera pose.
[0083] Therefore, the captured image corresponds to the image captured by a camera in a pose of the camera configuration for a given model / scene.
[0084] The model storage unit 101 is also coupled to a depth generation circuit 107, which is arranged to generate model depth data for the captured image. The model depth data is generated based on the model and not on the captured image or its content. Specifically, the model depth data can be arranged by determining the distance from each pixel of the captured image to the nearest object represented by the image. Therefore, the model depth data can be generated by evaluating the geometric properties of the model, and can, for example, be determined as part of a ray tracing algorithm that generates the captured image.
[0085] Therefore, model depth data represents the actual depth of the content of the captured image in the model, and can therefore be considered as real-world depth data, that is, it can be considered as highly accurate depth data.
[0086] Depth generation circuit 107 and capture circuit 105 are coupled to a first synthesis circuit 109, which is arranged to perform processing on the virtual capture image based on model depth data to generate a first view image of multiple test poses in a region of the scene.
[0087] Therefore, the first synthesis circuit 109 may include functionality for synthesizing view images for multiple test poses based on captured images and model depth data (i.e., based on real-world depth data). This synthesis may include view shifting, as is known to those skilled in the art.
[0088] Additionally, while in some embodiments the first synthesis circuit 109 may consist only of synthesis operations, in many embodiments the process may also include multiple functions or operations as part of a processing or allocation path for the evaluated application / system. For example, as will be described in more detail later, the process may include encoding, decoding, compression, decompression, view selection, communication error introduction, etc.
[0089] The first synthesis circuit 109 can therefore generate an image that can be synthesized based on the captured image and an assumed real-world depth. Thus, the resulting image can reflect the effects of processing and specific capture configurations.
[0090] The model storage unit 101 is also coupled to a depth estimation circuit 111, which is arranged to generate estimated depth data for the virtual capture image based on the virtual capture image. Therefore, in contrast to the depth generation circuit 107, which determines the depth based on the model itself, the depth estimation circuit 111 determines the depth data based on the capture image.
[0091] Specifically, the depth estimation circuit 111 can perform depth estimation based on a depth estimation technique to be used in the application / system being evaluated. For example, depth estimation can be performed by detecting corresponding image objects in different captured images and determining the differences between these corresponding image objects. Such differences can provide a depth estimate.
[0092] Estimated depth data can therefore represent the depth estimate generated through practical applications and processing, and thus will reflect the defects, errors, and artifacts introduced by that depth estimate. Estimated depth data may be considered less accurate than model depth data, but may actually be a better depth estimate determined and used in the application / system being evaluated.
[0093] Depth estimation circuit 111 and capture circuit 105 are coupled to a second synthesis circuit 113, which is arranged to perform processing of the virtual capture image based on the estimated depth data to generate a second view image for multiple test poses.
[0094] Therefore, the second synthesis circuit 113 may include functionality for synthesizing view images for multiple test poses based on captured images and estimated depth data (i.e., based on expected depth data generated by the evaluated application). This synthesis may include view shifting, as is known to those skilled in the art.
[0095] Additionally, while in some embodiments the second synthesis circuit 113 may consist only of synthesis operations, like the first synthesis circuit 109, in many embodiments the process may also include multiple functions or operations as part of the processing or allocation path for the evaluated application / system, such as encoding, decoding, compression, decompression, view selection, communication error introduction, etc.
[0096] Therefore, the second synthesis circuit 113 can generate an image that can be synthesized based on the captured image itself. The resulting image can reflect the effects of processing and specific capture configurations. In addition, the second view image can reflect the effects of suboptimal depth estimation and can directly reflect the image expected to be generated for the end user in the application and system being evaluated.
[0097] Reference circuit 103, first synthesis circuit 109 and second synthesis circuit 113 are coupled to quality circuit 115, which is arranged to generate a first quality metric in response to a comparison of a first view image, a second view image and a reference image.
[0098] In particular, a quality metric can be determined to reflect the degree of similarity between different images. Specifically, in many embodiments, the quality metric can reflect the improved quality resulting from a reduction in the difference between the first view image, the second view image, and the reference image (for the same test pose and according to any suitable measure or metric of difference).
[0099] A quality metric can reflect the attributes of the camera configuration and the characteristics of the processing performed (for both the first-view image and the second-view image). Therefore, a quality metric can be generated to reflect the impact of at least one of the camera configuration, the processing used to generate the first-view image, and the processing used to generate the second-view image. Typically, this metric can be generated to reflect the impact of all of these items.
[0100] Therefore, this device can provide an efficient and accurate method for evaluating the impact of different camera configurations and / or different processing on quality without having to perform complex, expensive and / or difficult tests and captures.
[0101] This method can provide particularly advantageous evaluations, and in particular, considering both view images generated from real-world data and view images generated from actual estimated data can provide especially valuable information. This is further amplified compared to reference images that do not depend on any capture. For example, by comparing the real image with the reference image, it is possible not only to evaluate how much a particular method affects the quality, but also to determine whether significant improvements can be achieved by improving depth estimation. The evaluation and differentiation of the impact of depth estimation deficiencies and / or their dependence on capture configuration are typically very complex, and the current method can provide efficient and useful evaluations that would otherwise be extremely difficult.
[0102] In particular, the ability to detect whether depth estimation or view shift (causing occlusion) leads to lower quality for a given capture configuration is useful. For example, if both true depth and estimated depth result in poor quality, either the capture configuration needs more cameras, or the view shift is too simple, and either more references need to be included (to handle occlusion), or a more complex prediction method is required.
[0103] It will be appreciated that, depending on the specific preferences and requirements of individual embodiments, different quality metrics and the algorithms and processes used to determine such metrics may be used in different embodiments. In particular, the determination of quality metrics may depend on the exact camera configuration and the processing of image and depth data, including the specific depth estimation and image synthesis methods used.
[0104] In many embodiments, the reference image can be considered the "correct" image, and two quality metrics can be generated by comparing the first and second view images to the "ideal" reference image, respectively. A partial quality metric for each view image can be determined based on the difference between each view image and the reference image for the same test pose. The component quality metrics can then be combined (e.g., summed or averaged) to provide a quality metric for each item in the first and second view image sets, respectively. This quality metric can be generated to include two quality metrics (and thus the quality metric can include multiple components).
[0105] In many embodiments, the quality circuit 115 may be arranged to generate a signal-to-noise ratio (SNR) metric for each view image in the first view image group, and may generate a quality metric in response to these SNR metrics for the first view images. For example, multiple SNR metrics may be combined into a single metric (e.g., by averaging multiple SNR metrics).
[0106] Similarly, in many embodiments, the quality circuitry can be arranged to generate a signal-to-noise ratio (SNR) metric for each view image in the second view image group, and can generate a quality metric in response to these SNR metrics for the second view images. For example, multiple SNR metrics can be combined into a single metric (e.g., by averaging multiple SNR metrics).
[0107] As a specific example, peak signal-to-noise ratio (PSNR) can be used, for example:
[0108]
[0109] MSE is the mean squared error of the view image across the RGB color channels. While PSNR may not be considered the optimal metric for absolute video quality in all cases, the inventors recognized that PSNR is important for... Figure 1 Comparison and evaluation are particularly useful in systems, Figure 1 In this context, providing references within a single dataset is useful.
[0110] The processing performed by the first synthesis circuit 109 and the second synthesis circuit 113 can, as previously described, exist solely within the view synthesis operation, which uses a suitable viewpoint shifting algorithm to synthesize view images for other poses based on the captured image and associated depth data (real-world data and estimated depth data, respectively). Such a method can, for example, generate a quality metric that provides a reasonable assessment of the impact of the particular camera configuration being evaluated on quality. For instance, it can be used in evaluating multiple camera configurations to determine the appropriate camera configuration for capturing a real-world scene.
[0111] However, in many embodiments, the system may include evaluations of other aspects, such as specific processes for the allocation and processing of image capture and image rendering.
[0112] Figure 3 An example of a process that can be included in the processes of the first synthesis circuit 109 and the second synthesis circuit 113 is illustrated.
[0113] In this example, the captured image is fed to image encoding function 301, and the depth data is fed to depth encoding function 303, which performs encoding on the captured image and the associated depth data, respectively. In particular, the encoding performed by the first synthesis circuit 109 and the second synthesis circuit 113 can be exactly the same as the encoding algorithm used in the system being evaluated.
[0114] Importantly, the encoding performed on the captured image data and depth data can be lossy encoding, in which information contained in the captured image and / or depth is lost when the captured image data and depth data are encoded into a suitable data stream. Therefore, in many embodiments, the encoding of the image / depth data also includes compression of the image / depth data. It is often difficult to evaluate the effects of (especially lossy) encoding and compression because it interacts with other effects and processes, and therefore the resulting impact often depends on characteristics other than the encoding itself. However, Figure 1 The device allows for the evaluation and consideration of such effects.
[0115] It will be appreciated that encoding may include any aspect of converting an image / frame / depth into a bitstream for allocation, and decoding may include any processing or operation required to recover the image / frame / depth from the bitstream. For example, encoding and decoding may include a range of operations, including scaling down the image or depth, packing the image and depth together into a single texture (image), bitstream formatting, compression, etc. The exact operations that will be evaluated and implemented by the first synthesis circuit 109 and the second synthesis circuit 113 will depend on the preferences and requirements of the particular embodiment.
[0116] In a typical distribution system, encoded data can usually be transmitted in a single data stream, including both encoded captured image data and depth data. The first synthesis circuit 109 and / or the second synthesis circuit 113 may also accordingly include processing that reflects this communication. This can be achieved through a communication function 305, which may, for example, introduce delays and / or communication errors.
[0117] The first synthesis circuit 109 and / or the second synthesis circuit 113 may further include decoding functions 307, 309 for capturing image data and depth data, respectively. These decoding functions 307, 309 may correspond accordingly to decoding performed at the client / receiver end of the distribution system being evaluated. They may typically be complementary to encoding performed by encoders 301, 303.
[0118] The decoded image and depth data are then used by an image synthesizer, which is configured to synthesize an image for the test pose.
[0119] Therefore, the processing of the first synthesis circuit 109 and the second synthesis circuit 113 can include not only image synthesis itself, but also some, or virtually all, aspects of communication / assignment from camera image capture to the presentation of a view for the test pose. Furthermore, this processing can be matched with the processing used in the real-world system being evaluated, and can virtually use the exact same algorithms, procedures, and code. Thus, this device not only provides an effective means of evaluating camera configurations, but also allows for accurate evaluation of all processing and functions that may be involved in the assignment and processing used to generate the view image.
[0120] A particular advantage of this method is that it can be tailored to accurately include functions and features deemed relevant and appropriate. Furthermore, the process can include algorithms and functions identical to those used in the system being evaluated, thus providing an accurate indication of the quality achievable in the system.
[0121] It will be appreciated that many variations and algorithms for encoding, decoding, communicating, and generally processing images and depths are known, and any suitable method can be used. It will also be appreciated that in other embodiments, the processing performed by the first synthesis circuit 109 and / or the second synthesis circuit 113 may include more or fewer functions. For example, the processing may include functions for selecting between different captured images during view compositing, or image manipulation (e.g., spatial filtering) may be applied before encoding, and sharpening processing may be performed after decoding, etc.
[0122] It will also be realized that, although Figure 3The diagram illustrates the essentially the same processing applied to capturing image and depth data; however, this is not mandatory or essential and can depend on the specific implementation. For example, if the depth data is in the form of a depth map, functions similar to those used for image data processing can often be used, while if the depth data is represented, for example, by a 3D mesh, there can be significant differences in the processing of depth and image data.
[0123] Similarly, in most embodiments, the processing of the first synthesis circuit 109 and the second synthesis circuit 113 is substantially the same, or even possibly identical. In many embodiments, the only difference is that one synthesis circuit uses real-world depth data, while the other uses estimated depth data. However, it will be appreciated that in other embodiments, the processing of the first synthesis circuit 109 and the second synthesis circuit 113 may differ. This can, for example, reduce computational burden, or can, for example, reflect scenarios where real-world depth data and estimated depth data are provided in different formats.
[0124] A particular advantage of this method is that it can be easily adapted for different depth representations and different depth processing procedures when, for example, compositing new views.
[0125] In particular, in some embodiments, at least one of the ground truth depth data and the estimated depth data can be represented by a depth map model, which can be a depth map for each captured image. Such depth maps can typically be encoded and decoded using algorithms also used for image data.
[0126] In such an embodiment, the image compositing function performed by the first compositing circuit 109 and the second compositing circuit 113 can use a depth map model to perform a view shift from a virtual captured image to a test pose. Specifically, pixels of the captured image can be shifted by an amount depending on the depth / parallax indicated by that pixel in the image. When deocclusion occurs, this may result in holes in the generated image. As those skilled in the art will know, such holes can be filled, for example, by padding or interpolation.
[0127] Using depth map models can be advantageous in many systems, and they can be adapted. Figure 1 The device is designed to accurately reflect such processing.
[0128] In other embodiments, other depth data may be used, and other image synthesis algorithms may be employed.
[0129] For example, in many embodiments, depth can be represented using a single 3D model generated from multiple captured images. The 3D model can be represented, for example, by multiple 3D points in space. The color for each of the 3D points can be determined by combining multiple captured images. Since the 3D point model exists in world space, any view can be synthesized from the 3D point model. For example, one approach is to project each 3D point according to a test pose and form an image. This process uses point projection, preserving the depth order and mapping the color corresponding to a given 3D point to the projected pixel location in a virtual camera image of a given test pose. Preserving the depth order ensures that only visible surfaces appear in the image. When a point covers a portion of a target pixel, the contribution of the point can be measured using a method known as snowballing.
[0130] Other variations and options can also be easily adjusted using this method. Figure 1 The device, and Figure 1 The apparatus can provide a particularly attractive evaluation of such a method. In many embodiments, such a complex method can be evaluated along with the rest of the process simply by including the same code / algorithm in the processing performed by the first synthesis circuit 109 and / or the second synthesis circuit 113.
[0131] As previously stated, this method allows for an accurate and reliable evaluation of the quality achievable with a given camera configuration and processing. It enables quality assessment of camera configurations (or a range of camera configurations) without requiring any complex physical setup or measurements. Furthermore, the system can provide quality assessments of various functions involved in the processing, allocation, and compositing of captured images to generate view images. In practice, the device can provide favorable quality assessments of camera configurations, image / depth processing (including, for example, communication), or both.
[0132] This device is particularly useful for selecting between different possible camera configurations. Performing specialized physical measurements and tests to select between different camera configurations would be cumbersome and expensive, but... Figure 1 The device allows for accurate quality assessments that can be used to compare different camera configurations.
[0133] In other embodiments, specific camera configurations may be used, and the device may be used, for example, to compare different algorithm or parameter settings for one or more processing steps included in the processing performed by the first synthesis circuit 109 and / or the second synthesis circuit 113. For example, when choosing between two alternative depth estimation techniques, Figure 1 The device can be used to determine the quality metric for both depth estimation techniques and to select the best depth estimation technique.
[0134] A significant advantage of this approach is that the specific feature being evaluated can be assessed based on multiple aspects of the system. For example, a simple comparison of the captured image or depth estimation itself may lead to a relatively inaccurate evaluation because it does not include, for example, the interactions between different functions.
[0135] In many embodiments, it is particularly advantageous to use three types of synthetic images (i.e., a reference image generated without considering the captured image, a first view image generated considering the true depth, and a second view image generated considering the estimated depth).
[0136] In particular, scene-based model-based evaluation systems allow for highly accurate baselines to be used to evaluate view images synthesized from captured images. Reference images provide a reliable reference for what is considered the "correct" image or view seen from the test pose. Therefore, comparison with such reference images can provide a highly reliable and accurate indication of how closely the view image approximates the image actually seen / captured from the test pose.
[0137] Furthermore, generating synthetic view images based on real-world depth data and estimated depth data provides additional information that is particularly useful in evaluating the quality impact. Of course, assessing the quality and quality impact of the depth estimation algorithm used is especially helpful. Therefore, choosing between different depth estimation algorithms is highly advantageous.
[0138] However, considering the advantages of both types of depth data can also provide valuable information for evaluating other factors related to processing or camera configuration. For example, multiple cameras often mean too many pixels and too high a bit rate. Therefore, image / depth packing and compression are often necessary. To determine whether image / depth packing and compression dominate error performance, it is possible to completely disregard packing and compression to provide a clear comparison.
[0139] In fact, even if perfect depth can be obtained for one or more nearby captured images, it is still impossible to perfectly synthesize images from different viewpoints. Obvious reasons for this include occlusion artifacts and lighting variations (both of which increase as the angle with the reference view increases). This type of error or degradation can be referred to as modeling error or view synthesis error.
[0140] Depth estimation introduces another uncertainty, and in fact, the error can be very large at some locations, and the entire synthesis may actually be interrupted due to depth estimation errors.
[0141] Determining quality metrics (e.g., PSNR) for both ground truth depth and estimated depth allows for better judgment on how to update camera configurations and whether maximum quality has been achieved. For example, if the PSNR for ground truth depth is not significantly better than the PSNR for estimated depth, adding more capture poses or physical cameras would be pointless.
[0142] As mentioned earlier, this method can be used to select between different camera configurations. For example, a range of possible camera configurations can be considered, and selection can be made through... Figure 1 The apparatus determines the quality metric for all possible camera configurations. A camera configuration that achieves the optimal trade-off between the complexity of the camera configuration (e.g., expressed in terms of the number of cameras) and the quality obtained is selectable.
[0143] In many embodiments, as previously described, by Figure 1 The quality metrics generated by the device may include both a first quality metric reflecting the degree of matching between the first view image and the reference image, and a second quality metric reflecting the degree of matching between the second view image and the reference image.
[0144] In many such embodiments, the selection of a given camera configuration may conform to a first quality metric and a second quality metric that satisfy a criterion. For example, the criterion may require that both quality metrics are above a threshold, i.e., the difference between the view image and the reference image is below a threshold.
[0145] However, it is also possible to require that the first and second quality metrics be sufficiently close to each other; that is, the difference between them can be required to be below a given threshold. This requirement provides additional consideration because a sufficiently accurate depth estimate ensures that depth-related estimation errors are unlikely to cause quality issues when a given capture configuration is deployed in practice.
[0146] As a specific example, the device can be used to select between different possible camera configurations. Camera configurations can be evaluated individually and sequentially according to their preferred state. For example, camera configurations can be evaluated in order of complexity; for instance, if the camera configurations correspond to linear arrangements of 3, 5, 7, and 9 cameras respectively, the device can first evaluate the 3-camera configuration, then the 5-camera configuration, then the 7-camera configuration, and finally the 9-camera configuration. The device can evaluate these camera configurations sequentially until a camera configuration is determined to meet the following criteria: a first quality metric satisfies a first standard (e.g., it is above a threshold); a second quality metric satisfies a second standard (e.g., it is above a threshold); and a third standard is met for the difference between the first and second quality metrics (specifically, the difference is below a threshold).
[0147] This selection criterion can be particularly advantageous because both the first and second quality metrics indicate that the synthetic quality is sufficient, and since the differences are small, we believe that the depth estimation will not be interrupted, as it will yield synthetic results similar to those used in the real case.
[0148] In some embodiments, the difference between the first quality metric and the second quality metric can be indirectly calculated by determining the PSNR (or other suitable signal-to-noise ratio) between the first (synthesized) view image and the second (synthesized) view image. This can provide useful additional information. For example, if the PSNR of both the first and second view images is high compared to the reference image, but low compared to each other, then the confidence in the particular configuration / depth estimation algorithm is also lower compared to the case where the PSNR between the first and second view images is also low.
[0149] This method can specifically utilize computer graphics (CG) models and image simulations to compare different camera capture configurations and / or image processing for 6DoF (degrees of freedom) video capture purposes. Given a predefined viewing area and a set of sampling positions / test poses, a single (potentially composite) quality metric can be computed for each capture configuration, and these quality metrics can be used to select the optimal camera configuration, thereby avoiding, for example, the need to actually build and test each system for performance evaluation.
[0150] Competing methods for 6DoF video capture / creation and compression are often compared visually, and quantitatively in the case of compression. However, quality is generally determined more by the type of camera sensor, the spatial configuration of the camera sensors (e.g., spacing), and camera parameters. Comparing such capture configurations is often expensive, as it involves costly instrumentation and the creation of labor-intensive setups. This method and Figure 1 The device can solve these problems.
[0151] Specifically, to compare two or more potential capture configurations (and / or processing methods), a suitable CG scene (e.g., a football field) can be used and represented by the model. A set of sample test poses (typically on a mesh) can then be defined within the boundaries of a predefined 6DoF viewing area. Virtual capture images, in the form of, for example, photorealistic images, can be drawn for each camera pose and each capture configuration to be evaluated. Necessary processing (e.g., depth estimation and compression) is then applied to the drawn capture images using both estimated depth data and ground truth data. As a next step, view images for a set of test poses within the 6DoF viewing area are predicted / synthesized. The results can be compared to reference images, and for each capture configuration, a single quality metric (e.g., the maximum prediction error across all samples) can be calculated. Finally, the quality metrics of all capture configurations can be compared, and the configuration with the minimum error can be selected.
[0152] This method can be particularly cost-effective when evaluating different camera configurations and associated processing. It allows for the evaluation of system performance without the need to purchase and install expensive camera equipment (e.g., around a stadium). Instead, the evaluation can be based on, for example, a realistic CG soccer model (including the field and players). Ray-traced images can also be used to estimate depth, allowing for reasonably low computational quality.
[0153] A specific example will be described in more detail below. In this example, Figure 1 The device can specifically provide a quality assessment method that uses ray-traced images of a virtual scene to simulate acquisition for a given camera capture configuration. The images are then fed into real-time depth estimation and view synthesis software. A view is then synthesized for a test pose with a preset viewing area, and the resulting image is compared to the ray-traced image (reference image). By comparing the image synthesized based on the real-world depth and the image synthesized based on the estimated depth with the ray-traced image, modeling errors can be isolated from depth estimation errors.
[0154] Creating live 6DoF video requires video capture using multiple cameras, real-time depth estimation, compression, streaming, and playback. All these components are under development, and off-the-shelf solutions are difficult to find. To make the right choices during development, it's desirable to be able to predict in advance the impact of system parameters (e.g., baseline distances between cameras) and depth estimation algorithms on image quality. In this specific example, Figure 1 The device can solve such problems and provide an effective method for quality evaluation.
[0155] This example is based on a real-world evaluation using a model powered by Blender, a graphics rendering engine commonly used in filmmaking and game development. In this example, a Python interface (e.g., version 2.79) is used to create a ray-traced image for a camera located in a regular grid of 15×15 anchor points, with the anchor points spaced 3 cm apart. The resulting viewing area allows the observer to move their head forward, backward, left, and right (see [link to example]). Figure 2 Specifically, the viewing area for a standing person allows for limited head motion parallax. The quality of the view synthesis based on a given set of capture camera poses is evaluated on a uniform test pose grid.
[0156] Python was used to automatically perform 15×15 Blender ray tracing on the images to generate a reference image and a capture image for the test pose. In a specific example, a sample spacing of 3 cm was used in both the x and y directions for the test pose. A key parameter that needs to be investigated in advance when designing the capture equipment is the camera spacing (baseline). Using ray-traced images to generate capture images allows finding the optimal baseline for a given minimum quality level within the expected viewing area. As representative scenarios, the specific method was analyzed considering the capture of a human scene (hereinafter referred to as "human") built using MakeHuman software and a car scene (hereinafter referred to as "car") based on a Blender demo file.
[0157] To facilitate a simple comparison of performance and system parameters, peak signal-to-noise ratio was used.
[0158]
[0159] Here, MSE is the mean squared error in the RGB color channels. Additionally, a visual comparison is made between the synthesized view image based on the estimated depth and a synthesized image generated using images from the real-world context.
[0160] This example is based on... Figure 4 The evaluation of the system shown is in Figure 4 In the system shown, the associated processing is implemented by the first synthesis circuit 109 and the second synthesis circuit 113.
[0161] Figure 4The algorithmic blocks from capture to rendering on the client device are illustrated. For live streaming, depth estimation and multi-view registration may include: calibration of intrinsic and extrinsic parameters for paired or camera systems, followed by a multi-camera pose refinement step. Specifically, the process may include disparity estimation, followed by a classifier that determines the probability that the estimated disparity is correct or incorrect. This processing can be implemented on a GPU to achieve real-time performance of 30Hz. A temporal bilateral filter ensures that the depth map changes smoothly as a function of time, making depth errors at least temporally unaffected.
[0162] Figure 5 An example of an experimental setup is shown, comprising a capture device 501, a processing unit 503, and a display 505. The capture device 501 has a camera configuration corresponding to six cameras arranged in a row. The system processes 6-camera feeds at a resolution of 640×1080, calculates six depth maps, and combines the six images with the six depth maps. Figure 1 The entire process is packaged into a single 4K video frame and encoded, all in real-time at 30fps. This system thus forms a scalable, low-cost (consumer hardware) solution for live streaming: depending on the target resolution, two, four, or six cameras can be attached to a PC, and the output of each PC can be streamed to a shared server. Frame synchronization of multiple videos is handled on the capture side. The 4K output of each PC is encoded using an encoding chip located on the graphics card. The system can output standard H.264 or HEVC video, or directly generate HLS / MPEG-DASH video clips to allow for adaptive streaming.
[0163] On the client side, the view is received as a packaged video and decoded using a platform-specific hardware decoder. Decoding is followed by unpacking, where the required reference capture view and depth map are extracted from the packaged frames. The depth map is then converted into a mesh using a vertex shader.
[0164] Run stream selection at the client device to select a subset of streams corresponding to the capture pose used to generate a view for a specific pose. See [link to relevant documentation]. Figure 6 For example, it is possible to assume that the client has a model matrix M that can be used as metadata for a reference view i. i The flow selection uses a 4×4 view matrix V. 左 and V 右 Choose two nearest reference viewpoints for each eye. The formula for calculating the nearest viewpoint is as follows:
[0165]
[0166] Among them, M i This is the model matrix for view i, with homogeneous coordinates p = (0,0,0,1).t V is the view matrix for the left or right eye.
[0167] This essentially corresponds to a method where, for each eye, the eye image is predicted using the most recent captured images (typically two images) with associated depth information. The absolute value notation converts the vectors to scalar distances in 3D space. Matrix V describes the position and orientation of the eye, and matrix M describes the position and orientation of each reference view. Argmini simply represents the minimum distance obtained across all reference cameras.
[0168] In this example, processing and view composition can be based on a 3D mesh. Specifically, a fixed-size, regular triangular mesh is created during initialization. Through sampling of the depth map, the vertex shader directly transforms each vertex of the mesh into a homogeneous output location in clip space:
[0169]
[0170] Among them, D i (u,v) is the disparity derived from the depth map at the input texture coordinates (u,v), Q i It is the disparity of the depth matrix, and PV eye M i It is the product of the model, view, and projection matrix for a given eye. For a specific example, a simple fragment shader could be used, but in principle, it can be used for more advanced occlusion handling and / or blending to improve image quality. The nearest reference view can be blended with the second nearest reference view to predict the final image. This, in principle, allows for scalable solutions for 6DoF video, where only a very large subset of views can be streamed to the user while the user is moving. For example, the blending might depend solely on the proximity of the reference views:
[0171]
[0172] Here, x1 and x2 are the distances along the x-axis to the nearest and second nearest captured view / image. This simple blending equation represents the trade-off between a perceived smooth transition between views and slightly lower view composition accuracy in occluded areas.
[0173] For example, Figure 7 and Figure 8 The PSNR variation is shown in the viewing area for three different camera baselines (distances between camera capture poses of 12cm, 6cm, and 3cm).
[0174] Figure 7The PSNR [dB] of a scene within the viewing area is shown using the change in real-world parallax (top line) relative to the estimated parallax / depth (bottom line) at the camera baseline, within a zoom range of 30–50 dB. Circles indicate the camera position.
[0175] Figure 8 The PSNR [dB] of the scene car within the view area is shown as the change in real-world parallax (top line) using the camera baseline compared to the estimated parallax / depth (bottom line). Circles indicate camera positions.
[0176] Therefore, the top row of each image was generated using the true depth map, while the bottom row was generated using the estimated depth map. The true and estimated depth maps showed a similar pattern: the smaller the baseline, the higher the PSNR in the viewing area. The table below summarizes the results for both scenarios, reporting the minimum PSNR over a 24×24cm area.
[0177]
[0178] It can be seen that the PSNR values for car scenes are systematically lower compared to human scenes. This is due to transparent objects (windows) in cars, for which a model with a single depth value per pixel is clearly too simplistic. For car scenes, depth estimators may fail for shiny and / or transparent parts of objects.
[0179] This method allows for a direct comparison between the actual depth and the estimated depth. Figure 9 This illustrates such a comparison for humans. Figure 9 The diagram illustrates the relationship between the actual baseline and the estimated disparity / depth for different cameras. Scaling is applied to compensate for baseline differences in order to generate the estimated image. Errors at larger baselines disappear at smaller baselines.
[0180] It can be seen that a smaller baseline causes less disparity estimation error. This is understandable because the synthesis occurs based on captured views at a smaller spatial distance, and the occlusion / illumination differences are smaller for a smaller baseline.
[0181] Since real-world images of ray tracing are available, visual comparisons can be made between ray-traced images (reference images) for the test pose, real-world synthetic images, and depth-estimated synthetic images. Figure 10 Such a comparison is shown for a car scene, and in particular, a visual comparison is shown between a ray-traced reference image, a view image synthesized using real-world depth, and a view image synthesized using estimated depth (B = 0.03 m) for different locations in the viewing area (i.e., for different reference poses).
[0182] It can be seen that when using the true depth, there is almost no visible difference between the ray-traced image (reference image) and the synthetic image. When using the estimated depth, some image blurring occurs.
[0183] The apparatus of this example allows for the prediction of the quality of, for example, a 6DoF video broadcasting system based on simulation methods such as ray-traced images. Errors or degradations may occur due to factors such as camera spacing, real-time depth estimation, and view synthesis, and the described method can evaluate all of these aspects.
[0184] This method allows for the separation / isolation of modeling errors from estimation errors, which is useful when attempting to improve depth estimation and view synthesis. It can also be used to design more complex (360-degree) capture rigs or potentially very large camera arrays.
[0185] It should be understood that, for clarity, the above description refers to different functional circuits, units, and processors to describe embodiments of the invention. However, it should be understood that any suitable functional distribution among different functional circuits, units, or processors can be used without departing from the invention. For example, a function illustrated as being performed by a separate processor or controller can be performed by the same processor or controller. Therefore, references to specific functional units or circuits are considered merely as references to suitable modules used to provide the described functions, and not as indications of a strict logical or physical structure or organization.
[0186] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination of these items. Optionally, the invention can be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Therefore, the invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.
[0187] While the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is limited only by the claims. Furthermore, although features may be described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, terms include but do not exclude the presence of other elements or steps.
[0188] Furthermore, although listed separately, multiple modules, elements, circuits, or method steps can be implemented, for example, by a single circuit, unit, or processor. Additionally, while individual features may be included in different claims, these features may also be advantageously combined, and the inclusion of features in different claims does not mean that such a combination is not feasible and / or advantageous. Moreover, the inclusion of a feature in a claim of one type does not mean that the feature is limited to that type, but rather indicates that the feature is equally applicable to other types of claims where appropriate. Furthermore, the order of features in a claim does not imply that the features must operate in any particular order, and in particular, the order of steps in a method claim does not imply that the steps must be performed in that order. Rather, the steps can be performed in any suitable order. Additionally, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. The reference numerals provided in the claims are for clarification purposes only and should not be construed as limiting the scope of the claims in any way.
Claims
1. An apparatus for evaluating the quality of image capture, the apparatus comprising: Storage unit (101), which is used for models of storage scenarios; A capture circuit (105) for generating virtual capture images of multiple camera poses for a camera configuration, the capture circuit (105) being arranged to generate the virtual capture images by drawing images of the camera poses based on the model; A depth generation circuit (107) is used to generate model depth data for the virtual capture image based on the model; A first synthesis circuit (109) is used to process the virtual capture image based on the model depth data to generate a first view image of multiple test poses in a region of the scene; A depth estimation circuit (111) is used to generate estimated depth data for the virtual capture image based on the virtual capture image; A second synthesis circuit (113) is used to process the virtual capture image based on the estimated depth data to generate a second view image for the plurality of test poses; Reference circuit (103) is used to generate a reference image for the plurality of test poses by drawing images for the plurality of test poses based on the model; A quality circuit (115) is configured to generate at least one of the following in response to a comparison of the first view image, the second view image, and the reference image: a quality metric for the camera configuration, the process for generating the first view image, and the process for generating the second view image. The quality circuit (115) is arranged to determine the quality metric, including a first quality metric for the first view image and a second quality metric for the second view image. The quality circuit (115) is also arranged to determine quality metrics for a plurality of camera configurations and to select among the plurality of camera configurations in response to both the first and second quality metrics.
2. The apparatus according to claim 1, wherein, The quality circuit (115) is arranged to select a camera configuration among the plurality of camera configurations in response to at least the following: The first quality metric meets the first standard; The second quality metric meets the second standard; and The difference between the first quality metric and the second quality metric meets the third standard.
3. The apparatus according to claim 1, wherein, The quality circuit (115) is arranged to generate a signal-to-noise ratio metric for each second view image and to generate the quality metric in response to the signal-to-noise ratio metric for the second view image.
4. The apparatus according to claim 1, 2 or 3, wherein, The processing of at least one of the first synthesis circuit (109) and the second synthesis circuit (113) includes encoding and decoding the virtual capture image based on the encoded virtual capture image and the decoded virtual capture image before image synthesis.
5. The apparatus according to claim 1, 2 or 3, wherein, The processing of at least one of the first synthesis circuit (109) and the second synthesis circuit (113) includes encoding and decoding at least one of the depth data associated with the virtual capture image based on at least one of the model depth data and the estimated depth data prior to image synthesis.
6. The apparatus according to claim 4, wherein, The encoding includes lossy encoding.
7. The apparatus according to claim 5, wherein, The encoding includes lossy encoding.
8. The apparatus according to claim 1, 2 or 3, wherein, At least some camera poses are the same as at least some test poses.
9. The apparatus according to claim 1, 2 or 3, wherein, The test pose should be no less than 10 times the camera pose.
10. The apparatus according to claim 1, 2 or 3, wherein, The camera positions form a one-dimensional arrangement, while the test positions form a two-dimensional or three-dimensional arrangement.
11. A method for evaluating the quality of image capture, the method comprising: A model for storage scenarios; Virtual capture images for the multiple camera poses are generated by drawing images of multiple camera poses for the camera configuration based on the model. Based on the model, generate model depth data for the virtual captured image; The virtual capture image is processed based on the model depth data to generate a first view image of multiple test poses in the region of the scene; Based on the virtual capture image, generate estimated depth data for the virtual capture image; The virtual capture image is processed based on the estimated depth data to generate a second view image for the multiple test poses; Reference images for the multiple test poses are generated by drawing images for the multiple test poses based on the model; In response to a comparison of the first view image, the second view image, and the reference image, a quality metric is generated for at least one of the following: the processing for generating the first view image and the processing for generating the second view image. The at least one of processing the virtual capture image based on the model depth data and processing the virtual capture image based on the estimated depth data includes encoding and decoding the virtual capture image based on an encoded virtual capture image and a decoded virtual capture image prior to image synthesis. Furthermore, a quality metric for a plurality of camera configurations is determined using a quality circuit, and a selection is made among the plurality of camera configurations in response to both a first quality metric of the first view image and a second quality metric of the second view image.
12. A computer program product comprising a computer program code module, wherein when the program is run on a computer, the computer program code module is adapted to perform all the steps of claim 11.
Citation Information
Patent Citations
Configuration settings of a digital camera for depth map generation
CN105721853A
Stereo Matching for 3D Encoding and Quality Assessment
US20140015923A1