Light field camera system and method for setting baseline and convergence distances
By dynamically adjusting the baseline and convergence distance of a multi-view camera rig based on 3D content, the system enhances 3D rendering quality and immersion in multi-view image capture systems.
Patent Information
- Application Number
- JP2023548559
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-11
- Filing Date
- 2022-01-31
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-01-31
AI Technical Summary
Existing multi-view image capture systems struggle with setting optimal baseline and convergence distances, leading to suboptimal 3D rendering and viewer experience.
Dynamically adjust the baseline and convergence distance of a multi-view camera rig based on the 3D content captured, using either physical or virtual cameras, to enhance the 3D rendering quality.
Improves the 3D rendering quality by adapting to changing depths in the scene, providing a more realistic and immersive 3D experience.
Smart Images

Figure 0007733122000005 
Figure 0007733122000006 
Figure 0007733122000007
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 148,587, filed February 11, 2021, which is incorporated herein by reference in its entirety.
[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT none [Background technology]
[0003] A scene in three-dimensional (3D) space may be viewed from multiple perspectives depending on the viewing angle. Additionally, when viewed in stereoscopic video, multiple views representing different perspectives of the scene may be perceived simultaneously, effectively creating a sense of depth that can be perceived by the viewer. A multi-view display presents an image with multiple views to represent how a scene is perceived in a 3D world. A multi-view display renders different views simultaneously to provide a realistic experience to the user. A multi-view image may be dynamically generated and processed by software. Capturing a multi-view image may require multiple cameras or camera positions. Summary of the Invention
[0004] Various features of examples and embodiments according to the principles described herein may be more readily understood by reference to the following detailed description, taken in conjunction with the accompanying drawings, in which like reference numerals designate like structural elements and in which: [Brief explanation of the drawings]
[0005] [Figure 1A] 1 shows a perspective view of an example multi-view display, according to an embodiment consistent with principles described herein.
[0006] [Figure 1B]1 illustrates displaying a multi-view image using a multi-view display in an example, according to one embodiment consistent with principles described herein.
[0007] [Figure 2] 1 illustrates capturing multi-view images of a three-dimensional (3D) scene in an example, according to an embodiment consistent with principles described herein.
[0008] [Figure 3] 1 shows a flowchart of a method for setting a baseline and convergence distance for a multi-view camera rig in one example, according to one embodiment of the principles described herein.
[0009] [Figure 4] 1 illustrates ray casting in an example, according to one embodiment of the principles described herein.
[0010] [Figure 5] 1 shows a block diagram of an example light field camera system, according to one embodiment of the principles described herein;
[0011] [Figure 6A] 1 illustrates a cross-sectional view of an example multi-view display, according to an embodiment consistent with principles described herein.
[0012] [Figure 6B] 10 shows a cross-sectional view of another example multi-view display according to an embodiment consistent with principles described herein.
[0013] [Figure 6C] 1 shows a perspective view of an example multi-view display, according to an embodiment consistent with principles described herein.
[0014] [Figure 7]1 illustrates a block diagram of an example client device, according to an embodiment consistent with principles described herein. DETAILED DESCRIPTION OF THE INVENTION
[0015] Some examples and embodiments may have other features in addition to or instead of the features shown in the above-referenced figures. These and other features are described in more detail below with reference to the above-referenced figures.
[0016] Examples and embodiments according to principles described herein provide techniques for setting both the baseline and convergence distance of a camera rig used to capture multi-view images of a three-dimensional (3D) scene. In particular, according to various embodiments, the baseline and convergence distance of a multi-view camera rig may be determined based on three-dimensional (3D) content captured by the multi-view camera rig. The baseline and convergence distance may then be dynamically adjusted based on the 3D content seen by one or more cameras of the multi-view camera rig as the 3D content changes. According to various embodiments, setting the baseline and convergence distance can use either physical cameras for capturing light field images and video or virtual cameras such as those found in any of a variety of rendering engines (e.g., 3D modeling / animation software, game engines, video editing tools).
[0017] As described below, rather than using preset baselines and convergence distances or manually adjusting the baselines and convergence distances, embodiments are directed to modifying these parameters in response to different depths of a scene within the view of a camera, or more particularly, one or more cameras of a multi-lens camera rig. The one or more cameras may be either virtual cameras or real or physical cameras. When implemented within a renderer, the virtual camera may be positioned to capture a 3D scene of a portion of a 3D model. The renderer may be a game engine, a 3D model player, a video player, or other software environment that positions a virtual camera to capture a 3D model.
[0018] According to various embodiments, a camera generally has a specific position (e.g., coordinates) and orientation to capture a view representing a 3D scene of a 3D model. In this regard, there are multiple depths (sample point depths) between the camera and various surfaces of the 3D scene within the camera's view. To generate a multi-view image, multiple cameras or a "multi-eye camera rig" capture various overlapping views of the 3D scene. In some embodiments using virtual cameras, a virtual camera (e.g., a reference camera) may be duplicated (e.g., generated, copied) to generate a multi-eye camera rig that captures varying overlapping views of the 3D scene from the 3D model, for example. In other embodiments, the multi-eye camera rig includes physical or real cameras configured to capture images representing different views of a physical or real 3D scene. In these embodiments, depth refers to actual or physical depth within the physical 3D scene.
[0019] 1A illustrates a perspective view of an example multi-view display 100 (or the multi-view mode of a multi-mode display), according to one embodiment consistent with principles described herein. As shown in FIG. 1A, the multi-view display 100 comprises a screen configured to display a multi-view image 110 for viewing.100 provides different views 112 of the multiview image in different view directions 120 relative to the screen of the multiview display 100. View Directions 12 0 are shown as arrows extending in various different principal angular directions from the screen, and the different views 112 are shown as shaded polygonal boxes at the ends of the arrows (i.e., depicting the view directions 120), and while only four views 112 and four view directions 120 are shown, all are by way of example and not limitation. While the different views 112 are shown as being above the screen in FIG. 1A, when the multi-view image 110 is displayed on the multi-view display 100, the views 112 may be positioned in different directions. 112 1A above the screen of multi-view display 100 is for ease of explanation only and is meant to represent viewing multi-view display 100 from each one of view directions 120 corresponding to a particular view 112. As shown, multi-view display 100 configured to display multi-view image 110 may be or function as a display (e.g., display screen) of a telephone (e.g., mobile phone, smartphone, etc.), a computer monitor of a tablet computer, laptop computer, desktop computer, a camera display, or an electronic display of virtually any other device, according to various embodiments.
[0020] A light beam having a view direction, in other words a direction corresponding to a view direction of a multi-view display, generally has, by definition herein, a principal angular direction given by angular components {θ, φ}. The angular component θ is referred to herein as the "elevation component" or "elevation angle" of the light beam. The angular component φ is referred to herein as the "seat component" or "seat angle" of the light beam. By definition, the elevation angle θ is an angle in a vertical plane (e.g., perpendicular to the plane of the screen of the multi-view display) and the seat angle φ is an angle in a horizontal plane (e.g., parallel to the plane of the screen of the multi-view display).
[0021] FIG. 1B illustrates an example of a multi-view image 110 being displayed using a multi-view display 100, according to an embodiment consistent with principles described herein. The multi-view image 110 has multiple views 112. Each of the views 112 corresponds to a different view direction 120 or viewpoint of a scene. The views 112 are rendered for display by the multi-view display 100. As such, each view 112 represents a different viewing angle of the multi-view image 110. Thus, the different views 112 have a level of parallax relative to each other. In some embodiments, a viewer may perceive one view 112 with their right eye and a different view 112 with their left eye. This allows the viewer to perceive different views simultaneously, resulting in a stereoscopic image. In other words, the different views 112 create a three-dimensional (3D) effect.
[0022] In some embodiments, as a viewer physically changes the viewing angle relative to the multi-view display 100, the viewer's eyes may encounter different views 112 of the multi-view image 110 at different times as the viewing angle changes. As a result, the viewer can interact with the multi-view display 100 by changing the viewing angle to see different views 112 of the multi-view image 110. For example, as the viewer moves to the left, the viewer can see more of the left side of an object in the multi-view image 110. According to various embodiments, the multi-view image 110 may have multiple views 112 along the horizontal plane or axis, providing a so-called "horizontal parallax only" (HPO) 3D multi-view image, while in other embodiments, the multi-view image 110 may have multiple views 112 along both the horizontal and vertical planes or axes, resulting in a so-called "full parallax" 3D multi-view image. Thus, as the user changes the viewing angle to see different views 112, the viewer can obtain more visual detail in the multi-view image 110. Once processed for display, the multi-view image 110 is stored as data in a format that records the different views 112 according to various embodiments.
[0023] As used herein, a "two-dimensional display" or "2D display" is defined as a display configured to provide substantially the same view of an image regardless of the direction from which the image is viewed (i.e., within a predefined viewing angle or range of the 2D display). Conventional liquid crystal displays (LCDs) found on many smartphones and computer monitors are examples of 2D displays. In contrast, as used herein, a "multi-view display" is defined as an electronic display or display system configured to provide different views of a multi-view image in or from different viewing directions simultaneously from a user's perspective. In particular, the different views 112 can represent different perspectives of the multi-view image 110.
[0024] As described in more detail below, the multi-view display 100 may be implemented using various techniques to adapt the presentation of different image views so that they are perceived simultaneously. One example of a multi-view display is a display that uses a diffraction grating to control the primary angular direction of the different views 112. According to some embodiments, the multi-view display 100 may be a light field display, which is a display that presents multiple light beams of different colors and directions corresponding to different views. In some examples, the light field display is a so-called "naked-eye" three-dimensional (3D) display that can use a diffraction grating or multi-beam element to provide an autostereoscopic representation of the multi-view image without requiring special eyewear to perceive depth. In some embodiments, the multi-view display 100 may require glasses or other eyewear to control which view 112 is perceived by each of the user's eyes.
[0025] In some embodiments, the multi-view display 100 is part of a multi-view display system that renders multi-view and 2D images. In this regard, the multi-view display system may include multiple backlights operating in different modes. For example, the multi-view display system may be configured to provide wide-angle emitted light during the 2D mode using a wide-angle backlight. In addition, the multi-view display system may be configured to provide directional emitted light during the multi-view mode using a multi-view backlight having an array of multi-beam elements, the directional emitted light including multiple directional light beams provided by each multi-beam element of the multi-beam element array. The multi-view display system may be configured to time-multiplex the 2D and multi-view modes using a mode controller to sequentially activate the wide-angle backlight during first sequential time intervals corresponding to the 2D mode and the multi-view backlight during second sequential time intervals corresponding to the multi-view mode. The directional light beam directions of the directional light beams may correspond to different viewing directions of the multi-view image.
[0026] For example, in 2D mode, a wide-angle backlight can generate an image such that the multi-view display system operates like a 2D display. By definition, "wide-angle" emitted light is defined as light that has a cone angle that is greater than the cone angle of the multi-view image or view of the multi-view display. Specifically, in some embodiments, the wide-angle emitted light can have a cone angle that is greater than about 20 degrees (e.g., >±20°). In other embodiments, the wide-angle emitted light cone angle is greater than about 30 degrees (e.g., >±30°), or greater than about 40 degrees (e.g., >±40°), or about It may be greater than 50 degrees (e.g., >±50°). For example, the cone angle of a wide-angle emitted light is approximately 60 degrees. Greater than (e.g., >±60°) There are cases .
[0027] The multi-view mode can use a multi-view backlight instead of a wide-angle backlight. The multi-view backlight can have an array of multi-beam elements that scatter light into multiple directional light beams with different principal angular directions. For example, if the multi-view display 100 operates in the multi-view mode to display a multi-view image having four views, the multi-view backlight can scatter light into four directional light beams, each corresponding to a different view. The mode controller can continuously switch between the 2D mode and the multi-view mode such that the multi-view image is displayed in first sequential time intervals using the multi-view backlight and the 2D image is displayed in second sequential time intervals using the wide-angle backlight.
[0028] In some embodiments, the multi-view display system is configured to guide light within a light guide as guided light. A "light guide" is defined herein as a structure that uses total internal reflection, or "TIR," to guide light within the structure. Specifically, a light guide may include a core that is substantially transparent at the light guide's operating wavelength. In various examples, the term "light guide" generally refers to a dielectric light guide that uses total internal reflection to guide light at the interface between the light guide's dielectric material and the material or medium surrounding the light guide. By definition, the condition for total internal reflection is that the refractive index of the light guide is greater than the refractive index of the surrounding medium adjacent the surface of the light guide material. In some embodiments, the light guide may include a coating in addition to or instead of the aforementioned refractive index difference to further facilitate total internal reflection. The coating may be, for example, a reflective coating. The light guide may be any of several light guides, including, but not limited to, one or both of a plate or slab guide and a strip guide. The light guide may be shaped like a plate or slab. The light guide may be edge-lit by a light source (e.g., a light-emitting device).
[0029] In some embodiments, the multi-view display system is configured to scatter portions of the guided light as directional radiation using multi-beam elements of a multi-beam element array, each multi-beam element of the multi-beam element array including one or more of a diffraction grating, a micro-refractive element, and a micro-reflective element. In some embodiments, the diffraction grating of the multi-beam element may include multiple individual sub-gratings. In some embodiments, the micro-reflective element is configured to reflectively combine or scatter the guided light portions as multiple directional light beams. The micro-reflective element may have a reflective coating to control how the guided light is scattered. In some embodiments, the multi-beam element includes a micro-refractive element configured to refractionally combine or scatter the guided light portions as multiple directional light beams (i.e., refractively scatter the guided light portions) by or using refraction.
[0030] FIG. 2 illustrates capturing a multi-view image of a three-dimensional (3D) scene 200 in one example, according to one embodiment consistent with principles described herein. As shown in FIG. 2, the 3D scene 200 includes various objects (e.g., physical or virtual), such as a tree 202 and a rock 204 on a ground 208. The tree 202, the rock 204, and the ground 208 may be referred to as objects, and together they form at least a portion of the 3D scene 200. The multi-view image of the 3D scene 200 may be displayed and viewed in a manner similar to that described with respect to FIGS. 1A-1B. To capture the multi-view image, a camera 210 may be used. In some embodiments, the camera 210 may include one or more physical cameras. For example, the physical camera includes a lens for capturing light and recording the light as an image. Multiple physical cameras may be used to capture different views of a scene to create the multi-view image. For example, each physical camera may be spaced a defined distance apart to allow different perspectives of objects in the scene to be captured. The distance between different physical cameras enables the ability to capture depth in the 3D scene 200 in the same way that the distance between the viewer's eyes enables 3D vision.
[0031] The camera 210 may also represent one or more virtual (e.g., simulated or imaginary) cameras, as opposed to a physical camera. The 3D scene 200 may be generated using computer graphics techniques that manipulate computer-generated information. In this example, the camera 210 is implemented as a virtual camera having a viewpoint for capturing the 3D scene 200. The virtual camera may be defined in terms of a viewing angle and coordinates within a 3D model. The 3D model may define various objects (e.g., tree 202, rock 204, and ground 208) that are captured by the virtual camera.
[0032] When using the camera 210 to generate or capture a view of a scene, the camera may be configured according to a convergence plane 230. A "convergence plane" or "convergence plane" is defined as a number of positions at which different views align so that there is little parallax between the different views. The convergence plane 230 occurs in front of the camera 210. Objects between the camera 210 and the convergence plane 230 appear closer to the viewer, while objects behind the convergence plane 230 appear farther away from the viewer. In this regard, the degree of parallax between different views increases the further an object is positioned from the convergence plane 230. Objects along the convergence plane 230 appear in focus to the viewer. The distance between the camera 210 and the convergence plane 230 is referred to as the convergence distance or convergence offset. As the camera 210 changes position or orientation, or as the scene changes, the convergence distance is dynamically updated as described herein.
[0033] Camera 210 captures the scene that falls within camera 210's frustum 220. Frustum 220 is shown to have upper and lower limits that define the viewing angle range of 3D scene 200. In Figure 2, convergence plane 230 intersects (relative to camera 210) the bottom of tree 202 and the back surface of tree 202. As a result, the bottom of tree 202 will appear as the primary point of interest to the viewer, as it appears to be in focus and located on the display. Additionally, rock 204 may appear in the foreground in front of tree 202.
[0034] As used herein, "parallax" is defined as the difference between at least two views of a multi-view image at corresponding positions. For example, in the context of stereoscopic video, the left and right eyes may see the same object, but at slightly different positions due to differences in the viewing angles between the eyes. This difference may be quantified as parallax. Changes in parallax across a multi-view image convey a sense of depth.
[0035] The term "baseline" or "camera baseline" is defined as the distance between two cameras capturing corresponding views of a multi-view image. For example, in the context of stereoscopic video, the baseline is the distance between the left eye and the right eye. A larger baseline can result in increased parallax and improve the 3D effect of a multi-view image. Scaling the baseline or baseline scaling refers to modifying or adjusting the baseline according to a multiplier to decrease or increase the baseline. As used herein, pairs of cameras in a multi-view camera rig are, by definition, separated from each other by a baseline. In some embodiments, a common baseline is used between each pair of cameras in a multi-view camera rig.
[0036] As used herein, "convergence distance" or "convergence offset" refers, by definition, to the distance between the camera and a point along the convergence plane. Modifying the convergence offset changes the location of the convergence plane to refocus the multi-view image onto new objects at different depths.
[0037] Further, as used herein, a "3D scene" refers to a scene containing one or more 3D objects that may exist in physical space or may be represented virtually as a 3D model or environment. A physical 3D scene may be captured by a physical camera, and a virtual 3D scene may be captured by a virtual camera.
[0038] Furthermore, as used herein, the article "a" is intended to have its ordinary meaning in the patent art, i.e., "one or more." For example, "camera" means one or more cameras, so "camera" herein means "multiple cameras." Also, any reference herein to "top," "bottom," "upper," "lower," "up," "down," "front," "back," "first," "second," "left," or "right" is not intended to be limiting. As used herein, when the term "about" is applied to a value, unless otherwise specified, it generally means within the tolerance of the equipment used to generate the value, or may mean ±10%, ±5%, or ±1%. Furthermore, as used herein, the term "substantially" means the majority, or almost all, or all, or an amount in the range of about 51% to about 100%. Additionally, the examples herein are intended to be illustrative only and are presented for purposes of explanation and not limitation.
[0039] According to some embodiments of the principles described herein, a method for setting a baseline and convergence distance of a multi-eye camera rig is provided. FIG. 3 shows a flowchart of an example method 300 for setting a baseline and convergence distance of a multi-eye camera rig, according to one embodiment of the principles described herein. In some embodiments, the method 300 for setting a baseline and convergence distance of a multi-eye camera rig may be used to dynamically adjust both the baseline and convergence distance of the multi-eye camera rig. For example, according to some embodiments, baseline adjustment or baseline scaling along with convergence distance adjustment may be performed in real time.
[0040] As shown, the method 300 for setting baseline and convergence distances for a multi-view camera rig includes determining 310 a set or plurality of sample point depths. According to various embodiments, the sample point depths represent a collection or plurality of distances between the multi-view camera rig and a plurality of sample points within a three-dimensional (3D) scene. For example, the 3D scene may be a scene within the field of view of the multi-view camera rig (e.g., as seen by the cameras of the multi-view camera rig), and the distances may represent distances to various objects or points of interest within the 3D scene.
[0041] In some embodiments, determining 310 the depths of a plurality of sample points may include performing ray casting within the 3D scene. For example, a grid of ray casts may be generated outward from a multi-eye camera rig into the 3D scene. Hit distances to various collision objects within the 3D scene are then recorded for ray casts within the ray cast grid. The hit distances correspond to the depths of the various collision objects within the 3D scene.
[0042] FIG. 4 illustrates ray casting in an example, in accordance with one embodiment of the principles described herein. As illustrated, a 3D scene 402 includes multiple objects 404, and a multi-view camera rig 410 is positioned to capture multi-view images of the 3D scene 402. Ray casting involves generating multiple, or a grid of, rays 420, as indicated by the arrows in FIG. 4, and directing the rays into the 3D scene 402 where each ray 420 encounters a point on the object 404. As illustrated, the hit distance is the length of the ray 420 extending between the origin of the multi-view camera rig 410 and the point where the ray 420 encounters (i.e., terminates at) a particular object 404. The hit distance of the grid of rays provided by the ray casting then determines 310 multiple sample point depths (or distances). FIG. 4 also illustrates the baseline b between adjacent cameras of the multi-view camera rig 410.
[0043] In other embodiments, determining 310 the depths of the plurality of sample points may include calculating the depth from a disparity map of the 3D scene. In some embodiments, the disparity map may be provided along with the image of the scene. In other embodiments, the disparity map may be calculated from visual disparity between images recorded by different cameras of a multi-lens camera rig. In particular, calculating the depth may include using image disparity between images captured by different cameras of the multi-lens camera rig. For example, It was captured 3D scene picture A depth buffer associated with the image may be consulted. In some embodiments, calculating the depth may further include forming a disparity map of the 3D scene from the image disparity.
[0044] In yet other embodiments, determining 310 the plurality of sample point depths may include using a depth sensor to measure the distance between the multi-lens camera rig and an object in the 3D scene. In various embodiments, the object may correspond to a sample point of the plurality of sample points, and the depth sensor may include any of a variety of depth sensors. For example, the depth sensor may be a laser distance sensor, including but not limited to a laser detection and ranging (LIDAR) system. In another example, the depth sensor may be a time-of-flight distance sensor. In yet another example, a distance or depth measurement system using either sound waves (e.g., a voice navigation and ranging, or "SONAR," system) or structured light. For example, an image having different colors at different heights may be projected onto the scene, and the image of the scene captured by the camera may then be used by an algorithm to generate a depth map by assigning depths based on the color of each pixel. Even a robot with 3D tracking may be used to effectively explore or "roll" through a scene, recording heights, distances, etc., and may be used to determine sample point depths.
[0045] 3, the method 300 for setting the baseline and convergence distance for a multi-lens camera rig further includes a step 320 of setting the convergence distance to be the average depth of the plurality of sample point depths. sample point depth in An average of the sample point depths is calculated. The convergence distance is then set equal to the calculated average of the sample point depths. In some embodiments, the plurality of sample point depths includes the depths of all sample points. In other embodiments, the plurality of sample point depths may include a selection of sample point depths that is smaller than the total number of sample points; for example, the set may include only sample points that are considered relevant or important, such as sample points associated with key objects or collision objects in the 3D scene. In another embodiment (e.g., for large 3D models), the plurality of sample point depths may be determined for a subset of the plurality of vertices (e.g., every other or every third vertex). Referring again to FIG. 4, the average sample depth (Equation 1) is shown as the average raycast hit distance.
[0046]
number
[0047] In some embodiments, the average depth of the multiple sample points is a weighted average. According to various embodiments, the weighted average (Equation 2) may be calculated using Equation (1):
[0048]
number
[0049]
number
[0050] According to some embodiments, the weights w i is a set of various sample points s from a given location or a particular object in a 3D scene. i For example, weights w i is the i-th sample point s from the center of the scene i In another example, the weight w i is the i-th sample point s from the focal point of the camera in the multi-camera rig. i In yet another example, the weights w i is the specific sample point s * (e.g., a sample point on or associated with an object of interest in a 3D scene) to the ith sample point s i The nodes may be selected or assigned based on distance.
[0051] In some embodiments, the weights w in the weighted average (Equation 2) are i is the distance, e.g., to the scene center, the focal point, or a particular sample point s * In some embodiments, the weight w i The decrease in may have either a linear distribution or a non-linear function (e.g., exponential distribution) as a function of distance. In some embodiments, the weights w i may be assigned according to a Gaussian distribution, e.g., at the scene center, the focal point, or a particular sample point s * It may be centered on.
[0052] Referring again to FIG. 3 , the method 300 for setting a baseline and convergence distance for a multi-view camera rig shown in FIG. 3 further includes step 330 of determining a minimum sample point depth of a plurality of sample point depths. Herein, the minimum sample point depth is generally defined as the sample point having the smallest depth or smallest distance from the cameras of the multi-view camera rig. The minimum sample point depth may be determined at 330, for example, by examining the plurality of sample point depths and finding or identifying the sample point having the smallest value. In other examples, the minimum sample point depth may be determined at 330 by identifying a group of sample points having the lowest depth (or distance) and then setting the minimum sample point depth to be equal to the average of the sample point depths within the group of sample points with the lowest depth. In some examples, the group of sample points with the lowest depth may include a percentage of sample points with the lowest depth or distance, for example, approximately five percent (5%) or approximately ten percent (10%) of the sample points with the lowest depth or distance. Referring back to FIG. 4, the minimum sample point depth z represents the minimum hit distance of ray 420. min is shown.
[0053] Returning again to FIG. 3, according to various embodiments, a method 300 for setting a baseline and convergence distance for a multi-lens camera rig includes: Negative further comprising the step 340 of setting the baseline to be the reciprocal;
[0054]
number
[0055] According to some embodiments, a multi-eye camera rig may include multiple virtual cameras. For example, the 3D scene may be a 3D model, and the multiple virtual cameras may be cameras associated with or used to capture the 3D model. In some embodiments, the virtual cameras of the multiple virtual cameras may be virtual cameras managed by a renderer. For example, the multi-eye camera rig may be associated with a renderer that captures a virtual 3D scene using the virtual cameras of the multi-eye camera rig. In another embodiment, the multi-eye camera rig may include multiple physical cameras. For example, the 3D scene may be or represent a physical scene captured by cameras of the multiple physical cameras. In yet another embodiment, the multi-eye camera rig may include cameras (e.g., one or more cameras) that move between multiple positions to capture images that form the 3D scene. In some embodiments, the sample point depth may represent a depth or distance relative to one camera of the multi-eye camera rig (e.g., a reference camera), while in other embodiments, the sample point depth may be a distance relative to the multi-eye camera rig as a whole. As mentioned above, FIG. 4 also shows the baseline b between adjacent cameras of the multi-lens camera rig 410.
[0056] In another embodiment according to the principles described herein, a light field camera system is provided. In some embodiments, the light field camera system has or can provide automatic baseline and convergence distance determination. Figure 5 shows a block diagram of an example light field camera system 500 according to one embodiment of the principles described herein. As shown in Figure 5, the light field camera system 500 comprises a multi-lens camera rig 510. The multi-lens camera rig 510 comprises multiple cameras separated from each other by a baseline b, as shown.
[0057] Light field camera system 500 further comprises processor 520 and memory 530. Memory 530 is configured to store instructions that, when executed by processor 520, perform determining a set or plurality of sample point depths representing distances between a multi-view camera rig and a plurality of sample points in a three-dimensional (3D) scene within a field of view of the multi-view camera rig. In some embodiments, determining the plurality of sample point depths may be substantially similar to step 310 of determining sample point depths, as described above with respect to method 300 of setting baseline and convergence distances for a multi-view camera rig.
[0058] In particular, in some embodiments, the plurality of sample point depths are calculated from a depth map of an image representing a 3D scene using a disparity map, and Inside performing ray casting to determine sample point depths within the 3D scene. In other embodiments, the plurality of sample point depths may be determined at 310 by using a depth sensor to measure a distance between the multi-eye camera rig and an object within the 3D scene, the object corresponding to a sample point among the plurality of sample points. In some embodiments, the average depth of the plurality of sample point depths is a weighted average, with weights in the weighted average assigned according to a decreasing function of distance from the focal point of the 3D scene.
[0059] In some embodiments, a camera of the plurality of cameras is a virtual camera and the 3D scene is a 3D model. In some embodiments, a camera of the plurality of cameras of a multi-lens camera rig includes a physical camera and the 3D scene represents a physical scene imaged by the physical camera.
[0060] When executed by the processor 520 shown in FIG. 5, the instructions stored in the memory 530 further perform setting the convergence distance and baseline of the multi-camera rig. RevenueThe convergence distance may be the average depth of multiple sample point depths, and the baseline may be the minimum sample point depth minus the convergence distance, according to various embodiments. Negative In some embodiments, setting the convergence distance and baseline may be substantially similar to step 320 of setting the convergence distance and step 340 of step 300 of setting the baseline and convergence distance for a multi-eye camera rig, described above.
[0061] In some implementations, an application executed by the processor 520 can generate a 3D model using computer graphics techniques for 3D modeling. A 3D model is a mathematical representation of various surfaces and textures of different objects, and may include spatial relationships between the objects. The application may include a renderer that generates and updates the 3D model in response to user input. User input may include navigating the 3D model by clicking or dragging a cursor, pressing directional buttons, converting the user's physical location to a virtual location within the 3D model, etc. The 3D model may be loaded into memory 530 and subsequently updated. The 3D model may be converted into a multi-view image that reveals a window into the 3D model. The window may be defined by multiple virtual cameras, and the multi-view camera rig 510 has coordinates and orientations within the 3D model. In some embodiments, the baseline and convergence distance of the virtual cameras may be dynamically updated in response to movement of the virtual cameras or changes in the 3D scene.
[0062] In some embodiments (e.g., as shown in FIG. 5), the light field camera system 500 further comprises a multi-view display 540. In these embodiments, the convergence distance may correspond to the zero parallax plane of the multi-view display 540.
[0063] 6A shows a cross-sectional view of an example multi-view display 600, according to one embodiment consistent with principles described herein. FIG. 6B shows a cross-sectional view of another example multi-view display 600, according to one embodiment consistent with principles described herein. In particular, FIG. 6A shows multi-view display 600 during or in a first mode or two-dimensional (2D) mode. 6 6B illustrates the multi-view display 600 during or in accordance with a second or multi-view mode. FIG. 6C illustrates a perspective view of the multi-view display 600 in an example, according to an embodiment consistent with principles described herein. The multi-view display 600 is shown, by way of example and not limitation, during the multi-view mode. 6 6B. Further, according to various embodiments, the 2D mode and the multi-view mode may be time-multiplexed in a time-sequential or time-interlaced manner to provide the 2D mode and the multi-view mode in alternating first and second time intervals (e.g., alternating between FIG. 6A and FIG. 6B). As such, the multi-view display 600 may also be referred to as a "time-multiplexed mode-switching" multi-view display.
[0064] As shown, the multi-view display 600 is configured to provide or emit light as emitted light 602. According to various examples and embodiments, the emitted light 602 may be used to illuminate and provide images using the multi-view display 600. For example, the emitted light 602 may be used to illuminate an array of light valves (e.g., light valves 630, described below) of the multi-view display 600.
[0065] According to various embodiments, directional emitted light 602″ may be provided during a multi-view mode, including a plurality of directional light beams having directions corresponding to different view directions of the multi-view image. Conversely, according to various embodiments, wide-angle emitted light 602′ may be provided during a 2D mode, which is largely omnidirectional and, more generally, has a cone angle that is larger than the cone angle of a view of the multi-view image or multi-view display associated with the multi-view display 600. The wide-angle emitted light 602′ is shown in FIG. 6A as a dashed arrow for ease of illustration. However, the dashed arrow representing the wide-angle emitted light 602′ does not imply any particular directionality of the emitted light 602, but instead merely represents, for example, the emission and propagation of light from the multi-view display 600. Similarly, FIGS. 6B and 6C illustrate the directional light beams of the directional emitted light 602″ as a plurality of diverging arrows. In various embodiments, the directional light beams may be or represent a light field.
[0066] As shown in FIGS. 6A-6C, the time-multiplexed multimode display 600 includes a wide-angle backlight 610. The illustrated wide-angle backlight 610 has a planar or substantially planar light-emitting surface configured to provide wide-angle emitted light 602′ (see, for example, FIG. 6A). According to various embodiments, the wide-angle backlight 610 may be virtually any backlight having a light-emitting surface 610′ configured to provide light to illuminate an array of light valves in the display. For example, the wide-angle backlight 610 may be a direct-emitting or directly-illuminated flat backlight. Direct-emitting or directly-illuminated flat backlights include, but are not limited to, backlight panels that use a planar array of cold cathode fluorescent lamps (CCFLs), neon lamps, or light-emitting diodes (LEDs) configured to directly illuminate the planar light-emitting surface 610′ and provide wide-angle emitted light 602′. An electroluminescent panel (ELP) is another non-limiting example of a direct-emitting flat backlight. In other examples, wide-angle backlight 610 may comprise a backlight that uses indirect light sources. Such indirectly lit backlights may include, but are not limited to, various forms of edge-coupled or so-called "edge-lit" backlights.
[0067] The multi-view display 600 shown in Figures 6A-6C further includes a multi-view backlight 620. As shown, the multi-view backlight 620 includes an array of multi-beam elements 622. According to various embodiments, the multi-beam elements 622 of the multi-beam element array are spaced apart from one another across the multi-view backlight 620. Different types of multi-beam elements 622 may be utilized within the multi-view backlight 620, including, but not limited to, active emitters and various scattering elements. According to various embodiments, each multi-beam element 622 of the multi-beam element array is configured to provide multiple directional light beams having directions corresponding to different viewing directions of the multi-view image during a multi-view mode.
[0068] In some embodiments (e.g., as shown), the multi-view backlight 620 further comprises a light guide 624 configured to guide the light as guided waves. In some embodiments, the light guide 624 may be a plate light guide. According to various embodiments, the light guide 624 is configured to guide the guided light along the length of the light guide 624 according to total internal reflection. The schematic propagation direction of the guided light within the light guide 624 is shown by the bold arrow in FIG. 6B. In some embodiments, the guided light may be directed in the propagation direction at a non-zero propagation angle and may include collimated light that is collimated according to a predetermined collimation factor σ, as shown in FIG. 6B.
[0069] In embodiments including a light guide 624, the multi-beam elements 622 of the array of multi-beam elements may be configured to scatter a portion of the guided light from within the light guide 624 and direct the scattered portion away from the emitting surface to provide directional emitted light 602″, as shown in FIG. 6B. For example, the guided light portion may be scattered by the multi-beam elements 622 through a first surface. Further, as shown in FIGS. 6A-6C, according to various embodiments, a second surface of the multi-view backlight 620 opposite the first surface may be adjacent to the planar emitting surface of the wide-angle backlight 610. Furthermore, the multi-view backlight 620 may be substantially transparent (e.g., at least in a 2D mode) to allow wide-angle emitted light 602′ from the wide-angle backlight 610 to pass through or penetrate the thickness of the multi-view backlight 620, as shown in FIG. 6A by the dashed arrows originating in the wide-angle backlight 610 and subsequently passing through the multi-view backlight 620.
[0070] In some embodiments (e.g., as shown in FIGS. 6A-6C ), the multi-view backlight 620 may further comprise a light source 626. As such, the multi-view backlight 620 may be, for example, an edge-lit backlight. According to various embodiments, the light source 626 is configured to provide light that is guided within the light guide 624. In various embodiments, the light source 626 may include substantially any light source (e.g., light emitter), including, but not limited to, one or more light-emitting diodes (LEDs) or lasers (e.g., laser diodes). In some embodiments, the light source 626 may include a light emitter configured to generate substantially monochromatic light having a narrowband spectrum denoted by a particular color. In particular, the color of the monochromatic light may be a primary color of a particular color space or color model (e.g., the red-green-blue (RGB) color model). In other examples, the light source 626 may be a substantially broadband light source configured to provide substantially broadband or polychromatic light. For example, the light source 626 may provide white light. In some embodiments, the light source 626 may include a plurality of different light emitters configured to provide light of different colors. The different light emitters may be configured to provide light having different color-specific non-zero propagation angles of the guided light corresponding to each of the different colors of light. As shown in FIG. 6B, operating the multi-view backlight 620 may include operating the light source 626.
[0071] According to some embodiments (e.g., shown in FIGS. 6A-6C ), the multi-beam elements 622 of the multi-beam element array may be disposed on a first surface of the light guide 624 (e.g., adjacent to a first surface of the multi-view backlight 620). In other embodiments (not shown), the multi-beam elements 622 may be disposed within the light guide 624. In yet other embodiments (not shown), the multi-beam elements 622 may be disposed on or at a second surface of the light guide 624 (e.g., adjacent to a second surface of the multi-view backlight 620). Furthermore, the size of the multi-beam elements 622 corresponds to the size of the light valves of the multi-view display 600. In some embodiments, the size of the multi-beam elements 622 may be about one-quarter to about two times the size of the light valves.
[0072] As described above and illustrated in FIGS. 6A-6C , the multi-view display 600 further comprises an array of light valves 630. In various embodiments, any of a variety of different types of light valves may be used as the light valves 630 of the light valve array, including, but not limited to, one or more of a liquid crystal light valve, an electrophoretic light valve, and a light valve based on or using electrowetting. Furthermore, as illustrated, there may be a unique plurality of light valves 630 for each multi-beam element 622 of the array of multi-beam elements. The unique plurality of light valves 630 may correspond, for example, to the multi-view pixels of the time-multiplexed multimode display 600. According to some embodiments, the comparable sizes of the multi-beam elements 622 and light valves 630 may be selected to reduce, or in some instances minimize, dark zones between views of the multi-view display, while also reducing, or in some instances minimize, overlap between views of the multi-view display or equivalent multi-view images.
[0073] According to various embodiments, the multi-beam element 622 of the multi-view backlight 620 may comprise any of several different structures configured to scatter portions of the guided light. For example, the different structures may include, but are not limited to, a diffraction grating, a micro-reflective element, a micro-refractive element, or various combinations thereof. In some embodiments, a multi-beam element 622 comprising a diffraction grating is configured to diffractively combine or scatter the guided light portions as directional emitted light 602″ comprising multiple directional light beams having different principal angular directions. In some embodiments, the diffraction grating of the multi-beam element may include multiple individual sub-gratings. In other embodiments, a multi-beam element 622 comprising a micro-reflective element is configured to reflectively combine or scatter the guided light portions as multiple directional light beams, or a multi-beam element 622 comprising a micro-refractive element is configured to combine or scatter the guided light portions as multiple directional light beams by or using refraction (i.e., refractively scatter the guided light portions).
[0074] In some embodiments, the light field camera system 500 of Figure 5 may be implemented in or using a client device. Figure 7 shows a block diagram of an example client device 700, according to an embodiment consistent with principles described herein. The light field camera system 500 may comprise, for example, the client device 700. For example, the processor 520 and memory 530 of the light field camera system 500 may be part of the client device 700.
[0075] As illustrated, client device 700 may comprise a system of components that perform various computing operations for a user of client device 700. Client device 700 may be a laptop, tablet, smartphone, touchscreen system, intelligent display system, or other client device. Client device 700 may include various components, such as a processor 710, memory 720, input / output (I / O) components 730, a display 740, and potentially other components. These components may be coupled to a bus 750 that functions as a local interface that allows the components of client device 700 to communicate with each other. While the components of client device 700 are shown as housed within client device 700, it should be appreciated that at least some of the components may be coupled to client device 700 via external connections. For example, components may be externally plugged into or otherwise connected to client device 700 via external ports, sockets, plugs, or connectors.
[0076] The processor 710 may be a central processing unit (CPU), a graphics processing unit (GPU), any other integrated circuit that performs computing operations, or any combination thereof. The processor 710 may include one or more processing cores. The processor 710 comprises circuitry that executes instructions. The instructions include, for example, computer code, programs, logic, or other machine-readable instructions that are received and executed by the processor 710 to perform the computing functions embodied in the instructions. The processor 710 may execute instructions to operate on data. For example, the processor 710 may receive input data (e.g., an image), process the input data according to an instruction set, and generate output data (e.g., a processed image). As another example, the processor 710 may receive instructions and generate new instructions for subsequent execution. The processor 710 may comprise hardware that implements a graphics pipeline that renders output from a renderer. For example, the processor 710 may comprise one or more GPU cores, vector processors, scalar processors, or hardware accelerators.
[0077] Memory 720 may include one or more memory components. Memory 720 is defined herein to include either or both volatile and non-volatile memory. Volatile memory components are memory components that do not retain information upon loss of power. Volatile memory may include, for example, random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), or other volatile memory structures. System memory (e.g., main memory, cache, etc.) may be implemented using volatile memory. System memory refers to high-speed memory that can temporarily store data or instructions for quick read and write access to support processor 710.
[0078] Nonvolatile memory components are memory components that retain information upon loss of power. Nonvolatile memory includes read-only memory (ROM), hard disk drives, solid-state drives, USB flash drives, memory cards accessed via memory card readers, floppy disks accessed via associated floppy disk drives, optical disks accessed via optical disk drives, and magnetic tapes accessed via appropriate tape drives. ROM may include, for example, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other similar memory devices. Storage memory may be implemented using nonvolatile memory to provide long-term retention of data and instructions.
[0079] Memory 720 may refer to a combination of volatile and non-volatile memory used to store instructions and data. For example, data and instructions may be stored in non-volatile memory and loaded into volatile memory for processing by processor 710. Execution of instructions may include, for example, a compiled program that is loaded from non-volatile memory into volatile memory and then converted into machine code in a format that can be executed by processor 710; source code that is converted into an appropriate format, for example, object code that can be loaded into volatile memory for execution by processor 710; or source code that is interpreted by another executable program to generate instructions in volatile memory and executed by processor 710. Instructions may be stored in or loaded into any portion or component of memory 720, including, for example, RAM, ROM, system memory, storage, or any combination thereof.
[0080] Although memory 720 is shown as separate from other components of client device 700, it should be appreciated that memory 720 may be incorporated or otherwise integrated, at least in part, into one or more components. For example, processor 710 may include on-board memory registers or cache for performing processing operations.
[0081] The I / O components 730 may include, for example, a touchscreen, speaker, microphone, buttons, switches, dials, cameras, sensors, accelerometers, or other components that receive user input or generate output directed to a user. The I / O components 730 may receive user input and convert it into data for storage in memory 720 or for processing by processor 710. The I / O components 730 may receive data output by memory 720 or processor 710 and convert it into a format that is perceivable by the user (e.g., sound, tactile response, visual information, etc.). The I / O components 730 may include one or more physical cameras coupled to the client device. The client device 700 may control the baselines of the cameras as well as their focusing capabilities.
[0082] A particular type of I / O component 730 is a display 740. The display 740 may include a multi-view display (e.g., multi-view display 100), a multi-view display combined with a 2D display, or any other display that presents images. A capacitive touchscreen layer functioning as the I / O component 730 may be layered within the display to allow a user to provide input while simultaneously perceiving visual output. The processor 710 may generate data formatted as an image for presentation on the display 740. The processor 710 may execute instructions to render images on the display for perception by the user.
[0083] Bus 750 facilitates communication of instructions and data between processor 710, memory 720, I / O components 730, display 740, and any other components of client device 700. Bus 750 may include address translators, address decoders, fabric, conductive traces, conductive wires, ports, plugs, sockets, and other connectors to enable communication of data and instructions.
[0084] The instructions in memory 720 may be embodied in various forms in a manner that implements at least a portion of a software stack. For example, the instructions may be embodied as an operating system 722, applications 724, device drivers (e.g., display driver 726), firmware (e.g., display firmware 728), or other software components. The operating system 722 is a software platform that supports basic functions of the client device 700, such as scheduling tasks, controlling I / O components 730, providing access to hardware resources, managing power, and supporting applications 724.
[0085] The applications 724 execute on the operating system 722 and can access the hardware resources of the client device 700 through the operating system 722. In this regard, the execution of the applications 724 is controlled at least in part by the operating system 722. The applications 724 may be user-level software programs that provide high-level functionality, services, and other features to a user. In some embodiments, the applications 724 may be dedicated “apps” that are downloadable or otherwise accessible to a user on the client device 700. A user may launch the applications 724 through a user interface provided by the operating system 722. The applications 724 may be developed by developers and defined in various source code formats. The applications 724 may be developed using several programming or scripting languages, such as C, C++, C#, Objective C, Java, Swift, JavaScript, Perl, PHP, Visual Basic, Python, Ruby, Go, or other programming languages. The applications 724 may be compiled into object code by a compiler or interpreted by an interpreter for execution by the processor 710. The application 724 may include a renderer or other graphics rendering engine.
[0086] Device drivers, such as display driver 726, contain instructions that enable operating system 722 to communicate with various I / O components 730. Each I / O component 730 may have its own device driver. Device drivers may be installed such that they are stored in storage and loaded into system memory. For example, upon installation, display driver 726 translates high-level display instructions received from operating system 722 into low-level instructions that are executed by display 740 to display images.
[0087] Firmware, such as display firmware 728, may include machine code or assembly code that enables I / O components 730 or display 740 to perform low-level operations. Firmware can translate the electrical signals of particular components into higher-level instructions or data. For example, display firmware 728 can control how display 740 activates individual pixels at a low level by adjusting voltage or current signals. Firmware may be stored in and executed directly from non-volatile memory. For example, display firmware 728 may be embodied in a ROM chip coupled to display 740 such that the ROM chip is isolated from other storage and system memory of client device 700. Display 740 may include processing circuitry for executing display firmware 728.
[0088] The operating system 722, applications 724, drivers (e.g., display driver 726), firmware (e.g., display firmware), and potentially other instruction sets may each include instructions executable by the processor 710 or other processing circuitry of the client device 700 to perform the functions and operations described above. While the instructions described herein may be embodied in software or code executed by the processor 710 as described above, the instructions may alternatively be embodied in dedicated hardware or a combination of software and dedicated hardware. For example, the functions and operations performed by the instructions described above may be implemented as circuits or state machines using any one or a combination of several technologies. These technologies may include, but are not limited to, discrete logic circuits having logic gates for performing various logical functions upon the application of one or more data signals, application specific integrated circuits (ASICs) having appropriate logic gates, field programmable gate arrays (FPGAs), or other components, etc.
[0089] In some embodiments of the described principles, a non-transitory computer-readable storage medium is provided that stores executable instructions that, when executed by a processor of a computer system, perform operations for determining a baseline and convergence distance of a multi-eye camera rig. In particular, instructions that perform the functions and operations described above may be embodied in a non-transitory computer-readable storage medium. For example, some embodiments may be directed to a non-transitory computer-readable storage medium that stores executable instructions that, when executed by a processor (e.g., processor 710) of a computing system (e.g., client device 700), cause the processor to perform various functions described above, including various operations for dynamically and automatically updating the convergence distance or baseline of a multi-eye camera rig.
[0090] In particular, the operations performed by the processor executing instructions stored on a non-transitory computer-readable storage medium may include determining a set or plurality of sample point depths representing distances between a multi-eye camera rig and a plurality of sample points in a three-dimensional (3D) scene within a field of view of the multi-eye camera rig, and the convergence distance is set as an average depth of the plurality of sample point depths. The operations may further include determining a minimum sample point depth of the plurality of sample point depths, and the baseline is a function of the difference between the minimum sample point depth and the convergence distance. Negative In some embodiments, determining the depths of the plurality of sample points includes calculating the depths from a depth map of an image representing the 3D scene using a disparity map, Inside and measuring a distance between the multi-eye camera rig and an object in the 3D scene using a depth sensor, the object corresponding to a sample point of the plurality of sample points. In some embodiments, the average depth of the plurality of sample point depths is a weighted average, and weights in the weighted average are assigned according to a decreasing function of distance from a focal point of the 3D scene.
[0091] As used herein, a "non-transitory computer-readable storage medium" is defined as any medium capable of containing, storing, or maintaining instructions described herein for use by or in connection with an instruction execution system. For example, a non-transitory computer-readable storage medium may store instructions for use by or in connection with the light field camera system 500 or the client device 700. Furthermore, a non-transitory computer-readable storage medium may or may not be part of the client device 700 described above (e.g., part of memory 720). Instructions stored by a non-transitory computer-readable storage medium may include, but are not limited to, statements, code, or declarations that may be fetched from the non-transitory computer-readable medium and executed by a processing circuit (e.g., processor 520 or processor 710). Furthermore, the term "non-transitory computer-readable storage medium," as defined herein, explicitly excludes transitory media, including, for example, carrier waves.
[0092] According to various embodiments, the non-transitory computer-readable medium may include any one of many physical media, such as magnetic, optical, or semiconductor media. More specific examples of suitable non-transitory computer-readable media may include, but are not limited to, magnetic tape, magnetic floppy diskettes, magnetic hard drives, memory cards, solid-state drives, USB flash drives, or optical disks. The non-transitory computer-readable medium may also be random access memory (RAM), including, for example, static random access memory (SRAM) and dynamic random access memory (DRAM), or magnetic random access memory (MRAM). In addition, the non-transitory computer-readable medium may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other types of memory devices.
[0093] The client device 700 may perform any of the operations or implement the functions described above. For example, the process flows described above may be performed by the client device 700 executing instructions and processing data. Although the client device 700 is shown as a single device, embodiments are not so limited. In some embodiments, the client device 700 may offload processing of instructions in a distributed manner such that multiple other client devices 700 or other computing devices work together to execute instructions that may be stored or loaded in a distributed arrangement. For example, at least some instructions or data may be stored, loaded, or executed within a cloud-based system operating in conjunction with the client device 700.
[0094] Thus, examples and embodiments have been described that apply to light field camera systems to set the baseline and convergence distance of a multi-view camera rig. In some embodiments, the baseline and convergence distance may be determined dynamically or in real time based on the depth of points within the camera views. It should be understood that the above examples are merely illustrative of some of the many specific examples that illustrate the principles described herein. Clearly, those skilled in the art can readily devise numerous other configurations without departing from the description provided herein. It should be noted that the present specification discloses the following aspects. [Aspect 1] 1. A method for establishing a baseline and convergence distance for a multi-lens camera rig, comprising: determining a plurality of sample point depths representing distances between the multi-view camera rig and a plurality of sample points in a three-dimensional scene within a field of view of the multi-view camera rig; setting the convergence distance to be an average sample point depth of the plurality of sample point depths; determining a minimum sample point depth of the plurality of sample point depths; setting the baseline to be the negative reciprocal of the difference between the minimum sample point depth and the convergence distance; A method for setting a baseline and a convergence distance, including: [Aspect 2] The method for setting a baseline and convergence distance according to aspect 1, wherein the step of determining the depths of the plurality of sample points includes the steps of performing ray casting within the three-dimensional scene and recording a hit distance for each sample point of the plurality of sample points. [Aspect 3] A method for setting a baseline and convergence distance according to aspect 1, wherein the step of determining the depths of the plurality of sample points includes a step of calculating the depths of the sample points from a disparity map of the three-dimensional scene for each sample point of the plurality of sample points. [Aspect 4] The method for setting a baseline and convergence distance of aspect 3, wherein the step of calculating sample point depth further includes the steps of using image disparity between images captured by different cameras of the multi-lens camera rig and forming a disparity map of the three-dimensional scene from the image disparity. [Aspect 5] The method for setting a baseline and convergence distance according to aspect 1, wherein the step of determining the depths of the plurality of sample points includes using a depth sensor to measure the distance between the multi-lens camera rig and an object in the three-dimensional scene, the object corresponding to a sample point among the plurality of sample points. [Aspect 6] A method for setting a baseline and convergence distance according to aspect 5, wherein the depth sensor comprises one of a laser distance sensor and a time-of-flight distance sensor. [Aspect 7] A method for setting a baseline and convergence distance according to aspect 1, wherein the average sample point depth is a weighted average having weights assigned according to a decreasing function of distance from a focal point within the three-dimensional scene. [Aspect 8] 8. The method for setting a baseline and a convergence distance according to aspect 7, wherein the weights of the weighted average are assigned according to a Gaussian distribution centered on the focal point. [Aspect 9] 2. The method for setting a baseline and convergence distance according to claim 1, wherein the multi-lens camera rig comprises a plurality of virtual cameras and the three-dimensional scene is a three-dimensional model. [Aspect 10] The method for setting a baseline and convergence distance according to aspect 1, wherein the multi-lens camera rig comprises a plurality of physical cameras, and the three-dimensional scene represents a physical scene captured by one of the plurality of physical cameras. [Aspect 11] The method for setting a baseline and convergence distance according to aspect 1, wherein the multi-camera rig comprises cameras that move between multiple positions to capture images that form the three-dimensional scene. [Aspect 12] a multi-lens camera rig having a plurality of cameras; a processor; When executed by the processor, determining a plurality of sample point depths representing distances between the multi-view camera rig and a plurality of sample points in a three-dimensional scene within a field of view of the multi-view camera rig; Set the convergence distance and baseline of the multi-lens camera a memory configured to store instructions; Equipped with The light field camera system, wherein the convergence distance is an average sample point depth of the plurality of sample point depths, and the baseline is a minimum sample point depth of the plurality of sample point depths minus a negative reciprocal of the convergence distance. [Aspect 13] A light field camera system as described in aspect 12, wherein the multiple sample point depths are determined by one or more of: calculating the sample point depths using a disparity map from a depth map of an image representing the three-dimensional scene; or performing ray casting within the three-dimensional scene to determine the sample point depths within the three-dimensional scene. [Aspect 14] A light field camera system as described in aspect 12, wherein the plurality of sample point depths are determined by using a depth sensor to measure the distance between the multi-lens camera rig and an object in the three-dimensional scene, and the object corresponds to a sample point among the plurality of sample points. [Aspect 15] A light field camera system as described in aspect 12, wherein the average sample point depth of the multiple sample point depths is a weighted average, and weights of the weighted average are assigned according to a decreasing function of distance from a focus of the three-dimensional scene. [Aspect 16] 13. The light field camera system of claim 12, wherein a camera among the plurality of cameras is a virtual camera and the three-dimensional scene is a three-dimensional model. [Aspect 17] 13. The light field camera system of claim 12, wherein a camera among the plurality of cameras of the multi-lens camera rig includes a physical camera, and the three-dimensional scene represents a physical scene captured by the physical camera. [Aspect 18] 13. The light field camera system of claim 12, further comprising a multi-view display, wherein the convergence distance corresponds to a zero parallax plane of the multi-view display. [Aspect 19] 1. A non-transitory computer-readable storage medium storing executable instructions that, when executed by a processor of a computer system, perform operations for determining a baseline and convergence distance of a multi-eye camera rig, the operations comprising: determining a plurality of sample point depths representing distances between the multi-eye camera rig and a plurality of sample points in a three-dimensional scene within a field of view of the multi-eye camera rig, wherein the convergence distance is set as an average sample point depth of the plurality of sample point depths; determining a minimum sample point depth of the plurality of sample point depths, the baseline being set as the negative reciprocal of the difference between the minimum sample point depth and the convergence distance; 1. A non-transitory computer-readable storage medium comprising: [Aspect 20] A non-transitory computer-readable storage medium as described in aspect 19, wherein determining the plurality of sample point depths includes one or more of: calculating the sample point depths using a disparity map from a depth map of an image representing the three-dimensional scene; performing ray casting within the three-dimensional scene to determine the sample point depths within the three-dimensional scene; and using a depth sensor to measure the distance between the multi-eye camera rig and an object within the three-dimensional scene, wherein the object corresponds to a sample point among the plurality of sample points. [Aspect 21] 20. The non-transitory computer-readable storage medium of claim 19, wherein the average sample point depth of the plurality of sample point depths is a weighted average, and weights of the weighted average are assigned according to a decreasing function of distance from a focus of the three-dimensional scene. [Explanation of symbols]
[0095] 100 Multi-View Displays 110 Multi-view images 112 views 120 view directions 200 three-dimensional (3D) scenes 202 Thu 204 Rock 208 Ground 210 Camera 220 frustum 230 Convergence Surface 300 ways 310 steps 320 steps 330 steps 340 steps 402 3D Scenes 404 Object 410 Multi-camera rig 420 Rays 500 Light Field Camera System 510 Multi-camera rig 520 processor 530 memory 540 Multi-View Display 600 Multiview Display 602 Synchrotron Radiation 602' Wide-angle Synchrotron Radiation 602” directional synchrotron radiation 610 Wide Angle Backlight 610' luminous surface 620 Multi-View Backlight 622 Multi-beam element 624 Light guide 626 light source 630 Light Bulb 700 client devices 710 processor 720 memory 722 Operating Systems 724 Applications 726 display driver 728 Display Firmware 730 I / O components 740 display 750 Bus b Baseline D conv Convergence Distance z min Minimum Sample Point Depth θ elevation angle component φ Sitting angle component σ Collimation Factor
Claims
1. 1. A method for establishing a baseline and convergence distance for a multi-lens camera rig, comprising: determining a plurality of sample point depths representing distances between the multi-view camera rig and a plurality of sample points in a three-dimensional scene within a field of view of the multi-view camera rig; setting the convergence distance to be an average sample point depth of the plurality of sample point depths; determining a minimum sample point depth of the plurality of sample point depths; setting the baseline to be the negative reciprocal of the difference between the minimum sample point depth and the convergence distance; Including, the plurality of sample points are all sample points in a three-dimensional scene; How to set the baseline and convergence distance.
2. 2. The method of claim 1, wherein determining the depths of the plurality of sample points comprises performing ray casting within the three-dimensional scene and recording a hit distance for each sample point of the plurality of sample points.
3. The method for setting a baseline and convergence distance according to claim 1 , wherein determining the plurality of sample point depths comprises calculating the sample point depth for each sample point of the plurality of sample points from a disparity map of the three-dimensional scene.
4. 4. The method of claim 3, wherein calculating sample point depths further comprises using image disparity between images captured by different cameras of the multi-camera rig, and forming a disparity map of the three-dimensional scene from the image disparity.
5. 2. The method for setting a baseline and convergence distance according to claim 1, wherein the step of determining the depths of the plurality of sample points includes using a depth sensor to measure a distance between the multi-lens camera rig and an object in the three-dimensional scene, the object corresponding to a sample point among the plurality of sample points.
6. The method of establishing a baseline and convergence distance of claim 5 , wherein the depth sensor comprises one of a laser distance sensor and a time-of-flight distance sensor.
7. The method of claim 1 , wherein the average sample point depth is a weighted average with weights assigned according to a decreasing function of distance from a focal point within the three-dimensional scene.
8. The method of claim 7 , wherein the weights of the weighted average are assigned according to a Gaussian distribution centered on the focal point.
9. The method of claim 1 , wherein the multi-camera rig comprises a plurality of virtual cameras and the three-dimensional scene is a three-dimensional model.
10. The method for establishing a baseline and convergence distance according to claim 1 , wherein the multi-camera rig comprises a plurality of physical cameras, and the three-dimensional scene represents a physical scene captured by one of the plurality of physical cameras.
11. The method of establishing a baseline and convergence distance of claim 1 , wherein the multi-camera rig comprises cameras that move between multiple positions to capture images forming the three-dimensional scene.
12. a multi-lens camera rig having a plurality of cameras; a processor; When executed by the processor, determining a plurality of sample point depths representing distances between the multi-view camera rig and a plurality of sample points in a three-dimensional scene within a field of view of the multi-view camera rig; Set the convergence distance and baseline of the multi-camera rig a memory configured to store instructions; Equipped with the convergence distance is an average sample point depth of the plurality of sample point depths, and the baseline is a minimum sample point depth of the plurality of sample point depths minus the negative reciprocal of the convergence distance; the plurality of sample points are all sample points in a three-dimensional scene; Light field camera system.
13. 13. The light field camera system of claim 12, wherein the plurality of sample point depths are determined by one or more of: calculating the sample point depths from a depth map of an image representing the three-dimensional scene using a disparity map; or performing ray casting within the three-dimensional scene to determine the sample point depths within the three-dimensional scene.
14. 13. The light field camera system of claim 12, wherein the plurality of sample point depths are determined by using a depth sensor to measure a distance between the multi-eye camera rig and an object in the three-dimensional scene, the object corresponding to a sample point among the plurality of sample points.
15. 13. The light field camera system of claim 12, wherein the average sample point depth of the plurality of sample point depths is a weighted average, weights of the weighted average being assigned according to a decreasing function of distance from a focal point of the three-dimensional scene.
16. The light field camera system of claim 12 , wherein a camera of the plurality of cameras is a virtual camera and the three-dimensional scene is a three-dimensional model.
17. The light field camera system of claim 12 , wherein a camera of the plurality of cameras of the multi-lens camera rig includes a physical camera, and the three-dimensional scene represents a physical scene imaged by the physical camera.
18. The light field camera system of claim 12 , further comprising a multi-view display, wherein the convergence distance corresponds to a zero parallax plane of the multi-view display.
19. 1. A non-transitory computer-readable storage medium storing executable instructions that, when executed by a processor of a computer system, perform operations for determining a baseline and convergence distance of a multi-eye camera rig, the operations comprising: determining a plurality of sample point depths representing distances between the multi-eye camera rig and a plurality of sample points in a three-dimensional scene within a field of view of the multi-eye camera rig, the convergence distance being set as an average sample point depth of the plurality of sample point depths; determining a minimum sample point depth of the plurality of sample point depths, the baseline being set as the negative reciprocal of the difference between the minimum sample point depth and the convergence distance; Including, the plurality of sample points are all sample points in a three-dimensional scene; A non-transitory computer-readable storage medium.
20. 20. The non-transitory computer-readable storage medium of claim 19, wherein determining the plurality of sample point depths includes one or more of: calculating the sample point depths using a disparity map from a depth map of an image representing the three-dimensional scene; performing ray casting within the three-dimensional scene to determine the sample point depths within the three-dimensional scene; and using a depth sensor to measure a distance between the multi-eye camera rig and an object within the three-dimensional scene, wherein the object corresponds to a sample point among the plurality of sample points.
21. 20. The non-transitory computer-readable storage medium of claim 19, wherein the average sample point depth of the plurality of sample point depths is a weighted average, and weights of the weighted average are assigned according to a decreasing function of distance from a focal point of the three-dimensional scene.
Citation Information
Patent Citations
Means for setting two-viewpoint camera
JP2003348621A
Three-dimensional image processing method and three-dimensional image processor
JP2005353047A
System and method for extracting depth from images using motion compensation
JP2011525657A
Three-dimensional video signal photography device
JP2012220603A
3D video photography control system, 3D video photography control method, and program
JP2013105002A