Multi-view image capture system and method
The multi-view camera system dynamically adjusts inter-camera focusing and field of view to optimize multi-view image capture and rendering, enhancing the immersive three-dimensional experience by ensuring optimal disparity and depth perception.
Patent Information
- Application Number
- JP2024510626
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-26
- Publication Date
- 2025-10-27
- Estimated Expiration
- 2041-08-26
AI Technical Summary
Existing multi-view display technologies struggle to optimize the capture and rendering of multi-view images to achieve a target disparity level, resulting in suboptimal viewing experiences.
A multi-view camera system dynamically adjusts the inter-camera focusing distance and field of view to match a target camera reference value, optimizing the capture process for a specific disparity level, and a multi-view display system uses a light guide and multibeam elements to render multiple views simultaneously without the need for special glasses.
The system provides a more immersive three-dimensional viewing experience by ensuring optimal parallax and depth perception, allowing viewers to interactively explore multi-view images with enhanced depth perception and clarity.
Smart Images

Figure 0007760707000003 
Figure 0007760707000004 
Figure 0007760707000005
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS Not applicable
[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT Not applicable [Background technology]
[0003] Multi-view display is an emerging display technology that offers a more immersive viewing experience compared to traditional 2D content. Multi-view content may include a still image or a series of video frames, with each multi-view image (e.g., picture or frame) containing a different view of a scene. A multi-view image may be composed of multiple different-view images, corresponding to different viewpoints of a scene. The different-view images may be captured by one or more cameras, or may be synthesized by extrapolating or interpolating from other view images.
[0004] Various features of examples and embodiments consistent with the principles described herein may be more readily understood by reference to the following detailed description in conjunction with the accompanying drawings, in which like numerals represent like structural elements. [Brief explanation of the drawings]
[0005] [Figure 1] FIG. 1 illustrates a multi-view image in an example according to an embodiment consistent with principles described herein. [Figure 2] FIG. 1 illustrates an example of a multi-view display, according to one embodiment consistent with principles described herein. [Figure 3]3A and 3B illustrate a multi-view imaging system including multiple cameras shown from a top view and a side view, respectively, according to one embodiment consistent with principles described herein. [Figure 4] FIG. 1 illustrates a computing system coupled to a camera system of a multi-view imaging system, according to one embodiment consistent with principles described herein. [Figure 5] 10A-10C illustrate examples of dynamic adjustment of inter-camera focusing distance by individually selecting a subset of cameras, according to one embodiment consistent with principles described herein. [Figure 6] FIG. 10 illustrates an example of dynamic adjustment of inter-camera focusing distance by individually moving each camera in a subset of cameras along a track, according to one embodiment consistent with principles described herein. [Figure 7] 1A-1C illustrate examples of modifying the field of view by applying a cropping window to each view image, according to one embodiment consistent with principles described herein. [Figure 8] A diagram showing an example in which data for a multi-view image is converted into frames of a multi-view video stream that are rendered in real time by a multi-view display device, according to one embodiment consistent with the principles described herein. [Figure 9] 1 is a flowchart providing an example of multi-view image capture according to one embodiment in accordance with principles described herein. [Figure 10] 1 illustrates an example of a multi-view image capture system, where the camera system comprises an unmanned aerial vehicle, according to one embodiment consistent with principles described herein. [Figure 11] FIG. 1 is a schematic block diagram illustrating an exemplary illustration of a computing system according to one embodiment in accordance with principles described herein. DETAILED DESCRIPTION OF THE INVENTION
[0006] Certain examples and embodiments have other features either in addition to or in place of the features shown in the above-referenced figures, which and other features are described in detail below with reference to the above-referenced figures.
[0007] Examples and embodiments consistent with the principles described herein provide techniques for capturing a scene as one or more multi-view images. When capturing the multi-view images, the techniques involve optimizing the capture process to achieve a target display level for a target multi-view display. Specifically, a multi-view display may be designed to render the multi-view images at a particular disparity level. The angular range of views presented by the multi-view display may define a preferred disparity level allowed by the multi-view display. To generate multi-view content tailored to the disparity level of the target multi-view display, a multi-view camera system may dynamically adjust the inter-camera shooting distance between cameras of a subset of cameras to match a target camera reference value, which provides a predetermined disparity level for the rendered multi-view images.
[0008] In some embodiments, a photographer may adjust various visual parameters (e.g., field of view, zoom, camera position, distance between the camera system and the object, etc.), and a target camera reference value is calculated and dynamically applied by the camera system to change the inter-camera focusing distance between the cameras of the camera system accordingly. For example, when a photographer provides input to cause the cameras of the camera system to zoom in on or out of an object being photographed, the inter-camera focusing distance dynamically changes to achieve a target camera reference value consistent with the optimal parallax level of the multi-view display. Some embodiments use multiple cameras that dynamically change the inter-camera focusing distance. Other embodiments are directed to a single camera that orbits an object and captures images of the object at multiple different locations or from multiple different viewpoints to generate a multi-view image. This single camera may be an unmanned aerial vehicle (e.g., a drone). The inter-camera focusing distance may be dynamically adjusted as images are captured by the single camera by determining a timestamp interval for each captured image.
[0009] The captured multi-view content (i.e., images of the set of images that make up the multi-view image) may be stored or streamed for rendering in real time on a target multi-view display. As used herein, "real time" is defined as the time it takes a system to process input and generate output. For example, storing data only for the purpose of processing it at a later time is not real time. As mentioned above, the camera reference values for the multi-view content are optimized for the preferred (e.g., typical or ideal) disparity level allowed by the receiving multi-view display.
[0010] FIG. 1 illustrates an example multi-view image according to one embodiment in accordance with the principles described herein. The multi-view image 103 may be a still image or a video frame from a multi-view video stream. The multi-view image 103 includes multiple view images 106 (e.g., views). Each view image 106 of the multiple view images corresponds to a view (e.g., left view, right view, etc.) of a scene or object from multiple different principal angular directions 109. The view images 106 are rendered on a multi-view display 112. Each view image 106 represents a different viewing angle of the scene or object in the multi-view image 103. Thus, the different view images 106 have some degree of parallax relative to each other. A viewer can perceive one view image 106 with their right eye and another view image 106 with their left eye simultaneously. This allows the viewer to perceive the different view images 106 simultaneously, thereby experiencing a three-dimensional (3D) effect. The multi-view display 112 may be capable of supporting a maximum or required disparity level, which may be defined as a percentage of the display width between adjacent views. For example, the multi-view display 112 may have a predetermined disparity level of approximately 1 percent. In other examples, the predetermined disparity level may be less than 1 percent (e.g., between 0.25 percent and 1 percent) or greater than 1 percent (e.g., between 1 percent and 10 percent). The predetermined disparity level may also be defined as the number of pixels that create disparity between views. For example, if a multi-view display is 1000 pixels wide and is capable of presenting views with a disparity of approximately 10 pixels, the multi-view display will have a predetermined disparity level of approximately 1 percent, where the pixel distance of the disparity is the pixel distance of the multi-view display.
[0011] In some embodiments, as a viewer physically changes their viewing angle relative to the multi-view display 112, the viewer's eyes may perceive different views of the multi-view image 103. As a result, the viewer can interact with the multi-view display 112 to see different view images 106 of the multi-view image 103. For example, as the viewer moves to the left, the viewer can see more of the left side of the scene in the multi-view image 103. The multi-view image 103 may have multiple view images 106 along a horizontal plane, or multiple view images 106 along a vertical plane, or both. Thus, as the user changes their viewing angle to see different view images 106, the viewer can obtain additional visual details of the scene represented in the multi-view image 103.
[0012] As discussed above, each view image 106 is presented by the multi-view display 112 with a different corresponding principal angular direction 109. When the multi-view images 103 are presented for display, the view images 106 may actually appear on or near the multi-view display 112. In this regard, the rendered multi-view content is sometimes referred to as light field content. One feature of viewing light field content is the ability to view different views simultaneously. Light field content includes visual images that may appear both in front of and behind the screen, conveying a sense of depth to the viewer.
[0013] A 2D display may be substantially similar to a multi-view display 112, except that a 2D display is generally configured to provide a single view (e.g., only one of the views) as opposed to multiple different views of a multi-view image 103. As used herein, a "two-dimensional display" or "2D display" is defined as a display configured to provide a view of an image that is substantially the same regardless of the direction from which the image is viewed (i.e., within a predetermined viewing angle or range of the 2D display). A conventional liquid crystal display (LCD) found on many smartphones and computer monitors is an example of a 2D display. In contrast, as used herein, a "multi-view display" is defined as an electronic display or display system configured to simultaneously provide different views of a multi-view image at or from different viewing directions from a user's perspective. In particular, the different view images 106 may represent how the multi-view image 103 appears from different perspectives.
[0014] The multi-view display 112 may be implemented using a variety of technologies capable of presenting multiple different image views so that the image views are perceived simultaneously. One example of a multi-view display utilizes a multi-beam element to scatter light in controlled principal angular directions corresponding to the viewing directions of the different view images 106. According to some embodiments, the multi-view display 112 may be a light field display, which is a display that presents multiple beams of light of different colors and directions corresponding to the different views. In some examples, the light field display is a so-called "glasses-free" three-dimensional (3D) display that may use a multi-beam element (e.g., a diffraction grating) to provide an autostereoscopic representation of the multi-view image without the need to wear special glasses for depth perception.
[0015] FIG. 2 illustrates an example of a multi-view display, according to one embodiment consistent with the principles described herein. The multi-view display 112 may generate multi-view content when operating in a multi-view mode. In some embodiments, the multi-view display 112 renders not only 2D images but also multi-view images depending on its operating mode. For example, the multi-view display 112 may include multiple backlights for operating in various modes. The multi-view display 112 may be configured to provide wide-angle illumination using a wide-angle backlight 115 during the 2D mode. Additionally, the multi-view display 112 may be configured to provide directional illumination using a multi-view backlight 118 having an array of multi-beam elements during the multi-view mode, the directional illumination including multiple directional light rays provided by each multi-beam element of the array of multi-beam elements. In some embodiments, multi-view display 112 may be configured to time-multiplex the 2D mode and the multi-view mode using mode controller 121 to sequentially enable wide-angle backlight 115 during a first consecutive time interval corresponding to the 2D mode and multi-view backlight 118 during a second consecutive time interval corresponding to the multi-view mode. The direction of each directional ray of the directional light rays may correspond to a different viewing direction of multi-view image 103. Mode controller 121 may generate mode selection signal 124 to enable wide-angle backlight 115 or multi-view backlight 118.
[0016] In 2D mode, a wide-angle backlight 115 may be used to generate images such that the multi-view display 112 operates like a 2D display. By definition, "wide-angle" illumination is defined as light having a cone angle greater than the cone angle of the multi-view image or view of the multi-view display. Specifically, in some embodiments, the wide-angle illumination may have a cone angle greater than about 20 degrees (e.g., >±20°). In other embodiments, the cone angle of the wide-angle illumination may be greater than about 30 degrees (e.g., >±30°), or greater than about 40 degrees (e.g., >±40°), or greater than about 50 degrees (e.g., >±50°). For example, the cone angle of the wide-angle illumination may be about 60 degrees (e.g., >±60°).
[0017] In the multi-view mode, a multi-view backlight 118 may be used instead of the wide-angle backlight 115. The multi-view backlight 118 may have an array of multi-beam elements on the top, bottom, or between the top and bottom surfaces of the multi-view backlight 118. The multi-beam elements are configured to scatter light into multiple directional light beams having different principal angular directions from each other. Furthermore, the directional light beams of the multiple directional light beams have directions that correspond to the viewing directions of the multi-view display 112, or equivalently, the multi-view image displayed by the multi-view display 112, during the multi-view mode. For example, if the multi-view display 112 operates in the multi-view mode to display a multi-view image having four views, the multi-view backlight 118 may scatter light including four directional light beams, each corresponding to a different one of the four views. The mode controller 121 may continuously switch between the 2D mode and the multi-view mode such that during a first consecutive time interval, the multi-view image is displayed using the multi-view backlight and during a second consecutive time interval, the 2D image is displayed using the wide-angle backlight. The directional light rays may be at a predetermined angle, and each directional light ray corresponds to a different view of the multi-view image.
[0018] In some embodiments, each backlight of the multi-view display 112 (i.e., the wide-angle backlight 115 and the multi-view backlight 118) is configured to direct light as guided rays in a light guide. That is, one or both of the wide-angle backlight 115 and the multi-view backlight 118 may include a light guide configured to direct light as guided rays. As used herein, a "light guide" is defined as a structure that directs light within the structure using total internal reflection or "TIR." Specifically, a light guide may include a core that is substantially transparent at the operating wavelength of the light guide. In various examples, the term "light guide" generally refers to a dielectric optical waveguide that utilizes total internal reflection to guide light at the interface between the light guide's dielectric material and the material or medium surrounding the light guide. By definition, a condition for total internal reflection is that the refractive index of the light guide is greater than the refractive index of the surrounding medium adjacent to the surface of the light guide. In some embodiments, the light guide may include a coating in addition to or instead of the refractive index difference described above to further promote total internal reflection. The coating may be, for example, a reflective coating. The light guide may be any of several types of light guides, including, but not limited to, a plate or flat guide and / or a strip guide. The light guide may be in the shape of a plate or flat plate. The light guide may be edge-lit by a light source (e.g., a light-emitting element).
[0019] As described above, the multi-view backlight 118 of the multi-view display 112 is configured to scatter a portion of the guided light as directional light using multibeam elements of an array of multibeam elements, each multibeam element of the array of multibeam elements including one or more of a diffraction grating, a micro-refractive element, and a micro-reflective element. In some embodiments, the diffraction grating of the multibeam element may include multiple individual sub-gratings, the micro-refractive element may include multiple individual refractive sub-elements, and the micro-reflective element may include multiple individual reflectors or reflective sub-elements. According to various embodiments, the multi-beam element including a diffraction grating is configured to diffract-couple or scatter the light-guided portion as multiple directional light rays. According to various embodiments, the micro-reflective element is configured to reflectively couple or scatter the light-guided portion as multiple directional light rays. The micro-reflective element may include a reflective coating that controls how the guided light is scattered from the light guide. According to various embodiments, the micro-refractive elements are configured to refractionally combine or scatter the light-guiding portion into multiple directional light rays (ie, refract and scatter the light-guiding portion) by or using refraction.
[0020] The multi-view display 112 may also include a light valve array disposed above the backlight (e.g., above the wide-angle backlight 115 and above the multi-view backlight 118). The light valves of the light valve array may be, for example, liquid crystal light valves, electrophoretic light valves, light valves based on or utilizing electrowetting, or any combination thereof. When operating in 2D mode, the wide-angle backlight 115 emits light toward the light valve array. Each light valve is controlled to realize a particular pixel valve that displays a 2D image when illuminated by the wide-angle light emitted by the wide-angle backlight 115. In this regard, each light valve corresponds to a single pixel. In this regard, a single pixel may include different colored pixels (e.g., red, green, blue) that make up a single pixel cell (e.g., an LCD cell).
[0021] When operating in multi-view mode, the multi-view backlight 118 emits directional light rays (e.g., as a bright field) to illuminate the light valve array. The light valves may be grouped together to form a multi-view pixel. For example, in a four-view multi-view format, a multi-view pixel may include four different pixels, each corresponding to a different view. Each pixel within a multi-view pixel may also include pixels of different colors.
[0022] Each light valve in the multiview pixel array may be illuminated by one of a plurality of directional light beams having a principal angular direction. Thus, a multiview pixel is a grouping of pixels that represent different pixels in different views of a multiview image. In some embodiments, each multibeam element of the multiview backlight 118 corresponds to a multiview pixel of the light valve array.
[0023] The multi-view display 112 provides a screen for displaying the multi-view image 103. The screen may be, for example, a display screen of a telephone (e.g., a mobile phone, a smartphone, etc.), a computer monitor of a tablet computer, a laptop computer, a desktop computer, a camera display, or an electronic display of virtually any other device.
[0024] As used herein, the article "a" is intended to have its ordinary meaning in the patent field, i.e., "one or more." For example, "a processor" means one or more processors, and therefore "the processor" herein means "one or more processors." Also, references herein to "top," "bottom," "upper," "lower," "top," "bottom," "front," "rear," "first," "second," "left," or "right" are not intended to be limiting. As used herein, the word "about," when applied to a value, generally means within the error range of the device used to generate the value, unless expressly stated otherwise, or may mean plus or minus 10%, or plus or minus 5%, or plus or minus 1%. Furthermore, as used herein, the word "substantially" means majority, or nearly all, or all, or an amount in the range of about 51% to about 100%. Furthermore, the examples herein are merely illustrative and are presented for purposes of illustration and not limitation.
[0025] As used herein, the term "field of view" (FoV) is defined to mean the range of angles a camera can observe through its lens. FoV may be defined as the angles forming a cone of light that is allowed to enter the camera lens. For example, a camera with a wide-angle lens may have an FoV ranging from 65° to 95°, while a camera with a telephoto lens may have a narrower FoV ranging from 10° to 30°.
[0026] 3A illustrates a multi-view imaging system including multiple cameras shown from a top view, according to one embodiment consistent with principles described herein. In some embodiments, the multi-view imaging system comprises a camera system 201 including multiple cameras 203 arranged along a circular arc 206, the multiple cameras 203 having coplanar orientations, each camera 203 having a common field of view (FoV) 207, and the camera system 201 configured to capture image data of an object 209 located at a common distance "d" from each of the cameras 203. The image data includes a multi-view image, which includes multiple different view images of the object. The camera 203 is a device that converts light received from the object 209 through a lens into an image, which may be a digital representation of the object based on the position, orientation, and configuration of the camera 203.
[0027] The cameras 203 are positioned along a circular arc 206. The arc 206 may be circular or partially circular such that the distance "d" that the cameras 203 are away from the object 209 is approximately the same. That is, the distance d may be the radius of the arc 206. It should be appreciated that some tolerance is allowed so that the distance "d" between the object 209 and each camera 203 is substantially approximately the same so as not to compromise the quality of the final multi-view image.
[0028] Additionally, the cameras 203 may have a coplanar orientation. By arranging the cameras 203 in a coplanar orientation, each camera is arranged along the same plane, with its field of view (FoV) along the same plane. In this manner, diametrically opposed cameras 203 are pointed towards each other. Cameras pointing in different directions are not coplanar. For example, a ring of cameras pointing down or up is not a coplanar orientation.
[0029] Each camera 203 may have a common FoV 207. In this regard, each camera 203 may have the same type of lens or the same lens configuration to allow light of the same angular range to enter the lens of each camera 203. The FoV 207 may be adjusted by zooming in or out. By configuring all cameras 203 to have a common FoV 207, each camera 203 globally changes its FoV 207 with user input that zooms in, zooms out, or otherwise changes the FoV 207. Thus, each camera 203 may dynamically adopt the same FoV 207 in response to user input. The FoV 207 is converted to a focal length, which is expressed as a percentage of the width of the recording surface. The focal length f may be determined according to the following equation (1):
[0030]
number
[0031] The camera system 201 may be configured to capture image data of an object 209 located at a common distance d from each of the cameras 203. The size of the object 209 may range from small items (e.g., shoes, watches) to large items (e.g., cars, buildings). In some embodiments, the cameras 203 may move in unison along an arc 206. In this regard, the cameras 203 may orbit around the object 209 to capture different perspectives of the object 209. Each camera 203 may be configured to place its focal plane 215 on the object such that each camera 203 focuses on the object 209 with the same focal length f.
[0032] In some embodiments, each camera 203 is mounted on a rotatable support member. The rotatable support member is configured to orbit the multiple cameras 203 uniformly along an arc 206 around the object 209. FIG. 3A shows a double-headed curved arrow illustrating how the cameras 203 may rotate in unison around the object 209 while maintaining a constant gap (e.g., camera reference) between each camera 203. Each camera 203 may be rigidly fixed to the rotatable support member, in which case the rotatable support member is shaped like a rigid ring. The rotatable support member may have tracks that allow each camera 203 to move uniformly and smoothly along the axis of rotation created by the rotatable support member. A photographer may provide user input to control the speed and direction of rotation so that the cameras 203 orbit around the object 209.
[0033] 3B illustrates a multi-view image capture system including multiple cameras shown from a side view, according to one embodiment consistent with the principles described herein. FIG. 3B shows how different cameras 203 are positioned along a circular arc 206 and aligned in coplanar orientations to capture multiple view images of an object.
[0034] FIG. 4 illustrates a computing system 221 coupled to a camera system 201 of a multi-view image capture system, according to one embodiment consistent with the principles described herein. The computing system 221 may be a processor-based system that receives input, executes program instructions, and generates output. An example of a computing system 221 is described in more detail in connection with FIG. 11 . FIG. 4 illustrates a computing system 221 coupled to the camera system 201 using a wired or wireless connection. The computing system 221 may send control signals 224 to the camera system 201. The control signals 224 may include instructions to control or configure the cameras of the camera system 201. For example, the control signals 224 may cause the camera system 201 to move its cameras, zoom in, zoom out, adjust the field of view, select a subset of cameras for image capture, deselect some cameras, start recording, stop recording, or other instructions for controlling the cameras and setting visual parameters.
[0035] The computing system 221 may also receive a feedback signal 227 from the camera system 201. The feedback signal 227 may include visual parameters of the camera (e.g., FoV, focal length, distance between the camera and the object), camera position, camera state, or other information related to the camera system 201. In some embodiments, the multi-view capture system may include a sensor configured to measure a common distance between the object and each camera of the set of cameras to generate a second value. The sensor may be a distance sensor attached to one or more cameras. The feedback signal 227 may include a second value indicative of the common distance measured by the sensor.
[0036] The camera system 201 may also capture image data 230 of the object and transmit the image data 230 to the camera system 201. The image data 230 includes different view images of the object and may be formatted as a multi-view image representing a video frame or a still image. The multi-view image includes multiple different view images of the object. The computing system 221 may receive the image data 230 and store it in memory. The computing system 221 may also process the image data 230 and stream it to a multi-view display device as live multi-view video for rendering in real time. For example, the multi-view image may be converted into a frame composed of tiles, with each tile representing a different view image.
[0037] Thus, according to an embodiment, computing system 221 is configured to receive a first value indicative of a common field of view and a second value indicative of a common distance. The first value or the second value may be part of feedback signal 227 or may be received by accessing memory of computing system 221. For example, a photographer may specify an FoV value (or focal length or other equivalent indicative of FoV) for camera system 201 as user input 232. The photographer may configure each camera with a common FoV and provide the common FoV as a first value that is an input to computing system 221. In this regard, camera system 201 may configure its cameras to adopt a user-specified FoV, which may be received by computing system 221 as a first value as part of a calculation to determine target camera reference value 233.
[0038] A second value indicating the common distance may be received as a distance input 243 provided to computing system 221. Distance input 243 may be specified by a user or may be determined automatically by a sensor configured to measure the distance between the object being photographed and the camera of camera system 201.
[0039] According to an embodiment, the computing system 221 is configured to calculate the target camera metric 233 based on a first value indicating a common FoV and a second value indicating a common distance. For example, the distance input 243 and the FoV may be used to determine an optimal target camera metric 233. The target camera metric 233 is a value indicating the distance between two viewpoints (e.g., camera positions) that provide corresponding views. The larger the target camera metric 233, the more parallax there is between different views. When the target camera metric 233 is zero, there is no parallax between the views, and therefore the multi-view image fits into a single view image.
[0040] In some embodiments, the target camera reference value 233 is further calculated to achieve a predetermined disparity level. The predetermined disparity level 246 may be a value that specifies a desired amount of disparity for the target multi-view display. The multi-view display may be similar to the multi-view display 112 of FIGS. 1 and 2. The predetermined disparity level 246 may be expressed as a percentage of the multi-view display width between adjacent views. In this regard, the target camera reference value 233 may be calculated to be optimized for the disparity level of the particular multi-view display. The predetermined disparity level 246 may be received as an input to the computing system 221 or otherwise stored in memory of the computing system 221.
[0041] The target camera reference value 233 may be calculated to be proportional to a predetermined disparity level 246, proportional to the common distance, and inversely proportional to the focal length f, which may be converted from the FoV as described above.
[0042] In some embodiments, computing system 221 is further configured to determine the presence of a textureless background. The presence of a textureless background may be used to further define how target camera metric 233 is calculated. For purposes of this definition, a "textureless background" may be a monochromatic background consisting of similar shades of color, or another background lacking diverse visual features. For example, a green screen or a blue screen may be a textureless background placed behind the object being photographed. The presence of a textureless background may be provided as background data 249 supplied as an input to computing system 221. For example, a user may specify whether there is a textureless background added behind the object. Background data 249 may be a flag or a binary indication of the presence of a textureless background. In some embodiments, background data 249 may be a user-specified input. In other embodiments, the presence of a textureless background may be determined automatically by a sensor. For example, a sensor may measure the degree of color variation in the background, which is compared to a threshold to determine whether it is textured or textureless.
[0043] In some embodiments, in response to the presence of a textureless background, the target camera reference values 233 are further calculated based on the depth of the scene. In this regard, in response to the presence of a textureless background, the target camera reference values are further calculated based on the distance between the subset of cameras and the front of the object and the distance between the subset of cameras and the rear of the object.
[0044] The common distance (distance between the camera and the object) may further be specified to include or otherwise be derived from two different values, Zmin and Zmax. Zmin is the minimum scene depth calculated by taking the distance between the camera and the front of the object, and Zmax is calculated by taking the distance between the camera and the back of the object. Thus, the common distance may be Zmin, Zmax, or an average of both. In response to the presence of a textureless background, the target camera reference value 233 may be calculated according to the following equation (2):
[0045]
number
[0046] In some embodiments, in response to the presence of a textured background, the target camera reference value is further calculated according to equation (3): B=D*d / f (3) where B is the target camera reference value 233, D is a predetermined disparity level, d is the common distance (e.g., the distance between the camera and the object), and f is the focal length expressed as a percentage of the width of the recording surface. Textured backgrounds contain color variations and visual artifacts that are intended to be part of the scene or otherwise the subject of the scene, while untextured backgrounds are intended to draw the viewer's attention away from the background and toward the object.
[0047] Thus, the presence of a textured or textureless background controls how the target camera reference 233 is calculated by factoring in the scene depth calculation, which, as explained above, is based on the change in depth between the front of an object and the back of an object for a subset of cameras.
[0048] Upon calculating the target camera reference value 233, the computing system 221 may be further configured to dynamically adjust an inter-camera focusing distance between the cameras of the subset of cameras to match the target camera reference value 233. The inter-camera focusing distance is the distance between the viewpoints of adjacent cameras. According to some embodiments, the inter-camera focusing distance may be the physical distance between the centers of adjacent camera lenses or an equivalent distance between the respective camera positions. Upon calculating the target camera reference value 233, the computing system 221 may generate control signals to control the camera system 201 to reconfigure the cameras to achieve the target camera reference value 223.
[0049] As described above, the inter-camera focusing distance may be dynamically changed. As the photographer zooms in or out, or changes the object position, or changes the focal length / FoV, or makes other similar changes, updated target camera reference values 233 may be calculated to be optimized for the particular multi-view display. The inter-camera focusing distance may be dynamically changed on the fly as the photographer changes various visual parameters (e.g., zoom, FoV, distance from object, etc.) to provide a parallax level optimized for the multi-view display.
[0050] A subset of cameras may be enabled or otherwise selected to capture multi-view content of an object. For example, if a photographer wants to capture multi-view content having four separate views, four different cameras may be selected to form the subset of cameras. The computing system 221 may select a number of cameras corresponding to the number of views of the multi-view image format. The multi-view image format may define a specific number of views and may be specified by the multi-view display. For multi-view displays capable of supporting a larger number of views, more cameras may be selected to capture the multi-view content.
[0051] FIG. 5 illustrates an example of dynamically adjusting the inter-camera focusing distance by individually selecting a subset of the cameras 203, according to one embodiment in accordance with the principles described herein. For example, the computing system 221 may be further configured to dynamically adjust the inter-camera focusing distance by individually selecting a subset of the cameras 203 from the plurality of cameras 203. FIG. 5 illustrates an example in which the captured multi-view content has four views. Thus, four cameras 203 are selected by the computing system. The selected cameras 203 are shown with solid lines, and the unselected cameras 203 are shown with dotted lines. Each camera 203 may be selected by the computing system 221 using a camera identifier. The camera identifier allows the computing system 221 to individually select a specific camera to capture a view image of an object. The four cameras 203 are selected so that the inter-camera focusing distance can be controlled to match a target camera reference value. In this embodiment, each camera 203 may have a fixed position relative to the other cameras. As a result, the inter-camera focusing distance is quantized and cameras 203 are selected to match the target camera reference as closely as possible. The cameras 203 may have fixed positions close to each other to accommodate small adjustments in the inter-camera focusing distance.
[0052] For example, cameras 203 that are closer to each other may be selected to decrease the calculated target camera reference value, while cameras 203 that are farther apart may be selected to increase the calculated target camera reference value. In this example of Figure 5, the inter-camera shooting distance may correspond to the number of cameras between the selected cameras.
[0053] FIG. 6 illustrates an example of dynamically adjusting the inter-camera focusing distance by individually moving each camera in the subset of cameras along a track, according to one embodiment consistent with the principles described herein. For example, FIG. 6 depicts an embodiment in which the computing system is further configured to dynamically adjust the inter-camera focusing distance by individually moving each camera in the subset of cameras along a track. The cameras 203 may be individually selected and moved along an arcing track. The cameras may move relative to one another. In this regard, each camera 203 may individually move toward or away from an adjacent camera 203. Additionally, each camera 203 may move uniformly to maintain approximately the same baseline across the subset of cameras. This allows a photographer to apply an inter-camera focusing distance across the subset of cameras 203a-d and orbit the cameras 203a-d around an object while maintaining the inter-camera focusing distance.
[0054] 6 illustrates an embodiment in which one camera in the subset of cameras is designated as a center-view camera. For example, the third camera 203c may be considered the center-view camera. The cameras to the left and right of center-view camera 203c may be moved closer or farther apart to achieve an inter-camera focusing distance that matches the target camera reference.
[0055] 7 illustrates an example of modifying the field of view by applying a cropping window to each view image, according to one embodiment consistent with principles described herein. For example, FIG. 7 depicts an embodiment in which the computing system is further configured to modify the common field of view by applying a cropping window to view images provided by a subset of the cameras.
[0056] 7 shows four cameras 203a-d selected to capture multi-view content of an object 209. The first camera 203a captures view image A, the second camera 203b captures view image B, the third camera 203c captures view image C, and the fourth camera 203d captures view image D. Each view image captured by the cameras 203a-d has a different perspective of the object 209 based on the position of the cameras 203a-d.
[0057] A user may change the field of view by providing user input that is received by the computing system. For example, the user may zoom in or zoom out. In response to the user input, the computing system may apply cropping windows 255a-d to the corresponding view images (e.g., view images A-D). A smaller cropping window zooms in and effectively narrows the field of view, whereas a larger cropping window zooms out and effectively widens the field of view. In this regard, a single user input that changes the field of view causes the computing system to apply cropping windows 255a-d to each view image of the multi-view image. In other words, the computing system applies cropping windows 255a-d globally for each selected camera to change the field of view in accordance with the user input.
[0058] FIG. 8 illustrates an example of data conversion of a multi-view image into a multi-view video stream rendered by a multi-view display device, according to one embodiment consistent with principles described herein. In some embodiments, the multi-view video stream may be rendered in real time. For example, FIG. 8 illustrates a case in which a multi-view image is converted into frames of a multi-view video stream rendered in real time by a multi-view display device, where the frames are configured to be streamed to the multi-view display device as multiple deinterlaced compressed-view images. Image data 230 constituting the multi-view image is generated by camera system 201 and transmitted to computing system 221. Image data 230 may include view images obtained from multiple cameras of a subset of the multiple cameras. The inter-camera focusing distance may be dynamically changed in response to user input (e.g., zooming in, zooming out, changing FoV, etc.). In response, additional image data 230 including additional multi-view images is captured using the dynamically updated inter-camera focusing distance.
[0059] The computing system 221 may generate the multi-view video stream 258 by formatting the multi-view images of the image data 230 into multiple tiled frames. FIG. 8 illustrates tiled frames ranging from frame A to frame n. Tiled frames are multi-view frames, and each multi-view image of a particular tiled frame is arranged as a tile within that tiled frame. As a result, each multi-view image (e.g., tile) within a tiled frame is spatially separated from other multi-view images (e.g., tiles). The example of FIG. 8 shows a tiled frame with four views, where view 1 is in the upper left quadrant, view 2 is in the upper right quadrant, view 3 is in the lower left quadrant, and view 4 is in the lower right quadrant.
[0060] The tiled multi-view frames may be said to be de-interlaced. Interlacing is the process of spatially multiplexing the pixels of each view image to form an interlaced multi-view image 259. The interlaced multi-view image 259 may represent a frame in the multi-view video stream 258, and the tiled multi-view frames are then interlaced and rendered on a target multi-view display device 260. The multi-view display device 260 may include a multi-view display, such as the multi-view display 112 of FIGS. 1 and 2. The multi-view display device 260 may be a computing system, such as the computing system described in FIG. 11.
[0061] As shown in FIG. 8, interlaced multiview image 259 has spatially multiplexed or otherwise interlaced views. Interlaced multiview image 259 includes an array of pixels, with each pixel corresponding to one of four views such that the pixels are interlaced (e.g., interleaved or spatially multiplexed). Pixels belonging to view 1 are represented by number 1, pixels belonging to view 2 are represented by number 2, pixels belonging to view 3 are represented by number 3, and pixels belonging to view 4 are represented by number 4. The views of interlaced multiview image 259 are interlaced pixel-by-pixel horizontally along each row. Interlaced multiview image 259 may be a pixel array having rows of pixels represented by uppercase letters A through E and columns of pixels represented by lowercase letters a through h. FIG. 8 shows the location of one multiview pixel 282 at row E, columns e through h. Multiview pixel 282 is an arrangement of pixels taken from each of the four views. In other words, multi-view pixel 282 is the result of interlacing the individual pixels of each of the four views in a spatially multiplexed manner. Although Figure 8 shows the pixels of different views interlaced horizontally, the pixels of different views may also be interlaced vertically, as well as both horizontally and vertically.
[0062] Assuming the multiview image format specifies four views, the interlaced multiview image 259 may be a multiview pixel 282 having one pixel from each of the four views. In some embodiments, the multiview pixels may be staggered in a particular direction, as shown in FIG. 8, where the multiview pixels are horizontally aligned while being staggered vertically. In other embodiments, the multiview pixels may be horizontally staggered and vertically aligned. The particular manner in which the multiview pixels are interlaced and staggered may depend on the design of the multiview display. The interlaced multiview image 259 may have its pixels interlaced and arranged as multiview pixels such that they can be mapped to physical pixels (e.g., a light valve array) of the multiview display. In other words, pixel coordinates of the interlaced multiview image 259 correspond to physical locations on the multiview display. The multiview pixel 282 has a mapping to a particular set of light valves in the light valve array. The light valve array is controlled to modulate light according to the interlaced multiview image 259. The light valves of the light valve array may be, for example, liquid crystal light valves, electrophoretic light valves, electrowetting-based or -utilizing light valves, or any combination thereof.
[0063] When generating the multi-view video stream 258, the computing system 221 may convert the multi-view images into multiple deinterlaced compressed-view images. The multi-view video stream 258 may include tiled multi-view frames that have been compressed by applying a compression process. Compression refers to the process of reducing the size (in bits) of video data while maintaining a minimum amount of video quality. Without compression, the time required to fully stream the video would increase or otherwise consume network bandwidth. Therefore, video compression may enable reduced video stream data, aiding in real-time video streaming, faster video streaming, or reduced buffering of the received video stream. The compression may be lossy, meaning that compressing and decompressing the input data causes some degradation in quality. The compressed video may be generated using a video encoder (e.g., compressor) (e.g., coder-decoder (CODEC)) that conforms to a compression specification, such as H.264 or other CODEC specifications. Compression may involve converting a series of frames into I-frames, P-frames, and B-frames as defined by the CODEC.
[0064] Multi-view display device 260 may be configured to receive multi-view video stream 258 (e.g., including deinterlaced compressed views). Multi-view display device 260 may then decompress multi-view video stream 258 using, for example, a CODEC. Multi-view display device 260 may interlace each view image of a frame to generate interlaced multi-view image 259. More efficient compression may be achieved by compressing tiled frames rather than compressing the interlaced multi-view image. As a result, rather than computing system 221 performing the interlacing processing, multi-view display device 260 may perform the interlacing processing.
[0065] 9 is a flowchart providing an example of multi-view image capture according to one embodiment in accordance with the principles described herein. The flowchart of FIG. 9 provides an example of operating a multi-view image capture system, where the multi-view image capture system includes a processor that executes a set of instructions to perform some of the operations.
[0066] In some embodiments, multi-view image capture includes arranging 304 cameras along a circular arc. For example, multi-view image capture includes arranging multiple cameras along a circular arc, where the multiple cameras have coplanar orientations and each camera has a common field of view. In some embodiments, the cameras may be attached to a rigid circular support so that they point toward the object of interest. The circular or arc-like orientation allows each camera to have a common distance to the object. An example of this camera arrangement is shown in Figures 3A and 3B.
[0067] In some embodiments, multi-view image capture involves rotating multiple cameras to orbit the object uniformly along a circular arc. A photographer may rotate a subset of the cameras along an arc around the object to provide new views of the object while maintaining a constant inter-camera focusing distance. Each camera may be mounted on a circular support that slides along a track to allow the multiple cameras to orbit the object uniformly.
[0068] In some embodiments, multi-view image capture includes capturing 307 multi-view images of an object. For example, this may involve capturing multi-view images of an object located at a common distance from a subset of cameras. Each camera may digitally record the object from its viewpoint, thereby creating a view image at a particular point in time. The view images across multiple cameras at a particular point in time result in a multi-view image of the object. The captured multi-view images may be transmitted to a computer system for processing and then streamed or stored.
[0069] In some embodiments, multi-view image capture includes receiving 310 a first value and a second value. For example, this may involve receiving a first value indicative of a common field of view and a second value indicative of a common distance. A computing system may receive the first and second values as input, and the first and second values may be specified by a user or may be received by accessing a memory that stores the first and second values.
[0070] In some embodiments, the multi-view image capture includes calculating 313 a target camera reference value. For example, this may involve calculating the target camera reference value based on a first value and a second value. A computing system may receive the first and second values as input and generate the target camera reference value using one or more of equations (1), (2), and (3) discussed above.
[0071] In some embodiments, the multi-view image capture further includes calculating a target camera metric to achieve a predetermined disparity level. For example, as an additional input, the computing system may receive a predetermined disparity level representing a predetermined optimal disparity level for a particular multi-view display. Accordingly, different multi-view displays may be designed to accommodate different disparity levels, and the calculation of the target camera metric results in a target camera metric that is optimized for the particular multi-view display. In some embodiments, the multi-view image capture system further includes calculating a target camera metric based on a depth of the scene in the presence of a textureless background. The presence of a textureless background creates a minimum and maximum depth associated with the multi-view image. That is, when using a textureless background, meaningful portions of the scene are limited to objects, and the depth of the scene may be determined accordingly and then used to calculate the target camera metric.
[0072] In some embodiments, the multi-view image capture includes dynamically adjusting 316 the inter-camera focusing distance. For example, this may involve dynamically adjusting the inter-camera focusing distance between the cameras of a subset of the cameras to match a target camera reference value. A computing system that calculates the target camera reference value may then control the camera systems to implement the target camera reference value by sending control signals to cause the camera systems to capture multi-view content according to the inter-camera focusing distance that matches the target camera reference value. Dynamically adjusting the inter-camera focusing distance may include one or more of selecting a subset of cameras from the plurality of cameras and moving each camera in the subset of cameras along a track.
[0073] The captured multi-view images may be stored for processing and rendering at a later time. Alternatively, multi-view image capture may include streaming the multi-view images to a multi-view display device in real time as multiple deinterlaced compressed-view images. An example of this is described above in connection with FIG. 8. A computing system may convert the multi-view images into tiled multi-view frames that are compressed and streamed to the multi-view display device in real time. The multi-view display device may decompress and interlace the tiled frames to render the streamed video content in real time.
[0074] The flowchart of FIG. 9 discussed above may illustrate a system or method for multi-view image capture. When implemented as software, and where applicable, boxes may represent modules, segments, or portions of code containing instructions that implement specified logical functions. The instructions may be embodied in the form of source code, including human-readable statements written in a programming language, object code compiled from source code, or machine code, including numerical instructions recognizable by a suitable execution system, such as a processor of a computing device. The machine code may be translated from source code, etc. When implemented as hardware, each block may represent a circuit or several interconnected circuits that implement the specified logical functions.
[0075] 9 shows a particular order, it is understood that the order may differ from that shown. For example, the order of two or more frames may be swapped relative to the order shown. Also, two or more frames shown may be performed concurrently or with partial concurrence. Furthermore, in some embodiments, one or more of the frames may be omitted or omitted.
[0076] As discussed above, embodiments include a multi-view image capture system including a camera system and a computing system. The camera system includes at least one camera and is configured to capture images of an object at multiple different locations along a circular arc, each image corresponding to a common field of view that includes the object. The image capture system also includes a computing system coupled to the camera system. The computing system and the camera system cooperate to dynamically adjust an inter-camera shooting distance to accommodate a required parallax level of the multi-view display, thereby generating multi-view content optimized for the target multi-view display. Additionally, the computing system may be configured to receive a first value indicative of the common field of view and a second value indicative of a distance between the camera system and the circular arc. The computing system may be configured to calculate a target camera reference value based on the first value indicative of the common field of view and the second value indicative of the common distance, the target camera reference value corresponding to the distance between the multiple different locations along the circular arc. Additionally, the computing system may be configured to dynamically adjust the inter-camera shooting distance of the camera system to match the target camera reference value. As discussed above, in some embodiments, the camera system includes multiple cameras aligned along a circular arc, the multiple cameras having coplanar orientations. For example, multiple cameras may be mounted on a circular rigid support member.
[0077] In other embodiments, the camera system comprises an unmanned aerial vehicle configured to travel along an arc, the unmanned aerial vehicle comprising a sensor configured to measure distance and generate a second value. For example, FIG. 10 illustrates an example of a multi-view image capture system 402 according to one embodiment in accordance with principles described herein, where the camera system 405 comprises an unmanned aerial vehicle. The unmanned aerial vehicle may be a drone or other autonomous vehicle remotely controlled or programmatically controlled to move along the arc 408 and capture one or more multi-view images of an object 411. At least one camera may be mounted on the unmanned aerial vehicle. As shown in FIG. 10, the camera system 405 is a single camera that moves counterclockwise along an arc (e.g., at least a portion of a circular path). The camera system 405 may have a fixed or adjustable field of view (FoV) 412 that captures an object 411 that falls within a focal plane 415 of the camera system 405.
[0078] The camera system 405 may have a communications link with a computing system 421 that receives the data multi-view images and may also control the camera system 405. The computing system 421 may receive various inputs, such as, for example, the FoV 412 (or focal length), the distance between the camera system 405 and the object 411, a predetermined parallax level of the multi-view display, or other visual parameters related to the object's position, view, zoom level, etc. In response, the computing system 421 may calculate a target camera reference value based on a first value indicating a common field of view and a second value indicating a common distance, and dynamically adjust the inter-camera distance to match the target camera reference value.
[0079] In some embodiments, the computing system 421 is configured to dynamically adjust the inter-view distance by determining a timestamp interval for each captured image. FIG. 10 shows timestamps (T1-T4) of different view images captured by the camera system 405. The timestamp interval may be expressed as time or distance as the camera system 405 moves along an arc. The larger the timestamp interval, the larger the inter-view distance. For example, four view images are captured at corresponding timestamps (T1-T4). The four view images constitute a multi-view image with a camera reference value determined by the time interval between successive timestamps.
[0080] FIG. 11 is a schematic block diagram depicting an exemplary illustration of a computing system according to one embodiment in accordance with the principles described herein. The computing system 1000 may include a system of components that perform various computing operations, such as receiving multi-view images, receiving input values, calculating camera system parameters, generating control signals for controlling the camera system, processing the multi-view images, and streaming them to a multi-view display device. The computing system 1000 may be a laptop, tablet, smartphone, touchscreen system, intelligent display system, or other client device. The computing system 1000 may include various components, such as a processor 1003, memory 1006, input / output (I / O) components 1009, a display 1012, and possibly other components. These components may connect to a bus 1015, which serves as a local interface allowing the components of the computing system 1000 to communicate with each other. While the components of the computing system 1000 are shown as being contained within the computing system 1000, it should be appreciated that at least some of the components may be connected to the computing system 1000 through external connections. For example, a component may plug into or otherwise connect externally to computing system 1000 via an external port, socket, plug, wireless link, or connector.
[0081] Computing system 1000 may implement computing system 221 of Figures 4 and 8 and computing system 421 of Figure 10. Additionally, in embodiments in which computing system 1000 includes a multi-view display, computing system 1000 may implement multi-view display device 260 of Figure 8.
[0082] The processor 1003 may be a central processing unit (CPU), a graphics processing unit (GPU), any other integrated circuit that performs computing operations, or any combination thereof. The processor 1003 may include one or more processing cores. The processor 1003 comprises circuitry that executes instructions. Instructions include, for example, computer code, programs, logic, or other machine-readable instructions that are received and executed by the processor 1003 to perform the computing function embodied in the instructions. The processor 1003 may execute instructions to operate on data. For example, the processor 1003 may receive input data (e.g., multi-view images, video streams, multi-view content, user input, etc.), process the input data according to an instruction set, and generate output data (e.g., process multi-view content for streaming, control signals, etc.). As another example, the processor 1003 may receive instructions and generate new instructions for later execution. The processor 1003 may comprise hardware that implements a graphics pipeline for processing and rendering multi-view content. For example, the processor 1003 may include one or more GPU cores, vector processors, scalar processes, or hardware accelerators.
[0083] The memory 1006 may include one or more memory components. The memory 1006 is defined herein to include either or both volatile and nonvolatile memory. A volatile memory component is a memory component that does not retain information when power is lost. Volatile memory may include, for example, random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), or other volatile memory structures. System memory (e.g., main memory, cache, etc.) may be implemented using volatile memory. System memory refers to high-speed memory that can temporarily store data and instructions for rapid read and write access to support the processor 1003.
[0084] Nonvolatile memory components are memory components that retain information when power is lost. Nonvolatile memory includes read-only memory (ROM), hard disk drives, solid-state drives, USB flash drives, memory cards accessed via memory card readers, floppy disks accessed via associated floppy disk drives, optical disks accessed via optical disk drives, and magnetic tapes accessed via suitable tape drives. ROM may include, for example, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other similar memory devices. Storage memory may be implemented using nonvolatile memory to provide long-term retention of data and instructions.
[0085] Memory 1006 may refer to a combination of volatile and non-volatile memory used to store instructions and data. For example, data and instructions may be stored in non-volatile memory and loaded into volatile memory for processing by processor 1003. Execution of instructions may include, for example, a compiled program that is translated from non-volatile memory into machine code in a format that can be loaded into volatile memory and then executed by processor 1003; source code that is converted into a suitable format such as object code that can be loaded into volatile memory for execution by processor 1003; or source code that is interpreted by another executable program to generate instructions in volatile memory and then executed by processor 1003. Instructions may be stored in or loaded into any portion or component of memory 1006, including, for example, RAM, ROM, system memory, storage, or any combination thereof.
[0086] Although memory 1006 is illustrated as separate from other components of computing system 1000, it should be appreciated that memory 1006 may be embedded in or otherwise at least partially integrated with one or more components. For example, processor 1003 may include internal memory registers or cache for performing processing operations. Device firmware or drivers may include instructions stored in dedicated memory devices.
[0087] I / O components 1009 include, for example, a touch screen, speakers, microphones, buttons, switches, dials, cameras, sensors, accelerometers, or other components that receive user input or generate output directed to a user. I / O components 1009 may receive user input and convert it into data for storage in memory 1006 or for processing by processor 1003. I / O components 1009 may receive data output by memory 1006 or processor 1003 and convert it into a form that is perceivable by the user (e.g., sound, tactile response, visual information, etc.).
[0088] A particular type of I / O component 1009 is a display 1012. The display 1012 may include a multi-view display (e.g., multi-view display 112), a multi-view display combined with a 2D display, or any other display that presents images. A capacitive touchscreen layer, acting as the I / O component 1009, may be layered within the display to allow a user to provide input while simultaneously perceiving visual output. The processor 1003 may generate data that is formatted as an image for presentation on the display 1012. The processor 1003 may execute instructions that cause the image to be rendered on the display for perception by the user.
[0089] The bus 1015 facilitates the communication of instructions and data between the processor 1003, the memory 1006, the I / O components 1009, the display 1012, and any other components of the computing system 1000. The bus 1015 may include address translators, address decoders, fabric, conductive wiring, conductive lines, ports, plugs, sockets, and other connectors that enable the communication of data and instructions.
[0090] The instructions in memory 1006 may be embodied in various forms to implement at least a portion of a software stack. For example, the instructions may be embodied as part of an operating system 1031, an application 1034, a device driver, firmware, other software components, or any combination thereof. The operating system 1031 is a software platform that supports basic functions of the computing system 1000, such as scheduling tasks, controlling I / O components 1009, providing access to hardware resources, power management, and supporting applications 1034.
[0091] Applications 1034 may execute on the operating system 1031 and may gain access to the hardware resources of the computing system 1000 through the operating system 1031. In this regard, the execution of applications 1034 is controlled, at least in part, by the operating system 1031. Applications 1034 are user-level software programs that provide high-level functionality, services, and other features to a user. In some embodiments, applications 1034 may be dedicated “apps” that are downloadable or otherwise accessible by a user to the computing system 1000. A user may launch applications 1034 through a user interface provided by the operating system 1031. Applications 1034 are developed and defined by developers in various source code formats. Applications 1034 may be developed using multiple programming or scripting languages, such as, for example, C, C++, C#, Objective C, Java, Swift, JavaScript, Perl, PHP, Visual Basic, Python, Ruby, Go, or other programming languages. The application 1034 may be compiled into object code by a compiler or interpreted by an interpreter for execution by the processor 1003. The application 1034 may implement at least some of the functionality described above. For example, the application 1034 may allow a user to control a camera system by providing various visual parameters.
[0092] The operating system 1031, applications 1034, drivers, firmware, and possibly other instruction sets may each comprise instructions executable by the processor 1003 or other processing circuitry of the computing system 1000 to perform at least some of the functions and operations described above. The instructions described herein may be embodied as software or code executed by the processor 1003 as described above, although the instructions may alternatively be embodied as dedicated hardware or a combination of software and dedicated hardware. For example, the functions and operations performed by the instructions described above may be implemented as circuits or state machines utilizing any one or a combination of several technologies. These technologies may include, but are not limited to, discrete logic circuits having logic gates for implementing various logical functions upon application of one or more data signals, application-specific integrated circuits (ASICs) having appropriate logic gates, field-programmable gate arrays (FPGAs), other components, and the like.
[0093] In some embodiments, instructions for implementing the functions and operations described above may be embodied in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may or may not be part of computing system 1000. The instructions may include, for example, statements, code, or declarations that can be retrieved from a computer-readable medium and processed by a processing circuit (e.g., processor 1003). As defined herein, a "non-transitory computer-readable storage medium" is defined as any medium that can contain, store, or maintain the instructions described herein for use by or in connection with an instruction execution system, such as computing system 1000, and further excludes transitory media such as, for example, a carrier wave.
[0094] The non-transitory computer-readable medium may comprise any one of many physical media, such as, for example, magnetic, optical, or semiconductor media. More specific examples of suitable non-transitory computer-readable media may include, but are not limited to, magnetic tape, magnetic floppy diskettes, magnetic hard drives, memory cards, solid-state drives, USB flash drives, or optical disks. The non-transitory computer-readable medium may also be, for example, random access memory (RAM), including static random access memory (SRAM) and dynamic random access memory (DRAM), or magnetic random access memory (MRAM). In addition, the non-transitory computer-readable medium may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other types of memory devices.
[0095] Computing system 1000 may perform any of the operations or implement functions described above. For example, the process flows discussed above may be performed by computing system 1000 executing instructions and processing data. Although computing system 1000 is illustrated as a single device, embodiments are not so limited. In some embodiments, computing system 1000 may offload processing of instructions in a distributed manner, such that multiple computing systems 1000 or other computing devices operate together to execute instructions that may be stored or loaded in a distributed arrangement. For example, at least some instructions or data may be stored, loaded, or executed within a cloud-based system operating in conjunction with computing system 1000.
[0096] Thus, examples and embodiments of a multi-view capture system and method for capturing multi-view images have been described. The camera system may capture various views of an object while dynamically changing its capture distance to correspond to a predetermined disparity level, simultaneously as a photographer changes various visual parameters. In this regard, the multi-view content generated by the camera system is optimized to achieve a particular disparity level for the target multi-view display. The photographer can freely change parameters such as field of view, focal length, zoom, etc., while ensuring that the rendering of the captured multi-view content achieves the predetermined disparity level. As will be apparent, those skilled in the art will readily understand the principles and advantages of the present invention as described in the accompanying patents. Numerous other arrangements can be readily devised without departing from the scope defined by the appended claims. It should be noted that the present specification discloses the following aspects. [Aspect 1] 1. A multi-view imaging system, comprising: a camera system comprising a plurality of cameras arranged along a circular arc, the cameras having coplanar orientations, each camera of the plurality of cameras having a common field of view, the camera system configured to capture a multi-view image of an object located a common distance from each of the plurality of cameras, the multi-view image including a plurality of different view images of the object; a computing system coupled to the camera system, calculating a target camera reference value based on a first value indicative of the common field of view and a second value indicative of the common distance; Dynamically adjusting inter-camera focusing distances between the cameras of the subset of cameras to match the target camera reference. and a computing system configured to: [Aspect 2] a sensor configured to measure the common distance between the object and each of the cameras of the subset of cameras to generate the second value; 2. The multi-view imaging system of embodiment 1, further comprising: [Aspect 3] A multi-view image capturing system as described in aspect 1, wherein each camera is mounted on a rotatable support member, and the rotatable support member is configured to uniformly rotate the multiple cameras around the object along the arc. [Aspect 4] 2. The multi-view image capture system of claim 1, wherein the target camera reference value is further calculated to achieve a predetermined disparity level. [Aspect 5] The multi-view image capture system of embodiment 4, wherein the computing system is further configured to determine the presence of a texture-free background. [Aspect 6] A multi-view image capturing system as described in aspect 5, wherein, in response to the presence of a textureless background, the target camera reference value is further calculated based on the distance between the subset of cameras and the front of the object and the distance between the subset of cameras and the rear of the object. [Aspect 7] A multi-view image capture system as described in aspect 1, wherein the computing system is further configured to dynamically adjust the inter-camera shooting distance by individually selecting cameras of the subset of cameras from the plurality of cameras. [Aspect 8] 2. The multi-view image capture system of claim 1, wherein the computing system is further configured to modify the common field of view by applying a cropping window to view images provided by the subset of the cameras. [Aspect 9] 2. The multi-view image capture system of claim 1, wherein the computing system is further configured to dynamically adjust the inter-camera shooting distance by individually moving each camera in the subset of cameras along a track. [Aspect 10] A multi-view image capturing system as described in aspect 1, configured to convert the multi-view images into frames of a multi-view video stream that are rendered in real time by a multi-view display device, and to stream the frames to the multi-view display device as multiple deinterlaced compressed view images. [Aspect 11] 1. A method for multi-view image capture, comprising: Arranging a plurality of cameras along an arc, the plurality of cameras being coplanar and each camera of the plurality of cameras having a common field of view; capturing a multi-view image of an object using a subset of cameras of the plurality of cameras, the multi-view image comprising a plurality of different view images of the object, the object being located at a common distance from the subset of cameras; receiving a first value indicative of the common field of view and a second value indicative of the common distance; calculating a target camera reference value based on the first value and the second value; and dynamically adjusting an inter-camera focusing distance between the cameras of the subset of cameras to match the target camera reference. [Aspect 12] 12. The method of multi-view image capture of claim 11, further comprising rotating the plurality of cameras to orbit the plurality of cameras uniformly along the arc around the object. [Aspect 13] 12. The method of multi-view image capture of claim 11, further comprising calculating the target camera reference value to achieve a predetermined disparity level. [Aspect 14] A method of multi-view image capture according to aspect 11, wherein the target camera reference values are calculated based on the distance between the subset of cameras and the front of the object and the distance between the subset of cameras and the rear of the object. [Aspect 15] selecting the subset of cameras from the plurality of cameras; and moving each camera in said subset of cameras individually along a track; 12. The method of multi-view image capture of claim 11, further comprising dynamically adjusting the inter-camera shooting distance by performing one or more of the following: [Aspect 16] 12. The method of multi-view image capture of claim 11, further comprising streaming the multi-view image as a plurality of de-interlaced compressed-view images to a multi-view display device in real time. [Aspect 17] 1. A multi-view imaging system, comprising: a camera system comprising at least one camera configured to capture images of an object at a plurality of different locations along an arc, each image corresponding to a common field of view that includes the object; a computing system coupled to the camera system, receiving a first value indicative of the common field of view and a second value indicative of a distance between the arc and the object; calculating a target camera reference value based on the first value and the second value, the target camera reference value corresponding to a distance between the plurality of different locations along the arc; dynamically adjusting the inter-camera distance of the camera system to match the target camera reference value; and a computing system configured to: [Aspect 18] A multi-view image capture system as described in aspect 17, wherein the camera system comprises a plurality of cameras arranged along the arc, the cameras having coplanar orientations. [Aspect 19] A multi-view image capturing system as described in aspect 17, wherein the camera system comprises an unmanned aerial vehicle configured to travel along the arc, the unmanned aerial vehicle comprising a sensor configured to measure the distance and generate the second value, and the at least one camera is mounted on the unmanned aerial vehicle. [Aspect 20] 20. The multi-view image capture system of claim 19, wherein the computing system is configured to dynamically adjust the inter-image distance by determining a timestamp interval for each captured image. [Explanation of symbols]
[0097] 103 Multi-view images 106 ViewsImage 109 Principal angular direction 112 Multi-view display 115 Wide-angle backlight 118 Multi-view backlight 121 Mode Controller 124 Mode selection signal 201 Camera System 203, 203a~d Camera 206 Arc 207 Field of view (FoV) 209 Object 215 focal plane 221 Computing Systems 224 Control Signal 227 Feedback Signal 230 Image Data 232 User Input 233 Target camera reference value 243 Distance input 246 Parallax Levels 249 Background Data 255a~d Cropping window 258 multi-view video streams 259 interlaced multiview images 260 Multi-view Display Device 282 multi-view pixels 402 Multi-view Imaging System 405 Camera System 408 Arc 411 Object 415 focal plane 421 Computing Systems 1000 Computing Systems 1003 processor 1006 memory 1009 I / O components 1012 display 1015 Bus 1031 Operating Systems 1034 Applications
Claims
1. 1. A multi-view imaging system, comprising: a camera system comprising a plurality of cameras arranged along a circular arc, the cameras having coplanar orientations, each camera of the plurality of cameras having a common field of view, the camera system configured to capture a multi-view image of an object located a common distance from each of the plurality of cameras, the multi-view image including a plurality of different view images of the object; a computing system coupled to the camera system, calculating a target camera reference value based on a first value indicative of the common field of view and a second value indicative of the common distance; Dynamically adjusting inter-camera focusing distances between the cameras of the subset of cameras to match the target camera reference. and a computing system configured to:
2. a sensor configured to measure the common distance between the object and each of the cameras of the subset of cameras to generate the second value; The multi-view imaging system of claim 1 further comprising:
3. 2. The multi-view imaging system of claim 1, wherein each camera is mounted on a rotatable support configured to uniformly orbit the plurality of cameras along the arc around the object.
4. The multi-view image capture system of claim 1 , wherein the target camera reference values are further calculated to achieve a predetermined disparity level.
5. The multi-view image capture system of claim 4 , wherein the computing system is further configured to determine the presence of a texture-free background.
6. 6. The multi-view image capture system of claim 5, wherein, in response to the presence of a textureless background, the target camera reference values are further calculated based on a distance between the subset of cameras and a front of the object and a distance between the subset of cameras and a rear of the object.
7. The multi-view image capture system of claim 1 , wherein the computing system is further configured to dynamically adjust the inter-camera focusing distance by individually selecting cameras of the subset of cameras from the plurality of cameras.
8. The multi-view image capture system of claim 1 , wherein the computing system is further configured to modify the common field of view by applying a cropping window to view images provided by the subset of cameras.
9. 2. The multi-view image capture system of claim 1, wherein the computing system is further configured to dynamically adjust the inter-camera focusing distance by individually moving each camera in the subset of cameras along a track.
10. 2. The multi-view image capture system of claim 1, wherein the multi-view images are converted into frames of a multi-view video stream that are rendered in real time by a multi-view display device, and the frames are configured to be streamed to the multi-view display device as multiple de-interlaced compressed-view images.
11. 1. A method for multi-view image capture, comprising: Arranging a plurality of cameras along an arc, the plurality of cameras being coplanar and each camera of the plurality of cameras having a common field of view; capturing a multi-view image of an object using a subset of cameras of the plurality of cameras, the multi-view image comprising a plurality of different view images of the object, the object being located at a common distance from the subset of cameras; receiving a first value indicative of the common field of view and a second value indicative of the common distance; calculating a target camera reference value based on the first value and the second value; and dynamically adjusting an inter-camera focusing distance between the cameras of the subset of cameras to match the target camera reference.
12. The method of claim 11 , further comprising rotating the multiple cameras to orbit the multiple cameras uniformly along the arc around the object.
13. The method of claim 11 , further comprising calculating the target camera reference values to achieve a predetermined disparity level.
14. 12. The method of claim 11, wherein the target camera reference values are calculated based on distances between the subset of cameras and the front of the object and distances between the subset of cameras and the rear of the object.
15. selecting the subset of cameras from the plurality of cameras; and moving each camera in said subset of cameras individually along a track; 12. The method of claim 11, further comprising dynamically adjusting the inter-camera shooting distance by performing one or more of:
16. The method of claim 11 , further comprising streaming the multi-view image as a plurality of de-interlaced compressed-view images to a multi-view display device in real time.
17. 1. A multi-view imaging system, comprising: a camera system comprising at least one camera configured to capture respective images of an object when the at least one camera is positioned at a plurality of different camera positions along a circular arc, the arc being centered on the object, and each image corresponding to a common field of view that includes the object; a computing system coupled to the camera system, receiving a first value indicative of the common field of view and a second value indicative of a distance between the arc and the object; calculating a target camera reference value based on the first value and the second value, the target camera reference value corresponding to a distance between the plurality of different camera positions along the arc; dynamically adjusting the inter-camera distance of the camera system to match the target camera reference value; and a computing system configured to:
18. 18. The multi-view imaging system of claim 17, wherein the camera system comprises a plurality of cameras aligned along the arc, the cameras having coplanar orientations.
19. 20. The multi-view image capture system of claim 17, wherein the camera system comprises an unmanned aerial vehicle configured to travel along the arc, the unmanned aerial vehicle comprising a sensor configured to measure the distance and generate the second value, and the at least one camera mounted on the unmanned aerial vehicle.
20. 20. The multi-view image capture system of claim 19, wherein the computing system is configured to dynamically adjust the inter-image distance by determining a timestamp interval for each captured image.
Citation Information
Patent Citations
Image processing apparatus, camera calibration processing apparatus and method, and computer program
JP2004088247A
Imaging apparatus
JP2008096721A
Image creation device
JP2008217243A
System for controlling a plurality of cameras
JP2010154052A
Image processor and method
JP2011227613A