Real-time three-dimensional imaging device and method
The HD3D imaging system addresses the challenge of combining real and computer-generated elements by using synchronized color and depth cameras to generate high-resolution, low-latency three-dimensional representations, enhancing applications in entertainment and other fields.
Patent Information
- Application Number
- JP2025531112
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-30
- Filing Date
- 2023-11-29
- Publication Date
- 2026-01-14
AI Technical Summary
Existing technologies face challenges in capturing reality in a way that allows for the effective combination of computer-generated elements with real-world elements in real time, particularly in industries like entertainment, due to limitations in depth information and practical implementation of RGBD cameras.
A system employing high-definition three-dimensional (HD3D) imaging using an invisible illumination source and two cameras, one for color and one for depth, generating synchronized digital maps that are processed to create a three-dimensional digital representation of a scene with high resolution, low latency, and adjustable frame rates.
Enables the seamless integration of computer-generated and real-world elements by providing high-resolution, low-latency three-dimensional representations suitable for various applications, including film production, live streaming, and volumetric capture.
Smart Images

Figure 2026501088000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-citation to related patent application(s) This U.S. provisional patent application is a national stage application claiming priority from and claiming the benefit of International Patent Application No. PCT / US2023 / 081670, filed November 29, 2022. This application also claims priority from U.S. Provisional Patent Application No. 63 / 429,135, filed November 30, 2022. Both of these patent applications are incorporated herein by reference in their entireties.
[0002] FIELD OF THE INVENTION
[0002] The present invention relates to the field of three-dimensional imaging and various parameters associated therewith.
[0003]
[0003] As should be apparent to those skilled in the art, digital images are employed in many industries for a variety of purposes.
[0004]
[0004] One industry that is experiencing a growing interest in and relying on three-dimensional digital images is the entertainment industry, particularly in the areas of live streaming, filmmaking, and gaming.
[0005]
[0005] There is a growing desire in these industries to combine computer-generated or digital elements with live or real elements or objects in real time. Many names are used to describe the ability to combine computer-generated or digital elements with live or real elements or objects, including augmented reality, mixed reality, visual effects, and virtual production, among others.
[0006]
[0006] However, combining raw and digital elements often presents many challenges and requires specialized tools and techniques and a great deal of manual effort to achieve the best results.
[0007]
[0007] For example, Microsoft has created the "Hololens" product, which displays computer-generated content on a transparent lens in front of a person's face, forming an overlay on top of the information perceived by the human eye in the scene. A heads-up display serves a similar purpose, but for simple computer-generated elements.
[0008]
[0008] In film and TV production, software tools allow computer-generated elements to be superimposed on or behind traditional images. These traditional software tools lack depth information and therefore cannot spatially blend the computer-generated elements with real-world elements.
[0009] Furthermore, in some implementations, a large display wall may be used to display computer-generated content. This large display wall is then filmed to provide a backdrop for the actors. While the display wall provides a proper view of the background scene to the film director, depth information is required to visualize the entire scene, including the foreground and any digital elements between the actors and the background display screen.
[0010]
[0010] Other approaches employ more complex solutions. For example, using color to separate the foreground from the background allows the actual action to be filmed in front of a blue or green screen. This technique is called chromakeying. In this implementation, a motion capture device can track the basic movements of real objects. Software is used to combine the recorded movements with computer-generated elements. Here, laser scanners can also be employed to digitize the positions of stationary objects. However, digitizing the positions of stationary objects requires a significant amount of time to obtain the necessary data.
[0011]
[0011] Some depth or "RGBD" (red, green, blue, depth) cameras have been employed to achieve real-time or near-real-time results, but these attempts have too many implementation constraints to be practical for anything other than limited experimentation.
[0012]
[0012] It has been unclear across many industries how to accomplish mixing real and computer-generated elements, even if such a combination would be useful. For example, Facebook promotes a cartoonish "metaverse," which does not appear to be at all realistic. Other companies have envisioned products in which both real and computer-generated elements appear realistic. Summary of the Invention [Problem to be solved by the invention]
[0013]
[0013] Capturing reality in an acceptable way (including 3D data at sufficiently high frame rates) is a major bottleneck to further development.
[0014]
[0014] Improved solutions are needed. [Means for solving the problem]
[0015] The present invention provides a system and method that addresses one or more deficiencies in the prior art.
[0016]
[0016] Generally speaking, the present invention employs high definition three-dimensional ("HD3D") imaging to enable the effective combination of computer-generated elements and real-world elements, either in real time or as part of a broader offline workflow.
[0017] In one contemplated embodiment, the present invention provides an optical system for generating a three-dimensional digital representation of a scene. The optical system includes an invisible illumination source that generates invisible light and illuminates the scene with the invisible light. The optical system also includes a first camera and a second camera. The first camera is configured to receive light from the scene and generate a first digital map therefrom. The first digital map includes a plurality of first camera pixels, the first camera pixels having a first resolution of approximately 100K to 400M. Each first camera pixel of the plurality of first camera pixels has an associated color. The second camera is configured to receive invisible light from the scene and generate a second digital map therefrom, the invisible light being generated from the illumination source. The second digital map includes a plurality of second camera pixels, the second camera pixels having a second resolution of approximately 100K to 400M. Each second camera pixel of the plurality of second camera pixels has an associated depth. The depth is determined as a function of light time-of-flight information for the invisible light. The system also includes a processor coupled to the first camera and the second camera. The first digital map and the second digital map are generated simultaneously to create a scene correlation therebetween. The processor receives the first digital map and the second digital map and combines the first digital map with the second digital map to generate a three-dimensional digital representation of the scene. The three-dimensional digital representation of the scene satisfies the following parameters: First, the three-dimensional digital representation of the scene has an image resolution of 900,000 image pixels or greater (or ≧0.9 megapixels, abbreviated as ≧0.9M), where the image pixel count includes the final number of pixels that make up the three-dimensional digital representation of the scene. Second, the three-dimensional digital representation of the scene has an image latency of 10 ms to 30 seconds, where the image latency includes the processing time between receipt of the first and second digital maps by the processor and generation of the three-dimensional digital representation of the scene. The three-dimensional digital representation of the scene requires a distance between 10 cm and 200 m, measured from the first and second cameras to the objects in the scene.
[0018] In another contemplated embodiment, the system also includes a display connected to the processor for displaying the three-dimensional representation.
[0019] Furthermore, the system of the present invention may also incorporate a memory coupled to the processor for storing a three-dimensional digital representation of the scene.
[0020]
[0020] The system of the present invention is also contemplated to encompass examples in which a continuous three-dimensional digital representation of a scene is assembled by a processor to generate video at frame rates between 5 frames per second ("fps") and 250 fps.
[0021]
[0021] When the system of the present invention is employed in film production, the 3D digital representation of the scene satisfies at least one of the following parameters: (1) the image resolution has an image pixel count of 0.9M or more, (2) the distance is 10cm to 50m, and (3) the frame rate is 23 to 250fps.
[0022]
[0022] When the system of the present invention is employed in pre-visualization (pre-visualization) of scenes for film production, the 3D digital representation of the scene also meets image latency of 10 ms to 10 s.
[0023]
[0023] When the system of the present invention is employed in the context of a live streaming event, the 3D digital representation of the scene satisfies at least one of the following: (1) an image latency of 10 ms to 1 s, (2) an image resolution of 0.9M or more image pixels, (3) a distance of 2 m to 200 m, and (4) a frame rate of 23 to 250 fps.
[0024]
[0024] When the system of the present invention is used for volumetric capture, the three-dimensional digital representation of the scene satisfies at least one of the following: (1) an image latency of 10 ms to 1 s, (2) an image resolution of 0.9M or more image pixels, (3) a distance of 1 m to 50 m, and (4) a frame rate of 5 to 250 fps.
[0025]
[0025] When the system of the present invention is employed for static capture, the 3D digital representation of the scene satisfies at least one of the following: (1) image latency is 10 ms to 1 s, (2) image resolution is 0.9M or more image pixels, (3) distance is 50 cm to 200 m, and (4) frame rate is 5 to 250 fps.
[0026]
[0026] Furthermore, the first camera and the second camera can be housed in a single housing.
[0027]
[0027] When the first sensor and the second sensor are housed within a single housing, it is assumed that the single housing includes a single optical element and, behind the optical element, a beam splitter that directs light to the first camera and invisible light to the second camera.
[0028]
[0028] The system of the present invention can also be constructed so that the first camera includes a plurality of first cameras, and further, the second camera can also include a plurality of second cameras.
[0029] Another contemplated embodiment of the present invention provides a system for generating a three-dimensional digital representation of a scene. The system includes an invisible light illumination source that generates invisible light and illuminates the scene with the invisible light, a first camera, and a second camera. The first camera is configured to receive light from the scene and generate a first digital map therefrom. The first digital map comprises a plurality of first camera pixels, each having a first resolution of approximately 100K to 400M. Each first camera pixel of the plurality of first camera pixels is associated with a color or grayscale. The second camera is configured to receive invisible light from the scene and generate a second digital map therefrom, the invisible light being generated by the invisible light illumination source. The second digital map comprises a plurality of second camera pixels, each having a second resolution of approximately 100K to 400M. Each second camera pixel of the plurality of second camera pixels is associated with a depth. Depth is determined as a function of optical time-of-flight information for the invisible light. The system also includes a processor coupled to the first camera and the second camera. The first digital map and the second digital map are generated simultaneously to create scene correlation therebetween. The processor receives the first digital map and the second digital map and combines the first digital map with the second digital map to generate a three-dimensional digital representation of the scene. The three-dimensional digital representation of the scene satisfies the following: (1) an image resolution of 0.9M or more image pixels, where the image pixel count includes the number of final pixels that make up the three-dimensional digital representation of the scene; (2) a distance of 10 cm to 200 m, where the distance is measured from the first and second cameras to objects in the scene; and (3) successive three-dimensional digital representations of the scene are assembled by the processor to generate a video having a frame rate of 5 fps to 250 fps.
[0030]
[0030] In this embodiment, the system may also include a display connected to the processor for displaying the three-dimensional representation.
[0031]
[0031] Again, a memory may be connected to the processor for storing a three-dimensional digital representation of the scene.
[0032]
[0032] In this embodiment, for movie production, the three-dimensional digital representation of the scene satisfies at least one of the following: a distance of 10 cm to 50 m, and a frame rate of 23 to 250 fps.
[0033]
[0033] For volumetric capture, the three-dimensional digital representation of the scene satisfies at least one of the following: a distance between 1 m and 50 m, and a frame rate between 5 and 250 fps.
[0034] For static capture, the three-dimensional digital representation of the scene satisfies at least one of the following: a distance of 50 cm to 200 m, and a frame rate of 5 to 250 fps.
[0035] In this embodiment, the first camera and the second camera can be housed in one housing.
[0036]
[0036] It is contemplated that this embodiment may include a beam splitter that directs light to a first camera and non-visible light to a second camera.
[0037]
[0037] As mentioned above, the first camera may include multiple first cameras, and the second camera may include multiple second cameras.
[0038]
[0038] It should be noted that the present invention is not limited to the features and aspects listed above. Further features and aspects of the present invention will become apparent from the following discussion.
[0039]
[0039] The present invention will now be described in connection with the drawings accompanying this specification. [Brief explanation of the drawings]
[0040] [Figure 1] 1 is a diagrammatic representation of a first possible embodiment of an optical system of the present invention. [Figure 2] 1 is a diagrammatic representation of a second embodiment of an optical system of the present invention. [Figure 3] 10 is a diagrammatic representation of a third embodiment of the optical system of the present invention. [Figure 4] 1 is a diagrammatic representation of one possible manipulation of digital information in generating a three-dimensional digital representation of a scene. [Figure 5A] 1 shows a comparison between a digital map produced by a prior art system and a digital map produced by the optical system of the present invention. [Figure 5B] 1 shows a comparison between a digital map produced by a prior art system and a digital map produced by the optical system of the present invention. [Figure 6] 1 shows various images illustrating aspects associated with manipulating a three-dimensional digital representation of a scene created in accordance with the present invention. [Figure 7] 1 shows various images illustrating aspects associated with manipulating a three-dimensional digital representation of a scene created in accordance with the present invention. [Figure 8] 1 shows various images illustrating aspects associated with manipulating a three-dimensional digital representation of a scene created in accordance with the present invention. [Figure 9] 1 shows various images illustrating aspects associated with manipulating a three-dimensional digital representation of a scene created in accordance with the present invention. [Figure 10] 1 shows various images illustrating aspects associated with manipulating a three-dimensional digital representation of a scene created in accordance with the present invention. [Figure 11A] 1 shows various images illustrating aspects associated with manipulating a three-dimensional digital representation of a scene created in accordance with the present invention. [Figure 11B] 1 shows various images illustrating aspects associated with manipulating a three-dimensional digital representation of a scene created in accordance with the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] Detailed Description of Embodiment(s) of the Invention
[0046] The present invention will now be described in connection with various examples and embodiments. The present invention should not be understood to be limited to only the examples and embodiments discussed. On the contrary, the discussion of selected examples and embodiments is intended to emphasize the breadth and scope of the invention, rather than to be limiting. As should be apparent to those skilled in the art, variations and equivalents of the described examples and embodiments may be employed without departing from the scope of the invention.
[0042]
[0047] Additionally, aspects of the present invention are discussed in connection with particular materials and / or components. These materials and / or components are not intended to limit the scope of the invention. As would be apparent to one skilled in the art, alternative materials and / or components may be employed without departing from the scope of the present invention.
[0043]
[0048] In the accompanying figures, for convenience and brevity, the same reference numbers are used to refer to like features in various examples and embodiments of the invention. The use of the same reference numbers for identical or similar structures and features is not intended to convey that each element bearing the same reference number is identical to all other elements bearing the same reference number. On the contrary, these elements may vary in various ways from embodiment to embodiment without departing from the scope of the invention.
[0044]
[0049] Furthermore, in the discussion that follows, the terms "first," "second," "third," etc. may be used to refer to similar elements. These terms are employed to distinguish similar elements from similar instances of the same element. For example, one fastener may be distinguished from other fasteners by calling it a "first" fastener. The other fastener may be called a "second fastener." The terms "first," "second," and "third" are not intended to convey any particular hierarchy between the elements so called.
[0045]
[0050] It should be noted that the use of "first," "second," and "third," etc., is intended to follow common grammatical conventions. Thus, while a component may be designated as "first" in one instance, in another instance, this same component may be referred to as "second," "third," etc. Thus, the use of "first," "second," and "third," etc., is not intended to limit the present invention.
[0046]
[0051] Before discussing the present invention, the following documents are incorporated herein by reference: U.S. Patent Nos. 8,471,895, 9,007,439, 10,218,962, 10,104,365, 10,437,082, 8,493,645, and 8,254,009.
[0047]
[0052] FIG. 1 is a diagrammatic representation of an optical system 10 according to a first embodiment of the present invention.
[0048]
[0053] It should be noted that the term "optical system" is employed because the system of the present invention involves light and optical components. The use of the word "optical" should not be understood to limit the scope of the present invention. As will be discussed in more detail in the following pages, the system may require components other than optical components.
[0049]
[0054] Optical system 10 combines various components that together generate a three-dimensional digital representation of scene 12. In Figure 1, scene 12 contains three representative objects: a person 14, a locomotive 16, and a tree 18. Optical system 10 of the present invention captures light reflected from scene 12 and renders a three-dimensional digital representation from this light.
[0050]
[0055] The light received by optical system 10 includes two components, which are described in more detail below. First, optical system 10 receives light from scene 12. This light is used to generate a color map, which is used to create a three-dimensional digital representation of scene 12. It should be noted that the light from scene 12 may be provided by natural and / or artificial light sources. For example, the light from scene 12 may combine natural sunlight with light provided by, for example, stage spotlights. Second, optical system 10 receives invisible light from scene 12. The invisible light is used to generate a depth map, which, together with the color map, is used to create a three-dimensional digital representation of scene 12.
[0051]
[0056] With respect to the color map portion of the three-dimensional digital representation of scene 12, note that this color encompasses two variations. First, the color map can encompass actual colors, meaning red, blue, and green ("RGB") components. Second, the color map can be a black-and-white color map, meaning that the color map encompasses a grayscale image. In the following paragraphs, references to an RGB color map and to a grayscale digital map will be used interchangeably.
[0052]
[0057] To generate invisible light, optical system 10 incorporates invisible light illumination source 20. Invisible light illumination source 20 generates invisible light 22 that is used to illuminate scene 12. Invisible light 22 may be any type of invisible light 22 from the electromagnetic spectrum. Invisible light 22 encompasses light that is outside of human vision. For optical system 10 described herein, it is assumed that invisible light 22 is either infrared ("IR") light or ultraviolet ("UV") light.
[0053]
[0058] The optical system 10 includes a first camera 24 and a second camera 26 .
[0054]
[0059] Referring to FIG. 4 , first camera 24 is a color camera incorporating first sensor 28 having a plurality of first camera pixels 30. In operation, each first camera pixel in the plurality of first camera pixels 30 is associated with a color or gray value. First sensor 28 can capture light 32 and convert this light into a first digital map 34. First digital map 34 is also referred to herein as a color digital map. As noted above, first digital map 34 may be an RGB map or a grayscale digital map. It should be noted that each first camera pixel in the plurality of first camera pixels 30 is associated with a color at each azimuthal and longitudinal position in the two-dimensional field of view of first camera 24.
[0055]
[0060] 4 , second camera 26 is a depth camera incorporating a second sensor 36 comprising a plurality of second camera pixels 38. In operation, each second camera pixel in the plurality of second camera pixels 38 is associated with a depth value. Second sensor 36 can capture invisible light 40 reflected from scene 12 and convert the invisible light 40 into a second digital map 42, also referred to herein as a depth digital map. It should be noted that each second camera pixel in the plurality of second camera pixels 38 is associated with a depth to a surface and further with the intensity of light reflected from that surface at each azimuthal and longitudinal position in the two-dimensional field of view of second camera 26.
[0056]
[0061] With respect to the depth values, it is noted that optical time-of-flight (“oTOF”) information associated with the invisible light 40 received by the second sensor 36 from the scene 12 is used to generate the depth values, as will be explained in more detail later in this specification.
[0057]
[0062] 4, the first digital map 34 and the second digital map 42 are input to a processor 46. The processor 46 combines the first digital map 34 and the second digital map 42 to create a three-dimensional digital representation 44 of the scene 12. The processor 46 should not be construed as a single processor. The processor 46 may include multiple components (e.g., different processors) without departing from the scope of the present invention. When there are multiple processors 46, the multiple processors 46 may operate at different times, as discussed in more detail below.
[0058]
[0063] The plurality of first camera pixels 30 in the first camera 24 generates a first pixel map 34 at a first resolution of approximately 100K-400M. As should be apparent to one skilled in the art, this indicates that there are approximately 100K-400M first camera pixels that make up the plurality of first camera pixels 30.
[0059]
[0064] Similarly, the plurality of second camera pixels 38 in the second camera 26 generates a second pixel map 42 at a second resolution of approximately 100K-400M. This therefore indicates that there are approximately 100K-400M second camera pixels making up the plurality of second camera pixels 38.
[0060]
[0065] It is noted that the first resolution may be lower than, equal to, or higher than the second resolution. In one contemplated embodiment, the first resolution is equal to the second resolution, although this is not intended to be a limitation of the present invention.
[0061]
[0066] It is assumed that the first camera 24 and the second camera 26 operate synchronously. Specifically, the cameras 24, 26 function such that for each three-dimensional digital representation 44 generated, the first digital map 34 and the second digital map 42 have a scene correlation therebetween. It is desirable that the information recorded in the first digital map 34 be at least substantially identical to the information recorded in the second digital map 42 in terms of what action is occurring in the scene 12. Because the information in the digital maps 34, 42 correlates with each other, the first digital map 34 and the second digital map 42 can be processed synchronously (or nearly synchronously) to create the three-dimensional digital representation 44 of the scene 12. From one perspective, it can be understood that because the first digital map 34 is generated at approximately (or nearly simultaneously) the same time as the second digital map 42, both digital maps 34, 42 effectively capture the same information from the scene 12. Looking at this from a slightly different perspective, it can be seen that the delay between the generation of the first digital map 34 and the generation of the second digital map 42 is less than about one-half (half) the frame length.
[0062]
[0067] 1, the processor 46 in the optical system 10 is connected to the first camera 24 via a first communication link 48. The processor 46 is connected to the second camera 26 via a second communication link 50. And, as shown, the processor 46 is connected to the invisible light illumination source 20 via a third communication link 52.
[0063]
[0068] 1, processor 46 may also be connected to a display 54 via a fourth communication link 56. Processor 46 may also be connected to a memory 58 via a fifth communication link 60.
[0064]
[0069] It is contemplated that processor 46 may be of any type capable of executing instructions in the form of software to generate three-dimensional digital representation 44 by combining first digital map 34 with second digital map 42. Any suitable processor 46 may be employed for this purpose.
[0065]
[0070] It should be noted that processor 46, in this embodiment, is assumed to generate instructions to invisible light illumination source 20 to generate invisible light 22 used to illuminate scene 12. Alternatively, instructions may be issued to invisible light illumination source 20 through other means that should be apparent to one skilled in the art.
[0066]
[0071] The three-dimensional digital representation 44 is assumed to satisfy at least one of the following parameters:
[0067]
[0072] First, it is assumed that the three-dimensional digital representation 44 has an image resolution of ≥ 0.9M image pixels, which is understood to be appropriate for the final pixel display. More specifically, the image pixel count may match the display pixel count associated with the display 54, although pixel matching is not required for the practice of the present invention.
[0068]
[0073] As should be apparent to one skilled in the art, a standard commercially available display 54 may have a display pixel count (or "pixel density") ranging anywhere from 720p to 8K. Currently, these resolutions include 720p (1280 x 720 pixels, 921,600 pixels total), 1080p (1920 x 1080 pixels, 2,073,600 pixels total), 1440p (2560 x 1440 pixels, 3,686,400 pixels total), 2K (2048 x 1080 pixels, 2,211,840 pixels total), 4K (3480 x 2160 pixels, 7,516,800 pixels total), 5K (5120 x 2880 pixels, 14,745,600 pixels total), and 8K (7860 x 4320 pixels, 33,955,200 pixels total). It is assumed that the image resolution of three-dimensional digital representation 44 matches the display resolution of display 54 so that three-dimensional digital representation 44 can be easily displayed on display 54. Alternatively, the image resolution of the three-dimensional digital representation 44 may be higher or lower than the display resolution, in which case appropriate corrections may be made to display the three-dimensional digital representation 44 on the display 54, as would be apparent to one skilled in the art.
[0069]
[0074] Second, the three-dimensional digital representation 44 is assumed to have an image latency of 10 ms to 30 seconds, which constitutes the processing time between receipt of the first and second digital maps 34, 42 by the processor 46 and generation of the three-dimensional digital representation 44 of the scene 12.
[0070]
[0075] It should be noted that the present invention does not limit image latency to a range of 10 ms to 30 seconds. In the context of the present invention, image latency can be defined as any range defined by 10 ms endpoints from 10 ms to 30 seconds. For example, in one contemplated embodiment, image latency can range from 500 ms to 1 second. In other embodiments, image latency can range from 10 ms to 1900 ms. Additionally, the present invention encompasses any specific image latency equal to each 10 ms endpoint from 10 ms to 30 seconds. In this context, specific values for image latency can be, by way of example, 10 ms, 20 ms, 30 ms, 40 ms, 50 ms, and 1010 ms.
[0071]
[0076] Third, the three-dimensional digital representation 44 is assumed to incorporate depth information encompassing distances between 10 cm and 200 m. Distances are measured from the first and second cameras 24, 26 to objects in the scene 12 (e.g., people 14, locomotive 14, and / or trees 18). In other words, objects in the scene 12 are assumed to be at distances between 1 and 200 m from the first and second cameras 24, 26.
[0072]
[0077] It should be noted that the present invention is not limited to distances ranging from 10 cm to 200 m. In the context of the present invention, distances may be defined as any range from 10 cm to 200 m, defined by 10 cm endpoints. For example, in one contemplated embodiment, the distance may be from 10 cm to 5 m. In other embodiments, the distance may range from 20 m to 100 m. Additionally, the present invention encompasses any specific distance equal to each 10 cm endpoint from 10 cm to 200 m. In this context, specific values for image latency may be, by way of example, 10 cm, 20 cm, 1 m, 5 m, 10 m, and 100 m.
[0073]
[0078] 1, a first object distance 62 is shown from the first camera 24 and the second camera 26 to the person 12. A second object distance 64 indicates the separation between the locomotive 16 and the first and second cameras 24, 26. Finally, a third object distance 66 identifies the separation between the tree 18 and the first and second cameras 24, 26. Each of these distances 62, 64, 66 is expected to fall within a distance range of 10 cm to 200 m.
[0074]
[0079] 1 also shows a camera separation distance 68. The camera separation distance 68 is the distance that separates the first camera 24 from the second camera 26.
[0075]
[0080] 1, it is contemplated that first camera 24 need only be separated from second camera 26 by a distance identified as camera separation distance 68. It is contemplated that the camera separation distance may be in the range of approximately 1 cm to 2 m. Furthermore, it is contemplated that first camera 24 and second camera 26 may be combined into a single unit that detects both light 32 and invisible light 40.
[0076]
[0081] It should be noted that the present invention does not limit the camera separation distance to a range of 1 cm to 2 m. In the context of the present invention, the camera separation distance can be defined as any range from 1 cm to 2 m, defined by 1 cm endpoints. For example, in one contemplated embodiment, the camera separation distance may be from 1 cm to 1 m. In other embodiments, a range of 50 cm to 1.5 m is contemplated for the camera separation distance. Additionally, the present invention encompasses any specific camera separation distance equal to each 1 cm endpoint from 1 cm to 2 m. In this context, specific values for image latency can be, for example, 1 cm, 2 cm, 5 cm, 10 cm, and 1 m.
[0077]
[0082] In one embodiment of the present invention, as noted above, it is also contemplated that optical system 10 includes a display 54. As shown, display 54 is contemplated to be connected to processor 46 via a fourth communications link 56. If provided, display 54 displays three-dimensional digital representation 44. It is noted that display 54 may be omitted without departing from the scope of the present invention.
[0078]
[0083] 1, the optical system 10 may also include a memory 58 connected to the processor 46 via a fifth communication link 60. The memory 58 is intended to satisfy one or more operating parameters. As should be apparent, the memory 58 provides a location where the three-dimensional digital representation 44 may be stored. Furthermore, the memory 58 may also store software for execution by the processor 46. It should be noted that the memory 58 may be omitted without departing from the scope of the present invention.
[0079]
[0084] As noted above, one embodiment of the present invention contemplates that display 54 is configured with a display resolution that matches the image resolution of three-dimensional digital representation 44. However, this configuration need not remain within the scope of the present invention. Other embodiments contemplate that the image resolution of three-dimensional digital representation 44 may differ from the display resolution.
[0080]
[0085] Up to this point, three-dimensional digital representation 44 of scene 12 has been described as being a single frame, or still image, of scene 12 .
[0081]
[0086] When successive three-dimensional digital representations 44 of scene 12 are collected sequentially, they form a video. The present invention contemplates that successive three-dimensional digital representations 44 of scene 12 may be collected sequentially by processor 46 to generate a video. When doing so, it is contemplated that the video will have a frame rate of between about 5 and 250 frames per second ("fps").
[0082]
[0087] It should be noted that the present invention does not limit the frame rate to a range between 5 and 250 fps. In the context of the present invention, the frame rate can be defined as any range defined by endpoints in 5 fps increments from 5 fps to 250 fps. For example, in one contemplated embodiment, the frame rate can be from 5 fps to 50 fps. In other embodiments, the frame rate can range from 15 fps to 25 fps. Additionally, the present invention encompasses any specific camera separation distance equal to each 5 fps endpoint from 5 fps to 250 fps. In this context, specific values for image latency can be, for example, 5 fps, 10 fps, 15 fps, 50 fps, and 100 fps. The present invention also encompasses the standard frame rate of 24 fps.
[0083]
[0088] FIG. 2 shows a diagrammatic representation of a second embodiment of an optical system 70 contemplated by the present invention.
[0084]
[0089] In this embodiment, optical system 70 shares many of the features described in connection with optical system 10 shown in FIG. 1. In this embodiment of optical system 70, invisible light illumination source 20, first camera 24, and second camera 26 are disposed within a first housing 72. First housing 72 includes a first optical element 74 that directs light 32 to first camera 24. First housing 72 also includes a second optical element 76 that directs invisible light 40 to second camera 26. First and second optical elements 74, 76 may be, for example, lenses. Alternatively, first and second optical elements 74, 76 may be apertures in first housing 72. Furthermore, first and second optical elements 74, 76 may be a combination of lenses and / or apertures, as needed and / or desired.
[0085]
[0090] FIG. 3 is a diagrammatic representation of a third embodiment of an optical system 78 according to the present invention.
[0086]
[0091] In optical system 78, first camera 24 and second camera 26 are enclosed in second housing 80 such that first camera 24 is positioned perpendicular to second camera 26. It should be noted that the perpendicular arrangement of first camera 24 relative to second camera 26 is merely an example and is not a limitation of the present invention. Any suitable spatial positioning of first camera 24 relative to second camera 26 is considered to fall within the scope of the present invention.
[0087]
[0092] In this embodiment, light 32 and invisible light 40 pass through a third optical element 82 and enter second housing 80. Third optical element 82 may be a lens or multiple lenses. Alternatively, as before, third optical element 82 may be an aperture in second housing 80. Furthermore, third optical element 82 may be a combination of lenses and / or apertures as needed and / or desired.
[0088]
[0093] In optical system 78, light 32 and invisible light 40 pass through fourth optical element 84. At fourth optical element 84, light 32 and invisible light 40 are separated from each other. For this reason, fourth optical element 84 is also referred to as a light splitter. In the illustrated embodiment, light splitter 84 passes light 32 to reach first camera 24. Since invisible light 40 is redirected perpendicular to light 32, invisible light 40 is directed to second camera 26.
[0089]
[0094] In the illustrated embodiment, it is assumed that the fourth optical element 84 is an optical component called a beam splitter, however, any other optical component that splits light into light 32 and invisible light 40 may be employed without departing from the scope of the present invention.
[0090]
[0095] Optical system 70 and optical system 78 are assumed to operate in a similar manner as discussed in connection with optical system 10 .
[0091]
[0096] The present invention contemplates that optical systems 10, 70, 78 may operate in at least one of five possible configurations, including, but not limited to, (1) film production, (2) "previs" film production, (3) live streaming production, (4) volumetric capture, and (5) static capture. Each of these five configurations is discussed below.
[0092]
[0097] In a first configuration, the optical systems 10, 70, 78 are intended to operate for film production.
[0093]
[0098] Filmmaking involves the creation of video, capturing visual information from, for example, actors positioned within a scene 12. When optical systems 10, 70, 78 are employed in a filmmaking context, they capture "raw data" of the performances performed by the actors within the scene 12. The "raw data" can include, for example, actors performing in front of a blue or green screen, which is a blue or green background employed by those skilled in the art. Once the actors perform in front of the blue or green screen, background elements are inserted at a later date, for example, during editing. Alternatively, actors may perform in front of one or more light-emitting diode ("LED") walls (otherwise referred to as "light-emitting walls") comprised of digitally generated backgrounds. Both approaches are commonly found in virtual production for movies and feature films, for cinema, TV, or personal device viewing.
[0094]
[0099] To operate for film production, it is assumed that the optical systems 10, 70, 78 operate according to the following parameters: First, the image resolution has an image pixel count of 0.9 megapixels ("M") ("0.9M") or greater. This refers to the total number of pixels in the three-dimensional digital representation 44. Second, the distance from the first and second cameras 24, 26 to the scene 12 is between approximately 10 cm and 50 m. Third, the video frame rate is between approximately 23 and 250 fps.
[0095]
[0100] Film production also encompasses post-production processing to create a final film ready for projection to an audience. Post-production editing and manipulation of the action captured during the film production phase can take days, weeks, months, or even years to generate the three-dimensional digital representation 44, where latency is not a parameter. Additionally, processor 46 can encompass various processors operating at various points after production.
[0096]
[0101] In a second configuration, the optical systems 10, 70, 78 are intended to operate for "pre-visual" film production.
[0097]
[0102] "Previs" filmmaking differs from cinematic production in that the optical system 10, 70, 78 incorporates the display 54 and the three-dimensional digital representation 44 incorporates at least a roughly rendered environment into which actors are inserted. The term "previs" means "pre-imaged" and indicates that the information generated by the optical system 10, 70, 78 comes after the film production stage, but before any editing and production stages in which the final environment is rendered.
[0098]
[0103] It is intended that the previs video will be shown to the director and / or producer of the film immediately after the film is made so that the producer and / or director can judge whether the actors have performed satisfactorily. In effect, "previs" filmmaking may be understood as a preview of the final film after editing is complete.
[0099]
[0104] To operate for previsualization filmmaking, the optical systems 10, 70, 78 are expected to operate according to the same parameters identified for filmmaking. For previsualization filmmaking, the optical systems 10, 70, 78 are also expected to meet image latency of approximately 10 ms to 10 s.
[0100]
[0105] In a third configuration, the optical systems 10, 70, 78 are intended to operate in conjunction with a live-streamed production. This differs from film production and pre-visualization film production in that the action is streamed live, as would be expected for a live sports event or music concert, for example. There is expected to be a delay between the capture of the live image and the distribution of this live image to the audience, but this delay is short.
[0101]
[0106] To operate for live streaming production, the optical system 10, 70, 78 is assumed to operate according to the following parameters: First, the image resolution is 0.9M or more image pixels. Second, the distance from the first and second cameras 24, 26 to the scene 12 is between approximately 2m and 200m. Third, the video frame rate is between approximately 23 and 250fps. Fourth, the latency is between approximately 10ms and 1s.
[0102]
[0107] In a fourth configuration, the optical systems 10, 70, 78 are assumed to operate for volumetric capture.
[0103]
[0108] Volumetric capture refers to the capture of images and / or video within a defined volumetric space, such as a sound stage. Volumetric capture is employed when the elements and actors of a scene 12 are positioned within a predetermined space. Volumetric capture may require generating a three-dimensional digital representation 44 from multiple viewpoints and / or multiple angles with respect to the volumetric space.
[0104]
[0109] Volumetric capture requires at least two images taken from different viewpoints to generate a three-dimensional digital representation.
[0105]
[0110] To operate for volumetric capture, it is assumed that the optical system 10, 70, 78 operates according to the following parameters: First, the image resolution has an image pixel count of 0.9M or more. Second, the distance from the first and second cameras 24, 26 to the scene 12 is between approximately 1 m and 50 m. Third, the video frame rate is between approximately 5 and 250 fps. Fourth, the latency is between approximately 10 ms and 1 s.
[0106]
[0111] In a volumetric capture variant with post-production processing, the three-dimensional digital representation 44 may take hours, days, weeks, months, or even years to produce, depending on the amount of post-production processing required and / or desired. Therefore, latency is not a parameter in this variant. Therefore, in this post-production still camera capture environment, the optical system 10, 70, 78 operates according to the following parameters: First, the image resolution is 0.9M image pixels or greater. Second, the distance from the first and second cameras 24, 26 to the scene 12 is between approximately 1 m and 50 m. Third, the video frame rate is between approximately 5 and 250 fps.
[0107]
[0112] In a fifth configuration, the optical systems 10, 70, 78 are assumed to operate for static capture. Static capture refers to the capture of images and / or video from a single perspective, such as would exist if a person were to take a photograph using the camera on their cell phone. Note that static capture encompasses two distinct conditions. In the first condition, referred to as “static scene capture,” the cameras 24, 26 move relative to the scene 12, and the objects 14, 16, 18 in the scene 12 are stationary. In the second condition, referred to as “static camera capture,” the cameras 24, 26 are stationary relative to the scene 12, and the objects 14, 16, 18 move within the scene 12. Both conditions are intended to be encompassed by the use of the term “static capture.”
[0108]
[0113] For both still scene capture and still camera capture, it is assumed that the optical system 10, 70, 78 operates according to the following parameters: First, the image resolution is 0.9M image pixels or greater. The distance from the first and second cameras 24, 26 to the scene 12 is between approximately 50 cm and 200 m. Third, the video frame rate is between approximately 5 and 250 fps. Fourth, the latency is between approximately 10 ms and 1 s.
[0109]
[0114] In a variant of still camera capture with post-production processing, the three-dimensional digital representation 44 may take hours, days, weeks, months, or even years to produce, depending on the amount of post-production processing required and / or desired. Therefore, latency is not a parameter in this variant. Therefore, in this post-production still camera capture environment, the optical system 10, 70, 78 operates according to the following parameters: First, the image resolution is 0.9M image pixels or greater. Second, the distance from the first and second cameras 24, 26 to the scene 12 is between approximately 50 cm and 200 m. Third, the video frame rate is between approximately 5 and 250 fps.
[0110]
[0115] To facilitate a better understanding of the optical systems 10, 70, 78 discussed in connection with Figures 1-4, the following additional information is provided and is intended to apply to each of the optical systems 10, 70, 78.
[0111]
[0116] The optical systems 10, 70, 78 of the present invention encompass systems and methods that enable the generation of high-resolution images of a scene, including wide-field-of-view scenes. Specifically, these systems and methods contemplate simultaneously recording three-dimensional position information for multiple objects in a scene, with high spatial and distance resolution, along with luminous intensity (grayscale or color) information about the scene. Both color and luminous intensity information are recorded for every pixel in the pixel array for each image. The luminous intensity and position information are combined into a single three-dimensional image that approximates the scene as seen by humans, which is referred to herein as the three-dimensional digital representation 44.
[0112]
[0117] The availability of dense depth data can be used to accomplish many things that are currently done manually or with many complex steps and real scenes. This depth information can be in the form of a 2D depth map, giving distance values to the surface imaged by that pixel. Or, the depth information can be in the form of a more extensive 3D representation of the locations of surfaces within an area or scene, such as a point cloud or surface mesh or voxel grid or similar way of representing such 3D location information. Finally, such a representation can also describe how such location information changes over time.
[0113] General Description
[0118] Capturing reality in 3D has been known since almost the beginning of photography.
[0114]
[0119] There are many ways in which people have attempted to determine the 3D position of a surface or a segment of a surface.
[0115]
[0120] Photogrammetry is an example of a broader class of techniques (stereo, structured light, structure from motion) that use geometry to determine the position of points, objects, or surfaces. Photogrammetry is similar to triangulation in navigation. Microsoft, Apple, and Intel, as well as other companies, have produced products based on this type of technique, but with poor performance.
[0116]
[0121] Furthermore, geometry-based methods are known to have limited operating ranges, typically only a few meters. Furthermore, geometry-based methods require significant computational power to correctly calculate position solutions for pixel densities suitable for imaging, e.g., >100,000 points. Such products have been used to capture 3D information for many applications and use cases, from robotics to film. However, the performance of these products is so poor that many have concluded that such 3D cameras or capture devices are not suitable for these use cases.
[0117]
[0122] Other traditional 3D capture techniques use electronic means to detect the time it takes for illuminating light to hit a camera or sensor and return. This is commonly referred to as time-of-flight ("TOF"), and has many subcategories, including direct TOF, indirect TOF, phase-based TOF, frequency-modulated continuous wave ("FMCW"), amplitude-modulated continuous wave ("AMCW"), linear-mode and Geiger-mode avalanche photodiodes ("APDs"), and single-photon avalanche diodes ("SPADs"), among others. All of these techniques have in common some form of electronic detector that measures the phase change or time-of-arrival of the returning light.
[0118]
[0123] For example, single-point scanners can achieve very high accuracy and point density over long ranges, up to 1 km or more, but they require minutes to hours to collect a reasonable number of points.
[0119]
[0124] Multi-point scanners typically have less accuracy and lower point density than single-point scanners, and also take many seconds to collect a reasonable number of points.
[0120]
[0125] Imaging arrays capable of capturing image-like depth data are limited to short distances, e.g., < a few meters, or very low point densities, e.g., < 20,000 points or pixels, or both. While these imaging arrays can achieve extended distances, they also suffer from very high costs, e.g., > $10,000. While imaging arrays have been tried for many use cases, with the exception of static area scanning where high cost is acceptable, their performance has been so poor that they have not been adopted, and many experts have concluded that such technology is not suitable for these use cases.
[0121]
[0126] A critical class of 3D capture technologies is based on optics, using the properties of light itself to determine distance changes. Traditionally, these 3D capture technologies have included interferometry techniques and coherent holography. Such systems are excessively expensive (e.g., over $100,000). While highly accurate, these technologies do not work well outside of a laboratory and have limitations in operating range and / or point density. They are not compatible with the broader commercial use cases described herein.
[0122]
[0127] As described in detail herein, the optical systems 10, 70, 78 of the present invention can provide new approaches to solving the problem of 3D capture in a practical way, enabling the use of 3D information to expand and improve these use cases. Some of the key elements required to achieve improvements in these use cases include resolution, operating range (e.g., distance), and speed (in terms of latency and capture rate).
[0123]
[0128] As noted above, the optical systems 20, 70, 78 of the present invention rely on all-optical TOF ("oTOF") techniques to capture the information used to generate the second digital map.
[0124]
[0129] In one contemplated embodiment, oTOF uses an external modulation device in front of a conventional sensor (e.g., second camera 26) coupled with pulsed illumination (e.g., by invisible light illumination source 20) to create a modulated image and an unmodulated, or reference, image. The ratio of these two images, multiplied by an appropriate coefficient, is a direct measure of the time it takes for the illumination light to return to the camera. This is achieved without measuring time using any electronic means (e.g., a second digital map). The combination of these variables results in a device that can simultaneously capture three key elements in one device, e.g., optical system 10, 70, 78. An added benefit is that this technology reduces costs similar to any 2D camera, making it affordable for almost any industry or use case.
[0125] resolution
[0130] The lateral resolution, or transverse resolution, or point density of a 3D point relates to the smallest feature size that can be distinguished, manipulated, or measured. Just like with regular 2D images, higher density means finer detail, and more detail can be captured and / or used.
[0126]
[0131] While an estimated 20,000 or 100,000 pixel image may have been useful in the 1800s or early 1900s, all of today's applications expect something close to HD, defined as approximating human vision. It is also useful to note that even the upscaling and heavy filtering that some 3D camera manufacturers do to achieve the 300,000 or even 1,000,000 pixel publishing specifications does not achieve the required performance.
[0127]
[0132] For example, Figures 5A and 5B show some comparisons between images 86, 88, 90 (Figure 5A) produced by current (prior art) devices and images 92, 94, 96 (Figure 5B) of the optical systems 10, 70, 78 of the present invention.
[0128]
[0133] Referring to the top images in Figures 5A and 5B, a depth map (a grayscale image representing distance per point) from Microsoft Kinect Azure, which extends up to 5 meters, is compared with a 700,000-point oTOF depth map (produced by the optical system 10, 70, 78 of the present invention), which extends up to 10 meters. Image 86 illustrates the Microsoft Kinect Azure approach. As shown, the actual resolution (e.g., the size of features that can be easily detected or identified) of the optical system 10, 70, 78 is much higher than the prior art. For example, fingers, the brim of a hat, and the spokes on a bench cart wheel are clearly visible in the 3D information, e.g., the second digital map generated by the optical system 10, 70, 78. This is image 92.
[0129]
[0134] Continuing with reference to Figures 5A and 5B, a second image compares a typical (prior art) point cloud (a collection of 3D points rendered in 3D coordinates) from a photogrammetric solution at scales up to 50 meters (Figure 5A, image 88) with a point cloud created in accordance with the present invention from an oTOF video feed out to 20 meters in the field (Figure 5B, image 94). As can be seen, image 94 produced in accordance with the present optical system 10, 70, 78 shows branches and leaves even at these long distances.
[0130]
[0135] It has been shown that expensive precision laser scanners can provide resolution similar to that of the oTOF point cloud, e.g., the present three-dimensional digital representation 44. However, such laser scanners take much longer to generate an output than optical systems 10, 70, 78.
[0131]
[0136] Referring again to Figures 5A and 5B, the bottom images compare a depth map from a multi-point LIDAR scanner (prior art, image 90) with a 700,000-point oTOF depth map (e.g., three-dimensional digital representation 44) (present invention, image 96). Color in these images represents the distance at each point, or pixel. In the image on the right, the tricycle is approximately 30 meters from the camera, or capture device. As should be apparent from these three comparisons, the present optical systems 10, 70, and 78 produce three-dimensional digital representations 44 with improved resolution and point density. This improvement is evidenced by the ability to identify features or objects, which is important in the present invention.
[0132]
[0137] As noted above, for some projects, it may be desirable to obtain depth maps with pixel counts greater than about 100,000 points per pixel (with spatial frequency performance appropriate for such pixel counts, instead of Kinect prior art devices). For other projects, depth maps with greater than about 300,000 points per pixel (VGA equivalent), greater than about 500,000 points per pixel, greater than about 700,000 points per pixel, or greater than about 1,000,000 points per pixel may be desired. Still further, it may be desirable to produce formats roughly equivalent to 480p, 720p, 1080p, 2K, 4K, or 8K images. Still other formats between these formats, or even larger formats, may be desired. The optical systems 10, 70, 78 of the present invention, unlike the prior art, are able to achieve these objectives.
[0133] range
[0138] The range of the optical systems 10, 70, 78 of the present invention also provides an improvement over the prior art, and range is discussed herein as the distances 62, 64, 66 from the first and second cameras 24, 26 to the objects 14, 16, 18 in the scene 12. The difference in operating range can also be compared with reference to Figures 5A and 5B.
[0134]
[0139] The Microsoft Kinect system at the top of Figure 5A cannot achieve a distance / range of 20-30 meters. Even at 5 meters, there are many shortcomings that make this product a poor solution for many of the use cases described herein.
[0135]
[0140] The use cases described below may involve indoor sets where objects may be within 1 meter of the camera, within 3 meters of the camera, or may be placed further away. For example, objects may be placed >3 meters, >10 meters, >20 meters, >30 meters, or even >100 meters from the first and second cameras 24, 26. Depending on the set or project, objects may be stationary or moving, or may be displayed as moving or still images on a video display. These objects may be located between 1 meter and 10 meters from the first and second cameras 24, 26, or between 3 meters and 30 meters from the first and second cameras 24, 26, or between 1 meter and 30 meters, or between 2 meters and 20 meters, or in other location-specified ranges, as appropriate. Sometimes objects may move outside of these ranges / distances. For outdoor use, objects may be positioned in the same manner as indoors. They may also be positioned between 10 m and 100 m from the first and second cameras 24, 26, between 1 m and 100 m from the first and second cameras 24, 26, between 10 m and 50 m, between 3 m and 50 m, or any other convenient location range. The optical systems 10, 70, 78 of the present invention can accommodate these distances. Furthermore, depth cameras such as Kinect cameras exhibit degradation in ranging and accuracy in outdoor scenes. The optical systems 10, 70, 78 of the present invention can accommodate outdoor scenes or indoor scenes with time-varying lighting.
[0136] Speed / Latency
[0141] Speed refers to how long it takes to acquire and use 3D data of an object. For example, how long it takes to acquire the equivalent of a frame of depth data (or a depth map or point cloud). This refers to the latency, or delay, from a physical event or time until the 3D data is available for use in a workflow. This refers to how quickly the 3D data is continuously available (e.g., frame rate). In the context of the optical systems 10, 70, 78 of the present invention, latency encompasses the time between when images are captured by the first and second cameras 24, 26 and the generation of the three-dimensional digital representation 44 by the processor 46.
[0137]
[0142] Furthermore, large studio photogrammetry solutions (such as those built by Canon, Microsoft, or Intel) can take hours or even days to perform all the calculations necessary to compute one second of 3D data. Even smaller multi-camera volumetric capture solutions can take days to compute 3D data. Using prior art devices, such large time scales are impractical and too expensive for these use cases.
[0138]
[0143] Referring again to Figures 5A and 5B, it should be noted that for the second image in Figure 5B, the point cloud image, prior art precision laser scanners can achieve resolutions approaching those of the present invention. However, to do so, precision laser scanners known in the prior art require minutes to tens of minutes per scan to collect. Furthermore, avoiding shadows or occlusions typically requires multiple scans. In other words, the latency of the prior art is simply prohibitive for generating three-dimensional digital representations similar to those of the present invention.
[0139]
[0144] For purposes of the present invention, the optical systems 10, 70, 78 operate with latencies of 5 seconds or less, <1 second, <200 ms, <100 ms, or <50 ms (or the equivalent in frames or other metrics). In other contemplated embodiments of the optical systems 10, 70, 78 of the present invention, latencies may be <10 minutes, <5 minutes, <1 minute, or <30 seconds.
[0140]
[0145] The present invention is also contemplated to be suitable for projects requiring frame rates of 10 fps or greater, including >1 fps, >20 fps, about 24 fps, about 30 fps, >30 fps, about 48 fps, about 60 fps, >90 fps, about 98 or 100 fps, about 120 fps, or any other convenient frame rate, where it may be necessary to synchronize these frames with other systems or cameras.
[0141] Lattice or mesh generation
[0146] As is apparent from the above discussion, the second camera 26 of the optical system 10, 70, 78 of the present invention first creates a 2D array of distance measurements, corresponding to each pixel of the image sensor(s) used in the 3D camera. This depth map (i.e., second digital map 42) can be used in several applications, such as projecting information onto a 2D monitor or display using a virtual camera or viewpoint.
[0142]
[0147] In another use, points in a depth map (e.g., second digital map 42) can be used as vertices to create a mesh of polygons, e.g., triangles or quadrilaterals, by connecting these points. The method of connecting these points can be simple or more elaborate, in which case similar polygons are combined into larger polygons. If there are no large distance variations within a large area, this larger polygon represents the surface positions over that area.
[0143]
[0148] In some cases, the depth map (second digital map 42) may be used to calculate other vertices that are attached to a fixed grid or pattern in a spatial volume, such as a voxel grid. Such calculation methods may be based on a reliability weighting factor for distance in the depth map.
[0144]
[0149] An oTOF camera, such as second camera 26, records or provides IR intensity as a luminosity map or IR image. This comes from the same array of pixels, so each IR pixel corresponds to a depth value in the depth map. This information can be used to colorize the black-and-white values or to associate black-and-white or color values with each point or voxel. Alternatively, texture coordinates or other equivalent representations can be generated, which show the correlation between the 3D mesh and the portion of the image or texture that corresponds to its 3D surface. The texture and 3D mesh can then be rendered in appropriate software, such as a game engine.
[0145] Combination with RGB or other 2D cameras
[0150] In some cases, it may be desirable to coordinate color or other 2D images (e.g., first digital map 34) or other information with 3D data (e.g., second digital map 42) from a 3DoF camera (e.g., second camera 26). Figures 1 and 2 show examples of how the first and second cameras 24, 26 may be oriented next to each other and physically separated by a distance. As should be clear from the above discussion, the word "camera" may also refer to a separately housed and lensed system, or a module integrated within a common housing, or sensors that share an optical system in an appropriate manner, as described in other chapters.
[0146] Obscura In Fill
[0151] A problem that arises when combining images from multiple sources is parallax, as shown in Figures 1 and 2. This parallax results in a misalignment of the location of each pixel's field of view ("iFOV") between the first and second cameras 24, 26 and / or the images they generate (e.g., first digital map 34 and second digital map 42). This misalignment depends on the separation between the first and second cameras 24, 26 and the distances 62, 64, 66 of the objects 14, 16, 18 from the first and second cameras 24, 26. Another effect is that surfaces behind the foreground objects 14, 16, 18 may be obscured from the view of one or more of the first and second cameras 24, 26. For example, vertically displacing the first and second cameras 24, 26 may cause the lower camera to see an area above the foreground object that is obscured from the view of the background object.
[0147]
[0152] To reduce or eliminate the impact of these effects, different arrangements can be used.
[0148]
[0153] First, the optical axes of two imaging lenses (e.g., first optic 74 and second optic 76) are positioned so that they are collinear. A method for achieving this alignment is shown, for example, in FIG. 3. As shown in FIG. 3, the first and second cameras 24, 26 are mechanically positioned with six degrees of freedom (“DOF”). The precision of this alignment can be varied depending on the requirements of a particular application. For example, the transverse positioning can be offset by one pixel or less, up to five pixels, up to 20 pixels, up to 100 pixels, or even more. One image can be rotated by a similar amount relative to the other image. The parallelism error of the optical axes can be <1 microradian, <20 microradians, <100 microradians, <1 milliradian, or even more. The position (depth) along the optical axis can be adjusted depending on the settings of each lens. Additionally, software or mathematical coordinate transformations can be used in conjunction with mechanical positioning to improve alignment accuracy between images or between corresponding pixels (or the images derived therefrom) on each sensor or set of sensors.
[0149]
[0154] While the mechanical positioning of a stereo camera pair (as shown in Figures 1 and 2) is similar, the setup and desired alignment are quite different. With a stereo camera, the first and second cameras 24, 26 must be offset to match the distance between the viewer's eyes, and the convergence of the first and second cameras 24, 26 must be adjusted to the desired distance from the first and second cameras 24, 26. These alignments are not necessary in this case.
[0150]
[0155] As discussed in connection with FIG. 3 , the first and second cameras 24, 26 are positioned so that the IR light of the second (depth) camera 26 is reflected back to the second (depth) camera 26, while the visible light (e.g., light 32) is transmitted to a color or black-and-white camera, in this case, the first camera 24. The shutter timing of the first and second cameras 24, 26 can be synchronized to ensure correspondence between the images from each camera. The spectral characteristics of the splitter optic (e.g., fourth optic 84) can be reversed to transmit the IR light (e.g., invisible light 40) to the depth camera (second camera 26). If desired, additional cameras can be added to this combination by using appropriate optical components to split the beam into the appropriate paths.
[0151]
[0156] In an embodiment of the present invention, the invisible light illumination source 20 is positioned to be aimed at the scene 12 and is matched, possibly by beam shaping optics, with one or other of the fields of view of the first and second cameras 24, 26. Specifically, the illumination pattern may be aligned to match the field of view (“FOV”) of the second (depth) camera 26. The timing of the invisible light illumination source 20 is then synchronized with the timing of the second (depth) camera 26.
[0152]
[0157] Another approach to reducing the effects of parallax is to use three or more cameras (not shown). In an embodiment using three or more cameras, two second (depth) cameras 26 are employed and positioned on either side of the first (color or black and white) camera 24. The two second (depth) cameras 26 may be positioned symmetrically around the first (color) camera 24. In addition, the two second (depth) cameras 26 may be positioned physically close to the first (color) camera 24. For example, the second optics 76 (e.g., lens) of the second (depth) camera 26 may be separated from the first optics 74 (e.g., lens) of the first (color) camera 24 by a separation distance of 3 cm, 10 cm, 20 cm, 30 cm, more than 30 cm, more than 50 cm, more than 1 m, more than 2 m, or even greater. The axes of the lenses (first and second optics 74, 76) may be roughly parallel and focused on a defined point or area. The optical axes of the three lenses (e.g., first and second optics 74, 76) may lie in one plane or all in different planes, as desired for the particulars of use. They may also be positioned on the second (depth) camera 26, which is not symmetrical about the first (color) camera 24.
[0153]
[0158] As shown in FIGS. 1 and 2, the invisible light illumination source 20 may be provided by a single illumination source or may be the result of two or more illumination sources. When there is more than one illumination source, the illumination patterns between these illumination sources may be positioned to overlap, parallel, partially overlap, or minimally overlap. Timing and synchronization may be set so that the two second (depth) cameras 26 are approximately coincident in time, offset by a known value, or set to minimize any overlap (e.g., offset by the shutter length of the second (depth) camera 26, between 1X and 2X the shutter length, or between 2X and 4X the shutter length). Other values for shutter or illumination spacing may be used as desired.
[0154]
[0159] Software can also be used to further refine the depth resolution provided by the second (depth) cameras 26. Each second (depth) camera 26 generates a 2D depth map (second digital map 42) and other data. For example, a photogrammetric solution based on triangulation of common surface locations can be calculated separately from the unique depth map, and the two can be mathematically combined to increase the precision and accuracy of the 3D location. Alternatively, one or more other solutions can be used, such as weighting factors or guidelines to increase the overall speed of the 3D location resolution. Additionally, other techniques can be used to combine multiple measurements of the same surface so that the resulting 3D location values are more precise or accurate, or can be effectively represented by a smaller data set size.
[0155]
[0160] Other embodiments may have more than two second (depth) cameras 26 arranged as described above.
[0156]
[0161] Other contemplated variations of the optical system 10, 70, 78 of the present invention may employ more than one first (color) camera 24. There may be more first (color) cameras 24 than second (depth) cameras 26. Alternatively, the arrangement of the first (color) camera(s) 24 relative to the second (depth) camera(s) 26, as previously described, may be reversed.
[0157]
[0162] Any of these configurations may have their output used in a real-time workflow, with any software computations being performed on appropriate computer hardware (e.g., CPU, GPU, FPGA, ASIC, ISP, DSP, or other similar computing platform, or any combination of these) with a latency that may be <1 ms, <10 ms, <50 ms, <200 ms, <1 second, <5 seconds, <30 seconds), or the output may be used to save data to disk or other storage options for further use at a later time (e.g., memory 58).
[0158]
[0163] For all of these scenarios, the relative timing between the first and second cameras 24, 26 can be set in a variety of ways. Synchronizing the timing of the second (depth) camera 26 allows the use of the same lighting pattern or even common lighting. They can also be synchronized so that the lighting and camera shutters operate without overlap. The second (depth) camera 26 can be timed to operate before, during, or at the end of the other camera's shutter, or at any other time position. The timing of these cameras can also be timed to synchronize with other tracking or LED display systems, or out of sync to minimize interference as desired.
[0159] Registration
[0164] When more than one first (color) camera 24 and second (depth) camera 26 are used, a process can be employed to measure the relative position and orientation (six degrees of freedom, or 6DOF) of each camera 24, 26. This process can be performed once per configuration, or can be updated frequently or regularly based on other available information.
[0160]
[0165] The result of this process is a mathematical correlation between the pixels of the two cameras 24,26 or sensors, which correlation depends on the lens properties, the spatial orientation of the two cameras, and the distance of the cameras 24,26 to the actual surface.
[0161]
[0166] This relationship can be used to transform or map a depth map (e.g., second digital map 42) onto an equivalent grid corresponding to the pixels of another camera, such as first camera 24. Or vice versa, pixels of another camera, such as first (color) camera 24, can be mapped onto the pixels or depth map from a 3DoF camera (second (depth) camera 26). As high point density depth measurements exist over a wide range, this process becomes more robust and accurate than existing products or solutions. Also, because the 3DoF camera (e.g., second (depth) camera 26) also provides information, this process works well over a wide volumetric space, for example, >3 m wide, >5 m wide, >10 m wide, >15 m wide, >20 m wide, >30 m wide, or even wider.
[0162] Multiple cameras (volumetric capture and constrained photogrammetry)
[0167] The concept of using multiple cameras 24, 26 to obtain more information about the scene 12 or about the action within the scene 12 can ultimately be expanded to allow for the capture of all surface position information for any object 14, 16, 18 or surface in an area or volumetric space. The end product of such an endeavor is often referred to in some marketing materials as volumetric video or holographic imagery (although such results are very different from true holograms).
[0163]
[0168] However, this is difficult to achieve with prior art. For example, the film "The Matrix" used a specialized setup consisting of over 60 cameras to record images in a single plane and calculate 3D position information. The process used for "The Matrix" is commonly referred to as "bullet time." Intel built a special studio for this process. The studio was 10,000 square feet and used 100 cameras arranged around the circumference. This required supercomputer processing, and a 10-second movie clip took a week to process.
[0164]
[0169] Other major companies such as Microsoft and Canon Inc. have also built similar studios to create color volumetric video or 3D information.
[0165]
[0170] Other smaller solutions have been built and are in use, but they all use approximately 100–200 cameras and require hours or days of processing to obtain even short clips. These miniaturized sets limit the scope of any scene to one to two meters in diameter and one or perhaps two people. Attempts to improve performance using 3D depth cameras or reduce the number of cameras required have not been successful. All of today's video-oriented 3D camera solutions have range and resolution limitations, resulting in poor and unstable results even for very small single-person volumetric spaces. Volumetric video capture using 3D cameras is considered unfeasible with current solutions and techniques due to low performance or slow processing speeds, limited resolution for important features, or limited working distance constraints on the volumetric space of the stage / scene.
[0166]
[0171] The increased resolution, range, and adequate capture speed of oTOF 3D cameras, such as the cameras 24 and 26 of the present invention, now make it possible to develop practical volumetric capture solutions. The capabilities described in the immediately preceding sections can be expanded to include additional cameras within or around an area or volume to capture 3D positional information for all surfaces in that area or volume. The increased operating distance of oTOFs can make it possible to create volumetric stages or areas of, for example, 3 meters or more, or 5 meters or more, or 10 meters or more, or 50 meters or more, or even 100 meters or more. The increased resolution available with oTOFs, such as those embodied in the optical systems 10, 70, and 78 of the present invention, is also significant because the projected area of a pixel grows as the square of the distance. As distance increases, more pixels or a denser density of points are required to achieve reasonable performance for a given object size, such as an arm, leg, finger, or similarly sized non-human object.
[0167]
[0172] In implementation, the first and second cameras 24, 26 are positioned around the circumference of an area or volume. Depending on the project, the first and second cameras 24, 26 may be positioned in volumetric space to capture a specific area. For projects with a small number of objects 14, 16, 18, or with large spaces between them (separated by >5%, >10%, or >20% of the diameter or lateral distance of the volume), four, five, or six cameras 24, 26 may be positioned around the area or volume. Each camera 24, 26 may be positioned to minimize or minimize obstruction of surfaces relative to the planned location and range of movement. The cameras 24, 26 may be positioned to focus their field of view ("FOV") near a preferred plane through the volume, such as a horizontal plane at a specific height above the floor. Alternatively, the cameras 24, 26 may be positioned approximately uniformly around a partial sphere around the volume of interest. Alternatively, the cameras 24, 26 may be placed approximately at the center of the volume, or at a different radius (or equivalent) from approximately the target location, or in other locations that reduce the possibility of occlusion or allow for increased resolution or point density for segments of the volume that are of more interest to the project.
[0168]
[0173] If there are several hidden surfaces, software and even other information from the 2D camera can be used to fill in location and color information about the hidden surfaces. The color or image information (which can be black and white, or other parts of the electromagnetic spectrum) can be correlated to depth coordinates, just as with the single camera and color camera described above. This correlation can be tracked in software to correlate the combined volumetric spatial location data with at least one pixel, and possibly more, from the color image or other images.
[0169]
[0174] Software can be used to incorporate the 3D positional data and images from all cameras into a common frame of reference, select or mathematically combine (e.g., average, median, or otherwise) multiple measurements of a point or region to arrive at a single value (or 3D triplet value) for the positional information for that region. The same can be done to track correlated image information for use in later display of the image information in 3D rendering software (e.g., texture mapping the image information onto a 3D mesh generated from the positional information, or displaying 2D images associated with a particular viewpoint, or creating a composite image based on a set of correlated images that represent what would have been observable from a particular viewpoint in 3D space).
[0170]
[0175] In other projects or stages, additional cameras 24, 26 can be used to enhance the system's capabilities and capture 3D position data for objects 14, 16, 18 that are closer to one another in the scene 12, or to capture 3D position data when there are more objects 14, 16, 18 in the scene 12. For example, the spacing between objects may be <2%, <5%, <10%, or <20% of the area diameter, and the number of objects may occupy >0.1% of the total horizontal area (or an equivalent metric along any other plane through the volumetric space), or >0.5%, >1%, >2%, or >5% of that area. In all cases, the total number of cameras 24, 26 required to achieve the desired level of 3D position point density, capture volume coverage, and operating distance is far less than any current approach (e.g., any approach offered by the prior art).
[0171]
[0176] Additionally, oTOF 3D camera systems (e.g., the optical systems 10, 70, 78 of the present invention) provide dense 3D position measurements, as described in other sections. However, the placement of the cameras 24, 26 also allows for different measurements from different positions, and therefore allows for the calculation of a photogrammetric 3D position solution. This photogrammetric calculation is much faster than current photogrammetric solutions because the oTOF depth values provide a starting point. These values can also be used as weights or thresholds to speed up the photogrammetric calculation. Combining 3D position determinations improves overall system performance, increasing the stereoscopic working envelope or reducing the minimum feature size that can be accommodated by any volumetric system using oTOF cameras 24, 26.
[0172] Depth keying
[0177] It is desirable to select elements of a scene 12 and display or composite them with other separately recorded or computer-generated (CG) elements. The act of isolating or segmenting such elements can be called keying. This has traditionally been done by manually rotoscoping these elements (tracing the outline of the object(s) of interest on film, either physically or digitally) or by using a green or blue background (as shown in Figure 8). The color is then used as a key to determine the foreground or background (known as chromakey). However, this requires the extra expense of building a specialized setup or environment, as well as additional equipment and software to automate the segmentation process. It can also involve extra cost and complexity to remove green or blue hues from a color image or to match multiple colors captured at different times or in different settings. It is also difficult to do well with low latency.
[0173]
[0178] Instead of using color or other manual processes to extract elements of interest, depth (or 3D position) can be used to provide a key (a "depth key") that distinguishes between objects to keep or remove. Depth or 3D data can be used as a key in a corresponding 2D image or other data / information in multiple ways. For example, a single depth value can be used. In a 2D image, any pixels with a corresponding depth value greater than a certain value can be designated as background or made transparent or otherwise differentiated (e.g., using an alpha channel in a computer display) for later processing. Alternatively, depth or 3D position values can be compared to a plane or other geometric shape, and a keying, mask, or matte value can be assigned depending on whether the 3D position is on one side or the other of the geometric shape.
[0174]
[0179] Figure 6 illustrates this. The top image 98 shows a color image with associated depth or 3D position information (such as a mesh). The correspondence between these can be handled by UV coordinates (as is well known in the art of computer graphics). The second image 100 in Figure 6 shows the same color image, but the 3D position information has been used to compare it with two planes placed slightly in front of two walls in the color image. Behind these planes, the color pixels with 3D data are not displayed, so the walls disappear. These "foreground objects" can then be displayed within the CG environment as if they were CG objects. The third image 102 in Figure 6 shows the selected elements superimposed in front of a new background.
[0175]
[0180] The image in Figure 7 is image 104, which shows a more complex result, where depth positions and planes are used to isolate two actors and a chair and show them in a CG environment, along with a CG table and the items on the table. The real table (seen in the inset of the grayscale depth map) is isolated using a series of parallelepiped shapes to segment the table top and table legs so that they do not appear in the final image. Real cups can then be placed on the real table, and in the final output they will appear to be resting on the digital table.
[0176]
[0181] The software and hardware used to create the final product can vary depending on the needs of the project. This can be done in real time using fast solutions like games, or rendering engines like Unity or Unreal, or other custom software. The information can be transmitted over a computer network or saved and read from a computer file. This combination can be done using computer software like Nuke, Maya, or Houdini, or other similar software that manipulates or combines images or 3D information.
[0177] Background Replacement
[0182] One use of keying, which has become increasingly common in recent years with the advent of large LED walls, is to provide a controllable and steerable background during recording or streaming of images and video. However, this displayed background may need to be corrected for a variety of reasons, which may be known before it is displayed or may be discovered after the fact. It may be desirable to replace all or a portion(s) of the background image during filming or recording, or to do so subsequently.
[0178]
[0183] An interlaced green frame can be displayed to display the background and then chromakeyed to cut it out from the foreground. This requires more expensive equipment and is still prone to errors with more moving objects. This results in an expensive manual process that may require more time than is available.
[0179]
[0184] Instead of requiring an interlaced green frame, the depth key described above can be used.
[0180]
[0185] Depth keys can also be used regardless of the background or surroundings present during filming or recording. As shown in the illustrations in this document, which are screen captures of live displays, the keys can be applied in real time and with low latency. Latencies of less than 100 ms have been achieved. For various projects, achieving depth key substitution in less than 5 seconds is sufficient. Alternative latencies include <2 s, <1 s, <500 ms, <200 ms, or <100 ms. Finally, depth keys can be used to enable the use of lower-cost, coarse-pitch display segments in a background display wall, where the computer-generated (CG) background in the display virtual camera is simply the original CG data, and the LED wall primarily provides the illumination. Alternatively, with additional software functionality, the in-camera visual effects enabled by the LED wall can be entirely driven by the depth key from a 3DoF camera (e.g., second camera 26), eliminating the need for a physical display.
[0181]
[0186] The CG backgrounds available through the depth key do not have to correspond to the size of any LED wall or to the physical FOV of a real camera. For example, the background of the illustration in the previous chapter is an entirely 3D world, only a small portion of which is visible to the display virtual camera at a time. The amount of this visible portion of a large CG background can be controlled by a computer control that digitally simulates a zoom function. The background can be smaller than, similar in size to, or larger than the FOV of the physical camera.
[0182] Real-time preview ("Simulcam+")
[0187] It is often desirable to be able to see some representation that is close to the final output. Optical or digital viewfinders provide this functionality in traditional and current cameras, and the output of digital cameras can often be viewed on a smartphone, tablet, computer, or other remote viewing device. This becomes challenging when computer-generated content is mixed in some fashion with live or real-world elements, especially in real time. For example, it is difficult to display the mixed output so that a real object is behind a computer-generated (CG) element. Often, the real object must be placed in a specific position and carefully measured to determine which surfaces to display.
[0183]
[0188] In film and television production, a technique known as "simulcam," or more loosely as pre-imaging, has been developed to allow computer-generated elements such as digital set extensions and skins to be displayed on a monitor alongside live action. However, the CG elements generally must always be displayed (or always be visible) on top of the real elements. This means that many effects or elements cannot be displayed with this type of solution. It is also difficult to consistently establish a ratio between real and CG elements, because the position of the real elements relative to the CG elements is usually not known.
[0184]
[0189] Depth maps (or other 3D position information of pixel surfaces) (e.g., second digital map 42) can also be used to improve the preview solution. The 3D surface positions provide the data necessary to determine whether to display CG or real surfaces whenever there is no overlap. This position information also provides the information necessary to determine the relative and absolute proportions of CG and real elements.
[0185]
[0190] When coupled with software capable of tracking predetermined or automatically determined elements such as hands, arms, legs, feet, fingers, head, face, or the like, CG skin can be placed over real objects and follow these objects during any image or video capture. In fact, the low latency (less than a few frames after an event occurs) depth data allows essentially all visual effects to be displayed in real time or near real time. This capability can also be called Simulcam+.
[0186]
[0191] Previous depth capture solutions have been ineffective at accomplishing this, potentially leading to the conclusion that this is not a desirable approach. They have been too slow, too coarse, or only effective at very short working distances. However, the present invention uses a much denser depth map that can be captured at any point in a typical working volumetric space to achieve a viable performance level. Figure 8 shows image 106, which shows two real actors and two real props inserted directly into a CG environment, with a CG object inserted between the body and hand of one of the actors. The real object appears correctly in front of the CG elements. Shadows and other desired effects are also correctly cast from the digital lighting. Any other digital effects can also be added to this display.
[0187]
[0192] Timing and other parameters may also be set as previously described.
[0188] Relighting, color grating, and shadow generation
[0193] There are times when you may want to change either the colors, lighting, or shadows in the final output to be different from what was recorded on the real set. Because things can change or there are CG elements, it becomes important to adjust how the real objects look, or the real objects may affect the desired appearance of the CG objects. The term for this is color grading or relighting, or other similar terms.
[0189]
[0194] How lighting and reflections or shadows from other objects affect the appearance of an object varies greatly depending on the location or proximity of the object and light source. Therefore, changing lighting or shadows, or including or changing CG elements, is often a laborious manual process unless the locations of all surfaces and objects are known.
[0190]
[0195] Today's state of the art technology uses precision laser scanners to capture the positions of stationary elements in a set or scene. However, these objects often move during a project, or new objects may be added or removed. Moving objects, such as actors, balls, cars, or other objects, are, by definition, changing and cannot be captured in this process. Still images or other auxiliary 2D cameras can provide information for photogrammetric positioning or as guidance for manual processes. However, no robust solutions exist today.
[0191]
[0196] A 3D Time-of-Flight camera (e.g., second camera 26) can be used to solve this problem. Because the positional information correlated to all or nearly all visible surfaces in the image, video, or video stream is known or can be inferred by software (as discussed above), the 3D positional data can be used in real time or during any post-recording operation. For example, in Figures 9A and 9B, the first image 108 shows a daytime scene in the shadow of a CG tree, i.e., a digital tree. The second image 110 shows a nighttime scene lit by candlelight. The long range and high resolution (combined with the video rate) make it possible to apply the correct lighting in all of these situations. In these two images, the actual studio lighting was not altered in any way (and, of course, no CG elements were present).
[0192]
[0197] In this example, the candle illuminates the CG table and the real female actor because she is very close to the candle and faces the display virtual camera. The male actor is not illuminated by the candle because he is far away and there is little light on his side, making him invisible to the display virtual camera.
[0193]
[0198] Figure 8 shows shadows cast by real objects in a CG environment. These shadows move predictably as a digital light source is moved, based on ray tracing or similar calculations. These calculations require 3D position information, either as points with small areas associated with them or as a polygonal mesh. Other representations of 3D position data can also be used. Similarly, CG objects cast shadows onto real objects, and the shape of the shadows is determined by the position of the CG objects and the shape or contour of the real objects. High-resolution 3D position data makes this possible with this level of fidelity, and for large areas or scenes.
[0194]
[0199] For example, oTOF3D position data, such as that provided by the second digital map 42, also provides the information needed to create the appropriate lighting or shadow appearance when additional lighting is added to the scene. This lighting can be a standard light source, such as a light bulb, or something more exotic, such as a fairy light bulb or a shooting star, or any other light source. A video rate suitable for high-resolution 3D data allows lighting effects to be applied correctly even when CG lighting or real objects are moving.
[0195]
[0200] In other projects, creatively, you may want to change or adjust the lighting on real objects after capturing an image or video, or even live. Using the 3D data in software, you can easily change the lighting (or color) characteristics of an area so that distance-dependent effects of light appear accordingly. For example, a light in the foreground will illuminate objects closer to the light more brightly than objects further away in the background.
[0196]
[0201] This process and results can be done more quickly than current solutions due to the resolution, range, and speed of the 3DoTOF camera.
[0197] Interaction between reality and CG
[0202] Rendering software, such as game engines or similar software, provides mechanisms for checking whether different meshes collide or overlap. Combining these features with 3D position information from a 3DoF camera (e.g., a second camera 26) can create interactive effects. The low latency and high speed of the 3D data, along with its range and high resolution, are valuable for performing such effects in real time. Current solutions are severely limited to simple, well-prepared movements for specific objects and positions, and often fail despite this. It is impractical to consistently achieve such effects using current solutions.
[0198]
[0203] For example, if there is a CG element in a displayed scene, it will have an associated mesh of positional data. The 3DoT positional data of real objects, such as paddles or actors, can also be represented by a mesh of polygons or other similar representations. Each mesh is associated with a displayed image or texture by UV coordinates or similar techniques (or colored points can be used). A mesh collision detection algorithm or similar method can be used to detect when two meshes begin to overlap. Software can then be used to move the CG mesh according to a set rule (e.g., move the CG mesh away from the real object mesh at the same speed as moving the real object mesh), apply some mathematical model of physical forces and reactions, apply random direction and velocity vectors, apply an acceleration vector according to the mesh overlap length, or real object mesh velocity vector, or weight factor, or mass factor, or similar. A CG object can disappear from view, or catch fire, or perform other types of effects.
[0199]
[0204] More generally, various mathematical or physical models can be applied to govern what happens to the CG meshes, or even the appearance of real objects, when two meshes begin to overlap, or overlap according to some rule, resulting in the ability for real objects to interact with CG objects in a way that appears realistic or fantastical, as desired.
[0200] 3D effects (e.g. digital smoke)
[0205] High depth accuracy means that things like smoke and fog will obscure some objects but not others, based on optical path length. Figure 10 shows image 112, which illustrates an example where digital fog emerges from the background and obscures a woman before a man in the foreground. These effects, like the relighting described above, depend on the 3D positions of various objects and the desired effect.
[0201] Digital Focusing
[0206] Another example of an effect that can be easily achieved using refinement of 3D position data from a 3D TOF camera (e.g., second camera 26) is shown in Figures 11A and 11B. Image 114 shows the as-captured color image extracted by depth and placed in a CG world. Image 116 shows the same data, but with digital blur applied to the displayed virtual camera depending on the distance from the virtual camera. This simulates the effect of narrowing the depth of field of a lens in a high-speed camera. This is possible in this example because the oTOF camera provides the necessary distance information for every frame with low latency, a wide range (5-7 m in this example), and high resolution.
[0202]
[0207] For example, such digital blurring can be applied to objects in any type of virtual or real space. CG objects can be selected to be affected differently than real objects. Certain objects can also be selected to be affected differently from the rest (e.g., never blurred). This effect can be applied to all objects (real or CG), objects between 2m and 20m, objects between 3m and 30m, objects between 2m and 10m, objects over 10m, objects over 20m, objects between 2m and 5m, and objects between 5m and 10m.
[0203]
[0208] As discussed above, the embodiments of the present invention are merely illustrative and are not intended to limit the present invention. As should be apparent to those skilled in the art, features from one embodiment can be interchanged with other embodiments. Therefore, modifications and equivalents of the embodiments described herein are also intended to fall within the scope of the claims appended hereto.
Claims
1. 1. A system for generating a three-dimensional digital representation of a scene, comprising: an invisible light illumination source that generates invisible light and illuminates the scene with the invisible light; A first camera, the first camera is configured to receive light from the scene and generate a first digital map from the light; the first digital map comprises a plurality of first camera pixels, the first camera pixels having a first resolution of approximately 100K to 400M; each first camera pixel among the plurality of first camera pixels is associated with a color or shade; A first camera; a second camera, the second camera is configured to receive invisible light from the scene, the invisible light being generated by the invisible light illumination source, and to generate a second digital map from the invisible light; the second digital map comprises a plurality of second camera pixels, the second camera pixels having a second resolution of approximately 100K to 400M; each second camera pixel in the plurality of second camera pixels is associated with a depth; the depth is determined as a function of optical time-of-flight information for the invisible light; A second camera; the first digital map and the second digital map are generated synchronously to create a scene correlation between the first digital map and the second digital map; a processor connected to the first camera and the second camera, the processor receives a first digital map and a second digital map; the processor combines the first digital map with the second digital map to generate a three-dimensional digital representation of the scene. a processor, a three-dimensional digital representation of the scene; an image resolution of 0.9M image pixels or greater, said image resolution including the number of final pixels that make up the three-dimensional digital representation of the scene; an image latency of 10 ms to 30 seconds, the image latency including the processing time between receipt of the first digital map and the second digital map by the processor and generation of the three-dimensional digital representation of the scene; a distance between 10 cm and 200 m measured between the first and second cameras and an object in the scene; Satisfy the system.
2. 10. The system of claim 1, further comprising: A system comprising a display connected to the processor for displaying the three-dimensional representation.
3. 10. The system of claim 1, further comprising: A system comprising a memory coupled to said processor for storing a three-dimensional digital representation of said scene.
4. 10. The system of claim 1, wherein the continuous three-dimensional digital representation of the scene is assembled by the processor to generate a video having a frame rate between 5 fps and 250 fps.
5. 5. The system of claim 4, wherein for movie production, the three-dimensional digital representation of the scene comprises: The distance is between 10 cm and 50 m, and The frame rate is between 23 and 250 fps. A system that satisfies at least one of the following:
6. 3. The system of claim 2, wherein for pre-visualization (pre-visualization) of a scene for a motion picture, a three-dimensional digital representation of the scene is the image latency is between 10 ms and 10 seconds; the distance is between 10 cm and 50 m; and The frame rate is 23 to 250 fps. That also satisfies the system.
7. 5. The system of claim 4, for a live-streamed event, wherein the three-dimensional digital representation of the scene comprises: the image latency is between 10 ms and 1 second; the distance is between 2 m and 200 m; and The frame rate is 23 to 250 fps. A system that satisfies at least one of the following:
8. 10. The system of claim 1, wherein for volumetric capture, the three-dimensional digital representation of the scene comprises: the image latency is between 10 ms and 1 second; the distance is between 1 m and 50 m; and The frame rate is about 5 to 250 fps. A system that satisfies at least one of the following:
9. 10. The system of claim 1, wherein, for static capture, the three-dimensional digital representation of the scene comprises: the image latency is between 10 ms and 1 second; the distance is between 50 cm and 200 m; and The frame rate is about 5 to 250 fps. A system that satisfies at least one of the following:
10. 10. The system of claim 1, wherein the first camera and the second camera are housed within a single housing.
11. 11. The system of claim 10, wherein the one housing includes a beam splitter that directs the light to the first camera and directs the invisible light to the second camera.
12. The system of claim 1 , wherein the first camera comprises a plurality of first cameras.
13. The system of claim 1 , wherein the second camera comprises a plurality of second cameras.
14. 1. A system for generating a three-dimensional digital representation of a scene, comprising: an invisible light illumination source that generates invisible light and illuminates the scene with the invisible light; A first camera, the first camera is configured to receive light from the scene and generate a first digital map from the light; the first digital map comprises a plurality of first camera pixels, and a first resolution of the first camera pixels is between about 100K and 400M; each first camera pixel among the plurality of first camera pixels is associated with a color or shade; A first camera; a second camera, the second camera is configured to receive invisible light from the scene, the invisible light being generated by the invisible light illumination source, and to generate a second digital map from the invisible light; the second digital map comprises a plurality of second camera pixels, and a second resolution of the second camera pixels is between about 100K and 400M; each second camera pixel in the plurality of second camera pixels is associated with a depth; the depth is determined as a function of optical time-of-flight information for invisible light; A second camera; the first digital map and the second digital map are generated synchronously to create a scene correlation between the first digital map and the second digital map; a processor connected to the first camera and the second camera, the processor receives a first digital map and a second digital map; the processor combines the first digital map with the second digital map to generate a three-dimensional digital representation of the scene. a processor, a three-dimensional digital representation of the scene; an image resolution of 0.9M image pixels or greater, said image pixel count including the number of final pixels that make up the three-dimensional digital representation of said scene; a distance between 10 cm and 200 m, the distance being measured from the first and second cameras to an object in the scene; Fulfilling A continuous three-dimensional digital representation of the scene is assembled by the processor to generate video at a frame rate between 5 fps and 250 fps.
15. 15. The system of claim 14, further comprising: A system comprising a display connected to the processor for displaying the three-dimensional representation.
16. 15. The system of claim 14, further comprising: A system comprising a memory coupled to said processor for storing a three-dimensional digital representation of said scene.
17. 17. The system of claim 16, wherein for movie production, the three-dimensional digital representation of the scene comprises: the distance is between 10 cm and 50 m; and The frame rate is 23 to 250 fps. A system that satisfies at least one of the following:
18. 15. The system of claim 14, wherein for volumetric capture, the three-dimensional digital representation of the scene comprises: the distance is between 1 m and 50 m; and the frame rate is 5 to 250 fps; A system that satisfies at least one of the following:
19. 15. The system of claim 14, wherein, for static capture, the three-dimensional digital representation of the scene comprises: the distance is between 50 cm and 200 m; and the frame rate is 5 to 250 fps; A system that satisfies at least one of the following:
20. 15. The system of claim 14, wherein the first camera and the second camera are housed within a single housing.
21. 21. The system of claim 20, wherein the one housing includes a beam splitter that directs the light to the first camera and directs the invisible light to the second camera.
22. 15. The system of claim 14, wherein the first camera comprises a plurality of first cameras.
23. 15. The system of claim 14, wherein the second camera comprises a plurality of second cameras.
Citation Information
Patent Citations
Imaging apparatus and imaging method
JP2021068929A
Light-receiving element and light-receiving device
JP2021097215A
Imaging apparatus, imaging method, and information processing apparatus
JP2022146976A