System and method for capturing and generating panoramic three-dimensional images
By combining a housing, a wide-angle lens, an image capture device, and a LiDAR device, the problem of generating 3D rendering images under bright light conditions is solved, achieving efficient panoramic 3D image capture suitable for both indoor and outdoor environments.
Patent Information
- Application Number
- CN202411959198.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-30
- Filing Date
- 2020-12-30
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2040-12-30
AI Technical Summary
Existing technologies struggle to capture and generate high-quality 3D rendered images under bright lighting conditions, especially in areas with windows or bright illumination. Outdoor environments also present challenges for structured light capture, and existing technologies are time-consuming.
The device, consisting of a housing, a wide-angle lens, an image capture device, and a LiDAR device, generates panoramic 3D images by combining depth data generated by LiDAR with motor-driven rotation and exposure control.
It generates high-quality 3D rendered images under bright lighting conditions, improves capture efficiency, reduces post-processing time, and is suitable for both indoor and outdoor environments.
Smart Images

Figure CN119520961B_ABST
Abstract
Description
[0001] This application is a divisional of Chinese Patent Application 202080084506.9 (PCT / US2020 / 067474), filed December 30, 2020, entitled “System and Method for Capturing and Generating Panoramic Three-Dimensional Images.” TECHNICAL FIELD
[0002] Embodiments of the present invention generally relate to capturing and stitching panoramic images of a scene in a physical environment. BACKGROUND
[0003] The proliferation of providing three-dimensional (3D) panoramic images of the physical world has resulted in many solutions having the ability to capture multiple two-dimensional (2D) images and generate 3D images based on the captured 2D images. There exist hardware solutions and software applications (or “apps”) that are capable of capturing multiple 2D images and stitching them into a panoramic image.
[0004] There exist techniques for capturing and generating 3D data from a building. However, existing techniques generally fail to capture and generate 3D renderings of areas in bright light conditions. Areas of a floor or wall in bright light conditions or windows through which sunlight shines are often represented as holes in the 3D renderings, which can require additional post-production work to fill in. This increases the turnaround time and realism of the 3D renderings. Furthermore, outdoor environments also present challenges for many existing 3D capturing devices, as structured light can not be used to capture 3D images.
[0005] Other limitations of existing techniques for capturing and generating 3D data include the amount of time required to capture and process the digital images needed to produce 3D panoramic images. SUMMARY
[0006] An example device includes a housing, a mount configured to be coupled to a motor to move the device horizontally, a wide-angle lens coupled to the housing positioned above the mount along an axis of rotation about which the device rotates when coupled to the motor, an image capture device within the housing configured to receive a two-dimensional image of an environment through the wide-angle lens, and a LiDAR device within the housing configured to generate depth data based on the environment.
[0007] An image capture device can include a housing, a first motor, a wide-angle lens, an image sensor, a mount, a LiDAR, a second motor, and a mirror. The housing can have a front side and a back side. The first motor can be coupled to the housing at a first position between the front side and the back side of the housing, the first motor configured to turn the image capture device horizontally about a vertical axis substantially 270 degrees. The wide-angle lens can be coupled to the housing along the vertical axis at a second position between the front side and the back side of the housing, the second position being a parallax-free point, the wide-angle lens having a field of view away from the front side of the housing. The image sensor can be coupled to the housing and configured to generate image signals from light received by the wide-angle lens. The mount can be coupled to the first motor. The LiDAR can be coupled to the housing at a third position, the LiDAR configured to generate laser pulses and generate depth signals. The second motor can be coupled to the housing. The mirror can be coupled to the second motor, the second motor can be configured to rotate the mirror about a horizontal axis, the mirror including an angled surface configured to receive the laser pulses from the LiDAR and direct the laser pulses about the horizontal axis.
[0008] In some embodiments, the image sensor is configured to generate a first plurality of images at different exposures when the image capture device is stationary and pointed in a first direction. The first motor can be configured to turn the image capture device about the vertical axis after the first plurality of images are generated. In various embodiments, the image sensor does not generate images when the first motor turns the image capture device, and wherein the LiDAR generates depth signals based on the laser pulses when the first motor turns the image capture device. The image sensor can be configured to generate a second plurality of images at the different exposures when the image capture device is stationary and pointed in a second direction, and the first motor is configured to turn the image capture device about the vertical axis 90 degrees after the second plurality of images are generated. The image sensor can be configured to generate a third plurality of images at the different exposures when the image capture device is stationary and pointed in a third direction, and the first motor is configured to turn the image capture device about the vertical axis 90 degrees after the third plurality of images are generated. The image sensor can be configured to generate a fourth plurality of images at the different exposures when the image capture device is stationary and pointed in a fourth direction, and the first motor is configured to turn the image capture device about the vertical axis 90 degrees after the fourth plurality of images are generated.
[0009] In some embodiments, the system can also include a processor configured to blend frames of the first plurality of images prior to the image sensor generating the second plurality of images. A remote digital device can be in communication with the image capture device and configured to generate a 3D visualization based on the first, second, third, and fourth plurality of images and the depth signal, the remote digital device configured to generate the 3D visualization using no more than an image of the first, second, third, and fourth plurality of images. In some embodiments, the first, second, third, and fourth plurality of images are generated between turns that combine turns that turn the image capture device 270 degrees around the vertical axis. A speed or rotation of the mirror around the horizontal axis increases as the first motor turns the image capture device. The angled surface of the mirror can be 90 degrees. In some embodiments, the LiDAR emits the laser pulses in a direction opposite the front side of the housing.
[0010] An example method includes receiving light from a wide-angle lens of an image capture device, the wide-angle lens coupled to a housing of the image capture device, the light received at a field of view of the wide-angle lens that extends away from a front side of the housing, generating a first plurality of images by an image sensor of the image capture device using the light from the wide-angle lens, the image sensor coupled to the housing, the first plurality of images at different exposures, turning the image capture device horizontally around a vertical axis substantially 270 degrees by a first motor, the first motor coupled to the housing in a first position between the front side and a back side of the housing, the wide-angle lens at a second position along the vertical axis, the second position being a parallax-free point, rotating a mirror having an angled surface around a horizontal axis by a second motor, the second motor coupled to the housing, generating laser pulses by a LiDAR, the LiDAR coupled to the housing at a third position, the laser pulses directed to the rotating mirror as the image capture device is turned horizontally, and generating a depth signal by the LiDAR based on the laser pulses.
[0011] Generating the first plurality of images by the image sensor can occur prior to the image capture device being turned horizontally. In some embodiments, the image sensor does not generate images as the first motor turns the image capture device, and wherein the LiDAR generates the depth signal based on the laser pulses as the first motor turns the image capture device.
[0012] The method can also include generating, by the image sensor, a second plurality of images at the different exposure when the image capture device is stationary and pointed in a second direction, and after generating the second plurality of images, rotating the image capture device 90 degrees about the vertical axis by the first motor.
[0013] In some embodiments, the method can also include generating, by the image sensor, a third plurality of images at the different exposure when the image capture device is stationary and pointed in a third direction, and after generating the third plurality of images, rotating the image capture device 90 degrees about the vertical axis by the first motor. The method can also include generating, by the image sensor, a fourth plurality of images at the different exposure when the image capture device is stationary and pointed in a fourth direction. The method can include generating a 3D visualization using the first, second, third, and fourth plurality of images and based on the depth signal, the generating the 3D visualization not using any other images.
[0014] In some embodiments, the method can also include blending frames of the first plurality of images prior to the image sensor generating the second plurality of images. The first, second, third, and fourth plurality of images can be generated between rotations that combine to rotate the image capture device 270 degrees about the vertical axis. In some embodiments, the speed or rotation of the mirror about the horizontal axis increases as the first motor rotates the image capture device. BRIEF DESCRIPTION OF DRAWINGS
[0015] FIG. 1a depicts a toy house view of an example environment, such as a house, in accordance with some embodiments.
[0016] FIG. 1b depicts a floor plan view of a first floor of a house in accordance with some embodiments.
[0017] Figure 2 depicts an example eye level view of a living room, which can be part of a virtual walkthrough.
[0018] Figure 3 depicts one example of an environment capture system in accordance with some embodiments.
[0019] Figure 4 depicts a perspective view of an environment capture system in some embodiments.
[0020] Figure 5 is a depiction of laser pulses from a LiDAR around an environment capture system in some embodiments.
[0021] Figure 6A depicts a side view of an environment capture system.
[0022] Figure 6B A view from above the environment capture system in some embodiments is depicted.
[0023] Figure 7 A perspective view of components of one example of an environment capture system according to some embodiments is depicted.
[0024] Figure 8A Example lens dimensions in some embodiments are depicted.
[0025] Figure 8B Example lens design specifications in some embodiments are depicted.
[0026] FIG. 9A depicts a block diagram of an example of an environment capture system according to some embodiments.
[0027] FIG. 9B depicts a block diagram of an example SOM PCBA of an environment capture system according to some embodiments.
[0028] Figures 10a-10c A process for taking images of an environment capture system in some embodiments is depicted.
[0029] Figure 11 A block diagram of an example environment capable of capturing and stitching images to form a 3D visualization according to some embodiments is depicted.
[0030] Figure 12 A block diagram of an example of an alignment and stitching system according to some embodiments is depicted.
[0031] Figure 13 A flowchart of a 3D panoramic image capture and generation process according to some embodiments is depicted.
[0032] Figure 14 A flowchart of a 3D and panoramic capture and stitching process according to some embodiments is depicted.
[0033] Figure 15 A flowchart depicting further details of one step of the 3D and panoramic capture and stitching process of Figure 14 is depicted.
[0034] Figure 16 A block diagram of an example digital device according to some embodiments is depicted. DETAILED DESCRIPTION
[0035] Many of the innovations described herein are made with reference to the drawings. Like numerals are used to refer to like elements throughout. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding. It can be apparent, however, that the different innovations can be practiced in different embodiments without all these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to facilitate describing the innovations.
[0036] Various embodiments of the device provide users with 3D panoramic images of indoor as well as outdoor environments. In some embodiments, the device can efficiently and quickly provide users with 3D panoramic images of indoor and outdoor environments using a single wide field of view (FOV) lens and a single light detection and ranging sensor (LiDAR sensor).
[0037] The following is an example use case of the example device described herein. The following use case is one of the embodiments. Different embodiments of the device as discussed herein can include one or more features and capabilities similar to those of this use case.
[0038] Figure la depicts a dollhouse view 100 of an example environment (e.g., a house) in accordance with some embodiments. The dollhouse view 100 gives an overall view of the example environment captured by the environment capture system (discussed herein). A user can interact with the dollhouse view 100 on the user system by switching between different views of the example environment. For example, the user can interact with the region 110 to trigger a floor plan view of the first floor of the house, as shown in Figure lb. In some embodiments, the user can interact with icons (such as icons 120, 130, and 140) in the dollhouse view 100 to provide a walkthrough view (e.g., for 3D walkthrough), a floor plan view, or a measurement view, respectively.
[0039] Figure lb depicts a floor plan view of the first floor of the house in accordance with some embodiments. The floor plan view is a top view of the first floor of the house. The user can interact with a region (such as region 150) of the floor plan view to trigger an eye level view of a particular portion (such as the living room) of the floor plan. An example of the eye level view of the living room can be found in Figure 2 which can be part of a virtual walkthrough.
[0040] A user can interact with a portion of the floor plan 200 corresponding to the area 150 of FIG. lb. The user can move the view around the room as if the user were actually in the living room. In addition to the horizontal 360° view of the living room, the user can also view or navigate the floor or ceiling of the living room. Furthermore, the user can move through the living room to other portions of the house by interacting with particular areas of the portion of the floor plan 200, such as areas 210 and 220. When the user interacts with area 220, the environment capture system can provide a walkthrough transition between the area of the house substantially corresponding to the portion of the house depicted by area 150 to the area of the house substantially corresponding to the portion of the house depicted by area 220.
[0041] Figure 3 One example of an environment capture system 300 is depicted in accordance with some embodiments. The environment capture system 300 includes a lens 310, a housing 320, a mounting attachment 330, and a movable cover 340.
[0042] When in use, the environment capture system 300 can be positioned in an environment, such as a room. The environment capture system 300 can be positioned on a support (e.g., a tripod). The movable cover 340 can be moved to reveal the LiDAR and the rotatable mirror. Once activated, the environment capture system 300 can take a burst of images and then rotate using the motor. The environment capture system 300 can open the mounting attachment 330. While rotating, the LiDAR can take measurements (the environment capture system can not take images while rotating). Once pointed in a new direction, the environment capture system can take another burst of images before rotating to the next direction.
[0043] For example, once positioned, a user can command the environment capture system 300 to begin a sweep. The sweep can be as follows:
[0044] (1) Exposure estimation, and then take HDR RGB images
[0045] Rotate 90 degrees to capture depth data
[0046] (2) Exposure estimation, and then take HDR RGB images
[0047] Rotate 90 degrees to capture depth data
[0048] (3) Exposure estimation, and then take HDR RGB images
[0049] Rotate 90 degrees to capture depth data
[0050] (4) Exposure estimation, and then take HDR RGB images
[0051] Capture depth data 90 degrees of rotation (360 degrees total)
[0052] For each burst, there can be any number of images at different exposures. The environment capture system can blend together any number of images of a burst while waiting for another frame and / or waiting for the next burst.
[0053] The housing 320 can protect the electronic components of the environment capture system 300 and can provide an interface for user interaction with power buttons, scan buttons, etc. For example, the housing 320 can include a movable cover 340, which can be movable to expose a LiDAR. Further, the housing 320 can include electronic interfaces such as a power adapter and indicator lights. In some embodiments, the housing 320 is a molded plastic housing. In various embodiments, the housing 320 is a combination of one or more of plastic, metal, and polymer.
[0054] The lens 310 can be part of a lens assembly. Further details of the lens assembly can be described in the description of Figure 7 The lens 310 is strategically placed at the center of the rotation axis 305 of the environment capture system 300. In this example, the rotation axis 305 is on the x-y plane. By placing the lens 310 at the center of the rotation axis 305, parallax effects can be eliminated or reduced. Parallax is an error that results from the rotation of an image capture device around a point that is not the no-parallax point (NPP). In this example, the NPP can be found at the center of the entrance pupil of the lens.
[0055] For example, assume that a panoramic image of a physical environment is generated based on four images captured by the environment capture system 300, with 25% overlap between the images of the panoramic image. If there were no parallax, then 25% of one image can overlap completely with another image of the same area of the physical environment. Eliminating or reducing parallax effects of multiple images captured by the image sensor through the lens 310 can facilitate stitching the multiple images into a 2D panoramic image.
[0056] The lens 310 can include a large field of view (e.g., the lens 310 can be a fisheye lens). In some embodiments, the lens can have a horizontal FOV (HFOV) of at least 148 degrees and a vertical FOV (VFOV) of at least 94 degrees.
[0057] The mounting attachment 330 can allow the environment capture system 300 to be attached to a mount. The mount can allow the environment capture system 300 to be coupled with a tripod, a flat surface, or a motorized mount (e.g., to move the environment capture system 300). In some embodiments, the mount can allow the environment capture system 300 to be rotated along a horizontal axis.
[0058] In some embodiments, the environment capture system 300 can include a motor for rotating the environment capture system 300 horizontally about the mounting attachment 330.
[0059] In some embodiments, the motorized mount can move the environment capture system 300 along a horizontal axis, a vertical axis, or both. In some embodiments, the motorized mount can rotate or move in an x-y plane. The use of the mounting attachment 330 can allow the environment capture system 300 to be coupled to a motorized mount, a tripod, or the like, to stabilize the environment capture system 300 to reduce or minimize shaking. In another example, the mounting attachment 330 can be coupled to a motorized mount that allows the 3D and environment capture system 300 to rotate at a known speed that is stable, which helps the LiDAR determine the (x, y, z) coordinates of each laser pulse of the LiDAR.
[0060] Figure 4 Renderings of the environment capture system 400 in some embodiments are depicted. The renderings show the environment capture system 400 (which can be an example of the environment capture system 300 of the first example) from various views, such as a front view 410, a top view 420, a side view 430, and a back view 440. In these renderings, the environment capture system 400 can include an optional hollowed-out portion, depicted in the side view 430. Figure 3
[0061] In some embodiments, the environment capture system 400 has a width of 75 mm, a height of 180 mm, and a depth of 189 mm. It should be understood that the environment capture system 400 can have any width, height, or depth. In various embodiments, the ratio of the width to the height to the depth in the first example is maintained regardless of the specific measurements.
[0062] The housing of the 3D and environment capture system 400 can protect the electronic components of the environment capture system 400 and can provide an interface for user interaction (e.g., a screen on the back view 440). In addition, the housing can include electronic interfaces, such as a power adapter and indicator lights. In some embodiments, the housing is a molded plastic housing. In various embodiments, the housing is a combination of one or more of plastic, metal, and polymer. The environment capture system 400 can include a removable cover that can be removable to expose the LiDAR and protect the LiDAR from elements when not in use.
[0063] The lens depicted on the front view 410 can be part of a lens assembly. Similar to the environment capture system 300, the lens of the environment capture system 400 is strategically placed at the center of the axis of rotation. The lens can include a large field of view. In various embodiments, the lens depicted on the front view 410 is concave and the housing is flared such that the wide angle lens is directly at the no-parallax point (e.g., directly above the midpoint of the mount and / or motor), but can still take images without interference from the housing.
[0064] The mount attachment at the base of the environment capture system 400 can allow the environment capture system to be attached to a mount. The mount can allow the environment capture system 400 to be coupled with a tripod, a flat surface, or a motorized mount (e.g., to move the environment capture system 400). In some embodiments, the mount can be coupled to an internal motor for rotating the environment capture system 400 about the mount.
[0065] In some embodiments, the mount can allow the environment capture system 400 to rotate along a horizontal axis. In various embodiments, the motorized mount can move the environment capture system 400 along a horizontal axis, a vertical axis, or both. The use of a mount attachment can allow the environment capture system 400 to be coupled to a motorized mount, a tripod, etc., to stabilize the environment capture system 400 to reduce or minimize shaking. In another example, the mount attachment can be coupled to a motorized mount that allows the environment capture system 400 to rotate at a known, stable speed, which helps the LiDAR determine the (x, y, z) coordinates of each laser pulse of the LiDAR.
[0066] In the view 430, the mirror 450 is revealed. The LiDAR can emit laser pulses (in a direction opposite to the lens view) to the mirror. The laser pulses can hit the mirror 450, which can be angled (e.g., at a 90 degree angle). The mirror 450 can be coupled to an internal motor that rotates the mirror, such that the laser pulses of the LiDAR can be emitted and / or received at many different angles around the environment capture system 400.
[0067] Figure 5 is a depiction of laser pulses from the LiDAR around the environment capture system 400 in some embodiments. In this example, the laser pulses are emitted at the rotating mirror 450. The laser pulses can be emitted and received perpendicular to the horizontal axis 602 (see FIG. 6) of the environment capture system 400. The mirror 450 can be angled such that the laser pulses from the LiDAR are directed away from the environment capture system 400. In some examples, the angle of the angled surface of the mirror can be 90 degrees, or at or between 60 degrees to 120 degrees.
[0068] In some embodiments, when the environment capture system 400 is stationary and in operation, the environment capture system 400 can take a burst of images through the lens. The environment capture system 400 can turn on the horizontal motor between the burst of images. As the mount is turned, the LiDAR of the environment capture system 400 can emit and / or receive laser pulses that hit the rotating mirror 450. The LiDAR can generate depth signals and / or generate depth data from the received laser pulse reflections.
[0069] In some embodiments, the depth data can be associated with coordinates with respect to the environment capture system 400. Similarly, pixels or portions of the images can be associated with coordinates with respect to the environment capture system 400 to enable the generation of 3D visualizations (e.g., images from different perspectives, 3D walkthroughs, etc.) using the images and depth data.
[0070] As shown, LiDAR pulses can be blocked by the bottom of the environment capture system 400. It should be understood that the mirror 450 can rotate consistently while the environment capture system 400 is moving around the mount, or the mirror 450 can rotate more slowly when the environment capture system 400 starts moving and in addition when the environment capture system 400 slows to a stop (e.g., maintaining a constant speed between the start and stop of the mount motor). Figure 5
[0071] The LiDAR can receive depth data from the pulses. Due to the movement of the environment capture system 400 and / or the increase or decrease in speed of the mirror 450, the density of the depth data with respect to the environment capture system 400 can be inconsistent (e.g., more dense in some areas and less dense in other areas).
[0072] FIG. 6a depicts a side view of the environment capture system 400. In this view, the mirror 450 is depicted, and the mirror 450 can rotate around a horizontal axis. Pulses 604 can be emitted by the LiDAR at the rotating mirror 450, and can be emitted perpendicular to the horizontal axis 602. Similarly, the pulses 604 can be received by the LiDAR in a similar manner.
[0073] Although the LiDAR pulses are discussed as being perpendicular to the horizontal axis 602, it should be understood that the LiDAR pulses can be at any angle with respect to the horizontal axis 602 (e.g., the mirror angle can be at any angle including between 60 degrees to 120 degrees). In various embodiments, the LiDAR emits pulses opposite the front side (e.g., front side 604) of the environment capture system 400 (e.g., in a direction opposite the center of the field of view of the lens or towards the back side 606).
[0074] As discussed herein, the environment capture system 400 can be turned about the vertical axis 608. In various embodiments, the environment capture system 400 takes images and then turns 90 degrees, taking a fourth set of images when the environment capture system 400 has completed a 270 degree turn from the original starting position of taking the first set of images. Thus, the environment capture system 400 can generate four sets of images between a total of 270 degrees of turns (e.g., assuming the first set of images was taken prior to the initial turn of the environment capture system 400). In various embodiments, the images from a single sweep of the environment capture system 400 (e.g., four sets of images) (e.g., taken in a single full rotation or 270 degree rotation about the vertical axis) are sufficient to generate a 3D visualization along with depth data acquired during the same sweep without any additional sweeps or turns of the environment capture system 400.
[0075] It should be appreciated that in this example, the LiDAR pulses are emitted and directed by a rotating mirror in a position away from the point of rotation of the environment capture system 400. In this example, the distance from the point of rotation of the mount is 608 (e.g., the lens can be at the no-parallax point, while the lens can be in a position behind the lens relative to the front of the environment capture system 400). Because the LiDAR pulses are directed by the mirror 450 at a position offset from the point of rotation, the LiDAR cannot receive depth data from a cylinder that extends from above the environment capture system 400 to below the environment capture system 400. In this example, the radius of the cylinder (e.g., the cylinder without depth information) can be measured from the center of the point of rotation of the motor mount to the point at which the LiDAR pulses are directed by the mirror 450.
[0076] Further, in FIG. 6b, a cavity 610 is depicted. In this example, the environment capture system 400 includes a rotating mirror within the body of the housing of the environment capture system 400. There is a cutout section from the housing. Laser pulses can be reflected out of the housing by the mirror, and then the reflections can be received by the mirror and directed back to the LiDAR to enable the LiDAR to produce depth signals and / or depth data. The base of the body of the environment capture system 400 below the cavity 610 can block some laser pulses. The cavity 610 can be defined by the base of the environment capture system 400 and the rotating mirror. As Figure 6B As shown, there can still be space between the edge of the angled mirror and the housing of the environment capture system 400 containing the LiDAR.
[0077] In various embodiments, the LiDAR is configured to stop emitting laser pulses if the rotational speed of the mirror falls below a rotational safety threshold (e.g., if there is a failure of the motor that causes the mirror to rotate or the mirror is held in place). In this way, the LiDAR can be configured for safety and to reduce the likelihood that laser pulses will continue to be emitted in the same direction (e.g., at a user’s eye).
[0078] Figure 6b depicts a view from above the environment capture system 400 in some embodiments. In this example, the front of the environment capture system 400 is depicted, with the lens being concave and directly above the center of the rotation point (e.g., above the center of the mount). The front of the camera is concave for the lens, and the front of the housing is flared to allow the field of view of the image sensor to not be blocked by the housing. The mirror 450 is depicted as pointing upward.
[0079] Figure 7 A perspective view of components of one example of an environment capture system 300 according to some embodiments is depicted. The environment capture system 700 includes a front cover 702, a lens assembly 704, a structural frame 706, a LiDAR 708, a front housing 710, a mirror assembly 712, a GPS antenna 714, a rear housing 716, a vertical motor 718, a display 720, a battery pack 722, a mount 724, and a horizontal motor 726.
[0080] In various embodiments, the environment capture system 700 can be configured to scan, align, and produce 3D meshes outdoors in full sunlight as well as indoors. This removes the barrier to adoption of other systems that are only indoor tools. The environment capture system 700 can be able to scan large spaces faster than other devices. In some embodiments, the environment capture system 700 can provide improved depth accuracy by improving single-scan depth accuracy at 90m.
[0081] In some embodiments, the environment capture system 700 can weigh 1 kg or about 1 kg. In one example, the environment capture system 700 can weigh 1-3 kg.
[0082] The front cover 702, the front housing 710, and the rear housing 716 make up part of the housing. In one example, the front cover can have a width w of 75 mm.
[0083] The lens assembly 704 can include a camera lens that focuses light onto an image capture device. The image capture device can capture images of a physical environment. A user can place the environment capture system 700 to capture a portion of a floor of a building, such as the second building 422 of FIG. 1, to obtain a panoramic image of the portion of the floor. The environment capture system 700 can be moved to another portion of the floor of the building to obtain a panoramic image of the other portion of the floor. In one example, the depth of field of the image capture device is 0.5 meters to infinity. Figure 8A Example lens dimensions in some embodiments are depicted.
[0084] In some embodiments, the image capture device is a complementary metal-oxide-semiconductor (CMOS) image sensor (e.g., a Sony IMX283 ~ 20 Megapixel CMOS MIPI sensor with NVidia Jetson Nano SOM). In various embodiments, the image capture device is a charge-coupled device (CCD). In one example, the image capture device is a red-green-blue (RGB) sensor. In one embodiment, the image capture device is an infrared (IR) sensor. The lens assembly 704 can give the image capture device a wide field of view.
[0085] Image sensors can have many different specifications. In one example, the image sensor includes the following:
[0086] Number of pixels per column Pixels 5496 Number of pixels per row Pixels 3694 Resolution MP >20 Image circle diameter mm 15.86 mm Pixel pitch um 2.4 um Pixels per degree (PPD) PPD >37 Chief ray angle at full height of sensor Degrees 3.0° Output interface - MIPI Green sensitivity V / lux*s >1.7 SNR (100 lux, 1x gain) dB >65 Dynamic range dB >70
[0087] Example specifications can be as follows:
[0088] Image circle diameter mm 15.86 Minimum object distance mm 500 Maximum object distance mm Infinity Chief ray angle at full height of sensor Degrees 3.0 L1 diameter mm <60 Lens total length (TTL) mm <=80 Back focal length (BFL) mm - Effective focal length (EFL) mm - Relative illumination % >50 Maximum distortion % <5 52 lp / mm (on axis) % >85 104 lp / mm (on axis) % >66 208 lp / mm (on axis) % >45 52 lp / mm (83% field) % >75 104 lp / mm (83% field) % >41 208 lp / mm (83% field) % >25
[0089] In various embodiments, when observing the MTF at the relative field of view F0 (i.e., center), the focus shift can vary from +28 microns at 0.5m to -25 microns at infinity for a total through focus shift of 53 microns.
[0090] Figure 8B Example lens design specifications in some embodiments are depicted.
[0091] In some examples, the lens assembly 704 has an HFOV of at least 148 degrees and a VFOV of at least 94 degrees. In one example, the lens assembly 704 has a field of view of 150°, 180°, or in the range of 145° to 180°. In one example, image capture of a 360° view around the environment capture system 700 can be obtained with three or four separate image captures from the image capture devices of the environment capture system 700. In various embodiments, the image capture devices can have a resolution of at least 37 pixels per degree. In some embodiments, the environment capture system 700 includes a lens cap (not shown) to protect the lens assembly 704 when not in use. The output of the lens assembly 704 can be a digital image of a region of the physical environment. The images captured by the lens assembly 704 can be stitched together to form a 2D panoramic image of the physical environment. A 3D panorama can be generated by combining depth data captured by the LiDAR 708 with the 2D panoramic image generated by stitching together multiple images from the lens assembly 704. In some embodiments, the images captured by the environment capture system 402 are stitched together by the image processing system 406. In various embodiments, the environment capture system 402 generates a "preview" or "thumbnail" version of the 2D panoramic image. The preview or thumbnail version of the 2D panoramic image can be rendered on a user system 1110 such as an iPad, a personal computer, a smartphone, etc. In some embodiments, the environment capture system 402 can generate a mini map of the physical environment representing a region of the physical environment. In various embodiments, the image processing system 406 generates a mini map representing the region of the physical environment.
[0092] The images captured by the lens assembly 704 can include capture device location data identifying or indicating a capture location of the 2D images. For example, in some implementations, the capture device location data can include global positioning system (GPS) coordinates associated with the 2D images. In other implementations, the capture device location data can include location information indicating a relative location of the capture device (e.g., a camera and / or 3D sensor) relative to its environment, such as a relative or calibrated location of the capture device relative to an object in the environment, another camera in the environment, another device in the environment, etc. In some implementations, this type of location data can be determined by the capture device (e.g., a camera and / or a device including positioning hardware and / or software operably coupled to the camera) associated with the capture of the image and received with the image. The placement of the lens assembly 704 is not merely by design. By placing the lens assembly 704 at or substantially at the center of the axis of rotation, parallax effects can be reduced.
[0093] In some embodiments, the structural frame 706 holds the lens assembly 704 and the LiDAR 708 in a particular position and can help protect the components of an example of the environment capture system. The structural frame 706 can be used to help rigidly mount the LiDAR 708 and place the LiDAR 708 in a fixed position. Further, the fixed position of the lens assembly 704 and the LiDAR 708 enables a fixed relationship to align depth data with image information to help produce 3D images. The 2D image data and depth data captured in a physical environment can be aligned with respect to a common 3D coordinate space to generate a 3D model of the physical environment.
[0094] In various embodiments, the LiDAR 708 captures depth information of a physical environment. When a user places the environment capture system 700 in a portion of a floor of a second building, the LiDAR 708 can obtain depth information of objects. The LiDAR 708 can include an optical sensing module that is able to measure the distance to a target or object in a scene by illuminating the target or scene with a pulse from a laser and measuring the time it takes for a photon to travel to the target and back to the LiDAR 708. The measurements can then be transformed into a grid coordinate system by using information derived from the horizontal drive train of the environment capture system 700.
[0095] In some embodiments, the LiDAR 708 can return a depth data point with a timestamp (of an internal clock) every 10 microseconds. The LiDAR 708 can sample a partial sphere (with small holes at the top and bottom) every 0.25 degrees. In some embodiments, with one data point every 10 microseconds and 0.25 degrees, there can be 14.40 milliseconds and 1440 disks per point “disk” to form a nominal 20.7 second sphere. Because each disk is captured front and back, the sphere can be captured with a 180° scan.
[0096] In one example, the LiDAR 708 specifications can be as follows:
[0097]
[0098]
[0099] One advantage of utilizing a LiDAR is that in the case of a LiDAR at a lower wavelength (e.g., 905 nm, 900-940 nm, etc.), it can allow the environment capture system 700 to determine depth information of outdoor environments or indoor environments with bright lights.
[0100] The placement of the lens assembly 704 and the LiDAR 708 can allow the environment capture system 700 or a digital device in communication with the environment capture system 700 to use depth data from the LiDAR 708 and the lens assembly 704 to generate 3D panoramic images. In some embodiments, 2D and 3D panoramic images are not generated on the environment capture system 402.
[0101] The output of the LiDAR 708 can include attributes associated with each laser pulse sent by the LiDAR 708. The attributes include intensity of the laser pulse, number of returns, current number of returns, classification point, RGC value, GPS time, scan angle, scan direction, or any combination thereof. The depth of field can be (0.5m; infinity), (1m; infinity), etc. In some embodiments, the depth of field is 0.2m to 1m and infinity.
[0102] In some embodiments, the environment capture system 700 uses the lens assembly 704 to capture four separate RBG images while the environment capture system 700 is stationary. In various embodiments, the LiDAR 708 captures depth data in four different instances while the environment capture system 700 is in motion as it moves from one RBG image capture position to another. In one example, a 3D panoramic image is captured with a 360° rotation of the environment capture system 700 (this can be referred to as a sweep). In various embodiments, a 3D panoramic image is captured with less than a 360° rotation of the environment capture system 700. The output of a sweep can be a sweep list (SWL) that includes image data from the lens assembly 704 and depth data from the LiDAR 708 as well as the nature of the sweep, including GPS location and a timestamp of when the sweep occurred. In various embodiments, a single sweep (e.g., a single 360 degree turn of the environment capture system 700) captures enough image and depth information to generate a 3D visualization (e.g., by a digital device in communication with the environment capture system 700 that receives image and depth data from the environment capture system 700 and produces a 3D visualization using only image and depth data from the environment capture system 700 captured in a single sweep).
[0103] In some embodiments, images captured by the environment capture system 402 can be blended, stitched together, and combined with depth data from the LiDAR 708 by the image stitching and processing systems discussed herein.
[0104] In various embodiments, the environment capture system 402 and / or an application on the user system 1110 can generate a preview or thumbnail version of the 3D panoramic image. The preview or thumbnail version of the 3D panoramic image can be rendered on the user system 1110 and can have a lower image resolution than the 3D panoramic image generated by the image processing system 406. After the lens assembly 704 and the LiDAR 708 capture image and depth data of a physical environment, the environment capture system 402 can generate a mini map representing an area of the physical environment that has been captured by the environment capture system 402. In some embodiments, the image processing system 406 generates a mini map representing the area of the physical environment. After capturing image and depth data of a living room of a home using the environment capture system 402, the environment capture system 402 can generate an overhead view of the physical environment. The user can use this information to determine areas of the physical environment that the user has not yet captured or generated a 3D panoramic image.
[0105] In one embodiment, the environment capture system 700 can interleave image capture with the image capture device utilizing the lens assembly 704 and depth information capture with the LiDAR 708. For example, the image capture device can capture an image of a segment 1605 (as seen in FIG. 16A) of the physical environment with the image capture device and then the LiDAR 708 obtains depth information from the segment 1605. Once the LiDAR 708 obtains the depth information from the segment 1605, the image capture device can continue to move to capture an image of another segment 1610 and then the LiDAR 708 obtains depth information from the segment 1610, interleaving image capture and depth information capture. Figure 16
[0106] In some embodiments, the LiDAR 708 can have a field of view of at least 145°, and depth information for all objects in a 360° view of the environment capture system 700 can be obtained by the environment capture system 700 in three or four scans. In another example, the LiDAR 708 can have a field of view of at least 150°, 180°, or between 145° and 180°.
[0107] The increase in the field of view of the lens reduces the amount of time required to obtain visual and depth information of the physical environment surrounding the environment capture system 700. In various embodiments, the LiDAR 708 has a minimum depth range of 0.5m. In one embodiment, the LiDAR 708 has a maximum depth range of greater than 8 meters.
[0108] LiDAR 708 can utilize a mirror assembly 712 to direct the laser at different scan angles. In one embodiment, an optional vertical motor 718 has the ability to move the mirror assembly 712 vertically. In some embodiments, the mirror assembly 712 can be a dielectric mirror with a hydrophobic coating or layer. The mirror assembly 712 can be coupled to the vertical motor 718, which rotates the mirror assembly 712 when in use.
[0109] The mirror of the mirror assembly 712 may, for example, include the following specifications for materials and coatings:
[0110]
[0111] The mirror of the mirror assembly 712 may, for example, include the following specifications for materials and coatings:
[0112] S1L1 Material Dielectric S1L2 Material Hydrophobic S2L1 Material Black paint emulsion Powder suspended in paint Substrate Material Schott B270I
[0113] The hydrophobic coating of the mirror of the mirror assembly 712 may, for example, include a contact angle number of > 105.
[0114] The mirror of the mirror assembly 712 can include the following quality specifications:
[0115]
[0116]
[0117] The vertical motor may, for example, include the following specifications:
[0118]
[0119] Due to the RGB capture device and the LiDAR 708, the environment capture system 700 can capture images outdoors in bright sunlight, or indoors in the presence of bright lights or glare from windows. In systems that utilize different devices (e.g., structured light devices), they can not be able to operate in bright environments, either indoors or outdoors. These devices are typically limited to use only indoors and only during dawn or dusk to control the light. Otherwise, bright spots in the room create artifacts or “holes” in the image that must be filled or corrected. However, the environment capture system 700 can be used indoors and outdoors in bright sunlight. The capture device and the LiDAR 708 can be able to capture images and depth data in bright environments without artifacts or holes caused by glare or bright lights.
[0120] In one embodiment, the GPS antenna 714 receives global positioning system (GPS) data. The GPS data can be used to determine the location of the environment capture system 700 at any given time.
[0121] In various embodiments, the display 720 allows the environment capture system 700 to provide the current status of the system, such as updates, warm-up, scanning, scan complete, errors, etc.
[0122] The battery pack 722 provides power to the environment capture system 700. The battery pack 722 can be removable and rechargeable, allowing a user to put in a new battery pack 722 while the depleted battery pack is being charged. In some embodiments, the battery pack 722 can allow for at least 1000 SWL or at least 250 SWL of continuous use before recharging. The environment capture system 700 can utilize a USB-C plug for recharging.
[0123] In some embodiments, the mount 724 provides a connector for the environment capture system 700 to connect to a platform such as a tripod or mount. The horizontal motor 726 can rotate the environment capture system 700 about the x-y plane. In some embodiments, the horizontal motor 726 can provide information to a grid coordinate system to determine the (x, y, z) coordinates associated with each laser pulse. In various embodiments, the horizontal motor 726 can enable the environment capture system 700 to scan quickly due to the wide field of view of the lens, the positioning of the lens about the axis of rotation, and the LiDAR device.
[0124] In one example, the horizontal motor 726 can have the following specifications:
[0125]
[0126] In various embodiments, the mount 724 can include a quick release adapter. The holding torque can be, for example, > 2.0 Nm, and the durability of the capture operation can be up to or over 70,000 cycles.
[0127] For example, the environment capture system 700 can enable the construction of a 3D grid of a standard home, where the distance between sweeps is greater than 8m. The time to capture, process, and align indoor sweeps can be under 45 seconds. In one example, the time frame from the start of sweep capture to when the user can move the environment capture system 700 can be less than 15 seconds.
[0128] In various embodiments, these components provide the ability for the environment capture system 700 to align scan locations both outdoors and indoors and thus produce a seamless roaming experience between indoors and outdoors (which can be a high priority for hotels, vacation rentals, real estate, architectural documentation, CRE, and as-built modeling and verification). The environment capture system 700 can also produce an "outdoor toy box" or outdoor mini map. As shown herein, the environment capture system 700 can also improve the accuracy of 3D reconstruction primarily from a measurement perspective. The ability for a user to tune the scan density can also be a benefit. These components can also enable the environment capture system 700 to capture wide open spaces (e.g., longer ranges). To generate a 3D model of a wide open space, the environment capture system can need to scan and capture 3D data and depth data from a larger range of distances than to generate a 3D model of a smaller space.
[0129] In various embodiments, these components enable the environment capture system 700 to align SWL and reconstruct 3D models in a similar fashion for both indoors and outdoors. These components can also enable the environment capture system 700 to perform geolocation of 3D models (which can be easily integrated into Google Street View and help align outdoor panoramas if desired).
[0130] The image capture device of the environment capture system 700 can be capable of providing DSLR class images with quality that can be printed at 8.5" x 11" for a 70° VFOV and RGB image type.
[0131] In some embodiments, the environment capture system 700 can take an RGB image with the image capture device (e.g., using a wide angle lens) and then move the lens (using a motor a total of four times) before taking the next RGB image. While the horizontal motor 726 rotates the environment capture system 90 degrees, the LiDAR 708 can capture depth data. In some embodiments, the LiDAR 708 includes an APD array.
[0132] In some embodiments, the image and depth data can then be sent to a capture application (e.g., a device in communication with the environment capture system 700, such as a smart device or an image capture system on a network). In some embodiments, the environment capture system 700 can send the image and depth data to the image processing system 406 for processing and generating a 2D panoramic image or a 3D panoramic image. In various embodiments, the environment capture system 700 can generate a swathe of captured RGB images and depth data from a 360 degree rotation of the environment capture system 700. The swathe can be sent to the image processing system 406 for stitching and alignment. The output of the swathe can be a SWL that includes image data from the lens assembly 704 and depth data from the LiDAR 708 as well as the nature of the swathe, including GPS location and timestamp when the swathe occurred.
[0133] In various embodiments, the LiDAR, vertical mirror, RGB lens, tripod mount, and horizontal drive are rigidly mounted within the housing to allow the housing to be opened without needing to recalibrate the system.
[0134] Figure 9a A block diagram 900 of an example of an environment capture system according to some embodiments is depicted. The block diagram 900 includes a power source 902, a power converter 904, an input / output (I / O) printed circuit board assembly (PCBA), a system on module (SOM) PCBA, a user interface 910, a LiDAR 912, a mirror brushless direct current (BLCD) motor 914, a drivetrain 916, a wide FOV (WFOV) lens 918, and an image sensor 920.
[0135] The power source 902 can be Figure 7 a battery pack 722 of 4x 18650 lithium ion batteries. The power source can be a removable, rechargeable battery, such as a lithium ion battery (e.g., 4x 18650 lithium ion batteries), that is capable of providing power to the environment capture system.
[0136] The power converter 904 can change the voltage level from the power source 902 to a lower or higher voltage level so that it can be utilized by the electronic components of the environment capture system. The environment capture system can utilize 4x 18650 lithium ion batteries in a 4S1P configuration or a four connected in series and one connected in parallel configuration.
[0137] In some embodiments, the I / O PCBA 906 can include elements that provide an IMU, Wi-Fi, GPS, Bluetooth, inertial measurement unit (IMU), motor drivers, and a microcontroller. In some embodiments, the I / O PCBA 906 includes a microcontroller for controlling the horizontal motor and encoding the horizontal motor controls and controlling the vertical motor and encoding the vertical motor controls.
[0138] The SOM PCBA 908 can include a central processing unit (CPU) and / or a graphics processing unit (GPU), memory, and a mobile interface. The SOM PCBA 908 can control the LiDAR 912, the image sensor 920, and the I / O PCBA 906. The SOM PCBA 908 can determine (x, y, z) coordinates associated with each laser pulse of the LiDAR 912 and store the coordinates in a memory component of the SOM PCBA 908. In some embodiments, the SOM PCBA 908 can store the coordinates in an image processing system of the environment capture system 400. In addition to the coordinates associated with each laser pulse, the SOM PCBA 908 can determine additional attributes associated with each laser pulse, including intensity of the laser pulse, number of returns, current number of returns, classification point, RGC value, GPS time, scan angle, and scan direction.
[0139] In some embodiments, the SOM PCBA 908 includes an Nvidia SOM PCBA w / CPU / GPU, DDR, eMMC, Ethernet.
[0140] The user interface 910 can include physical buttons or switches that a user can interact with. The buttons or switches can provide functionality such as turning the environment capture system on and off, scanning a physical environment, and the like. In some embodiments, the user interface 910 can include a display, such as the display 720 of FIG. 7. Figure 7
[0141] In some embodiments, the LiDAR 912 captures depth information of a physical environment. The LiDAR 912 includes an optical sensing module that is capable of measuring the distance to a target or object in a scene by illuminating the target or scene with light using pulses from a laser. The optical sensing module of the LiDAR 912 measures the time it takes for a photon to travel to the target or object and back to a receiver in the LiDAR 912 after reflecting, giving the distance of the LiDAR from the target or object. Along with the distance, the SOM PCBA 908 can determine (x, y, z) coordinates associated with each laser pulse. The LiDAR 912 can fit within a width of 58 mm, a height of 55 mm, and a depth of 60 mm.
[0142] The LiDAR 912 can include a range of 90 m (10% reflectivity), a range of 130 m (20% reflectivity), a range of 260 m (100% reflectivity), a range accuracy of 2 cm (1σ@900 m), a wavelength of 1705 nm, and a beam divergence of 0.28 x 0.03 degrees.
[0143] The SOM PCBA 908 can determine coordinates based on the position of the drive train 916. In various embodiments, the LiDAR 912 can include one or more LiDAR devices. Multiple LiDAR devices can be utilized to increase LiDAR resolution.
[0144] The mirror brushless direct current (BLCD) motor 914 can control Figure 7 the mirror assembly 712.
[0145] In some embodiments, the drive train 916 can include Figure 7 a horizontal motor 726. When the environment capture system is mounted on a platform such as a tripod, the drive train 916 can provide rotation of the environment capture system. The drive train 916 can include a stepper motor Nema 14, a worm and plastic wheel drive train, a clutch, a bushing bearing, and a backlash prevention mechanism. In some embodiments, the environment capture system can be capable of completing a scan in less than 17 seconds. In various embodiments, the drive train 916 has a maximum speed of 60 degrees / second, a maximum acceleration of 300 degrees / second 2 , a maximum torque of 0.5 nm, an angular position accuracy of less than 0.1 degrees, and an encoder resolution of approximately 4096 counts per revolution.
[0146] In some embodiments, the drive train 916 includes a vertical monogon mirror and motor. In this example, the drive train 916 can include a BLDC motor, an external Hall effect sensor, a magnet (paired with the Hall effect sensor), a mirror holder, and a mirror. In this example, the drive train 916 can have a maximum speed of 4,000 rpm and a maximum acceleration of 300 degrees / second2. In some embodiments, the monogon mirror is a dielectric mirror. In one embodiment, the monogon mirror includes a hydrophobic coating or layer.
[0147] The placement of the components of the environment capture system is such that the lens assembly and LiDAR are placed substantially at the center of the axis of rotation. This can reduce image parallax that occurs when the image capture system is not placed at the center of the axis of rotation.
[0148] In some embodiments, the WFOV lens 918 can be Figure 7The WFOV lens 918 focuses light onto an image capture device. In some embodiments, the WFOV lens can have a FOV of at least 145 degrees. With such a wide FOV, image capture of a 360-degree view around the environment capture system can be obtained with three separate image captures by the image capture device. In some embodiments, the WFOV lens 918 can be about ~60mm diameter and ~80mm total lens length (TTL). In one example, the WFOV lens 918 can include a horizontal field of view greater than or equal to 148.3 degrees and a vertical field of view greater than or equal to 94 degrees.
[0149] The image capture device can include the WFOV lens 918 and an image sensor 920. The image sensor 920 can be a CMOS image sensor. In one embodiment, the image sensor 920 is a charge-coupled device (CCD). In some embodiments, the image sensor 920 is a red-green-blue (RGB) sensor. In one embodiment, the image sensor 920 is an IR sensor. In various embodiments, the image capture device can have a resolution of at least 35 pixels per degree (PPD).
[0150] In some embodiments, the image capture device can include an F number of f / 2.4, an image circle diameter of 15.86mm, a pixel pitch of 2.4um, an HFOV of >148.3°, a VFOV of >94.0°, a number of pixels per degree of >38.0 PPD, a chief ray angle at full height of 3.0°, a minimum object distance of 1300mm, a maximum object distance of infinity, a relative illumination of >130%, a maximum distortion of <90%, and a spectral transmission variation of <=5%.
[0151] In some embodiments, the lens can include an F number of 2.8, an image circle diameter of 15.86mm, a number of pixels per degree of >37, a chief ray angle at full height of 3.0, an LI diameter of <60mm, a TTL of <80mm, and a relative illumination of >50%.
[0152] The lens can include a 52 lp / mm (on axis) of >85%, a 104 lp / mm (on axis) of >66%, a 1308 lp / mm (on axis) of >45%, a 52 lp / mm (83% field) of >75%, a 104 lp / mm (83% field) of >41%, and a 1308 lp / mm (83% field) of >25%.
[0153] The environment capture system can have a resolution of >20MP, a green sensitivity of >1.7V / lux*s, an SNR (100 lux, lx gain) of >65dB, and a dynamic range of >70dB.
[0154] Figure 9bA block diagram of an example SOM PCBA 908 of an environment capture system according to some embodiments is depicted. The SOM PCBA 908 can include a communication component 922, a LiDAR control component 924, a LiDAR position component 926, a user interface component 928, a classification component 930, a LiDAR data store 932, and a captured image data store 934.
[0155] In some embodiments, the communication component 922 can send and receive requests or data between any of the components of the SOM PCBA 1008 and Figure 9a the components of the environment capture system of FIG. 1.
[0156] In various embodiments, the LiDAR control component 924 can control various aspects of the LiDAR. For example, the LiDAR control component 924 can send a control signal to the LiDAR 912 to begin emitting laser pulses. The control signal sent by the LiDAR control component 924 can include instructions regarding the frequency of the laser pulses.
[0157] In some embodiments, the LiDAR position component 926 can utilize GPS data to determine the position of the environment capture system. In various embodiments, the LiDAR position component 926 utilizes the position of the mirror assembly to determine a scan angle and (x, y, z) coordinates associated with each laser pulse. The LiDAR position component 926 can also utilize an IMU to determine the orientation of the environment capture system.
[0158] The user interface component 928 can facilitate user interaction with the environment capture system. In some embodiments, the user interface component 928 can provide one or more user interface elements with which a user can interact. The user interface provided by the user interface component 928 can be sent to the user system 1110. For example, the user interface component 928 can provide a visual representation of one area of a floor plan of a building to a user system (e.g., a digital device). The environment capture system can generate the visual representation of the floor plan as the user places the environment capture system in different portions of the floor of the building to capture and generate 3D panoramic images. The user can place the environment capture system in one area of a physical environment to capture and generate 3D panoramic images in that area of the house. Once the 3D panoramic images of the area have been generated by the image processing system, the user interface component can update the floor plan view with an overhead view of the living room area depicted in FIG. lb. In some embodiments, the floor plan view 200 can be generated by the user system 1110 after a second sweep of the same home or floor of a building has been captured.
[0159] In various embodiments, the classification component 930 can classify the type of physical environment. The classification component 930 can analyze the objects in the image or the objects in the image to classify the type of physical environment captured by the environment capture system. In some embodiments, the image processing system can be responsible for classifying the type of physical environment captured by the environment capture system 400.
[0160] The LiDAR data store 932 can be any one structure and / or plurality of structures suitable for captured LiDAR data (e.g., active database, relational database, self-referential database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, and / or the like). The image data store 408 can store captured LiDAR data. However, in the event that the communication network 404 is not functioning, the LiDAR data store 932 can be used to cache captured LiDAR data. For example, in the event that the environment capture system 402 and the user system 1110 are in a remote location without cellular network or in an area without Wi-Fi, the LiDAR data store 932 can store captured LiDAR data until they can be transmitted to the image data store 934.
[0161] Similar to the LiDAR data store, the captured image data store 934 can be any one structure and / or plurality of structures suitable for captured images (e.g., active database, relational database, self-referential database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, and / or the like). The image data store 934 can store captured images.
[0162] Figures 10a-10c A process for taking images by the environment capture system 400 in some embodiments is depicted. As shown, the environment capture system 400 can take a burst of images at different exposures. The burst of images can be a set of images, each image having a different exposure. The first image burst occurs at time 0.0. The environment capture system 400 can receive a first frame and then evaluate the first frame while waiting for a second frame. Figures 10a-10c Indicate blending the first frame before the second frame arrives. In some embodiments, the environment capture system 400 can process each frame to identify pixels, colors, etc. Once the next frame arrives, the environment capture system 400 can process the most recently received frame and then blend the two frames together. Figure 10a
[0163] In various embodiments, the ambient capture system 400 performs image processing to blend a sixth frame and further evaluate the pixels in the blended frame (e.g., a frame that may include elements from any number of frames taken in a burst of images). Optionally, the ambient capture system 400 may transfer the blended image from the graphics processing unit to CPU memory before or during the final step of movement (e.g., rotation) of the ambient capture system 400.
[0164] This process is in Figure 10b Continued in... Figure 10b At the start, the environment capture system 400 takes another burst of shots. The environment capture system 400 can use JxR to compress all or part of the mixed frames and / or captured frames. Similar to... Figure 10a A burst of images can be a group of images, each with a different exposure (the exposure length of each frame in the group can be...). Figure 10a and 10c The other burst shots covered in the same sequence are identical. The second burst shot occurs at 2 seconds. The ambient capture system 400 can receive the first frame and then evaluate it while waiting for the second frame. Figure 10b The instruction is to blend the first frame before the second frame arrives. In some embodiments, the ambient capture system 400 can process each frame to identify pixels, colors, etc. Once the next frame arrives, the ambient capture system 400 can process the most recently received frame and then blend the two frames together.
[0165] In various embodiments, the ambient capture system 400 performs image processing to blend a sixth frame and further evaluate the pixels in the blended frame (e.g., a frame that may include elements from any number of frames taken in a burst of images). Optionally, the ambient capture system 400 may transfer the blended image from the graphics processing unit to CPU memory before or during the final step of movement (e.g., rotation) of the ambient capture system 400.
[0166] After the rotation, the ambient capture system 400 can continue the process by taking another color burst shot at approximately 3.5 seconds (e.g., after rotating 180 degrees). The ambient capture system 400 can use JxR to compress all or part of the mixed frames and / or captured frames. The burst of images can be a set of images, each with a different exposure (the exposure length of each frame in the set can be...). Figure 10a and 10c (The other burst shots covered are the same and in the same order). The environment capture system 400 can receive the first frame and then evaluate that first frame while waiting for the second frame. Figure 10bThe first frame is instructed to be blended before the second frame arrives. In some embodiments, the environment capture system 400 can process each frame to identify pixels, colors, etc. Once the next frame arrives, the environment capture system 400 can process the most recently received frame and then blend the two frames together.
[0167] In various embodiments, the environment capture system 400 performs image processing to blend the sixth frame and further evaluate pixels in the blended frame (e.g., a frame that can include elements from any number of frames of an image burst). During the last step before or during movement (e.g., a turn) of the environment capture system 400, the environment capture system 400 can optionally transfer the blended image from the graphics processing unit to the CPU memory.
[0168] In Figure 10c the last burst occurs at time 5 seconds. The environment capture system 400 can use JxR to compress all or portions of the blended frame and / or the captured frames. The burst of images can be a set of images, each image having a different exposure (the length of exposure for each frame of the set can be different than the length of exposure for the other frames of the set). The environment capture system 400 can receive the first frame and then evaluate the first frame while waiting for the second frame. Figure 10a and 10b covered in and in the same order as other bursts covered in Figure 10c The first frame is instructed to be blended before the second frame arrives. In some embodiments, the environment capture system 400 can process each frame to identify pixels, colors, etc. Once the next frame arrives, the environment capture system 400 can process the most recently received frame and then blend the two frames together.
[0169] In various embodiments, the environment capture system 400 performs image processing to blend the sixth frame and further evaluate pixels in the blended frame (e.g., a frame that can include elements from any number of frames of an image burst). During the last step before or during movement (e.g., a turn) of the environment capture system 400, the environment capture system 400 can optionally transfer the blended image from the graphics processing unit to the CPU memory.
[0170] The dynamic range of an image capture device is a measure of how much light an image sensor can capture. The dynamic range is the difference between the darkest and brightest areas of an image. There are many ways to increase the dynamic range of an image capture device, one of which is to capture multiple images of the same physical environment using different exposures. An image captured with a short exposure will capture brighter areas of the physical environment, while a long exposure will capture darker physical environment areas. In some embodiments, an environment capture system can capture multiple images with six different exposure times. Some or all of the images captured by the environment capture system are used to generate a 2D image with high dynamic range (HDR). One or more of the captured images can be used for other functions, such as ambient light detection, flicker detection, etc.
[0171] A 3D panoramic image of a physical environment can be generated based on four separate image captures by an image capture device and four separate depth data captures by a LiDAR device of an environment capture system. Each of the four separate image captures can include a series of image captures with different exposure times. A blending algorithm can be used to blend the series of image captures with different exposure times to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environment capture system can be used to capture a 3D panoramic image of a kitchen. An image of one wall of the kitchen can include a window, an image with a shorter exposure capture can provide a view outside the window, although it can underexpose the rest of the kitchen. Conversely, another image captured with a longer exposure can provide a view of the interior of the kitchen. A blending algorithm can generate a blended RGB image by blending the view outside the kitchen window from one image with the rest of the kitchen view from another image.
[0172] In various embodiments, a 3D panoramic image can be generated based on three separate image captures by an image capture device and four separate depth data captures by a LiDAR device of an environment capture system. In some embodiments, the number of image captures and the number of depth data captures can be the same. In one embodiment, the number of image captures and the number of depth data captures can be different.
[0173] After capturing a first of a series of images with one exposure time, a blending algorithm receives the first of the series of images, calculates an initial intensity weight for the image, and sets the image as a baseline image for combining subsequently received images. In some embodiments, the blending algorithm can utilize a graphics processing unit (GPU) image processing routine, such as a “blend_kernel” routine. The blending algorithm can receive a subsequent image that can be blended with a previously received image. In some embodiments, the blending algorithm can utilize a variation of the Blend_Kernel GPU image processing routine.
[0174] In one embodiment, the blending algorithm utilizes other methods of blending multiple images, such as determining a difference between the darkest and brightest portions of a baseline image or a contrast of the baseline image to determine whether the baseline image can be overexposed or underexposed. For example, a contrast value that is less than a predetermined contrast threshold value means that the baseline image is overexposed or underexposed. In one embodiment, the contrast of the baseline image can be calculated by taking an average of the light intensities of the image or a subset of the image. In some embodiments, the blending algorithm calculates an average light intensity of each row or column of the image. In some embodiments, the blending algorithm can determine a histogram of each image received from the image capture device and analyze the histogram to determine the light intensities of the pixels that make up each image.
[0175] In various embodiments, blending can involve sampling colors within two or more images of the same scene, including along objects and seams. If there are significant differences in color between the two images (e.g., within predetermined thresholds of color, hue, brightness, saturation, and / or the like), the blending module (e.g., on the environment capture system 400 or the user device 1110) can blend the two images of a predetermined size along where the differences exist. In some embodiments, the greater the difference in color or image at a location in the image, the greater the amount of space that can be blended around or near the location.
[0176] In some embodiments, after blending, the blending module (e.g., on the environment capture system 400 or the user device 1110) can rescan and sample colors along the image(s) to determine whether there are other differences in the image or color that exceed predetermined thresholds of color, hue, brightness, saturation, and / or the like. If so, the blending module can identify portions within the image(s) and continue blending that portion of the image. The blending module can continue resampling the image along the seam until there are no other portions of the image to blend (e.g., any differences in color are below the predetermined threshold(s)).
[0177] Figure 11 A block diagram of an example environment 1100 that is capable of capturing and stitching images to form 3D visualizations is depicted in accordance with some embodiments. The example environment 1100 includes a 3D and panoramic capture and stitching system 1102, a communication network 1104, an image stitching and processor system 1106, an image data store 1108, a user system 1110, and a first scene of a physical environment 1112. The 3D and panoramic capture and stitching system 1102 and / or the user system 1110 can include an image capture device (e.g., the environment capture system 400) that can be used to capture images of an environment (e.g., the physical environment 1112).
[0178] 3D and panoramic capture and stitching system 1102 and image stitching and processor system 1106 can be part of the same system (e.g., part of one or more digital devices) communicatively coupled to environment capture system 400. In some embodiments, one or more functions of components of 3D and panoramic capture and stitching system 1102 and image stitching and processor system 1106 can be performed by environment capture system 400. Similarly or alternatively, 3D and panoramic capture and stitching system 1102 and image stitching and processor system 1106 can be performed by user system 1110 and / or image stitching and processor system 1106.
[0179] A user can utilize 3D panoramic capture and stitching system 1102 to capture multiple 2D images of an environment, such as an interior of a building and / or an exterior of a building. For example, a user can utilize 3D and panoramic capture and stitching system 1102 to capture multiple 2D images of a first scene of physical environment 1112 provided by environment capture system 400. 3D and panoramic capture and stitching system 1102 can include alignment and stitching system 1114. Alternatively, user system 1110 can include alignment and stitching system 1114.
[0180] Alignment and stitching system 1114 can be software, hardware, or a combination of both, configured to provide guidance to a user of an image capture system (e.g., on 3D and panoramic capture and stitching system 1102 or user system 1110) and / or process images to enable improved panoramic pictures to be made (e.g., by stitching, aligning, cropping, etc.). Alignment and stitching system 1114 can be on a computer readable medium (described herein). In some embodiments, alignment and stitching system 1114 can include a processor for performing functions.
[0181] An example of a first scene of physical environment 1112 can be any room, real estate, etc. (e.g., a representation of a living room). In some embodiments, 3D and panoramic capture and stitching system 1102 is used to generate a 3D panoramic image of an indoor environment. In some embodiments, 3D panoramic capture and stitching system 1102 can be about Figure 4 Environment capture system 400 discussed.
[0182] In some embodiments, 3D panoramic capture and stitching system 1102 can communicate with a device and software (e.g., environment capture system 400) for capturing images and depth data. All or part of the software can be installed on 3D panoramic capture and stitching system 1102, user system 1110, environment capture system 400, or both. In some embodiments, a user can interact with 3D and panoramic capture and stitching system 1102 via user system 1110.
[0183] The 3D and panoramic capture and stitching system 1102 or the user system 1110 can obtain multiple 2D images. The 3D and panoramic capture and stitching system 1102 or the user system 1110 can obtain depth data (e.g., from a LiDAR device, etc.).
[0184] In various embodiments, an application on the user system 1110 (e.g., a user’s smart device, such as a smartphone or tablet) or an application on the environment capture system 400 can provide visual or audible guidance to the user for taking images with the environment capture system 400. The graphical guidance can include, for example, a floating arrow on the display of the environment capture system 400 (e.g., on a viewfinder or LED screen on the back of the environment capture system 400) to guide the user where to position and / or point the image capture device. In another example, the application can provide audio guidance about where to position and / or point the image capture device.
[0185] In some embodiments, the guidance can allow the user to capture multiple images of a physical environment without the aid of a stabilizing platform (e.g., a tripod). In one example, the image capture device can be a personal device, such as a smartphone, tablet, media tablet, laptop, etc. The application can provide a direction for each sweep about a location to approximate a parallax-free point based on the location of the image capture device, location information from the image capture device, and / or previous images of the image capture device.
[0186] In some embodiments, the visual and / or audible guidance enables capturing images that can be stitched together to form a panorama without a tripod and without camera positioning information (e.g., indicating the orientation, location, and / or position of the camera from sensors, GPS devices, etc.).
[0187] The alignment and stitching system 1114 can align or stitch 2D images (e.g., captured by the user system 1110 or the 3D panoramic capture and stitching system 1102) to obtain a 2D panoramic image.
[0188] In some embodiments, the alignment and stitching system 1114 aligns or stitches multiple 2D images into a 2D panoramic image with a machine learning algorithm. The parameters of the machine learning algorithm can be managed by the alignment and stitching system 1114. For example, the 3D and panoramic capture and stitching system 1102 and / or the alignment and stitching system 1114 can identify objects within the 2D images to help align the images into a 2D panoramic image.
[0189] In some embodiments, the alignment and stitching system 1114 can utilize depth data and 2D panoramic images to obtain 3D panoramic images. The 3D panoramic images can be provided to the 3D and panoramic stitching system 1102 or the user system 1110. In some embodiments, the alignment and stitching system 1114 determines 3D / depth measurements associated with recognized objects within the 3D panoramic images and / or sends one or more 2D images, depth data, 2D panoramic image(s), 3D panoramic image(s) to the image stitching and processor system 106 to obtain 2D panoramic images or 3D panoramic images having greater pixel resolution than the 2D panoramic images or 3D panoramic images provided by the 3D and panoramic capture and stitching system 1102.
[0190] The communication network 1104 can represent one or more computer networks (e.g., LAN, WAN, etc.) or other transmission mediums. The communication network 1104 can provide communication between the systems 1102, 1106-1110 and / or other systems described herein. In some embodiments, the communication network 104 includes one or more digital devices, routers, cables, buses, and / or other network topologies (e.g., mesh, etc.). In some embodiments, the communication network 1104 can be wired and / or wireless. In various embodiments, the communication network 1104 can include the Internet, one or more wide area networks (WANs) or local area networks (LANs), can be public, private, IP-based, non-IP-based, one or more networks, etc.
[0191] The image stitching and processor system 1106 can process 2D images captured by image capture devices (e.g., the environment capture system 400 or user devices such as smartphones, personal computers, media tablets, etc.) and stitch them into 2D panoramic images. The 2D panoramic images processed by the image stitching and processor system 106 can have higher pixel resolution than the panoramic images obtained by the 3D and panoramic capture and stitching system 1102.
[0192] In some embodiments, the image stitching and processor system 1106 receives and processes 3D panoramic images to produce 3D panoramic images having higher pixel resolution than the pixel resolution of the received 3D panoramic images. The higher pixel resolution panoramic images can be provided to output devices having higher screen resolution than the user system 1110, such as computer screens, projector screens, etc. In some embodiments, the higher pixel resolution panoramic images can provide the panoramic images to the output devices in more detail and can be zoomed in.
[0193] Image data store 1108 can be any one structure and / or plurality of structures (e.g., active database, relational database, self-referential database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar and / or the like) suitable for captured image and / or depth data. Image data store 1108 can store images captured by image capture devices of user system 1110. In various embodiments, image data store 1108 stores depth data captured by one or more depth sensors of user system 1110. In various embodiments, image data store 1108 stores properties associated with image capture devices or properties associated with each of a plurality of image captures or depth captures used to determine a 2D or 3D panoramic image. In some embodiments, image data store 1108 stores a panoramic 2D or 3D panoramic image. The 2D or 3D panoramic image can be determined by 3D and panoramic capture and stitching system 1102 or image stitching and processor system 106.
[0194] User system 1110 can communicate between a user and other related systems. In some embodiments, user system 1110 can be or include one or more mobile devices (e.g., smart phone, cellular phone, smart watch, etc.).
[0195] User system 1110 can include one or more image capture devices. The one or more image capture devices can include, for example, RGB cameras, HDR cameras, video cameras, IR cameras, etc.
[0196] 3D and panoramic capture and stitching system 1102 and / or user system 1110 can include two or more capture devices that can be arranged in positions opposite each other on or within the same mobile housing such that their collective field of view spans up to 360°. In some embodiments, pairs of image capture devices capable of generating stereoscopic pairs of images (e.g., with slightly offset but partially overlapping fields of view) can be used. User system 1110 can include two image capture devices with vertically stereoscopic offset fields of view capable of capturing vertically stereoscopic pairs of images. In another example, user system 1110 can include two image capture devices with vertically stereoscopic offset fields of view capable of capturing vertically stereoscopic pairs of images.
[0197] In some embodiments, the user system 1110, the environment capture system 400, or the 3D and panoramic capture and stitching system 1102 can generate and / or provide image capture location and orientation information. For example, the user system 1110 or the 3D and panoramic capture and stitching system 1102 can include an inertial measurement unit (IMU) to help determine location data associated with one or more image capture devices that capture the plurality of 2D images. The user system 1110 can include a global positioning sensor (GPS) to provide GPS coordinate information associated with the plurality of 2D images captured by the one or more image capture devices.
[0198] In some embodiments, a user can interact with the alignment and stitching system 1114 using a mobile application installed in the user system 1110. The 3D and panoramic capture and stitching system 1102 can provide images to the user system 1110. The user can utilize the alignment and stitching system 1114 on the user system 1110 to view the images and previews.
[0199] In various embodiments, the alignment and stitching system 1114 can be configured to provide or receive one or more 3D panoramic images from the 3D and panoramic capture and stitching system 1102 and / or the image stitching and processor system 1106. In some embodiments, the 3D and panoramic capture and stitching system 1102 can provide to the user system 1110 a visual representation of a portion of a floor plan of a building that has been captured by the 3D and panoramic capture and stitching system 1102.
[0200] A user of the system 1110 can navigate the space around the area and view different rooms of the house. In some embodiments, as the image stitching and processor system 1106 completes the generation of the 3D panoramic image, the user of the user system 1110 can display the 3D panoramic image, such as the example 3D panoramic image. In various embodiments, the user system 1110 generates a preview or thumbnail of the 3D panoramic image. The preview 3D panoramic image can have a lower image resolution than the 3D panoramic image generated by the 3D and panoramic capture and stitching system 1102.
[0201] Figure 12 is a block diagram of an example of the alignment and stitching system 1114 in accordance with some embodiments. The alignment and stitching system 1114 includes a communication module 1202, an image capture location module 1204, a stitching module 1206, a cropping module 1208, a graph cut module 1210, a blending module 1211, a 3D image generator 1214, a captured 2D image data store 1216, a 3D panoramic image data store 1218, and a guidance module 220. It can be appreciated that there can be any number of modules of the alignment and stitching system 1114 that perform one or more different functions as described herein.
[0202] In some embodiments, the alignment and stitching system 1114 includes an image capture module configured to receive images from one or more image capture devices (e.g., cameras). The alignment and stitching system 1114 can also include a depth module configured to receive depth data from depth devices, such as LiDAR, if available.
[0203] The communication module 1202 can send and receive requests, images, or data between any module or data store of the alignment and stitching system 1114 and Figure 11 components of the example environment 1100. Similarly, the alignment and stitching system 1114 can send and receive requests, images, or data to any device or system over the communication network 1104.
[0204] In some embodiments, the image capture location module 1204 can determine image capture device location data for an image capture device (e.g., a camera, which can be a standalone camera, a smartphone, a media tablet, a laptop, etc.). The image capture device location data can indicate the position and orientation of the image capture device and / or lens. In one example, the image capture location module 1204 can utilize an IMU of the user system 1110, a camera, a digital device with a camera, or the 3D and panoramic capture and stitching system 1102 to generate position data for the image capture device. The image capture location module 1204 can determine the current direction, angle, or tilt of one or more image capture devices (or lenses). The image capture location module 1204 can also utilize a GPS of the user system 1110 or the 3D and panoramic capture and stitching system 1102.
[0205] For example, when a user wants to capture a 360° view of a physical environment, such as a living room, using the user system 1110, the user can hold the user system 1110 in front of them at eye level to start capturing one of a plurality of images that will eventually become a 3D panoramic image. To reduce the amount of parallax in the images and capture images that are more suitable for stitching and generating a 3D panoramic image, it can be preferable if the one or more image capture devices are rotated at the center of the axis of rotation. The alignment and stitching system 1114 can receive position information (e.g., from the IMU) to determine the position of the image capture device or lens. The alignment and stitching system 1114 can receive and store the field of view of the lens. The guidance module 1220 can provide visual and / or audio information about the recommended initial position of the image capture device. The guidance module 1220 can make recommendations for positioning the image capture device for subsequent images. In one example, the guidance module 1220 can provide guidance to the user to rotate and position the image capture device so that the image capture device is rotated close to the center of rotation. In addition, the guidance module 1220 can provide guidance to the user to rotate and position the image capture device so that subsequent images are substantially aligned based on the field of view and / or characteristics of the image capture device.
[0206] The guidance module 1220 can provide visual guidance to the user. For example, the guidance module 1220 can place a marker or arrow in a viewer or display on the user system 1110 or the 3D and panoramic capture and stitching system 1102. In some embodiments, the user system 1110 can be a smartphone or tablet computer with a display. The guidance module 1220 can position one or more markers (e.g., different colored markers or the same marker) on the output device and / or in the viewfinder when taking one or more pictures. The user can then use the markers on the output device and / or viewfinder to align the next image.
[0207] There are many techniques for guiding a user of the user system 1110 or the 3D and panoramic capture and stitching system 1102 to take a plurality of images to facilitate stitching the images together into a panorama. When a panorama is derived from a plurality of images, the images can be stitched together. To improve the time, efficiency, and effectiveness of stitching the images together while reducing the need to correct for artifacts or misalignments, the image capture position module 1204 and the guidance module 1220 can help the user take a plurality of images in positions that improve the quality, time efficiency, and effectiveness of the image stitching of the desired panorama.
[0208] For example, after taking the first picture, the display of the user system 1110 can include two or more objects, such as circles. The two circles can appear to be stationary relative to the environment, and the two circles can move with the user system 1110. When the two stationary circles align with the two circles that move with the user system 1110, the image capture device and / or the user system 1110 can align for the next image.
[0209] In some embodiments, after the image capture device takes an image, the image capture position module 1204 can obtain sensor measurements of the position of the image capture device (e.g., including orientation, tilt, etc.). The image capture position module 1204 can determine one or more edges of the taken image by calculating the position of the edges of the field of view based on the sensor measurements. Additionally or alternatively, the image capture position module 1204 can determine one or more edges of the image by scanning the image taken by the image capture device, identifying objects within the image (e.g., using the machine learning models discussed herein), determining one or more edges of the image, and positioning the objects (e.g., circles or other shapes) at the edges of the display on the user system 1110.
[0210] The image capture position module 1204 can display two objects within the display of the user system 1110 that indicate the positioning of the field of view for the next picture. The two objects can indicate the position in the environment that represents the edge of the last image. The image capture position module 1204 can continue to receive sensor measurements of the position of the image capture device and calculate two additional objects in the field of view. The two additional objects can be spaced apart the same width as the first two objects. While the first two objects can represent the edge of the taken image (e.g., the rightmost edge of the image), the next two additional objects that represent the edge of the field of view can be on the opposite edge (e.g., the leftmost edge of the field of view). By having the user physically align the first two objects on the edge of the image with the additional two objects on the opposite edge of the field of view, the image capture device can be positioned to take another image that can be more effectively stitched together without a tripod. This process can continue for each image until the user determines that the desired panorama has been captured.
[0211] While multiple objects are discussed herein, it should be understood that the image capture position module 1204 can calculate the position of one or more objects for positioning the image capture device. The objects can be any shape (e.g., circular, elliptical, square, emoji, arrow, etc.). In some embodiments, the objects can have different shapes.
[0212] In some embodiments, there can be a distance between objects representing edges of captured images, and there can be a distance between objects in the field of view. The user can be directed to move forward to move away to enable sufficient distance between objects. Alternatively, the size of objects in the field of view can change to match the size of objects representing edges of captured images as the image capture device approaches the correct position (e.g., by moving closer or further away to a position that would enable the next image to be taken in a position that would improve image stitching).
[0213] In some embodiments, the image capture position module 1204 can utilize objects in images captured by the image capture device to estimate the position of the image capture device. For example, the image capture position module 1204 can utilize GPS coordinates to determine a geographic location associated with an image. The image capture position module 1204 can use the position to identify landmarks that can be captured by the image capture device.
[0214] The image capture position module 1204 can include a 2D machine learning model that converts 2D images to 2D panoramic images. The image capture position module 1204 can include a 3D machine learning model that converts 2D images to 3D representations. In one example, a 3D representation can be utilized to display a three-dimensional walkthrough or visualization of an interior and / or exterior environment.
[0215] The 2D machine learning model can be trained to stitch or help stitch together two or more 2D images to form a 2D panoramic image. The 2D machine learning model can be, for example, a neural network trained with 2D images that include physical objects in the images as well as object recognition information that trains the 2D machine learning model to identify objects in subsequent 2D images. The objects in the 2D images can help determine the location(s) within the 2D images to help determine edges of the 2D images, warping in the 2D images, and help alignment of the images. In addition, the objects in the 2D images can help determine artifacts in the 2D images, mixing of artifacts or boundaries between two images, locations to cut images and / or crop images.
[0216] In some embodiments, the 2D machine learning model can be, for example, a neural network trained with 2D images that include depth information of the environment (e.g., from a LiDAR device or structured light device of the user system 1110 or the 3D and panoramic capture and stitching system 1102) as well as include physical objects in the images to identify the physical objects, locations of the physical objects, and / or locations of the image capture device / field of view. The 2D machine learning model can identify the physical objects as well as their depth relative to other aspects of the 2D images to help alignment and positioning of two 2D images for stitching (or stitching two 2D images).
[0217] The 2D machine learning models can include any number of machine learning models (e.g., any number of models generated by neural networks, etc.).
[0218] The 2D machine learning models can be stored on the 3D and panoramic capture and stitching system 1102, the image stitching and processor system 1106, and / or the user system 1110. In some embodiments, the 2D machine learning models can be trained by the image stitching and processor system 1106.
[0219] The image capture position module 1204 can estimate the position of the image capture device (the position of the field of view of the image capture device) based on the seam between the two or more 2D images from the stitching module 1206, the image warping from the cropping module 1208, and / or the graphical cut from the graphical cut module 1210.
[0220] The stitching module 1206 can combine the two or more 2D images to generate a 2D panorama. Based on the seam between the two or more 2D images from the stitching module 1206, the image warping from the cropping module 1208, and / or the graphical cut, it has a field of view that is greater than the field of view of each of the two or more images.
[0221] The stitching module 1206 can be configured to align or“stitch together” two different 2D images that provide different perspectives of the same environment to generate a panoramic 2D image of the environment. For example, the stitching module 1206 can employ information about the capture position and orientation of the respective 2D images that is known or derived (e.g., using the techniques described herein) to help stitch the two images together.
[0222] The stitching module 1206 can receive two 2D images. The first 2D image can be taken immediately before or within a predetermined time period of the second image. In various embodiments, the stitching module 1206 can receive positioning information for the image capture device associated with the first image and then receive positioning information associated with the second image. The positioning information can be associated with the images based on positioning data from an IMU, GPS, and / or information provided by a user at the time the images were taken.
[0223] In some embodiments, the stitching module 1206 can utilize a 2D machine learning module to scan the two images to recognize objects within the two images, including objects (or portions of objects) that can be shared by the two images. For example, the stitching module 1206 can identify a corner shared at the relative edges of the two images, a pattern on a wall, furniture, etc.
[0224] The stitching module 1206 can align the edges of the two 2D images based on the positioning of the shared object (or portion of the object), the positioning data from the IMU, the positioning data from the GPS, and / or the information provided by the user, and then combine the two edges of the images (i.e., “stitch” them together). In some embodiments, the stitching module 1206 can identify a portion of the two 2D images that overlap each other and stitch the images at the location of the overlap (e.g., using the positioning data and / or the results of the 2D machine learning model).
[0225] In various embodiments, the 2D machine learning model can be trained to combine or stitch the two edges of the images using the positioning data from the IMU, the positioning data from the GPS, and / or the information provided by the user. In some embodiments, the 2D machine learning model can be trained to identify common objects in the two 2D images to align and position the 2D images, and then combine or stitch the two edges of the images. In further embodiments, the 2D machine learning model can be trained to align and position the 2D images using the positioning data and object recognition, and then stitch the two edges of the images together to form all or a portion of a panoramic 2D image.
[0226] The stitching module 1206 can utilize the depth information of the respective images (e.g., pixels in the respective images, objects in the respective images, etc.) to facilitate aligning the respective 2D images with each other in association with generating a single 2D panoramic image of the environment.
[0227] The cropping module 1208 can address issues with two or more 2D images in situations where the image capture device was not held in the same position when capturing the 2D images. For example, while capturing an image, the user can position the user system 1110 in a vertical position. However, while capturing another image, the user can position the user system at an angle. The resulting images can be misaligned and can have a parallax effect. The parallax effect occurs when the foreground and background objects are not aligned in the same way in the first image and the second image.
[0228] The cropping module 1208 can utilize the 2D machine learning model (by applying the positioning information, the depth information, and / or the object recognition) to detect the change in position of the image capture device in the two or more images, and then measure the amount of change in the position of the image capture device. The cropping module 1208 can warp one or more of the 2D images such that when the images are stitched, the images can be able to align together to form a panoramic image, while at the same time preserving certain characteristics of the images, such as keeping straight lines straight.
[0229] The output of the cropping module 1208 can include the number of pixel columns and rows to offset each pixel of the image to straighten the image. The amount of offset for each image can be output in the form of a matrix representing the number of pixel columns and rows to offset each pixel of the image.
[0230] In some embodiments, the cropping module 1208 can determine an amount of image warping to perform on one or more of the plurality of 2D images captured by the image capture device of the user system 1110 based on one or more image capture locations from the image capture location module 1204, a seam between two or more 2D images from the stitching module 1206, a graphical cut from the graphical cut module 1210, or a blending of colors from the blending module 1211.
[0231] The graphical cut module 1210 can determine where to cut or slice one or more 2D images captured by the image capture device. For example, the graphical cut module 1210 can utilize a 2D machine learning model to identify objects in two images and determine that they are the same object. The image capture location module 1204, the cropping module 1208, and / or the graphical cut module 1210 can determine that the two images cannot be aligned even when warped. The graphical cut module 1210 can utilize information from the 2D machine learning model to identify sections of the two images that can be stitched together (e.g., by cutting out a portion of one or both images to help with alignment and positioning). In some implementations, the two 2D images can overlap with at least a portion of the physical world represented in the images. The graphical cut module 1210 can identify objects in both images, such as the same chair. However, even after image warping by the image capture location and cropping module 1208, the images of the chair cannot be aligned to generate an undistorted panorama and properly represent a portion of the physical world. The graphical cut module 1210 can select one of the two images of the chair as the correct representation (e.g., based on the misalignment, positioning, and / or artifacts of one image when compared to the other image) and cut out the chair from the image with the misalignment, positioning error, and / or artifacts. The stitching module 1206 can subsequently stitch the two images together.
[0232] The graphical cut module 1210 can attempt both combinations, e.g., cutting the image of the chair from the first image and stitching the first image with the chair subtracted to the second image, to determine which graphical cut generates a more accurate panoramic image. The output of the graphical cut module 1210 can be the location to cut one or more of the plurality of 2D images corresponding to the graphical cut that generates a more accurate panoramic image.
[0233] The graphic cut module 1210 can determine how to cut or slice one or more 2D images captured by an image capture device based on one or more image capture locations from the image capture location module 1204, stitching or seams between two or more 2D images from the stitching module 1206, image warping from the cropping module 1208, and graphic cuts from the graphic cut module 1210.
[0234] The blending module 1211 can color at a seam (e.g., stitching) between two images so that the seam is not visible. Changes in lighting and shading can cause the same object or surface to output in slightly different colors or shades. The blending module can determine the amount of color blending needed based on one or more image capture locations from the image capture location module 1204, stitching, image color along the seam from the two images, image warping from the cropping module 1208, and / or graphic cuts from the graphic cut module 1210.
[0235] In various embodiments, the blending module 1211 can receive a panorama from the combination of two 2D images and then sample colors along the seam of the two 2D images. The blending module 1211 can receive seam location information from the image capture location module 1204 to enable the blending module 1211 to sample colors along the seam and determine differences. If there are significant differences in color along the seam between the two images (e.g., within a predetermined threshold of color, hue, brightness, saturation, and / or the like), the blending module 1211 can blend the two images along the seam at the location where the difference exists for a predetermined size. In some embodiments, the greater the difference in color or image along the seam, the greater the amount of space along the seam of the two images that can be blended.
[0236] In some embodiments, after blending, the blending module 1211 can rescan and sample colors along the seam to determine if there are other differences in the images or colors that exceed a predetermined threshold of color, hue, brightness, saturation, and / or the like. If so, the blending module 1211 can identify the portion along the seam and continue blending that portion of the image. The blending module 1211 can continue to resample the image along the seam until there are no other portions of the image to blend (e.g., any differences in color are below the predetermined threshold).
[0237] The 3D image generator 1214 can receive the 2D panoramic images and generate 3D representations. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to transform the 2D panoramic images into 3D representations. The 3D machine learning model can be trained using the 2D panoramic images and depth data (e.g., from LiDAR sensors or structured light devices) to produce 3D representations. The 3D representations can be tested and reviewed for curation and feedback. In some embodiments, the 3D machine learning model can be used with the 2D panoramic images and depth data to generate 3D representations.
[0238] In various embodiments, the accuracy, rendering speed, and quality of the 3D representations generated by the 3D image generator 1214 are greatly improved by utilizing the systems and methods described herein. For example, by rendering 3D representations from 2D panoramic images that have been aligned, positioned, and stitched using the methods described herein (e.g., by alignment and positioning information provided by hardware, by improved positioning caused by guidance provided to the user during image capture, by cropping and changing the distortion of the images, by cutting the images to avoid artifacts and overcome distortion, by blending the images, and / or any combination), the accuracy, rendering speed, and quality of the 3D representations are improved. Further, it should be appreciated that training of the 3D machine learning model can be greatly improved (e.g., in speed and accuracy) by utilizing 2D panoramic images that have been aligned, positioned, and stitched using the methods described herein. Further, in some embodiments, the 3D machine learning model can be smaller and less complex due to the reduction in processing and learning to overcome misalignment, positioning errors, distortion, poor graphical cutting, poor blending, artifacts, and the like to generate reasonably accurate 3D representations.
[0239] The trained 3D machine learning model can be stored in the 3D and panoramic capture and stitching system 1102, the image stitching and processor system 106, and / or the user system 1110.
[0240] In some embodiments, a 3D machine learning model can be trained using multiple 2D images and depth data from image capture devices of user system 1110 and / or 3D and panoramic capture and stitching system 1102. Additionally, 3D image generator 1214 can be trained using image capture location information associated with each of the multiple 2D images from image capture location module 1204, seam location of each of the multiple 2D images aligned or stitched from stitching module 1206, pixel offset(s) of each of the multiple 2D images from cropping module 1208, and / or graphic cuts from graphic cutting module 1210. In some embodiments, the 3D machine learning model can be used with 2D panoramic images, depth data, image capture location information associated with each of the multiple 2D images from image capture location module 1204, seam location of each of the multiple 2D images aligned or stitched from stitching module 1206, pixel offset(s) of each of the multiple 2D images from cropping module 1208, and / or graphic cuts from graphic cutting module 1210 to generate a 3D representation.
[0241] Stitching module 1206 can be a part of a 3D model that converts multiple 2D images into a 2D panoramic or 3D panoramic image. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D. Cropping module 1208 can be a part of a 3D model that converts multiple 2D images into a 2D panoramic or 3D panoramic image. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D. Graphic cutting module 1210 can be a part of a 3D model that converts multiple 2D images into a 2D panoramic or 3D panoramic image. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D. Hybrid module 1211 can be a part of a 3D machine learning model that converts multiple 2D images into a 2D panoramic or 3D panoramic image. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D.
[0242] 3D image generator 1214 can generate a weight for each of image capture location module 1204, cropping module 1208, graphic cutting module 1210, and hybrid module 1211, which can represent the reliability or “strength” or “weakness” of the module. In some embodiments, the sum of the weights of the modules equals 1.
[0243] In instances where depth data is not available for multiple 2D images, the 3D image generator 1214 can determine depth data for one or more objects in the multiple 2D images captured by the image capture devices of the user system 1110. In some embodiments, the 3D image generator 1214 can derive depth data based on images captured by a stereo image pair. Rather than determining depth data from a passive stereo algorithm, the 3D image generator can evaluate a stereo image pair to determine data about photometric matching quality between images at various depths (more intermediate results).
[0244] The 3D image generator 1214 can be part of a 3D model that converts multiple 2D images into a 2D panorama or a 3D panoramic image. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D.
[0245] The captured 2D image data store 1216 can be any one structure and / or plurality of structures suitable for captured images and / or depth data (e.g., active database, relational database, self-referential database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar and / or the like). The captured 2D image data store 1216 can store images captured by the image capture devices of the user system 1110. In various embodiments, the captured 2D image data store 1216 stores depth data captured by one or more depth sensors of the user system 1110. In various embodiments, the captured 2D image data store 1216 stores image capture device parameters associated with the image capture devices, or capture properties associated with each of the multiple image captures, or depth captures used to determine a 2D panoramic image. In some embodiments, the image data store 1108 stores a panoramic 2D panoramic image. The 2D panoramic image can be determined by the 3D and panoramic capture and stitching system 1102 or the image stitching and processor system 106. The image capture device parameters can include lighting, color, image capture lens focal length, maximum aperture, tilt angle, and the like. The capture properties can include pixel resolution, lens distortion, lighting, and other image metadata.
[0246] 3D panoramic image data store 1218 can be any one structure and / or plurality of structures suitable for 3D panoramic images (e.g., active database, relational database, self-referential database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar and / or the like). 3D panoramic image data store 1218 can store 3D panoramic images generated by 3D and panoramic capture and stitching system 1102. In various embodiments, 3D panoramic image data store 1218 stores properties associated with the image capture device or properties associated with each of the plurality of image captures or depth captures used to determine the 3D panoramic image. In some embodiments, 3D panoramic image data store 1218 stores 3D panoramic images. 2D or 3D panoramic images can be determined by 3D and panoramic capture and stitching system 1102 or image stitching and processor system 106.
[0247] Figure 13 A flowchart 1300 of a 3D panoramic image capture and generation process is depicted in accordance with some embodiments. In step 1302, an image capture device can capture a plurality of 2D images using image sensor 920 and WFOV lens 918 of FIG. 9. A wider FOV means that the environment capture system 402 will need fewer scans to obtain a 360° view. WFOV lens 918 can also be wider horizontally as well as vertically. In some embodiments, image sensor 920 captures RGB images. In one embodiment, image sensor 920 captures black and white images.
[0248] In step 1304, the environment capture system can send the captured 2D images to image stitching and processor system 1106. Image stitching and processor system 1106 can apply a 3D modeling algorithm to the captured 2D images to generate a panoramic 2D image. In some embodiments, the 3D modeling algorithm is a machine learning algorithm for stitching the captured 2D images into a panoramic 2D image. In some embodiments, step 1304 can be optional.
[0249] In step 1306, LiDAR 912 and WFOV lens 918 of FIG. 9 can capture LiDAR data. A wider FOV means that the environment capture system 400 will need fewer scans to obtain a 360° view.
[0250] In step 1308, the LiDAR data can be sent to image stitching and processor system 1106. Image stitching and processor system 1106 can input the LiDAR data and the captured 2D images into a 3D modeling algorithm to generate a 3D panoramic image. The 3D modeling algorithm is a machine learning algorithm.
[0251] In step 1310, the image stitching and processor system 1106 generates a 3D panoramic image. The 3D panoramic image can be stored in the image data store 408. In one embodiment, the 3D panoramic image generated by the 3D modeling algorithm is stored in the image stitching and processor system 1106. In some embodiments, as the environment capture system is used to capture various portions of the physical environment, the 3D modeling algorithm can generate a visual representation of a floor plan of the physical environment.
[0252] In step 1312, the image stitching and processor system 1106 can provide at least a portion of the generated 3D panoramic image to the user system 1110. The image stitching and processor system 1106 can provide a visual representation of a floor plan of the physical environment.
[0253] The order of one or more steps of the flowchart 1300 can be changed without affecting the end result of the 3D panoramic image. For example, the environment capture system can interleave image capture with the image capture device with LiDAR data or depth information capture with the LiDAR 912. For example, the image capture device can capture an image of a section of the physical environment with the image capture device, and then the LiDAR 912 obtains depth information from the section 1605. Once the LiDAR 912 obtains the depth information from the section, the image capture device can continue to move to capture an image of another section, and then the LiDAR 912 obtains depth information from the section, interleaving the image capture and depth information capture.
[0254] In some embodiments, the devices and / or systems discussed herein employ one image capture device to capture 2D input images. In some embodiments, the one or more image capture devices 1116 can represent a single image capture device (or image capture lens). According to some of these embodiments, a user of a mobile device housing the image capture device can be configured to rotate about an axis to generate images at different capture orientations relative to an environment, with a common field of view of the images spanning up to 360° horizontally.
[0255] In various embodiments, the devices and / or systems discussed herein can employ two or more image capture devices to capture 2D input images. In some embodiments, two or more image capture devices can be arranged in opposing positions to one another on or within the same mobile housing such that their collective field of view spans up to 360°. In some embodiments, pairs of image capture devices capable of generating stereoscopic pairs of images (e.g., with slightly offset but partially overlapping fields of view) can be used. For example, user system 1110 (e.g., a device including one or more image capture devices for capturing 2D input images) can include two image capture devices with horizontally stereoscopic offset fields of view capable of capturing stereoscopic pairs of images. In another example, user system 1110 can include two image capture devices with vertically stereoscopic offset fields of view capable of capturing stereoscopic pairs of images. According to either of these examples, each camera can have a field of view spanning up to 360°. In this regard, in one embodiment, user system 1110 can employ two panoramic cameras with vertical stereoscopic offset capable of capturing pairs of panoramic images forming stereoscopic pairs (with vertical stereoscopic offset).
[0256] Positioning component 1118 can include any hardware and / or software configured to capture user system position data and / or user system orientation data. For example, positioning component 1118 includes an IMU to generate user system 1110 position data associated with the one or more image capture devices of user system 1110 used to capture a plurality of 2D images. Positioning component 1118 can include a GPS unit to provide GPS coordinate information associated with a plurality of 2D images captured by the one or more image capture devices. In some embodiments, positioning component 1118 can associate position data and orientation data of a user system with respective images captured using the one or more image capture devices of user system 1110.
[0257] Various embodiments of the device provide a user with 3D panoramic images of indoor and outdoor environments. In some embodiments, the device can efficiently and quickly provide a user with 3D panoramic images of indoor and outdoor environments using a single wide field of view (FOV) lens and a single light and detection and ranging sensor (LiDAR sensor).
[0258] The following is an example use case for the example device described herein. The following use case is one of many embodiments. Different embodiments of the device as discussed herein can include one or more features and capabilities similar to those of this use case.
[0259] Figure 14 A flowchart of a 3D and panoramic capture and stitching process 1400 is depicted in accordance with some embodiments. Figure 14The flowchart of FIG. 1 refers to the 3D and panoramic capture and stitching system 1102 as including image capture devices, but in some embodiments, the data capture devices can be the user system 1110.
[0260] In step 1402, the 3D and panoramic capture and stitching system 1102 can receive a plurality of 2D images from at least one image capture device. The image capture device of the 3D and panoramic capture and stitching system 1102 can be or include a complementary metal-oxide-semiconductor (CMOS) image sensor. In various embodiments, the image capture device is a charge-coupled device (CCD). In one example, the image capture device is a red-green-blue (RGB) sensor. In one embodiment, the image capture device is an IR sensor. Each of the plurality of 2D images can have a field of view that partially overlaps with at least one other image of the plurality of 2D images. In some implementations, at least some of the plurality of 2D images combine to produce a 360° view of a physical environment (e.g., an indoor, outdoor, or both).
[0261] In some embodiments, all of the plurality of 2D images are received from the same image capture device. In various implementations, at least a portion of the plurality of 2D images are received from two or more image capture devices of the 3D and panoramic capture and stitching system 1102. In one example, the plurality of 2D images includes a set of RGB images and a set of IR images, where the IR images provide depth data to the 3D and panoramic capture and stitching system 1102. In some embodiments, each 2D image can be associated with depth data provided by a LiDAR device. In some implementations, each 2D image can be associated with localization data.
[0262] In step 1404, the 3D and panoramic capture and stitching system 1102 can receive capture parameters and image capture device parameters associated with each of the received plurality of 2D images. The image capture device parameters can include illumination, color, image capture lens focal length, maximum aperture, field of view, etc. The capture properties can include pixel resolution, lens distortion, illumination, and other image metadata. The 3D and panoramic capture and stitching system 1102 can also receive localization data and depth data.
[0263] In step 1406, the 3D and panoramic capture and stitching system 1102 can take the information received from steps 1402 and 1404 to stitch the 2D images to form a 2D panoramic image. Further discussion of the process of stitching 2D images is discussed with respect to Figure 15 The flowchart of FIG. 1 refers to the 3D and panoramic capture and stitching system 1102 as including image capture devices, but in some embodiments, the data capture devices can be the user system 1110.
[0264] In step 1408, the 3D and panoramic capture and stitching system 1102 can apply a 3D machine learning model to generate a 3D representation. The 3D representation can be stored in a 3D panoramic image datastore. In various embodiments, the 3D representation is generated by the image stitching and processor system 1106. In some embodiments, as the environment capture system is used to capture various portions of the physical environment, the 3D machine learning model can generate a visual representation of a floorplan of the physical environment.
[0265] In step 1410, the 3D and panoramic capture and stitching system 1102 can provide at least a portion of the generated 3D representation or model to the user system 1110. The user system 1110 can provide a visual representation of a floorplan of the physical environment.
[0266] In some embodiments, the user system 1110 can send the plurality of 2D images, capture parameters, and image capture parameters to the image stitching and processor system 1106. In various embodiments, the 3D and panoramic capture and stitching system 1102 can send the plurality of 2D images, capture parameters, and image capture parameters to the image stitching and processor system 1106.
[0267] The image stitching and processor system 1106 can process the plurality of 2D images captured by the image capture devices of the user system 1110 and stitch them into a 2D panoramic image. The 2D panoramic image processed by the image stitching and processor system 1106 can have a higher pixel resolution than the 2D panoramic image obtained by the 3D and panoramic capture and stitching system 1102.
[0268] In some embodiments, the image stitching and processor system 106 can receive the 3D representation and output a 3D panoramic image having a higher pixel resolution than the pixel resolution of the received 3D panoramic image. The higher pixel resolution panoramic image can be provided to an output device having a higher screen resolution than the user system 1110, such as a computer screen, a projector screen, etc. In some embodiments, the higher pixel resolution panoramic image can provide the panoramic image to the output device in more detail and can be zoomed in.
[0269] Figure 15 depicts a flowchart illustrating a method 1400 for generating a 3D representation of a physical environment in accordance with some embodiments Figure 14A flowchart providing further details of one step in the 3D and panoramic capture and stitching process. In step 1502, the image capture position module 1204 may determine image capture device position data associated with each image captured by the image capture device. The image capture position module 1204 may utilize the IMU of the user system 1110 to determine position data of the image capture device (or the field of view of the lens of the image capture device). The position data may include the orientation, angle, or tilt of one or more image capture devices when capturing one or more 2D images. One or more of the cropping module 1208, the graphic cutting module 1210, or the blending module 1212 may utilize the orientation, angle, or tilt associated with each of the multiple 2D images to determine how to warp, cut, and / or blend the images.
[0270] In step 1504, the cropping module 1208 can distort one or more of the plurality of 2D images such that two images can be aligned together to form a panoramic image while preserving specific characteristics of the images, such as keeping straight lines straight. The output of the cropping module 1208 may include the number of pixel columns and rows to offset each pixel of the image to straighten it. The offset of each image may be output as a matrix representing the number of pixel columns and rows offset per pixel of the image. In this embodiment, the cropping module 1208 may determine the desired amount of distortion for each of the plurality of 2D images based on an image capture pose estimation of each of the plurality of 2D images.
[0271] In step 1506, the image cutting module 1210 determines where to cut or slice one or more of the plurality of 2D images. In this embodiment, the image cutting module 1210 may determine where to cut or slice each of the plurality of 2D images based on image capture pose estimation and image distortion of each of the plurality of 2D images.
[0272] In step 1508, the stitching module 1206 can stitch two or more images together using the edges and / or cuts of the images. The stitching module 1206 can align and / or position the images based on objects detected within the images, image distortions, cuts, etc.
[0273] In step 1510, the blending module 1212 can adjust the color at the seam (e.g., the stitching of two images) or at a location on one image that contacts or connects to another image. The blending module 1212 can determine the desired amount of color blending based on one or more image capture positions from the image capture position module 1204, image distortion from the cropping module 1208, and graphic cutting from the graphic cutting module 1210.
[0274] The order of one or more steps in the 3D and panoramic capture and stitching process 1400 can be changed without affecting the final result of the 3D panoramic image. For example, an environment capture system can interleave image capture using an image capture device with LiDAR data or depth information capture. For example, an image capture device can use an image capture device to capture the physical environment. Figure 16 The image capture device captures an image of segment 1605, and then the LiDAR 612 obtains depth information from segment 1605. Once the LiDAR has obtained depth information from segment 1605, the image capture device can continue to move to capture an image of another segment 1610, and then the LiDAR 612 obtains depth information from segment 1610, thus interleaving image capture and depth information capture.
[0275] Figure 16 A block diagram of an example digital device 1602 according to some embodiments is depicted. Any of the user system 1110, the 3D panoramic capture and stitching system 1102, and the image stitching and processor system may include an instance of the digital device 1602. The digital device 1602 includes a processor 1604, a memory 1606, a storage device 1608, an input device 1610, a communication network interface 1612, an output device 1614, an image capture device 1616, and a positioning component 1618. The processor 1604 is configured to execute executable instructions (e.g., a program). In some embodiments, the processor 1604 includes circuitry or any processor capable of processing executable instructions.
[0276] Memory 1606 stores data. Some examples of memory 1606 include storage devices such as RAM, ROM, RAM cache, virtual memory, etc. In various embodiments, working data is stored in memory 1606. Data in memory 1606 can be erased or eventually transferred to storage device 1608.
[0277] Storage device 1608 includes any storage device configured to retrieve and store data. Some examples of storage device 1608 include flash drives, hard disk drives, optical disk drives, and / or magnetic tapes. Each of memory 1606 and storage device 1608 includes a computer-readable medium storing instructions or programs executable by processor 1604.
[0278] Input device 1610 is any device that inputs data (e.g., touch keyboard, stylus). Output device 1614 outputs data (e.g., speaker, display, virtual reality headset). It should be understood that storage device 1608, input device 1610, and output device 1614. In some embodiments, output device 1614 is optional. For example, a router / switch can include processor 1604 and memory 1606 and devices that receive and output data (e.g., communication network interface 1612 and / or output device 1614).
[0279] Communication network interface 1612 can be coupled to a network (e.g., communication network 104) via communication network interface 1612. Communication network interface 1612 can support communication over Ethernet connections, serial connections, parallel connections, and / or ATA connections. Communication network interface 1612 can also support wireless communication (e.g., 802.16a / b / g / n, WiMAX, LTE, Wi-Fi). It will be apparent that communication network interface 1612 can support many wired and wireless standards.
[0280] A component can be hardware or software. In some embodiments, a component can configure one or more processors to perform the functions associated with the component. Although different components are discussed herein, it should be understood that a server system can include any number of components performing any or all of the functions discussed herein.
[0281] Digital device 1602 can include one or more image capture devices 1616. The one or more image capture devices 1616 can include, for example, RGB cameras, HDR cameras, video cameras, and the like. According to some embodiments, the one or more image capture devices 1616 can also include video cameras capable of capturing video. In some embodiments, the one or more image capture devices 1616 can include image capture devices that provide a relatively standard field of view (e.g., approximately 75°). In other embodiments, the one or more image capture devices 1616 can include cameras that provide a relatively wide field of view (e.g., from approximately 120° up to 360°), such as fisheye cameras and the like (e.g., digital device 1602 can include or be included in environment capture system 400).
[0282] A component can be hardware or software. In some embodiments, a component can configure one or more processors to perform the functions associated with the component. Although different components are discussed herein, it should be understood that a server system can include any number of components performing any or all of the functions discussed herein.
Claims
1. A method of an apparatus for capturing image and depth information in a 360-degree scene, wherein the apparatus comprises a first motor, an image capturing device, and a depth information capturing device, the method comprising: capturing, by the image capturing device pointing in a first direction, a plurality of images at different exposures in a first field of view of the 360-degree scene, wherein the first field of view, first FOV, the image capturing device comprises a lens coupled to a frame of the apparatus at a first position between a front side and a back side of the frame of the apparatus; rotating, by the first motor, the depth information capturing device and the image capturing device about a first substantially vertical axis until the image capturing device points in a second direction, and capturing, by the depth information capturing device that is rotating, depth information of a first portion of the 360-degree scene; capturing, by the image capturing device pointing in the second direction, a plurality of images at different exposures in a second FOV that overlaps the first FOV of the 360-degree scene; rotating, by the first motor, the depth information capturing device and the image capturing device about the first axis until the image capturing device points in a third direction, and capturing, by the depth information capturing device that is rotating, depth information of a second portion of the 360-degree scene; and capturing, by the image capturing device pointing in the third direction, a plurality of images at different exposures in a third FOV that overlaps the second FOV of the 360-degree scene.
2. The method of claim 1, further comprising: rotating, by the first motor, the depth information capturing device and the image capturing device about the first axis until the image capturing device points in a fourth direction; and capturing, by the image capturing device pointing in the fourth direction, a plurality of images at different exposures in a fourth FOV that overlaps the first FOV and the third FOV of the 360-degree scene.
3. The method of claim 2, further comprising: capturing, by the depth information capturing device, depth information of a third portion of the 360-degree scene while rotating, by the first motor, the depth information capturing device and the image capturing device about the first axis until the image capturing device points in the fourth direction.
4. The method of claim 3, wherein capturing, by the depth information capturing device, depth information of the first portion, the second portion, and the third portion of the 360-degree scene comprises capturing, by the depth information capturing device, depth information of a first plurality of segments, a second plurality of segments, and a third plurality of segments of the 360-degree scene.
5. The method of claim 1, further comprising: stitching, by an image stitching module, one or more of the plurality of images in the first FOV, the second FOV, and the third FOV together into a panoramic image of the 360-degree scene; and combining, by a 3D image generation module, the depth information of the first portion, the second portion, and the third portion of the 360-degree scene with the panoramic image of the 360-degree scene to produce a 3D panoramic image of the 360-degree scene.
6. The method of claim 5, further comprising: mixing, by a mixing module, two or more of the plurality of images in the first FOV, the second FOV, and the third FOV into first, second, and third mixed images of the first FOV, the second FOV, and the third FOV, wherein stitching, by the image stitching module, one or more of the plurality of images in the first FOV, the second FOV, and the third FOV together into the panoramic image of the 360-degree scene comprises stitching the first, second, and third mixed images together into the panoramic image of the 360-degree scene.
7. The method of claim 6, wherein mixing, by the mixing module, two or more of the plurality of images in the first FOV, the second FOV, and the third FOV into the first, second, and third mixed images of the first FOV, the second FOV, and the third FOV comprises mixing two or more of the plurality of images in the first FOV, the second FOV, and the third FOV into first, second, and third high dynamic range images of the first FOV, the second FOV, and the third FOV, wherein the first, second, and third high dynamic range images are first, second, and third HDR images; and wherein stitching, by the image stitching module, one or more of the plurality of images in the first FOV, the second FOV, and the third FOV together into the panoramic image of the 360-degree scene comprises stitching the first, second, and third HDR images in the first FOV, the second FOV, and the third FOV together into the panoramic image of the 360-degree scene.
8. The method of claim 6, wherein the apparatus further comprises a second motor, wherein the depth information capturing device comprises a light detection and ranging (LiDAR) device and a mirror, and the method further comprises rotating the mirror about a second substantially horizontal axis by the second motor and emitting a plurality of laser pulses by the LiDAR device to the mirror, which in turn emits the plurality of laser pulses in one or more revolutions about the second axis and receives a corresponding plurality of reflected laser pulses, and the LiDAR device in turn receives the plurality of reflected laser pulses and generates therefrom the depth information, wherein the first, second and third mixed images of the first, second and third FOVs each comprise a plurality of pixels, wherein each of the plurality of pixels is associated with a digital coordinate identifying a location of that pixel in the scene, and wherein each of the reflected laser pulses is similarly associated with a corresponding digital coordinate identifying a location of the depth information generated therefrom, and the combining of the depth information of the 360-degree scene with the panoramic image of the 360-degree scene by the 3D image generation module to produce the 3D panoramic image of the 360-degree scene comprises combining the depth information at one location in the 360-degree scene with the panoramic image at the same location in the 360 scene to produce the 3D panoramic image of the 360-degree scene.
9. The method of claim 1, wherein the image capturing device comprises a lens located at a parallax-free point substantially at a center of the first axis, such that as the first motor rotates the image capturing device about the first axis to capture the plurality of images at different exposures in respective first, second and third overlapping FOVs while the image capturing device is pointing in the first, second and third directions, parallax effects between the plurality of images in the first, second and third overlapping FOVs are reduced or eliminated.
10. The method of claim 1, wherein the device further comprises a second motor, the depth information capturing device comprises a light detection and ranging (LiDAR) device and a mirror, and the method further comprises rotating the mirror about a second substantially horizontal axis by the second motor and emitting a plurality of laser pulses by the LiDAR device to the mirror, which in turn emits the plurality of laser pulses in one or more revolutions about the second axis and receives a corresponding plurality of reflected laser pulses, and the LiDAR device in turn receives the plurality of reflected laser pulses and generates therefrom the depth information, such that as the image capturing device and the LiDAR device are rotated about the first axis by the first motor and the mirror is rotated about the second axis by the second motor, the LiDAR device captures the depth information of the 360 scene in three dimensions.
11. A device that acquires images and depth information for use in a 360 degree scene, comprising: a first motor; an image capturing device that captures a plurality of images at different exposures in a first field of view of the 360 degree scene while pointing in a first direction, wherein the first field of view (FOV), the image capturing device comprises a lens coupled to a frame of the device at a first location between a front side and a back side of the frame of the device; and a depth information capturing device, wherein the depth information capturing device and the image capturing device are rotated about a first substantially vertical axis by the first motor until the image capturing device points in a second direction, the rotating depth information capturing device captures depth information of a first portion of the 360 degree scene, the image capturing device pointing in the second direction captures a plurality of images at different exposures in a second FOV that overlaps the first FOV of the 360 degree scene, the depth information capturing device and the image capturing device are rotated about the first axis by the first motor until the image capturing device points in a third direction, the rotating depth information capturing device captures depth information of a portion of the 360 degree scene, and the image capturing device pointing in the third direction captures a plurality of images at different exposures in a third FOV that overlaps the second FOV of the 360 degree scene.
12. The device of claim 11, wherein the depth information capturing device and the image capturing device are rotated about the first axis by the first motor until the image capturing device points in a fourth direction; and while pointing in the fourth direction, the image capturing device captures a plurality of images at different exposures in a fourth FOV that overlaps the first FOV and the third FOV of the 360 degree scene.
13. The apparatus of claim 12, wherein the depth information capture device captures depth information of a third portion of the 360-degree scene while being rotated by the first motor about the first axis until the image capture device is pointed in the fourth direction.
14. The apparatus of claim 13, wherein the depth information capture device being rotated captures depth information of the first portion, second portion, and third portion of the 360-degree scene comprises the depth information capture device being rotated capturing depth information of a first plurality of segments, a second plurality of segments, and a third plurality of segments of the 360-degree scene.
15. The apparatus of claim 11, further comprising: an image stitching module that stitches one or more of the plurality of images in the first FOV, the second FOV, and the third FOV together into a panoramic image of the 360-degree scene; and a 3D image generation module that combines the depth information of the first portion, the second portion, and the third portion of the 360-degree scene with the panoramic image of the 360-degree scene to produce a 3D panoramic image of the 360-degree scene.
16. The apparatus of claim 15, further comprising: a blending module that blends two or more of the plurality of images in the first FOV, the second FOV, and the third FOV into first, second, and third blended images of the first FOV, the second FOV, and the third FOV; wherein the image stitching module stitches one or more of the plurality of images in the first FOV, the second FOV, and the third FOV together into the panoramic image of the 360-degree scene comprises the image stitching module stitching the first, second, and third blended images together into the panoramic image of the 360-degree scene.
17. The apparatus of claim 16, wherein the mixing of two or more of the plurality of images in the first FOV, the second FOV, and the third FOV into the first mixed image, the second mixed image, and the third mixed image of the first FOV, the second FOV, and the third FOV by the mixing module comprises the mixing of two or more of the plurality of images in the first FOV, the second FOV, and the third FOV into a first high dynamic range image, a second high dynamic range image, and a third high dynamic range image of the first FOV, the second FOV, and the third FOV by the mixing module, wherein the first high dynamic range image, the second high dynamic range image, and the third high dynamic range image are first, second, and third high dynamic range images, HDR images, respectively; and the stitching of one or more of the plurality of images in the first FOV, the second FOV, and the third FOV together into the panoramic image of the 360-degree scene by the image stitching module comprises the stitching of the first HDR image, the second HDR image, and the third HDR image in the first FOV, the second FOV, and the third FOV together into the panoramic image of the 360-degree scene by the image stitching module.
18. The apparatus of claim 16, further comprising a second motor, wherein the depth information capturing device comprises a light detection and ranging device, LiDAR device, and a mirror, the second motor rotates the mirror about a second substantially horizontal axis while the LiDAR device emits a plurality of laser pulses to the mirror, the mirror in turn emits the plurality of laser pulses and receives a corresponding plurality of reflected laser pulses in one or more revolutions about the second axis, and the LiDAR device in turn receives the plurality of reflected laser pulses and generates the depth information therefrom, wherein the first mixed image, the second mixed image, and the third mixed image of the first FOV, the second FOV, and the third FOV each comprise a plurality of pixels, wherein each of the plurality of pixels is associated with a digital coordinate identifying a location of that pixel in the scene, and wherein each of the reflected laser pulses is similarly associated with a corresponding digital coordinate identifying a location of the depth information generated therefrom, and the combining of the depth information of the 360-degree scene with the panoramic image of the 360-degree scene by the 3D image generation module to produce the 3D panoramic image of the 360-degree scene comprises the combining of the depth information at one location in the 360-degree scene with the panoramic image at the same location in the 360 scene to produce the 3D panoramic image of the 360-degree scene.
19. The apparatus of claim 11, wherein the image capture device comprises a lens positioned at a parallax-free point substantially at a center of the first axis, such that as the first motor rotates the image capture device about the first axis to capture the multiple images at different exposures in respective first, second, and third overlapping FOVs while the image capture device is pointed in the first, second, and third directions, parallax effects between the multiple images in the first, second, and third overlapping FOVs are reduced or eliminated.
20. The apparatus of claim 11, further comprising a second motor, wherein the depth information capture device comprises a light detection and ranging (LiDAR) device and a mirror, the mirror is rotated about a second substantially horizontal axis by the second motor while the LiDAR device emits a plurality of laser pulses to the mirror, the mirror in turn emits the plurality of laser pulses and receives a corresponding plurality of reflected laser pulses in one or more revolutions about the second axis, the LiDAR device in turn receives the plurality of reflected laser pulses and generates the depth information therefrom, such that as the image capture device and the LiDAR device are rotated about the first axis by the first motor and the mirror is rotated about the second axis by the second motor, the LiDAR device captures the depth information of the 360 scene in three dimensions.
Citation Information
Patent Citations
Image capturing apparatus for enabling generation of data of panoramic image with wide dynamic range
CN102739955A
Apparatus and method for capturing an area in 3D
US20100134596A1
Capturing and aligning panoramic image and depth data
US20190394441A1