System and method of capturing and generating panoramic three-dimensional images
Patent Information
- Application Number
- HK42026126871
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-30
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2040-12-29
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202511699345.5 (22) Application Date 2020.12.30 (30) Priority Data 62 / 955,414 2019.12.30 US (62) Divisional Application Data 202080084506.9 2020.12.30 (71) Applicant Matt Porter Inc. Address California, USA (72) Inventors D.A. Gossebeck K. Stromberg L.D. Marzano D. Proctor N. Sakakibara S. Deliyu K. Kane S. Wayne (74) Patent Agency Beijing Jikai Intellectual Property Agency Co., Ltd. 11245 Patent Attorney Xu Dongsheng (51) Int.Cl. H04N 23 / 58 (2023.01) H04N 23 / 55(2023.01) H04N 13 / 271(2018.01) H04N 13 / 254(2018.01) H04N 13 / 221(2018.01) H04N 23 / 54(2023.01) H04N 23 / 51(2023.01) G03B 37 / 00(2021.01) G03B 37 / 02(2021.01) G01S 17 / 89(2020.01) G01S 17 / 86(2020.01) G01S 17 / 42(2006.01) H04N 23 / 695(2023.01) G01S 7 / 481(2006.01) G01S 17 / 931(2020.01) G01S 17 / 894(2020.01) G03B 17 / 56(2021.01) H04N 23 / 698(2023.01) (54) Invention Title System and Method for Capturing and Generating Panoramic 3D Images (57) Abstract A system and method for capturing and generating panoramic 3D images are disclosed. A device includes a housing, a mounting, a wide-angle lens, an image capture device, and a LiDAR device. The mounting is configured to be coupled to a motor for horizontal movement of the device. The wide-angle lens is coupled to the housing and positioned above the mounting along a rotation axis, which is the axis of rotation of the device. The image capture device is located within the housing and configured to receive a two-dimensional image of the environment through the wide-angle lens. The LiDAR device is located within the housing and configured to generate depth data based on the environment. (Claims: 2 pages; Description: 35 pages; Drawings: 16 pages; CN 121531224 A 2026.02.13 CN 1 21 53)12 24 A 1. An apparatus for acquiring images and depth information used in an environment, the apparatus comprising: a first motor; an image capturing device configured to capture, while pointing in a first direction within the environment, a first plurality of images at different exposures within the field of view of the image capturing device; and a depth information capturing device; the first motor being configured to rotate the depth information capturing device and the image capturing device about a first axis until the image capturing device points in a second direction, the depth information capturing device being configured to capture first depth information of a first portion of the environment while being rotated, the image capturing device being further configured to capture, while pointing in the second direction within the environment, a second plurality of images at different exposures that at least partially overlap with one or more of the first plurality of images, the first motor being further configured to rotate the depth information capturing device and the image capturing device about the first axis until the image capturing device points in a third direction, the depth information capturing device being further configured to capture, while being rotated, second depth information of a second portion of the environment, and the image capturing device being further configured to capture, while pointing in the third direction, a third plurality of images at different exposures that at least partially overlap with one or more of the second plurality of images. 2. The device of claim 1, wherein the depth information capturing device and the image capturing device are configured to rotate about the first axis by the first motor until the image capturing device points to a fourth direction, and while pointing to the fourth direction, the image capturing device is configured to capture a fourth plurality of images at different exposures in the fourth direction, the fourth plurality of images at different exposures captured in the fourth direction at least partially overlapping one or more of the first plurality of images and the third plurality of images. 3. The device of claim 2, wherein the depth information capturing device is configured to capture third depth information of a third portion of the environment while rotating about the first axis by the first motor until the image capturing device points to the fourth direction. 4. The device of claim 3, wherein the depth information capturing device that captures depth information of the first, second, and third portions of the environment while being rotated includes a depth information capturing device that captures depth information of a first plurality of segments, a second plurality of segments, and a third plurality of segments of the environment while being rotated. 5. The device of claim 1, further comprising: a communication module configured to provide one or more of the plurality of images of the first, second, and third portions of the environment to a digital device to form a panoramic image of the environment, the digital device being configured to provide one or more of the plurality of images of the first plurality of images, the second plurality of images, and the third plurality of images to a digital device to form a panoramic image of the environment.6. The device of claim 5, further comprising: the digital device being configured to blend two or more of the first plurality of images, the second plurality of images, and the third plurality of images to generate a first blended image, a second blended image, and a third blended image, wherein the digital device is configured to blend one or more of the first plurality of images, the second plurality of images, and the third plurality of images to form the panoramic image of the environment, including blending the first blended image, the second blended image, and the third blended image to form the panoramic image of the environment. 7. The device according to claim 1, further comprising: a second motor, wherein the depth information acquisition device includes a photodetector and ranging device, i.e., a LiDAR device, and a reflector, the second motor being configured to cause the reflector to rotate about a second axis while the LiDAR device emits a plurality of laser pulses toward the reflector, the reflector being configured to provide the plurality of laser pulses and receive corresponding plurality of reflected laser pulses in one or more rotations about the second axis, and the LiDAR device being further configured to receive the plurality of reflected laser pulses and generate the depth information therefrom. 8. The device of claim 6, further comprising: a second motor, wherein the depth information capturing device includes a light detection and ranging device, i.e., a LiDAR device, and a reflector, the second motor being configured to cause the reflector to rotate about a second axis while the LiDAR device emits a plurality of laser pulses toward the reflector, the reflector being configured to provide the plurality of laser pulses and receive corresponding plurality of reflected laser pulses in one or more rotations about the second axis, and the LiDAR device being further configured to receive the plurality of reflected laser pulses and generate the depth information therefrom, wherein the first mixed image, the second mixed image, and the third mixed image each include a plurality of pixels, wherein each of the plurality of pixels is associated with digital coordinates of the position of that pixel identifying the environment, and wherein each of the reflected laser pulses is similarly associated with corresponding digital coordinates of the position identifying the depth information generated therefrom, and the digital device being further configured to combine the depth information of the environment with the panoramic image of the environment to generate the 3D panoramic image of the environment using at least one set of the digital coordinates of the plurality of pixels and at least one set of the digital coordinates of the depth information. 9. The device of claim 1, wherein the image capturing device includes a lens substantially located within the...A parallax-free point at or near the center of the first axis, such that as the first motor rotates the image capturing device about the first axis to capture images while the image capturing device is pointing towards the first direction, the second direction, and the third direction, parallax effects between the first plurality of images, the second plurality of images, and the third plurality of images are reduced or eliminated. 10. The device of claim 1, further comprising: a second motor, wherein the depth information capturing device includes a light detection and ranging device, i.e., a LiDAR device, and a mirror, the second motor being configured to cause the mirror to rotate about a second axis while the LiDAR device emits a plurality of laser pulses toward the mirror, the mirror being configured to provide the plurality of laser pulses and receive corresponding plurality of reflected laser pulses in one or more rotations about the second axis, and the LiDAR device being further configured to receive the plurality of reflected laser pulses and generate the depth information therefrom, the device being further configured such that as the image capturing device and the LiDAR device rotate about the first axis via the first motor and the mirror rotates about the second axis via the second motor, the LiDAR device captures the depth information in three dimensions of the environment. Claims 2 / 2 Page 3 CN 121531224 A System and Method for Capturing and Generating Panoramic 3D Images
[0001] This application is a divisional application of Chinese Patent Application 202411960746.7, entitled "System and Method for Capturing and Generating Panoramic 3D Images," filed on December 30, 2020. Chinese Patent Application 202411960746.7 is a divisional application of Chinese Patent Application 202080084506.9 (PCT / US2020 / 067474), also entitled "System and Method for Capturing and Generating Panoramic 3D Images," filed on December 30, 2020. Technical Field
[0002] Embodiments of the present invention generally relate to capturing and stitching panoramic images of scenes in a physical environment. Background Art
[0003] The proliferation of providing three-dimensional (3D) panoramic images of the physical world has led to numerous solutions with the ability to capture multiple two-dimensional (2D) images and generate 3D images based on the captured 2D images. Hardware solutions and software applications (or "apps") exist that can capture multiple 2D images and stitch them together to form a panoramic image.
[0004] Techniques exist for capturing and generating 3D data from buildings. However, existing techniques typically cannot capture and generate 3D renderings of areas under bright lighting conditions. Areas of windows through which sunlight shines or floors or walls under bright lighting conditions often appear as holes in 3D renderings, which may require additional post-production work to fill in. This increases the complexity of the 3D rendering.Turnaround time and authenticity. Furthermore, outdoor environments present challenges for many existing 3D capture devices because structured light may not be used to capture 3D images.
[0005] Other limitations of existing technologies for capturing and generating 3D data include the amount of time required to capture and process the digital images needed to generate 3D panoramic images. Summary of the Invention
[0006] An example device includes a housing, a mounting, a wide-angle lens, an image capture device, and a LiDAR device, the mounting being configured to be coupled to a motor to move the device horizontally, the wide-angle lens being coupled to the housing and positioned above the mounting along a rotation axis, the rotation axis being the axis along which the device rotates when coupled to the motor, the image capture device being within the housing and configured to receive a two-dimensional image of the environment through the wide-angle lens, and the LiDAR device being within the housing and configured to generate depth data based on the environment.
[0007] An image capture device may include a housing, a first motor, a wide-angle lens, an image sensor, a mounting, a LiDAR, a second motor, and a reflector. The housing may have a front side and a rear side. A first motor may be coupled to the housing at a first position between the front and rear sides, the first motor being configured to rotate the image capturing device horizontally about a vertical axis by approximately 270 degrees. A wide-angle lens may be coupled to the housing at a second position along the vertical axis between the front and rear sides, the second position being a parallax-free point, the wide-angle lens having a field of view remote from the front side of the housing. An image sensor may be coupled to the housing and configured to generate an image signal from light received by the wide-angle lens. A mounting may be coupled to the first motor. A LiDAR may be coupled to the housing at a third position, the LiDAR being configured to generate laser pulses and generate a depth signal. A second motor may be coupled to the housing. The reflector may be coupled to a second motor as described on page 1 / 35 of the specification, CN 121531224 A, the second motor being configured to rotate the reflector about a horizontal axis. The reflector includes an angled surface configured to receive the laser pulses from the LiDAR and guide the laser pulses about the horizontal axis.
[0008] In some embodiments, the image sensor is configured to generate a first plurality of images with different exposures when the image capturing device is stationary and pointing in a first direction. The first motor may be configured to rotate the image capturing device about the vertical axis after the first plurality of images are generated. In various embodiments, the image sensor does not generate an image when the first motor rotates the image capturing device, and wherein when the first motor rotates the image capturing device...When the image capturing device is described, the LiDAR generates a depth signal based on the laser pulse. The image sensor can be configured to generate a second plurality of images with different exposures when the image capturing device is stationary and pointing in a second direction, and the first motor is configured to rotate the image capturing device 90 degrees about the vertical axis after generating the second plurality of images. The image sensor can be configured to generate a third plurality of images with different exposures when the image capturing device is stationary and pointing in a third direction, and the first motor is configured to rotate the image capturing device 90 degrees about the vertical axis after generating the third plurality of images. The image sensor can be configured to generate a fourth plurality of images with different exposures when the image capturing device is stationary and pointing in a fourth direction, and the first motor is configured to rotate the image capturing device 90 degrees about the vertical axis after generating the fourth plurality of images.
[0009] In some embodiments, the system may further include a processor configured to mix frames of the first plurality of images before the image sensor generates the second plurality of images. A remote digital device (RDD) can communicate with the image capture device and is configured to generate a 3D visualization based on the first, second, third, and fourth plurality of images and the depth signal. The RDD is configured to generate the 3D visualization using no more than the first, second, third, and fourth plurality of images. In some embodiments, the first, second, third, and fourth plurality of images are generated between rotations, which combine rotations that cause the image capture device to rotate 270 degrees about the vertical axis. The speed or rotation of the reflector about the horizontal axis increases as the first motor rotates the image capture device. The angled surface of the reflector can be 90 degrees. In some embodiments, the LiDAR emits the laser pulses in a direction opposite to the front side of the housing.
[0010] An example method includes receiving light from a wide-angle lens of an image capturing device, the wide-angle lens being coupled to a housing of the image capturing device, the light being received at the field of view of the wide-angle lens extending away from the front side of the housing, generating a first plurality of images by an image sensor of the image capturing device using the light from the wide-angle lens, the image sensor being coupled to the housing, the first plurality of images being at different exposures, rotating the image capturing device horizontally about a vertical axis by a first motor approximately 270 degrees, the first motor being coupled to the housing in a first position between the front and rear sides of the housing, the wide-angle lens being in a second position along the vertical axis, the second position being parallax-free, and rotating a mirror having an angled surface about a horizontal axis by a second motor.The image capture device is rotated, the second motor is coupled to the housing, a laser pulse is generated by the LiDAR at a third position coupled to the housing, the laser pulse is guided to the rotating mirror as the image capture device rotates horizontally, and a depth signal is generated by the LiDAR based on the laser pulse.
[0011] The generation of the first plurality of images by the image sensor may occur before the image capture device rotates horizontally. In some embodiments, the image sensor does not generate images when the first motor rotates the image capture device, and wherein the LiDAR generates the depth signal based on the laser pulse when the first motor rotates the image capture device. Specification 2 / 35 pages 5 CN 121531224 A
[0012] The method may further include generating a second plurality of images by the image sensor with the different exposures when the image capture device is stationary and pointing in a second direction, and after generating the second plurality of images, rotating the image capture device 90 degrees about the vertical axis by the first motor.
[0013] In some embodiments, the method may further include generating a third plurality of images by the image sensor at the different exposures when the image capturing device is stationary and pointing in a third direction, and after generating the third plurality of images, rotating the image capturing device 90 degrees about the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor at the different exposures when the image capturing device is stationary and pointing in a fourth direction. The method may include generating a 3D visualization using the first, second, third, and fourth plurality of images and based on the depth signal, wherein generating the 3D visualization does not use any other images.
[0014] In some embodiments, the method may further include blending frames of the first plurality of images before the image sensor generates the second plurality of images. The first, second, third, and fourth plurality of images may be generated between rotations that combine a rotation of the image capturing device about 270 degrees about the vertical axis. In some embodiments, the speed or rotation of the reflector about the horizontal axis increases as the first motor rotates the image capturing device. Brief Description of the Drawings
[0015] FIG1a depicts a dollhouse view of an example environment (such as a house) according to some embodiments.
[0016] FIG1b depicts a floor plan view of the first floor of a house according to some embodiments.
[0017] FIG2 depicts an example eye-level view of a living room, which may be part of a virtual walkthrough.
[0018] FIG3 depicts an example of an environmental capture system according to some embodiments.
[0019] FIG4 depicts a perspective view of an environmental capture system according to some embodiments.
[0020] Figure 5 is a depiction of laser pulses from a LiDAR around an environment capture system in some embodiments.
[0021] Figure 6A depicts a side view of the environment capture system.
[0022] Figure 6B depicts a view from above the environment capture system in some embodiments.
[0023] Figure 7 depicts a perspective view of components of an example environment capture system according to some embodiments.
[0024] Figure 8A depicts example lens dimensions in some embodiments.
[0025] Figure 8B depicts example lens design specifications in some embodiments.
[0026] Figure 9A depicts a block diagram of an example environment capture system according to some embodiments.
[0027] Figure 9B depicts a block diagram of an example SOM PCBA of an environment capture system according to some embodiments.
[0028] Figures 10a-10c depict the process of an environment capture system for capturing images in some embodiments.
[0029] Figure 11 depicts a block diagram of an example environment capable of capturing and stitching images to form a 3D visualization according to some embodiments.
[0030] Figure 12 is a block diagram of an example alignment and stitching system according to some embodiments.
[0031] FIG13 depicts a flowchart of a 3D panoramic image capture and generation process according to some embodiments.
[0032] FIG14 depicts a flowchart of a 3D and panoramic capture and stitching process according to some embodiments.
[0033] FIG15 depicts a flowchart showing further details of one step of the 3D and panoramic capture and stitching process of FIG14.
[0034] FIG16 depicts a block diagram of an example digital device according to some embodiments. Specification 3 / 35 pages 6 CN 121531224 A Detailed Description
[0035] Many innovations described herein are made with reference to the accompanying drawings. The same reference numerals are used to refer to the same elements. In the following description, many specific details are set forth for purposes of explanation in order to provide a thorough understanding. However, it will be apparent, however, that different innovations may be practiced without these specific details. In other instances, well-known structures and components are shown in block diagram form to facilitate the description of innovations.
[0036] Various embodiments of the device provide users with 3D panoramic images of indoor and outdoor environments. In some embodiments, the device can efficiently and quickly provide users with 3D panoramic images of indoor and outdoor environments using a single wide field of view (FOV) lens and a single light detection and ranging sensor (LiDAR sensor).
[0037] The following are example uses of the example device described herein. The following use case is one of the embodiments. Different embodiments of the device discussed herein may include one or more features and capabilities similar to those of this use case.
[0038] FIG1a depicts a dollhouse view 100 of an example environment (e.g., a house) according to some embodiments. Dollhouse viewFigure 100 shows an overall view of an example environment captured by an environment capture system (discussed herein). A user can interact with the dollhouse view 100 on a user's system by switching between different views of the example environment. For example, a user can interact with area 110 to trigger a floor plan view of the first floor of the house, as shown in Figure 1b. In some embodiments, a user can interact with icons in the dollhouse view 100 (such as icons 120, 130, and 140) to provide a walkthrough view (e.g., for 3D walkthrough), a floor plan view, or a measurement view, respectively.
[0039] Figure 1b depicts a floor plan view of the first floor of a house according to some embodiments. The floor plan view is a top view of the first floor of the house. A user can interact with areas of the floor plan view (such as area 150) to trigger an eye-level view of a specific portion of the floor plan (such as the living room). An example of an eye-level view of the living room can be found in Figure 2, which can be part of a virtual walkthrough.
[0040] A user can interact with a portion of the floor plan 200 corresponding to area 150 of Figure 1b. Users can move the view around the room as if they were actually in the living room. In addition to a horizontal 360° view of the living room, users can also view or navigate the floor or ceiling of the living room. Furthermore, users can traverse the living room to reach other parts of the house by interacting with specific areas (such as areas 210 and 220) of this portion of floor plan 200. When a user interacts with area 220, the environment capture system can provide a walkable transition between areas of the house that substantially correspond to blocks of the house depicted by area 150 and areas of the house that substantially correspond to blocks of the house depicted by area 220.
[0041] Figure 3 depicts an example of an environment capture system 300 according to some embodiments. The environment capture system 300 includes a lens 310, a housing 320, a mounting attachment 330, and a movable cover 340.
[0042] When in use, the environment capture system 300 can be positioned in an environment such as a room. The environment capture system 300 can be positioned on a support (e.g., a tripod). The movable cover 340 can be moved to expose the LiDAR and the rotatable mirror. Once activated, the ambient capture system 300 can take a burst of images and then rotate using a motor. The ambient capture system 300 can open the mounting attachment 330. The LiDAR can make measurements while rotating (the ambient capture system can not take images while rotating). Once pointed in a new direction, the ambient capture system can take another burst of images before rotating to the next direction.
[0043] For example, once positioned, the user can command the ambient capture system 300 to start scanning. The scanning can be as follows: (1) Exposure estimation, and then taking HDR RGB images; rotate 90 degrees to capture depth data; (2) Exposure estimation, and then taking HDR images.RGB Image Manual 4 / 35 Page 7 CN 121531224 A Rotate 90 degrees to capture depth data (3) Exposure estimation, and then capture HDR RGB image Rotate 90 degrees to capture depth data (4) Exposure estimation, and then capture HDR RGB image Rotate 90 degrees (360 degrees in total) to capture depth data For each burst, there can be any number of images at different exposures. The ambient capture system can blend any number of images from the burst while waiting for another frame and / or waiting for the next burst.
[0044] The housing 320 can protect the electronic components of the ambient capture system 300 and can provide an interface for the user to interact with the power button, scan button, etc. For example, the housing 320 may include a removable cover 340, which may be removable to expose the LiDAR. In addition, the housing 320 may include electronic interfaces such as a power adapter and indicator lights. In some embodiments, the housing 320 is a molded plastic housing. In various embodiments, the housing 320 is one or more combinations of plastic, metal and polymer.
[0045] The lens 310 may be part of a lens assembly. Further details of the lens assembly can be described in the description of Figure 7. Lens 310 is strategically placed at the center of the rotation axis 305 of the environment capture system 300. In this example, the rotation axis 305 is in the x-y plane. By placing lens 310 at the center of the rotation axis 305, parallax effects can be eliminated or reduced. Parallax is an error caused by the rotation of the image capture device around a point that is not a non-parallax point (NPP). In this example, the NPP can be found at the center of the entrance pupil of the lens.
[0046] For example, suppose a panoramic image of the physical environment is generated based on four images captured by the environment capture system 300, wherein the images of the panoramic image have 25% overlap. If there is no parallax, then 25% of one image can completely overlap with another image of the same area of the physical environment. Eliminating or reducing the parallax effect of multiple images captured by the image sensor through lens 310 can help stitch multiple images into a 2D panoramic image.
[0047] Lens 310 may include a large field of view (e.g., lens 310 may be a fisheye lens). In some embodiments, the lens may have a horizontal field of view (FOV) of at least 148 degrees and a vertical field of view (FOV) of at least 94 degrees.
[0048] The mounting attachment 330 may allow the environmental capture system 300 to be attached to a mounting. The mounting may allow the environmental capture system 300 to be coupled to a tripod, a flat surface, or a motorized mounting (e.g., to move the environmental capture system 300). In some embodiments, the mounting may allow the environmental capture system 300 to rotate along a horizontal axis.
[0049] In some embodiments, the environment capture system 300 may include a motor for horizontally rotating the environment capture system 300 about a mounting attachment 330.
[0050] In some embodiments, the motorized mount may move the environment capture system 300 along a horizontal axis, a vertical axis, or both. In some embodiments, the motorized mount may rotate or move in the x-y plane. The use of the mounting attachment 330 may allow the environment capture system 300 to be coupled to the motorized mount, tripod, etc., to stabilize the environment capture system 300 to reduce or minimize swaying. In another example, the mounting attachment 330 may be coupled to a motorized mount that allows the 3D and environment capture system 300 to rotate at a stable, known speed, which helps the LiDAR determine the (x, y, z) coordinates of each laser pulse of the LiDAR.
[0051] Figure 4 depicts a rendered image of the environment capture system 400 in some embodiments. The renderings show the environment capture system 400 (which may be an example of the environment capture system 300 of FIG. 3) from various views such as front view 410, top view 420, side view 430, and rear view 440. In these renderings, the environment capture system 400 may include an optional hollow portion depicted in side view 430.
[0052] In some embodiments, the environment capture system 400 has a width of 75 mm, a height of 180 mm, and a depth of 189 mm. It should be understood that the environment capture system 400 may have any width, height, or depth. In various embodiments, the width-to-height-to-depth ratio of the first example is maintained regardless of the specific measurements.
[0053] The housing of the 3D and environment capture system 400 may protect the electronic components of the environment capture system 400 and may provide an interface / interface for user interaction (e.g., a screen on rear view 440). In addition, the housing may include electronic interfaces such as a power adapter and indicator lights. In some embodiments, the housing is a molded plastic housing. In various embodiments, the housing is a combination of one or more of plastic, metal, and polymer. The environmental capture system 400 may include a removable cover that is movable to expose the LiDAR and protect the LiDAR from components when not in use.
[0054] The lens depicted in front view 410 may be part of a lens assembly. Similar to the environmental capture system 300, the lens of the environmental capture system 400 is strategically positioned at the center of the rotation axis. The lens may include a large field of view. In various embodiments, the lens depicted in front view 410 is recessed, and the housing is flared, such that the wide-angle lens is directly above the parallax-free point (e.g., directly above the midpoint of the mount and / or motor), but can still capture images without interference from the housing.
[0055] A mounting attachment at the base of the environment capture system 400 allows the environment capture system to be attached to a mounting. The mounting allows the environment capture system 400 to be coupled to a tripod, a flat surface, or a motorized mount (e.g., to move the environment capture system 400). In some embodiments, the mounting may be coupled to an internal motor for rotating the environment capture system 400 about the mounting.
[0056] In some embodiments, the mounting allows the environment capture system 400 to rotate along a horizontal axis. In various embodiments, the motorized mount can move the environment capture system 400 along a horizontal axis, a vertical axis, or both. The use of the mounting attachment allows the environment capture system 400 to be coupled to a motorized mount, tripod, etc., to stabilize the environment capture system 400 to reduce or minimize swaying. In another example, the mounting attachment may be coupled to a motorized mount that allows the environment capture system 400 to rotate at a stable, known speed, which helps the LiDAR determine the (x, y, z) coordinates of each laser pulse of the LiDAR.
[0057] In view 430, a reflector 450 is exposed. The LiDAR can emit laser pulses to the reflector (in a direction opposite to the lens view). The laser pulses can strike the reflector 450, which can be angled (e.g., at a 90-degree angle). The reflector 450 can be coupled to an internal motor that rotates the reflector, such that the LiDAR's laser pulses can be emitted and / or received around the environment capture system 400 at many different angles.
[0058] FIG5 is a depiction of laser pulses from a LiDAR around the environment capture system 400 in some embodiments. In this example, the laser pulses are emitted at the rotating reflector 450. The laser pulses can be emitted and received perpendicular to the horizontal axis 602 of the environment capture system 400 (see FIG6). The reflector 450 can be angled such that the laser pulses from the LiDAR are directed away from the environment capture system 400. In some examples, the angle of the angled surface of the reflector may be 90 degrees, or between 60 and 120 degrees.
[0059] In some embodiments, when the environment capture system 400 is stationary and in operation, the environment capture system 400 may capture a burst of images through the lens. The environment capture system 400 may activate a horizontal motor between bursts of images. As it rotates along the mounting, the LiDAR of the environment capture system 400 may emit and / or receive laser pulses that strike the rotating reflector 450. The LiDAR may reflect the received laser pulses to generate a depth signal and / or generate depth data.
[0060] In some embodiments, the depth data may be associated with coordinates relative to the environment capture system 400. Similarly, pixels or portions of an image may be associated with coordinates relative to the environment capture system 400 to enable 3D visualization (e.g., to...The generation of images (from different directions, 3D walkthroughs, etc.) can be achieved using image and depth data.
[0061] As shown in FIG5, the LiDAR pulse can be blocked by the bottom of the environment capture system 400. It should be understood that while the environment capture system 400 moves around the mount, the reflector 450 may rotate uniformly, or the reflector 450 may rotate more slowly when the environment capture system 400 starts to move and, in addition, when the environment capture system 400 slows down to stop (e.g., maintaining a constant speed between the start and stop of the mount motor).
[0062] The LiDAR can receive depth data from the pulse. Due to the movement of the environment capture system 400 and / or the increase or decrease in the speed of the reflector 450, the density of the depth data of the environment capture system 400 may be inconsistent (e.g., denser in some areas and less dense in others).
[0063] FIG6a depicts a side view of the environment capture system 400. In this view, a reflector 450 is depicted, and the reflector 450 can rotate about a horizontal axis. A pulse 604 can be emitted by the LiDAR at the rotating reflector 450 and can be emitted perpendicular to the horizontal axis 602. Similarly, the pulse 604 can be received by the LiDAR in a similar manner.
[0064] Although the LiDAR pulse is discussed as being perpendicular to the horizontal axis 602, it should be understood that the LiDAR pulse can be at any angle relative to the horizontal axis 602 (e.g., the reflector angle can be any angle including 60 degrees to 120 degrees). In various embodiments, the LiDAR emits a pulse opposite to the front side (e.g., front side 604) of the environment capture system 400 (e.g., in a direction opposite to the center of the lens's field of view or toward the rear side 606).
[0065] As discussed herein, the environment capture system 400 can rotate about a vertical axis 608. In various embodiments, the environment capture system 400 captures images and then rotates 90 degrees, so that a fourth set of images is captured when the environment capture system 400 has completed a 270-degree rotation from the original starting position from which the first set of images was captured. Therefore, the environment capture system 400 can generate four sets of images between a total of 270 degrees of rotation (e.g., assuming the first set of images was captured before the initial rotation of the environment capture system 400). In various embodiments, images from a single sweep of the environment capture system 400 (e.g., four sets of images) (e.g., captured in a single full rotation or 270-degree rotation about a vertical axis) together with depth data acquired during the same sweep are sufficient to generate a 3D visualization without any additional sweeps or rotations of the environment capture system 400.
[0066] It should be understood that, in this example, the LiDAR pulses are generated from a position remote from the rotation point of the environment capture system 400.A rotating reflector emits and guides the pulse. In this example, the distance from the rotation point of the mount is 608 (e.g., the lens may be at a parallax-free point, while the lens may be positioned behind the lens relative to the front of the environment capture system 400). Since the LiDAR pulse is guided by the reflector 450 at a position off the rotation point, the LiDAR cannot receive depth data from a cylinder extending from above to below the environment capture system 400. In this example, the radius of the cylinder (e.g., a cylinder without depth information) can be measured from the center of the rotation point of the motor mount to the point where the reflector 450 guides the LiDAR pulse.
[0067] Furthermore, in FIG. 6b, a cavity 610 is depicted. In this example, the environment capture system 400 includes a rotating reflector within the body of the housing of the environment capture system 400. There is a cutout section from the housing. The laser pulse can be reflected out of the housing by the reflector, and the reflection can then be received by the reflector and guided back to the LiDAR so that the LiDAR can generate a depth signal and / or depth data. The base of the body of the environment capture system 400 below cavity 610 can block some laser pulses. Cavity 610 can be defined by the base of the environment capture system 400 and a rotating mirror. As shown in FIG6B, space may still exist between the edge of the angled mirror and the housing of the environment capture system 400 containing the LiDAR.
[0068] In various embodiments, the LiDAR is configured to stop emitting laser pulses if the rotational speed of the mirror drops below a rotational safety threshold (e.g., if there is a malfunction of the motor that rotates the mirror or the mirror is held in place). In this way, the LiDAR can be configured for safety and to reduce the likelihood that laser pulses will continue to be emitted in the same direction (e.g., at the user's eye).
[0069] FIG6b depicts a view from above the environment capture system 400 in some embodiments. In this example, the front of the environment capture system 400 is depicted, where the lens is recessed and directly above the center of the rotation point (e.g., above the center of the mounting specification page 7 / 35 of 10 CN 121531224 A). The front of the camera is recessed for the lens, and the front of the housing is flared to allow the field of view of the image sensor to be unobstructed. The reflector 450 is depicted pointing upwards.
[0070] FIG7 depicts a perspective view of components of an example of an environmental capture system 300 according to some embodiments. The environmental capture system 700 includes a front cover 702, a lens assembly 704, a structural frame 706, a LiDAR 708, a front housing 710, a reflector assembly 712, a GPS antenna 714, a rear housing 716, a vertical motor 718, a display 720, a battery pack 722, a mounting bracket 724, and a horizontal motor 726.
[0071] In various embodiments, the environmental capture system 700 can be configured to scan, align, and generate 3D meshes both outdoors and indoors under full sunlight. This eliminates the barriers of other systems that are used as indoor-only tools. The environmental capture system 700 can be able to scan large spaces faster than other devices. In some embodiments, the environmental capture system 700 can provide improved depth accuracy by improving single-scan depth accuracy at 90m.
[0072] In some embodiments, the environmental capture system 700 can weigh 1kg or about 1kg. In one example, the environmental capture system 700 can weigh 1-3kg.
[0073] The front cover 702, the front housing 710, and the rear housing 716 form part of the housing. In one example, the front cover can have a width w of 75mm.
[0074] The lens assembly 704 can include a camera lens that focuses light onto the image capture device. The image capture device can capture images of the physical environment. A user can place the environmental capture system 700 to capture a portion of a floor of a building (such as the second building 422 of FIG. 1) to obtain a panoramic image of that portion of the floor. The environmental capture system 700 can be moved to another part of a building floor to obtain a panoramic image of that other part of the floor. In one example, the depth of field of the image capture device is from 0.5 meters to infinity. Figure 8A depicts example lens sizes in some embodiments.
[0075] In some embodiments, the image capture device is a complementary metal-oxide-semiconductor (CMOS) image sensor (e.g., a Sony IMX283 ~ 20 Megapixel CMOS MIPI sensor with an NVIDIA Jetson Nano SOM). In various embodiments, the image capture device is a charge-coupled device (CCD). In one example, the image capture device is a red-green-blue (RGB) sensor. In one embodiment, the image capture device is an infrared (IR) sensor. The lens assembly 704 can give the image capture device a wide field of view.
[0076] The image sensor can have many different specifications. In one example, the image sensor includes the following: Specification 8 / 35 page 11 CN 121531224 A Example specifications may be as follows: Specification 9 / 35 page 12 CN 121531224 A In various embodiments, when observing the MTF at the F0 relative field (i.e., the center), for a total through focus shift of 53 micrometers, the focus shift can vary from +28 micrometers at 0.5m to -25 micrometers at infinity.
[0077] Figure 8B depicts example lens design specifications in some embodiments.
[0078] In some examples, the lens assembly 704 has an HFOV of at least 148 degrees and a VFOV of at least 94 degrees. In one exampleFor example, lens assembly 704 has a field of view of 150°, 180°, or within the range of 145° to 180°. In one example, an image capture of a 360° view around environment capture system 700 can be obtained using three or four separate image captures from the image capture device of environment capture system 700. In various embodiments, the image capture device may have a resolution of at least 37 pixels per degree. In some embodiments, environment capture system 700 includes a lens cap (not shown) to protect lens assembly 704 when not in use. The output of lens assembly 704 may be a digital image of a region of the physical environment. Images captured by lens assembly 704 can be stitched together to form a 2D panoramic image of the physical environment. A 3D panorama can be generated by combining depth data captured by LiDAR 708 with a 2D panoramic image generated by stitching together multiple images from lens assembly 704. In some embodiments, images captured by environment capture system 402 are stitched together by image processing system 406. In various embodiments, the environment capture system 402 generates a “preview” or “thumbnail” version of the 2D panoramic image. The preview or thumbnail version of the 2D panoramic image can be displayed on a user system 1110 such as an iPad, personal computer, smartphone, etc. In some embodiments, the environment capture system 402 can generate a mini-map of the physical environment representing an area of the physical environment. In various embodiments, the image processing system 406 generates the mini-map representing that area of the physical environment.
[0079] The image captured by the lens assembly 704 can include capture device location data that identifies or indicates the capture location of the 2D image. For example, in some embodiments, the capture device location data can include Global Positioning System (GPS) coordinates associated with the 2D image. In other embodiments, the capture device location data can include location information indicating the relative position of the capture device (e.g., a camera and / or 3D sensor) relative to its environment, such as the relative or calibrated position of the capture device relative to objects in the environment, another camera in the environment, another device in the environment, etc. In some implementations, this type of location data can be determined by a capture device associated with image capture (e.g., a camera and / or a device operatively coupled to the camera, including positioning hardware and / or software) and received along with the image. The placement of the lens assembly 704 is not merely a matter of design. By placing the lens assembly 704 at or substantially at the center of the rotation axis, parallax effects can be reduced.
[0080] In some embodiments, the structural frame 706 holds the lens assembly 704 and the LiDAR 708 in a specific position, andComponents that can help protect the example of the environment capture system. A structural frame 706 can be used to help rigidly mount the LiDAR 708 and place it in a fixed position. Furthermore, the fixed position of the lens assembly 704 and the LiDAR 708 allows the fixed relationship to align depth data with image information to help generate 3D images. 2D image data and depth data captured in the physical environment can be aligned relative to a common 3D coordinate space to generate a 3D model of the physical environment.
[0081] In various embodiments, the LiDAR 708 captures depth information of the physical environment. When a user places the environment capture system 700 in a section of a second building, the LiDAR 708 can obtain depth information of objects. The LiDAR 708 may include an optical sensing module capable of measuring the distance to a target or object in a scene by illuminating the target or scene with pulses from a laser, and measuring the time it takes for photons to travel to the target and return to the LiDAR 708. The measurement results can then be transformed into a grid coordinate system using information derived from the horizontal drive system of the environment capture system 700.
[0082] In some embodiments, the LiDAR 708 can return depth data points with timestamps (internal clock) every 10 microseconds. The LiDAR 708 can sample a portion of the sphere (the small holes at the top and bottom) every 0.25 degrees. In some embodiments, with one data point every 10 microseconds and 0.25 degrees, each point "disk" can have 14.40 milliseconds and 1440 disks to form a sphere with a nominal duration of 20.7 seconds. Because each disk is captured before and after, the sphere can be captured by scanning at 180°.
[0083] In one example, the LiDAR 708 specification can be as follows: Specification 11 / 35 pages 14 CN 121531224 A One advantage of using LiDAR is that, in the case of LiDAR at lower wavelengths (e.g., 905nm, 900-940nm, etc.), it allows the environmental capture system 700 to determine depth information of outdoor or indoor environments with bright light.
[0084] The placement of the lens assembly 704 and the LiDAR 708 allows the environment capture system 700 or a digital device communicating with the environment capture system 700 to use depth data from the LiDAR 708 and the lens assembly 704 to generate a 3D panoramic image. In some embodiments, as described on pages 12 / 35 of the specification 15 CN 121531224 A, 2D and 3D panoramic images are not generated on the environment capture system 402.
[0085] The output of the LiDAR 708 may include properties associated with each laser pulse transmitted by the LiDAR 708. PropertiesThis includes the intensity of the laser pulse, the number of returns, the current number of returns, the classification point, the RGC value, the GPS time, the scanning angle, the scanning direction, or any combination thereof. The depth of field can be (0.5m; infinity), (1m; infinity), etc. In some embodiments, the depth of field is 0.2m to 1m and infinity.
[0086] In some embodiments, while the environment capture system 700 is stationary, the environment capture system 700 uses the lens assembly 704 to capture four separate RBG images. In various embodiments, while the environment capture system 700 is in motion, moving from one RBG image capture location to another, the LiDAR 708 captures depth data in four different scenarios. In one example, a 360° rotation of the environment capture system 700 (which may be referred to as sweep) is used to capture a 3D panoramic image. In various embodiments, a rotation of less than 360° of the environment capture system 700 is used to capture a 3D panoramic image. The output of a sweep can be a sweep list (SWL), which includes image data from lens assembly 704 and depth data from LiDAR 708, as well as the nature of the sweep, including GPS location and a timestamp when the sweep occurred. In various embodiments, a single sweep (e.g., a single 360-degree rotation of environment capture system 700) captures sufficient image and depth information to generate a 3D visualization (e.g., via a digital device communicating with environment capture system 700, which receives image and depth data from environment capture system 700 and uses only the image and depth data from environment capture system 700 captured in a single sweep to generate a 3D visualization).
[0087] In some embodiments, images captured by environment capture system 402 can be blended, stitched together, and combined with depth data from LiDAR 708 using an image stitching and processing system discussed herein.
[0088] In various embodiments, applications on environment capture system 402 and / or user system 1110 can generate preview or thumbnail versions of 3D panoramic images. A preview or thumbnail version of the 3D panoramic image can be displayed on the user system 1110 and can have a lower image resolution than the 3D panoramic image generated by the image processing system 406. After the lens assembly 704 and LiDAR 708 capture image and depth data of the physical environment, the environment capture system 402 can generate a minimap representing an area of the physical environment that has been captured by the environment capture system 402. In some embodiments, the image processing system 406 generates a minimap representing that area of the physical environment. After capturing image and depth data of the living room of a home using the environment capture system 402, the environment capture system 402 can generate a top-down view of the physical environment. The user can use this information to determine areas of the physical environment that the user has not yet captured or generated a 3D panoramic image of.
[0089] In one embodiment, the environment capture system 700 may interleave image capture using an image capture device with lens assembly 704 with depth information capture using LiDAR 708. For example, the image capture device may capture an image of segment 1605 of the physical environment (as seen in FIG. 16), and then LiDAR 708 obtains depth information from segment 1605. Once LiDAR 708 has obtained depth information from segment 1605, the image capture device may continue to move to capture an image of another segment 1610, and then LiDAR 708 obtains depth information from segment 1610, thereby interleaving image capture and depth information capture.
[0090] In some embodiments, LiDAR 708 may have a field of view of at least 145°, and depth information of all objects in the 360° view of the environment capture system 700 may be obtained by the environment capture system 700 in three or four scans. In another example, the LiDAR 708 may have a field of view of at least 150°, 180°, or between 145° and 180°.
[0091] The increased field of view of the lens reduces the amount of time required to obtain visual and depth information about the physical environment surrounding the environment capture system 700. In various embodiments, the LiDAR 708 has a minimum depth range of 0.5 m. In one embodiment, the LiDAR 708 has a maximum depth range greater than 8 meters.
[0092] The LiDAR 708 may utilize the mirror assembly 712 to guide the laser at different scanning angles. In one embodiment, as described on pages 13 / 35 of CN 121531224 A, an optional vertical motor 718 has the capability to move the mirror assembly 712 vertically. In some embodiments, the mirror assembly 712 may be a dielectric mirror with a hydrophobic coating or layer. The mirror assembly 712 may be coupled to the vertical motor 718, which rotates the mirror assembly 712 during use.
[0093] The reflector of the reflector assembly 712 may, for example, include the following specifications: The reflector of the reflector assembly 712 may, for example, include the following specifications for materials and coatings: The hydrophobic coating of the reflector of the reflector assembly 712 may, for example, include a contact angle number >105.
[0094] The reflector of the reflector assembly 712 may include the following quality specifications: Specification 14 / 35 pages 17 CN 121531224 A The vertical motor may include, for example, the following specifications: Specification 15 / 35 pages 18 CN 121531224 A Due to the RGB capture device and LiDAR 708, the ambient capture system 700 can capture images outdoors in bright sunlight, or indoors in the presence of bright light or sunlight glare from windows. When using different devices (e.g.For example, in systems of structured light devices, they may not be able to operate in bright environments, whether indoors or outdoors. These devices are typically limited to indoor use only and only during dawn or sunset to control light. Otherwise, bright spots in the room produce artifacts or “holes” in the image that must be filled or corrected. However, the ambient capture system 700 can be used indoors and outdoors in bright sunlight. The capture device and LiDAR 708 can be able to capture images and depth data in bright environments without artifacts or holes caused by glare or bright light.
[0095] In one embodiment, the GPS antenna 714 receives Global Positioning System (GPS) data. The GPS data can be used to determine the location of the ambient capture system 700 at any given time.
[0096] In various embodiments, the display 720 allows the ambient capture system 700 to provide the current status of the system, such as update, warm-up, scan, scan complete, error, etc.
[0097] The battery pack 722 provides power to the ambient capture system 700. Battery pack 722 may be removable and rechargeable, allowing a user to insert a new battery pack 722 while charging a depleted battery pack. In some embodiments, battery pack 722 may allow continuous use of at least 1000 SWLs or at least 250 SWLs before recharging. The environment capture system 700 may utilize a USB-C plug for recharging.
[0098] In some embodiments, mounting 724 provides a connector for connecting the environment capture system 700 to a platform such as a tripod or mounting. Horizontal motor 726 may rotate the environment capture system 700 about the x-y plane. In some embodiments, horizontal motor 726 may provide information to a grid coordinate system to determine the (x, y, z) coordinates associated with each laser pulse. In various embodiments, horizontal motor 726 may enable the environment capture system 700 to scan rapidly due to the wide field of view of the lens, the positioning of the lens about the axis of rotation, and the LiDAR device.
[0099] In one example, the horizontal motor 726 may have the following specifications: Specification 16 / 35 pages 19 CN 121531224 A In various embodiments, the mounting component 724 may include a quick-release adapter. The holding torque may be, for example, > 2.0 Nm, and the durability of the capture operation may be up to or exceed 70,000 cycles.
[0100] For example, the environmental capture system 700 can realize the construction of a 3D grid of a standard home, wherein the distance between sweeps is greater than 8 m. The time for capturing, processing, and aligning indoor sweeps can be less than 45 seconds. In one example, the time frame from the start of sweep capture to when the user can move the environmental capture system 700 can be less than 15 seconds.
[0101] In various embodiments, these components provide the environmental capture system 700 with scan positions aligned with the outdoors and indoors.This enables the ability to create a seamless roaming experience between indoors and outdoors (which can be a high priority for hotels, vacation rentals, real estate, architectural documentation, CRE, and as-built modeling and verification). The environment capture system 700 can also generate an “outdoor dollhouse” or outdoor mini-map. As shown herein, the environment capture system 700 can also improve the accuracy of 3D reconstruction primarily from a measurement perspective. The user’s ability to tune the scan density can also be a favorable factor. These components also enable the environment capture system 700 to capture wide blank spaces (e.g., longer ranges). To generate a 3D model with a wide blank space, the environment capture system may need to scan and capture 3D data and depth data from a larger range of distances than generating a 3D model with a smaller space.
[0102] In various embodiments, these components enable the environment capture system 700 to align the SWL and reconstruct the 3D model in a similar manner for both indoor and outdoor use. These components also enable the environment capture system 700 to perform geolocation of the 3D model (which can be easily integrated into Google Street View and help align outdoor panoramas if needed).
[0103] The image capture device of the ambient capture system 700 can provide DSLR-class images with printable quality at 8.5" x 11" for 70° VFOV and RGB image types.
[0104] In some embodiments, the ambient capture system 700 can capture RGB images using the image capture device (e.g., using a wide-angle lens) and then move the lens (a total of four moves using a motor) before capturing the next RGB image. While the horizontal motor 726 rotates the ambient capture system 90 degrees, the LiDAR 708 can capture depth data. In some embodiments, the LiDAR 708 includes an APD array.
[0105] In some embodiments, the image and depth data can then be sent to a capture application (e.g., a device communicating with the ambient capture system 700, such as a smart device or an image capture system on a network). In some embodiments, the ambient capture system 700 can send the image and depth data to an image processing system 406 for processing and generating 2D or 3D panoramic images. In various embodiments, the environment capture system 700 can generate a sweep list of captured RGB images and depth data from a 360-degree rotation of the environment capture system 700. The sweep list can be sent to the image processing system 406 for stitching and alignment. The output of the sweep can be a sweep list (SWL), which includes image data from the lens assembly 704 and depth data from LiDAR 708 specification page 17 / 35 20 CN 121531224 A, as well as the nature of the sweep, including GPS location and a timestamp when the sweep occurred.
[0106] In various embodiments, the LiDAR, vertical mirror, RGB lens, tripod mount, and horizontal actuator are rigidly mounted within the housing to allow the housing to be opened without requiring system recalibration.
[0107] FIG9a depicts a block diagram 900 of an example environmental capture system according to some embodiments. Block diagram 900 includes a power supply 902, a power converter 904, an input / output (I / O) printed circuit board assembly (PCBA), a system-on-module (SOM) PCBA, a user interface 910, a LiDAR 912, a mirror brushless DC (BLCD) motor 914, a drivetrain 916, a wide FOV (WFOV) lens 918, and an image sensor 920.
[0108] The power supply 902 may be the battery pack 722 of FIG7. The power supply may be a removable, rechargeable battery, such as a lithium-ion battery (e.g., 4 x 18650 lithium-ion batteries), capable of providing power to the environmental capture system.
[0109] The power converter 904 can change the voltage level from the power supply 902 to a lower or higher voltage level so that it can be utilized by the electronic components of the environmental capture system. The environmental capture system can utilize 4 x 18650 lithium-ion batteries in a 4S1P configuration or a configuration of four series connections and one parallel connection.
[0110] In some embodiments, the I / O PCBA 906 may include elements providing an IMU, Wi-Fi, GPS, Bluetooth, an inertial measurement unit (IMU), a motor driver, and a microcontroller. In some embodiments, the I / O PCBA 906 includes a microcontroller for controlling a horizontal motor and encoding horizontal motor controls, and for controlling a vertical motor and encoding vertical motor controls.
[0111] The SOM PCBA 908 may include a central processing unit (CPU) and / or a graphics processing unit (GPU), memory, and a mobile interface. The SOM PCBA 908 can control the LiDAR 912, the image sensor 920, and the I / O PCBA 906. The SOM PCBA 908 can determine the (x, y, z) coordinates associated with each laser pulse of the LiDAR 912 and store the coordinates in a memory component of the SOM PCBA 908. In some embodiments, the SOM PCBA 908 can store the coordinates in the image processing system of the environment capture system 400. In addition to the coordinates associated with each laser pulse, the SOM PCBA 908 can also determine other attributes associated with each laser pulse, including the intensity of the laser pulse, the number of returns, the current number of returns, the classification point, the RGC value, the GPS time, the scan angle, and the scan direction.
[0112] In some embodiments, the SOM PCBA 908 includes an Nvidia SOM PCBA with CPU / GPU, DDR, eMMC, and Ethernet.
[0113] User interface 910 may include physical buttons or switches that a user can interact with. The buttons or switches may provide functions such as turning the environmental capture system on and off, scanning the physical environment, etc. In some embodiments, user interface 910 may include a display, such as display 720 of FIG. 7.
[0114] In some embodiments, LiDAR 912 captures depth information of the physical environment. LiDAR 912 includes an optical sensing module capable of measuring the distance to a target or object in a scene by illuminating the target or scene with light using pulses from a laser. The optical sensing module of LiDAR 912 measures the time it takes for photons to travel to the target or object and return to the receiver in LiDAR 912 after reflection, thus giving the distance of LiDAR to the target or object. Along with the distance, SOM PCBA 908 can determine the (x, y, z) coordinates associated with each laser pulse. LiDAR 912 can be mounted within a width of 58 mm, a height of 55 mm, and a depth of 60 mm.
[0115] The LiDAR 912 may include a range of 90m (10% reflectivity), a range of 130m (20% reflectivity), a range of 260m (100% reflectivity), a range accuracy of 2cm (1σ@900m), a wavelength of 1705nm, and a beam divergence of 0.28 × 0.03 degrees.
[0116] The SOM PCBA 908 may determine coordinates based on the position of the drivetrain 916. In various embodiments, the LiDAR 912 may include one or more LiDAR devices. Multiple LiDAR devices may be used to increase LiDAR resolution.
[0117] The mirror brushless DC (BLCD) motor 914 may control the mirror assembly 712 of FIG. 7.
[0118] In some embodiments, the drivetrain 916 may include the horizontal motor 726 of FIG. 7. When the environmental capture system is mounted on a platform such as a tripod, the drivetrain 916 may provide rotation of the environmental capture system. The drivetrain 916 may include a stepper motor Nema 14, a worm gear and plastic wheel drivetrain, a clutch, bushing bearings, and a backlash prevention mechanism. In some embodiments, the environmental capture system may be able to complete a scan in less than 17 seconds. In various embodiments, the drivetrain 916 has a maximum speed of 60 degrees / second, a maximum acceleration of 300 degrees / second², a maximum torque of 0.5 nm, an angular position accuracy of less than 0.1 degrees, and an encoder resolution of approximately 4096 counts per revolution.
[0119] In some embodiments, the drivetrain 916 includes a vertical monogon reflector and a motor. In this example, the drivetrain 916 may include a BLDC motor, an external Hall effect sensor, a magnet (paired with the Hall effect sensor), and a reflector.Mirror holder and reflector. In this example, the drivetrain 916 may have a maximum speed of 4,000 rpm and a maximum acceleration of 300 degrees / second^2. In some embodiments, the one-way reflector is a dielectric reflector. In one embodiment, the one-way reflector includes a hydrophobic coating or layer.
[0120] The components of the environmental capture system are positioned such that the lens assembly and the LiDAR are substantially centered on the axis of rotation. This can reduce image parallax that occurs when the image capture system is not centered on the axis of rotation.
[0121] In some embodiments, the WFOV lens 918 may be a lens of the lens assembly 704 of FIG. 7. The WFOV lens 918 focuses light onto the image capture device. In some embodiments, the WFOV lens may have an FOV of at least 145 degrees. With such a wide FOV, an image capture of a 360-degree view around the environmental capture system can be obtained using three separate image captures of the image capture device. In some embodiments, the WFOV lens 918 may have a diameter of approximately ~60 mm and a total lens length (TTL) of approximately ~80 mm. In one example, the WFOV lens 918 may include a horizontal field of view greater than or equal to 148.3 degrees and a vertical field of view greater than or equal to 94 degrees.
[0122] The image capturing device may include the WFOV lens 918 and an image sensor 920. The image sensor 920 may be a CMOS image sensor. In one embodiment, the image sensor 920 is a charge-coupled device (CCD). In some embodiments, the image sensor 920 is a red-green-blue (RGB) sensor. In one embodiment, the image sensor 920 is an IR sensor. In various embodiments, the image capturing device may have a resolution of at least 35 pixels per degree (PPD).
[0123] In some embodiments, the image capturing device may include an F-number of f / 2.4, an image circle diameter of 15.86 mm, a pixel pitch of 2.4 μm, an HFOV of >148.3°, a VFOV of >94.0°, a pixel per degree of >38.0 PPD, a principal ray angle of 3.0° at full height, a minimum object distance of 1300 mm, a maximum object distance of infinity, a relative illumination of >130%, a maximum distortion of <90%, and a spectral transmittance variation of <=5%.
[0124] In some embodiments, the lens may include an F-number of 2.8, an image circle diameter of 15.86 mm, a pixel per degree of >37, a principal ray angle of 3.0 at full height, an L1 diameter of <60 mm, a TTL of <80 mm, and a relative illumination of >50%.
[0125] The lens may include > 85% of 52 lp / mm (on-axis), > 66% of 104 lp / mm (on-axis), > 45% of 1308 lp / mm (on-axis), > 75% of 52 lp / mm (83% field), >41% of 104 lp / mm (83% field) and > 25% of 1308 lp / mm (83% field).
[0126] The environmental capture system may have a resolution of >20 MP, a green sensitivity of >65 dB (100 lux, 1x gain), and a dynamic range of >70 dB.
[0127] FIG9b depicts a block diagram of an example SOM PCBA 908 of an environmental capture system according to some embodiments. The SOM PCBA 908 may include a communication component 922, a LiDAR control component 924, a LiDAR positioning component 926, a user interface component 928, a classification component 930, a LiDAR data storage area 932, and a captured image data storage area 934.
[0128] In some embodiments, the communication component 922 may send and receive requests or data between any component of the SOM PCBA 1008 and the components of the environmental capture system of FIG9a.
[0129] In various embodiments, the LiDAR control component 924 can control various aspects of the LiDAR. For example, according to the LiDAR control specification 19 / 35 pages 22 CN 121531224 A, component 924 can send control signals to the LiDAR 912 to begin emitting laser pulses. The control signals sent by the LiDAR control component 924 may include instructions regarding the frequency of the laser pulses.
[0130] In some embodiments, the LiDAR positioning component 926 can utilize GPS data to determine the location of the environment capture system. In various embodiments, the LiDAR positioning component 926 utilizes the position of the mirror assembly to determine the scan angle and (x, y, z) coordinates associated with each laser pulse. The LiDAR positioning component 926 can also utilize the IMU to determine the orientation of the environment capture system.
[0131] The user interface component 928 can facilitate user interaction with the environment capture system. In some embodiments, the user interface component 928 can provide one or more user interface elements with which a user can interact. The user interface provided by the user interface component 928 can be sent to the user system 1110. For example, user interface component 928 can provide a user system (e.g., a digital device) with a visual representation of an area of a building's floor plan. The environmental capture system can generate a visual representation of the floor plan when the user places it in different sections of a building's floors to capture and generate 3D panoramic images. The user can place the environmental capture system in an area of the physical environment to capture and generate a 3D panoramic image of that area of the house. Once the 3D panoramic image of that area has been generated by the image processing system, the user interface component can update the floor plan view using a top view of the living room area depicted in Figure 1b. In some embodiments, the floor plan...The surface view 200 can be generated by the user system 1110 after a second scan of the same household or floor of the building that has already been captured.
[0132] In various embodiments, the classification component 930 can classify the type of physical environment. The classification component 930 can analyze objects in the image or objects in the image to classify the type of physical environment captured by the environment capture system. In some embodiments, the image processing system can be responsible for classifying the type of physical environment captured by the environment capture system 400.
[0133] The LiDAR data storage area 932 can be any and / or multiple structures suitable for the captured LiDAR data (e.g., active database, relational database, self-referenced database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, and / or similar structures). The image data storage area 408 can store the captured LiDAR data. However, in the absence of communication network 404, the LiDAR data storage area 932 can be used to cache the captured LiDAR data. For example, in cases where the environment capture system 402 and user system 1110 are in remote locations without cellular networks or in areas without Wi-Fi, the LiDAR data storage area 932 can store the captured LiDAR data until it can be transferred to the image data storage area 934.
[0134] Similar to the LiDAR data storage area, the captured image data storage area 934 can be any and / or multiple structures suitable for capturing images (e.g., active database, relational database, self-referenced database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, and / or similar structures). The image data storage area 934 can store the captured images.
[0135] Figures 10a-10c depict the process of an environment capture system 400 for capturing images in some embodiments. As shown in Figures 10a-10c, the environment capture system 400 can capture images in bursts at different exposures. A burst of images can be a set of images, each with a different exposure. The first burst of images occurs at time 0.0. The environment capture system 400 may receive a first frame and then evaluate the first frame while waiting for a second frame. Figure 10a illustrates the blending of the first frame before the arrival of the second frame. In some embodiments, the environment capture system 400 may process each frame to identify pixels, colors, etc. Once the next frame arrives, the environment capture system 400 may process the most recently received frame and then blend the two frames together.
[0136] In various embodiments, the environment capture system 400 performs image processing to blend a sixth frame and further evaluate the blend.Pixels in a combined frame (e.g., a frame that may include elements from any number of frames from a burst of images). During the final step before or during movement (e.g., rotation) of the ambient capture system 400, the ambient capture system 400 may optionally transfer the blended image from the graphics processing unit to the CPU memory.
[0137] This process continues in FIG. 10b. At the beginning of FIG. 10b, the ambient capture system 400 performs another burst of images. The ambient capture system 400 may use JxR to compress all or part of the blended frame and / or the captured frame. Similar to FIG. 10a, the burst of images may be a set of images, each with a different exposure (the exposure length of each frame in this set may be the same as and in the same order as the other bursts covered in FIG. 10a and in FIG. 10c). The second burst of images occurs at time 2 seconds. The ambient capture system 400 may receive the first frame and then evaluate the first frame while waiting for the second frame. FIG. 10b illustrates blending the first frame before the arrival of the second frame. In some embodiments, the ambient capture system 400 may process each frame to identify pixels, colors, etc. Once the next frame arrives, the ambient capture system 400 may process the most recently received frame and then blend the two frames together.
[0138] In various embodiments, the ambient capture system 400 performs image processing to blend a sixth frame and further evaluate the pixels in the blended frame (e.g., a frame that may include elements from any number of frames from a burst of images). The ambient capture system 400 may optionally transfer the blended image from the graphics processing unit to CPU memory before or during the final step of a movement (e.g., rotation) of the ambient capture system 400.
[0139] After a rotation, the ambient capture system 400 may continue the process by taking another color burst at approximately 3.5 seconds (e.g., after a 180-degree rotation). The ambient capture system 400 may use JxR to compress all or part of the blended frame and / or the captured frame. A burst of images can be a set of images, each with a different exposure (the exposure length of each frame in this set can be the same as and in the same order as the other bursts covered in Figures 10a and 10c). The ambient capture system 400 can receive a first frame and then evaluate that first frame while waiting for a second frame. Figure 10b illustrates blending the first frame before the second frame arrives. In some embodiments, the ambient capture system 400 can process each frame to identify pixels, colors, etc. Once the next frame arrives, the ambient capture system 400 can process the most recently received frame and then blend the two frames together.
[0140] In various embodiments, the ambient capture system 400 performs image processing to blend a sixth frame and further evaluate the pixels in the blended frame (e.g., a frame that may include elements from any number of frames from the image burst). In the ambient capture systemBefore or during the final step of movement (e.g., rotation) of system 400, the ambient capture system 400 may optionally transfer the blended image from the graphics processing unit to the CPU memory.
[0141] In FIG. 10c, the last burst occurs at time 5 seconds. The ambient capture system 400 may use JxR to compress all or part of the blended frame and / or the captured frame. The burst of images may be a set of images, each with a different exposure (the exposure length of each frame in this set may be the same as and in the same order as the other bursts covered in FIG. 10a and 10b). The ambient capture system 400 may receive a first frame and then evaluate that first frame while waiting for the second frame. FIG. 10c illustrates blending the first frame before the second frame arrives. In some embodiments, the ambient capture system 400 may process each frame to identify pixels, colors, etc. Once the next frame arrives, the ambient capture system 400 may process the most recently received frame and then blend the two frames together.
[0142] In various embodiments, the ambient capture system 400 performs image processing to blend a sixth frame and further evaluate the pixels in the blended frame (e.g., a frame that may include elements from any number of frames from a burst of images). The ambient capture system 400 may optionally transfer the blended image from the graphics processing unit to CPU memory before or during the final step of movement (e.g., rotation) of the ambient capture system 400.
[0143] The dynamic range of an image capture device is a measure of how much light an image sensor can capture. Dynamic range is the difference between the darkest and brightest areas of an image. There are many ways to increase the dynamic range of an image capture device, one method being to use different exposures to capture multiple images of the same physical environment. Images captured with short exposures will capture brighter areas of the physical environment, while long exposures will capture darker areas of the physical environment. In some embodiments, the ambient capture system may capture multiple images with six different exposure times. Some or all of the images captured by the ambient capture system are used to generate a 2D image with high dynamic range (HDR). One or more of the captured images can be used for other functions, such as ambient light detection, flicker detection, etc.
[0144] A 3D panoramic image of the physical environment can be generated based on four separate image captures by the image capture device and four separate depth data captures by the LiDAR device of the environment capture system. Each of the four separate image captures may include a series of image captures with different exposure times. A mixing algorithm can be used to mix a series of image captures with different exposure times to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, the environment capture system can be used to capture a 3D panoramic image of a kitchen. An image of one wall of the kitchen may include a window with a shorter exposure time.An image captured by light can provide a view outside the window, but may underexpose the rest of the kitchen. Conversely, another image captured with a longer exposure can provide a view inside the kitchen. The blending algorithm can generate a blended RGB image by blending the view outside the kitchen window from one image with the rest of the kitchen view from another image.
[0145] In various embodiments, a 3D panoramic image can be generated based on three separate image captures by an image capture device and four separate depth data captures by a LiDAR device of an environment capture system. In some embodiments, the number of image captures and the number of depth data captures can be the same. In one embodiment, the number of image captures and the number of depth data captures can be different.
[0146] After capturing the first of a series of images with one exposure time, the blending algorithm receives the first of the series of images, calculates an initial intensity weight for the image, and sets the image as a baseline image for combining subsequently received images. In some embodiments, the blending algorithm may utilize graphics processing unit (GPU) image processing routines, such as the “blend_kernel” routine. The blending algorithm may receive subsequent images that can be blended with previously received images. In some embodiments, the blending algorithm may utilize a variation of the Blend_Kernel GPU image processing routine.
[0147] In one embodiment, the blending algorithm utilizes other methods of blending multiple images, such as determining the difference between the darkest and brightest parts of a baseline image or the contrast of the baseline image, to determine whether the baseline image may be overexposed or underexposed. For example, a contrast value less than a predetermined contrast threshold means that the baseline image is overexposed or underexposed. In one embodiment, the contrast of the baseline image may be calculated by taking the light intensity of the image or the average of a subset of the images. In some embodiments, the blending algorithm calculates the average light intensity of each row or column of the image. In some embodiments, the blending algorithm may determine a histogram of each image received from an image capture device and analyze the histogram to determine the light intensity of the pixels that constitute each image.
[0148] In various embodiments, blending may involve sampling colors within two or more images of the same scene, including along objects and seams. If there is a significant difference in color between two images (e.g., within a predetermined threshold for color, hue, brightness, saturation, and / or similar amounts), then (e.g., on the ambient capture system 400 or user device 1110) the blending module can blend the two images of a predetermined size along the location where the difference exists. In some embodiments, the greater the color or image difference at a location in the images, the greater the amount of space that can be blended around or near that location.
[0149] In some embodiments, after blending, (e.g., on the ambient capture system 400 or user device 1110)The blending module can rescan and sample colors along one or more images to determine if there are other differences in the images or colors that exceed predetermined thresholds for color, hue, brightness, saturation, and / or similar amounts. If so, the blending module can identify portions within one or more images and continue blending those portions of the image. The blending module can continue resampling the image along the seams until there are no other portions of the image to blend (e.g., any color differences are below one or more predetermined thresholds as per specification page 22 / 35, CN 121531224 A).
[0150] Figure 11 depicts a block diagram of an example environment 1100 capable of capturing and stitching images to form a 3D visualization according to some embodiments. Example environment 1100 includes a 3D and panoramic capture and stitching system 1102, a communication network 1104, an image stitching and processing system 1106, an image data storage area 1108, a user system 1110, and a first scene 1112 of the physical environment. The 3D and panoramic capture and stitching system 1102 and / or the user system 1110 may include an image capture device (e.g., environment capture system 400) that can be used to capture images of the environment (e.g., physical environment 1112).
[0151] The 3D and panoramic capture and stitching system 1102 and the image stitching and processing system 1106 may be part of the same system (e.g., part of one or more digital devices) communicatively coupled to the environment capture system 400. In some embodiments, one or more functions of the components of the 3D and panoramic capture and stitching system 1102 and the image stitching and processing system 1106 may be performed by the environment capture system 400. Similarly or alternatively, the 3D and panoramic capture and stitching system 1102 and the image stitching and processing system 1106 may be executed by the user system 1110 and / or the image stitching and processing system 1106.
[0152] A user may utilize the 3D panoramic capture and stitching system 1102 to capture multiple 2D images of an environment such as the interior and / or exterior of a building. For example, a user may utilize the 3D and panoramic capture and stitching system 1102 to capture multiple 2D images of a first scene of a physical environment 1112 provided by the environment capture system 400. The 3D and panoramic capture and stitching system 1102 may include an alignment and stitching system 1114. Alternatively, the user system 1110 may include an alignment and stitching system 1114.
[0153] The alignment and stitching system 1114 can be software, hardware, or a combination of both, configured to provide guidance and / or process images to a user of an image capture system (e.g., on the 3D and panoramic capture and stitching system 1102 or user system 1110) to enable the creation of improved panoramic images (e.g., through stitching, alignment, cropping, etc.). Alignment and stitching system1114 may be on a computer-readable medium (described herein). In some embodiments, the alignment and stitching system 1114 may include a processor for performing functions.
[0154] An example of the first scene of the physical environment 1112 may be any room, real estate, etc. (e.g., a representation of a living room). In some embodiments, the 3D and panoramic capture and stitching system 1102 is used to generate a 3D panoramic image of an indoor environment. In some embodiments, the 3D panoramic capture and stitching system 1102 may be the environment capture system 400 discussed with respect to FIG. 4.
[0155] In some embodiments, the 3D panoramic capture and stitching system 1102 may communicate with devices for capturing image and depth data and software (e.g., the environment capture system 400). All or part of the software may be installed on the 3D panoramic capture and stitching system 1102, the user system 1110, the environment capture system 400, or both. In some embodiments, a user may interact with the 3D and panoramic capture and stitching system 1102 via the user system 1110.
[0156] The 3D and panoramic capture and stitching system 1102 or the user system 1110 can acquire multiple 2D images. The 3D and panoramic capture and stitching system 1102 or the user system 1110 can acquire depth data (e.g., from LiDAR devices, etc.).
[0157] In various embodiments, an application on the user system 1110 (e.g., a user's smart device, such as a smartphone or tablet computer) or an application on the environment capture system 400 can provide the user with visual or auditory guidance for capturing images using the environment capture system 400. Graphical guidance may include, for example, floating arrows on the display of the environment capture system 400 (e.g., on a viewfinder or LED screen on the back of the environment capture system 400) to guide the user in positioning and / or pointing the image capture device. In another example, the application may provide audio guidance on positioning and / or pointing the image capture device.
[0158] In some embodiments, the guidance may allow the user to capture multiple images of the physical environment without the aid of a stabilizing platform (e.g., a tripod). In one example, the image capturing device can be a personal device, such as a smartphone, tablet, media tablet, laptop, etc. The application can provide orientation about the location for each sweep to approximate parallax-free based on the location of the image capturing device (pages 23 / 35 of the specification, CN 121531224 A), location information from the image capturing device, and / or previous images from the image capturing device.
[0159] In some embodiments, visual and / or auditory guidance enables the capture of images that can be stitched together to form a panorama without a tripod and without camera positioning information (e.g., indicating the orientation, position, and / or orientation of the camera from sensors, GPS devices, etc.).
[0160] Alignment and stitching system 1114 can align or stitch 2D images (e.g., captured by user system 1110 or 3D panoramic capture and stitching system 1102) to obtain a 2D panoramic image.
[0161] In some embodiments, alignment and stitching system 1114 utilizes machine learning algorithms to align or stitch multiple 2D images into a 2D panoramic image. The parameters of the machine learning algorithm can be managed by alignment and stitching system 1114. For example, 3D and panoramic capture and stitching system 1102 and / or alignment and stitching system 1114 can identify objects within 2D images to help align the images into a 2D panoramic image.
[0162] In some embodiments, alignment and stitching system 1114 can utilize depth data and 2D panoramic images to obtain a 3D panoramic image. The 3D panoramic image can be provided to 3D and panoramic stitching system 1102 or user system 1110. In some embodiments, the alignment and stitching system 1114 determines 3D / depth measurements associated with an identified object within a 3D panoramic image and / or sends one or more 2D images, depth data, (one or more) 2D panoramic images, and (one or more) 3D panoramic images to the image stitching and processor system 106 to obtain a 2D panoramic image or 3D panoramic image with a larger pixel resolution than the 2D panoramic image or 3D panoramic image provided by the 3D and panoramic capture and stitching system 1102.
[0163] The communication network 1104 may represent one or more computer networks (e.g., LAN, WAN, etc.) or other transmission media. The communication network 1104 can provide communication between systems 1102, 1106-1110 and / or other systems described herein. In some embodiments, the communication network 104 includes one or more digital devices, routers, cables, buses and / or other network topologies (e.g., mesh, etc.). In some embodiments, the communication network 1104 may be wired and / or wireless. In various embodiments, the communication network 1104 may include the Internet, one or more wide area networks (WANs) or local area networks (LANs), and may be public, private, IP-based, non-IP-based, or one or more networks, etc.
[0164] The image stitching and processing system 1106 may process 2D images captured by an image capture device (e.g., an environment capture system 400 or a user device such as a smartphone, personal computer, media tablet, etc.) and stitch them into a 2D panoramic image. The 2D panoramic image processed by the image stitching and processing system 106 may have a higher pixel resolution than the panoramic image obtained by the 3D and panoramic capture and stitching system 1102.
[0165] In some embodiments, the image stitching and processing system 1106 receives and processes a 3D panoramic image to produce a 3D panoramic image with a higher pixel resolution than the received 3D panoramic image. Higher pixel resolution can be achieved by...A high-resolution panoramic image is provided to an output device, such as a computer screen, projector screen, etc., with a higher screen resolution than the user system 1110. In some embodiments, a higher pixel resolution panoramic image can be provided to the output device in more detail and can be magnified.
[0166] The image data storage area 1108 can be any and / or multiple structures suitable for capturing image and / or depth data (e.g., active database, relational database, self-referenced database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, and / or similar structures). The image data storage area 1108 can store images captured by the image capture device of the user system 1110. In various embodiments, the image data storage area 1108 stores depth data captured by one or more depth sensors of the user system 1110. In various embodiments, the image data storage area 1108 stores properties associated with the image capture device or properties associated with each of multiple image captures or depth captures used to determine a 2D or 3D panoramic image. In some embodiments, on pages 24 / 35 of CN 121531224 A, image data storage area 1108 stores panoramic 2D or 3D panoramic images. The 2D or 3D panoramic images may be determined by 3D and panoramic capture and stitching system 1102 or image stitching and processor system 106.
[0167] User system 1110 can communicate between a user and other associated systems. In some embodiments, user system 1110 may be or include one or more mobile devices (e.g., smartphones, cellular phones, smartwatches, etc.).
[0168] User system 1110 may include one or more image capture devices. The one or more image capture devices may include, for example, RGB cameras, HDR cameras, video cameras, IR cameras, etc.
[0169] 3D and panoramic capture and stitching system 1102 and / or user system 1110 may include two or more capture devices that may be arranged in opposite positions on or within the same mobile housing, such that their common field of view spans up to 360°. In some embodiments, paired image capture devices capable of generating stereoscopic image pairs (e.g., fields of view with slight offsets but partial overlap) may be used. User system 1110 may include two image capture devices having a vertically offset field of view capable of capturing vertical stereoscopic image pairs. In another example, user system 1110 may include two image capture devices having a vertically offset field of view capable of capturing vertical stereoscopic image pairs.
[0170] In some embodiments, user system 1110, environment capture system 400, or 3D and panoramic capture and stitching system1102 can generate and / or provide image capture location and orientation information. For example, user system 1110 or 3D and panoramic capture and stitching system 1102 may include an inertial measurement unit (IMU) to help determine location data associated with one or more image capture devices that capture multiple 2D images. User system 1110 may include a global positioning sensor (GPS) to provide GPS coordinate information associated with multiple 2D images captured by one or more image capture devices.
[0171] In some embodiments, a user may interact with alignment and stitching system 1114 using a mobile application installed in user system 1110. 3D and panoramic capture and stitching system 1102 may provide images to user system 1110. The user can utilize alignment and stitching system 1114 on user system 1110 to view and preview images.
[0172] In various embodiments, alignment and stitching system 1114 may be configured to provide or receive one or more 3D panoramic images from 3D and panoramic capture and stitching system 1102 and / or image stitching and processor system 1106. In some embodiments, the 3D and panoramic capture and stitching system 1102 may provide the user system 1110 with a visual representation of a portion of a floor plan of a building that has been captured by the 3D and panoramic capture and stitching system 1102.
[0173] A user of system 1110 may navigate the space around the area and view different rooms of the house. In some embodiments, as the image stitching and processor system 1106 completes the generation of the 3D panoramic image, a user of user system 1110 may display the 3D panoramic image, such as the example 3D panoramic image. In various embodiments, user system 1110 generates a preview or thumbnail of the 3D panoramic image. The preview 3D panoramic image may have a lower image resolution than the 3D panoramic image generated by the 3D and panoramic capture and stitching system 1102.
[0174] FIG12 is a block diagram of an example of an alignment and stitching system 1114 according to some embodiments. The alignment and stitching system 1114 includes a communication module 1202, an image capture position module 1204, a stitching module 1206, a cropping module 1208, a graphic cutting module 1210, a mixing module 1211, a 3D image generator 1214, a captured 2D image data storage area 1216, a 3D panoramic image data storage area 1218, and a guidance module 220. It is understood that any number of modules performing one or more of the different functions described herein may exist in the alignment and stitching system 1114.
[0175] In some embodiments, the alignment and stitching system 1114 includes an image capture module configured to receive images from one or more image capture devices (e.g., cameras). The alignment and stitching system 1114 may also include a depth module configured to receive depth data from a depth device such as LiDAR (if available).
[0176] The communication module 1202 can send and receive requests, images, or data between any module or data storage area of the alignment and stitching system 1114 and components of the example environment 1100 of FIG11. Similarly, the alignment and stitching system 1114 can send and receive requests, images, or data to any device or system on the communication network 1104 on pages 25 / 35 of the specification, CN 121531224 A.
[0177] In some embodiments, the image capture location module 1204 can determine image capture device location data (e.g., a camera with a stand-alone camera, a smartphone, a media tablet, a laptop, etc.). The image capture device location data can indicate the position and orientation of the image capture device and / or lens. In one example, the image capture location module 1204 can utilize the user system 1110, a camera, a digital device with a camera, or the IMU of the 3D and panoramic capture and stitching system 1102 to generate image capture device location data. The image capture location module 1204 can determine the current orientation, angle, or tilt of one or more image capture devices (or lenses). Image capture positioning module 1204 may also utilize GPS from user system 1110 or 3D and panoramic capture and stitching system 1102.
[0178] For example, when a user wants to use user system 1110 to capture a 360° view of a physical environment (such as a living room), the user can hold user system 1110 in front of them at eye level to begin capturing one of several images that will eventually become a 3D panoramic image. To reduce the amount of parallax in the images and capture images that are more suitable for stitching and generating a 3D panoramic image, it is preferable if one or more image capture devices are rotated at the center of the rotation axis. Alignment and stitching system 1114 may (e.g., from IMU) receive position information to determine the position of the image capture device or lens. Alignment and stitching system 1114 may receive and store the field of view of the lens. Guidance module 1220 may provide visual and / or audio information about a recommended initial position for the image capture device. Guidance module 1220 may make recommendations for positioning the image capture device for subsequent images. In one example, the guidance module 1220 can provide guidance to the user to rotate and position the image capture device such that the image capture device rotates closer to the center of rotation. Furthermore, the guidance module 1220 can provide guidance to the user to rotate and position the image capture device such that subsequent images are substantially aligned based on the field of view and / or the characteristics of the image capture device.
[0179] The guidance module 1220 can provide visual guidance to the user. For example, the guidance module 1220 can place markers or arrows in a viewer or display on the user system 1110 or the 3D and panoramic capture and stitching system 1102. In some implementations...For example, user system 1110 may be a smartphone or tablet computer with a display. When one or more images are captured, guidance module 1220 may position one or more markers (e.g., markers of different colors or the same markers) on the output device and / or in the viewfinder. The user can then use the markers on the output device and / or viewfinder to align the next image.
[0180] There are many techniques for guiding the user of user system 1110 or 3D and panoramic capture and stitching system 1102 to capture multiple images in order to stitch the images into a panorama. When a panorama is obtained from multiple images, the images can be stitched together. In order to improve the time, efficiency and effectiveness of stitching images together while reducing the need to correct artifacts or misalignment, image capture positioning module 1204 and guidance module 1220 can help the user capture multiple images in a location that improves the quality, time efficiency and effectiveness of image stitching of the desired panorama.
[0181] For example, after capturing the first image, the display of user system 1110 may include two or more objects, such as circles. The two circles may appear stationary relative to the environment, and the two circles may move with the user system 1110. When the two stationary circles are aligned with the two circles moving with the user system 1110, the image capturing device and / or the user system 1110 may be aligned for the next image.
[0182] In some embodiments, after the image capturing device captures an image, the image capturing position module 1204 may acquire sensor measurements of the position of the image capturing device (e.g., including orientation, tilt, etc.). The image capturing position module 1204 may determine one or more edges of the captured image by calculating the position of the edge of the field of view based on the sensor measurements. Alternatively or additionally, the image capturing position module 1204 may determine one or more edges of the image by scanning the image captured by the image capturing device, identifying objects within the image (e.g., using machine learning models discussed herein), determining one or more edges of the image, and positioning the objects (e.g., circles or other shapes) at the edges of the display on the user system 1110. Instruction manual, pages 26 / 35, 29, CN 121531224 A
[0183] The image capture position module 1204 can display two objects within the display of the user system 1110, indicating the location of the field of view for the next image. These two objects can indicate the location of the edge in the environment representing the presence of the last image. The image capture position module 1204 can continue to receive sensor measurements of the position of the image capture device and calculate two additional objects in the field of view. The two additional objects can be spaced apart by the same width as the first two objects. While the first two objects can represent the edges of the captured image (e.g., the rightmost edge of the image), the next two objects representing the edges of the field of view...Another object can be on the opposite edge (e.g., the leftmost edge of the field of view). By physically aligning the first two objects on the edge of the image with the other two objects on the opposite edge of the field of view, the image capturing device can be positioned to capture another image that can be stitched together more efficiently without a tripod. This process can continue for each image until the user determines that the desired panorama has been captured.
[0184] Although multiple objects are discussed herein, it should be understood that the image capturing position module 1204 can calculate the position of one or more objects for positioning the image capturing device. Objects can be of any shape (e.g., circles, ellipses, squares, emojis, arrows, etc.). In some embodiments, objects can have different shapes.
[0185] In some embodiments, there can be a distance between objects representing the edges of the captured image and between objects in the field of view. The user can be guided forward to move away so that sufficient distance can exist between objects. Alternatively, as the image capturing device approaches the correct position (e.g., by moving closer or further away from a position that would allow the next image to be captured in a position that would improve image stitching), the size of the objects in the field of view can be changed to match the size of the objects representing the edges of the captured image.
[0186] In some embodiments, the image capture location module 1204 may use objects in an image captured by the image capture device to estimate the location of the image capture device. For example, the image capture location module 1204 may use GPS coordinates to determine the geographic location associated with the image. The image capture location module 1204 may use this location to identify landmarks that can be captured by the image capture device.
[0187] The image capture location module 1204 may include a 2D machine learning model for converting 2D images into 2D panoramic images. The image capture location module 1204 may include a 3D machine learning model for converting 2D images into 3D representations. In one example, the 3D representation may be used to display a three-dimensional walkthrough or visualization of an interior and / or exterior environment.
[0188] The 2D machine learning model may be trained to stitch together or help stitch together two or more 2D images to form a 2D panoramic image. The 2D machine learning model may be, for example, a neural network trained with 2D images, which include physical objects in the images and object recognition information that trains the 2D machine learning model to recognize objects in subsequent 2D images. Objects in a 2D image can help determine one or more locations within the 2D image, aiding in the identification of edges, distortions, and alignment. Furthermore, objects in a 2D image can help determine artifacts, blending of artifacts or boundaries between two images, and the location of image cropping and / or truncating.
[0189] In some embodiments, the 2D machine learning model may be, for example, a neural network trained with 2D images, wherein the 2D images include depth information of the environment (e.g., from a LiDAR device or structured light device or 3D and panoramic capture and stitching system 1102 from user system 1110) and include physical objects in the image to identify the physical objects, the location of the physical objects, and / or the location of the image capture device / field of view. The 2D machine learning model may identify the depth of the physical objects and their other aspects relative to the 2D images to aid in the alignment and positioning of the two 2D images for stitching (or stitching two 2D images).
[0190] The 2D machine learning model may include any number of machine learning models (e.g., any number of models generated by neural networks, etc.).
[0191] The 2D machine learning model may be stored on the 3D and panoramic capture and stitching system 1102, the image stitching and processor system 1106, and / or the user system 1110. In some embodiments, the 2D machine learning model may be trained by the image stitching and processor system 1106.
[0192] The image capture position module 1204 can estimate the position of the image capture device (the position of the field of view of the image capture device) based on the seams between two or more 2D images from the stitching module 1206, the image distortion from the cropping module 1208, and / or the graphic cutting from the graphic cutting module 1210.
[0193] The stitching module 1206 can combine two or more 2D images to generate a 2D panorama. Based on the seams between the two or more 2D images from the stitching module 1206, the image distortion from the cropping module 1208, and / or the graphic cutting, it has a larger field of view than each of the two or more images.
[0194] The stitching module 1206 can be configured to align or “stitch together” two different 2D images that provide different perspectives of the same environment to generate a panoramic 2D image of the environment. For example, the stitching module 1206 can use known or (e.g., derived) information about the capture position and orientation of the respective 2D images to help stitch the two images together.
[0195] The stitching module 1206 can receive two 2D images. The first 2D image may have been captured just before the second image or within a predetermined time period. In various embodiments, the stitching module 1206 may receive positioning information of the image capture device associated with the first image, and then receive positioning information associated with the second image. The positioning information may be associated with the image based on positioning data from an IMU, GPS, and / or information provided by the user at the time the image was captured.
[0196] In some embodiments, the stitching module 1206 may utilize a 2D machine learning module to scan the two images for identification.Objects within the two images include objects (or portions of objects) that can be shared by the two images. For example, the stitching module 1206 can identify corners, patterns on walls, furniture, etc., that are shared at the opposite edges of the two images.
[0197] The stitching module 1206 can align the edges of the two 2D images based on the location of the shared object (or portion of the object), location data from the IMU, location data from the GPS, and / or information provided by the user, and then combine the two edges of the images (i.e., "stitch" them together). In some embodiments, the stitching module 1206 can identify portions of the two 2D images that overlap each other and stitch the images at the overlapping locations (e.g., using location data and / or the results of a 2D machine learning model).
[0198] In various embodiments, the 2D machine learning model can be trained to combine or stitch the two edges of an image using location data from the IMU, location data from the GPS, and / or information provided by the user. In some embodiments, the 2D machine learning model can be trained to identify common objects in two 2D images to align and locate the 2D images, and then combine or stitch the two edges of the images. In another embodiment, a 2D machine learning model can be trained to align and position 2D images using localization data and object recognition, and then stitch the two edges of the images together to form all or part of a panoramic 2D image.
[0199] The stitching module 1206 can utilize depth information of the respective images (e.g., pixels in the respective images, objects in the respective images, etc.) to align the respective 2D images with each other in association with a single 2D panoramic image of the generated environment.
[0200] The cropping module 1208 can address the problem of two or more 2D images when the image capturing devices are not held in the same position while capturing 2D images. For example, while capturing an image, the user can position the user system 1110 in a vertical position. However, while capturing another image, the user can position the user system at an angle. The resulting images may be misaligned and may exhibit parallax effects. Parallax effects occur when foreground and background objects are not aligned in the same way in the first and second images.
[0201] The cropping module 1208 can utilize a 2D machine learning model (by applying positioning information, depth information, and / or object recognition) to detect changes in the position of the image capturing device in two or more images, and then measure the amount of change in the position of the image capturing device. The cropping module 1208 can distort one or more 2D images such that when the images are stitched together, the images can be aligned to form a panoramic image, while preserving certain characteristics of the images, such as keeping straight lines straight.
[0202] The output of the cropping module 1208 may include the number of pixel columns and rows to straighten the image by offsetting each pixel of the image. The offset of each image may be output as a matrix representing the number of pixel columns and rows offsetting each pixel of the image.
[0203] In some embodiments, the cropping module 1208 may determine the amount of image distortion to be performed on one or more of the plurality of 2D images captured by the image capture device of the user system 1110 based on one or more image capture locations from the image capture location module 1204 or the gap between two or more 2D images from the stitching module 1206, the graphic cutting from the graphic cutting module 1210, or the color mixing from the mixing module 1211.
[0204] The graphic cutting module 1210 may determine where to cut or slice the one or more 2D images captured by the image capture device. For example, the graphic cutting module 1210 may use a 2D machine learning model to identify objects in two images and determine that they are the same objects. Image capture positioning module 1204, cropping module 1208, and / or image cutting module 1210 can determine that two images cannot be aligned even when distorted. Image cutting module 1210 can utilize information from a 2D machine learning model to identify segments of the two images that can be stitched together (e.g., by cutting off a portion of one or both images to aid alignment and positioning). In some implementations, the two 2D images can overlap with at least a portion of the physical world represented in the images. Image cutting module 1210 can identify objects in the two images, such as the same chair. However, even after image distortion by image capture positioning and cropping module 1208, the images of the chair cannot be aligned to generate a distortion-free panorama and will not correctly represent a portion of the physical world. Image cutting module 1210 can select one of the two images of the chair as the correct representation (e.g., based on misalignment, positioning errors, and / or artifacts in one image compared to the other) and cut the chair from the image with misalignment, positioning errors, and / or artifacts. Stitching module 1206 can then stitch the two images together.
[0205] The graphic cutting module 1210 can try two combinations, for example, cutting an image of a chair from a first image and stitching the first image minus the chair to a second image, to determine which graphic cutting generates a more accurate panoramic image. The output of the graphic cutting module 1210 can be the position of cutting one or more of the multiple 2D images corresponding to the graphic cutting, which generates a more accurate panoramic image.
[0206] The graphic cutting module 1210 can be based on one or more image capture positions from the image capture position module 1204, the stitching or seam between two or more 2D images from the stitching module 1206, and the image from the cropping module 1208.Image distortion and graphic cutting from the graphic cutting module 1210 determine how to cut or slice one or more 2D images captured by the image capture device.
[0207] The blending module 1211 can color at the seam (e.g., stitching) between two images, making the seam invisible. Variations in lighting and shadows can cause the same object or surface to be output in slightly different colors or shadows. The blending module can determine the desired amount of color blending based on one or more image capture positions from the image capture position module 1204, stitching, image colors along the seam from the two images, image distortion from the cropping module 1208, and / or graphic cutting from the graphic cutting module 1210.
[0208] In various embodiments, the blending module 1211 can receive a panorama from a combination of two 2D images and then sample colors along the seam between the two 2D images. The blending module 1211 can receive seam position information from the image capture position module 1204 so that the blending module 1211 can sample colors along the seam and determine differences. If there is a significant difference in color along the seam between two images (e.g., within a predetermined threshold for color, hue, brightness, saturation, and / or similar amounts), the blending module 1211 can blend two images of a predetermined size along the seam at the location where the difference exists. In some embodiments, the greater the difference in color or image along the seam, the greater the amount of space that can be blended along the seam between the two images.
[0209] In some embodiments, after blending, the blending module 1211 can rescan and sample the color along the seam to determine if there are other differences in the image or color exceeding a predetermined threshold for color, hue, brightness, saturation, and / or similar amounts. If so, the blending module 1211 can identify the portion along the seam and continue blending that portion of the image. The blending module 1211 can continue resampling the image along the seam until there are no other portions of the image to blend (e.g., any color difference is below a predetermined threshold).
[0210] The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to transform a 2D panoramic image into a 3D representation. The 3D machine learning model can be trained using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or structured light device) to produce the 3D representation. The 3D representation can be tested and reviewed against curation and feedback. In some embodiments, the 3D machine learning model can be used together with the 2D panoramic image and depth data to generate the 3D representation.
[0211] In various embodiments, the accuracy, rendering speed, and quality of the 3D representation generated by the 3D image generator 1214 are significantly improved by utilizing the systems and methods described herein. For example, the accuracy, rendering speed, and quality of the 3D representation are improved by rendering a 3D representation from a 2D panoramic image that has already been aligned, positioned, and stitched using the methods described herein (e.g., by alignment and positioning information provided by hardware, by improved positioning resulting from guidance provided to the user during image capture, by cropping and altering image distortion, by cutting the image to avoid artifacts and overcome distortion, by blending images, and / or any combination thereof). Furthermore, it should be understood that by utilizing a 2D panoramic image that has already been aligned, positioned, and stitched using the methods described herein, the training of 3D machine learning models can be significantly improved (e.g., in terms of speed and accuracy). Moreover, in some embodiments, the 3D machine learning model can be smaller and less complex due to the reduction in processing and learning required to overcome misalignment, positioning errors, distortion, poor graphic cutting, poor blending, artifacts, etc., to generate a reasonably accurate 3D representation.
[0212] The trained 3D machine learning model can be stored in the 3D and panoramic capture and stitching system 1102, the image stitching and processing system 106, and / or the user system 1110.
[0213] In some embodiments, a 3D machine learning model can be trained using multiple 2D images and depth data from the image capture devices of the user system 1110 and / or the 3D and panoramic capture and stitching system 1102. Additionally, the 3D image generator 1214 can be trained using image capture location information associated with each of the multiple 2D images from the image capture location module 1204, seam positions for aligning or stitching each of the multiple 2D images from the stitching module 1206, pixel offsets (one or more) of each of the multiple 2D images from the cropping module 1208, and / or graphic cuts from the graphic cutting module 1210. In some embodiments, the 3D machine learning model may be used with 2D panoramic images, depth data, image capture location information associated with each of a plurality of 2D images from image capture location module 1204, seam positions for aligning or stitching each of a plurality of 2D images from stitching module 1206, pixel offsets (one or more) of each of a plurality of 2D images from cropping module 1208, and / or graphic cuts from graphic cutting module 1210 to generate a 3D representation.
[0214] Stitching module 1206 may be part of a 3D model that converts a plurality of 2D images into a 2D panoramic or 3D panoramic image. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D. Cropping module 1208 may be part of a 3D model that converts a plurality of 2D images into a 2D panoramic or 3D panoramic image. In some embodimentsIn this context, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D. The image cropping module 1210 may be part of a 3D model that converts multiple 2D images into 2D panoramas or 3D panoramic images. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D, as described on pages 30 / 35 of the specification (CN 121531224 A). The blending module 1211 may be part of a 3D machine learning model that converts multiple 2D images into 2D panoramas or 3D panoramic images. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D.
[0215] The 3D image generator 1214 may generate weights for each of the image capture location module 1204, the cropping module 1208, the image cropping module 1210, and the blending module 1211, which may represent the reliability or "strength" or "weakness" of the module. In some embodiments, the sum of the weights of the modules is equal to 1.
[0216] In cases where depth data is not available for multiple 2D images, the 3D image generator 1214 may determine depth data for one or more objects in a plurality of 2D images captured by the image capture device of the user system 1110. In some embodiments, the 3D image generator 1214 may derive depth data based on images captured by a stereo image pair. The 3D image generator may evaluate the stereo image pair to determine data (a more intermediate result) regarding the photometric matching quality between images at various depths, rather than determining depth data based on a passive stereo algorithm.
[0217] The 3D image generator 1214 may be part of a 3D model that converts multiple 2D images into 2D panoramas or 3D panoramic images. In some embodiments, the 3D model is a machine learning algorithm, such as a predictive neural network model from 2D to 3D.
[0218] The captured 2D image data storage area 1216 can be any and / or multiple structures suitable for the captured image and / or depth data (e.g., active database, relational database, self-referenced database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, and / or similar structures). The captured 2D image data storage area 1216 can store images captured by the image capture device of the user system 1110. In various embodiments, the captured 2D image data storage area 1216 stores depth data captured by one or more depth sensors of the user system 1110. In various embodiments, the captured 2D image data storage area 1216 stores image capture device parameters associated with the image capture device, or capture properties associated with each of a plurality of image captures, or depth captures used to determine the 2D panoramic image. In some embodiments, the image data storage area 1108 stores panoramic 2D images.Scene image. The 2D panoramic image can be determined by the 3D and panoramic capture and stitching system 1102 or the image stitching and processor system 106. Image capture device parameters may include illumination, color, image capture lens focal length, maximum aperture, tilt angle, etc. Capture properties may include pixel resolution, lens distortion, illumination, and other image metadata.
[0219] The 3D panoramic image data storage area 1218 may be any and / or multiple structures suitable for 3D panoramic images (e.g., active database, relational database, self-referenced database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, and / or similar structures). The 3D panoramic image data storage area 1218 may store 3D panoramic images generated by the 3D and panoramic capture and stitching system 1102. In various embodiments, the 3D panoramic image data storage area 1218 stores properties associated with the image capture device or properties associated with each of multiple image captures or depth captures used to determine the 3D panoramic image. In some embodiments, the 3D panoramic image data storage area 1218 stores 3D panoramic images. 2D or 3D panoramic images can be determined by the 3D and panoramic capture and stitching system 1102 or the image stitching and processor system 106.
[0220] FIG13 depicts a flowchart 1300 of a 3D panoramic image capture and generation process according to some embodiments. In step 1302, the image capture device can use the image sensor 920 of FIG9 and the WFOV lens 918 to capture multiple 2D images. A wider FOV means that the environment capture system 402 will require fewer scans to obtain a 360° view. The WFOV lens 918 can also be wider horizontally and vertically. In some embodiments, the image sensor 920 captures RGB images. In one embodiment, the image sensor 920 captures black and white images.
[0221] In step 1304, the environment capture system can send the captured 2D images to the image stitching and processor system 1106. The image stitching and processing system 1106 can apply a 3D modeling algorithm to the captured 2D images to generate a panoramic 2D image. In some embodiments, the 3D modeling algorithm is a machine learning algorithm for stitching captured 2D images into a panoramic 2D image. In some embodiments, step 1304 may be optional.
[0222] In step 1306, the LiDAR 912 and WFOV lens 918 of FIG9 can capture LiDAR data. A wider FOV means that the environment capture system 400 will require fewer scans to obtain a 360° view.
[0223] In step 1308, the LiDAR data can be sent to the image stitching and processing system 1106. Image stitching and...Processor system 1106 can input LiDAR data and captured 2D images into a 3D modeling algorithm to generate a 3D panoramic image. The 3D modeling algorithm is a machine learning algorithm.
[0224] In step 1310, image stitching and processor system 1106 generates a 3D panoramic image. The 3D panoramic image can be stored in image data storage area 408. In one embodiment, the 3D panoramic image generated by the 3D modeling algorithm is stored in image stitching and processor system 1106. In some embodiments, as the environment capture system is used to capture various parts of the physical environment, the 3D modeling algorithm can generate a visual representation of a plan view of the physical environment.
[0225] In step 1312, image stitching and processor system 1106 can provide at least a portion of the generated 3D panoramic image to user system 1110. Image stitching and processor system 1106 can provide a visual representation of a plan view of the physical environment.
[0226] The order of one or more steps in flowchart 1300 can be changed without affecting the final result of the 3D panoramic image. For example, an environment capture system can interleave image capture using an image capture device with LiDAR data or depth information capture using LiDAR 912. For example, the image capture device can capture an image of a segment of the physical environment, and then LiDAR 912 obtains depth information from segment 1605. Once LiDAR 912 has obtained depth information from the segment, the image capture device can continue to move to capture an image of another segment, and then LiDAR 912 obtains depth information from the segment, thus interleaving image capture and depth information capture.
[0227] In some embodiments, the devices and / or systems discussed herein employ a single image capture device to capture a 2D input image. In some embodiments, the one or more image capture devices 1116 may represent a single image capture device (or image capture lens). According to some of these embodiments, a user of a moving device housing an image capture device can be configured to rotate about an axis to generate images with different capture orientations relative to the environment, wherein the common field of view of the images horizontally spans up to 360°.
[0228] In various embodiments, the devices and / or systems discussed herein may employ two or more image capture devices to capture 2D input images. In some embodiments, two or more image capture devices may be arranged in opposite positions on or within the same movable housing such that their common field of view spans up to 360°. In some embodiments, paired image capture devices capable of generating stereoscopic image pairs (e.g., with slightly offset but partially overlapping fields of view) may be used. For example, user system 1110 (e.g., a device including one or more image capture devices for capturing 2D input images).It may include two image capture devices having a horizontal stereo offset field of view capable of capturing pairs of stereo images. In another example, user system 1110 may include two image capture devices having a vertical stereo offset field of view capable of capturing pairs of vertical stereo images. According to any of these examples, each camera may have a field of view spanning up to 360 degrees. In this regard, in one embodiment, user system 1110 may employ two panoramic cameras with vertical stereo offset capable of capturing paired panoramic images of stereo pairs (with vertical stereo offset).
[0229] Positioning component 1118 may include any hardware and / or software configured to capture user system position data and / or user system orientation data. For example, positioning component 1118 includes an IMU to generate user system 1110 position data associated with the one or more image capture devices of user system 1110 for capturing multiple 2D images. Positioning component 1118 may include a GPS unit to provide GPS coordinate information associated with multiple 2D images captured by one or more image capture devices. In some embodiments, positioning component 1118 may associate location and orientation data of the user system with corresponding images captured by the one or more image capture devices of the user system 1110.
[0230] Various embodiments of the device provide users with 3D panoramic images of indoor and outdoor environments. In some embodiments, the device may efficiently and quickly provide users with 3D panoramic images of indoor and outdoor environments using a single wide field of view (FOV) lens and a single light and detection and ranging sensor (LiDAR sensor).
[0231] The following are example uses of the example devices described herein. The following use case is one of the embodiments. Different embodiments of the device discussed herein may include one or more features and capabilities similar to those of this use case.
[0232] FIG14 depicts a flowchart of a 3D and panoramic capture and stitching process 1400 according to some embodiments. The flowchart in Figure 14 refers to the 3D and panoramic capture and stitching system 1102 as including an image capture device, but in some embodiments, the data capture device may be the user system 1110.
[0233] In step 1402, the 3D and panoramic capture and stitching system 1102 may receive multiple 2D images from at least one image capture device. The image capture device of the 3D and panoramic capture and stitching system 1102 may be or include a complementary metal-oxide-semiconductor (CMOS) image sensor. In various embodiments, the image capture device is a charge-coupled device (CCD). In one example, the image capture device is a red-green-blue (RGB) sensor. In one embodiment, the image capture device is an IR sensor.Each of the plurality of 2D images may have a field of view that overlaps with at least one other image portion of the plurality of 2D images. In some embodiments, at least some of the plurality of 2D images are combined to produce a 360° view of a physical environment (e.g., indoor, outdoor, or both).
[0234] In some embodiments, all of the plurality of 2D images are received from the same image capture device. In various embodiments, at least a portion of the plurality of 2D images is received from two or more image capture devices of the 3D and panoramic capture and stitching system 1102. In one example, the plurality of 2D images includes a set of RGB images and a set of IR images, wherein the IR images provide depth data to the 3D and panoramic capture and stitching system 1102. In some embodiments, each 2D image may be associated with depth data provided by a LiDAR device. In some embodiments, each 2D image may be associated with positioning data.
[0235] In step 1404, the 3D and panoramic capture and stitching system 1102 may receive capture parameters and image capture device parameters associated with each of the received plurality of 2D images. Image capture device parameters may include illumination, color, image capture lens focal length, maximum aperture, field of view, etc. Capture properties may include pixel resolution, lens distortion, illumination, and other image metadata. The 3D and panoramic capture and stitching system 1102 may also receive positioning data and depth data.
[0236] In step 1406, the 3D and panoramic capture and stitching system 1102 may acquire the information received from steps 1402 and 1404 for stitching 2D images to form a 2D panoramic image. The process of stitching 2D images is further discussed with reference to the flowchart of FIG. 15.
[0237] In step 1408, the 3D and panoramic capture and stitching system 1102 may apply a 3D machine learning model to generate a 3D representation. The 3D representation may be stored in a 3D panoramic image data storage area. In various embodiments, the 3D representation is generated by the image stitching and processor system 1106. In some embodiments, as the environment capture system is used to capture various parts of the physical environment, the 3D machine learning model may generate a visual representation of a planar view of the physical environment.
[0238] In step 1410, the 3D and panoramic capture and stitching system 1102 may provide at least a portion of the generated 3D representation or model to the user system 1110. The user system 1110 may provide a visual representation of a planar view of the physical environment.
[0239] In some embodiments, the user system 1110 may send multiple 2D images, capture parameters, and image capture parameters to the image stitching and processing system 1106. In various embodiments, the 3D and panoramic capture and stitching system 1102 may send multiple 2D images, capture parameters, and image capture parameters to the image stitching and processing system 1106. Specification 33 / 35 pages 36 CN121531224 A
[0240] The image stitching and processing system 1106 can process multiple 2D images captured by the image capture device of the user system 1110 and stitch them into a 2D panoramic image. The 2D panoramic image processed by the image stitching and processing system 1106 can have a higher pixel resolution than the 2D panoramic image obtained by the 3D and panoramic capture and stitching system 1102.
[0241] In some embodiments, the image stitching and processing system 106 can receive a 3D representation and output a 3D panoramic image with a higher pixel resolution than the received 3D panoramic image. The panoramic image with higher pixel resolution can be provided to an output device with a higher screen resolution than the user system 1110, such as a computer screen, projector screen, etc. In some embodiments, the panoramic image with higher pixel resolution can be provided to the output device in more detail and can be magnified.
[0242] FIG15 depicts a flowchart showing further details of one step of the 3D and panoramic capture and stitching process of FIG14. In step 1502, the image capture position module 1204 may determine image capture device position data associated with each image captured by the image capture device. The image capture position module 1204 may utilize the IMU of the user system 1110 to determine position data of the image capture device (or the field of view of the lens of the image capture device). The position data may include the orientation, angle, or tilt of one or more image capture devices when capturing one or more 2D images. One or more of the cropping module 1208, the graphic cutting module 1210, or the blending module 1212 may utilize the orientation, angle, or tilt associated with each of the plurality of 2D images to determine how to warp, cut, and / or blend the images.
[0243] In step 1504, the cropping module 1208 may warp one or more of the plurality of 2D images such that two images can be aligned together to form a panoramic image while preserving specific characteristics of the images, such as keeping straight lines straight. The output of the cropping module 1208 may include the number of pixel columns and rows to offset each pixel of the image to straighten the image. The offset of each image can be output as a matrix representing the number of pixel columns and pixel rows offsetting each pixel of the image. In this embodiment, the cropping module 1208 can determine the required amount of distortion for each of the plurality of 2D images based on the image capture pose estimation of each of the plurality of 2D images.
[0244] In step 1506, the graphics cutting module 1210 determines where to cut or slice one or more of the plurality of 2D images. In this embodiment, the graphics cutting module 1210 can determine where to cut or slice each of the plurality of 2D images based on the image capture pose estimation and image distortion of each of the plurality of 2D images.
[0245] In step 1508, the stitching module 1206 may stitch two or more images together using the edges and / or cuts of the images. The stitching module 1206 may align and / or position the images based on objects detected within the images, image distortion, cuts, etc.
[0246] In step 1510, the blending module 1212 may adjust the colors at seams (e.g., the stitching of two images) or at locations on one image that contact or connect to another image. The blending module 1212 may determine the desired amount of color blending based on one or more image capture locations from the image capture location module 1204, image distortion from the cropping module 1208, and graphic cuts from the graphic cutting module 1210.
[0247] The order of one or more steps in the 3D and panoramic capture and stitching process 1400 may be changed without affecting the final result of the 3D panoramic image. For example, an environmental capture system may interleave image capture using an image capture device with LiDAR data or depth information capture. For example, an image capture device can capture an image of segment 1605 of FIG16 of the physical environment, and then LiDAR 612 obtains depth information from segment 1605. Once LiDAR obtains depth information from segment 1605, the image capture device can continue to move to capture an image of another segment 1610, and then LiDAR 612 obtains depth information from segment 1610, thereby interleaving image capture and depth information capture.
[0248] FIG16 depicts a block diagram of an example digital device 1602 according to some embodiments. Any of the user system 1110, the 3D panoramic capture and stitching system 1102, and the image stitching and processor system may include an example of digital device 1602. (See page 34 / 35 of the specification, CN 121531224 A) Digital device 1602 includes processor 1604, memory 1606, storage device 1608, input device 1610, communication network interface 1612, output device 1614, image capture device 1616, and positioning component 1618. Processor 1604 is configured to execute executable instructions (e.g., a program). In some embodiments, processor 1604 includes circuitry or any processor capable of processing executable instructions.
[0249] Memory 1606 stores data. Some examples of memory 1606 include storage devices such as RAM, ROM, RAM cache, virtual memory, etc. In various embodiments, working data is stored in memory 1606. Data in memory 1606 can be erased or eventually transferred to storage device 1608.
[0250] Storage device 1608 includes any storage device configured to retrieve and store data. Storage device 1608...Examples include flash drives, hard disk drives, optical disk drives, and / or magnetic tapes. Each of the memory 1606 and storage device 1608 includes a computer-readable medium storing instructions or programs executable by the processor 1604.
[0251] Input device 1610 is any device that inputs data (e.g., a touch keyboard, stylus). Output device 1614 outputs data (e.g., a speaker, display, virtual reality headset). It should be understood that storage device 1608, input device 1610, and output device 1614 are included. In some embodiments, output device 1614 is optional. For example, a router / switch may include processor 1604 and memory 1606, as well as devices for receiving and outputting data (e.g., communication network interface 1612 and / or output device 1614).
[0252] Communication network interface 1612 may be coupled to a network (e.g., communication network 104). Communication network interface 1612 may support communication via Ethernet connection, serial connection, parallel connection, and / or ATA connection. The communication network interface 1612 may also support wireless communication (e.g., 802.16 a / b / g / n, WiMAX, LTE, Wi-Fi). It will be apparent that the communication network interface 1612 may support many wired and wireless standards.
[0253] Components may be hardware or software. In some embodiments, a component may be configured with one or more processors to perform functions associated with the component. Although different components are discussed herein, it should be understood that a server system may include any number of components performing any or all of the functions discussed herein.
[0254] Digital device 1602 may include one or more image capture devices 1616. One or more image capture devices 1616 may include, for example, an RGB camera, an HDR camera, a video camera, etc. According to some embodiments, one or more image capture devices 1616 may also include a video camera capable of capturing video. In some embodiments, one or more image capture devices 1616 may include an image capture device that provides a relatively standard field of view (e.g., approximately 75°). In other embodiments, one or more image capture devices 1616 may include cameras that provide a relatively wide field of view (e.g., from about 120° to 360°), such as fisheye cameras (e.g., digital device 1602 may include or be included in environment capture system 400).
[0255] Components may be hardware or software. In some embodiments, a component may be configured with one or more processors to perform functions associated with the component. Although different components are discussed herein, it should be understood that a server system may include any number of components that perform any or all of the functions discussed herein. Specification 35 / 35 pages 38 CN 121531224 A Figure 1a Figure 1b DescriptionFigure 1 / 16, page 39, CN 121531224 A; Figure 2, Figure 3, Figure 41, CN 121531224 A; Figure 5, Figure 6A, Figure 6B, Figure 7, Figure 8A, Figure 8B, Figure 9A, Figure 10a, Figure 10b, Figure 10c, Figure 10b, Figure 10c Figure 11 (Page 10 / 16, CN 121531224 A) Figure 12 (Page 12 / 16, CN 121531224 A) Figure 13 (Page 13 / 16, CN 121531224 A) Figure 14 (Page 14 / 16, CN 121531224 A) Figure 15 (Page 15 / 16, CN 121531224 A) Figure 16 (Page 16 / 16, CN 121531224 A) Abstract: A system and method of capturing and generating paramagnetic three-dimensional images is disclosed. An apparatus comprising a housing, a mount configured to be coupled to a motor to horizontally move the apparatus, a wide-angle lens coupled... to the housing, the wide-angle lens being positioned above the mount thereby being along an axis of rotation,the axis of rotation being the axis along which the apparatus rotates, an image capture device within the housing, the image capture device configured to receive two-dimensional images through the wide-angle lens of environment, and a LiDAR device within the housing, the LiDAR device configured to generate depth data based on the environment.
Claims
1. A device for acquiring image and depth information used in an environment, the device comprising: First motor; An image capturing device configured to capture a first plurality of images at different exposures within the field of view of the image capturing device while pointing in a first direction in the environment; as well as Depth information capture device; A first motor is configured to rotate the depth information capturing device and the image capturing device about a first axis until the image capturing device points in a second direction. The depth information capturing device is configured to capture first depth information of a first portion of the environment while being rotated. The image capturing device is also configured to capture, within the environment and pointing in the second direction, a second plurality of images at different exposures that at least partially overlap with one or more of the first plurality of images. The first motor is further configured to rotate the depth information capturing device and the image capturing device about the first axis until the image capturing device points to a third direction. The depth information capturing device is also configured to capture second depth information of a second portion of the environment while being rotated. The image capturing device is also configured to capture a third plurality of images at different exposures that at least partially overlap with one or more of the second plurality of images while pointing in the third direction.
2. The device of claim 1, wherein the depth information capturing device and the image capturing device are configured to rotate about the first axis by the first motor until the image capturing device points to a fourth direction, and while pointing to the fourth direction, the image capturing device is configured to capture a fourth plurality of images at different exposures in the fourth direction, the fourth plurality of images at different exposures captured in the fourth direction at least partially overlapping one or more of the first plurality of images and the third plurality of images.
3. The device of claim 2, wherein the depth information capturing device is configured to capture third depth information of a third portion of the environment while the image capturing device is rotated about the first axis by the first motor until it points to the fourth direction.
4. The device of claim 3, wherein the depth information capturing device that captures depth information of the first, second, and third portions of the environment while being rotated includes the depth information capturing depth information of the first plurality of segments, the second plurality of segments, and the third plurality of segments of the environment while being rotated.
5. The device according to claim 1, further comprising: A communication module is configured to provide one or more of the plurality of images of the first, second, and third portions of the environment to a digital device to form a panoramic image of the environment. The digital device is configured to stitch one or more of the plurality of images of the first, second, and third portions together to form a panoramic image of the environment and to combine the depth information of the first, second, and third portions of the environment with the panoramic image of the environment to generate a 3D image of the environment.
6. The device according to claim 5, further comprising: The digital device is further configured to mix two or more of the first plurality of images, the second plurality of images, and the third plurality of images to generate a first mixed image, a second mixed image, and a third mixed image. The digital device is configured to stitch together one or more of the first plurality of images, the second plurality of images, and the third plurality of images to form the panoramic image of the environment, including stitching together the first composite image, the second composite image, and the third composite image to form the panoramic image of the environment.
7. The device according to claim 1, further comprising: Second motor, The depth information acquisition device includes a light detection and ranging device, namely a LiDAR device, and a reflector. The second motor is configured to cause the reflector to rotate about a second axis while the LiDAR device emits multiple laser pulses toward the reflector. The reflector is configured to provide the multiple laser pulses and receive corresponding multiple reflected laser pulses in one or more rotations about the second axis. The LiDAR device is also configured to receive the multiple reflected laser pulses and generate the depth information therefrom.
8. The device according to claim 6, further comprising: Second motor, The depth information acquisition device includes a light detection and ranging device, namely a LiDAR device, and a reflector. The second motor is configured to rotate the reflector about a second axis while the LiDAR device emits multiple laser pulses toward the reflector. The reflector is configured to provide and receive the multiple laser pulses during one or more rotations about the second axis, and the LiDAR device is further configured to receive the multiple reflected laser pulses and generate the depth information therefrom. The first, second, and third blended images each include a plurality of pixels, wherein each of the plurality of pixels is associated with digital coordinates of the location of that pixel identifying the environment, and wherein each of the reflected laser pulses is similarly associated with corresponding digital coordinates of the location of the depth information generated therefrom. The digital device is further configured to combine the depth information of the environment with the panoramic image of the environment to generate the 3D panoramic image of the environment using at least one set of digital coordinates from the plurality of pixels and at least one set of digital coordinates from the depth information.
9. The device of claim 1, wherein the image capturing device includes a lens located substantially at or near the center of the first axis at a parallax-free point, such that as the first motor rotates the image capturing device about the first axis to capture images while the image capturing device is pointing toward the first direction, the second direction, and the third direction, parallax effects between the first plurality of images, the second plurality of images, and the third plurality of images are reduced or eliminated.
10. The device according to claim 1, further comprising: Second motor, The depth information capturing device includes a light detection and ranging device, namely a LiDAR device, and a reflector. The second motor is configured to rotate the reflector about a second axis while the LiDAR device emits multiple laser pulses toward the reflector. The reflector is configured to provide the multiple laser pulses and receive corresponding multiple reflected laser pulses in one or more rotations about the second axis. The LiDAR device is also configured to receive the multiple reflected laser pulses and generate the depth information therefrom. The device is further configured such that as the image capturing device and the LiDAR device rotate about the first axis via the first motor and the reflector rotates about the second axis via the second motor, the LiDAR device captures the depth information in three dimensions of the environment.