System and method for capturing and generating panoramic three-dimensional image

The device addresses the challenge of capturing 3D renderings in bright and outdoor environments by using a wide-angle lens and LiDAR to efficiently generate 3D renderings with reduced post-processing, enhancing reliability and speed.

JP2025111555APending Publication Date: 2025-07-30MATTERPORT INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025068080
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2025-04-17
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Existing technologies struggle to capture and generate 3D renderings of areas with bright light, such as windows with sunlight or illuminated floors and walls, which appear as holes and require additional post-processing, increasing turnaround time and affecting reliability, and cannot effectively handle outdoor environments without structured lighting.

Method used

A device comprising a housing with a wide-angle lens and a LiDAR device, coupled to motors for horizontal movement and rotation, captures 2D images and generates depth data to create 3D renderings efficiently, using a combination of image capture and LiDAR to handle bright light and outdoor environments.

Benefits of technology

The device effectively captures and generates 3D renderings of both indoor and outdoor environments with bright light, reducing the need for post-processing and improving reliability by integrating image capture and LiDAR technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111555000001_ABST
    Figure 2025111555000001_ABST
Patent Text Reader

Abstract

To enable capture and generation for 3D rendering of structures located in illuminated areas.SOLUTION: A device includes an image capture device and a LiDAR device in a housing. The image capture device is configured to be positioned above a mount configured to be coupled to a motor to move the device horizontally. The device rotates along a rotary shaft. The image capture device is configured to receive a two-dimensional image of the environment through a wide-angle lens. The LiDAR device is configured to generate depth data on the basis of the environment.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the capture and stitching of panoramic images of scenes in a physical environment. Stitching.

Background Art

[0002] Due to the popularity of providing 3D panoramic images of the real world, many solutions have been created that have the function of capturing 2D images and creating 3D images based on the captured 2D images. There are hardware solutions and software applications (i.e., "apps") that can capture multiple 2D images and stitch them into a panoramic image. Create. Capture. Hardware solution. And software applications (i.e., "apps") exist.

Summary of the Invention

Problems to be Solved by the Invention

[0003] There are technologies for capturing and generating 3D data from buildings. However, existing technologies generally cannot capture and generate 3D renderings of areas with bright light. Windows where sunlight shines in, or areas of the floor or wall that are illuminated by bright light, usually appear as holes in the 3D rendering, and additional post-processing work may be required to fill them. This increases the turnaround time of the 3D rendering and improves reliability. Furthermore, since structured lighting cannot be used for capturing 3D images, outdoor environments also pose problems for many existing 3D capture devices. Generate. Area. Appear as holes in the 3D rendering, and additional post-processing work may be required to fill them. This increases the turnaround time of the 3D rendering and improves reliability. Furthermore, since structured lighting cannot be used for capturing 3D images, outdoor environments also pose problems for many existing 3D capture devices. Bring problems.

[0004] Another limitation of existing technologies for capturing and generating 3D data is that 3D panorama. The amount of time required for capturing and processing digital images necessary for generating the ma image is mentioned 。

Means for Solving the Problem

[0005] One exemplary device includes: a housing, and a mount configured to be coupled to a motor for horizontally moving the device ; a wide-angle lens coupled to the housing, the wide-angle lens being positioned above the mount and thus along the axis of rotation, the axis of rotation being the axis about which the device rotates when coupled to the motor ; an image capture device within the housing, the image capture device being configured to receive a two-dimensional image of the environment through the wide-angle lens ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment ; an image capture device within the housing, the image capture device being configured to receive a two-dimensional image of the environment through the wide-angle lens ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment DAR device; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment device.

[0006] The image capture device may include a housing, a first motor, a wide-angle lens, an image sensor, a mount, LiDAR, a second motor, and a mirror. The housing may have a front and a back. The first motor may be coupled to the housing at a first position between the front and the back of the housing, and the first motor is configured to turn the image capture device horizontally by approximately 270° around a vertical axis ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment ; and a LiDAR device within the housing, the LiDAR device being configured to generate depth data based on the environment The wide-angle lens has a field of view away from the front surface of the housing. An image sensor may be coupled to the housing and adapted to capture an image from light received by the wide-angle lens. The mount may be coupled to the first motor and configured to generate an image signal. The LiDAR may be coupled to the housing at a third position. The LiDAR is configured to generate laser pulses and generate a depth signal. The second motor may be coupled to the housing. The mirror may be The second motor may be coupled to a second motor, and the second motor may rotate the mirror about a horizontal axis. the mirror may be configured to receive the laser pulses from the LiDAR. and an angled surface configured to direct the laser pulses about the horizontal axis. nothing.

[0007] In some embodiments, the image sensor detects when the image capture device is stationary. The camera is configured to generate a first plurality of images at different exposures while facing a first direction. The first motor drives the image capture device after generating the first plurality of images. The device may be configured to turn about the vertical axis. wherein the image sensor detects when the first motor turns the image capture device. The LiDAR does not generate images while the first motor is moving, and the LiDAR While the device is turning, a depth signal is generated based on the laser pulses. The image sensor captures the image when the image capture device is stationary and facing a second direction. and generating a second plurality of images at the different plurality of exposures, The first motor is configured to turn the image capture device 90° around the vertical axis after the generation of the second plurality of images. The image sensor may be configured to generate a third plurality of images with different exposures when the image capture device is stationary and facing a third direction, and the first motor is configured to turn the image capture device 90° around the vertical axis after the generation of the third plurality of images. The image sensor may be configured to generate a fourth plurality of images with different exposures when the image capture device is stationary and facing a fourth direction, and the first motor is configured to turn the image capture device 90° around the vertical axis after the generation of the fourth plurality of images. In some embodiments, the system may further comprise a processor configured to blend frames of the first plurality of images before the image sensor generates the second plurality of images. The remote digital device may communicate with the image capture device and may be configured to generate a 3D visualization based on the first, second, third, and fourth pluralities of images and the depth signal. The remote digital device may be configured to generate the 3D visualization without using images other than the first, second, third, and fourth pluralities of images. In some embodiments, the first, second, third, and fourth pluralities of images are generated during turns that combine a plurality of turns that turn the image capture device 270° around the vertical axis. The water

[0008] The speed or rotation of the mirror about the horizontal axis increases as the first motor performs the image capture while turning the vice. The angled surface of the mirror may be 90°. In some embodiments, the LiDAR emits the laser pulses in a direction opposite to the front face of the housing. [[ID=⑧]]

[0009] An exemplary method includes: receiving light from a wide-angle lens of an image capture device, wherein the wide-angle lens is coupled to a housing of the image capture device, and the light is received within a field of view of the wide-angle lens that extends away from the front face of the housing; generating, using the light from the wide-angle lens, a first plurality of images with an image sensor of the image capture device, wherein the image sensor is coupled to the housing and the first plurality of images are at different exposures; turning the image capture device substantially 270° horizontally about a vertical axis by a first motor, wherein the first motor is coupled to the housing at a first position between the front face and the back face of the housing, the wide-angle lens is at a second position along the vertical axis, and the second position is the nodal point; rotating, by a second motor, a mirror having an angled surface about a horizontal axis, wherein the second motor is coupled to the housing; generating laser pulses by LiDAR, wherein the LiDAR is coupled to the housing at a third position, and the laser pulses are directed at the rotating mirror while the image capture device is turning horizontally; and the laser ​ Including the step of generating a depth signal by the LiDAR based on a pulse.

[0010] The step of generating the first plurality of images by the image sensor may be performed before the image capture device turns horizontally. In some embodiments, the image sensor does not generate an image while the first motor is turning the image capture device, and the LiDAR generates the depth signal based on the laser pulse while the first motor is turning the image capture device. The method may further include: when the image capture device is stationary and facing a second direction, generating a second plurality of images by the image sensor with different exposures; and after generating the second plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of The method may further include: when the image capture device is stationary and facing a second direction, generating a second plurality of images by the image sensor with different exposures; and after generating the second plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of

[0011] The method may further include: when the image capture device is stationary and facing a second direction, generating a second plurality of images by the image sensor with different exposures; and after generating the second plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of The method may further include: when the image capture device is stationary and facing a second direction, generating a second plurality of images by the image sensor with different exposures; and after generating the second plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of

[0012] In some embodiments, the method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of In some embodiments, the method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of In some embodiments, the method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of In some embodiments, the method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of In some embodiments, the method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of In some embodiments, the method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of In some embodiments, the method may further include: when the image capture device is stationary and facing a third direction, generating a third plurality of images by the image sensor with different exposures; and after generating the third plurality of images, turning the image capture device 90° around the vertical axis by the first motor. The method may further include generating a fourth plurality of images by the image sensor with different exposures when the image capture device is stationary and facing a fourth direction. The method includes the first, second, third, and fourth pluralities of Using the image and based on the depth signal, steps for generating 3D visualization may be included, and the steps for generating the 3D visualization do not use any other images. In some embodiments, the method may further include blending frames of the first plurality of images before the image sensor generates the second plurality of images. The first, second, third, and fourth pluralities of images can be generated during turns that combine a plurality of turns of rotating the image capture device 270° around the vertical axis. In some embodiments, the speed or rotation of the mirror around the horizontal axis increases when the first motor turns the image capture device.

[0013] In some embodiments, the method may further include blending frames of the first plurality of images before the image sensor generates the second plurality of images. The first, second, third, and fourth pluralities of images can be generated during turns that combine a plurality of turns of rotating the image capture device 270° around the vertical axis. In some embodiments, the speed or rotation of the mirror around the horizontal axis increases when the first motor turns the image capture device. The first, second, third, and fourth pluralities of images can be generated during turns that combine a plurality of turns of rotating the image capture device 270° around the vertical axis. In some embodiments, the speed or rotation of the mirror around the horizontal axis increases when the first motor turns the image capture device. The first, second, third, and fourth pluralities of images can be generated during turns that combine a plurality of turns of rotating the image capture device 270° around the vertical axis. In some embodiments, the speed or rotation of the mirror around the horizontal axis increases when the first motor turns the image capture device. In some embodiments, the speed or rotation of the mirror around the horizontal axis increases when the first motor turns the image capture device. The first, second, third, and fourth pluralities of images can be generated during turns that combine a plurality of turns of rotating the image capture device 270° around the vertical axis. In some embodiments, the speed or rotation of the mirror around the horizontal axis increases when the first motor turns the image capture device.

Brief Description of the Drawings

[0014]

Figure 1a

Figure 1b

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6a

Figure 6b

Figure 7

Figure 8a

Figure 8b

Figure 9a

Figure 9b

Figure 10a - 10c

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

DETAILED DESCRIPTION OF THE INVENTION

[0015] Many of the innovations described in this specification are made with reference to the drawings. Like reference numbers are used to refer to like elements. In the following description, for purposes of explanation, a number of specific details are set forth in order to provide a thorough understanding. However, it may be apparent that different innovations may be practiced without these specific details. In other instances, well-known structures and components are shown in block diagram form in order to facilitate the description of the innovations.

[0016] Various embodiments of the apparatus provide 3D panoramic images of indoor and outdoor environments to a user. In some embodiments, the apparatus can efficiently and quickly provide 3D panoramic images of indoor and outdoor environments to a user using a single wide field-of-view (FOV ) lens and a single light and detection and ranging sensor (LiDAR sensor).

[0017] The following is an exemplary usage example of an exemplary apparatus described in this specification. The following usage example is one of a number of embodiments. As described in this specification, different embodiments of the apparatus may include one or more features and functions similar to this usage example.

[0018] FIG. 1a is a dollhouse view 100 of an exemplary environment, such as a house, according to some embodiments. The dollhouse view 100 provides an overall view of the exemplary environment captured by an (environmental capture system described in this specification). A user can switch between a plurality of different views of this exemplary environment to view the dollhouse view on a user system. One can interact with 100. For example, a user can interact with area 110 to trigger a floor plan of a house as shown in FIG. 1b. In some embodiments, the user can interact with the icons within the dollhouse view 100, such as icons 120, 130, 140, to provide a walkthrough view, a floor plan, or a measurement view (for example, for a 3D walkthrough), respectively. FIG. 1b shows a floor plan of the first floor of a house according to some embodiments. This floor plan is a top - down view of the first floor of the house. The user can interact with an area of this floor plan, such as area 150, to trigger

[0019] a view from eye level of a specific part of this floor plan, such as the living room. An example of a view from eye level of the living room can be confirmed in FIG. 2, which can be part of a virtual walkthrough. The user can interact with a part of the floor plan 200 corresponding to area 150 of FIG. 1b. The user can move the view around the room as if the user were actually inside this living room. In addition to a horizontal 360° view of the living room, the user can also view or operate on the floor or ceiling of the living room. Further, the user can interact with specific areas of the above - mentioned part of the floor plan 200, such as areas 210, 220, to pass through the living room towards other parts of the house. When the user interacts with area 220, the environmental capture system can provide a walking - like transition between an area of the house that roughly corresponds to the area of the house indicated by area 150 and an area of the house that roughly corresponds to the area of the house indicated by area 220.

[0020] The user may interact with a portion of the floor plan 200 corresponding to area 150 of FIG. 1b. The user can move the view around the room as if the user were actually inside this living room. In addition to a horizontal 360° view of the living room, the user can also view or operate on the floor or ceiling of the living room. Further, the user can interact with specific areas of the above - mentioned portion of the floor plan 200, such as areas 210, 220, to pass through the living room towards other parts of the house. When the user interacts with area 220, the environmental capture system can provide a walking - like transition between an area of the house that roughly corresponds to the area of the house indicated by area 150 and an area of the house that roughly corresponds to the area of the house indicated by area 220. The user can interact with specific areas of the above - mentioned portion of the floor plan 200, such as areas 210, 220, to pass through the living room towards other parts of the house. When the user interacts with area 220, the environmental capture system can provide a walking - like transition between an area of the house that roughly corresponds to the area of the house indicated by area 150 and an area of the house that roughly corresponds to the area of the house indicated by area 220. When the user interacts with area 220, the environmental capture system can provide a walking - like transition between an area of the house that roughly corresponds to the area of the house indicated by area 150 and an area of the house that roughly corresponds to the area of the house indicated by area 220. The user can interact with area 220, and the environmental capture system can provide a walking - like transition between an area of the house that roughly corresponds to the area of the house indicated by area 150 and an area of the house that roughly corresponds to the area of the house indicated by area 220. [[ID=...]]

[0021] [[ID=...]] ​​Figure 3 shows an example of an environmental capture system 300 according to some embodiments. The environment capture system 300 includes a lens 310, a housing 320, a mount attachment tor 330, and a movable cover 340.

[0022] In use, the environmental capture system 300 may be positioned in an environment such as a room. The environmental capture system 300 may be positioned on a support (e.g., a tripod). The movable cover 340 may be moved to expose the LiDAR and the mirror that can rotate at high speed. When activated the environmental capture system 300 can take a burst of images and then turn using a motor The environmental capture system 300 can turn on the mount attachment 330 When turning, the LiDAR may perform measurements (during turning, the environmental capture system cannot take images). When facing a new direction, the environmental capture system can take a burst of images and then turn to the next direction.

[0023] For example, after positioning, the user may instruct the environmental capture system 300 to start a sweep The sweep may be as follows: (1) Estimation of exposure and subsequent capture of HDR RGB images 90° rotation, capture of depth data (2) Estimation of exposure and subsequent capture of HDR RGB images 90° rotation, capture of depth data (3) Estimation of exposure and subsequent capture of HDR RGB images 90° rotation, capture of depth data (4) Estimation of exposure and subsequent capture of HDR RGB images 90° rotation (total 360°), capture of depth data

[0024] For each burst, there may be any number of images with different exposures. The environmental capture system can blend any number of images of one burst into one while waiting for another frame and / or waiting for the next burst.

[0025] The housing 320 may protect the electronic components of the environmental capture system 300 and may also be provided with a power button, a scan button, etc. for interface with the user. For example, the housing 320 may include a movable cover 340, which may be movable to remove the cover of the LiDAR. Further, the housing 320 may include electronic interfaces such as a power adapter and an indicator light. In some embodiments, the housing 320 is a molded plastic housing. In various embodiments, the housing 320 is a combination of one or more of plastic, metal, and polymer.

[0026] The lens 310 may be part of a lens assembly. Further details of the lens assembly can be described in the description of FIG. 7. The lens 310 is strategically positioned at the center of the rotation axis 305 of the environmental capture system 300. In this example, the rotation axis 305 is in the x-y plane. By positioning the lens 310 at the center of the rotation axis 305, the parallax effect can be eliminated or reduced. Parallax is an error caused by the rotation of the image capture device around the non-parallax point (NPP). In this example, the NPP can be identified at the center of the entrance pupil of the lens.

[0027] For example, assume that a panoramic image of the physical environment is generated using four images captured by the environmental capture system 300, where there is a 25% overlap between the images of the panoramic image. Here, when there is no parallax, 25% of an image can be accurately overlapped with another image of the same area of this physical environment. Eliminating or reducing the parallax effect of a plurality of images captured by the image sensor via the lens 310 can assist in stitching the plurality of images into a 2D panoramic image. For example, assume that a panoramic image of the physical environment is generated using four images captured by the environmental capture system 300, where there is a 25% overlap between the images of the panoramic image. Here, when there is no parallax, 25% of an image can be accurately overlapped with another image of the same area of this physical environment. Eliminating or reducing the parallax effect of a plurality of images captured by the image sensor via the lens 310 can assist in stitching the plurality of images into a 2D panoramic image. For example, assume that a panoramic image of the physical environment is generated using four images captured by the environmental capture system 300, where there is a 25% overlap between the images of the panoramic image. Here, when there is no parallax, 25% of an image can be accurately overlapped with another image of the same area of this physical environment. Eliminating or reducing the parallax effect of a plurality of images captured by the image sensor via the lens 310 can assist in stitching the plurality of images into a 2D panoramic image. For example, assume that a panoramic image of the physical environment is generated using four images captured by the environmental capture system 300, where there is a 25% overlap between the images of the panoramic image. Here, when there is no parallax, 25% of an image can be accurately overlapped with another image of the same area of this physical environment. Eliminating or reducing the parallax effect of a plurality of images captured by the image sensor via the lens 310 can assist in stitching the plurality of images into a 2D panoramic image. For example, assume that a panoramic image of the physical environment is generated using four images captured by the environmental capture system 300, where there is a 25% overlap between the images of the panoramic image. Here, when there is no parallax, 25% of an image can be accurately overlapped with another image of the same area of this physical environment. Eliminating or reducing the parallax effect of a plurality of images captured by the image sensor via the lens 310 can assist in stitching the plurality of images into a 2D panoramic image. For example, assume that a panoramic image of the physical environment is generated using four images captured by the environmental capture system 300, where there is a 25% overlap between the images of the panoramic image. Here, when there is no parallax, 25% of an image can be accurately overlapped with another image of the same area of this physical environment. Eliminating or reducing the parallax effect of a plurality of images captured by the image sensor via the lens 310 can assist in stitching the plurality of images into a 2D panoramic image.

[0028] The lens 310 may include a wide field of view (for example, the lens 310 may be a fisheye lens). In some embodiments, the lens may have a horizontal FOV (HFOV) of at least 148° and a vertical FOV (VFOV) of at least 94°. The lens 310 may include a wide field of view (for example, the lens 310 may be a fisheye lens). In some embodiments, the lens may have a horizontal FOV (HFOV) of at least 148° and a vertical FOV (VFOV) of at least 94°. The lens 310 may include a wide field of view (for example, the lens 310 may be a fisheye lens). In some embodiments, the lens may have a horizontal FOV (HFOV) of at least 148° and a vertical FOV (VFOV) of at least 94°.

[0029] The mount attachment 330 can enable the environmental capture system 300 to be attached to a mount. The above mount can enable the environmental capture system 300 to be coupled to a tripod, a flat surface, or a motorized mount (for example, for moving the environmental capture system 300). In some embodiments, the above mount can enable the environmental capture system 300 to rotate along a horizontal axis. The mount attachment 330 can enable the environmental capture system 300 to be attached to a mount. The above mount can enable the environmental capture system 300 to be coupled to a tripod, a flat surface, or a motorized mount (for example, for moving the environmental capture system 300). In some embodiments, the above mount can enable the environmental capture system 300 to rotate along a horizontal axis. The mount attachment 330 can enable the environmental capture system 300 to be attached to a mount. The above mount can enable the environmental capture system 300 to be coupled to a tripod, a flat surface, or a motorized mount (for example, for moving the environmental capture system 300). In some embodiments, the above mount can enable the environmental capture system 300 to rotate along a horizontal axis. The mount attachment 330 can enable the environmental capture system 300 to be attached to a mount. The above mount can enable the environmental capture system 300 to be coupled to a tripod, a flat surface, or a motorized mount (for example, for moving the environmental capture system 300). In some embodiments, the above mount can enable the environmental capture system 300 to rotate along a horizontal axis. The mount attachment 330 can enable the environmental capture system 300 to be attached to a mount. The above mount can enable the environmental capture system 300 to be coupled to a tripod, a flat surface, or a motorized mount (for example, for moving the environmental capture system 300). In some embodiments, the above mount can enable the environmental capture system 300 to rotate along a horizontal axis. The mount attachment 330 can enable the environmental capture system 300 to be attached to a mount. The above mount can enable the environmental capture system 300 to be coupled to a tripod, a flat surface, or a motorized mount (for example, for moving the environmental capture system 300). In some embodiments, the above mount can enable the environmental capture system 300 to rotate along a horizontal axis.

[0030] In some embodiments, the environmental capture system 300 may include a motor for horizontally turning the environmental capture system 300 around the mount attachment 330. In some embodiments, the environmental capture system 300 may include a motor for horizontally turning the environmental capture system 300 around the mount attachment 330. In some embodiments, the environmental capture system 300 may include a motor for horizontally turning the environmental capture system 300 around the mount attachment 330.

[0031] In some embodiments, a motorized mount turns the environmental capture system 300 around a horizontal axis , may be moved along the vertical axis, or both. In some embodiments, the above-mentioned electric mount can rotate or move within the x-y plane. Using the mount attachment 330 , the environmental capture system 300 can be coupled to an electric mount, a tripod, etc. to reduce or minimize shaking by stabilizing the environmental capture system 300. In another example, the mount attachment 330 may be coupled to an electric mount that can rotate the 3D environmental capture system 300 at a stable known speed, which supports the determination of the (x, y, z) coordinates of each laser pulse of the LiDAR.

[0032] FIG. 4 shows a schematic view of an environmental capture system 400 in some embodiments. This schematic view shows the environmental capture system 400 (which can be an example of the environmental capture system 300 in FIG. 3) from various views, such as a front view 410, a top view 420, a side view 430, and a rear view 440. In these schematic views, the environmental capture system 400 may include any hollow portion shown in the side view 430.

[0033] In some embodiments, the environmental capture system 400 has a width of 75 mm, a height of 180 mm , and a depth of 189 mm. It will be understood that the environmental capture system 400 may have any width, height, or depth. In various embodiments, the ratio of the width to the depth in the first example is maintained regardless of the specific measurements.

[0034] The housing of the 3D environmental capture system 400 is for the environmental capture system 400 The electronic components can be well protected, and an interface for interaction with the user (e.g., the rear view 4 0's screen) can be provided. Further, the housing may include electronic interfaces such as a power adapter and an indicator la ight. In some embodiments, the housing is a molded plastic housing. In various embodiments, the housing is a combina tion of one or more of plastic, metal, and polymer. The environmental capture system 400 may include a movable cover, which may be movable to release the LiDAR cover and to protect the LiDAR from multiple elements when not in use. The lens shown in the front view 410 may be part of a lens assembly. Similar to the environmental capture system 300, the lens of the environmental capture system 400 is strategically located at the center of the rotation axis 3

[0035] 05. The lens may include a wide field of view. In various embodiments, the lens shown in the front view 410 is concave, and the housing is flared such that the wide-angle lens is exactly at the blind spot (e.g., directly above the midpoint of the mount and / or motor), yet an image can be taken without interference from the housing. The mount attachment at the base of the environmental capture system 400 can enable the environmental capture system to be attached to the mount. The mount can enable the environmental capture system 400 to be coupled to a tripod, a flat surface, or a motorized mount (e.g., for moving the environmental capture system 400). In some

[0036] embodiments, the mount can enable the environmental capture system 400 to be mounted to the mount The mount can enable the environmental capture system 400 to be coupled to a tripod, a flat surface, or a motorized mount (e.g., for moving the environmental capture system 400). In some embodiments, the mount can enable the environmental capture system 400 to be mounted to the mount It may be coupled to an internal motor for turning around.

[0037] In some embodiments, the mount can enable the environmental capture system 400 to rotate along a horizontal axis. In various embodiments, an electric mount can move the environmental capture system 400 along a horizontal axis, a vertical axis, or both. Using a mount attachment, the environmental capture system 400 can be coupled to an electric mount, a tripod, etc. to stabilize the environmental capture system 400, thereby reducing or minimizing shaking. In another example, the mount attachment may be coupled to an electric mount that can rotate the environmental capture system 400 at a stable known speed, which assists the LiDAR in determining the (x, y, z) coordinates of each laser pulse of the LiDAR. In view 430, the mirror 450 is exposed. The LiDAR may emit laser pulses towards the mirror (in a direction opposite to the view of the lens). The laser pulses can hit the mirror 450, which may be angled (e.g., at an angle of 90°). The mirror 450 may be coupled to an internal motor that turns the mirror, whereby the laser pulses of the LiDAR can be emitted and / or received at a number of different angles around the environmental capture system 400. Figure 5 is a diagram of laser pulses from a LiDAR around the environmental capture system 400 in some embodiments. In this example, the laser pulses are from a rapidly rotating mirror.

[0038]

[0039] It is emitted at 450. The laser pulse may be emitted and received perpendicular to the horizontal axis 6 of the environmental capture system 400 02 (see FIG. 6). The laser pulses from the LiDAR may be directed in a direction away from the environmental capture system 400. An angle may be provided to the mirror 450 such that the laser pulses from the LiDAR are directed away from the environmental capture system 400. In some examples, the angle of the angled surface of the mirror may be 90° or may be 60°, 120°, or 60° - 120°. [[ID=TO]]

[0040] In some embodiments, when the environmental capture system 400 is stationary and operating, the environmental capture system 400 can capture a burst of images through a lens The environmental capture system 400 may turn on a horizontal motor between bursts of images. While turning along the mount, the LiDAR of the environmental capture system 400 may emit and / or receive laser pulses that hit the fast - rotating mirror 450. The LiDAR may generate a depth signal from the reflection of the received laser pulses and / or generate depth data.

[0041] In some embodiments, the depth data may be associated with coordinates related to the environmental capture system 400. Similarly, by associating pixels or portions of an image with coordinates related to the environmental capture system 400 3D visualization (such as images from multiple different directions, 3D walk - throughs, etc.) using the image and depth data can be created

[0042] As shown in FIG. 5, the pulses of the LiDAR are of the environmental capture system 400 It can be blocked by the bottom. While the environmental capture system 400 moves around the mount during, the mirror 450 can rotate continuously at high speed, or when the environmental capture system 400 starts to move, and when the environmental capture system 400 decelerates and stops again, the mirror 450 can rotate at high speed more slowly (for example, a constant speed can be maintained during the start and stop of the mount motor) will be understood.

[0043] The LiDAR can receive depth data from the above pulses. Due to the movement of the environmental capture system 40 0 and / or the increase and decrease of the speed of the mirror 450, the density of the depth data regarding the environmental capture system 400 is not consistent (for example, the density is high in some areas and low in other areas).

[0044] FIG. 6a shows a side view of the environmental capture system 400. The mirror 450 is shown in this figure, and this mirror 450 can rotate at high speed around the horizontal axis. The pulse 604 can be emitted by the LiDAR in the mirror 450 that rotates at high speed, and can also be emitted perpendicular to the horizontal axis 60 2. Similarly, the pulse 604 can be received by the LiDAR in a similar manner as well.

[0045] Although the LiDAR pulse is described as being perpendicular to the horizontal axis 602, the LiDAR pulse can be at any angle with respect to the horizontal axis 602 (for example, the angle of the mirror can be any angle including 60 - 120°) will be understood. In various embodiments, the LiDAR is on the front side of the environmental capture system 400 (for example, the front side side On the opposite side of (604) (e.g., in a direction opposite to the center of the lens's field of view or towards the back surface 606), a pulse is emitted. in a direction), a pulse is emitted.

[0046] As described herein, the environmental capture system 400 may turn around the vertical axis 608. In various embodiments, the environmental capture system 400 turns 90° after taking an image. When the environmental capture system 400 completes a 270° turn from the original starting position where the first set of images was taken, a fourth set of images is taken. Thus, the environmental capture system 400 can generate four sets of images during a total of 270° of multiple turns (assuming, for example, that the first set of images was taken before the first turn of the environmental capture system 400). In various embodiments, the images (e.g., four sets of images) from a single sweep of the environmental capture system 400 (e.g., taken during one full rotation or a 270° rotation around the vertical axis) are sufficient to generate 3D visualization without using further sweeps or turns of the environmental capture system 400, together with depth data acquired during the same sweep. After taking an image, the environmental capture system 400 turns 90°. When the environmental capture system 400 completes a 270° turn from the original starting position where the first set of images was taken, a fourth set of images is taken. Thus, the environmental capture system 400 can generate four sets of images during a total of 270° of multiple turns (assuming, for example, that the first set of images was taken before the first turn of the environmental capture system 400). In various embodiments, the images (e.g., four sets of images) from a single sweep of the environmental capture system 400 (e.g., taken during one full rotation or a 270° rotation around the vertical axis) are sufficient to generate 3D visualization without using further sweeps or turns of the environmental capture system 400, together with depth data acquired during the same sweep. It should be understood that in this example, the LiDAR pulse is emitted and redirected by a mirror that rotates at high speed at a position away from the rotation point of the environmental capture system 400. In this example, the distance from the rotation point of the mount is 608 (e.g., the lens may be at the nodal point, or the lens may be at a position behind the front of the environmental capture system 400). The LiDAR pulse is redirected by a mirror 450 at a position away from the rotation point. In various embodiments, the images (e.g., four sets of images) from a single sweep of the environmental capture system 400 (e.g., taken during one full rotation or a 270° rotation around the vertical axis) are sufficient to generate 3D visualization without using further sweeps or turns of the environmental capture system 400, together with depth data acquired during the same sweep. In various embodiments, the images (e.g., four sets of images) from a single sweep of the environmental capture system 400 (e.g., taken during one full rotation or a 270° rotation around the vertical axis) are sufficient to generate 3D visualization without using further sweeps or turns of the environmental capture system 400, together with depth data acquired during the same sweep. In various embodiments, the images (e.g., four sets of images) from a single sweep of the environmental capture system 400 (e.g., taken during one full rotation or a 270° rotation around the vertical axis) are sufficient to generate 3D visualization without using further sweeps or turns of the environmental capture system 400, together with depth data acquired during the same sweep. In various embodiments, the images (e.g., four sets of images) from a single sweep of the environmental capture system 400 (e.g., taken during one full rotation or a 270° rotation around the vertical axis) are sufficient to generate 3D visualization without using further sweeps or turns of the environmental capture system 400, together with depth data acquired during the same sweep. In various embodiments, the images (e.g., four sets of images) from a single sweep of the environmental capture system 400 (e.g., taken during one full rotation or a 270° rotation around the vertical axis) are sufficient to generate 3D visualization without using further sweeps or turns of the environmental capture system 400, together with depth data acquired during the same sweep. In various embodiments, the images (e.g., four sets of images) from a single sweep of the environmental capture system 400 (e.g., taken during one full rotation or a 270° rotation around the vertical axis) are sufficient to generate 3D visualization without using further sweeps or turns of the environmental capture system 400, together with depth data acquired during the same sweep.

[0047] In this example, it will be understood that the LiDAR pulse is emitted and redirected by a mirror that rotates at high speed at a position away from the rotation point of the environmental capture system 400. In this example, it will be understood that the LiDAR pulse is emitted and redirected by a mirror that rotates at high speed at a position away from the rotation point of the environmental capture system 400. In this example, the distance from the rotation point of the mount is 608 (e.g., the lens may be at the nodal point, or the lens may be at a position behind the front of the environmental capture system 400). The LiDAR pulse is redirected by a mirror 450 at a position away from the rotation point. In this example, the distance from the rotation point of the mount is 608 (e.g., the lens may be at the nodal point, or the lens may be at a position behind the front of the environmental capture system 400). The LiDAR pulse is redirected by a mirror 450 at a position away from the rotation point. In this example, the distance from the rotation point of the mount is 608 (e.g., the lens may be at the nodal point, or the lens may be at a position behind the front of the environmental capture system 400). The LiDAR pulse is redirected by a mirror 450 at a position away from the rotation point. Since the direction is changed, the LiDAR cannot receive depth data from the cylinder extending from above the environmental capture system 400 to below the environmental capture system 400. In this example, the radius of the above cylinder (for example, this cylinder lacks depth information) can be measured from the center of the rotation point of the motor mount to the point where the mirror 450 deflects the LiDAR pulse.

[0048] Furthermore, Figure 6a shows a cavity 610. In this example, the environmental capture system 400 includes a rapidly rotating mirror within the housing of the environmental capture system 400. There is a notch section from the housing. After reflecting the laser pulse out of the housing by the mirror, the reflection can be received by the mirror and deflected back to the LiDAR, enabling the LiDAR to generate a depth signal and / or depth data. The base of the body of the environmental capture system 400 below the cavity 610 can block a part of the laser pulse. The cavity 610 can be defined by the base of the environmental capture system 400 and the rotating mirror. As shown in Figure 6b, there may still be a space between the angled mirror and the housing of the environmental capture system 400 including the LiDAR.

[0049] In various embodiments, the LiDAR is configured to stop emitting laser pulses when the rotation speed of the mirror drops below a rotation safety threshold (for example, when the motor that rotates the mirror at high speed fails, or when the mirror is held in a predetermined position). This way, the LiDAR can be configured for safety, and the laser pulse is in the same direction (for example, towards the user's It is possible to reduce the possibility of continuous emission into the eyes.

[0050] Figure 6b shows a top view of an environmental capture system 400 according to some embodiments. In this example, the front of the environmental capture system 400 is shown concave with a lens directly above the center of the rotation point (e.g., directly above the center of the mount). The front of the camera is concave for the lens, and the front of the housing is flared so that the housing does not block the field of view of the image sensor. The mirror 450 is shown facing upward.

[0051] Figure 7 shows a schematic diagram of the components of an environmental capture system 300 according to some embodiments. The environmental capture system 700 includes a front cover 702, a lens assembly 704, a structural frame 706, a LiDAR 708, a front housing 710, a mirror assembly 712, a GPS antenna 714, a rear housing 716, a vertical motor 718, a display 720, a battery pack 722, a mount 724, and a horizontal motor 726.

[0052] In various embodiments, the environmental capture system 700 can be configured to perform 3D mesh scanning, alignment, and creation outdoors and indoors on sunny days. This eliminates the barriers when adopting other systems that are indoor-only tools. The environmental capture system 700 can scan a large space more quickly than other devices. In some embodiments, the environmental capture system 700 can provide improved depth accuracy by improving the single scan depth accuracy at 90m.

[0053] In some embodiments, the weight of the environmental capture system 700 is 1 kg or about 1 kg and may be. In one example, the weight of the environmental capture system 700 may be 1 to 3 kg .

[0054] The front cover 702, the front housing 710, and the rear housing 716 constitute part of the housing. In one example, the width w of the front cover may be 75 mm .

[0055] The lens assembly 704 may include a camera lens that focuses light onto the image capture device . The image capture device can capture an image of the physical environment. The user may place the environmental capture system 700 to capture a portion of the floor of a building such as the second building 422 in FIG. 1 and obtain a panoramic image of the above portion of the above floor . By moving the environmental capture system 700 to another portion of the above floor of the above building , a panoramic image of another portion of the above floor can be obtained. In one example , the depth of field of the image capture device is from 0.5 meter to infinity. FIG. 8 a shows the exemplary lens dimensions in some embodiments .

[0056] In some embodiments, the image capture device is a complementary metal-oxide-semiconductor (co mplementary metal-oxide-semiconductor: CM OS) image sensor (e.g., Sony IMX283 ~20 Megapixel CMOS MIPI sensor equipped with an Nvidia Jetson Nano SOM ). In various embodiments, the image capture device is a charge-coupled device (charged ). coupled device:CCD). In one example, the image capture device is a red - green - blue (RGB) sensor. In one embodiment the image capture device is an infrared (IR) sensor. The lens assembly 70 4 can provide a wide field of view to the image capture device.

[0057] The image sensor may have many different specifications. In one example, the image sensor includes the following:

[0058]

Table 1

[0059] Exemplary specifications may be as follows:

[0060]

Table 2

[0061] In various embodiments, looking at the MTF at the F0 relative field of view (i.e., the center), the focus shift can vary from +28 micrometers at 0.5 m to -25 micrometers at infinity, and the overall focus shift is 53 micrometers.

[0062] Figure 8b shows exemplary lens design specifications in some embodiments.

[0063] In some examples, the lens assembly 704 has an HFOV of at least 148° and a VFOV of at least 94°. In one example, the lens assembly 704 has a field of view of 150 °, 180°, or 145° - 180°. The environmental capture system 700 In one example, the 360° view image capture around can be obtained by three or four separate image captures from 700 image capture devices. In various embodiments, the image capture device may have a resolution of at least 37 pixels per degree. In some embodiments, the environmental capture system 700 may include a lens cap (not shown) for protecting the lens assembly 704 when not in use. The output of the lens assembly 704 may be a digital image of an area of the physical environment. By stitching together a plurality of images captured by the lens assembly 704, a 2D panoramic image of the physical environment can be formed. A 3D panorama can be generated by combining depth data captured by the LiDAR 708 with a 2D panoramic image generated by stitching together a plurality of images from the lens assembly 704. In some embodiments, a plurality of images captured by the environmental capture system 402 are stitched together by the image processing system 406. In various embodiments, the environmental capture system 402 generates a "preview" or "thumbnail" version of the 2D panoramic image. The preview or thumbnail version of the 2D panoramic image can be presented on the user system 1110 such as an iPad®, personal computer, smartphone, etc. In some embodiments, the environmental capture system 402 may generate a minimap of the physical environment representing an area of the physical environment. In various embodiments, the image processing system 406 generates a minimap representing an area of the physical environment. In some embodiments, the environmental capture system 402 generates a "preview" or "thumbnail" version of the 2D panoramic image. The preview or thumbnail version of the 2D panoramic image can be presented on the user system 1110 such as an iPad®, personal computer, smartphone, etc.

[0064] The image captured by the lens assembly 704 identifies the location where the 2D image was captured. For example, in some implementations, the capture device location data may include capture device location data that indicates or indicates the capture device location. The capture device location data is stored in the Global Positioning System (GPS) associated with the 2D image. Global Positioning System (GPS) coordinates. In other implementations, the capture device location data is stored in the capture device (e.g., the position of the capture device (camera and / or 3D sensor) relative to its environment, e.g. The relative position of the device to an object in the environment, another device in the environment, etc. In some implementations, this type of The location data of the capture device (e.g., camera and / or positioning hardware) a device operably coupled to a camera with hardware and / or software Thus, it can be determined in connection with the capture of an image and received along with the image. The placement of the lens assembly 704 is not just a matter of design. By placing it at or near the center, the parallax effect can be reduced.

[0065] In some embodiments, the structural frame 706 supports the lens assembly 704 and the LiD The AR708 is held in a specific position to protect the components of this example environmental capture system. The structural frame 706 provides a solid installation for the LiDAR 708. It can assist in the placement of the LiDAR 708 in a fixed location. The fixed position of the lens assembly 704 and LiDAR 708 allows for the capture of depth data. Enabling a fixed relationship to assist in creating a 3D image aligned with image information . By aligning the 2D image data and depth data captured in the physical environment with a common 3D coordinate space, a 3D model of the physical environment can be generated.

[0066] In various embodiments, LiDAR 708 captures depth information of the physical environment . When the user places the environment capture system 700 on a part of a floor with a second building , LiDAR 708 can obtain depth information of the object. LiDAR 70 8 may include an optical sensing module, which uses pulses from a laser to target or irradiate a scene and measures the time it takes for photons to travel to the target and return to LiDAR 708 , thereby measuring the distance to an object in the target or scene. Subsequently , the measurement values may be converted to a grid coordinate system using information derived from the horizontal drive train of the environment capture system 700 .

[0067] In some embodiments, LiDAR 708 can return depth data points every 10 microseconds with a time stamp (of the internal clock) every 10 microseconds . LiD AR 708 can sample a partial sphere (with small holes at the top and bottom) every 0.25° with a ring. With data points every 10 microseconds and every 0.25°, in some embodiments , each "disk" of multiple points can be 14.40 milliseconds, and there can be 1440 disks to form a sphere that is nominally 20.7 seconds . Since each disk captures before and after, the sphere can be captured with a 180° sweep .

[0068] In one example, the specifications of the LiDAR 708 may be as follows:

[0069]

Table 3

[0070] One advantage of using LiDAR is that by using LiDAR at a relatively low wavelength (e.g., 905 nm , 900 - 940 nm, etc.), the environmental capture system 700 can determine depth information regarding the outdoor environment or a brightly lit indoor environment.

[0071] Depending on the arrangement of the lens assembly 704 and the LiDAR 708, the environmental capture system 700 or a digital device can be made to communicate with the environmental capture system 700 to generate a 3D panoramic image using the depth data from the LiDA R 708 and the lens assembly 704. In some embodiments, the 2D and 3D panoramic images are not generated by the environmental capture system 402.

[0072] The output of the LiDAR 708 may include attributes associated with each laser pulse transmitted by the LiDAR 708. Such attributes may include: the intensity of the laser pulse; the number of returns; the current return number; classification point; RGC value; GPS time; scan angle; scan direction; or any combination of these. The depth of field may be (0.5 m; infinity), (1 m; infinity ), etc. In some embodiments, the depth of field is 0.2 m - 1 m and infinity.

[0073] In some embodiments, the environmental capture system 700 uses the lens assembly 704 to capture four separate RBG images while the environmental capture system 700 is stationary.​ Capture. In various embodiments, LiDAR 708 captures environmental capture system while the system 700 is moving and moving from one RBG image capture position to another RBG image capture position, capturing four different instances of depth data. In one example, a 3D panoramic image is captured by a 360° rotation of the image capture system 700, which may be referred to as a sweep in some cases. In various embodiments, a 3D panoramic image is captured by a rotation of less than 360° of the environmental capture system 700. The output of the sweep is often referred to as a sweep list (SWL), which includes the image data from the lens assembly 704, the depth data from LiDAR 708, the GPS location, and the sweep characteristics including the timestamp when the sweep was performed. In various embodiments, a single sweep (e.g., a single 360° turn of the environmental capture system 700) is (e.g., by a digital device communicating with the environmental capture system 700 that receives images and depth data from the environmental capture system 700 and creates a 3D visualization using only the above-described images and depth data captured in a single sweep) sufficient to capture image and depth information to generate a 3D visualization. is. and, depth data from LiDAR 708, and the GPS location and timestamp at the time the sweep was performed. In various embodiments, a single sweep (e.g., a single 360° turn of the environmental capture system 700) is (e.g., by a digital device communicating with the environmental capture system 700 that receives images and depth data from the environmental capture system 700 and creates a 3D visualization using only the above-described images and depth data captured in a single sweep) sufficient to capture image and depth information to generate a 3D visualization of the environmental capture system 700 captured in a single sweep. 0) is used to generate a 3D visualization. 0 and receives images and depth data, and uses only the above-described images and depth data captured in a single sweep from the environmental capture system 700 captured in a single sweep to create a 3D visualization. 3D visualization. Capture image and depth information sufficient to generate a 3D visualization. .

[0074] In some embodiments, the environmental capture system 402 captures a plurality of images captured by the environmental capture system 402 and stitches them together into one, and can be combined with the depth data from LiDAR 708 by the image stitching and processing system described below. stitched and combined with the depth data from LiDAR 708. ​​​

[0075] In various embodiments, the environmental capture system 402, and / or the user system 1110 may generate a preview or thumbnail version of the 3D panoramic image. The preview or thumbnail version of the 3D panoramic image may be presented on the user system 1110 and may have a lower image resolution than the 3D panoramic image generated by the image processing system 406. After the lens assembly 704 and the LiDAR 708 capture the physical environmental image and depth data, the environmental capture system 402 may generate a minimap representing an area of the physical environment captured by the environmental capture system 402. In some embodiments, the image processing system 406 generates a minimap representing an area of the physical environment. Using the environmental capture system 402, after capturing the image and depth data of the living room of a house, the environmental capture system 402 may generate a top view of the physical environment. The user may use this information to determine areas of the physical environment where the user has not captured or generated a 3D panoramic image. In one embodiment, the environmental capture system 700 may sandwich the depth information capture by the LiDAR 708 between image captures by the image capture device of the lens assembly 704. For example, the image capture device may capture an image of a section 1605 of the physical environment as seen in FIG. 16, and then the LiDAR 70 8 obtains depth information from the section 1605. After the LiDAR 708 obtains the depth information from the section 1605, the image capture device may capture another image of the section 1605. After obtaining the depth information from the section 1605, the image capture device may capture another image of the section 1605. The user can use this information to determine areas of the physical environment where the user has not captured or generated a 3D panoramic image.

[0076] In one embodiment, the environmental capture system 700 may sandwich the depth information capture by the LiDAR 708 between image captures by the image capture device of the lens assembly 704. For example, the image capture device may capture an image of a section 1605 of the physical environment as seen in FIG. 16, and then the LiDAR 70 8 obtains depth information from the section 1605. 8 obtains depth information from the section 1605. After the LiDAR 708 obtains the depth information from the section 1605, the image capture device may capture another image of the section 1605. 8 obtains depth information from the section 1605. After the LiDAR 708 obtains the depth information from the section 1605, the image capture device may capture another image of the section 1605. ​​​​When obtaining depth information, the image capture device may move to capture an image of another section 1610, and subsequently, LiDAR 708 obtains depth information from section 1610. In this way, image capture and depth information capture are performed alternately. In some embodiments, LiDAR 708 may have a field of view of at least 145°, and depth information of all objects in the 360° view of the environmental capture system 700 can be obtained by the environmental capture system 700 in three or four scans. In another example, LiDAR 708 may have a field of view of at least 150°, 180°, or 145° - 180°.

[0077] By increasing the field of view of the lens, the amount of time required to obtain visual and depth information of the physical environment around the environmental capture system 700 is reduced. In various embodiments, LiDAR 708 has a minimum depth range of 0.5 m. In one embodiment, LiDAR 708 has a maximum depth range of more than 8 meters. The visual field of the lens and depth information of all objects in the 360° view of the environmental capture system 700 can be obtained by the environmental capture system 700 in three or four scans. In another example, LiDAR 708 may have a field of view of at least 150°, 180°, or 145° - 180°. In some embodiments, LiDAR 708 may have a field of view of at least 145°, and depth information of all objects in the 360° view of the environmental capture system 700 can be obtained by the environmental capture system 700 in three or four scans. In another example, LiDAR 708 may have a field of view of at least 150°, 180°, or 145° - 180°.

[0078] By increasing the field of view of the lens, the amount of time required to obtain visual and depth information of the physical environment around the environmental capture system 700 is reduced. In various embodiments, LiDAR 708 has a minimum depth range of 0.5 m. In one embodiment, LiDAR 708 has a maximum depth range of more than 8 meters. In various embodiments, LiDAR 708 has a minimum depth range of 0.5 m. In one embodiment, LiDAR 708 has a maximum depth range of more than 8 meters. LiDAR 708 can use the mirror assembly 712 to direct the laser at different scan angles. In one embodiment, any vertical motor 718 has the function of vertically moving the mirror assembly 712. In some embodiments, the mirror assembly 712 may be a dielectric mirror having a hydrophobic coating or layer. The mirror assembly 712 may be coupled to a vertical motor 718 that rotates the mirror assembly 712 during use. LiDAR 708 can use the mirror assembly 712 to direct the laser at different scan angles. In one embodiment, any vertical motor 718 has the function of vertically moving the mirror assembly 712. In some embodiments, the mirror assembly 712 may be a dielectric mirror having a hydrophobic coating or layer. The mirror assembly 712 may be coupled to a vertical motor 718 that rotates the mirror assembly 712 during use.

[0079] LiDAR 708 can use the mirror assembly 712 to direct the laser at different scan angles. In one embodiment, any vertical motor 718 has the function of vertically moving the mirror assembly 712. In some embodiments, the mirror assembly 712 may be a dielectric mirror having a hydrophobic coating or layer. The mirror assembly 712 may be coupled to a vertical motor 718 that rotates the mirror assembly 712 during use. In one embodiment, any vertical motor 718 has the function of vertically moving the mirror assembly 712. In some embodiments, the mirror assembly 712 may be a dielectric mirror having a hydrophobic coating or layer. The mirror assembly 712 may be coupled to a vertical motor 718 that rotates the mirror assembly 712 during use. In some embodiments, the mirror assembly 712 may be a dielectric mirror having a hydrophobic coating or layer. The mirror assembly 712 may be coupled to a vertical motor 718 that rotates the mirror assembly 712 during use. In some embodiments, the mirror assembly 712 may be a dielectric mirror having a hydrophobic coating or layer. The mirror assembly 712 may be coupled to a vertical motor 718 that rotates the mirror assembly 712 during use. The mirror assembly 712 may be coupled to a vertical motor 718 that rotates the mirror assembly 712 during use. The mirror assembly 712 may be coupled to a vertical motor 718 that rotates the mirror assembly 712 during use.

[0080] The mirror of the mirror assembly 712 may have, for example, the following specifications:

[0081]

Table 4

[0082] The mirror of the mirror assembly 712 may have, for example, the following specifications with respect to materials and coatings:

[0083]

Table 5

[0084] The hydrophobic coating of the mirror of the mirror assembly 712 may have, for example, a contact angle exceeding 105°

[0085] The mirror of the mirror assembly 712 may have the following quality specifications:

[0086]

Table 6

[0087] The vertical motor may have, for example, the following specifications:

[0088]

Table 7

[0089] The environmental capture system 700 can capture images outdoors on a sunny day or indoors where the light is bright or the sunlight from the window is dazzling by the RGB capture device and the LiDAR 708 ​These devices are often restricted to use only indoors and only during dawn or sunset in order to control light. Otherwise, artifacts or "holes" are created in the image by bright indoor spots, and it is necessary to fill or correct these. However, the environmental capture system 700 can be utilized both indoors and outdoors, even in bright sunlight. The capture device and LiDAR 708 can capture image and depth data in a bright environment without artifacts or holes caused by glare or bright light. In one embodiment, the GPS antenna 714 receives Global Positioning System (GPS) data. Using the GPS data, the location of the environmental capture system 700 at any given point in time can be determined. In various embodiments, the display 720 can provide the current state of the system, such as during an update, warm-up, scan, scan completion, error, etc., for the environmental capture system 700. The battery pack 722 supplies power to the environmental capture system 700. The battery pack 722 may be removable and rechargeable, allowing the user to insert a new battery pack 722 while charging a depleted one. In some embodiments, the battery pack 722 may be capable of continuous use of at least 1000 SWL or at least 250 SWL before recharging. The environmental capture system 700 may utilize a USB-C plug for recharging.

[0090]

[0091]

[0092]

[0093] ​​​​​​​​​​​​ In some embodiments, the mount 724 can be used to mount the environmental capture system 700 on a tripod or provides a connector for connecting to a platform such as a mount. 6 can rotate the environment capture system 700 about the x-y plane. In some embodiments, the horizontal motor 726 controls the (x, y) , z) coordinates, the information may be provided in a grid coordinate system. This allows for a wide field of view for the lens, positioning of the lens around the axis of rotation, and the ability to mount the lens on a LiDAR device. Thus, the horizontal motor 726 allows the environmental capture system 700 to perform scans quickly. It can be made possible.

[0094] The horizontal motor 726 may have the following specifications, by way of example:

[0095] [Table 8]

[0096] In various embodiments, the mount 724 may include a quick release adapter. The holding torque can be, for example, over 2.0 Nm, and the durability of the capture operation can be up to 70, 000 cycles, or more than 70,000 cycles.

[0097] For example, the Environmental Capture System 700 captures the entire image of a standard home with sweep-to-sweep distances of over 8 m. It may be possible to capture, process, and analyze indoor sweeps. The time for alignment can be less than 45 seconds. From the start of the capture, the user can move the environmental capture system 700. The time window can be less than 15 seconds.

[0098] In various embodiments, these components are used to provide environmental capture system 700 with By aligning multiple scan positions both indoors and outdoors, the system can be It provides the ability to create seamless walk-through experiences (this applies to hotels, vacation rentals, High-quality software for real estate and construction research, CRE, and as-built modeling and verification. The environment capture system 700 may be an "outdoor dollhouse" or An external minimap can also be created using the Environment Capture System 700, as shown here. It can also improve the accuracy of 3D reconstruction, mainly from a measurement perspective. Regarding density, it may also be a plus for users to be able to fine-tune this. The components also allow the environmental capture system 700 to capture large open spaces (e.g., relatively long It is possible to capture a 3D model of a large empty space. To generate it, the environment capture system needs to generate a 3D model of a smaller space. Need to scan and capture 3D and depth data from a large distance range? could be.

[0099] In various embodiments, these components are used to enable the environmental capture system 700 to capture images indoors. A similar method is used for indoor and outdoor use to align multiple SWLs and create 3D models. These components also allow the environment capture system 700 to: It is also possible to perform geolocation of the 3D model (this is done by Google Easily integrate into eStreet View and align multiple outdoor panoramas as needed that may be useful for doing so).

[0100] The image capture device of the environmental capture system 700 may be capable of providing an image such as a DSLR that has a quality printable at 8.5 inches by 11 inches with respect to a 70° VFOV, and an RGB image style. that may be useful for doing so). LR style.

[0101] In some embodiments, the environmental capture system 700 captures an RGB image by the image capture device (e.g., using a wide-angle lens), and after moving the lens, can capture the next RGB image (using a motor to move a total of four times). While the horizontal motor 726 rotates the environmental capture system by 90°, the LiDAR 708 can capture depth data. In some embodiments, the LiDAR 708 includes an APD array. that may be useful for doing so). that may be useful for doing so). that may be useful for doing so). that may be useful for doing so).

[0102] In some embodiments, the image and depth data may then be sent to a capture application (e.g., a device that communicates with the environmental capture system 700, such as a smart device or an image capture system on a network). In some embodiments, the environmental capture system 700 can send the image and depth data to the image processing system 406 for processing and generation of a 2D panoramic image or a 3D panoramic image. that may be useful for doing so). that may be useful for doing so). that may be useful for doing so). that may be useful for doing so). In various embodiments, the environmental capture system 700 may generate a sweep list of RGB images and depth data captured from a 360° rotation of the environmental capture system 700. This sweep list can be sent to the image processing system 406 for stitching and alignment. The output of the sweep may be an SWL, which that may be useful for doing so). that may be useful for doing so). that may be useful for doing so). Image data from the lens assembly 704, depth data from the LiDAR 708, and G including the characteristics of the sweep including the location of the GPS and the timestamp at the time the sweep was performed .

[0103] In various embodiments, the LIDAR, vertical mirror, RGB lens, tripod mount, and horizontal drive are securely installed within the housing such that the housing can be opened without the need for recalibration of the system .

[0104] FIG. 9a shows a block diagram 900 of an example of an environmental capture system according to some embodiments . The block diagram 900 includes a power supply 902, a power converter 904, an input / output (I / O) printed circuit board assembly (PCBA), a system on module (SOM) PCBA, a user interface 910, a LiDAR 912, a mirror brushless direct current (BLDC) motor 914, a drive train 916, a wide field of view (WFOV) lens 918, and an image sensor 920 . The power supply 902 may be the battery pack 722 of FIG. 7. The power supply may be a removable and rechargeable battery such as a lithium ion battery (e.g., 4 x 18650 Li-Ion battery) that can supply power to the environmental capture system . The power converter 904 can convert the voltage level from the power supply 902 to a lower or higher voltage such that the electronic components of the environmental capture system can utilize it .

[0105] The power supply 902 may be the battery pack 722 of FIG. 7. The power supply is a removable and rechargeable battery such as a lithium ion battery (e.g., 4 x 18650 Li-Ion battery) that can supply power to the environmental capture system .

[0106] The power converter 904 can convert the voltage level from the power supply 902 to a lower or higher voltage such that the electronic components of the environmental capture system can utilize it The environmental capture​ The system may utilize a 4S1P configuration, i.e., a configuration of four series connections and one parallel connection, of 4 × 18 650 Li-Ion batteries.

[0107] In some embodiments, the I / O PCBA 906 may include elements providing an IMU, Wi-Fi, GPS, Bluetooth, an inertial measurement unit (IMU), a motor drive, and a microcontroller. In some embodiments, the I / O PCBA 906 may include a microcontroller for controlling a horizontal motor and encoding the control of the horizontal motor, as well as for controlling a vertical motor and encoding the control of the vertical motor.

[0108] The SOM PCBA 908 may include a central processing unit (CPU) and / or a graphics processing unit (GPU), memory, and a mobile interface. The SOM PCBA 908 may control the LiDAR 912, the image sensor 920, and the I / O PCBA 906. The SOM PCBA 908 may determine the (x, y, z) coordinates associated with each laser pulse of the LiDAR 912 and store the coordinates in the memory components of the SOM PCBA 908. In some embodiments, the SOM PCBA 908 may store the coordinates in the image processing system of the environmental capture system 400. In addition to the coordinates associated with each laser pulse, the SOM PCBA 908 may store the intensity of the laser pulse, the number of returns, the current return number, classification points, RGC values, GPS time, scan angle, and scan direction Additional attributes associated with each laser pulse, including direction, may be determined.

[0109] In some embodiments, the SOM PCBA908 includes a Nvidia SOM PCBA with a CPU / GPU, DDR, eM MC, and Ethernet.

[0110] The user interface 910 may include physical buttons or switches with which the user can interact. The buttons or switches can provide functions such as turning the environmental capture system on and off and scanning the physical environment. In some embodiments, the user interface 910 may include a display such as the display 720 of FIG. 7.

[0111] In some embodiments, LiDAR912 captures depth information of the physical environment. LiDAR912 includes an optical sensing module that can measure the distance to an object within a target or scene by irradiating the target or scene with light. The optical sensing module of LiDAR912 measures the time it takes for photons to travel to the target or object and return to the receiver of LiDAR912 after reflection, thereby providing the distance of LiDAR from the target or object. The SOM PCBA908 can determine the (x, y, z) coordinates associated with each laser pulse along with the above distance. LiDAR912 can fit within a range of 58 mm in width, 55 mm in height, and 60 mm in depth.

[0112] LiDAR912 has a range (10% reflectivity) of 90 m, a range (20% reflectivity) of 130 m, a range (100% reflectivity) of 260 m, a range accuracy (1σ @ 900 m) of 2 cm, and a wavelength of It may be 1705 nm with a beam divergence of 0.28×0.03°.

[0113] SOM PCBA908 may determine coordinates based on the location of drive train 916. In various embodiments, LiDAR912 may include one or more LiDAR devices. By utilizing multiple LiDAR devices, the resolution of the LiDAR can be improved. This can be achieved.

[0114] The mirror brushless direct current (BLDC) motor 914 can control the mirror assembly 712 of FIG. 7. This can be achieved.

[0115] In some embodiments, drive train 916 may include the horizontal motor 726 of FIG. 7. The drive train 916 can provide rotation of the environmental capture system when the environmental capture system is installed on a platform such as a tripod. The drive train 916 may include a stepper motor Nema14, a worm and plastic gear drive train, a clutch, a bushing bearing gear, and a backlash prevention mechanism. In some embodiments, the environmental capture system can complete one scan in less than 17 seconds. In various embodiments, the drive train 916 has a maximum speed of 60° / second, a maximum acceleration of 300° / second a maximum torque of 0.5 nm, an angular position accuracy of less than 0.1°, and an encoder resolution of approximately 4096 counts per revolution. In some embodiments, the environmental capture system can complete one scan in less than 17 seconds. In various embodiments, the drive train 916 has a maximum speed of 60° / second, a maximum acceleration of 300° / second 2 a maximum torque of 0.5 nm, an angular position accuracy of less than 0.1°, and an encoder resolution of approximately 4096 counts per revolution. In some embodiments, drive train 916 includes a vertical single-gon mirror and a motor. In this example, drive train 916 includes a BLDC motor, an external hall effect sensor, (hall effect sensor

[0116] In some embodiments, drive train 916 includes a vertical single-gon mirror and a motor. In this example, drive train 916 includes a BLDC motor, an external hall effect sensor, (hall effect sensor The drive train 91 in this example may include a magnet (paired with the magnet), a mirror bracket, and a mirror. 6 may have a maximum speed of 4,000 RPM and a maximum acceleration of 300° / sec^2. In some embodiments, the monogon mirror is a dielectric mirror. The monogon mirror includes a hydrophobic coating or layer.

[0117] The environmental capture system components are arranged such that the lens assembly and LiDAR rotate. This allows the image capture system to rotate This reduces the parallax of the image that occurs when the camera is not positioned at the center of the rotation axis.

[0118] In some embodiments, the WFOV lens 918 is the lens assembly 704 of FIG. The WFOV lens 918 may be a lens that focuses light onto an image capture device. In some embodiments, the WFOV lens has a FOV of at least 145°. Such a wide FOV allows for a 360° view around the environmental capture system. The image capture is obtained by three separate image captures of the image capture device. In some embodiments, the WFOV lens 918 can have a diameter of about 60 mm, and and a total track length (TTL) of about 80 mm. Lens 918 may have a horizontal field of view of 148.3° or greater and a vertical field of view of 94° or greater.

[0119] The image capture device may include a WFOV lens 918 and an image sensor 920. Image sensor 920 may be a CMOS image sensor. In one embodiment, image sensor 920 is a charge coupled device (CCD). In some embodiments, image sensor 920 is is a red - green - blue (RGB) sensor. In one embodiment, the image sensor 920 is an IR sensor. In various embodiments, the image capture device may have a resolution of at least 35 pixels per degree (PPD).

[0120] In some embodiments, the image capture device has: an f - number of f / 2.4; an image circle diameter of 15.86 mm; a pixel pitch of 2.4 μm; a horizontal field of view (HFOV) greater than 148.3° ; a vertical field of view (VFOV) greater than 94.0°; a number of pixels per degree greater than 38.0 PPD; a chief ray incidence angle at the full height of 3.0° ; a minimum shooting distance of 1300 mm; a maximum shooting distance of infinity; a relative illuminance greater than 130% ; a maximum distortion less than 90%; and a spectral transmittance variation of 5% or less. may be.

[0121] In some embodiments, the lens has: an f - number of 2.8; an image circle diameter of 15.86 mm ; a number of pixels per degree greater than 37; a chief ray incidence angle in the full - height sensor of 3.0 ; an L1 diameter less than 60 mm; a total track length (TTL) less than 80 mm; and a relative illuminance greater than 50% may be.

[0122] The lens may have greater than 85% at 52 lp / mm (on - axis), greater than 66% at 104 lp / mm (on - axis) , greater than 45% at 208 lp / mm (on - axis), greater than 75% at 52 lp / mm (83% of the field of view), greater than 41% at 104 lp / mm (83% of the field of view), and greater than 25% at 208 lp / mm (83% of the field of view).

[0123] The environmental capture system has a resolution greater than 20 MP, a green sensitivity greater than 1.7 V / lux*sec , a signal - to - noise ratio (SNR) greater than 65 dB (100 lux, 1x gain), and a dynamic range greater than 70 dB It may have a range.

[0124] Figure 9b shows a block diagram of an example of the SOM PCBA908 of an environmental capture system according to some embodiments. The SOM PCBA908 may include communication components 922, Li DAR control components 924, LiDAR placement components 926, user interface components 928, classification components 930, LiDAR data store 932, and captured image data store 934.

[0125] In some embodiments, the communication components 922 can send and receive requests or data between any of the components of the SOM PCBA1008 and the components of the environmental capture system of FIG. 9a.

[0126] In various embodiments, the LiDAR control components 924 can control various aspects of the LiDAR. For example, the LiDAR control components 924 may send a control signal to the LiDAR912 to start sending laser pulses. The control signal sent by the LiDAR control components 9 24 may include instructions for the frequency of the laser pulses.

[0127] In some embodiments, the LiDAR placement components 926 can use GPS data to determine the location of the environmental capture system. In various embodiments, the LiDAR placement components 926 can use the position of the mirror assembly to determine the scan angle and (x, y, z) coordinates associated with each laser pulse. The LiDAR placement components 926 can also use an IMU to determine the orientation of the environmental capture system.

[0128] The user interface component 928 can facilitate the user's interaction with the environmental capture system. In some embodiments, the user interface component 92 8 may provide one or more user interface elements with which the user can interact. The user interface provided by the interface component 928 can be sent to the user system 11 10. For example, the user interface component 928 can provide a visual representation of an area with the floor plan of a building to the user system (such as a digital device). When the user places the environmental capture system in multiple different parts of one floor of a building and captures and generates a 3D panoramic image, the environmental capture system can generate a visual representation of the floor plan. When the user places the environmental capture system in an area of a physical environment and captures and generates a 3D panoramic image of that area of the house. After the 3D pano rama image of the area is generated by the image processing system, the user interface component can update the floor plan using a view from above of the living room area as shown in FIG. 1b. In some embodiments, the floor plan 200 can be generated by the user system 1110 after capturing a second sweep of one house or a certain floor of a building.

[0129] In various embodiments, the classification component 930 can classify the type of the physical environment. The classification component 930 can analyze objects within the image or objects in the image to classify the type of the physical environment captured by the environmental capture system. In some embodiments, the image processing system is by the environmental capture system 400 It can serve the role of classifying the type of the captured physical environment.

[0130] The LiDAR data store 932 can be any structure and / or multiple structures suitable for the captured LiDAR data (for example, an active database, a relational database system, a self-referential database, a table, a matrix, an array, a flat file, a document-oriented storage system, a non-relational No-SQL system, a FTS management system such as Lucene / Solar, etc.). The image data store 408 can store the captured LiDAR data. However, the LiDAR data store 932 can be used to cache the captured LiDAR data when the communication network 404 is not functioning. For example, when the environmental capture system 40 2 and the user system 1110 are in a remote location without a cellular network or in an area without Wi-Fi, the LiDAR data store 932 can store the captured LiDAR data until it can be transferred to the image data store 934. For example, when the environmental capture system 402 and the user system 1110 are in a remote location without a cellular network or in an area without Wi-Fi, the LiDAR data store 932 can store the captured LiDAR data until it can be transferred to the image data store 934. store the captured LiDAR data until it can be transferred to the image data store 934. AR data can be stored until it can be transferred to the image data store 934.

[0131] Similar to the LiDAR data store, the captured image data store 934 can be any structure and / or multiple structures suitable for the captured images (for example, an active database, a relational database system, a self-referential database, a table, a matrix, an array, a flat file, a document-oriented storage system, a non-relational No-SQL system, a FTS management system such as Lucene / Solar, etc.). The image data store 934 can store the captured images. can store the captured images.

[0132] 10a-10c illustrate an environmental capture system for capturing images in some embodiments. 10a-10c show the process of the environmental chimney system 400. The capture system 400 can take a burst of images at different exposures. A burst can be a set of multiple images, each with a different exposure. is at time 0.0. The environment capture system 400 receives the first frame. This frame can be evaluated while waiting for the second frame. This shows that the first frame after the arrival of the frame is blended. The environment capture system 400 processes each frame to identify pixels, colors, etc. When the next frame arrives, the environment capture system 400 The 3D image processing unit 102 may process the resulting frame and blend the two frames into one.

[0133] In various embodiments, the environment capture system 400 performs image processing to The blended frames (e.g., any number of image frames) are then blended. The environment capture evaluates the pixels in the frame (which may contain elements from the previous frame). During this final step, before or during a movement (e.g., a turn) of the capture system 400: The environment capture system 400 optionally receives the blended image from the image processor. May be transferred to CPU memory.

[0134] The process continues in Figure 10b. At the beginning of Figure 10b, the environmental capture system 40 0 performs another burst. The environment capture system 400 You may compress all or part of the frames captured, using JXR. Similar to FIG. 10a, a burst of images may be a set of multiple images with different exposures (the exposure length of each frame in the set may be the same and may be in the same order as other bursts included in FIGS. 10a and 10c). The second image burst is at the 2-second mark. The environmental capture system 400 may receive the first frame and evaluate this frame while waiting for the second frame. FIG. 10b shows that the first frame is blended after the arrival of the second frame. In some embodiments, the environmental capture system 400 may process each frame to identify pixels, colors, etc. When the next frame arrives, the environmental capture system 400 may process the most recently received frame and blend the two frames into one. In various embodiments, the environmental capture system 400 performs image processing to blend the sixth frame and further evaluates the pixels in the blended frame (which may include elements from frames of any number of image bursts). During this last step, before or during a turn (e.g., a turn) of the environmental capture system 400, the environmental capture system 400 may optionally transfer the blended image from the image processing device to the CPU memory. After turning, the environmental capture system 400 may continue the process by performing another color burst at approximately the 3.5-second mark (e.g., after a 180° turn).

[0135]

[0136] It is. The environmental capture system 400 may compress all or part of the blended frame and / or the captured frame using JXR. An image burst may be a set of multiple images with different exposures (the exposure lengths of each frame in the set may be the same and may be in the same order as other bursts included in FIGS. 10a and 10c). The environmental capture system 400 may receive a first frame and evaluate this frame while waiting for a second frame. FIG. 10b shows that the first frame is blended after the arrival of the second frame. In some embodiments, the environmental capture system 400 may process each frame to identify pixels, colors, etc. When the next frame arrives, the environmental capture system 400 may process the most recently received frame and blend the two frames into one. In various embodiments, the environmental capture system 400 may perform image processing to blend the sixth frame and further evaluate the pixels in the blended frame (which may include elements from frames of any number of image bursts). During this last step, before or during the movement (e.g., turn) of the environmental capture system 400, the environmental capture system 400 may optionally transfer the blended image from the image arithmetic processing unit to the CPU memory. The last burst occurs at the 5 - second mark in FIG. 10c. The environmental capture system 400 may compress all or part of the blended frame and / or the captured frame using JXR. The environmental capture system 400 may receive a first frame and evaluate this frame while waiting for a second frame. FIG. 10b shows that the first frame is blended after the arrival of the second frame. In some embodiments, the environmental capture system 400 may process each frame to identify pixels, colors, etc. When the next frame arrives, the environmental capture system 400 may process the most recently received frame and blend the two frames into one. In various embodiments, the environmental capture system 400 may perform image processing to blend the sixth frame and further evaluate the pixels in the blended frame (which may include elements from frames of any number of image bursts). During this last step, before or during the movement (e.g., turn) of the environmental capture system 400, the environmental capture system 400 may optionally transfer the blended image from the image arithmetic processing unit to the CPU memory. The last burst occurs at the 5 - second mark in FIG. 10c. The environmental capture system 400 may compress all or part of the blended frame and / or the captured frame using JXR. The environmental capture system 400 may receive a first frame and evaluate this frame while waiting for a second frame.

[0137] In some embodiments, the environmental capture system 400 may process each frame to identify pixels, colors, etc. When the next frame arrives, the environmental capture system 400 may process the most recently received frame and blend the two frames into one. In various embodiments, the environmental capture system 400 may perform image processing to blend the sixth frame and further evaluate the pixels in the blended frame (which may include elements from frames of any number of image bursts). During this last step, before or during the movement (e.g., turn) of the environmental capture system 400, the environmental capture system 400 may optionally transfer the blended image from the image arithmetic processing unit to the CPU memory. The last burst occurs at the 5 - second mark in FIG. 10c. The environmental capture system 400 may compress all or part of the blended frame and / or the captured frame using JXR.

[0138] The last burst occurs at the 5 - second mark in FIG. 10c. The environmental capture system 400 may compress all or part of the blended frame and / or the captured frame using JXR. The environmental capture system 400 may receive a first frame and evaluate this frame while waiting for a second frame. It may be compressed using JXR. A burst of images may be a set of a plurality of images with different exposures (the exposure lengths of each frame of the above set may be the same and may be in the same order as other bursts included in FIGS. 10a and 10b). The environmental capture system 400 may receive the first frame and evaluate this frame while waiting for the second frame. FIG. 10c shows that the first frame is blended after the arrival of the second frame. In some embodiments, the environmental capture system 400 may process each frame to identify pixels, colors, etc. When the next frame arrives, the environmental capture system 400 may process the most recently received frame and blend the two frames into one. It may be a set (the exposure length of each frame of the above set may be the same and may be in the same order as other bursts included in FIGS. 10a and 10b). 10a, 10b). The environmental capture system 400 may receive the first frame and evaluate this frame while waiting for the second frame. FIG. 10c shows that the first frame is blended after the arrival of the second frame. It can be evaluated. FIG. 10c shows that the first frame is blended after the arrival of the second frame. In some embodiments, the environmental capture system 400 may process each frame to identify pixels, colors, etc. When the next frame arrives, the environmental capture system 400 may process the most recently received frame and blend the two frames into one. It may identify pixels, colors, etc. When the next frame arrives, the environmental capture system 400 may process the most recently received frame and blend the two frames into one. The environmental capture system 400 may process the most recently received frame and blend the two frames into one.

[0139] In various embodiments, the environmental capture system 400 may perform image processing to blend the sixth frame and further evaluate the pixels in the blended frame (for example, a frame that may include elements from frames of any number of image bursts). During this last step, before or during the movement (for example, turn) of the environmental capture system 400, The environmental capture system 400 may optionally transfer the blended image from the image arithmetic processing device to the CPU memory. It may include elements from frames of any number of image bursts). During this last step, before or during the movement (for example, turn) of the environmental capture system 400, The environmental capture system 400 may optionally transfer the blended image from the image arithmetic processing device to the CPU memory.

[0140] The dynamic range of an image capture device is a measure of the amount of light that the image sensor can capture. The dynamic range is the difference between the darkest area and the brightest area of the image. There are many ways to improve the dynamic range of an image capture device, One of which is to use a plurality of images of the same physical environment with different exposures. It is the difference between the darkest area and the brightest area of the image. There are many ways to improve the dynamic range of an image capture device, One of which is to use a plurality of images of the same physical environment with different exposures. ​​It is to capture. An image captured with a short exposure will capture the brightest area of the physical environment, and a long exposure will capture a darker area of the physical environment. In some embodiments, the environmental capture system may capture multiple images at six different exposure times. Using some or all of the images captured by the environmental capture system, a 2D image with a high dynamic range (HDR) is generated. One or more of the captured images may be used for other functions such as light detection and flicker detection. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. In some embodiments, the environmental capture system may capture multiple images at six different exposure times. Using some or all of the images captured by the environmental capture system, a 2D image with a high dynamic range (HDR) is generated. One or more of the captured images may be used for other functions such as light detection and flicker detection. In some embodiments, the environmental capture system may capture multiple images at six different exposure times. Using some or all of the images captured by the environmental capture system, a 2D image with a high dynamic range (HDR) is generated. One or more of the captured images may be used for other functions such as light detection and flicker detection. In some embodiments, the environmental capture system may capture multiple images at six different exposure times. Using some or all of the images captured by the environmental capture system, a 2D image with a high dynamic range (HDR) is generated. One or more of the captured images may be used for other functions such as light detection and flicker detection. In some embodiments, the environmental capture system may capture multiple images at six different exposure times. Using some or all of the images captured by the environmental capture system, a 2D image with a high dynamic range (HDR) is generated. One or more of the captured images may be used for other functions such as light detection and flicker detection.

[0141] A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. A 3D panoramic image of the physical environment can be generated based on four separate image captures of the image capture device and four separate depth data of the LiDAR device of the environmental capture system. Each of the four separate image captures may include a series of image captures with different exposure times. Using a blending algorithm, the series of image captures with different exposure times can be blended to generate one of the four RGB image captures, which can be used to generate a 2D panoramic image. For example, an environmental capture system may be used to capture a 3D panoramic image of a kitchen. An image of one wall of this kitchen may include a window. An image captured with a short exposure can provide a view outside the window, but the rest of the kitchen may remain underexposed. Symmetrically, another image captured with a long exposure can provide a view inside the kitchen. The blending algorithm can combine the view outside the kitchen window from one image with the view inside the kitchen from another image. It is possible to generate a blended RGB image by blending the remaining views of the chin.

[0142] In various embodiments, the 3D panoramic image can be based on three separate image captures of an image capture device and four separate depth data from a LiDAR device of an environmental capture system. In some embodiments, the number of image captures and the number of depth data captures may be the same. In one embodiment, the number of

[0143] image captures and the number of depth data captures may be different. After capturing a first series of images at a certain exposure time, the blending algorithm receives the first series of images, calculates an initial intensity weight for the images, and sets the images as baseline images for combining subsequent received images. In some embodiments, the blending algorithm may utilize an image processing routine of a graphics processing unit (GPU), such as a "blend_kernel" routine. The blending algorithm can receive subsequent images, which can be blended with previously received images. In some embodiments, the blending algorithm may utilize variations of the blend_kernel GPU

[0144] image processing routine. In one embodiment, the blending algorithm determines the difference, i.e., the contrast, between the darkest and brightest parts of the baseline image, and uses other methods of blending multiple images, such as determining whether the baseline image is overexposed Means being an exposed over or an exposed under. In one embodiment, the baseline image contrast can be calculated by obtaining the average of the light intensity of the image or a subset of the image. In some embodiments, the blending algorithm calculates the average light intensity for each row or column of the image. In some embodiments, the blending algorithm determines the histogram of each image received from the image capture device and analyzes this histogram to determine the light intensity of the pixels that make up each image.

[0145] In various embodiments, blending may include sampling the colors within two or more images of the same scene, including those along objects and seams. (For example within a predetermined threshold such as color, hue, luminance, chroma, etc.) If there is a significant difference in color between two images the blending module (e.g., on the environmental capture system 400 or the user device 1110) may blend both images of a predetermined size along the location where the difference exists. In some embodiments, the greater the color or image difference at a location in the image, the more space in the vicinity of that location may be blended.

[0146] In some embodiments, after blending, the blending module (e.g., on the environmental capture system 400 or the user device 1110) may re-sample and sample the colors along one or more images to determine whether there are other differences in the image or color that exceed the above predetermined threshold values such as color, hue, luminance, chroma, etc. If so, the blending module may identify the portion within the one or more images and continue to blend the portion of the image. The blend module will blend when there are no more parts of the image to blend (e.g. color Continue resampling the image along the seam until the difference is less than one or more predetermined thresholds. It is okay.

[0147] FIG. 11 illustrates a method for capturing and stitching images to create a 3D 11 shows a block diagram of an example environment 1100 in which a visualization can be created. The illustrative environment 1100 includes a 3D and panoramic capture and stitching system 1102. , a communication network 1104, an image stitching processor system 1106, an image A first sequence of a data store 1108, a user system 1110, and a physical environment 1112. 3D and panoramic capture and stitching system 1102 and / or The user system 1110 is used to capture images of an environment (e.g., a physical environment 1112). The environmental capture system 400 may include an image capture device capable of capturing an image.

[0148] 3D and Panorama Capture Stitching System 1102 and Image Stitching The imaging processor system 1106 is communicatively coupled to the environment capture system 400. The digital information may be part of a single integrated system (e.g., part of one or more digital devices). In some embodiments, the 3D and panoramic capture and stitching system 110 2 and some of the functions of the components of the image stitching processor system 1106 At least one of these may be implemented by the environmental capture system 400. 3D and panoramic capture stitching system 1102 and image stitching The ring - processor system 1106 can be implemented by the user system 1110 and / or the image stitching system 1106.

[0149] The user can use the 3D panorama capture - stitching system 1102 to capture a plurality of 2D images of an environment such as the inside and / or the outside of a building. For example, the user can use the 3D and panorama capture - stitching system 1102 to capture a plurality of 2D images of the first scene of the physical environment 1112 provided by the environment capture system 400. The 3D and panorama capture - stitching system 1102 may include an alignment - stitching system 1114. Alternatively, the user system 1110 may include the alignment - stitching system 1114.

[0150] The alignment - stitching system 1114 is software, hardware, or a combination of both configured to provide guidance to the user of the image capture system (e.g., for the 3D and panorama capture - stitching system 1102, or for the user system 1110) and / or to enable the creation of an improved panorama photograph (by stitching, alignment, cropping, etc.). The alignment - stitching system 1114 may be on a computer - readable medium (described herein). In some embodiments, the alignment - stitching system 1114 may include a processor for implementing the functions.

[0151] An example of a first scene of the physical environment 1112 may be some room, real estate, etc. (e.g., a living room representation). In some embodiments, a 3D and panoramic capture stitching system 1102 is utilized to generate a 3D panoramic image of the indoor environment. In some embodiments, the 3D panoramic capture stitching system 1102 may be the environment capture system 400 described in connection with FIG. 4.

[0152] In some embodiments, the 3D capture stitching system 1102 can communicate with a device for capturing images and depth data, as well as software (e.g., the environment cap ture system 400). All or part of the software can be installed in the 3D panoramic capture stitching system 1102, the user system 1110, the environment capture system 400, or all of them. In some embodiments, the user can interact with the 3D and panoramic capture stitching system 1102 via the user system 1110.

[0153] The 3D and panoramic capture stitching system 1102, or the user system 1110 can obtain a plurality of 2D images. The 3D and panoramic capture stitching system 1102, or the user system 1110 can obtain depth data (e.g., from LiDAR devices etc.).

[0154] In various embodiments, an application on the user system 1110 (e.g., a user's smart device such as a smartphone or a tablet computer), or The application on the environmental capture system 400 can provide visual or auditory guidance to the user for taking pictures using the environmental capture system 4 00. Examples of graphic guidance include, for example, free-moving arrows on the display of the environmental capture system 4 00 (e.g., on the finder or LED screen on the back of the environmental capture system 400) to guide the user on where to position and / or aim the image capture device. In another example, the above application can provide voice guidance regarding where to position and / or aim the image capture device. In some embodiments, with the above guidance, the user can capture multiple images of the physical environment without the help of a stabilization platform such as a tripod. In one example, the image capture device can be a personal device such as a smartphone, tablet, media tablet, laptop, etc. The above application can provide a direction regarding the position of each sweep to approximate the vanishing point based on the position of the image capture device, location information from the image capture device, and / or past images of the image capture device. In some embodiments, with visual and / or auditory guidance, it is possible to capture multiple images that can be stitched together into a panorama without using a tripod and without using camera position information (e.g., indicating the location, position, and / or orientation of the camera from sensors, GPS devices, etc.).

[0155]

[0156]

[0157] ​​​​​​​​​​​ The alignment and stitching system 1114 can align or stitch (for example, captured by the 3D panoramic capture and stitching system 1102 in the user system 1110 or otherwise) 2D images to obtain a 2D panoramic image.

[0158] In some embodiments, the alignment and stitching system 1114 utilizes a machine learning algorithm to align or stitch a plurality of 2D images into a 2D panorama image. The parameters of the machine learning algorithm can be managed by the alignment and stitching sys tem 1114. For example, the 3D and panoramic capture and stitching system 1102 and / or the alignment and stitching system 1114 can assist in aligning these images into a 2D panorama image by recognizing objects within the 2D images.

[0159] In some embodiments, the alignment and stitching system 1114 can obtain a 3D panoramic image using depth data and a 2D panoramic image. The 3D panoramic image may be provided to the 3D and panoramic stitching system 1102 or the user system 111 0. In some embodiments, the alignment and stitching system 1 114 determines 3D-depth measurement values associated with objects recognized within the 3D panoramic image, and / or sends one or more 2D images, depth data, one or more 2D panoramic images and one or more 3D panoramic images to the image stitching processor system 1106, whereby the 3D and panoramic capture and stitching system 1102 to obtain a 2D panoramic image or a 3D panoramic image having a higher pixel resolution than the provided 2D panoramic image or 3D panoramic image. to obtain a 2D panoramic image or a 3D panoramic image having a higher pixel resolution than the provided 2D panoramic image or 3D panoramic image.

[0160] Communication network 1104 may represent one or more computer networks (e.g., LAN, WAN, etc.) or other transmission media. Communication network 1104 provides communication between systems 11 02, 1106 - 1110, and / or other systems described herein. In some embodiments, communication network 104 includes one or more digital data buses, routers, cables, buses, and / or other network topologies (e.g., mesh etc.). In some embodiments, communication network 1104 may be wired and / or wireless. In various embodiments, communication network 1104 may be: the Internet; one or more wide area networks (WANs) or local area networks (LANs); one or more networks that may be public, private, IP - based, non - IP - based, etc. including. or local area network (LAN); one or more networks that may be public, private, IP - based, non - IP - based, etc. including. including.

[0161] The image stitching processor system 1106 processes 2D images captured by an image capture device (e.g., the environmental capture system 400, or user devices such as smartphones, personal computers, media tablets, etc.) and stitches them into 2D panoramic images. The 2D panoramic images processed by the image stitching processor system 1106 may have a higher pixel resolution than the panoramic images obtained by the 3D and panoramic capture and stitching system 1102. system 1102.

[0162] In some embodiments, the image stitching processor system 1106 receives and processes a 3D panoramic image and creates a 3D panoramic image having a higher pixel resolution than the received 3D panoramic image. This panoramic image with a higher pixel resolution can be fed to an output device, such as a computer screen, a projector screen, etc., having a screen resolution higher than that of the user system 1110. In some embodiments, this panoramic image with a higher pixel resolution can provide a more detailed panoramic image to the output device and is also scalable.

[0163] The image data store 1108 can be any suitable structure and / or structures (e.g., an active database, a relational database, a self-referential database, a table, a matrix, an array, a flat file, a document-oriented storage system, a non-relational No-SQL system, an FTS management system such as Lucene / Solar, etc.) suitable for storing captured images and / or depth data. The image data store 1108 can store the images captured by the image capture device of the user system 1110. In various embodiments, the image data store 1108 stores the depth data captured by one or more depth sensors of the user system 1110. In various embodiments, the image data store 1108 stores the characteristics associated with the image capture device, or the characteristics associated with each of the plurality of image captures or depth captures used in the determination of 2D or 3D panoramic images. In some embodiments, the image The data store 1108 stores 2D or 3D panoramic images. The 2D or 3D panoramic image can be determined by the 3D and panoramic capture and stitching system 1102 or the image sti tching processor system 106.

[0164] The user system 1110 can communicate between the user and other associated systems. In some embodiments, the user system 1110 may be one or more mobile devices (e.g., smartphones, mobile phones, smartwatches, etc.) or may include these .

[0165] The user system 1110 may include one or more image capture devices. The one or more image capture devices can include, for example, RGB cameras, HDR cameras, video cameras, IR cameras, etc.

[0166] The 3D and panoramic capture and stitching system 1102 and / or the user system 1110 may include two or more capture devices, which may be arranged in relative positions to each other on the same mobile housing or within the same mobile housing such that the combined field of view covers 360°. In some embodiments, a plurality of pairs of image capture devices (e.g., having a partially overlapping field of view, although slightly offset) that can generate a pair of stereo images can be used. The user system 1110 may include two image capture devices with a vertical stereo offset field of view that can capture a pair of vertical stereo images. In another example, the user system 1110 may include two image capture devices with a vertical stereo offset field of view that can capture a pair of vertical stereo images. . offset field of view... Two image capture devices with a fuselage view can be provided.

[0167] In some embodiments, the user system 1110, the environment capture system 400, or the 3D and panoramic capture and stitching system 1102 can generate and / or provide image capture position and location information. For example, the user system 1110 or the 3 D and panoramic capture and stitching system 1102 can assist in determining position data associated with one or more image capture devices that capture a plurality of 2D images and may include an inertial measurement unit (IMU). The user system 1110 may include a global positioning sensor (GPS) to provide GPS coordinate information associated with a plurality of 2D images captured by one or more image capture devices.

[0168] In some embodiments, the user may interact with the alignment and stitching system 1114 using a mobile application installed on the user system 1110. The 3D and panoramic capture and stitching system 1102 may provide images to the user system 1110. The user may utilize the alignment and stitching system 1114 on the user system 1110 to view images and previews.

[0169] In various embodiments, the alignment and stitching system 1114 is configured to transmit and receive one or more 3D panoramic images to and from the 3D and panoramic capture and stitching system 1102 and / or the image stitching and processing system 1106. may be made. In some embodiments, the 3D and panoramic capture and stitching system 1102 may provide a visual representation of a portion of the floor plan of a building captured by the 3D and panoramic capture and stitching system 110 2 to the user system 1110.

[0170] The user of the system 1110 can navigate the space around the area described above to view different rooms of the house. In some embodiments, the user of the user system 1110 can cause the image stitching processor system 1106 to display a 3D panoramic image, such as an exemplary 3D panoramic image, upon completion of the generation of the 3D panoramic image. In various embodiments, the user system 1110 generates a preview or thumbnail of the 3D panoramic image. The preview of the 3D panoramic image may have a lower image resolution than the 3D panoramic image generated by the 3D and panoramic capture and stitching system 1102.

[0171] FIG. 12 is a block diagram of an example of an alignment and stitching system 1114 according to some embodiments. The alignment and stitching system 1114 includes a communication module 1202, an image capture position module 1204, a stitching module 120 6, a crop module 1208, an image cutout module 1210, a blend module 1211, a 3D image generator 1214, a captured 2D image data store 1216, a 3D panoramic image data store 1218, and a guidance module 220. The alignment and stitching performs one or more different functions as described in this specification. It can be understood that any number of modules of the system 1114 may exist.

[0172] In some embodiments, the alignment and stitching system 1114 includes an image capture module configured to receive images from one or more image capture devices (e.g., cameras). The alignment and stitching system 1114 may include a depth module configured to receive depth data from a depth device such as LiDAR, if available. The communication module 1202 can send and receive requests, images, or data between any of the modules or data stores of the alignment and stitching system 1114 and the components of the exemplary environment 1100 of FIG. 11. Similarly, the alignment and stitching system 1114 can send and receive requests, images, or data to and from any device or system via the communication network 1104. In some embodiments, the image capture position module 1204 can determine the image capture device position data of an image capture device (e.g., a camera, which may be a stand-alone camera, a smartphone, a media tablet, a laptop, etc.). The image capture device position data may indicate the position and orientation of the image capture device and / or the lens. In one example, the image capture position module 1204 utilizes the IMU of the user system 1110, the camera, the digital device equipped with the camera, or the 3D and

[0173] panorama capture and stitching system 1102 to capture an image. between any of the modules or data stores of the alignment and stitching system 1114 and the components of the exemplary environment 1100 of FIG. 11. In the same way, the alignment and stitching system 1114 can send and receive requests, images, or data to and from any device or system via the communication network 1104. The alignment and stitching system 1114 can send and receive requests, images, or data to and from any device or system via the communication network 1104.

[0174] In some embodiments, the image capture position module 1204 can determine the image capture device position data of an image capture device (e.g., a camera, which may be a stand-alone camera, a smartphone, a media tablet, a laptop, etc.). The image capture device position data may indicate the position and orientation of the image capture device and / or the lens. In one example, the image capture position module 1204 utilizes the IMU of the user system 1110, the camera, the digital device equipped with the camera, or the 3D and panorama capture and stitching system 1102 to capture an image. The image capture device position data may indicate the position and orientation of the image capture device and / or the lens. In one example, the image capture position module 1204 utilizes the IMU of the user system 1110, the camera, the digital device equipped with the camera, or the 3D and panorama capture and stitching system ११०२ to capture an image. ​The position data of the camera device can be generated. The image capture position module 1204 can determine the current direction, angle, or tilt of one or more image capture devices (or lenses). The image capture position module 1204 may utilize the GPS of the user system 1110 or the 3D and panoramic capture and stitching system 1102. For example, when the user wants to use the user system 1110 to capture a 360° view of a physical environment such as a living room, the user may hold the user system 1110 at eye level in front of himself or herself and initiate the capture of one of a plurality of images that will ultimately form one 3D panoramic image. To reduce the amount of parallax for the images and capture images suitable for the stitching and generation of 3D panoramic images, it may be preferable for one or more image capture devices to rotate about the center of the axis of rotation. The alignment and stitching system 1114 can receive position information (e.g., from an IMU) and determine the position of the image capture device or lens. The alignment and stitching system 1114 can receive and store the field of view of the lens. The guidance module 1220 can provide visual and / or audio information regarding the recommended initial position of the image capture device. The guidance module 1220 can make recommendations for the positioning of the image capture device for subsequent images. In one example, the guidance module 1220 can provide guidance to the user to rotate and position the image capture device such that the image capture device rotates near the center of rotation.

[0175] For example, when the user wants to use the user system 1110 to capture a 360° view of a physical environment such as a living room, the user may hold the user system 1110 at eye level in front of himself or herself and initiate the capture of one of a plurality of images that will ultimately form one 3D panoramic image. To reduce the amount of parallax for the images and capture images suitable for the stitching and generation of 3D panoramic images, it may be preferable for one or more image capture devices to rotate about the center of the axis of rotation. The alignment and stitching system 1114 can receive position information (e.g., from an IMU) and determine the position of the image capture device or lens. The alignment and stitching system 1114 can receive and store the field of view of the lens. The guidance module 1220 can provide visual and / or audio information regarding the recommended initial position of the image capture device. The guidance module 1220 can make recommendations for the positioning of the image capture device for subsequent images. In one example, the guidance module 1220 can provide guidance to the user to rotate and position the image capture device such that the image capture device rotates near the center of rotation. The alignment and stitching system 1114 can receive position information (e.g., from an IMU) and determine the position of the image capture device or lens. The alignment and stitching system 1114 can receive and store the field of view of the lens. The guidance module 1220 can provide visual and / or audio information regarding the recommended initial position of the image capture device. The guidance module 1220 can make recommendations for the positioning of the image capture device for subsequent images. In one example, the guidance module 1220 can provide guidance to the user to rotate and position the image capture device such that the image capture device rotates near the center of rotation. The alignment and stitching system 1114 can receive position information (e.g., from an IMU) and determine the position of the image capture device or lens. The alignment and stitching system 1114 can receive and store the field of view of the lens. The rule 1220 determines whether subsequent images are captured based on the field of view and / or characteristics of the image capture device. Rotate and position the image capture devices so that they are roughly aligned. This can provide guidance to users on how to

[0176] The guidance module 1220 may provide visual guidance to the user. The guidance module 1220 may be configured to provide guidance to the user system 1110 or the 3D and panoramic capture system. The viewer or display on the stitching system 1102 displays a marker or In some embodiments, the user system 1110 may place an arrow on the display. The camera may be a smartphone or tablet computer equipped with a QR code. When capturing a shot, the guidance module 1220 may display one or more markers (e.g., different colored marker or identical marker) on the output device and / or in the viewfinder. The user may then use these marks on the output device and / or viewfinder. The car may be used to align the next image.

[0177] User system 1110 or 3D and panoramic capture and stitching system 1 102 guides users to easily scan multiple images into a single panorama. There are many techniques for creating panoramas from multiple images. When the images are acquired, they may be stitched together. This reduces the time, efficiency, and effectiveness of stitching images together while reducing the need for correction. To improve the accuracy of the image capture, the image capture position module 1204 and the guidance module 1220 can assist the user in taking multiple images at positions that improve the quality, time efficiency, and effectiveness of image stitching for a desired panorama.

[0178] For example, after taking the first photo, the display of the user system 1110 may include two or more objects such as circles. The two circles may appear to be stationary with respect to the environment, and the two circles may be movable with the user system 1110. When aligning the two stationary circles with the circles that move with the user system 1110, the image capture device and / or the user system 1110 can be aligned for the next image.

[0179] In some embodiments, after taking an image with an image capture device, the image capture position module 1204 can obtain sensor measurements (including, for example, orientation, tilt, etc.) of the position of the image capture device. The image capture position module 1204 can determine one or more edges of the captured image by calculating the location of the edge of the field of view based on the above-mentioned sensor measurements. Further, or alternatively, the image capture position module 1204 can scan the image captured by the image capture device, identify the objects in the image (using, for example, the machine learning model described herein), determine one or more edges of the image, and determine one or more edges of the image by positioning the objects (such as circles or other shapes) at the edges of the display on the user system 1110.

[0180] The image capture position module 1204 indicates the positioning of the field of view for the next photo to the user ​​​​​​​​​​​​​Two objects can be displayed within the display of the user system 1110. These two objects can indicate positions representing where the edges of the last image in the environment are located. The image capture position module 1204 can continue to receive sensor measurements of the position of the image capture device and calculate two additional objects within the field of view. These two additional objects may be separated by the same width as the previous two objects. The first two objects may represent an edge of the captured image (e.g., the right edge of the image), but the next two additional objects representing the edge of the field of view may be on the opposite edge (e.g., the left edge of the field of view). By physically aligning the user with the first two objects at the edge of the image and the two additional objects on the opposite edge of the field of view, the image capture device can be positioned to take another image more effectively without using a tripod, sticking them together. This process can continue for each image until the user determines that the desired panorama has been captured. Although multiple objects have been described herein, it will be understood that the image capture position module 1204 can calculate the positions of one or more objects for positioning the image capture device. The above objects can be of any shape (e.g., circle, ellipse, square, emoji, arrow, etc.). In some embodiments, the above objects

[0181] can be of different shapes. In some embodiments, the distance between the objects representing the edges of the captured image can be any (e.g., circle, ellipse, square, emoji, arrow, etc.). In some embodiments, the above objects can be of different shapes.

[0182] In some embodiments, the distance between the objects representing the edges of the captured image may exist, and there may be a distance between the objects in the field of view. The user can be guided to move forward away so as to create a sufficient distance between the objects. Alternatively, the size of the objects in the field of view may change to match the size of the objects representing the edges of the captured images (e.g., getting closer to or farther from a position that allows the next image to be captured at a position that improves the stitching of the images) so that the image capture device is in the correct position.

[0183] In some embodiments, the image capture position module 1204 can estimate the position of the image capture device using the objects in the image captured by the image capture device. For example, the image capture position module 1204 may use GPS coordinates to determine the geographical location associated with the image. The image capture position module 1204 can use this position to identify landmarks that can be captured by the image capture device.

[0184] The image capture position module 1204 may include a 2D machine learning model for converting a 2D image into a 2D panoramic image. The image capture position module 1204 may include a 3D machine learning model for converting a 2D image into a 3D representation. In one example, the 3D representation can be used to display a three-dimensional walkthrough or visualization of the indoor and / or outdoor environment.

[0185] The 2D machine learning model may be for 2D panorama by stitching two or more 2D images. It may be trained to form or assist in forming a Raman image. 2D machine learning models can be trained using, for example, 2D images that include physical objects in the image and object identification information, whereby the 2D machine learning model is trained to identify objects in subsequent 2D images. Objects in the 2D image can assist in determining the position of one or more positions within the 2D image, thereby assisting in determining the edges of this 2D image, warping within this 2D image, and image alignment. Further, objects in the 2D image can assist in determining artifacts in the 2D image, blending of artifacts or boundaries between two image views, determining where to crop the image, and / or determining where to crop the image to assist. In some embodiments, the 2D machine learning model may be, for example, a neural network trained on 2D images, where the 2D images include depth information of the environment (from, for example, the LiDAR device of the user system 1110 or the 3 D and panoramic capture and stitching system 1102) and include physical objects in the image, thereby identifying physical objects, the positions of the physical objects, and / or the positions of the image capture device / field of view. The 2D machine learning model can assist in aligning and positioning two 2D images for stitching (or stitching two images) by identifying the physical object and the depth of the physical object relative to other aspects of the 2D image.

[0186] The 2D machine learning model can be any number of machine learning models (e.g., any number of neural networks), where the 2D images include depth information of the environment (from, for example, the LiDAR device of the user system 1110 or the 3 D and panoramic capture and stitching system 1102) and include physical objects in the image, thereby identifying physical objects, the positions of the physical objects, and / or the positions of the image capture device / field of view. The 2D machine learning model can assist in aligning and positioning two 2D images for stitching (or stitching two images) by identifying the physical object and the depth of the physical object relative to other aspects of the 2D image. In some embodiments, the 2D machine learning model may be, for example, a neural network trained on 2D images, where the 2D images include depth information of the environment (from, for example, the LiDAR device of the user system 1110 or the 3 D and panoramic capture and stitching system 1102) and include physical objects in the image, thereby identifying physical objects, the positions of the physical objects, and / or the positions of the image capture device / field of view. The 2D machine learning model can assist in aligning and positioning two 2D images for stitching (or stitching two images) by identifying the physical object and the depth of the physical object relative to other aspects of the 2D image. By identifying the physical object and the depth of the physical object relative to other aspects of the 2D image, the 2D machine learning model can assist in aligning and positioning two 2D images for stitching (or stitching two images). In some embodiments, the 2D machine learning model may be, for example, a neural network trained on 2D images, where the 2D images include depth information of the environment (from, for example, the LiDAR device of the user system 1110 or the 3 D and panoramic capture and stitching system 1102) and include physical objects in the image, thereby identifying physical objects, the positions of the physical objects, and / or the positions of the image capture device / field of view. The 2D machine learning model can assist in aligning and positioning two 2D images for stitching (or stitching two images) by identifying the physical object and the depth of the physical object relative to other aspects of the 2D image.

[0187] The 2D machine learning model can be any number of machine learning models (e.g., any number of neural may include a model generated by a neural network or the like.

[0188] The 2D machine learning model may be stored in the 3D and panoramic capture stitching system 110 2, the image stitching processor system 1106, and / or the user system 11 10. In some embodiments, the 2D machine learning model may be trained by the image stitching processor system 1106.

[0189] The image capture position module 1204 can estimate the position of the image capture device (a part of the field of view of the image capture device) based on the seams between two or more 2D images from the stitching module 1206, the warping of the images from the crop module 1208, and / or the image cropping from the image cropping module 1210. The stitching module 1206 can combine two or more 2D images to generate a 2D panorama based on the seams between two or more 2D images from the stitching module 1206, the warping of the images from the crop module 1208, and / or the image cropping, which has a larger field of view than the field of view of each of the two or more images above.

[0190] The stitching module 1206 can align two different 2D images that provide different viewpoints of the same environment or "stitch together" to generate a panoramic 2D image of the environment. For example, the stitching module 1206 can determine the capture position and orientation of each 2D image and combine two or more 2D images to generate a 2D panorama based on the seams between two or more 2D images from the stitching module 1206, the warping of the images from the crop module 1208, and / or the image cropping, which has a larger field of view than the field of view of each of the two or more images above. The stitching module 1206 may be configured to generate a panoramic 2D image of the environment by aligning two different 2D images that provide different viewpoints of the same environment or "stitching them together". For example, the stitching module 1206 can determine the capture position and orientation of each 2D image

[0191] The stitching module 1206 may be configured to generate a panoramic 2D image of the environment by aligning two different 2D images that provide different viewpoints of the same environment or "stitching them together". For example, the stitching module 1206 can determine the capture position and orientation of each 2D image and combine two or more 2D images to generate a 2D panorama based on the seams between two or more 2D images from the stitching module 1206, the warping of the images from the crop module 1208, and / or the image cropping, which has a larger field of view than the field of view of each of the two or more images above. For example, the stitching module 1206 can determine the capture position and orientation of each 2D image and combine two or more 2D images to generate a 2D panorama based on the seams between two or more 2D images from the stitching module 1206, the warping of the images from the crop module 1208, Using known information or information derived (e.g., using the techniques of this specification), 2 can assist in stitching two images into one.

[0192] The stitching module 1206 may receive two 2D images. The first 2D image may have been taken immediately before or within a predetermined period of the second 2D image. In various embodiments, the stitching module 1206 may receive positioning information of the image capture device associated with the first image, and positioning information associated with the second image. These positioning information can be associated with the above images based on positioning data from the IMU, GPS, and / or information provided by the user at the time of image capture.

[0193] In some embodiments, the stitching module 1206 can utilize 2D machine learning rules to scan both images and recognize objects within both images, where the above objects include objects (or parts of objects) that the two images may share. For example, the stitching module 1206 can identify corners, wall patterns, furniture, etc. that are shared at the opposing edges of both images.

[0194] Based on the positioning of the shared object (or part of the object), positioning data from the IMU, positioning data from the GPS, and / or information provided by the user, the stitching module 1206 aligns the edges of the two 2D images and combines the above two edges of these images (i.e., "stitches" these together into one). can be "stitched". In some embodiments, the stitching module 120 6 identifies a portion of two overlapping 2D images and stitches these images at the overlapping positions (e.g., using positioning data and / or the results of a 2D machine learning model).

[0195] In various embodiments, the 2D machine learning model may be trained to combine or stitch the two edges of an image using positioning data from an IMU, positioning data from a GPS, and / or information provided by a user. In some embodiments, the 2D machine learning model may be trained to align and position the two 2D images by identifying common objects within both 2D images and to combine or stitch the two edges of these images. In further embodiments, the 2D machine learning model may be trained to align and position 2D images using positioning data and object recognition to stitch the two edges of these images together to form all or part of a panoramic 2D image.

[0196] The stitching module 1206 can facilitate the alignment of each 2D image relative to each other for the generation of a single 2D panoramic image of the environment by utilizing depth information about each image (e.g., pixels within each image, objects within each image, etc.).

[0197] The crop module 1208 can solve problems with two or more 2D images when the image capture device was not held in the same position during the capture of the 2D images. For example ​​​​​​​​​​​​ Well, during the capture of a certain image, the user can position the user system 1110 at a certain vertical position. However, during the capture of another image, the user may position the user system at a certain angle. As a result, the obtained images may not be aligned and there is a risk of being troubled by the parallax effect. The parallax effect may occur when the foreground object and the background object are not aligned in the same way in the first image and the second image.

[0198] The crop module 1208 can detect the change in the position of the image capture device in two or more images by using a 2D machine learning model (by applying positioning information, depth information, and / or object recognition), and measure the amount of change in the position of the image capture device. The crop module 1208 can warp one or more 2D images so that these images can be arranged in a row to form a single panoramic image when these images are stitched, and at the same time, specific characteristics of the image can be preserved, such as keeping the straight line straight.

[0199] The output of the crop module 1208 may include the number of pixel columns and rows for offsetting each pixel of the image to make the image straight. The amount of offset for each image can be output in the form of a matrix representing the number of pixel columns and pixel rows for offsetting each pixel of the image.

[0200] In some embodiments, the crop module 1208 is implemented for one or more of the plurality of 2D images captured by the image capture device of the user system 1110. ​​​​​​​​​​​​ The amount of warping of the image to be performed can be determined based on one or more image capture positions from the image capture position module 1204, or seams between two or more 2D images from the stitching module 1206, image cropping from the image cropping module 1210, or color blending from the blending module 1211.

[0201] The image cropping module 1210 can determine the positions at which one or more of the 2D images captured by the image capture device should be cropped or sliced. For example, the image cropping module 1210 may utilize a 2D machine learning model to identify objects within both images and determine that they are the same object. The image capture position module 1204, the crop module 1208, and / or the image cropping module 1210 may determine that these two images cannot be aligned even if warped. The image cropping module 1210 may utilize information from the 2D machine learning model to identify sections of the two images that can be stitched together (e.g., by cropping a portion of one or both images to assist in alignment and positioning). In some embodiments, the two 2D images may overlap in at least a portion of the real-world scene represented within the images. The image cropping module 1210 can identify an object, e.g., the same chair, within both images. However, the images of this chair do not line up in a row to generate an undistorted panorama even after warping of the images by the image capture positioning and crop module 1208. There may be cases where it fails to correctly represent the above part of the real world. Image cropping module - Rule 1210 selects one of the two images of the chair as the correct representation (e.g., based on the displacement, positioning, and / or artifacts of one image compared to the other image) and can crop the chair from the image with displacement, positioning errors, and artifacts. The stitching module 1206 can then stitch the two images into one. The image cropping module 1210 can try to crop an image of the chair from the first image and stitch the remaining part after removing the chair from the first image to the second image, and determine which image cropping can generate a more precise panoramic image . The output of the image cropping module 1210 can be the location to crop one or more of the multiple 2D images corresponding to the image cropping that generates a more precise panoramic image

[0202] The image cropping module 1210 can determine how to crop or stitch one or more of the 2D images captured by the image capture device based on one or more image capture positions from the image capture position module 1204; the stitching or seam between two or more 2D images from the stitching module 1206; the warp transformation of the image from the crop module 1208; and the image cropping from the image cropping module 1210 . The blending module 1211 can blend the seam (e.g., stitching) between the two images .

[0203] The image cropping module 1210 can determine how to crop or stitch one or more of the 2D images captured by the image capture device based on one or more image capture positions from the image capture position module 1204; the stitching or seam between two or more 2D images from the stitching module 1206; the warp transformation of the image from the crop module 1208; and the image cropping from the image cropping module 1210 . The blending module 1211 can blend the seam (e.g., stitching) between the two images .

[0204] The blending module 1211 can blend the seam (e.g., stitching) between the two images ​​​The seam can be colored so that it is not visible. Objects or surfaces may appear in slightly different colors or shades. Blending The module includes: one or more image capture locations from the image capture location module 1204; Position; Stitching; Image color along the seam from two images; Crop module warp transformation of the image from module 1208; and / or image cropping module 1210 Based on the image crop, you can determine the amount of color blending required.

[0205] In various embodiments, the blending module 1211 performs the blending of two 2D images. The system may receive a panorama from the camera and sample colors along the seam between the two 2D images. The blend module 1211 extracts the seam location from the image capture location module 1204. The blending module 1211 may receive information about the location of the Colors can be sampled to determine differences (e.g., predetermined thresholds of color, hue, luminance, saturation, etc.) If there is a significant difference in color along the seam between the two images (within The rule 1211 is a predetermined size along the seam at the position where the difference exists. The images may be blended. In some embodiments, the difference in color or image along the seam is large. The larger the blending width, the more space along the seam between the two images is allowed to blend.

[0206] In some embodiments, after blending, the blending module 1211 Rescan and sample the colors to add color, hue, brightness, saturation, etc. to the image or color. It may be determined whether other differences exceeding the predetermined threshold exist. If so, The domodule 1211 may identify the portion along the seam and continue to blend the portion of the image. The blend module 1211 may continue to resample the image along the seam until there are no further portions of the image to be blended (e.g., the color difference is less than one or more predetermined thresholds). The blend module 1211 may continue to resample the image along the seam until there are no further portions of the image to be blended (e.g., the color difference is less than one or more predetermined thresholds). The blend module 1211 may continue to resample the image along the seam until there are no further portions of the image to be blended (e.g., the color difference is less than one or more predetermined thresholds). The blend module 1211 may continue to resample the image along the seam until there are no further portions of the image to be blended (e.g., the color difference is less than one or more predetermined thresholds).

[0207] The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to convert the 2D panoramic image into a 3D representation. The 3D machine learning model may be trained to create a 3D representation using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or a structured illumination device). The 3D representation may be tested and reviewed for curation and feedback. In some embodiments, a 3D representation can be generated by using the 3D machine learning model together with the 2D panoramic image and depth data. The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to convert the 2D panoramic image into a 3D representation. The 3D machine learning model may be trained to create a 3D representation using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or a structured illumination device). The 3D representation may be tested and reviewed for curation and feedback. In some embodiments, a 3D representation can be generated by using the 3D machine learning model together with the 2D panoramic image and depth data. The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to convert the 2D panoramic image into a 3D representation. The 3D machine learning model may be trained to create a 3D representation using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or a structured illumination device). The 3D representation may be tested and reviewed for curation and feedback. In some embodiments, a 3D representation can be generated by using the 3D machine learning model together with the 2D panoramic image and depth data. The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to convert the 2D panoramic image into a 3D representation. The 3D machine learning model may be trained to create a 3D representation using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or a structured illumination device). The 3D representation may be tested and reviewed for curation and feedback. In some embodiments, a 3D representation can be generated by using the 3D machine learning model together with the 2D panoramic image and depth data. The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to convert the 2D panoramic image into a 3D representation. The 3D machine learning model may be trained to create a 3D representation using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or a structured illumination device). The 3D representation may be tested and reviewed for curation and feedback. In some embodiments, a 3D representation can be generated by using the 3D machine learning model together with the 2D panoramic image and depth data. The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to convert the 2D panoramic image into a 3D representation. The 3D machine learning model may be trained to create a 3D representation using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or a structured illumination device). The 3D representation may be tested and reviewed for curation and feedback. In some embodiments, a 3D representation can be generated by using the 3D machine learning model together with the 2D panoramic image and depth data. The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to convert the 2D panoramic image into a 3D representation. The 3D machine learning model may be trained to create a 3D representation using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or a structured illumination device). The 3D representation may be tested and reviewed for curation and feedback. In some embodiments, a 3D representation can be generated by using the 3D machine learning model together with the 2D panoramic image and depth data. The 3D image generator 1214 can receive a 2D panoramic image and generate a 3D representation. In various embodiments, the 3D image generator 1214 utilizes a 3D machine learning model to convert the 2D panoramic image into a 3D representation. The 3D machine learning model may be trained to create a 3D representation using the 2D panoramic image and depth data (e.g., from a LiDAR sensor or a structured illumination device). The 3D representation may be tested and reviewed for curation and feedback. In some embodiments, a 3D representation can be generated by using the 3D machine learning model together with the 2D panoramic image and depth data.

[0208] In various embodiments, the accuracy, rendering speed, and quality of the 3D representation generated by the 3D image generator 1214 are significantly improved by utilizing the systems and methods described herein. For example, by using the methods described herein (e.g., by alignment and positioning information provided by hardware; by improved positioning resulting from guidance provided to the user during image capture; by cropping of the image and modification of the warp transformation; by cropping of the image to avoid artifacts and overcome the warp transformation; by blending of the image; and / or combinations thereof). In various embodiments, the accuracy, rendering speed, and quality of the 3D representation generated by the 3D image generator 1214 are significantly improved by utilizing the systems and methods described herein. For example, by using the methods described herein (e.g., by alignment and positioning information provided by hardware; by improved positioning resulting from guidance provided to the user during image capture; by cropping of the image and modification of the warp transformation; by cropping of the image to avoid artifacts and overcome the warp transformation; by blending of the image; and / or combinations thereof). In various embodiments, the accuracy, rendering speed, and quality of the 3D representation generated by the 3D image generator 1214 are significantly improved by utilizing the systems and methods described herein. For example, by using the methods described herein (e.g., by alignment and positioning information provided by hardware; by improved positioning resulting from guidance provided to the user during image capture; by cropping of the image and modification of the warp transformation; by cropping of the image to avoid artifacts and overcome the warp transformation; by blending of the image; and / or combinations thereof). In various embodiments, the accuracy, rendering speed, and quality of the 3D representation generated by the 3D image generator 1214 are significantly improved by utilizing the systems and methods described herein. For example, by using the methods described herein (e.g., by alignment and positioning information provided by hardware; by improved positioning resulting from guidance provided to the user during image capture; by cropping of the image and modification of the warp transformation; by cropping of the image to avoid artifacts and overcome the warp transformation; by blending of the image; and / or combinations thereof). In various embodiments, the accuracy, rendering speed, and quality of the 3D representation generated by the 3D image generator 1214 are significantly improved by utilizing the systems and methods described herein. For example, by using the methods described herein (e.g., by alignment and positioning information provided by hardware; by improved positioning resulting from guidance provided to the user during image capture; by cropping of the image and modification of the warp transformation; by cropping of the image to avoid artifacts and overcome the warp transformation; by blending of the image; and / or combinations thereof). In various embodiments, the accuracy, rendering speed, and quality of the 3D representation generated by the 3D image generator 1214 are significantly improved by utilizing the systems and methods described herein. For example, by using the methods described herein (e.g., by alignment and positioning information provided by hardware; by improved positioning resulting from guidance provided to the user during image capture; by cropping of the image and modification of the warp transformation; by cropping of the image to avoid artifacts and overcome the warp transformation; by blending of the image; and / or combinations thereof). In various embodiments, the accuracy, rendering speed, and quality of the 3D representation generated by the 3D image generator 1214 are significantly improved by utilizing the systems and methods described herein. For example, by using the methods described herein (e.g., by alignment and positioning information provided by hardware; by improved positioning resulting from guidance provided to the user during image capture; by cropping of the image and modification of the warp transformation; by cropping of the image to avoid artifacts and overcome the warp transformation; by blending of the image; and / or combinations thereof). By alignment, positioning, and stitching the 2D panoramic images, By rendering a 3D representation, the accuracy of the 3D representation, the speed of rendering, and The quality is improved. Further, by using the alignment, positioning, and stitched 2D panoramic images described herein, the training of the 3D machine learning model It will be understood that can be significantly improved (e.g., in terms of speed and accuracy). Further, in some Embodiments, the 3D machine learning model can be made smaller and less complex. This is to overcome misalignment, positioning errors, warp distortion, insufficient image cropping, Insufficient blending, artifacts, etc. and to generate a 3D representation of reasonable accuracy, because the processing and learning are reduced. The trained 3D machine learning model can be stored in the 3D and panoramic capture and stitching system 1102, the image stitching processor system 106, and / or the user system 1110. In some embodiments, the 3D machine learning model may be trained using a plurality of 2D images and depth data from the image capture devices of the user system 1110 and / or the 3D and panoramic capture and stitching system 1102. Further, the 3D image generator 1214: image capture position information associated with each of the plurality of 2D images from the image capture position module 1204; the seam locations for aligning or stitching each of the plurality of 2D images from the stitching module 1206

[0209]

[0210] ; One of the pixels for each of the plurality of 2D images from the crop module 1208 The above offsets; and / or used with the image cropping from the image cropping module 1210 It may be trained. In some embodiments, the 3D machine learning model is: 2D pano rama image; depth data; image capture position information associated with each of the plurality of 2D images from the image capture position module 1204; stitching module 1206 For each of the plurality of 2D images from, the seams for aligning or stitching each of the plurality of 2D images Location; one or more offsets of the pixels for each of the plurality of 2D images from the crop module 1208; and / or image cropping from the image cropping module 1210 Can be used together to generate a 3D representation.

[0211] The stitching module 1206 may be part of a 3D model that converts a plurality of 2D images into a 2D panorama or a 3D panorama image. In some embodiments, the 3D model Is a machine learning algorithm such as a 3D-from-2D prediction neural network model. The crop module 1208 may be part of a 3D model that converts a plurality of 2D images into a 2 D panorama or a 3D panorama image. In some Embodiments, the 3D model is a machine learning algorithm such as a 3D-from-2D prediction neural network model. The image cropping module 1210 may be part of a 3D model that converts a plurality of 2D images Into a 2D panorama or a 3D panorama image. In some Embodiments, the 3D model is a machine learning algorithm such as a 3D-from-2D prediction neural network model. The blend module 1211 is a plurality of 2D images May be part of a 3D model that converts into a 2D panorama or a 3D panorama image. In some Embodiments, the 3D model is a machine learning algorithm such as a 3D-from-2D prediction neural network model. The blend module 1211 is a plurality of 2D images may be part of a 3D machine learning model that converts it into a 2D panorama or 3D panorama image . In some embodiments, the 3D model is a machine learning algorithm such as a 3D-from-2D prediction neural network model.

[0212] The 3D image generator 1214 may generate weights for the image capture position module 1204, the crop module 1208, the image cutout module 1210, and the blend module 1211 respectively, which may represent the confidence of the modules, i.e., "strength" or "weakness". In some embodiments, the sum of the weights of these modules is equal to 1.

[0213] If depth data is not available for a plurality of 2D images, the 3D image generator 1214 may determine depth data for one or more objects in a plurality of 2D images captured by the image capture device of the user system 1110. In some embodiments, the 3D image generator 1214 may derive depth data based on images captured by a stereo image pair. The 3D image generator may not determine depth data from a passive stereo algorithm, but may evaluate the stereo image pair to determine data regarding photometric consistency (more intermediate results) between the images at various depths.

[0214] The 3D image generator 1214 may be part of a 3D model that converts a plurality of 2D images into a 2D panorama or 3D panorama image. In some embodiments, the 3D model is a machine learning algorithm such as a 3D-from-2D prediction neural network model.

[0215] The captured 2D image data store 1216 can be any suitable structure and / or structures (e.g., an active database, a relational database, a self - referencing database, a table, a matrix, an array, a flat file, a document - oriented storage system, a non - relational No - SQL system, an FTS management system such as Lucene / Solar, etc.) for the captured images and / or depth data. The captured 2D image data store 1216 can store the images captured by the image capture device of the user system 1110. In various embodiments, the captured 2D image data store 1216 can store the depth data captured by one or more depth sensors of the user system 1110. In various embodiments, the captured 2D image data store 1216 can store the image capture device characteristics associated with the image capture device, or the capture characteristics associated with each of a plurality of image captures or depth captures used in the determination of a 2D panoramic image. In some embodiments, the image data store 1108 can store a 2D panoramic image. The 2D panoramic image can be determined by the 3D and panoramic capture - stitching system 1102 or the image stitching - processing system 106. Image capture device parameters include illumination, color, focal length of the image capture lens, maximum aperture, tilt angle, etc. Capture characteristics include pixel resolution, lens distortion, illumination, and other image metadata.

[0216] The 3D panoramic image data store 1218 can have any structure suitable for 3D panoramic images and / or multiple structures (e.g., active database, relational database, self-referential database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, etc.). The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. and / or multiple structures (e.g., active database, relational database, self-referential database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, etc.). The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. and / or multiple structures (e.g., active database, relational database, self-referential database, table, matrix, array, flat file, document-oriented storage system, non-relational No-SQL system, FTS management system such as Lucene / Solar, etc.). The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store ၁၂၁၈ stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106. The 3D panoramic image data store 1218 can store 3D panoramic images generated by the 3D and panoramic capture / stitching system 1102. In various embodiments, the 3D panoramic image data store 1218 stores characteristics associated with the image capture device or characteristics associated with each of multiple image captures or depth captures used in the determination of 3D panoramic images. In some embodiments, the 3D panoramic image data store 1218 stores 3D panoramic images. The 2D or 3D panoramic images can be determined by the 3D and panoramic capture / stitching system 1102 or the image stitching / processor system 106.

[0217] Figure 13 shows a flowchart 1300 of a 3D panoramic image capture / generation process according to some embodiments. In step 1302, the image capture device can capture a plurality of 2D images using the image sensor 920 and the WFOV lens 918 of FIG. 9. A wider FOV means that the environmental capture system 402 requires fewer scans to obtain a 360° view. The WFOV lens 918 can also be wider horizontally and vertically. In some embodiments, the image sensor 920 Figure 13 shows a flowchart 1300 of a 3D panoramic image capture / generation process according to some embodiments. In step 1302, the image capture device can capture a plurality of 2D images using the image sensor 920 and the WFOV lens 918 of FIG. 9. A wider FOV means that the environmental capture system 402 requires fewer scans to obtain a 360° view. The WFOV lens 918 can also be wider horizontally and vertically. In some embodiments, the image sensor 920 Figure 13 shows a flowchart 1300 of a 3D panoramic image capture / generation process according to some embodiments. In step 1302, the image capture device can capture a plurality of 2D images using the image sensor 920 and the WFOV lens 918 of FIG. 9. A wider FOV means that the environmental capture system 402 requires fewer scans to obtain a 360° view. The WFOV lens 918 can also be wider horizontally and vertically. In some embodiments, the image sensor 920 Figure 13 shows a flowchart 1300 of a 3D panoramic image capture / generation process according to some embodiments. In step 1302, the image capture device can capture a plurality of 2D images using the image sensor 920 and the WFOV lens 918 of FIG. 9. A wider FOV means that the environmental capture system 402 requires fewer scans to obtain a 360° view. The WFOV lens 918 can also be wider horizontally and vertically. In some embodiments, the image sensor 920 Figure 13 shows a flowchart 1300 of a 3D panoramic image capture / generation process according to some embodiments. In step 1302, the image capture device can capture a plurality of 2D images using the image sensor 920 and the WFOV lens 918 of FIG. 9. A wider FOV means that the environmental capture system 402 requires fewer scans to obtain a 360° view. The WFOV lens 918 can also be wider horizontally and vertically. In some embodiments, the image sensor 920 Figure 13 shows a flowchart 1300 of a 3D panoramic image capture / generation process according to some embodiments. In step 1302, the image capture device can capture a plurality of 2D images using the image sensor 920 and the WFOV lens 918 of FIG. 9. A wider FOV means that the environmental capture system 402 requires fewer scans to obtain a 360° view. The WFOV lens 918 can also be wider horizontally and vertically. In some embodiments, the image sensor 920 Capture an RGB image. In one embodiment, the image sensor 920 captures a black image and a white image.

[0218] In step 1304, the environmental capture system may send the captured 2D image to the image stitching processor system 1106. The image stitching processor system 1106 can obtain a panoramic 2D image by applying a 3D modeling algorithm to the captured 2D image. In some embodiments, the 3D modeling algorithm is a machine learning algorithm for stitching the captured 2D images into one panoramic 2D image. In some embodiments, step 1304 may be optional.

[0219] In step 1306, the LiDAR 912 and the WFOV lens 918 in FIG. 9 may capture LiDAR data. The wider FOV means that the environmental capture system 400 requires fewer scans to obtain a 360° view.

[0220] In step 1308, the LiDAR data may be sent to the image stitching processor system 1106. The image stitching processor system 1106 inputs the LiDAR data and the captured 2D image into a 3D modeling algorithm to generate a 3D panoramic image. The 3D modeling algorithm is a machine learning algorithm.

[0221] In step 1310, the image stitching processor system 1106 is a 3D pano​​​​​​​​​​​​​Generate a panoramic image. The 3D panoramic image may be stored in the image data store 408. In one embodiment, the 3D panoramic image generated by the 3D modeling algorithm is stored in the image stitching processor system 1106. In some embodiments the 3D modeling algorithm can utilize an environmental capture system to capture various parts of the physical environment, thus generating a visual representation of the floor plan of the physical environment.

[0222] In step 1312, the image stitching processor system 1106 may provide at least a portion of the generated 3D panoramic image to the user system 1110. The image stitching processor system 1106 can provide a visual representation of the floor plan of the physical environment. It can be provided.

[0223] The order of one or more steps of the flowchart 1300 can be changed without affecting the final product of the 3D panoramic image. For example, the environmental capture system can sandwich LiDAR data captured by the LiDAR 912 or depth information capture between image captures by the image capture device. For example, the image capture device may capture an image of a section of the physical environment, and then the LiDAR 912 obtains depth information from this section 1605. After the LiDAR 912 obtains depth information from this section, the image capture device may move to capture an image of another section, and subsequently the LiDAR 912 obtains depth information from this section. In this way, image capture and depth information capture are performed alternately.

[0224] ​In some embodiments, the devices and / or systems described herein capture a 2D input image using one image capture device. In some embodiments, one or more image capture devices 1116 can represent a single image capture device (or image capture lens). According to some of these embodiments, a user of a mobile device housing the image capture device can rotate about an axis and be configured to generate images at different capture orientations with respect to the environment, and the combined field of view of these images extends up to 360° horizontally.

[0225] In various embodiments, the devices and / or systems described herein may capture a 2D input image using two or more image capture devices. In some embodiments, two or more image capture devices can be arranged at relative positions with respect to each other on or within the same mobile housing such that their combined field of view covers 360°. In some embodiments, a plurality of pairs of image capture devices (e.g., having partially overlapping fields of view that are slightly offset) can be used to generate pairs of stereo images. For example, a user system 1110 (e.g., a device having one or more image capture devices used to capture a 2D input image) can include two image capture devices with a horizontal stereo offset field of view that can capture a pair of stereo images. In another example, the user system 1110 can include two image capture devices with a vertical stereo offset field of view that can capture a pair of vertical stereo images. Two image capture devices with a fuselage view can be provided. These examples According to any of them, each camera can have a field of view covering 360°. In this regard In one embodiment, the user system 1110 can capture a pair of panoramic images that form a stereo pair (having a vertical stereo offset ). Two panoramic cameras with a vertical stereo offset can be used.

[0226] The positioning component 1118 can include any hardware and / or software configured to capture user system position data and / or user system location data. For example, the positioning component 1118 includes an IMU to generate position data of the user system 1110 associated with one or more image capture devices of the user system 1110 used to capture a plurality of 2D images . The positioning component 1118 may include a GPS unit to provide GPS coordinate information associated with a plurality of 2D images captured by one or more image capture devices. In some embodiments, the positioning component 1118 can correlate the position data and location data of the user system with each image captured using one or more image capture devices of the user system 1110 . The positioning component 1118 can include a GPS unit to provide GPS coordinate information associated with a plurality of 2D images captured by one or more image capture devices. In some embodiments, the positioning component 1118 can correlate the position data and location data of the user system with each image captured using one or more image capture devices of the user system 1110 . In some embodiments, the positioning component 1118 can correlate the position data and location data of the user system with each image captured using one or more image capture devices of the user system 1110 .

[0227] Various embodiments of the device provide 3D panoramic images of indoor and outdoor environments to the user. In some embodiments, the device can efficiently and quickly provide 3D panoramic images of indoor and outdoor environments to the user using a single wide field of view (FOV) lens and a single optical detection and ranging sensor (LiDAR sensor) . ​

[0228] The following is an exemplary usage example of an exemplary device described in this specification. The following usage example is one of a plurality of embodiments. As described in this specification, different embodiments of the above device may include one or more features and functions similar to this usage example.

[0229] FIG. 14 shows a flowchart of a 3D and panoramic capture and stitching process 1400 according to some embodiments. The flowchart of FIG. 14 shows the 3D and panoramic capture and stitching system 1102 as including an image capture device, although in some embodiments, the data capture device may be the user system 1110.

[0230] ​​​​​​​​​​​​​​​​

[0231] In some embodiments, the multiple 2D images are all captured from the same image capture device. In various embodiments, at least a portion of the plurality of 2D images are received. Two or more image capture devices of the panoramic capture and stitching system 1102 In one example, the plurality of 2D images may include a set of RGB images and an IR image. Includes a set of IR image, 3D and panoramic capture stitching system 1 102. In some embodiments, each 2D image is captured by a LiDAR device. In some embodiments, the depth data provided by the device can be correlated with the Each 2D image can be associated with positioning data.

[0232] Step 1404: The 3D and panoramic capture and stitching system 1102 , capture parameters and image capture parameters associated with each of the received 2D images. The image capture device parameters may be received as These include lighting, color, focal length of the image capture lens, maximum aperture, field of view, etc. Image characteristics include pixel resolution, lens distortion, lighting, and other image metadata. The 3D and panoramic capture and stitching system 1102 is The sensor may also receive position and depth data.

[0233] In step 1406, the 3D and panoramic capture and stitching system 110 2. Stitching the above 2D image using the information received from steps 1402 and 1404 This can be used to stitch 2D images together to form a 2D panoramic image. This process is further described in connection with the flowchart of FIG.

[0234] In step 1408, the 3D and panoramic capture and stitching system 110 2 may apply a 3D machine learning model to generate a 3D representation. The 3D representation may be a 3D panoramic view. In various embodiments, the 3D representation may be stored in an image data store. generated by the switching processor system 1106. In some embodiments, ,3D machine learning models utilize an environment capture system to capture,various parts of the physical environment. To capture this, a visual representation of the floor plan of the physical environment can be generated.

[0235] In step 1410, the 3D and panoramic capture and stitching system 110 2 provides at least a portion of the generated 3D representation or model to a user system 1110. The user system 1110 may provide a visual representation of the floor plan of the physical environment. .

[0236] In some embodiments, the user system 1110 may include a plurality of 2D images, a capture pattern, and a parameters, and image capture parameters, image stitching processor system In various embodiments, 3D and panoramic capture stations may be used. The imaging system 1102 includes a plurality of 2D images, capture parameters, and an image capture The parameters may be sent to the image stitching processor system 1106.

[0237] The image stitching processor system 1106 processes the image of the user system 1110. Process multiple 2D images captured by the capture device and combine them into a 2D panorama. It may be stitched to an image. The image stitching processor system 1106 The processed 2D panoramic image may have a higher pixel resolution than the 2D panoramic image obtained by the 3D and panoramic capture stitching system 1 102.

[0238] In some embodiments, the image stitching processor system 106 may receive a 3D representation and output a 3D panoramic image having a higher pixel resolution than the received 3D panoramic image. This panoramic image with a higher pixel resolution can be supplied to an output device having a higher screen resolution than the user system 11 10, such as a computer screen, a projector screen, etc. In some embodiments, this panoramic image with a higher pixel resolution can provide a more detailed panoramic image to the output device and is also enlargeable. 10, such as a computer screen, a projector screen, etc. In some embodiments, this panoramic image with a higher pixel resolution can provide a more detailed panoramic image to the output device and is also enlargeable. enlargeable.

[0239] FIG. 15 shows a flowchart showing further details of one step of the 3D and panoramic capture stitching process of FIG. 14. In step 1502, the image capture position module 1204 may determine the image capture device position data associated with each image captured by the image capture device. The image capture position module 1204 may utilize the IMU of the user system 1110 to determine the position data of the image capture device (or the field of view of the lens of the image capture device). The above position data may include the direction, angle, or tilt of one or more image capture devices during the capture of one or more 2D images. The crop module 1208, the image cropping module 1 1204 may utilize the IMU of the user system 1110 to determine the position data of the image capture device (or the field of view of the lens of the image capture device). The above position data may include the direction, angle, or tilt of one or more image capture devices during the capture of one or more 2D images. The crop module 1208, the image cropping module 1 angle, or tilt of one or more image capture devices during the capture of one or more 2D images. The crop module 1208, the image cropping module 1 1208, the image cropping module 1 ​​​One or more of 210 and blend module 1212 may determine how to warp, crop, and / or blend these images using the directions, angles, or inclinations associated with each of the plurality of 2D images.

[0240] In step 1504, crop module 1208 may warp one or more of the plurality of 2D images so that these two images can be arranged in a row to form a single panoramic image, and at the same time, certain characteristics of the image can be preserved, such as keeping the straight line straight. The output of crop module 1208 may include the number of pixel columns and rows for straightening the image by offsetting each pixel of the image. The amount of offset for each image can be output in the form of a matrix representing the number of pixel columns and pixel rows for offsetting each pixel of the image. In this embodiment, crop module 12 08 may determine the amount of warping required for each of the plurality of 2D images based on the estimated image capture poses of each of the plurality of 2D images.

[0241] In step 1506, image cropping module 1210 determines the position where one or more of the plurality of 2D images should be cropped or sliced. In this embodiment, image cropping module 1210 may determine the position where each of the plurality of 2D images should be cropped or sliced based on the estimated image capture poses and image warping of each of the plurality of 2D images.

[0242] In step 1508, stitching module 1206 stitches the edges and / or ​​​​​​​​​​​​​Two or more images may be stitched into one using image cropping. Stitching module 1206 may align and / or position the images based on objects detected in the image, warp transformation, image cropping, etc.

[0243] In step 1510, blend module 1212 may adjust the seam (e.g., stitching of two images), or the location of an image that touches or connects to another image. Blend module 1212 can determine the amount of color blend needed based on: one or more image capture positions from image capture position module 1204; warp transformation of the image from crop module 1208; and / or image cropping from image cropping module 1210.

[0244] The order of one or more steps of the 3D and panoramic capture and stitching process 1400 can be changed without affecting the final product of the 3D panoramic image. For example, an environmental capture system can sandwich LiDAR data or depth information capture between image captures by an image capture device. For example, the image capture device may capture an image of section 1605 of the physical environment in FIG. 16, and then LiDAR 612 obtains depth information from section 1605. When LiDAR obtains depth information from section 1605, the image capture device may move to capture an image of another section 1610, and subsequently LiDAR 612 obtains depth information from section 1610. In this way, image capture and depth information capture are performed alternately.

[0245] ​​​​​​​​​​​​​​​ FIG. 16 is a block diagram of an exemplary digital device 1602 according to some embodiments is shown. Any of user system 1110, 3D panoramic capture and stitching system 1 102, and the image stitching and processor system may include an instance of digital device 1602. The digital device 1602 includes a processor 1604, a memory 1606, a storage 1608, an input device 1610, a communication network interface 1612, an output device 1614, an image capture device 1616 , and a positioning component 1618. The processor 1604 is configured to execute executable instructions ( e.g., a program). In some embodiments, processor 1 604 includes a circuit or any processor capable of processing executable instructions.

[0246] The memory 1606 stores data. Some examples of the memory 1606 include storage devices such as RAM, ROM, RAM cache, virtual memory, and the like. In various embodiments, the working data is stored in the memory 1606. The data in the memory 1606 may be cleared or ultimately transferred to the storage 1608.

[0247] The storage 1608 includes any storage configured to acquire and store data device. Some examples of the storage 1608 include flash drives, hard drives , optical drives, and / or magnetic tapes. The memory 1606 and the storage 1608 each include a computer-readable medium, which stores executable instructions or programs for the processor 1604 to execute.

[0248] The input device 1610 is any device that inputs data (e.g., a touch keyboard, a stylus). The output device 1614 outputs data (e.g., a speaker, a display, a virtual reality headset). Storage 1608, input device 1610, and output device 1614 will be understood. In some embodiments, the output device 1614 can be anything. For example, a router / switcher can include a processor 1604 and memory 1606, as well as a device for receiving and outputting data (e.g., a communication network interface 1612 and / or output device 1614).

[0249] The communication network interface 1612 may be coupled to a network (e.g., communication network 104) via the communication network interface 161 2. The communication network interface 1612 can support communication via Ethernet connection, serial connection, parallel connection, and / or ATA connection. The communication network interface 1612 can also support wireless communication (e.g., 802.16 a / b / g / n, WiMAX, LTE, W i-Fi). It will be apparent that the communication network interface 1612 can support both wired and wireless standards.

[0250] The components can be hardware or software. In some embodiments, the components can be configured with one or more processors to perform the functions associated with the components. Although various components are described herein, the server system can include any number of components that perform any of the functions described herein. ​​​It will be understood.

[0251] The digital device 1602 may include one or more image capture devices 1616. The one or more image capture devices 1616 can include, for example, an RGB camera, an HDR camera, a video camera, etc. The one or more image capture devices 1616 can include, in accordance with some embodiments, a video camera that can capture video. In some embodiments, the one or more image capture devices 1616 can include an image capture device that provides a relatively standard field of view (e.g., about 75°). In other embodiments, the one or more image capture devices 1616 can include a camera that provides a relatively wide field of view (e.g., about 120° - 360°), such as a fish-eye camera ( for example, the digital device 1602 may include or be included in the environmental capture system 400).

[0252] The components may be hardware or software. In some embodiments, the components can be configured with one or more processors to perform the functions associated with the components. Although various components are described herein, the server system can include any number of components to perform any of the functions described herein. It will be understood.

Claims

1. An image capture device comprising: A housing having a front face and a rear face; A first motor coupled to the housing at a first position between the front face and the rear face of the housing, the first motor being configured to turn the image capture device substantially 270° horizontally about a vertical axis; A first motor coupled to the housing at a first position between the front face and the rear face of the housing, the first motor being configured to turn the image capture device substantially 270° horizontally about a vertical axis; A first motor configured to turn the image capture device substantially 270° horizontally about a vertical axis; At a second position between the front face and the rear face of the housing along the vertical axis, a wide-angle lens coupled to the housing, the second position being the nodal point, the wide-angle lens having a field of view away from the front face of the housing; At a second position between the front face and the rear face of the housing along the vertical axis, a wide-angle lens coupled to the housing, the second position being the nodal point, the wide-angle lens having a field of view away from the front face of the housing; A wide-angle lens having a field of view away from the front face of the housing; An image sensor coupled to the housing and configured to generate an image signal from the light received by the wide-angle lens; An image sensor; A mount coupled to the first motor; A LiDAR coupled to the housing at a third position, the LiDAR being configured to generate laser pulses and generate a depth signal; A LiDAR configured to generate laser pulses and generate a depth signal; A second motor coupled to the housing; and A mirror coupled to the second motor, the second motor being configured to rotate the mirror about a horizontal axis, the mirror including an angled surface configured to receive the laser pulses from the LiDAR and direct the laser pulses about the horizontal axis. A mirror including an angled surface configured to receive the laser pulses from the LiDAR and direct the laser pulses about the horizontal axis. A mirror including an angled surface configured to receive the laser pulses from the LiDAR and direct the laser pulses about the horizontal axis. An image capture device comprising a mirror including an angled surface configured to receive the laser pulses from the LiDAR and direct the laser pulses about the horizontal axis. An image capture device comprising a mirror including an angled surface configured to receive the laser pulses from the LiDAR and direct the laser pulses about the horizontal axis.

2. The image sensor of claim 1, wherein the image sensor is configured to generate a plurality of first images with different exposures when the image capture device is stationary and facing a first direction. The image sensor of claim 1, wherein the image sensor is configured to generate a plurality of first images with different exposures when the image capture device is stationary and facing a first direction. The image capture device of claim 1.

3. The first motor of claim 2, wherein the first motor is configured to turn the image capture device about the vertical axis after the plurality of first images are generated. ​ ​ ​ ​ ​ ​ ​ ​ ​ configured to generate a second plurality of images with the different plural exposures, and the first mo tor is configured to turn the image capture device 90° around the vertical axis after the generation of the second plurality of images, the image capture device according to claim 3.

6. The image sensor is configured to generate a third plurality of images with the different plural exposures when the image capture device is stationary and facing a third direction, and the first mo tor is configured to turn the image capture device 90° around the vertical axis after the generation of the third plurality of images, the image capture device according to claim 5.

7. The image sensor is configured to generate a fourth plurality of images with the different plural exposures when the image capture device is stationary and facing a fourth direction, and the first mo tor is configured to turn the image capture device 90° around the vertical axis after the generation of the fourth plurality of images, the image capture device according to claim 6.

8. The image capture device according to claim 7, further comprising a processor configured to blend frames of the first plurality of images before the image sensor generates the second plurality of images.

9. The image capture device according to claim 7, further comprising a remote digital device configured to communicate with the image capture device and generate 3D visualization based on the first, second, third, and fourth pluralities of images and the depth signal, the remote digital device being configured to generate the 3D visualization without using images other than the first, second, third, and fourth pluralities of images.

10. The first, second, third, and fourth pluralities of images are generated during a turn that combines plural turns of turning the image capture device 270° around the vertical axis, the image capture device according to claim 9.

11. The speed or rotation of the mirror around the horizontal axis increases when the first motor turns the image capture device, the image capture device according to claim 4.

12. The angled surface of the mirror is 90°, the image capture device according to claim 1.

13. ​ ​ ​ ​ ​ 。 ​ ​ ​ ​ The LiDAR emits the laser pulse in a direction opposite to the front surface of the housing. The image capture device according to claim 1.

14. A method comprising: Receiving light from a wide-angle lens of an image capture device, wherein the wide-angle lens is coupled to a housing of the image capture device, the light is received within a field of view of the wide-angle lens, and the field of view extends away from a front surface of the housing; Generating, by an image sensor of the image capture device, a first plurality of images using the light from the wide-angle lens, wherein the image sensor is coupled to the housing and the first plurality of images are at different exposures; Turning, by a first motor, the image capture device substantially 270° horizontally about a vertical axis, wherein the first motor is coupled to the housing at a first position between a front surface and a back surface of the housing, the wide-angle lens is at a second position along the vertical axis, and the second position is a nadir point; ; Rotating, by a second motor, a mirror having an angled surface about a horizontal axis, wherein the second motor is coupled to the housing; Generating, by a LiDAR, a laser pulse, wherein the LiDAR is coupled to the housing at a third position and the laser pulse is directed at the rotating mirror while the image capture device is turning horizontally; and Generating, by the LiDAR, a depth signal based on the laser pulse. The method of claim 14, wherein the step of generating the first plurality of images by the image sensor is performed before the image capture device turns horizontally.

15. The method of claim 14, wherein the image sensor does not generate an image while the first motor is turning the image capture device, and the LiDAR generates the depth signal based on the laser pulse while the first motor is turning the image capture device.

16. When the image capture device is stationary and facing a second direction, the image sensor... ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ generating, by the image sensor, a second plurality of images with the different plurality of exposures; and after generating the second plurality of images, turning, by the first motor, the image capture device 90° around the vertical axis The method according to claim 16, further comprising.

18. When the image capture device is stationary and facing a third direction, generating, by the image sensor, a third plurality of images with the different plurality of exposures; and after generating the third plurality of images, turning, by the first motor, the image capture device 90° around the vertical axis The method according to claim 17, further comprising.

19. When the image capture device is stationary and facing a fourth direction, further comprising generating, by the image sensor, a fourth plurality of images with the different plurality of exposures, The method according to claim 18.

20. further comprising generating 3D visualization using the first, second, third, and fourth pluralities of images and based on the depth signal, wherein the step of generating the 3D visualization does not use any other images, The method according to claim 19 。

21. The method according to claim 17, further comprising blending frames of the first plurality of images before the image sensor generates the second plurality of images.

22. The first, second, third, and fourth pluralities of images are generated during turns that combine a plurality of turns of turning the image capture device 270° around the vertical axis, The method according to claim 19.

23. The speed or rotation of the mirror around the horizontal axis increases when the first motor turns the image capture device, The method according to claim 1. ​ ​ ​

Citation Information

Patent Citations

  • Laser scanner with dynamic adjustment of angle scan speed

    JP2015535337A

  • Systems and methods for capturing and generating panoramic three-dimensional images

    JP2023509137A

  • Panoramic-imaging digital camera, and panoramic imaging system

    WO2014156747A1