Method and apparatus for precise mobile data acquisition
By employing image alignment techniques and multiple resampling interpolation operations, the problem of inaccurate trajectory estimation in existing 3D measurement systems in dynamic environments has been solved, enabling efficient and accurate trajectory estimation of sensors in the environment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FARO TECHNOLOGIES INC
- Filing Date
- 2024-10-03
- Publication Date
- 2026-04-21
AI Technical Summary
Existing 3D measurement systems struggle to accurately estimate sensor motion trajectories in environments, especially in dynamic scenes, when dealing with noise and uncertainties, resulting in inaccurate and unreliable measurement results.
Image alignment technology is used to estimate the trajectory of captured image sequences. The image sequence is formed by the sensor moving in the environment. The processing system performs photogrammetry, and the structured point cloud is generated by multiple resampling and interpolation operations, which reduces the computational requirements and improves the accuracy.
It enables reliable and accurate estimation of sensor trajectories in dynamic environments, reducing time, data storage, and bandwidth requirements while maintaining measurement accuracy and efficiency.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] Cross-references to related applications
[0002] This application relates to provisional patent application serial number 63 / 542,922 entitled "METHOD AND APPARATUS FOR ACCURATEMOBILE DATA CAPTURE SENSOR", filed on October 6, 2023, the contents of which are incorporated herein by reference. Background Technology
[0003] The subject matter disclosed herein relates to trajectory estimation, and more particularly to estimating the trajectory of a three-dimensional motion measurement device that moves through and / or relative to the environment.
[0004] Processing systems (e.g., smartphones, laptops, tablets, wearable computing devices, etc., including combinations and / or multiple such devices) may include sensors for capturing images of objects or environments. In some cases, images are processed, analyzed, or otherwise used for a purpose, such as measuring the environment or object. For example, photogrammetry is a technique for measuring objects using images, such as photographic images acquired by sensors of a processing system or other suitable sensors. Photogrammetry can perform three-dimensional (3D) measurements from two-dimensional (2D) images or photographs. In some cases, photogrammetry involves using triangulation to determine three-dimensional (3D) coordinates, which is based at least in part on common features or landmarks between two images.
[0005] Therefore, while existing 3D measurement systems are suitable for their intended purpose, this paper introduces a system with improved features. Summary of the Invention
[0006] According to one aspect of this disclosure, a system for scanning an environment is provided. The system includes a sensor that acquires (captures) multiple images as the system traverses the environment. A processing system is also provided, having at least one processor that executes at least one photogrammetric function (function or algorithm) to generate a data sequence and performs two passes of sampling and interpolation (interpolation or interpolation value) on the data sequence to generate a single (monolithic) structured point cloud.
[0007] In another embodiment, the method includes: performing edge filtering on the data of operation 1 using a moving window four times the size of the cell; and performing edge filtering on the data of operation 2 using a moving window four times the size of the cell.
[0008] The above-described features and advantages, as well as other features and advantages, of this disclosure will become apparent when taken in conjunction with the accompanying drawings and from the following detailed description. Attached Figure Description
[0009] From the following detailed description taken in conjunction with the accompanying drawings, the foregoing and other features and advantages of one or more embodiments described herein will be apparent:
[0010] Figure 1A-1C A system for collecting and / or displaying environmental images according to one or more embodiments described herein is shown;
[0011] Figure 2 This is a schematic diagram of a processing system for trajectory estimation using image sequences, according to one or more embodiments described herein;
[0012] Figure 3 This is an example of a movable window according to one or more embodiments described herein, wherein gray grid lines represent a 24MP grid and black borders represent a window twice the size of a grid cell;
[0013] Figure 4 It is a dot plot obtained at the end of resampling operations 1 and 2 below, according to one or more embodiments described herein;
[0014] Figure 5 The following diagram illustrates the movement of windows in operations 1 and 2, where some grid lines represent 24MP grids, others represent the movement of windows in operation 1, and the remaining grid lines represent the movement of windows in operation 2. The highlighted cells in the grid will be filled by the corresponding operation according to one or more embodiments described herein.
[0015] Figure 6 An example of a moving window centered on a blank grid cell being filled with interpolated (interpolated) point distances is shown, and a range of point distances that can be used to calculate interpolated point distances according to one or more embodiments described herein is further shown;
[0016] Figure 7 A method for interpolation using a moving window only in the horizontal offset direction is shown, as illustrated in Operation 3 below, which is performed according to one or more embodiments described herein;
[0017] Figure 8 A method for interpolation using a moving window only in the vertical offset direction is shown, as illustrated in Operation 4 below, which is performed according to one or more embodiments described herein;
[0018] Figure 9 and Figure 10 A mobile scanning device of one type utilizing the method (process) described herein is shown according to one or more embodiments; and
[0019] Figure 11It is the output generated by a mobile scanning device according to one or more embodiments.
[0020] The following detailed description illustrates embodiments, advantages, and features of this disclosure with reference to the examples in the accompanying drawings. Detailed Implementation
[0021] The embodiments described herein provide a method for trajectory estimation using a series of images acquired by a mobile sensor device as it moves through a target environment. Various embodiments offer advantages in reducing the computational requirements of one or more processors, thereby reducing time, data storage, and bandwidth requirements, while still determining a reliable and accurate trajectory for the sensor acquiring the images.
[0022] Trajectory estimation is the process of accurately determining the path and motion of an object or intelligent agent over time, given or relative to its environment. Trajectory estimation is extremely useful in many fields, such as robotics, computer vision, aerospace engineering, autonomous systems, and combinations thereof. It helps in understanding the behavior and dynamics of moving entities, thereby providing accurate predictions and control in dynamic scenarios.
[0023] Trajectory estimation plays a crucial role in a wide range of tasks. For example, in robotics, trajectory estimation is used for robot localization, mapping, and navigation. By accurately estimating a robot's trajectory, it can determine its current position, create a map of its surroundings, and plan the optimal path to its destination or to perform a specific task. In computer vision, trajectory estimation is commonly used to track the movement of objects in video sequences, enabling applications such as object tracking, surveillance, and activity recognition. In aerospace engineering, trajectory estimation is used for the navigation and control of spacecraft and aircraft. For example, it is used to maneuver aircraft / spacecraft, determine trajectories, and perform other tasks that can affect mission success.
[0024] Trajectory estimation can be extremely challenging due to various uncertainties and noise sources, such as sensor errors, occlusion, environmental changes, dynamic interactions with other objects (e.g., moving objects), and combinations and / or multiple combinations of these factors. To address these challenges, a variety of techniques can be employed, including sensor fusion, probabilistic filtering (e.g., Kalman filtering, particle filtering), machine learning, and optimization methods.
[0025] To address the shortcomings of existing technologies, one or more embodiments described herein provide a method for trajectory estimation of captured image sequences using image alignment techniques. For example, when a sensor moves in an environment, it captures images during this movement, forming an image sequence. Using one or more embodiments described herein, this image sequence can be used to estimate the sensor's motion trajectory in the environment during the capture process.
[0026] Referring now to FIG1, an embodiment of a system 100 for acquiring and / or displaying environmental images according to at least one embodiment described herein is shown. Specifically, FIG1A shows a system 100 for acquiring and / or displaying images of an environment or object, the system 100 having a sensor 104 and a processing system 102.
[0027] Processing system 102 can be any suitable processing system, such as a built-in processing unit, a smartphone, tablet, laptop, a node in a cloud computing system, etc. Although not shown, in various embodiments, processing system 102 includes other components, such as at least one processor or microprocessor for executing processing instructions, memory (e.g., a hard disk drive), random access memory, read-only memory, memory chips, etc., for storing processing instructions and data in a computer-readable format for electronic storage. In various embodiments, processing system 102 also includes a display for displaying a user interface, an input device for receiving input, an output device for generating output, a communication adapter for communicating with other devices (e.g., sensor 104), and / or similar components, including combinations and / or multiple components of the above.
[0028] Sensor 104 can capture one or more environmental images, such as panoramic images. For example, sensor 104 may be an ultra-wide-angle sensor. For example, sensor 104 may be an omnidirectional sensor, such as a RICO THETA or INSTA360 sensor. In one embodiment, sensor 104 includes at least one sensor 110 (FIG. 1B), which in various embodiments includes a photosensitive pixel array. Sensor 110 is configured to receive light from lens 112. In the illustrated embodiment, lens 112 is an ultra-wide-angle lens that (used in conjunction with sensor 110) provides a field of view θ between, for example, 100 degrees and 270 degrees. In one embodiment, the field of view θ is greater than 180 degrees and less than 270 degrees. It should be understood that although the embodiments herein describe lens 112 as a single lens, this is for illustrative purposes only, and in other embodiments, lens 112 comprises multiple optical elements.
[0029] In some embodiments, sensor 104 includes a pair of sensors 110A and 110B, which are arranged to receive light from ultra-wide-angle lenses 112A and 112B, respectively. Figure 1C In this disclosure, sensor 104 is sometimes referred to as a dual sensor because it has a pair of sensors 110A and 110B and corresponding lenses 112A and 112B, as shown. Sensors 110A and lens 112A are arranged to acquire images in a first direction, and sensors 110B and 110B are arranged to acquire images in a second direction. In the illustrated embodiment, the second direction is opposite to the first direction (e.g., separated by 180 degrees). A sensor having opposingly arranged sensors and lenses with a field of view of at least 180 degrees is sometimes referred to as an omnidirectional sensor, a 360-degree sensor, or a panoramic sensor because it can acquire images within a 360-degree range on at least one axis of the sensor.
[0030] Sensor 104, as further described herein, may also be referred to as a “rig” or simply a “suite”. In various embodiments, sensor 104 (e.g., the rig) moves through the environment to capture multiple images of the surrounding environment. These multiple images constitute an “image sequence” or “image sequence set,” which are referred to in various ways herein. Processing system 102 or other suitable processing system or device performs photogrammetric methods on the captured images constituting the image sequence to generate a three-dimensional representation or model of the image-captured environment. In some embodiments, it is necessary to estimate the trajectory of sensor 104 as it moves within or relative to the environment in order to determine the three-dimensional representation or model. Various trajectory estimation methods will be further described herein.
[0031] In various embodiments, more than two such images are captured at each point along the trajectory, and the sequence of images captured along the entire trajectory includes hundreds, thousands, or tens of thousands of images from one or two sensors of sensor 104.
[0032] Looking at it now Figure 2 The figure schematically illustrates a schematic diagram 200 of one embodiment of a processing system 102 for trajectory estimation using image sequences. Other useful embodiments are also readily understood. The processing system 102 can be any suitable computing device, such as a laptop, desktop computer, smartphone, tablet, etc., including combinations and / or multiple thereof. According to at least one embodiment described below, Figure 5 Processing system 500 is shown, which is another example of processing system 102. Return to... Figure 2The processing system 102 includes a processing device 202, a system memory 204, a network adapter 206, a data memory 208, a display 210, a sensor 211, an acquisition (or capture) engine 212, and an alignment and trajectory engine 214.
[0033] In various embodiments, the components, modules, engines, etc., described in FIG2 (e.g., capture engine 212 and alignment and trajectory engine 214) can be implemented as computer processing instructions stored on a computer-readable storage medium, hardware modules, special-purpose hardware (e.g., special-purpose hardware, application-specific integrated circuits (ASICs), special-purpose processors (ASSPs), field-programmable gate arrays (FPGAs), embedded controllers, hardwired circuits, etc.), or combinations of several of the above implementations. According to certain aspects of this disclosure, the engine described herein is a combination of hardware and a program. The program is embodied in the form of processor-executable instructions stored in tangible memory, and the hardware may include a processing device 202 for executing these instructions. Thus, system memory 204 can store program instructions that, when executed by processing device 202, implement the engine described herein. Other engines may also be used to implement other features and functions described in other examples herein. Network adapter 206 enables processing system 102 to send and / or receive data to other sources (e.g., sensor 104).
[0034] In various embodiments, sensor 104 is any device suitable for collecting (acquiring) images of object or environment 222. For example, processing system 102 receives data (e.g., one or more images of environment 222) directly from sensor 104 and / or via wired or wireless network 207. Data from sensor 104 (e.g., images) is stored as data 209a (also referred to as “image 209a”) in data memory 208 of processing system 102. According to at least one embodiment described herein, acquisition engine 212 causes sensor 104 to capture data 209a (e.g., images), for example by sending an activation signal to sensor 104. According to at least one embodiment described herein, sensor 104 automatically and / or in response to some other triggering event (e.g., user command) acquires data 209a and transmits data 209a to processing system 102 for processing and / or storage.
[0035] According to the embodiments described herein, the acquisition engine 212 uses sensor 211 to acquire additional data and / or images of the environment 222 as data 209b. According to the embodiments described herein, the processing system 102 generates a point cloud representation from the data 209a and / or data 209b and displays it on the display 210. Sensor 211 includes at least one sensor for acquiring image data and information about the environment 222. Sensor 211 may include at least one of the following: a sensor, an omnidirectional sensor, a panoramic sensor, a 360-degree sensor, a LiDAR sensor, etc., including combinations and / or multiple combinations thereof.
[0036] Network 207 represents various suitable communication networks (including combinations thereof), such as wired networks, public networks (e.g., the Internet), private networks, wireless networks, cellular networks, and any other suitable private and / or public networks. Furthermore, Network 207 can have any suitable communication range, such as a global network (e.g., the Internet), a metropolitan area network (MAN), a wide area network (WAN), a local area network (LAN), or a personal area network (PAN). Additionally, Network 207 can include any medium used for transmitting network traffic, including but not limited to coaxial cable, twisted pair, optical fiber, hybrid fiber-coaxial (HFC) media, microwave terrestrial transceivers, radio frequency communication media, satellite communication media, or any combination thereof.
[0037] During image acquisition, sensor 104 is positioned on, inside, and / or around environment 222 to acquire data / images of environment 222. For example, sensor 104 acquires a series of environmental images as it moves along a path.
[0038] According to the embodiments described herein, processing system 102 performs various photogrammetric functions using images received from sensor 104 (e.g., data 209a) and / or images captured by sensor 211 (e.g., data 209b). Photogrammetry is a technique for modeling objects using images (e.g., photographic images acquired by digital sensors). Photogrammetry can generate three-dimensional (3D) models from two-dimensional images or photographs. When multiple images are acquired at different locations with overlapping fields of view, common points or features (sometimes referred to herein as “connecting points”) are identified in each image. The three-dimensional coordinates of the feature / connecting point can be determined using triangulation or triangulation methods by projecting rays from the sensor location to the feature / connecting point on the object. In some examples, photogrammetry is based on markers / targets placed on objects in the environment (e.g., lights or reflective stickers) or on unique natural features of the object. To perform photogrammetry, in various cases, it is necessary to acquire images using sensors with one or more sensors (e.g., photosensitive arrays). By acquiring multiple images of an object or a portion thereof from different locations or orientations, the three-dimensional coordinates of points on the object can be determined based on common features or points and the sensor's position and orientation information (e.g., attitude or pitch angle, yaw angle, or roll angle) at each image acquisition. To reliably acquire the necessary information for determining the three-dimensional coordinates, these features need to be identified in two or more images. Because the images are acquired from different locations or orientations, common features lie in overlapping areas of the image field of view. It should be noted that a photogrammetric technique is described in U.S. Patent No. 10,477,180 to Wolke et al. (titled "PHOTOGRAMMETRY SYSTEM AND METHOD OF OPERATION"), the contents of which are incorporated herein by reference. Photogrammetric techniques determine the three-dimensional (3D) coordinates of ground features by capturing two or more images.
[0039] Different configurations can be employed, such as linear, circular, back-to-back, and combinations and / or stacking (multiplication) of various configurations. The specific arrangement of the sensors is designed to provide the required overlap and coverage of the object or environment being photographed. In different situations, depending on the specific requirements of the photogrammetric task being performed, the focal length and / or resolution of the sensors can be the same or different.
[0040] Each sensor is calibrated to determine its inherent parameters. Examples of inherent parameters include focal length, principal point, lens distortion coefficient, tilt, aspect ratio, and combinations and / or superpositions (multiples) of these. Sensor calibration employs an on-the-fly calibration method, requiring no additional data acquisition before or after image acquisition. As described herein, common points generated in images involving the same location on the environment or object through feature extraction and matching can be used in this process. In various embodiments, each sensor is calibrated individually using a selected portion of the image set used for sensor calibration. This paper considers some common features of bundle adjustment and root mean square error (RMSE) for image candidate objects. To obtain a sufficient number of common points suitable for sensor calibration and to establish a sufficiently large baseline between sensor poses, the sensors can be aligned perpendicular to the direction of motion.
[0041] Once the inherent parameters of each sensor are calculated, the relative orientation between the sensors can be determined. This requires calculating the translation and rotation between the sensor coordinate systems. Relative orientation calibration can be performed using a calibration object with known 3D points. This object is placed in the environment (e.g., environment 222), and its 3D points are captured in images taken by each sensor. By matching the 3D points in the images, the relative pose between the sensors can be calculated.
[0042] According to the embodiments described herein, another relative orientation calibration method is based on real-time system calibration. This method utilizes common image points in each image sequence and can be applied to applications such as trajectory estimation that do not require scale information.
[0043] According to the embodiments described herein, in certain configuration settings, existing information can be used as constraints to improve the reliability of system calibration. Existing information can be, for example, the back-to-back placement of sensors in a sensor system (when the distance between the two sensor locations is unknown), such as the sensor system shown and described herein (see...). Figure 1A-1C ), and / or similar systems, including combinations and / or superpositions (multiplications). The distance between the two sensors is unknown, for example, because their precise distance cannot be determined beforehand. As described herein, system calibration is used to determine this unknown distance, as well as certain other unknown parameters.
[0044] Referring again to FIG2, according to the embodiments described herein, the processing system 102 may use the alignment and trajectory engine 214 to perform trajectory estimation over the image sequence using images received from sensor 104 (e.g., data 209a or 209b) and / or images captured by sensor 211.
[0045] According to the embodiments described herein, the alignment and trajectory engine 214 performs image alignment for trajectory estimation. Image alignment is the process of determining the position and angular orientation of the sensor relative to a reference coordinate system. In various embodiments, image alignment is a prerequisite for achieving the required accuracy in three-dimensional scene reconstruction.
[0046] Alignment and trajectory engine 214 can perform at least one image alignment technique in photogrammetry, such as feature-based matching, direct image alignment, sequential alignment, bundle adjustment, global positioning system (GPS) and control points, and / or similar techniques, including combinations and / or multiples thereof (multiple stacking).
[0047] Feature-based matching refers to identifying and matching salient features in different images. These features can be key points, corners, or other easily detectable and describable salient points (i.e., connection points). Once feature matching between images is complete, the alignment and trajectory engine 214 calculates transformations (e.g., translation, rotation) to align the images.
[0048] Sequence image alignment refers to aligning images sequentially, where an already oriented image serves as a reference image, and subsequent images are aligned with that reference image.
[0049] Bundle adjustment is an optimization technique used to simultaneously optimize sensor parameters and 3D coordinates of points in a scene. In various implementations of this method, the alignment and trajectory engine 214 considers multiple images and features simultaneously, thereby achieving consistent and accurate alignment by reducing or minimizing reprojection errors (i.e., the difference between observed image points and projected 3D points).
[0050] If available, GPS and control points (depending on their accuracy) can provide initial image alignment. For example, the alignment and trajectory engine 214 uses GPS coordinates and ground control points (GCPs) to perform initial alignment after completing the geoaligning of the images and the entire photogrammetric network.
[0051] It should be understood that image alignment can be a computationally intensive process, especially when dealing with large numbers of images and complex scenes. Therefore, in some embodiments, specialized software and / or libraries are used to implement efficient image alignment algorithms.
[0052] Due to the inherent characteristics of data acquired by handheld or mobile scanners, it is noisier (e.g., may contain erroneous data) and less accurate compared to data acquired by fixed scanners. However, when scanning large environments, mobile data acquisition methods are much faster than using a series of fixed scanners. The process and apparatus described herein combine the speed and efficiency of mobile data acquisition methods while maintaining accuracy comparable to data from fixed scanners such as arm-mounted devices, thus enabling accurate data acquisition from large and complex environments in a short time. In addition to the various types of useful scanners and methods of use described herein, PCT Publication No. WO2020 / 234575 (titled "SENSOR SYNCHRONIZATION") and U.S. Patent No. 11,176,353 (titled "THREE-DIMENSIONAL DATASET AND TWO-DIMENSIONAL IMAGE LOCALIZATION") disclose various types of useful scanners and methods of use, the entire contents of which are incorporated herein by reference.
[0053] This document uses several terms, which are defined in the following paragraphs:
[0054] Spatial flash: A form of flash based on the geometry of point clouds rather than photogrammetry. Aspects of space flash are described in the following jointly owned U.S. patent applications: U.S. Patent Application No. 18 / 449943 entitled “CALIBRATING SYSTEM FOR COLORIZING POINT CLOUDS”; U.S. Patent Application No. 17 / 451946 entitled “THREE DIMENSIONAL MEASUREMENT DEVICE HAVING A CAMERA WITH A FISHEYE LENS”; U.S. Patent Application No. 17 / 492801 entitled “DYNAMIC SELF-CALIBRATING OF AUXILIARY CAMERA OF LASER SCANNER”; U.S. Patent Application No. 17 / 379268 entitled “Laser Scanner with Ultra Wide-Angle Lens Camera for Registration”; and U.S. Patent Application No. 17 / 678116 entitled “CALIBRATING SYSTEM FOR COLORIZING POINT-CLOUDS”, the entire contents of which are incorporated herein by reference.
[0055] Mobile data: Spatial data acquired using a handheld LiDAR scanner designed for mobile use in the scanning environment and processed via SLAM.
[0056] Static data: Spatial data collected from a single stationary location, which can be acquired by a fixed scanner or extracted from a single stationary time period in moving data.
[0057] SLAM: Simultaneous Localization and Mapping Algorithm, which uses trajectory and distance data collected by a mobile LiDAR scanner to gradually build a point cloud.
[0058] Point cloud: A collection of data points collected by a scanner.
[0059] Point: A single spatial data derived from a single ray emitted by a LiDAR scanner or similar device. It mainly includes three-dimensional coordinate data (i.e., X, Y, Z axes), as well as the reflectivity (intensity), color (RGB values of red, blue, and green) of the surface it represents, and the scanner's capture time.
[0060] Space: Operations are performed in physical three-dimensional space.
[0061] Two-pass method: A method for resampling and interpolating on a two-dimensional grid to fill blank cells. This method involves two stages of resampling and two stages of interpolation, each time filling a different spatial region of the grid.
[0062] Point distance: The distance from a point to the scanner's origin. This is the distance from the laser beam to the physical, real-world object with which it interacts.
[0063] Structured point cloud: A type of point cloud in which each Cartesian coordinate can be assigned row and column values in a grid format.
[0064] Figure 3 illustrates an example of a moving window 300, where grid lines 302 represent a 24-megapixel (MP) raster (or grating), and black boxes 304 represent windows twice the size of the raster cells, according to the embodiments described herein. One embodiment of the process disclosed herein (sometimes referred to herein as spatial flash) utilizes point data collected from a single fixed location during mobile data acquisition to create structured and precise data. For this purpose, three-dimensional (3D) points collected from the single fixed location are projected onto an equidistant cylindrical projection two-dimensional raster with a resolution of 24 MP. The boundaries of each cell in this raster represent the azimuth and elevation angles (elevation), and the values within each cell represent the distance to the scanner (point pitch).
[0065] At the end of the algorithm, each cell will contain a point distance. Using the azimuth and elevation angle boundaries of each cell, the central angle can be derived, and the Cartesian coordinates can be reconstructed by combining the point distances. This will generate a point cloud containing approximately 24 million points, uniformly distributed in a row / column structure, and averaged and smoothed to reduce outlier noise.
[0066] This paper develops and employs a two-pass re-sampling and interpolation method. An example of implementing this method is shown below:
[0067] Operation 1: Resample using a moving window that is twice the size of the cell.
[0068] Operation 2: Resample using a moving window that is twice the size of the cell and offset by one cell in both the horizontal and vertical directions.
[0069] Operation 3: Use the combined data from Operations 1 and 2 to perform interpolation (or interpolation) using a moving window that is twice the size of the cells and offset by only one cell horizontally.
[0070] Operation 4: Use the combined data from Operations 1 and 2 to perform interpolation (or interpolation) using a moving window that is twice the cell size and offset by only one cell in the vertical direction.
[0071] Operation 5: (Optional) Perform edge filtering on the data from Operation 1 using a moving window that is four times the size of the cell.
[0072] Operation 6: (Optional) Perform edge filtering on the data from Operation 2 using a moving window that is four times the size of the cell.
[0073] The results of operations 3, 4, 5, and 6 (or operations 1, 2, 3, and 4 without edge filtering) are combined to generate a single structured point cloud.
[0074] The aforementioned moving window (see Figure 3) is, in various embodiments, a virtual frame with a spatial size twice the size of a grid cell, centered on one of the cells. The window uses distance data from all points within its boundaries to calculate a (single) point distance value and assigns that value to the distance value of the cell where the window's center is located.
[0075] Figure 4 is a bitmap 400 obtained after completing resampling operations 1 and 2 as described herein, according to the embodiment. Resampling extracts a static point cloud from a static period of approximately 15 seconds in the moving point cloud. This data is then projected onto a 24MP raster. During this stage, each cell contains multiple point pitches, and a single point pitch value is exported for each cell using a moving window. In some cases, if the depth range between point pitches within a cell is large (i.e., greater than -0.05 meters), it is assumed that a geometric region resembling a corner has been found, and therefore the minimum point pitch (the point closest to the scanner) is selected to represent that cell. Otherwise, in some cases, the median point pitch is used.
[0076] In various embodiments, the spatial regions of these filled cells are then used as templates to retrieve all data located within these spatial regions and adjacent spatial regions (e.g., within 2 cells) from the complete moving point cloud, containing points whose distances are close to the values of the currently exported individual cells. The median distance of each cell is then recalculated, including data from outside the approximately 15-second static time period, which helps to smooth and fill gaps.
[0077] In various embodiments, operations 1 and 2 described above will follow the same resampling process, the only difference between them being the offset nature of the moving window (see, for example, Figures 4 and 5).
[0078] Figure 5 depicts the moving window 500 in the following operations 1 and 2, where the lighter (gray) grid lines 502 represent a 24MP grid, the darker (black) grid lines 504 represent the moving window of operation 1, the dashed grid lines 506 represent the moving window of operation 2, and the cells are the cells in the grid that will be filled by the corresponding operations according to the embodiments described herein.
[0079] Figure 6 is an example in which the moving window 600 is centered on an empty cell that is being filled with interpolation point distances and shows the range of point distances used to calculate the interpolation point distances according to the embodiments described herein.
[0080] Interpolation (or interpolation) uses the singularpoint distance values filled in each cell in operations 1 and 2 above, instead of the actual point data. The interpolated point distances are created by centering a moving window over each cell that needs interpolation. In various embodiments, the moving window is created over the combined data generated by operations 1 and 2. A new point distance is created only if the number of points within the moving window is greater than or equal to (>=) 3 and the range difference between point distances is less than or equal to (<=) 0.05 meters (m). In various embodiments, the new point distance is the average range value of all point distances within the window (see Figure 6).
[0081] In some cases, edge filtering uses the singularity distance values filled into each cell in Operations 1 and 2, instead of the actual point data. Edge filtering is an optional operation designed to further remove outliers, such as those scattered along the ray direction when a scanner beam sweeps across the edge or surface of an object at an angle. In various embodiments, edge filtering uses a larger moving window (4x4 grid cells) that covers each cell filled by Operations 1 (Operation 5 above) and Operation 2 (Operation 6 above). If the point distance collected within the window is greater than -0.05 meters, the data for that cell is purged; thus, in this case, a point is effectively removed from the final output.
[0082] Figure 7 illustrates the process (procedure) of interpolating 700 using a moving window only in the horizontal direction, as shown in Operation 3 above and described according to one or more embodiments described herein. Note that each point spaced apart in the longer rows shown in the figure is a reference point for various cases, including the highlighted point 702 and other points 704.
[0083] Figure 8 illustrates the process (procedure) of interpolating 800 using a moving window only in the vertical direction, as shown in Operation 4 above and described according to the embodiments described herein. Note that each point spaced apart in the alternating rows shown in the figure is a reference point for various cases, including the highlighted point 802 and other points 804.
[0084] Figure 9 illustrates a mobile scanning device 900 (sometimes referred to herein as ORBIS, manufactured by FARO Technologies, Lake Mary, Florida, USA) that utilizes the methods (processes) described herein, according to various embodiments. This handheld scanner can quickly and conveniently acquire 3D point cloud data by walking within a region of interest. In various embodiments, the scanner includes a 2D time-of-flight (TOF) laser rangefinder rigidly connected to an inertial measurement unit (IMU) mounted on a motor drive. The motion of the scanning head on the motor drive provides a third dimension for generating 3D information. A 3D simultaneous localization and mapping (SLAM) algorithm (i.e., FARO's 3D SLAM) combines the 2D laser scan data with the IMU data to generate an accurate 3D point cloud.
[0085] Mobile scanners can acquire raw laser ranging and inertial data in various embodiments. This data must be processed using 3D SLAM algorithms to convert the raw data into a 3D point cloud. Data processing can be performed using tools such as the FARO HUB processing application.
[0086] FARO STREAM is a utility mobile application developed by FARO Technologies Inc., based in Lake Mary, Florida, USA, that connects hardware with cloud applications and services. By combining hardware with cloud software, STREAM improves the efficiency of field acquisition workflows and imports acquired data directly into the FARO ecosystem, while providing real-time feedback and scan result display (see, for example, Figure 11). One or more aspects of STREAM have been described in U.S. Patent No. 11,340,058 entitled “REAL-TIME SCAN POINT HOMOGENIZATION FOR TERRESTRIAL LASERSCANNER,” the contents of which are incorporated herein by reference.
[0087] STREAM provides efficient on-site data acquisition for scanning operations in the fields of architecture, engineering, construction, and facilities management. Users can rest assured that they can acquire complete and successful scan data in real time without the need for additional on-site visits due to missing data, thereby accelerating project completion.
[0088] FARO Technologies' cloud applications, such as FARO SPHERE, provide a centralized, efficient, and collaborative environment encompassing point cloud applications and customer support tools to accelerate the acquisition, processing, and delivery of 3D data, all through secure single sign-on processes in various embodiments. With STREAM and SPHERE, remote colleagues can begin processing data or share data with end customers via WEBSHARE software or similar tools, including FARO's collaborative point cloud project management solutions. One or more aspects of SPHERE have been described in a jointly owned U.S. Patent Publication No. 2023-0119214, entitled "FOUR-DIMENSIONALDATA PLATFORM USING AUTOMATIC REGISTRATION FOR DIFFERENT DATA SOURCES," the contents of which are incorporated herein by reference.
[0089] To download the raw scan data, power on the scanner's data logger. Connect a Universal Serial Bus (USB) memory device or similar device to the USB port on the front panel of the data logger. While data is being transferred to the flash drive, the AUX indicator or similar indicator light will illuminate, for example, green. Do not remove the USB flash drive while the AUX indicator light is on. After a few seconds (the exact time depends on the size of the data file being transferred), the AUX indicator light will turn off. All remaining data will have been transferred, at which point the USB flash drive can be removed.
[0090] In various embodiments, the 3D SLAM algorithm used to process raw laser scan data into a 3D point cloud relies on the presence of features in the scanned environment that are repeatedly scanned as the operator moves through it. For a given feature, its size-to-distance ratio is approximately 1:10 (e.g., at a distance of 5 meters, a feature must be larger than 0.5 meters to be considered significant). Environments with "few features" include open spaces and passageways with smooth walls. In corridors with smooth walls, the number of features identifiable along the direction of travel is typically insufficient, making it impossible for the 3D SLAM algorithm to determine the direction of travel.
[0091] In many cases, 3D SLAM algorithms employ a method similar to traverse surveying in measurement practice, processing the raw scan data into a point cloud. This technique utilizes the scanner's previously known position to determine the current location. This method can amplify small measurement errors, causing the calculated position to "drift." Good measurement practice involves "closing the loop" by rescanning the known position, thus distributing the accumulated error throughout the loop.
[0092] At a minimum, measurements should begin and end at the same location to ensure at least one loop closure is completed. It is recommended to complete loop closures as frequently as possible to reduce or minimize errors and improve the accuracy of the generated point cloud.
[0093] Generally, it's best to perform a circular scan rather than simply repeating the process. This applies to both horizontal and vertical loops, so you should enter and exit through different doors and move between floors via different staircases.
[0094] In some cases, it is necessary to carefully scan closed-loop areas to ensure that key features are scanned from similar perspectives. In one embodiment, a user can turn around and return to an area from a different direction. This is especially important in environments with fewer features, such as corridors and passageways.
[0095] In some embodiments, to avoid scanner operators including data within a small range of values in the final point cloud, thus saving bandwidth and data storage space, data within that range is not processed by default. In various cases, proximity to walls and ceilings should be avoided.
[0096] In various embodiments, the scanner's maximum detection range is approximately (~) 100 meters. This distance is achievable under optimal conditions (indoors, with good target reflectivity). In most cases, the typical maximum detection range is 60-80 meters. It is recommended to keep the detection range within 50 meters whenever possible to ensure good point cloud density and to support the 3D SLAM algorithm.
[0097] In most cases, 3D SLAM algorithms can handle moving objects in the environment. To estimate sensor trajectories, the algorithm assumes that most of the environment is static. However, in some feature-sparse environments, moving objects can have a greater impact on the results due to the lack of three-dimensional structure in certain dimensions. The influence of moving objects should be avoided, especially in long tunnel-like environments (such as corridors), relatively open spaces, and when passing through doorways.
[0098] As shown in Figure 10, a scanning kit 1000 is provided in various embodiments. In this embodiment, the scanning kit 1000 includes a handheld scanner 1001, which comprises the following components (as shown): a data logger 1002, a battery 1003, a data cable, a battery charger and power supply 1005, a reference base plate 1006, a USB or other type of flash drive 1007, a backpack 1008 or similar transport case, and a shoulder strap 1009 for the data logger for mobility. The scanning kit 1000 provides a convenient way to carry the scanner 1001 and necessary accessories to perform scanning projects.
[0099] Figure 11 shows an output 1100 generated by a mobile scanner according to one or more embodiments described herein. To view the processed dataset, a 3D view of the dataset will open in a new view window when the "View" button 1102 or a similar button is selected next to the filename in the dataset list; one example is shown in Figure 11. In some cases, the displayed output will be presented in color (not shown) instead of in black and white as shown in the figure.
[0100] It should be understood that one or more embodiments described herein are embodied in the form of a system, method, or computer program product, and take the form of hardware embodiments, software embodiments (including firmware, resident software, microcode, etc.), or combinations thereof. Furthermore, one or more embodiments described herein are embodied in the form of a computer program product, which is embodied in one or more computer-readable media having computer-readable program code thereon.
[0101] The term “about” is intended to cover the range of error that may exist when measuring a specific quantity based on the equipment available at the time of application submission. For example, “about” could include a range of ±8%, 5%, or 2% for a given value.
[0102] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. Unless the context clearly requires otherwise, the singular forms “a,” “an,” and “the” used herein should also include the plural forms. Furthermore, it should be understood that the terms “comprising” and / or “including” as used herein mean the presence of the stated feature, integer, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integers, operations, elements, and / or components thereof.
[0103] While this disclosure has been provided in detail with reference to only a limited number of embodiments, it should be readily understood that this disclosure is not limited to these disclosed embodiments. Rather, this disclosure may be modified to incorporate any number of variations, alterations, substitutions, or equivalent arrangements not previously described but commensurate with the spirit and scope of this disclosure. Furthermore, although various embodiments of this disclosure have been described, it should be understood that the embodiments include only some of the aspects described.
[0104] Therefore, this disclosure should not be considered as being limited by the foregoing description.
Claims
1. A system for estimating the trajectory of movement through an environment, comprising: Sensors that acquire multiple images while traversing the environment; and At least one processor performs at least one photogrammetric function to generate a data sequence from multiple images received from a sensor, and performs two rounds of sampling and interpolation on the data sequence to generate a structured point cloud.
2. The system of claim 1, wherein the two-pass sampling and interpolation comprises: (i) The data sequence is first resampled using a moving window twice the size of the sensor grid cell; and (ii) The data sequence is resampled a second time using a moving window that is twice the size of the sensor grid cells and offset by one cell in both the horizontal and vertical directions.
3. The system of claim 2, wherein the two-pass sampling and interpolation further comprises performing a first interpolation using combined data from the first resampled data and the second resampled data, employing a moving window that is substantially twice the size of the cell and offset by only one cell in the horizontal direction.
4. The system of claim 3, wherein the two-pass sampling and interpolation further comprises a second interpolation using a combination of data from the first resampled data and the second resampled data, employing a moving window that is substantially twice the size of the cell and offset by only one cell in the vertical direction.
5. The system according to claim 4, wherein, Edge filtering is applied to the data from the first resampled sample.
6. The system according to claim 4, wherein, Edge filtering is applied to the data from the second resampled sample.
7. The system according to claim 4, wherein, The data from the first resampling, the second resampling, the first interpolation, and the second interpolation are combined to form a single structured point cloud.
8. The system according to claim 2, wherein, The at least one processor extracts the static point cloud during the stationary period when traversing the environment from the data generated by the first resampling and the second resampling.
9. The system according to claim 8, wherein, The static point cloud is projected onto a megapixel grid.
10. The system of claim 1, wherein the sensor is a camera.
11. The system of claim 1, wherein the sensor is a laser scanner.
12. The system according to claim 1, wherein, The sensor comprises two sensors arranged opposite each other, each sensor having a field of view of at least 180 degrees.
13. A method for estimating the trajectory of a scanning system moving through an environment, the method comprising: The scanning system uses sensors to acquire multiple images as the sensors move through the environment; Perform at least one photogrammetric function on at least one of the plurality of images to generate a data sequence; and Perform two rounds of sampling and interpolation on the data sequence to generate a single structured point cloud.
14. The method according to claim 13, wherein, Performing the two rounds of sampling and interpolation also includes: (i) The data sequence is first resampled using a first moving window that is essentially twice the size of a sensor grid cell; and (ii) The data sequence is resampled a second time using a second moving window that is twice the size of the cell and offset by one cell in both the horizontal and vertical directions.
15. The method according to claim 14, wherein, Performing the two rounds of sampling and interpolation also includes: The first interpolation is performed using a combination of data from the first and second resampling, which employs a third moving window that is twice the size of the cell and offset by only one cell in the horizontal direction.
16. The method according to claim 15, wherein, Performing the two rounds of sampling and interpolation also includes: The combined data is used for a second interpolation, which employs a fourth moving window that is twice the size of the cell and offset by only one cell in the vertical direction.
17. The method of claim 14, wherein, Performing the two rounds of sampling and interpolation also includes edge filtering of the data sequence from the first resample.
18. The method according to claim 14, wherein, Performing the two rounds of sampling and interpolation also includes edge filtering of the data sequence from the second resample.
19. The method of claim 16, further comprising: The data sequences obtained from the first resampling, the second resampling, the first interpolation, and the second interpolation are combined to form a single structured point cloud.
20. The method of claim 14, wherein, The processor extracts the static point cloud during the static period from the data sequence generated by the first and second resampling.
Citation Information
Patent Citations
Photogrammetry system and method of operation
US10477180B1
Three-dimensional dataset and two-dimensional image localization
US11176353B2
Real-time scan point homogenization for terrestrial laser scanner
US11340058B2
System and method for storing a database on flash memory or other degradable storage
US12141457B1
Three dimensional measurement device having a camera with a fisheye lens
US12535590B2