Method and apparatus for photogrammetry feature extraction and matching
The method improves photogrammetry by generating a database of 2D features and associated 3D coordinates, reducing computational costs and enhancing accuracy in determining 3D coordinates from 2D images, suitable for applications in construction sites and unmanned aerial vehicles.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2026-03-12
AI Technical Summary
Current photogrammetry methods are computationally expensive for extracting features and determining corresponding 3D coordinates from 2D images.
A method involving capturing data about an environment, generating a database of 2D features and associated 3D coordinates, extracting 2D features from images, identifying matches, and filtering out identical features to determine 3D coordinates using techniques like photogrammetry and laser scanning.
This approach reduces computational complexity and enhances the efficiency of feature extraction and 3D coordinate determination, enabling faster and more accurate processing of images, particularly in environments like construction sites and using unmanned aerial vehicles.
Smart Images

Figure 00000027_0000 
Figure 00000028_0000 
Figure 00000029_0000
Abstract
Description
Attorney Docket No. P1825-US-WO (138178-80220)METHOD AND APPARATUS FOR PHOTOGRAMMETRY FEATURE EXTRACTION AND MATCHINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application Number 63 / 690,977, filed September 5, 2024, the entire contents of which are incorporated by reference.BACKGROUND
[0002] The subject matter disclosed herein relates to photogrammetry feature extraction and methods for expediting the process of determining three-dimensional (3D) coordinates associated with features extracted from two-dimensional (2D) images captured by a camera. Current methods of photogrammetry oftentimes rely on computationally expensive methods of extracting features and determining corresponding 3D coordinates.BRIEF DESCRIPTION
[0003] In one exemplary embodiment, a method for determining three- dimensional (3D) coordinates in a scene based on two-dimensional (2D) features is provided. The method comprises capturing data about an environment. A database of 2D features and associated 3D coordinates are generated based at least in part on the data about the environment. 2D features are extracted from a 2D image. A match is identified between the 2D features of the 2D image and the 2D features in the database. At least one feature is filtered out from the 2D image, which is also associated with an image that is identical to the 2D image. 3D coordinates associated with a scene in the 2D image are determined based at least in part on the at least one feature filtered out from the 2D image.
[0004] Additional technical features and benefits are realized through the techniques of the present invention. Embodiments and aspects of the invention areMEl\57042703.vl 1Attorney Docket No. P1825-US-WO (138178-80220) described in detail herein and are considered a part of the claimed subject matter. For a better understanding, refer to the detailed description and to the drawings.BRIEF DESCRIPTION OF DRAWINGS
[0005] The subject matter, which is regarded as the disclosure, is particularly pointed out and distinctly claimed in the claims at the conclusion of the specification. The foregoing and other features, and advantages of the disclosure are apparent from the following detailed description taken in conjunction with the accompanying drawings in which:
[0006] FIG. 1 depicts a block diagram of a processing system for performing initialization and tracking according to embodiments described herein;
[0007] FIG. 2 depicts a flow diagram of a method for performing initialization according to embodiments described herein;
[0008] FIG. 3 depicts a flow diagram of a method for performing photogrammetry according to embodiments described herein;
[0009] FIG. 4 depicts a grid structure for feature extraction according to embodiments described herein;
[0010] FIG. 5 depicts a spherical grid structure layout for feature extraction for wide-angle (fish-eye) images according to embodiments described herein;
[0011] FIG. 6 depicts feature clustering, in which each section denotes a cluster of descriptors in n-dimensional space according to embodiments described herein;
[0012] FIG. 7 depicts a feature cross check validation diagram between two images, according to embodiments described herein; and
[0013] FIG. 8 depicts a k-dimensional (kD) tree illustrating image pairing, where each leaf corresponds to a single image.MEl\57042703.vl 2Attorney Docket No. P1825-US-WO (138178-80220)
[0014] The detailed description explains embodiments of the disclosure, together with advantages and features, by way of example with reference to the drawings.DETAILED DESCRIPTION
[0015] FIG. 1 depicts a block diagram of a processing system 100 for performing initialization and tracking according to embodiments described herein. The processing system 100 includes a processing device 102, a memory 104, a network adapter 106, a data store 108 for storing data 109a, 109b, and an initialization engine 110 configured and arranged as shown.
[0016] The various components, modules, engines, etc. described regarding the processing system 100 are, in various embodiments, implemented as instructions stored on a computer-readable storage medium, as hardware modules, as special-purpose hardware (e.g., application specific hardware, application specific integrated circuits (ASICs), application specific special processors (ASSPs), field programmable gate arrays (FPGAs), as embedded controllers, hardwired circuitry, etc.). In an embodiment, the components are a combination of the foregoing. According to aspects of the present disclosure, the engine(s) described herein are a combination of hardware and programming. The programming are processor executable instructions stored on a tangible memory, and the hardware includes the processing device 102 for executing those instructions. Thus, a system memory (e.g., the memory 104) stores program instructions that when executed by the processing device 102 implement the engines described herein. Other engines are also utilized to include other features and functionality described in other examples herein.
[0017] In an embodiment, the network adapter 106 provides for the processing system 100 to transmit data to and receive data from other sources, such as other processing systems, data repositories, and the like. As an example, the processing system 100 transmits data to and receives data from a camera 120, a scanner 130, and a user device 140 directly and via a network 150.MEl\57042703.vl 3Attorney Docket No. P1825-US-WO (138178-80220)
[0018] The network 150 represents any one of different types of suitable communications networks such as, for example, wired networks, public networks (e.g., the Internet), private networks, wireless networks, cellular networks, or any other suitable private and / or public networks. In an embodiment, the network 150 is a combination of the foregoing communications networks. Further, the network 150 has any suitable communication range associated therewith and includes, for example, global networks (e.g., the Internet), metropolitan area networks (MANs), wide area networks (WANs), local area networks (LANs), and personal area networks (PANs). In addition, the network 150 includes any type of medium over which network traffic is carried including, but not limited to, coaxial cable, twisted-pair wire, optical fiber, a hybrid fiber coaxial (HFC) medium, microwave terrestrial transceivers, radio frequency communication mediums, satellite communication mediums, and any combination thereof without limitation.
[0019] The camera 120 is a 2D camera or a 3D camera (red-green -blue-depth (RGBD) or time-of-fhght for example). The camera 120 captures an image (or multiple images), such as of an environment 160. The camera 120 transmits the image to the processing system 100. In some embodiments, the camera 120 encrypts the image before transmitting it to the processing system 100. Although not shown, the camera 120 includes components such as a processing device, a memory, a network adapter, and the like, which is functionally similar to those included in the processing system 100 as described herein.
[0020] In some examples, the camera 120 is portable or movable and may be moved about the environment 160. In some examples, the camera 120 is disposed coupled to an unmanned aerial vehicle (i.e. a drone). In some examples, the camera 120 is mounted to a fixture, which is user-configurable to rotate about a roll axis, a pan axis, and a tilt axis. In such examples, the camera 120 is mounted to the fixture to rotate about the roll axis, the pan axis, and the tilt axis. Other configurations of mounting options for the camera 120 also are possible.
[0021] A coordinate measurement device, such as scanner 130 for example, is any suitable device for measuring 3D coordinates or points in an environment, such asMEl\57042703.vl 4Attorney Docket No. P1825-US-WO (138178-80220) the environment 160, to generate data about the environment. A collection of 3D coordinate points within the environment 160 is sometimes referred to as a point cloud. According to some embodiments described herein, the scanner 130 is a three- dimensional (3D) laser scanner time-of-fhght (TOF) coordinate measurement device. It should be appreciated that while embodiments herein refer to a laser scanner, this is for example purposes and the claims should not be so limited. In other embodiments, other types of coordinate measurement devices or combinations of coordinate measurement devices are used, such as but not limited to triangulation scanners, structured light scanners, laser line probes, photogrammetry devices, and the like. A 3D TOF laser scanner steers a beam of light to a non-cooperative target such as a diffusely scattering surface of an object. A distance meter in the scanner 130 measures a distance to the object, and angular encoders measure the angles of rotation of two axles in the device. The measured distance and two angles enable a processor in the scanner 130 to determine the 3D coordinates of the target. In some embodiments, the camera 120 is optionally coupled to or otherwise integrated in operable communication with a scanner 130.
[0022] According to some embodiments described herein, the camera 120 captures 2D image(s) of the environment 160 and the scanner 130 captures 3D information of the environment 160. In some examples, the camera 120 and the scanner 130 are separate devices; however, in some examples, the camera 120 and the scanner 130 are integrated into a single device. For example, the camera 120 includes depth acquisition functionality and / or is used in combination with a 3D acquisition depth camera, such as a time of flight camera, a stereo camera, a triangulation scanner, LIDAR, and the like. In some examples, 3D information is measured / acquired / captured using a projected light pattern and a second camera (or the camera 120) using triangulation techniques for performing depth determinations. In some examples, a time-of-flight (TOF) approach is used to enable intensity information (2D) and depth information (3D) to be acquired / captured. The camera 120 is, in various instances, a stereo-camera to facilitate 3D acquisition. In some examples, a 2D image and 3D information (i.e., a 3D data set) is captured / acquired at the same time; however, the 2D image and the 3D information is obtained at different times. For clarity, as used hereinMEl\57042703.vl 5Attorney Docket No. P1825-US-WO (138178-80220) the term “scanner” includes a device consisting of only of at least one 2D camera and a device for determining a position and orientation (e.g. the pose) of the camera, such as a device used in photogrammetry for example. The device for determining the position of the of the cameras are in various embodiments active, such as with inertial measurement units, global positioning satellites (GPS), or based on cellular tower triangulation, for example. The device for determining the position of the camera is alternatively passive in some implementations, where position is determined by performing image analysis or via computer vision techniques using simultaneous localization and mapping (SLAM) algorithms.
[0023] The user device 140 (e.g., a smartphone, a laptop or desktop computer, a tablet computer, a wearable computing device, a smart display, and the like) is located within the environment 160. The user device 140 displays an image of the environment 160, such as on a display (not shown) of the user device 140. In some examples, the user device 140 includes components such as a processor, a memory, an input device (e.g., a touchscreen, a mouse, a microphone, etc.), an output device (e.g., a display, a speaker, etc.), and the like. As one example, the user device 140 includes computerexecutable instructions stored on a memory and executable by a processor to implement the process 200 in FIG. 2 and the process 300 in FIG. 3 on the user device 140.
[0024] The processing system 100 provides for performing initialization and tracking functions in various implementations. As an example of the initialization, the camera 120 and the scanner 130 capture data 109a about the environment 160. The processing system 100 uses the data 109a about the environment 160 to generate the data 109b, which is a database of 2D features and associated 3D coordinates. The processing system 100 then performs tracking, such as tracking the user device 140 as it moves through the environment 160. The tracking includes, for example, determining a position and an orientation (e.g. the pose) of the user device 140 within the environment 160, which is accomplished by comparing an image captured by the user device 140 to the data 109b, which is the database of 2D features and associated 3D coordinates.MEl\57042703.vl 6Attorney Docket No. P1825-US-WO (138178-80220)
[0025] The features and functionality of the processing system 100 are now described further with reference to FIGS. 2-3. In particular, FIG. 2 depicts a nonlimiting flow diagram of a method 200 for performing initialization according to an embodiment. The method 200 is performed by or implemented on any suitable processing system (e.g., the processing system 100 of FIG. 1, a cloud computing node (not shown)), any suitable processing device (e.g., the processing device 102 of FIG. 1), and combinations thereof. FIG. 2 is now described in more detail with reference to the elements of at least FIG. 1 but is not so limited.
[0026] At block 202, data about the environment 160 is captured, such as by acquiring a plurality of 2D images of the environment from different positions. At block 204 the initialization engine 110 extracts 2D features and descriptors using the 2D images. At block 206, the initialization engine 110 estimates 3D coordinates for the 2D features using a method such as photogrammetry. A scene of the environment 160 is pre-scanned by taking 2D images, which includes frame array or spherical images. The initialization engine 110 performs feature extraction on these images (e.g., the data 109a) and matches like features together from different images. Examples of feature extractors include, but are not limited to, the following:• Harris corner detectors;• Harris-Laplace or other scale-invariant versions of a Harris detector;• Multi-Scale Oriented Patches (MOPs) (including a feature descriptor);• Scale-Invariant Feature Transform (SIFT) (including a descriptor);• Speeded Up Robust Features (SURF) (including a descriptor);• Features from Accelerated Segment Test (FAST);• Binary Robust Invariant Scalable Keypoints (BRISK) (including a descriptor)• Oriented FAST and Rotated BRIEF (ORB) (including a descriptor);• KAZE with modified SURF (M-SURF) descriptor, which outperforms SIFT and SURF in some cases;• AKAZE - accelerated version of KAZE with M-LDB descriptor (modified fast binary descriptor); and• Learned Invariant Feature Transforms.MEl\57042703.vl 7Attorney Docket No. P1825-US-WO (138178-80220)
[0027] Matching the features includes recognizing a feature in multiple images and then estimating the 3D coordinates of that feature, such as using photogrammetry or laser scanning. In order to do the feature matching with greater accuracy, descriptors are defined for the extracted features. Some of the feature extractors include descriptor definitions like SIFT, SURF, BRISK, and ORB as identified above. In practice, any feature descriptor definition is useful for associating the extracted features. For example, the following descriptor definitions are possible, as are other descriptor definitions:• Normalized gradient;• Principle Component Analysis (PCA) transformed image patch;• Histogram of oriented gradients; and• Gradient Location and Orientation Histogram (GLOH), Local Energy -Based Shape Histogram (LESH), BRISK, ORB, Fast Retina Keypoint (FREAK), and Local Discriminant Bases (LDB).
[0028] Photogrammetry is a technique for measuring objects using images, such as photographic images acquired by a digital camera, for example. Photogrammetry techniques typically make 3D measurements from 2D images or photographs. When a plurality of images are acquired at different positions that have an overlapping field of view, common points or features are identified within each image. By projecting a ray from the camera location to the feature / point on the object, the 3D coordinate of the feature / point is determined using trigonometry or triangulation. In some examples, photogrammetry is based on markers / targets (e.g., lights or reflective stickers) or instead based on natural features in the environment 160. To perform photogrammetry, for example, images are captured, such as with a camera (e.g., the camera 120) having a sensor, such as a photosensitive array. By acquiring multiple images of at least a portion of an object from different positions or orientations, 3D coordinates of points on the object are determined based on common features or points, along with information on the position and orientation of the camera when each image was acquired. In order to obtain the desired information for determining 3D coordinates, the features are identified in at least two images. Since the images areMEl\57042703.vl 8Attorney Docket No. P1825-US-WO (138178-80220) acquired from different positions or orientations, the common features are located in overlapping areas of the field of view of the images. Process 300 describes the methodology of the photogrammetry techniques disclosed herein.
[0029] Starting at block 302, the process 300 extracts at least one 2D feature associated with the 2D images (e.g., data 109a) captured from the environment. In some embodiments, the feature extraction process includes splitting the 2D images into a plurality of segments, which increase the extracted features’ quality and distribution. The plurality of segments thus form a normalized grid structure containing an even distribution of features across the 2D images. By splitting the 2D images into a normalized grid structure, the number of features is normalized by counting the number of features in a preset grid size and normalizing the number of features across the image. In this stage, the extracted features are filtered down to a more desirable value for efficient processing.
[0030] In some embodiments, the 2D images are subdivided into smaller areas as shown in FIG. 4. The subdivision 400 ensures a desired level of distribution of the features across the 2D images despite having a normalized distribution of features. This is desirable, for example, when the feature count is very high in one area of the 2D images relative to other areas, which will bias the normalized distribution of the features. This subdivision also helps circumvent the memory limitations of when the 2D images increase in size and are typically unable to fit in a Graphics Processing Unit (GPU) memory buffer. By subdividing the 2D images into smaller chunks, the feature extraction process is more efficiently performed.
[0031] In other embodiments, the 2D images are subdivided using a dual-ring partitioning processes 500 as shown in FIG. 5. In an embodiment, this subdivision process is used with fisheye 2D images to increase the quality and the matching probability of extracted features. The middle of a fisheye 2D image typically contains the least distorted features in a 2D image, which increases the matching quality. Furthermore, the middle of the fisheye 2D image has a lower probability of a person being present within the middle of the image. As result, the feature budget in the middle of the fisheye 2D image is increased, while decreasing the feature budget in the area ofMEl\57042703.vl 9Attorney Docket No. P1825-US-WO (138178-80220) the fisheye 2D image surrounding the middle of the fisheye 2D image, in various instances.
[0032] At block 304 the process 300 identifies a match in a database of 2D features as well as associated 3D coordinates for each 2D feature. Process 300 identifies shared features within distinct 2D images. The k-nearest neighbors algorithm (kNN) accomplishes this by way of a nearest neighbor search. An accurate, albeit slow, technique for this task is a brute-force approach, which involves comparing each point in one 2D image to all other points in another 2D image in order to determine its closest neighbor. In most instances, photogrammetry involves thousands to billions of individual features, making brute-force or manual computation prohibitively timeconsuming. To reduce processing time, in one non-limiting embodiment, the descriptors in an n-dimensional space are clustered into Voronoi cells 600, an example of which is shown in FIG. 6. Each of the cells have a different color, shape or combination thereof in various instances. Each of the clusters encapsulates similar features of the 2D image. It should be noted that the terms clusters and cells are sometimes used interchangeably herein.
[0033] Each cluster is assigned a unique index value stored, for example, in an inverted file index. By assigning a unique index value to each cluster, swift localization of a Voronoi cluster likely to contain the nearest neighbor is achieved. A search is then conducted within the corresponding cluster. In some embodiments, this search scope expands to neighboring clusters, enhancing the probability of identifying the nearest neighbor. The speed of the search is increased to be several orders of magnitude faster than that of a conventional CPU by utilizing Compute Unified Device Architecture (CUD A) in various implementations. This particularly occurs if the search is executed in batches.
[0034] In an embodiment, Fast Library for Approximate Nearest Neighbors (FLANN) is combined with kD tree neighbor search, optimizing the selection of pairs to search and thereby reducing the number of potential pairs. This is visually represented in FIG. 8, which depicts kD tree 800. In this approach, the features in ImageMEl\57042703.vl 10Attorney Docket No. P1825-US-WO (138178-80220)4 are exclusively matched to Images 5, 2, and 1 therein, effectively reducing or minimizing the number of searches.
[0035] Batches of searches are created, while simultaneously querying thousands of features in typical implementations. Vectors describing the features are significantly compressed using a technique called product quantization. Consequently, a substantially higher number of features are accommodated in GPU memory compared to traditional methods. In this way, almost all the features are queried in parallel, speeding up the query time significantly. It is noted that in an embodiment, the distribution of the features into the Voronoi clusters is unique to each dataset, and therefore sample data from a subset of the features are first used to “learn” or train how to cluster the features. This takes a significant amount of computational time depending on the number of features, and in an embodiment, is done anytime the features differ significantly. Furthermore, moving the features onto the GPU is also time-consuming, even when built directly on the GPU, which is a significant portion of the search time. As such, the algorithm that is selected to conduct the search changes based on the size of the data set. In some embodiments, the process 300 implements a Hierarchical Navigable Small Worlds (HNSW) search algorithm. In such a situation, HNSW is used because creating and filling the searchable index takes longer than the actual search time, making HNSW more efficient in certain scenarios. After the process 300 identifies at least one feature in the database that match the features extracted from the 2D images, the process then progresses to block 306 and filters out features from the same image associated with the 2D features and associated 3D coordinates.
[0036] At block 306, the process filters out (i.e., excludes) features originating from the same 2D image. In one embodiment, a Lowe’s Ratio Test is performed by process 300 in order to evaluate a distance ratio to the current nearest neighbor of a feature and the distance to the subsequent nearest neighbor. If the ratio exceeds a predetermined threshold, the nearest neighbor is rejected. In another embodiment, a Grid-Motion based Statistics (GMS) filtering is employed. GMS filtering, also referred to as feature crosschecking, cross checks that a feature in a first 2D image is also present in a second 2D image to which the first 2D image is being compared. GMS filtering also cross checks that the feature in the second 2D image is also present in the first 2DMEl\57042703.vl 11Attorney Docket No. P1825-US-WO (138178-80220) image. If both of these conditions hold, then the feature in the first 2D image and the feature in the second 2D image is accepted. The cross-checking process 700 is illustrated in FIG. 7 where each feature in Image 1 that is connected to a feature in Image 2 by two distinct non-overlapping lines is a feature that is present in both images. As a result, those features are accepted as a feature pair that ensures that the features are from two distinct and different images. The process 300 removes any repeated or incorrect feature match identified in block 302 from the list of features before the features are passed along to a reconstruction process (not shown in FIG. 3).
[0037] In some embodiments, it should be appreciated that alternative photogrammetry techniques are used and described in commonly-owned U.S. Patent 10,659,753, the contents of which are incorporated by reference herein.
[0038] The above methods for processing images provide advantages in many applications, such as the processing of a plurality of images acquired by an unmanned aerial vehicle (i.e. a drone) for example. These methods also are used with hand-held devices that acquire images, such as at construction sites or other large scale environments. In an embodiment, this method is used to improve image processing of images described in commonly owned United States Patent 11,663,785 entitled “Augmented and Virtual Reality” filed on April 26, 2021, the contents of which are incorporated herein by reference.
[0039] It should be appreciated that embodiments herein may describe the acquisition of image as discrete / individual images, however the claims should not be so limited. In other embodiments, the images may be acquired from a video (e.g. an image capture device that acquires a plurality of sequential images, typically at a predetermined rate) without deviating from the teachings provided herein.
[0040] Still further embodiments may utilize the methods provided herein in combination with a portable / mobile scanning device or a three-dimensional coordinate measurement device. These devices typically use a method, such as simultaneous localization and mapping (SLAM), to track the position of the mobile scanning device in the environment. The methods provided herein may be used to improve theMEl\57042703.vl 12Attorney Docket No. P1825-US-WO (138178-80220) functionality of the SLAM method and decrease the computational processing time of tracking the location of the scanning device.
[0041] In addition to the features described herein, or as an alternative, further embodiments of the method include that generating the database of 2D features and associated 3D coordinates by performing photogrammetry using at least two images.
[0042] In addition to the features described herein, or as an alternative, further embodiments of the method include that the data about the environment is captured using a camera.
[0043] In addition to the features described herein, or as an alternative, further embodiments of the method include that the data about the environment includes a point cloud.
[0044] In addition to the features described herein, or as an alternative, further embodiments of the method include that the data about the environment is captured using a three-dimensional scanner.
[0045] In addition to the features described herein, or as an alternative, further embodiments of the method include that the data about the environment includes at least one image captured using a camera and a point cloud captured using a three- dimensional scanner.
[0046] In addition to the features described herein, or as an alternative, further embodiments of the method include that generating the database of 2D features and associated 3D coordinates includes: performing feature extraction; determining a first 3D coordinate for a first feature of a plurality of features using a laser scan; and responsive to determining that a second 3D coordinate cannot be determined for a second feature of the plurality of features using the laser scan, determining the second 3D coordinate for the second feature of the plurality of features using photogrammetry.
[0047] In addition to the features described herein, or as an alternative, further embodiments of the method include that generating the database of 2D features andMEl\57042703.vl 13Attorney Docket No. P1825-US-WO (138178-80220) associated 3D coordinates by performing feature extraction on the data about the environment.
[0048] In addition to the features described herein, or as an alternative, further embodiments of the method include that the data about the environment are images of the environment, and wherein generating the database of 2D features and associated 3D coordinates based at least in part on the data about the environment is performed using photogrammetry.
[0049] In addition to the features described herein, or as an alternative, further embodiments of the method include generating the database of 2D features and associated 3D coordinates based at least in part on the data about the environment by performing feature extraction.
[0050] In addition to the features described herein, or as an alternative, further embodiments of the method include that the data about the environment is captured using a three-dimensional scanner.
[0051] In addition, some embodiments described herein are associated with an “indication.” As used herein, the term “indication” is used to refer to any indicia and / or other information indicative of or associated with a subject, item, entity, and / or other object and / or idea. As used herein, the phrases “information indicative of’ and "indicia" is used to refer to any information that represents, describes, and / or is otherwise associated with a related entity, subject, or object. Indicia of information includes, for example, a code, a reference, a link, a signal, an identifier, and / or any combination thereof and / or any other informative representation associated with the information. In some embodiments, indicia of information (or indicative of the information) includes the information itself and / or any portion or component of the information. In some embodiments, an indication includes a request, a solicitation, a broadcast, and / or any other form of information gathering and / or dissemination.
[0052] Numerous embodiments are described herein and are presented for illustrative purposes only. The described embodiments are not, and are not intended to be, limiting in any sense. The presently disclosed invention(s) are widely applicable toMEl\57042703.vl 14Attorney Docket No. P1825-US-WO (138178-80220) numerous embodiments, as is readily apparent from the disclosure. One of ordinary skill in the art will recognize that the disclosed invention(s) are practicable with various modifications and alterations, such as structural, logical, software, and electrical modifications. Although particular features of the disclosed invention(s) are described with reference to one or more particular embodiments and / or drawings, it should be understood that such features are not limited to usage in the one or more particular embodiments or drawings with reference to which they are described, unless expressly specified otherwise.
[0053] Devices that are in communication with each other need not be in continuous communication with each other, unless expressly specified otherwise. On the contrary, such devices need only transmit to each other as necessary or desirable, and will actually refrain from exchanging data most of the time. For example, a machine in communication with another machine via the Internet will not transmit data to the other machine for weeks at a time. In addition, devices that are in communication with each other communicate directly or indirectly through at least one intermediary.
[0054] A description of an embodiment with several components / features does not imply that all or even any of such components and / or features are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the present invention(s). Unless otherwise specified explicitly, no component and / or feature is essential or required.
[0055] Further, although process steps, algorithms or the like are described in a sequential order, such processes are be configured to work in different orders. In other words, any sequence or order of steps that are explicitly described does not necessarily indicate a requirement that the steps be performed in that order. The steps of processes described herein are performed in any practical order. Further, some steps are performed simultaneously despite being described or implied as occurring non- simultaneously (e.g., because one step is described after the other step). Moreover, the illustration of a process by its depiction in a drawing does not imply that the illustrated process is exclusive of other variations and modifications thereto, does not imply thatMEl\57042703.vl 15Attorney Docket No. P1825-US-WO (138178-80220) the illustrated process or any of its steps are necessary to the invention, and does not imply that the illustrated process is preferred.
[0056] The term “determining” (and like terms) includes calculating, computing, deriving, looking up (e.g., in a table, database or data structure), ascertaining, and the like.
[0057] It will be readily apparent that the various methods and algorithms described herein are implemented by, e.g., appropriately and / or specially-programmed general purpose computers and / or computing devices. Typically, a processor (e.g., one or more microprocessors) will receive instructions from a memory or like device, and execute those instructions, thereby performing one or more processes defined by those instructions. Further, programs that implement such methods and algorithms are stored and transmitted using a variety of media (e.g., computer readable media) in a number of manners. In some embodiments, hard-wired circuitry or custom hardware are used in place of, or in combination with, software instructions for implementation of the processes of various embodiments. Thus, embodiments are not limited to any specific combination of hardware and software.
[0058] A “processor” generally means any of microprocessors, digital CPU devices, GPU devices, computing devices, microcontrollers, digital signal processors (DSPs), field programmable gate arrays (FPGAs), and like devices, as further described herein. A CPU typically performs a variety of tasks while a GPU is optimized to display or process images and 3D datasets.
[0059] Where databases are described, it will be understood by one of ordinary skill in the art that (i) alternative database structures to those described are readily employed, and (ii) other memory structures besides databases are readily employed. Any illustrations or descriptions of any sample databases presented herein are illustrative arrangements for stored representations of information. Any number of other arrangements are employed besides those suggested by, e.g., tables illustrated in drawings or elsewhere. Similarly, any illustrated entries of the databases represent exemplary information only; one of ordinary skill in the art will understand that theMEl\57042703.vl 16Attorney Docket No. P1825-US-WO (138178-80220) number and content of the entries is different from those described herein. Further, despite any depiction of the databases as tables, other formats (including relational databases, object-based models and / or distributed databases) could be used to store and manipulate the data types described herein. Likewise, object methods or behaviors of a database are used to implement various processes, such as the described herein. In addition, the databases are, in any known manner, stored locally or remotely from a device that accesses data in such a database.
[0060] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and components, but do not preclude the presence or addition of other features, integers, steps, operations, element components, and groups thereof.
[0061] Terms such as processor, controller, computer, DSP, FPGA are understood in this document to mean a computing device that are located within an instrument, distributed in multiple elements throughout an instrument, or placed externally to an instrument.
[0062] While the invention has been described in detail in connection with only a limited number of embodiments, it should be readily understood that the invention is not limited to such disclosed embodiments. Rather, the invention is modified to incorporate any number of variations, alterations, substitutions or equivalent arrangements not heretofore described, but which are commensurate with the spirit and scope of the invention. Additionally, while various embodiments of the invention have been described, it is to be understood that aspects of the invention include only some of the described embodiments. Accordingly, the invention is not to be seen as limited by the foregoing description, but is only limited by the scope of the appended claims.MEl\57042703.vl 17Attorney Docket No. P1825-US-WO (138178-80220)
[0063] The term “about” is intended to include the degree of error associated with measurement of the particular quantity based upon the equipment available at the time of filing the application. For example, “about” includes a range of ± 8% or 5%, or 2% of a given value.MEl\57042703.vl 18
Claims
Attorney Docket No. P1825-US-WO (138178-80220)CLAIMSWhat is claimed is:
1. A method comprising: capturing data about an environment by acquiring a plurality of two- dimensional images of a scene from different positions; extracting features from the plurality of two-dimensional images of the scene acquired from the different positions; storing the features extracted from the plurality of two-dimensional images of the scene acquired from the different positions in at least one memory; identifying matches between a plurality of features extracted from a plurality of first images of the plurality of two-dimensional images and a plurality of features extracted from a plurality of second images of the plurality of two-dimensional images, where the plurality of first images are distinct from the plurality of second images; filtering features from the plurality of two-dimensional images of the scene acquired from the different positions, by removing the features extracted from the plurality of first images that do not match the features extracted from the plurality of second images; determining three-dimensional coordinates of the scene based at least on the plurality of features extracted from the plurality of first images that match the plurality of features extracted from the plurality of second images; and excluding the features filtered from the plurality of two-dimensional images of the scene to determine the three-dimensional coordinates of the scene.
2. The method of claim 1, wherein identifying matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images comprises:MEl\57042703.vl 19Attorney Docket No. P1825-US-WO (138178-80220) generating clusters of the features extracted from the plurality of two- dimensional images of the scene; and searching, within the clusters of the features extracted from the plurality of two- dimensional images of the scene, for the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images.
3. The method of claim 1 , wherein identifying matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images comprises: generating clusters of the features extracted from the plurality of two- dimensional images of the scene; compressing the clusters of the features extracted from the plurality of two- dimensional images of the scene within the at least one memory; loading a portion of the compressed clusters of the features extracted from the plurality of two-dimensional images of the scene; decompressing the portion of the compressed clusters of the features extracted from the plurality of two-dimensional images of the scene to generate a portion of decompressed clusters of the features extracted from the plurality of two-dimensional images of the scene; and searching, within the portion of the decompressed clusters of the features extracted from the plurality of two-dimensional images of the scene, using a graphics processing unit, for the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images.
4. The method of claim 3, wherein the identifying of the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images further comprises:MEl\57042703.vl 20Attorney Docket No. P1825-US-WO (138178-80220) loading a next portion of the compressed clusters of the features extracted from the plurality of two-dimensional images of the scene while performing the searching, within the portion of the decompressed clusters of the features extracted from the plurality of two-dimensional images of the scene using the graphics processing unit, for the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images; decompressing the next portion of the compressed clusters of the features extracted from the plurality of two-dimensional images of the scene to generate a next portion of decompressed clusters of the features extracted from the plurality of two- dimensional images of the scene; and searching, within the next portion of the decompressed clusters of the features extracted from the plurality of two-dimensional images of the scene using the graphics processing unit, for the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images.
5. The method of claim 3, wherein the identifying of the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images further comprises: assigning, to each cluster of the features extracted from the plurality of two- dimensional images of the scene, a unique index value stored in an inverted filed index.
6. The method of claim 3, further comprising: training an algorithm for generating the clusters of the features extracted from the plurality of two-dimensional images of the scene, using sample data from a subset of the features extracted from the plurality of two-dimensional images of the scene.
7. The method of claim 3, wherein compressing the clusters of the features extracted from the plurality of two-dimensional images of the scene within the at least one memory comprises quantizing the features extracted from the plurality of two- dimensional images of the scene.MEl\57042703.vl 21Attorney Docket No. P1825-US-WO (138178-80220)8. The method of claim 1, wherein: the filtering of the features from the plurality of two-dimensional images of the scene acquired from the different positions comprises cross-checking that a feature of a first image of the plurality of first images matches a feature of a second image of the plurality of second images, and that the feature of the second image of the plurality of second images matches the feature of the first image of the plurality of first images, and the feature of the first image matches the feature of the second image of the plurality of second images when information from the feature of the second image comprises information from the feature of the first image, and the feature of the second image of the plurality of second images does not match the feature of the first image when the first image does not comprise information of the second image.
9. The method of claim 1, further comprising: subdividing the two-dimensional images into areas based on a predetermined distribution level of the features of the plurality of two-dimensional images, prior to extracting the features from the plurality of two-dimensional images of the scene acquired from the different positions.
10. The method of claim 1, further comprising: subdividing at least a portion of two-dimensional images using a dual-ring partitioning process prior to extracting the features from the plurality of two- dimensional images of the scene acquired from the different positions, wherein the portion of the two-dimensional images that are subdivided using the dual-ring partitioning process comprise fisheye two-dimensional images.
11. The method of claim 1, wherein determining the three-dimensional coordinates of the scene comprises performing photogrammetry using the plurality of two- dimensional images of the scene acquired from the different positions.
12. The method of claim 1, wherein the data about the environment is captured using at least one of a camera and a three-dimensional scanner.MEl\57042703.vl 22Attorney Docket No. P1825-US-WO (138178-80220)13. The method of claim 1, wherein the data about the environment comprises a point cloud.
14. The method of claim 1, wherein determining the three-dimensional coordinates of the scene comprises: determining a first three-dimensional coordinate for a first feature of the features extracted from the plurality of two-dimensional images of the scene using a laser scan; and responsive to determining that a second three-dimensional coordinate cannot be determined for a second feature of the features extracted from the plurality of two- dimensional images using the laser scan, determining the second three-dimensional coordinate for the second feature using photogrammetry.
15. A method comprising: extracting features from a plurality of two-dimensional images of a scene acquired from different positions; storing the features extracted from the plurality of two-dimensional images of the scene acquired from the different positions in at least one memory; generating clusters of the features extracted from the plurality of two- dimensional images of the scene; compressing the clusters of the features extracted from the plurality of two- dimensional images of the scene within the at least one memory; loading a portion of the compressed clusters of the features extracted from the plurality of two-dimensional images of the scene from the at least one memory; decompressing the portion of the compressed clusters of the features extracted from the plurality of two-dimensional images of the scene to generate a portion of decompressed clusters of the features extracted from the plurality of two-dimensional images of the scene;MEl\57042703.vl 23Attorney Docket No. P1825-US-WO (138178-80220) searching, within the portion of the decompressed clusters of the features extracted from the plurality of two-dimensional images of the scene using a graphics processing unit, for matches between the plurality of features extracted from a plurality of first images of the plurality of two-dimensional images and the plurality of features extracted from a plurality of second images of the plurality of two-dimensional images; and generating three-dimensional coordinates of the scene based on the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images.
16. The method of claim 15, wherein the identifying of the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images further comprises: loading a next portion of the compressed clusters of the features extracted from the plurality of two-dimensional images of the scene from the at least one memory while performing the searching, within the portion of the decompressed clusters of the features extracted from the plurality of two-dimensional images of the scene using the graphics processing unit, for the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images; decompressing the next portion of the compressed clusters of the features extracted from the plurality of two-dimensional images of the scene to generate a next portion of decompressed clusters of the features extracted from the plurality of two- dimensional images of the scene; and searching, within the next portion of the decompressed clusters of the features extracted from the plurality of two-dimensional images of the scene using the graphics processing unit, for the matches between the plurality of features extracted from the plurality of first images and the plurality of features extracted from the plurality of second images.MEl\57042703.vl 24Attorney Docket No. P1825-US-WO (138178-80220)17. The method of claim 15, wherein compressing the clusters of the features extracted from the plurality of two-dimensional images of the scene within the at least one memory further comprises quantizing the features extracted from the plurality of two- dimensional images of the scene.
18. The method of claim 15, further comprising: filtering features from the plurality of two-dimensional images of the scene acquired from the different positions, by removing the features extracted from the plurality of first images that do not match the features extracted from the plurality of second images; and determining to exclude the features filtered from the plurality of two- dimensional images of the scene to determine the three-dimensional coordinates of the scene.
19. A method comprising: capturing a plurality of two-dimensional images of an environment from different positions; generating a database of two-dimensional features based at least in part on the plurality of two-dimensional images of the environment; extracting two-dimensional features from a two-dimensional image; identifying matches between the two-dimensional features of the two- dimensional image and the two-dimensional features in the database; and determining three-dimensional coordinates of the environment based on the matches between the two-dimensional features of the two-dimensional image and the two-dimensional features in the database.
20. The method of claim 19, wherein the identifying of the matches between the two- dimensional features of the two-dimensional image and the two-dimensional features in the database further comprises:MEl\57042703.vl 25Attorney Docket No. P1825-US-WO (138178-80220) generating clusters of the two-dimensional features within the database; compressing the clusters of the two-dimensional features within at least one memory; loading a portion of the compressed clusters of the two-dimensional features from the at least one memory; decompressing the portion of the compressed clusters of the two-dimensional features to generate a portion of decompressed clusters of the two-dimensional features; and searching, within the portion of the decompressed clusters of the two- dimensional features using a graphics processing unit, for the matches between the two- dimensional features of the two-dimensional image and the two-dimensional features in the database.MEl\57042703.vl 26
Citation Information
Patent Citations
Photogrammetry system and method of operation
US10659753B2
Augmented and virtual reality
US11663785B2
Tracking with reference to a world coordinate system
US20220414925A1
US202463690977P