Methods and systems for landmark-based visual localization and pose estimation of large dynamic objects

WO2026207510A1PCT designated stage Publication Date: 2026-10-01NIKON CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/021387
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure US2026021387_01102026_PF_FP_ABST
    Figure US2026021387_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems for using landmark maps associated with an object to control a robotic operating apparatus are disclosed which obtain a landmark map associated with an object, determine a plurality of positions and orientations assumed by the object over time based on the landmark map, and determine operations to be performed on the object by the robotic operating apparatus based on the plurality of positions and orientations. Methods and systems for generating landmark maps associated with an object for use in controlling a robotic operating apparatus are disclosed which generate a landmark map associated with an object and transmit the landmark map for subsequent use such as: determining a plurality of positions and orientations assumed by the object over time based on the landmark map, and determining operations to be performed on the object by a robotic operating apparatus based on the plurality of positions and orientations.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR LANDMARK-BASED VISUAL LOCALIZATION AND POSE ESTIMATION OF LARGE DYNAMIC OBJECTSCROSS-REFERENCE

[0001] The present application claims priority to U.S. Provisional Patent Application No. 63 / 779,215 filed on March 27, 2025, entitled “METHODS AND SYSTEMS FOR LANDMARK-BASED VISUAL LOCALIZATION AND POSE ESTIMATION OF LARGE DYNAMIC OBJECTS,” which is incorporated herein by reference in its entirety for all purposes.BACKGROUND

[0002] Numerous applications in robotics technologies may require robots or robotic operating apparatuses to perform processing operations (such as manufacturing operations) on large objects as those large objects move relative to the robots or robotic operating apparatuses. For instance, large objects may move relative to the robots or robotic operating apparatuses on conveyor belts or other movement systems. During such movement, the objects may assume a number of different positions or orientations (also known as “poses”). Thus, in order for the robots or robotic operating apparatuses to perform the processing operations, the robots or robotic operating apparatuses must be capable of determining the object’s poses and adapt their own motions to changes in the object’s poses over time. Accordingly, presented herein are methods and systems that permit rapid and accurate estimation of poses of objects for use in controlling robots or robotic operating apparatuses.BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Various embodiments of the inventions are disclosed in the following detailed description and the accompanying drawings.

[0004] FIG. 1 A shows a flowchart depicting an exemplary method for using a landmark map associated with an object to control a robotic operating apparatus.

[0005] FIG. IB shows an example of how a landmark or key point identifies a location in an image that exhibits distinctive visual characteristics according to certain embodiments.

[0006] FIG. 1C depicts an example of generating a visual feature descriptor according to certain embodiments.

[0007] FIG. 2A shows a schematic depicting an exemplary system for using a landmark map associated with an object to control a robotic operating apparatus.

[0008] FIG. 2B shows a schematic depicting another exemplary system for using a landmarkmap associated with an object to control a robotic operating apparatus.

[0009] FIG. 3 shows a flowchart depicting an exemplary method for generating a landmark map associated with an object for use in controlling a robotic operating apparatus.

[0010] FIG. 4A shows a schematic depicting an exemplar}7system for generating a landmark map associated with an object for use in controlling a robotic operating apparatus.

[0011] FIG. 4B shows an example of an object and a set of landmarks or key points according to certain embodiments.

[0012] FIG. 4C shows an example of an object and keyframes containing keypoints according to certain embodiments.

[0013] FIG. 5 shows a flowchart depicting a first exemplary7method for processing a target obj ect with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device.

[0014] FIG. 6 shows a schematic depicting a first exemplary system for processing a target obj ect with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device.

[0015] FIG. 7 shows a flowchart depicting a second exemplary method for processing a target obj ect with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device.

[0016] FIG. 8 shows a schematic depicting a second exemplary system for processing a target obj ect with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device.

[0017] FIG. 9 shows a block diagram of a computer system for implementing the methods herein.

[0018] FIG. 10A shows a schematic overview of the data flow between modules in the systems described herein for using landmark maps associated with an object to control a robotic operating apparatus.

[0019] FIG. 10B shows an exemplary method for kinematic pose fusion for correcting a visual pose estimation of an object tracked by a robotic operating apparatus.

[0020] FIG. 10C shows another exemplary method for kinematic pose fusion for correcting a visual pose estimation of an object tracked by a robotic operating apparatus.

[0021] FIG. 11 shows an exemplary proportional integral controller with lead filter and feedforward components for use with the methods and systems described herein.

[0022] FIG. 12 shows an exemplary plot of the pose estimate obtained with a system using the proportional integral controller of FIG. 11 compared against a desired trajectory.

[0023] FIG. 13 shows another exemplary plot of the pose estimate obtained with a system using the proportional integral controller of FIG. 11 compared against a desired trajectory.DETAILED DESCRIPTION

[0024] The inventions can be implemented in numerous ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by a memory coupled to the processor. In this specification, these implementations, or any other form that the inventions may take, may be referred to as techniques. In general, the order of the steps of disclosed processes may be altered within the scope of the inventions. Unless stated otherwise, a component such as a processor or a memory described as being configured to perform a task may be implemented as a general component that is temporarily configured to perform the task at a given time or a specific component that is manufactured to perform the task. As used herein, the term “processor” refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.

[0025] A detailed description of one or more embodiments of the inventions is provided below along with accompanying figures that illustrate the principles of the inventions. The inventions are described in connection with such embodiments, but the inventions are not limited to any embodiment. The scope of the inventions is limited only by the claims and the inventions encompass numerous alternatives, modifications and equivalents. Numerous specific details are set forth in the following description in order to provide a thorough understanding of the inventions. These details are provided for the purpose of example and the inventions may be practiced according to the claims without some or all of these specific details. For the purpose of clarity, technical material that is known in the technical fields related to the inventions has not been described in detail so that the inventions are not unnecessarily obscured.

[0026] The innovations described herein can be embodied in a multitude of different ways, for example, as defined and covered by the claims. In this description, reference is made to the drawings wherein like reference numerals can indicate identical or functionally similar elements. It will be understood that elements illustrated in the figures are not necessarily drawn to scale. Moreover, it will be understood that certain embodiments can include more elements than illustrated in a drawing and / or a subset of the elements illustrated in a drawing. Further, some embodiments can incorporate any suitable combination of features from two or more.

[0027] As used herein, the term “or” shall convey both disjunctive and conjunctive meanings,unless otherwise indicated or impossible. For instance, the phrase “A or B” shall be interpreted to include element A alone, element B alone, and the combination of elements A and B. As another example, the phraseC'A, B, or C” shall be interpreted to include element A alone, element B alone, element C alone, the combination of elements A and B but not C, the combination of elements A and C but not B, the combination of elements B and C but not A, and the combination of elements, A, B, and C.

[0028] All references and publications referred to herein, such as but not limited to U.S. patents and published applications, are to be understood as being herein incorporated by reference in their entirety. Unless otherwise clear from the context, like reference characters identify like parts throughout.

[0029] Numerous applications in robotics technologies may require robots or robotic operating apparatuses to perform processing operations (such as manufacturing operations) on large objects as those large objects move relative to the robots or robotic operating apparatuses. For instance, large objects may move relative to the robots or robotic operating apparatuses on conveyor belts or other movement systems. During such movement, the objects may assume a number of different positions or orientations (also known as "poses”). Thus, in order for the robots or robotic operating apparatuses to perform the processing operations, the robots or robotic operating apparatuses must be capable of determining the object’s poses and adapt their own motions to changes in the object's poses over time.

[0030] Accordingly, presented herein are methods and systems that permit rapid and accurate estimation of poses of objects for use in controlling robots or robotic operating apparatuses. The methods and systems utilize multiple techniques such as construction of three-dimensional (3D) models from two-dimensional (2D) images, visual segmentation of the 3D models, landmark tracking, and pose estimation, all while considering constraints imposed by the robotic operating apparatus’s state.

[0031] In order to interact with an object, the robots or robotic operating apparatus must be able to determine where the object is located and how it is oriented. However, the robots or robotic operating apparatus can generally only observe a portion or section of the object. A global image or model is therefore required to identity' where this observed segment is located on the object. A scanning process may be used to create such a global image or model or a global reference which can be reused later during runtime (i.e. during an online operational phase where the system is actively tracking and processing a moving target object). In some cases, a 3D model or visual data (e.g. images of the object) can be used to observe the entire object. Visual features such as landmarks or keypoints may be extracted from this data acrossthe entire object. The extracted visual features such as landmarks or key points and their associated locations may be stored as a global reference or landmark map to be reused at a later time.

[0032] As used herein, a landmark or keypoint identifies a location in an image that exhibits distinctive visual characteristics. As an example, distinctive visual characteristics can include comers, edge intersections, or regions of significant intensity variation. Thus, a landmark or keypoint is identified by or characterized by visual characteristics that make the landmark or keypoint reliably detectable and re-identifiable across different images of an object taken from different viewpoints or at different times.

[0033] As used herein, a time step refers to an image capture time step. In other words, a time step as used herein is each time the imaging module or imaging device (e.g.. a camera) captures a new color and / or depth image.

[0034] Additionally, where reference is made to an “object” or a “target object,” the terms are synonymous in that they refer to an object being processed by the disclosed systems and methods as described herein.

[0035] During robotic operations, in some embodiments, a camera can observe a portion of the object to generate 3D visual information. These images (i.e. the generated 3D visual information) can be processed to remove unwanted information such as background or image noise. Visual landmarks can be extracted from across the observed portion of the object. These visual landmarks can be matched across consecutive frames and to the global reference or landmark map to estimate a relative pose of the object with respect to the robots or robotic operating apparatus. To refine the estimate further and reduce effects from image noise, a state (e.g. pose, velocity) of the robot or robotic operating apparatus can be used as a constraint in optimizing the estimate.

[0036] Performing operations on the object can require the robots or robotic operating apparatus to follow a specific trajectory path or operational path, in some cases with some variation of speed along the path. In some embodiments, operational constraints are encoded into a trajectory’ generator that produces waypoints along the desired operational path according to a required operational speed. The waypoints can be predefined offline based on the desired operational path over the target object.

[0037] As used herein, a waypoint is a discrete reference point along the operational path, defined by a target position (e g., x, y, z coordinates) and orientation (e.g., roll, pitch, yaw or other rotation representations) for an end-effector of a robot or robotic operating apparatus (e.g., a robot arm) relative to a target object surface. The ordered sequence of way pointscollectively defines the path a processing device (i.e., a device that operates to perform a process on the target object) follows over the target object during the operation. In some embodiments, the processing device is mounted on, attached to, or removably coupled to the end-effector of the robot or robotic operating apparatus.

[0038] In some cases, the waypoints are preprocessed before runtime and stored to provide object pose references for controlling the robot or robotic operating apparatus. The waypoints can include position information and posture information and are used by a trajectory generator as described herein to define where the robot or robotic operating apparatus should go.

[0039] Finally, in some cases, the desired trajectory path or desired operational path can be tracked using a proportional-integral (Pl)-lead-feedforward controller. The current state estimate of the object can be provided by a visual tracker or visual tracking device or module as the 3D relative pose of the object with respect to the robot or robotic operating apparatus. As an example, the visual tracker or visual tracking device can be a 3D tracking module or a position determination module as described herein. It receives or is configured to receive processed image data (e.g., color and / or depth images) and estimates a total apparent motion by matching keypoints against the landmark map and across consecutive frames.

[0040] Moreover, the desired object pose can be provided by the trajectory waypoints. The feedforward component may permit consistent tracking of the trajectory even under dynamic motion of the object in relation to the robot or robotic operating apparatus.Methods and systems for using landmark maps associated with an object to control a robot or robotic operating apparatus

[0041] Provided herein are methods and systems for using landmark maps associated with an object to control a robot or robotic operating apparatus. The methods and systems generally obtain a landmark map associated with an object, determine a plurality of positions and orientations assumed by the object over time based on the landmark map, and determine a trajectory of operations to be performed on the object by the robot or robotic operating apparatus based on the plurality7of positions and orientations assumed by the object over time.

[0042] FIG. 1A shows a flowchart depicting an exemplary method 100 for using a landmark map associated with an object to control a robot or robotic operating apparatus.

[0043] In the example shown, a landmark map associated with an object is obtained at 110.

[0044] In some embodiments, as shown for example in FIG. 1 A at operation 110, the landmark map is obtained by: (i) obtaining a plurality7of images of the object from a plurality of different angles; (ii) forming a three-dimensional (3D) model of the object from the plurality of images; (iii) identifying a plurality of landmarks (or keypoints) on the object based on the 3D model;and (iv) collating the plurality of landmarks (or key points) to form the landmark map.

[0045] FIG. IB shows an example of how a landmark or key point identifies a location in an image that exhibits distinctive visual characteristics. Distinctive visual characteristics can include comers, edge intersections, or regions of significant intensity variation. In the example of FIG. IB, a star-shaped object 140 has landmarks or keypoints at 141 identifying the points of the star and at 142 identifying comers of the star. A set of intersecting rods at 150 have landmarks or key points shown at 151 (the ends or tips of the rods) and at 152 (the intersection points of the rods). Finally, on a circular or spherical object 160, a landmark or keypoint is show n at 161 denoting an area or region of significant intensify variation at the center of the object 160. Such visual characteristics as shown in FIG. IB make the identified landmarks or key points reliably detectable and re-identifiable across different images of objects taken from different viewpoints or at different times.

[0046] As known in the art, examples of methods to find landmarks or keypoints include FAST (see, e.g. Rosten, E., & Drummond, T. (2006, May). Machine learning for high-speed comer detection. In European conference on computer vision (pp. 430-443). Berlin. Heidelberg: Springer Berlin Heidelberg); Harris comer (see, e.g., Hanis, C., & Stephens, M. (1988, August). A combined comer and edge detector. In Alvey vision conference (Vol. 15, No. 50, pp. 10-5244); SIFT (see, e.g., Lowe, D. G. (2004). Distinctive image features from scaleinvariant keypoints. International journal of computer vision, 60(2), 91-110; SURF (see, e.g., Bay, H., Tuytelaars. T.. & Van GooL L. (2006, May). Surf: Speeded up robust features. In European conference on computer vision (pp. 404-417) Berlin, Heidelberg: Springer Berlin Heidelberg. Patent - Funayama, R., Yanagihara, H., Van Gool, L., Tuytelaars, T., & Bay, H. (2012). U.S. Patent No. 8,165,401. Washington, DC: U.S. Patent and Trademark Office; and SuperPoint (see. e.g., deep learning based keypoint: DeTone. D., Malisiewicz, T.. & Rabinovich, A. (2018). Superpoint: Self-supervised interest point detection and description. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops (pp. 224-236). The disclosed techniques are not limited by any specific approach to finding landmarks or keypoints; any of the approaches mentioned herein and as known in the art may be used without impacting the scope of what is disclosed; and all of the aforementioned references are herein incorporated by reference for all purposes in their entirety.

[0047] Additionally, and as described in further detail herein with respect to FIGS. 1B-1C, descriptors such as visual feature descriptors can be associated with each landmark or keypoint. Visual feature descriptors can be numerical representations (e.g., binary vectors or floatingpoint vectors) that encode a local appearance pattern surrounding each landmark or keypoint.Examples of such visual feature descriptors include ORB, SIFT, SURF, and BRIEF descriptors. These descriptors can enable matching of landmarks or keypoints across different images by comparing their numerical similarity.

[0048] FIG. 1C depicts an example of generating a visual feature descriptor. In this example, a landmark or key point is shown at 141 identified on an object 140. Based on a local appearance pattern or neighborhood (indicated by a box 170) surrounding the key point 141, a visual feature descriptor is generated through a process 171 that encodes the local appearance pattern into encoded into a numerical representation 172 (i.e. the visual feature descriptor), which can be a binary vector or a floating-point vector.

[0049] In some embodiments, the plurality of images comprises at least about 2, 3, 4, 5. 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300. 400, 500, 600. 700, 800, 900, 1.000, or more images, at most about 1,000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, or 2, or a number of images that is within a range defined by any tw o of the preceding values. In some embodiments, the plurality' of images is obtained while the object is stationary. In some embodiments, each image of the plurality of images comprises a two-dimensional (2D) image or a 3D image. In some embodiments, at least some images of the plurality of images comprise a stereoscopic image, a time of flight image, or a structured light image. In some embodiments, the plurality of images are obtained from at least about 2, 3, 4, 5, 6, 7, 8. 9, 10, 20, 30, 40. 50. 60, 70, 80, 90, 100, 200, 300, 400. 500, 600, 700, 800, 900, 1,000, or more different angles, at most about 1,000. 900, 800, 700. 600, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, or 2 different angles, or a number of different angles that is within a range defined by any two of the preceding values.

[0050] In other embodiments, as shown for example in FIG. 1 A at operation 110, the landmark map is obtained by : (i) obtaining a plurality of images of the obj ect from a plurality of different angles; (ii) extracting texture-based landmark points from each image of the plurality of images, thereby forming a set of extracted landmark points; (iii) performing a segmentation procedure on the set of extracted landmark points to retain only landmark points on the object, thereby forming a set of object landmark points; (iv) extracting landmark points from consecutive frames of the set of object landmark points, thereby forming a set of consecutive object landmark points; and (v) collating the plurality of consecutive object landmark points to form the landmark map. In some embodiments, the plurality of images comprises any number of images described herein. In some embodiments, the plurality of images is obtained while the object is stationary. In some embodiments, each image of the plurality of images comprises a 2D image or a 3D image. In some embodiments, at least some images of the plurality ofimages comprise a stereoscopic image, a time of flight image, or a structured light image. In some embodiments, the plurality of images is obtained from any number of different angles described herein.

[0051] In the example shown, a plurality of positions and orientations assumed by the object over time is determined at 120. In some embodiments, the plurality of positions and orientations assumed by the object are determined based on the landmark map. In some embodiments, as shown for example in FIG. 1 A at operation 120, the plurality of positions and orientations are determined by: (i) obtaining a plurality of images of the object at a plurality of points in time; and (ii) determining a location and orientation of the object at each point in time based on the plurality' of images. In some embodiments, the plurality of images are obtained at a frame rate of at least about 1 hertz (Hz). 2 Hz. 3 Hz, 4 Hz, 5 Hz, 6 Hz, 7 Hz, 8 Hz. 9 Hz, 10 Hz, 15 Hz, 20 Hz, 25 Hz, 30 Hz, or more, at most about 30 Hz, 25 Hz, 30 Hz, 35 Hz, 40 Hz, 45 Hz, 50 Hz, 55 Hz, 60 Hz, 70 Hz, 80 Hz, 90 Hz, 100 Hz, 200 Hz, 300 Hz, 400 Hz, 500 Hz, or more, at most about 500 Hz, 400 Hz, 300 Hz, 200 Hz, 100Hz, 90 Hz, 80 Hz, 70 Hz, 60 Hz, 55 Hz, 50 Hz, 45 Hz, 40 Hz, 35 Hz. 30 Hz, 25 Hz, 20 Hz, 15 Hz, 10 Hz, 9 Hz, 8 Hz, 7 Hz, 6 Hz, 5 Hz, 4 Hz, 3 Hz, 2 Hz, 1 Hz, or a frame rate that is within a range defined by any two of the preceding values. In some embodiments, the plurality of images is obtained while the object is moving.

[0052] In some embodiments, the plurality of positions and orientations are further determined (for example, as shown in FIG. 1A at operation 120) by calculating an expected position and orientation to be assumed by the object based on the location and orientation of the object at a most recent point in time. That is, in some embodiments, the location and orientation of the object at a most recent point in time is determined and such information is used to calculate an expected location and orientation at a later point in time. In some embodiments, calculating the expected location and orientation at the later point in time permits updates to control of the robot or robotic operating apparatus (discussed below ) at a faster rate than would be permitted based on the plurality7of images (which is limited by the frame rate). In some embodiments, the expected position and orientation are calculated at a refresh rate of at least about 10 Hz, 20 Hz. 30 Hz. 40 Hz, 50 Hz, 60 Hz, 70 Hz. 80 Hz. 90 Hz, 100 Hz, 200 Hz, 300 Hz. 400 Hz, 500 Hz, 600 Hz, 700 Hz, 800 Hz, 900 Hz, , 1 kilohertz (kHz), 2 kHz, 3 kHz, 4 kHz, 5 kHz, 6 kHz, 7 kHz, 8 kHz, 9 kHz, 10 kHz, 20 kHz, 30 kHz, 40 kHz, 50 kHz, 60 kHz, 70 kHz, 80 kHz, 90 kHz, 100 kHz, or more, at most about 100 kHz, 90 kHz, 80 kHz, 70 kHz, 60 kHz, 50 kHz, 40 kHz, 30 kHz, 20 kHz, 10 kHz, 9 kHz, 8 kHz, 7 kHz, 6 kHz. 5 kHz. 4 kHz. 3 kHz, 2 kHz, 1 kHz, 900 Hz, 800 Hz, 700 Hz, 600 Hz, 500 Hz, 400 Hz, 300 Hz, 200 Hz, 100 Hz, 90 Hz, 80Hz, 70 Hz, 60 Hz, 50 Hz, 40 Hz, 30 Hz, 20 Hz, 10 Hz, or less, or a refresh rate that is within a range defined by any two of the preceding values.

[0053] In the example shown in FIG. 1, a trajectory of operations to be performed on the object by the robot or robotic operating apparatus is determined based on the plurality of positions and orientations assumed by the object at 130. In some embodiments, the trajectory of operations comprises a trajectory of positions and orientations to be assumed by the robotic operating system as it interacts with the object, a trajectory of robotic operations (such as robotic manufacturing operations, including but not limited to, tightening operations, loosening operations, cutting operations, joining operations, welding operations, riveting operations, and the like) to be performed by the robotic operating system as it interacts with the object, or a combination thereof. In some embodiments, the trajectory of operations comprises a constantspeed trajectory of operations.

[0054] In some embodiments, the object moves relative to the robot or robotic operating apparatus. For example, in some embodiments, the object moves along a conveyor belt. As another example, in some embodiments, the object and the robot or robotic operating apparatus are moved relative to one another. For instance, in some embodiments, the robot or robotic operating apparatus is moved relative to the object by a robotic transport system.

[0055] In some embodiments, although not shown in FIG. 1, the method 100 further comprises performing the trajectory of operations on the object.

[0056] FIG. 2A shows a schematic depicting an exemplary system 200 for using a landmark map associated with an object to control a robot or robotic operating apparatus.

[0057] In the example shown, the system 200 comprises a landmark map module 210. In some embodiments, the landmark map module 210 is configured to obtain a landmark map associated with an object. In some embodiments, the landmark map module 210 is configured to generate the landmark map, as described herein with respect to method 100 or operation 110 of FIG. 1.

[0058] In some embodiments, the landmark map module 210 comprises a first imaging module 211. In some embodiments, the first imaging module 211 comprises a camera. In some embodiments, the first imaging module 211 is mounted on, attached, or removably coupled to an operations module 250. In some embodiments, the operations module 250 comprises a robot arm having an end-effector, and the first imaging module 211 is mounted on, attached, or removably coupled to the end-effector (e.g. in a 3D eye-in-hand configuration). In some embodiments, the first imaging module 211 is positioned or disposed in a location some distance away from the landmark map module, in which case it is configured to observe thescene or the object from a fixed position located some distance away from the operations module 250.

[0059] In some embodiments, the first imaging module 211 is configured to obtain a plurality of images of the object from a plurality of different angles. In some embodiments, the plurality of images comprise any plurality7of images described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the plurality7of different angles comprises any plurality of different angles described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the landmark map module 210 comprises a first computer system 212. In some embodiments, the first computer system 212 is configured to form a 3D model of the object based on the plurality of images, to identify a plurality of landmarks (or key points) on the object based on the 3D model, and to collate the plurality of landmarks (or keypoints) to form the landmark map, as described herein with respect to method 100 or operation 110 of FIG. 1.

[0060] In some embodiments, the landmark map module 210 comprises a first imaging module 211. In some embodiments, the first imaging module 211 comprises a camera. In some embodiments, the first imaging module 211 is mounted on, attached, or removably coupled to an operations module 250. In some embodiments, the operations module 250 comprises a robot arm having an end-effector, and the first imaging module 211 is mounted on, attached, or removably coupled to the end-effector (e.g. in a 3D eye-in-hand configuration). In some embodiments, the first imaging module 211 is positioned or disposed in a location some distance away from the landmark map module, in which case it is configured to observe the scene or the object from a fixed position located some distance away from the operations module 250.

[0061] In some embodiments, the first imaging module 211 is configured to obtain a plurality of images of the object from a plurality7of different angles. In some embodiments, the plurality of images comprise any' plurality of images described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the plurality7of different angles comprises any plurality of different angles described herein with respect to method 100 or operation 110 of FIG. I. In some embodiments, the landmark map module 210 comprises a first computer system 212. In some embodiments, the first computer system 212 is configured to extract texture-based landmark points from each image of the plurality7of images, thereby forming a set of extracted landmark points, to perform a segmentation procedure on the set of extracted landmark points to retain only landmark points on the object, thereby forming a set of object landmark points, to extract landmark points from consecutive frames of the set of objectlandmark points, thereby forming a set of consecutive object landmark points, and to collate the plurality of consecutive object landmark points to form the landmark map, as described herein with respect to method 100 or operation 110 of FIG. 1.

[0062] In the example shown, the system 200 comprises a position determination module 220. In some embodiments, the position determination module 220 is configured to determine a plurality of positions and orientations assumed by the object over time based on the landmark map, as described herein with respect to method 100 or operation 120 of FIG. 1.

[0063] In some embodiments, the position determination module 220 comprises a second imaging module 221. In some embodiments, the second imaging module 221 comprises a camera. In some embodiments, the second imaging module 221 is mounted on, attached, or removably coupled to an operations module 250. In some embodiments, the operations module 250 comprises a robot arm having an end-effector, and the second imaging module 221 is mounted on, attached, or removably coupled to the end-effector (e.g. in a 3D eye-in-hand configuration). In some embodiments, the second imaging module 221 is positioned or disposed in a location some distance away from the landmark map module, in which case it is configured to observe the scene or the object from a fixed position located some distance away from the operations module 250.

[0064] In some embodiments, the second imaging module 221 is configured to obtain a plurality of images of the object at a plurality of points in time. In some embodiments, the plurality of images is any plurality of images described herein with respect to method 100 or operation 120 of FIG. 1. In some embodiments, the plurality of images is obtained at any frame rate described herein with respect to method 100 or operation 120 of FIG. 1. In some embodiments, the position determination module 220 comprises a second computer system 222. In some embodiments, the second computer system 222 is configured to determine a location and orientation of the object at each point in time based on the plurality of images, as described herein with respect to method 100 or operation 120 of FIG. 1.

[0065] In some embodiments, the second computer system 222 is further configured to: between successive points in time, calculating an expected position and orientation to be assumed by the object based on the location and orientation of the object at a most recent point in time, as described herein with respect to method 100 or operation 120 of FIG. 1. In some embodiments, the expected position and orientation are calculated at any refresh rate described herein with respect to method 100 or operation 120 of FIG. 1.

[0066] Accordingly, as described above, the position determination module 220 performs visual tracking of the object being processed thereby functioning as a visual tracker, visualtracking device, or 3D tracking module. As will be discussed with respect to FIGS. 10A-10C, in some embodiments, the position determination module 220 provides a current state estimate of the object as a 3D relative pose of the object with respect to the robot or robotic operating apparatus.

[0067] Although the system 200 is depicted in FIG. 2A as comprising a distinct landmark map module 210 (comprising a first imaging module 211 and first computer system 212) and a distinct position determination module 220 (comprising a second imaging module 221 and second computer system 222), the disclosure is not intended to be so limiting. For instance, in some embodiments, both the first imaging module 211 and the second imaging module 212 are contained in or comprised by a single imaging module, such as a single camera, (not shown in FIG. 2A) and the first computer system 212 and second computer system 222 are contained in or comprised by a single computer system (not shown in FIG. 2A). Thus, in some embodiments, the landmark map module 210 and the position determination module 220 may be regarded as a single element that performs both landmark map operations and position determination operations.

[0068] In the example shown, the system 200 comprises a trajectory generation module 230. In some embodiments, the trajectory generation module 230 is configured to determine a trajectory’ of operations to be performed on the object by a robot or robotic operating apparatus based on the plurality of positions and orientations assumed by the object, as described herein with respect to method 100 or operation 130 of FIG. 1. In some embodiments, the trajectory of operations comprises any trajectory of operations described herein with respect to method 100 or operation 130 of FIG. 1.

[0069] In the example shown, the system 200 comprises a conveyor belt 240. In some embodiments, the conveyor belt 240 is configured to move the object.

[0070] In the example shown, the system 200 comprises an operations module 250. In some embodiments, the operations module 250 is configured to perform the trajectory of operations on the object. In some embodiments, the operations module 250 comprises a robotic manufacturing module (such as, for example, a robot arm or other robot or robotic operating apparatus, robotic device, or machine that can move a processing device or tool, or an object being manufactured or processed). In some embodiments, the trajectory of operations comprises any trajectory7of robotic manufacturing operations described herein with respect to method 100 or operation 130 of FIG. 1.

[0071] Although the system 200 is depicted in FIG. 2A as having the landmark map module 210, the position determination module 220, and the trajectory' generation module 230 attachedor affixed to the operations module 250 (e.g., mounted on an end-effector of a robot arm in a 3D eye-in-hand configuration), the disclosure is not intended to be so limiting. In some embodiments, any 1, 2, or 3 of the landmark map module 210, the position determination module 220, and the trajectory generation module 230 are not attached or affixed to the operations module 250. For instance, in some embodiments, the landmark map module 210 (and any imaging modules included therein) are located a distance away from the operations module 250. In such embodiments, measurements made by the landmark map module 210 may be transmitted for processing by the position determination module 220 or the trajectory generation module 230, each of which may or may not be attached or affixed to the operations module 250. Similarly, the imaging module (e.g. imaging module 211) can comprise an external camera (not attached to the operations module 250). where the camera is configured to observe the scene or the object from a fixed position located some distance away from the operations module 250 or the robot or robotic operating apparatus.

[0072] For instance, as another example, FIG. 2B shows a schematic depicting an exemplary system 201 for using a landmark map associated with an object to control a robot or robotic operating apparatus. In the example shown, computation modules, in this case, the landmark map module 210, the position determination module 220, and the trajectory generation module 230, are not attached or affixed to the operations module 250 (comprising the robot or robotic operating apparatus). Rather, the operations module 250 in the embodiment shown is a separate component located a distance away from computation modules 210. 220, and 230. Such computation modules (i.e. the landmark map module 210, the position determination module 220, and the trajectory' generation module 230), as shown in FIG. 2B, can be operated or configured to operate on a separate computing device, such as an external PC or industrial controller connected (for example, via communication lines or network link 260) to the operations module 250. For instance, a wired communication link such as an Ethernet connection using a real-time communication protocol can also be used at 260. The operations module 250 can comprise a robotic manufacturing module, for example, a robot arm or other robotic device or machine or apparatus that operates or is configured to perform an operation or process on a part being processed or manufactured.

[0073] In the embodiment as shown in FIG. 2B, the operations module 250 comprises a robot arm 251 having an end-effector 252. In this case, an imaging module 253 comprising at least one imaging device, such as one or more cameras, can be mounted on the end-effector 252 of the robot arm 251 in a 3D eye-in-hand configuration as shown in FIG. 2B. A processing device or tool (not shown), such as a sealing applicator or other tool can also be mounted on, attached,or removably coupled to the end-effector 252 of the robot arm 251 (e.g. in a 3D eye-in-hand configuration), where the processing device or tool is configured to perform an operation on an object 256 being manufactured or processed by the robot or robotic operating apparatus. Examples of processing devices or tools that may be mounted on the end-effector include, but are not limited to: a sealing or adhesive dispensing nozzle, a welding torch, a painting or spray nozzle, a grinding or polishing tool, a cutting tool, a fastening or riveting tool, an inspection sensor or probe, or a pick-and-place gripper. The disclosed systems and methods not limited to any particular processing device or tool. In the disclosed techniques, visual tracking and trajectory’ control as described herein can be provided for any device or tool that needs to follow a defined path or a desired operating path relative to a moving or stationary target object.

[0074] In the example of FIG. 2B. a large object 256 is shown being moved on a conveyor belt 240 as it is being processed or operated on by the system 201. In contrast to the example of FIG. 2A, an imaging module 253 (in this case, providing the function of the first imaging module 211 and / or the second imaging module 221 as described with respect to FIG. 2A) can be mounted on, attached, or removably coupled to the end-effector 252 of the robot arm 251 (e.g. in a 3D eye-in-hand configuration), while the landmark map module 210, the position determination module 220, and the trajectory generation module 230 are not attached or affixed to the operations module 250.

[0075] As described above and as shown in FIG. 2B, the system 201, like the system 200 in FIG. 2A. comprises a landmark map module 210, a position determination module 220, and the trajectory' generation module 230. However, as shown in FIG. 2B, these computation modules 210, 220, and 230, can be operated or configured to operate on a separate computing device, such as an external PC or industrial controller. In this case, the landmark map module 210 is configured to obtain a landmark map associated with an object (e.g., the large object 256). In some embodiments, the landmark map module 210 is configured to generate the landmark map, as described herein with respect to method 100 or operation 110 of FIG. 1.

[0076] In some embodiments, as shown in FIG. 2B, the imaging module 253 comprises a first imaging module 271 (e.g.. a camera) configured to obtain a plurality’ of images of the object from a plurality of different angles. Additionally, or in the alternative, the landmark map module 210 can comprise a first imaging module 211 (e.g., a camera) configured to obtain a plurality' of images of the object from a plurality' of different angles. In such a case, the first imaging module 211, being a sub-component of the landmark map module 210, can comprise an external camera (not attached to the operations module 250), where the camera is configured to observe the scene or the object from a fixed position located some distance away from theoperations module 250 or the robotic operating apparatus. In some embodiments, the plurality of images comprise any plurality of images described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the plurality of different angles comprises any plurality of different angles described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the landmark map module 210 comprises a first computer system 212. In some embodiments, the first computer system 212 is configured to form a 3D model of the object based on the plurality of images, to identify a plurality of landmarks on the object based on the 3D model, and to collate the plurality of landmarks to form the landmark map, as described herein with respect to method 100 or operation 110 of FIG. 1.

[0077] In some embodiments, as shown in FIG. 2B, the imaging module 253 comprises a first imaging module 271 (e.g.. a camera) configured to obtain a plurality of images of the object from a plurality of different angles. Additionally, or in the alternative, the landmark map module 210 can comprise a first imaging module 211 (e.g., a camera) configured to obtain a plurality7of images of the object from a plurality of different angles. In such a case, the first imaging module 211, being a sub-component of the landmark map module 210, can comprise an external camera (not attached to the operations module 250), where the camera is configured to observe the scene or the object from a fixed position located some distance away from the operations module 250 or the robotic operating apparatus. In some embodiments, the plurality7of images comprise any plurality of images described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the plurality of different angles comprises any plurality of different angles described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the landmark map module 210 comprises a first computer system 212. In some embodiments, the first computer system 212 is configured to extract texture-based landmark points from each image of the plurality7of images, thereby forming a set of extracted landmark points, to perform a segmentation procedure on the set of extracted landmark points to retain only7landmark points on the object, thereby forming a set of object landmark points, to extract landmark points from consecutive frames of the set of object landmark points, thereby forming a set of consecutive object landmark points, and to collate the plurality of consecutive object landmark points to form the landmark map, as described herein with respect to method 100 or operation 110 of FIG. 1.

[0078] In the example shown in FIG. 2B, the system 201 comprises a position determination module 220. In some embodiments, the position determination module 220 is configured to determine a plurality of positions and orientations assumed by the object over time based on the landmark map, as described herein with respect to method 100 or operation 120 of FIG. 1.

[0079] The position determination module 220 can perform visual tracking of the object being processed thereby functioning as a visual tracker, visual tracking device, or 3D tracking module. As will be discussed with respect to FIGS. 10A-10C, in some embodiments, the position determination module 220 provides a current state estimate of the object as a 3D relative pose of the object with respect to the robot or robotic operating apparatus.

[0080] As shown in FIG. 2B. in some embodiments, the imaging module 253 comprises a second imaging module 281 (e.g. a camera) configured to obtain a plurality of images of the object from a plurality of different angles. Additionally, or in the alternative, the position determination module 220 can comprise a second imaging module 221 (e.g., a camera) configured to obtain a plurality of images of the object from a plurality of different angles. In such a case, the second imaging module 221, being a sub-component of the position determination module 220, can comprise an external camera (not attached to the operations module 250), where the camera is configured to observe the scene or the object from a fixed position located some distance away from the operations module 250 or the robotic operating apparatus. In some embodiments, the plurality of images is any plurality of images described herein with respect to method 100 or operation 120 of FIG. 1. In some embodiments, the plurality of images is obtained at any frame rate described herein with respect to method 100 or operation 120 of FIG. 1. In some embodiments, the position determination module 220 comprises a second computer system 222. In some embodiments, the second computer system 222 is configured to determine a location and orientation of the object at each point in time based on the plurality of images, as described herein with respect to method 100 or operation 120 of FIG. 1.

[0081] In some embodiments, the second computer system 222 as shown in FIG. 2B, is further configured to: between successive points in time, calculate an expected position and orientation to be assumed by the object based on the location and orientation of the object at a most recent point in time, as described herein with respect to method 100 or operation 120 of FIG. 1. In some embodiments, the expected position and orientation are calculated at any refresh rate described herein with respect to method 100 or operation 120 of FIG. 1.

[0082] Although the system 201 is depicted in FIG. 2B as comprising a distinct landmark map module 210 (that in some embodiments, comprises a first imaging system 211 or a first computer system 212) and a distinct position determination module 220 (that in some embodiments, comprises a second imaging system 221 or a second computer system 222), the disclosure is not intended to be so limiting.

[0083] For instance, in some embodiments, both the first imaging system 211 and secondimaging system 221 are contained in or comprised by a single imaging system (not shown in FIG. 2B). This single imaging system may be disposed or located on the operations module 250 or some distance away from the operations module, for example, as an external camera not attached to the operations module 2 0, where the camera is configured to observe the scene or the object from a fixed position located some distance away from the operations module 250 or the robotic operating apparatus.

[0084] Similarly, in some embodiments, both the first computer system 212 and second computer system 222 are contained in or comprised by a single computer system (not shown in FIG. 2B). Thus, in some embodiments, the landmark map module 210 and the position determination module 220 may be regarded as a single element that performs both landmark map operations and position determination operations.

[0085] In the example shown, the system 201 comprises a trajectory generation module 230. In some embodiments, the trajectory generation module 230 is configured to determine a trajectory of operations to be performed on the object by a robot or robotic operating apparatus based on the plurality of positions and orientations assumed by the object, as described herein with respect to method 100 or operation 130 of FIG. 1. In some embodiments, the trajectory of operations comprises any trajectory of operations described herein with respect to method 100 or operation 130 of FIG. 1.

[0086] In the example shown, the system 201 comprises a conveyor belt 240. In some embodiments, the conveyor belt 240 is configured to move the object 256.

[0087] In the example shown, the system 201 comprises an operations module 250. In some embodiments, the operations module 250 is configured to perform the trajectory of operations on the object. In some embodiments, the trajectory of operations comprises any trajectory of robotic manufacturing operations described herein with respect to method 100 or operation 130 of FIG. 1. In some embodiments, the operations module 250 comprises a robotic manufacturing module (such as, for example, a robot arm 251 or other robotic device or machine or robot or robotic operating apparatus that can operate on a part being manufactured). In some embodiments, the operations module 250 moves or is configured to move with respect to the object, which itself may be moving (e.g. on a conveyor belt) or stationary. For instance, the operations module 250 may be mounted on a transport system (e.g., a vehicle, mechanical device, moving mechanism, or conveyor belt) that moves or is configured to move the robot or robotic operating apparatus.

[0088] The various different modules and computer systems described herein, including but not limited to the landmark map module 210 (that in some embodiments, comprises a firstimaging module and a first computer system), the position determination module 220 (that in some embodiments, comprises a second imaging module and second computer system), the trajectory generation module 230, and the operations module 250 (that in some embodiments comprises a robotic manufacturing module) can comprise or be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. In this disclosure, the processor includes or may be implemented as one or more circuits or circuitry configured to perform the disclosed functions.Methods and systems for generating landmark maps associated with an object for use in controlling a robot or robotic operating apparatus

[0089] Provided herein are methods and systems for generating landmark maps associated with an object for use in controlling a robot or robotic operating apparatus. The methods and systems generally generate a landmark map associated with an object and transmit the landmark map for subsequent use. The subsequent use can include: determining a plurality of positions and orientations assumed by the object over time based on the landmark map, and determining a trajectory of operations to be performed on the object by a robot or robotic operating apparatus based on the plurality of positions and orientations assumed by the object.

[0090] FIG. 3 shows a flowchart depicting an exemplary method 300 for generating a landmark map associated with an object for use in controlling a robot or robotic operating apparatus.

[0091] In the example shown, a landmark map associated with an object is generated at 310.

[0092] In some embodiments, as shown for example in FIG. 3 at operation 310, the landmark map is generated by: (i) obtaining a plurality of images of the object from a plurality of different angles; (ii) forming a 3D model of the object based on the plurality of images; (iii) identifying a plurality7of landmarks (or keypoints) on the object based on the 3D model; and (iv) collating the plurality of landmarks (or key points) to form the landmark map, as described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the plurality of images comprises any plurality of images described herein with respect to method 100 or operation110 of FIG. 1. In some embodiments, the plurality of different angles comprises any plurality of different angles described herein with respect to method 100 or operation 110 of FIG. 1.

[0093] In some embodiments, as shown for example in FIG. 3 at operation 310, the landmark map is generated by: (i) obtaining a plurality of images of the object from a plurality of different angles; (ii) extracting texture-based landmark points from each image of the plurality of images, thereby forming a set of extracted landmark points; (iii) performing a segmentation procedure on the set of extracted landmark points to retain only landmark points on the object, thereby forming a set of object landmark points; (iv) extracting landmark points from consecutive frames of the set of object landmark points, thereby forming a set of consecutive object landmark points; and (v) collating the plurality of consecutive object landmark points to form the landmark map. as described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the plurality of images comprises any plurality of images described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the plurality7of different angles comprises any plurality7of different angles described herein with respect to method 100 or operation 110 of FIG. 1.

[0094] In the example of FIG. 3, the landmark map is transmitted for subsequent use at 320. In some embodiments, the subsequent use comprises: determining, based on the landmark map, a plurality7of positions and orientations assumed by the object over time, as described herein with respect to method 100 or operation 120 of FIG. 1. In some embodiments, the subsequent use comprises: determining, based on the plurality of positions and orientations assumed by the object, a trajectory of operations to be performed on the object by a robot or robotic operating apparatus, as described herein with respect to method 100 or operation 130 of FIG.1.

[0095] In some embodiments, the object moves relative to the robot or robotic operating apparatus. In some embodiments, the object moves along a conveyor belt.

[0096] FIG. 4A shows a schematic depicting an exemplary system 400 for generating a landmark map associated with an object for use in controlling a robot or robotic operating apparatus.

[0097] In the example shown, the system 400 comprises a landmark map module 410. In some embodiments, the landmark map module 410 is configured to generate a landmark map associated with an object. In some embodiments, the landmark map module 410 is configured to generate the landmark map, as described herein with respect to method 300 or operation 310 of FIG. 3.

[0098] In some embodiments, the landmark map module 410 comprises an imaging module411. In some embodiments, the imaging module 411 comprises a camera. In some embodiments, the imaging module 411 is mounted on, attached, or removably coupled to an operations module (not shown in FIG. 4A). In some embodiments, the operations module comprises a robot arm having an end-effector, and the imaging module 411 is mounted on, attached, or removably coupled to the end-effector (e.g. in a 3D eye-in-hand configuration). In some embodiments, the imaging module 411 is positioned or disposed in a location some distance away from the landmark map module, in which case it is configured to observe the scene or the object from a fixed position located some distance away from the operations module.

[0099] In some embodiments, the imaging module 411 is configured to obtain a plurality of images of the object from a plurality of different angles. In some embodiments, the plurality of images comprise any plurality of images described herein with respect to method 300 or operation 310 of FIG. 3. In some embodiments, the plurality of different angles comprises any plurality7of different angles described herein with respect to method 300 or operation 310 of FIG. 3. In some embodiments, the landmark map module 410 comprises a computer system 412. In some embodiments, the computer system 412 is configured to form a 3D model of the object based on the plurality of images, to identify a plurality of landmarks (or keypoints) on the object based on the 3D model, and to collate the plurality of landmarks (or keypoints) to form the landmark map, as described herein with respect to method 300 or operation 310 of FIG. 3.

[0100] In some embodiments, the landmark map module 410 comprises an imaging module 411. In some embodiments, the imaging module 411 comprises a camera. In some embodiments, the imaging module 411 is mounted on, attached, or removably coupled to an operations module (not shown in FIG. 4A). In some embodiments, the operations module comprises a robot arm having an end-effector, and the imaging module 411 is mounted on, attached, or removably coupled to the end-effector (e.g. in a 3D eye-in-hand configuration). In some embodiments, the imaging module 411 is positioned or disposed in a location some distance away from the landmark map module, in which case it is configured to observe the scene or the object from a fixed position located some distance away from the operations module.

[0101] In some embodiments, the imaging module 411 is configured to obtain a plurality of images of the object from a plurality of different angles. In some embodiments, the plurality of images comprise any plurality of images described herein with respect to method 100 or operation 110 of FIG. 1. In some embodiments, the plurality of different anglescomprises any plurality of different angles described herein with respect to method 300 or operation 310 of FIG. 3. In some embodiments, the landmark map module 410 comprises a computer system 412. In some embodiments, the computer system 412 is configured to extract texture-based landmark points from each image of the plurality of images, thereby forming a set of extracted landmark points, to perform a segmentation procedure on the set of extracted landmark points to retain only landmark points on the object, thereby forming a set of object landmark points, to extract landmark points from consecutive frames of the set of object landmark points, thereby forming a set of consecutive object landmark points, and to collate the plurality of consecutive object landmark points to form the landmark map, as described herein with respect to method 300 or operation 310 of FIG. 3.

[0102] In this example, the system 400 comprises a transmission module 420. In some embodiments, the transmission module 420 is configured to transmit the landmark map for subsequent use. In some embodiments, the subsequent use comprises determining a plurality of positions and orientations assumed by the object over time based on the landmark map, as described herein with respect to method 300 or operation 320 of FIG. 3. The subsequent use can further comprise determining a trajectory of operations to be performed on the object by a robot or robotic operating apparatus based on the plurality of positions and orientations assumed by the object, as described herein with respect to method 300 or operation 320 of FIG.3.

[0103] In some embodiments, the transmission module 420 comprises a file storage system 421 comprising a transmission storage device 422. In some cases, during an offline scanning or mapping phase performed to obtain or to generate a landmark map associated with an object (i.e. as described herein with respect to method 100 or operation 110 of FIG. 1), the landmark map module 210 writes the landmark map to a set of data files on the transmission storage device 422 (e.g., local disk or SSD for persistent file storage). In some embodiments, the set of data files on the transmission storage device 422 comprises: a keypoint positions file 423, a descriptors file 424, a map configuration file 425, and a keyframe file 426.

[0104] In some embodiments, the keypoint positions file 423 is configured to contain or used to store 3D information, for example, 3D coordinates (x, y, z) of each extracted landmark or keypoint (as described with respect to FIG. IB). In some embodiments, the descriptors file 424 is configured to contain or used to store visual feature descriptors associated with each landmark or keypoint. As described with respect to FIG. 1C, visual feature descriptors can be numerical representations (e.g., binary vectors or floating-point vectors) that encode a local appearance pattern surrounding each landmark or keypoint.Examples of such visual feature descriptors include ORB, SIFT, SURF, and BRIEF descriptors. These descriptors can enable matching of landmarks or keypoints across different images by comparing their numerical similarity.

[0105] In some embodiments, the map configuration file 425 is configured to contain or used to store a mapping origin position or reference coordinate frame and mapping parameters (e.g., feature extraction settings). The mapping origin position can be a single fixed reference point (coordinate frame) from which all landmark or keypoint positions in the landmark map are measured. This single fixed reference point or coordinate frame is typically defined as the imaging device's position and orientation at the start of the scanning / mapping process used to obtain or generate the landmark map (for example, as described by operations 110, 310. and 510 herein). All landmark or keypoint 3D coordinates (x. y, z) can be expressed relative to this mapping origin position. In other words, there is one mapping origin position or single fixed reference point for the entire landmark map from which all landmarks or keypoints are measured as an offset from the mapping origin position.

[0106] As an example, if the imaging module (e.g. the camera) such as imaging module 411 in FIG. 4A, begins scanning at a particular initial position over the object, that initial position corresponding to the starting camera pose becomes or designates the mapping origin position. The mapping origin position establishes the single fixed reference point such that the 3D position of every landmark or key point can be stored as an offset from the mapping origin position.

[0107] The map configuration file 425 is also configured to contain or used to store mapping parameters including, for example, feature extraction settings (e.g. number of landmarks or keypoints to extract, scale levels, patch size, detection threshold), feature matching method type, feature matching threshold, depth filtering parameters (e.g., minimum and maximum valid depth range for segmentation), camera intrinsic parameters (focal length, principal point, distortion coefficients), and image resolution settings used during scanning.

[0108] As a more specific example, a feature extraction setting for ORB features can include the maximum number of landmarks or keypoints to detect per image (e.g., 1000), the number of scale pyramid levels (e.g., 8). and the FAST comer detection threshold (e.g., 20).

[0109] In some embodiments, the keyframe file is used or configured to store or contain a list of keyframes, the landmarks or keypoints in the keyframes, and the keyframe position with respect to the mapping origin position or reference coordinate frame. In these embodiments, a keyframe is a selected image frame from a scanning sequence of image frames taken of an object that is retained as a representative snapshot of the landmark map at aparticular location. During runtime tracking (i.e., tracking during operation of the system 200 or the system 201), instead of matching a current live image against the entire landmark map (which may contain tens of thousands of landmarks or keypoints), the position determination module 220 matches against a nearest keyframe (which typically contains on the order of 1,000 keypoints). This improves computational efficiency while maintaining matching accuracy. Each keyframe stores its associated landmarks or keypoints, their associated descriptors (e.g. visual feature descriptors), and each keyframe's position and orientation relative to the mapping origin position.

[0110] In the example shown, the position determination module 220 performs the keyframe matching during runtime (e.g., where the system such as system 200 or 201 is performing robotic operations on an object). Additionally, the position determination module 220 receives the live image, matches it against the nearest keyframe, and computes the pose estimate.

[0111] FIGs. 4B-4C show an example of keyframes selected from a scanning sequence of image frames taken of an object 430. In particular. FIG. 4B shows the object 430 and a set of landmarks or keypoints at 441 and at 451. FIG. 4C shows the object 430 and a keyframe 440 that contains the keypoints 441. Also shown is a keyframe 450 that contains the keypoints 551. Other keyframes are also depicted, for example a set of keyframes is shown at 460.

[0112] In some embodiments, during runtime, the position determination module 220 comprises a map loader module (not shown) that can be used or configured to read the set of data files (e.g., the keypoints position file, the descriptors file, the map configuration file, and the keyframe file) from the transmission storage device and can load the landmark or keypoint positions (e.g., 3D information such as 3D coordinates (x, y, z) of each extracted landmark or keypoint), the visual feature descriptors associated with each landmark or keypoint, the keyframes, and the mapping origin position or reference coordinate frame into memory (i.e. working memory or RAM of the computing device that runs the position determination module 220) for use by the position determination module 220.

[0113] In other words, the map loader module (which can be a sub-component of the position determination module 220) can be used or configured to read the set of data files from the transmission storage device 422 and to load them into the computing device's working memory or RAM for computation to enable quick access by the position determination module 220 during runtime tracking. The position determination module 220 uses the landmark map, landmarks or keypoints, visual feature descriptors, and keyframes to perform matching and pose estimation. As described in further detail with respect to FIG. 10A, the trajectory generatormodule 230 (Trajectory Generator in FIG. 10A) uses pre-defined way points and generates a smooth trajectory, but does not use the landmark map.

[0114] In some embodiments, the transmission module 420 comprises the transmission storage device 422 (e.g., local disk or SSD) and a file I / O interface. The landmark map module 410 can be used or configured to write structured data files (e.g., the landmark map) to the transmission storage device 422, and the position determination module 220 can be used or configured to read the structured data files via a map loader module. In this manner, the transmission is file-based (i.e., files are being written and read or transmitted by the system) rather than network-based.

[0115] The various different modules and computer systems described herein, including but not limited to the landmark map module 410, the imaging module 411, the computer system 412, and the transmission module 420 can comprise or be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the processor described herein encompasses circuitry, including one or more circuits configured to execute the corresponding operations.

[0116] In some embodiments, the object moves relative to the robot or robotic operating apparatus. In some embodiments, the object moves along a conveyor belt.

[0117] Further disclosed herein are methods that combine all or a portion of method 100 described herein with respect to FIG. 1A and method 300 described herein with respect to FIG. 3.

[0118] Further disclosed herein are systems that combine all or a portion of system 200 described herein with respect to FIG. 2 and system 400 described herein with respect to FIG.4.

[0119] Although systems 100 and 300 described herein with respect to FIGs. 1 and 3, respectively, depict a stationary robot or robotic operating apparatus configured to track a moving object, the disclosure should not be interpreted as limited to such embodiments. For instance, in some embodiments, the robot or robotic operating apparatus moves with respect tothe object, which itself may be moving or stationary'. For instance, the robot or robotic operating apparatus (e.g., operations module 250) may be mounted on a transport system which moves the robot or robotic operating apparatus (e.g., operations module 250).Methods and systems for processing a target object with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device

[0120] Provided herein are first methods and systems for processing a target obj ect with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device. The methods and systems generally generate a landmark map of the target object based on 2D information associated with the target object and 3D information associated with the target object, generate control information for controlling the robot to process the target object with the processing device based on the landmark map and the image data, and output the control information.

[0121] FIG. 5 shows a flowchart depicting a first exemplary method 500 for processing a target object with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device.

[0122] In the example show n, a landmark map of the target object is generated based on 2D information associated with the target object and 3D information associated with the target object at 510. In some embodiments, the 2D information comprises 2D image data. In some embodiments, the 2D image data is captured by the imaging device. In some embodiments, the 3D information comprises measurement data. In some embodiments, the measurement data comprises distance information, depth information, three-dimensional coordinates, or any possible combination thereof. In some embodiments, each of the 2D information and the 3D information comprises a plurality of image data repeatedly captured by the imaging device while the imaging device is moving relative to the target object.

[0123] In some embodiments, the landmark map is generated using the 3D information and landmarks extracted from the 2D information.

[0124] In some embodiments, the landmark map is generated by deleting or invalidating landmarks of objects around the target object from the 2D information using the 3D information.

[0125] In some embodiments, the measurement data comprises stereo image data. In some embodiments, the stereo image data is captured by the imaging device. In some embodiments, the landmark map is generated using the stereo image data and landmarksextracted from the 2D information.

[0126] In some embodiments, each of the 2D information and the 3D information comprises a first set of image data and a second set of image data. In some embodiments, the first set of image data is repeatedly captured by the imaging device while the imaging device is moving relative to the target object in a first direction. In some embodiments, the second set of image data is repeatedly captured by the imaging device while the imaging device is moving relative to the target object in a second direction different from the first direction. In some embodiments, the landmark map is generated based on the first set of image data and the second set of image data.

[0127] In some embodiments, the landmark map is generated by extracting landmarks from the 2D information based on the 3D information and depth information.

[0128] In some embodiments, the plurality of image data included in the 2D information comprises a plurality of monocular image data and the plurality of image data included in the 3D information comprises a plurality of stereo image data. In some embodiments, the landmark map is generated by using the stereo image data to identify landmarks of the target object from among landmarks extracted from the monocular image data. In some embodiments, the landmark map is generated by using the stereo image data and set depth information to identify landmarks of the target object from among landmarks extracted from the monocular image data. In some embodiments, the depth information is set by a user.

[0129] In the example shown, control information for controlling the robot to process the target object with the processing device is generated based on the landmark map and the image data at 520.

[0130] In some embodiments, the control information is generated based on the landmarks of the target object extracted from the image data and the landmark map.

[0131] In some embodiments, the control information is generated to sequentially process different positions on the target object with the processing device based on each of the plurality of image data and the landmark map.

[0132] In the example shown, the control information is output for use in controlling the robot at 530.

[0133] In some embodiments, the control information further comprises an output signal. In some embodiments, the output signal is for displaying the generated landmark map of target object on a display device. In some embodiments, the output signal indicates position information of the target object from which non-extracted feature points are removed.

[0134] FIG. 6 shows a schematic depicting a first exemplary system 600 for processing a target object with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device.

[0135] In the example shown, the system 600 comprises a calculation device 610. In some embodiments, the calculation device 610 is configured to: generate a landmark map of the target object based on 2D information associated with the target object and 3D information associated with the target object; and generate control information for controlling the robot to process the target object with the processing device based on the landmark map and the image data. In some embodiments, the calculation device 610 is configured to implement any portion of method 500, operation 510, or operation 520 described herein with respect to FIG. 5.

[0136] In the example shown, the system 600 comprises an output device 620. In some embodiments, the output device 620 is configured to output the control information. In some embodiments, the output device 620 is configured to implement any portion of method 500 or operation 530 described herein with respect to FIG. 5.

[0137] Provided herein are second methods and systems for processing a target object with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device. The methods and systems generally generate an operating path along which the processing device passes based on a path condition for generating the operating path and period information based on at least one of a capturing period of the imaging device and a control period of the robot, generate control information for controlling the robot, such that the processing device moves along the operating path based on the image data, and output the control information.

[0138] FIG. 7 shows a flowchart depicting a second exemplary' method 700 for processing a target object with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device.

[0139] In the example shown, an operating path along which the processing device passes is generated at 710. In some embodiments, the operating path is generated based on a path condition and on period information. In some embodiments, the path condition is for generating the operating path. In some embodiments, the period information is based on at least one of a capturing period of the imaging device and a control period of the robot. In some embodiments, the processing device is configured to use the operating path to process the target object.

[0140] In some embodiments, the operating path is generated based on the period information and a motion result of the robot operating. Thus, in some embodiments, theprocessing devices moves based on the path condition.

[0141] In some embodiments, the operating path comprises position information or posture information of the robot or the processing device while the processing device is moving based on the path condition.

[0142] In some embodiments, the path condition comprises a path which the processing device passes over the target object. In some embodiments, the path is set by a user.

[0143] In some embodiments, the path comprises position information or posture information of the processing device or the robot at each position of the target object set by the user.

[0144] In some embodiments, the operating path has waypoints. A waypoint is a discrete reference point along the operational path, defined by a target position (e.g.. x, y, z coordinates) and orientation (e.g., roll, pitch, yaw or other rotation representations) for an endeffector of a robot or robotic operating apparatus (e.g., a robot arm) relative to a target object surface. The ordered sequence of waypoints collectively defines the path a processing device (i.e., a device that operates to perform a process on the target object) follows over the target object during the operation. In some embodiments, the processing device is mounted on, attached to, or removably coupled to the end-effector of the robot or robotic operating apparatus.

[0145] In some embodiments, each of the waypoints has position information or posture information. The waypoints can be predefined offline based on the desired operational path over the target object. In some embodiments, intervals between the plurality of waypoints are calculated such that a moving speed of the processing device over the target object is constant. In some embodiments, intervals between the plurality of way points are constant.

[0146] In some cases, the waypoints are preprocessed before runtime and stored to provide object pose references for controlling the robot or robotic operating apparatus. The way points are used by a trajectory generator as described herein to define where the robot or robotic operating apparatus should go.

[0147] In the example shown, control information for controlling the robot is generated at 720. In some embodiments, the control information causes the processing device to move along the operating path based on the image data.

[0148] In some embodiments, while the target object is moving, the processing device performs the processing on the target object while moving along the operating path based on the control information.

[0149] In the example shown, the control information is output for use in controllingthe robot at 730.

[0150] FIG. 8 shows a schematic depicting a second exemplary system 800 for processing a target object with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device.

[0151] In the example shown, the system 800 comprises a calculation device 810. In some embodiments, the calculation device 810 is configured to: generate an operating path along which the processing device passes based on a path condition for generating the operating path and period information based on at least one of a capturing period of the imaging device and a control period of the robot; and generate control information for controlling the robot, such that the processing device moves along the operating path based on the image data. In some embodiments, the calculation device 810 is configured to implement any portion of method 700, operation 710, or operation 720 described herein with respect to FIG. 7.

[0152] In the example shown, the system 800 comprises an output device 820. In some embodiments, the output device 820 is configured to output the control information. In some embodiments, the output device 820 is configured to implement any portion of method 700 or operation 730 described herein with respect to FIG. 7.

[0153] Further disclosed herein are methods that combine all or a portion of method 500 described herein with respect to FIG. 5 and method 700 described herein with respect to FIG. 7.

[0154] Further disclosed herein are systems that combine all or a portion of system 600 described herein with respect to FIG. 6 and system 800 described herein with respect to FIG.8.

[0155] The various systems and devices described herein with respect to FIGs. 6 and 8 for example and including but not limited to the system 600, the calculation device 610, the output device 620, the system 800, the calculation device 810, the output device 820, and the processing device (such as the processing device attached to the robot based on image data generated by capturing one or more images of the target object with an imaging device) can comprise or be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardw are components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combinationof a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the processor described herein encompasses circuitry, including one or more circuits configured to execute the corresponding operations.Computer systems

[0156] Additionally, systems are disclosed that can be used to perform any of the methods 100, 300, 500, and 700, respectively, of FIGs. 1, 3, 5, and 7, respectively, or any of operations 110, 120, 130, 310, 320, 510, 520, 530, 710, 720, and 730 described herein. In some embodiments, the systems comprise one or more processors and memory coupled to the one or more processors. In some embodiments, the one or more processors are configured to implement one or more operations of method 100, 300, 500, and 700, such as any of operations 110, 120, 130, 310, 320, 510, 520, 530, 710, 720, and 730. In some embodiments, the memory is configured to provide the one or more processors with instructions corresponding to the operations. In some embodiments, the instructions are embodied in a tangible computer readable storage medium

[0157] In some embodiments, the first computer system included in the landmark map module 210 is implemented by a computer system such as the computer system 900 shown in FIG. 9, including one or more processors and memory storing processor-executable program code. When executed by the one or more processors, the program code causes the first computer system to form a 3D model, identify landmarks, and generate the landmark map as described herein.

[0158] In some embodiments, the second computer system included in the position determination module 220 is implemented by a computer system such as the computer system 900 shown in FIG. 9. The second computer system executes processor-executable program code stored in memory7to determine a position and orientation of the object over time based on image data and the landmark map.

[0159] In some embodiments, the calculation device 610 as shown in FIG. 6 and the calculation device 810 as shown in FIG. 8 are implemented by one or more processors, such as the microprocessor subsystem 901 shown in FIG. 9, in combination with memory 904 that stores processor-executable instructions embodied in a tangible computer-readable storage medium. The instructions comprise program code executable by the one or more processors. When executed by the one or more processors, the program code causes the processors to perform operations including generating the operating path and generating the controlinformation, as described herein with respect to FIGs. 5 and 7.

[0160] In some embodiments, the output device 620 as shown in FIG. 6 and the output device 820 as shown in FIG. 8 are implemented by the microprocessor subsystem 901 executing processor-executable program code stored in memory 904, together with an output interface for transmitting control information to an external robot control system. The output device outputs the control information generated by the calculation device to control movement of the robot.

[0161] FIG. 9 is a block diagram of an exemplary computer system 900 used in some embodiments to perform portions of methods described herein. In some embodiments, the computer system may be utilized as a component in systems described herein. FIG. 9 illustrates one embodiment of a general-purpose computer system. Other computer system architectures and configurations can be used for carrying out the processing of the present inventions. Computer system 900, made up of various subsystems described below, includes at least one microprocessor subsystem 901. In some embodiments, the microprocessor subsystem comprises at least one central processing unit (CPU) or graphical processing unit (GPU). The microprocessor subsystem can be implemented by a single-chip processor or by multiple processors. In some embodiments, the microprocessor subsystem is a general-purpose digital processor which controls the operation of the computer system 900. Using instructions retrieved from memory 904, the microprocessor subsystem controls the reception and manipulation of input data, and the output and display of data on output devices.

[0162] The microprocessor subsystem 901 is coupled bi-directionally with memory 904, which can include a first primary storage, typically a random-access memory (RAM), and a second primary storage area, typically a read-only memory (ROM). As is well known in the art, primary’ storage can be used as a general storage area and as scratch-pad memory, and can also be used to store input data and processed data. It can also store programming instructions and data, in the form of data objects and text objects, in addition to other data and instructions for processes operating on microprocessor subsystem. Also, as well known in the art, primary storage typically includes basic operating instructions, program code, data and objects used by the microprocessor subsystem to perform its functions. Primary storage devices 904 may include any suitable computer-readable storage media, described below, depending on whether, for example, data access needs to be bi-directional or uni-directional. The microprocessor subsystem 901 can also directly and very' rapidly retrieve and store frequently needed data in a cache memory (not shown).

[0163] A removable mass storage device 905 provides additional data storage capacityfor the computer system 900, and is coupled either bi-directionally (read / write) or unidirectionally (read only) to microprocessor subsystem 901. Storage 905 may also include computer-readable media such as magnetic tape, flash memory, signals embodied on a carrier wave, PC-CARDS, portable mass storage devices, holographic storage devices, and other storage devices. A fixed mass storage 909 can also provide additional data storage capacity. The most common example of mass storage 909 is a hard disk drive. Mass storage 905 and 909 generally store additional programming instructions, data, and the like that typically are not in active use by the processing subsystem. It will be appreciated that the information retained within mass storage 905 and 909 may be incorporated, if needed, in standard fashion as part of primary storage 904 (e.g. RAM) as virtual memory.

[0164] In addition to providing processing subsystem 901 access to storage subsystems, bus 906 can be used to provide access other subsystems and devices as well. In the described embodiment, these can include a display monitor 908, a network interface 907, a keyboard 902, and a pointing device 903, as well as an auxiliary input / output device interface, a sound card, speakers, and other subsystems as needed. The pointing device 903 may be a mouse, stylus, track ball, or tablet, and is useful for interacting with a graphical user interface.

[0165] The network interface 907 allows the processing subsystem 901 to be coupled to another computer, computer network, or telecommunications network using a network connection as shown. Through the network interface 907, it is contemplated that the processing subsystem 901 might receive information, e.g., data objects or program instructions, from another network, or might output information to another network in the course of performing the above-described method steps. Information, often represented as a sequence of instructions to be executed on a processing subsystem, may be received from and outputted to another network, for example, in the form of a computer data signal embodied in a carrier wave. An interface card or similar device and appropriate software implemented by processing subsystem 901 can be used to connect the computer system 900 to an external network and transfer data according to standard protocols. That is, method embodiments of the present inventions may execute solely upon processing subsystem 901, or may be performed across a network such as the Internet, intranet networks, or local area networks, in conjunction with a remote processing subsystem that shares a portion of the processing. Additional mass storage devices (not shown) may also be connected to processing subsystem 901 through network interface 907.

[0166] An auxiliary' I / O device interface (not shown) can be used in conjunction with computer system 900. The auxiliary I / O device interface can include general and customized interfaces that allow the processing subsystem 901 to send and, more ty pically, receive datafrom other devices such as microphones, touch-sensitive displays, transducer card readers, tape readers, voice or handwriting recognizers, biometrics readers, cameras, portable mass storage devices, and other computers.

[0167] In addition, embodiments of the present inventions further relate to computer storage products with a computer readable medium that contains program code for performing various computer-implemented operations. The computer-readable medium is any data storage device that can store data which can thereafter be read by a computer system. The media and program code may be those specially designed and constructed for the purposes of the present inventions, or they may be of the kind well known to those of ordinary skill in the computer software arts. Examples of computer-readable media include, but are not limited to, all the media mentioned above: magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as CD-ROM disks; magneto-optical media such as floptical disks; and specially configured hardware devices such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), and ROM and RAM devices. The computer-readable medium can also be distributed as a data signal embodied in a carrier wave over a network of coupled computer systems so that the computer-readable code is stored and executed in a distributed fashion. Examples of program code include both machine code, as produced, for example, by a compiler, or files containing higher level code that may be executed using an interpreter. The computer system shown in FIG. 9 is but an example of a computer system suitable for use with the inventions. Other computer systems suitable for use with the inventions may include additional or fewer subsystems. In addition, bus 906 is illustrative of any interconnection scheme serving to link the subsystems. Other computer architectures having different configurations of subsystems may also be utilized.Overview of Data Flow Between Modules

[0168] FIG. 10A show s a schematic overview of the data flow betw een modules in the systems described herein for using landmark maps associated with an object to control a robot or robotic operating apparatus. The systems are used and configured to obtain a landmark map associated with an object, determine a plurality of positions and orientations assumed by the object over time based on the landmark map, determine a trajectory of operations to be performed on the object by the robot or robotic operating apparatus based on the plurality of positions and orientations assumed by the object over time, and provide output to a controller to execute the trajectory of operations.

[0169] As depicted in FIG. 10 A, modules described herein are used or configured toperform four main phases or steps shown as operational blocks at 1010 (Scanning Phase), 1020 (Visual Tracking Phase). 1030 (Trajectory Generating Phase) and 1040 (Control Phase).

[0170] First, in a Scanning Phase shown at 1010, a landmark map module (e.g., landmark map module 210 as described with respect to FIGS. 2A-2B or landmark map module 410 of FIG. 4 A) is configured to perform a scanning step, operating offline (i.e. prior to runtime) to obtain or generate a landmark map of an object at least in part by scanning the object. In some embodiments, a robot or robotic operating apparatus (e.g. operations module 250 as described with respect to FIGS. 2A-2B comprising a robot arm having an end-effector) moves an imaging device comprising a camera (e.g., first imaging device 211) that can be mounted on the robot arm end-effector along a predefined path over a stationary' reference instance of the object. At each position along the predefined path, the landmark map module scans the object at 1010 by using the imaging device to capture image data 1011. The image data 1011 captured by the imaging device can comprise color / monochrome images (2D information) and depth images (3D information). Additionally, or in the alternative, the scanning function can be done by capturing images and generating the map or directly from CAD or from a 3D model as shown at 1012.

[0171] In the example shown, visual features such as landmarks or keypoints and their associated visual feature descriptors can be extracted from the images at 1013. Additionally, 3D coordinates (x, y, z) of each keypoint can be determined using corresponding depth information from the depth image (3D information). A segmentation procedure can also be performed at 1013 using the depth information to remove landmarks belonging to background objects thereby retaining only the landmarks or key points located on the object surface. The retained landmarks or keypoints, their respective 3D positions, their respective visual feature descriptors, the mapping origin position, keyframes, and mapping parameters are stored as files constituting the landmark map at 1014 for subsequent use during runtime.

[0172] During runtime, in a Visual Tracking Phase shown at 1020, a second module for tracking and determining object position, shown here as 3D Visual Tracking Module 1021 (described as a position determination module 220 with respect to FIGS. 2A-2B) is configured to perform a visual tracking step, by visually tracking the position of a target object (e.g., object 256 on a conveyor belt 240 in FIG. 2B). This 3D Visual Tracking Module or position determination module determines where the target object currently is by matching live images against the landmark map.

[0173] Here, visual tracking of the object is performed using an imaging device (e.g., imaging device 253 shown in FIG. 2B mounted on the end-effector of the robot arm or roboticoperating apparatus 1060) to capture image data of a portion of the object as the object moves (e.g., object 256 on a conveyor belt 240 in FIG. 2B). The captured image data, shown at 1051, can comprise a color / monochrome image and a depth image. An image pre-process block at 1023 can receive the color and depth images 1051, and can perform segmentation to remove background information, retaining only the portion of the image data corresponding to the target object surface, as shown at 1052. Visual feature keypoints and their associated descriptors are then extracted from the segmented image data 1052 by the 3D Tracking Module 1021.

[0174] As shown in FIG. 10A, the 3D Tracking Module 1021 receives the landmark map 1014 (generated by the landmark map module in the scanning block 1010) and the segmented image data 1052 as output from the image pre-process block 1023, and matches the extracted keypoints from the segmented image data 1052 against the keyframes and landmarks in the landmark map to determine the position and orientation (pose) of the target object relative to the imaging device. Additionally, keypoints can be matched across consecutive image frames to track relative motion between frames. The images can be continuously streamed from the camera (for example at 100 frames per second), and two consecutive frames in time can be matched.

[0175] To refine the orientation or pose estimate and improve robustness to image noise, the 3D Tracking Module 1021 can receive a state of the robot arm or robotic operating apparatus 1060 (e.g., obtained from joint encoder data of the robot arm via end-effector pose data obtained from the robotic operating apparatus controller).

[0176] In some embodiments, the imaging device is mounted on an end-effector of a robot arm or robot or robotic operating apparatus (shown as an eye-in-hand configuration in FIG. 2B). In these embodiments, the observed visual motion of the object is a combination of the end-effector’s or the robot’s or the robotic operating apparatus's own motion and the target object's motion. As described in more detail with respect to FIGS. 10B-10C, the Inertial State 1054 along with a Measured state 1053 of the robot arm or the robotic operating apparatus 1060 can be used as a constraint in a pose optimization step to decouple the end-effector's or the robot’s or the robotic operating apparatus's self-motion from the target object's motion, thereby isolating the true pose change of the target object (i.e. the Relative Pose 1022). In the example show n, the resulting output of the 3D Tracking Module 1021 is a refined 3D pose estimate (the Relative Pose 1022) of the target object relative to the end-effector or robot or robotic operating apparatus 1060.

[0177] Returning to FIG. 10A, in a third step Traj ectory Generator or Generating Phaseat 1030, a Trajectory Generator (described as a trajectory generation module 230 with respect to FIGS. 2A-2B) generates a trajectory for the object by receiving a predefined operational path (e.g., a sealing path defined over the target object surface) to generate a sequence of Waypoints 1033 along that path. This ordered sequence of waypoints collectively defines the path the processing device follows over the target object during the operation (i.e. where the robot or robotic operating apparatus should go to perform the operation on the target object).

[0178] Each waypoint comprises position and orientation information for the robot’s or robotic operating apparatus's end-effector. The intervals between waypoints are calculated based on a required operational speed of the processing device (e.g., constant application speed for sealing) and a control period of the robot or robotic operating apparatus. The waypoints are preprocessed prior to runtime and stored as desired pose references for the Motion Controller 1041 shown in the Control Phase at 1050.

[0179] In a fourth step shown as a Control Phase 1050, a Motion Controller 1041, is used and configured to receive: (i) the current pose estimate (Relative Pose 1022) of the target object from the 3D Tracking Module 1021. (ii) the desired pose (Way points 1033) from the Trajectory Generator 1030, and (iii) a velocity estimate of the target object derived from the 3D Tracking Module 1021. The Motion Controller 1041 can comprise a proportional -integral (PI) controller (as shown in and described with respect to FIG. 11) with a lead filter and a feedforward component. The Motion Controller 1041 is used and configured to compute a position error between the desired pose 1033 and the current pose estimate 1022 and can generate a control signal 1059 for the robot or robotic operating apparatus 1060. In the example shown, the feedforward component uses the estimated velocity of the target object to compensate for latency in the image capture and processing pipeline, thereby reducing tracking lag. The feedforward component can also compensate for continued target object motion during the computation and actuation delay — i.e., the target object continues to move during the time the system (e.g., system 200 or 201 of FIGS. 2A-2B) captures an image, computes the pose, and executes the correction. As shown in FIG. 10 A, the Motion Controller 1041 generates a control signal 1059 which is used by the operations module (e.g. operations module 250) and output to the robot or robotic operating apparatus 1060 to adjust its motion such that the processing device (e.g., mounted at the end-effector) follows the target object along the desired operational path.

[0180] As described above and in particular, in an eye-in-hand configuration where an imaging device is mounted on an end-effector or a robot arm or robot or robotic operating apparatus or where the robot or robotic operating apparatus itself is moving (e.g., where it ismounted on a robotic transport system or other mechanical device), it becomes crucial to decouple the robot's or the robotic operating apparatus' self-motion from the target object's motion. An approach to addressing this problem is described in more detail below.

[0181] In an eye-in-hand configuration, or additionally or in the alternative, in the case where the robot or robot operating apparatus is mounted on a robotic transport system that moves the robot or robot operating apparatus, both the robot or robotic operating apparatus and the target object (e.g., object 256 on a conveyor belt 240 in FIG. 2B) are moving simultaneously. Standard visual SLAM (e.g., ORB-SLAM3) cannot separate robot-induced camera motion from object motion, which can result in large errors (e.g., ~23mm error). Thus, in the case where the imaging module or camera is mounted on or attached or coupled to the end-effector of the robot or robotic operating apparatus, the imaging module only sees relative motion.

[0182] In order to address this issue, a method for kinematic pose fusion is disclosed for using kinematic state data from the robot or robotic operating apparatus to constrain or correct the visual pose estimation of the object. In some embodiments, the method uses the robot’s or robotic operating apparatus’s own joint encoder data to decouple the robot’s or robotic operating apparatus’s self-motion from the target object's during a pose estimation operation. In some embodiments, kinematic pose fusion estimation provides the current realtime pose of the target object during runtime, which the Motion Controller 1041 compares against the Waypoints 1033 to compute tracking error and generate control signals (e.g. control signal 1059). The waypoints define where the robot should go and the kinematic pose fusion estimate tells the controller where the target object currently is. The control signal can be used by the operations module (e.g. operations module 250), which receives and executes the control signal to operate the robot or robotic operating apparatus 1060.

[0183] Referringto FIG. 10A, during the Visual Tracking Phase shown at 1020, the 3D Tracking Module 1021 performs kinetic pose fusion by combining two sources of information: (1) visual pose information derived from the Color and Depth image 1051, which is preprocessed by the Image Pre-process block 1023. and (2) kinematic state data from the robotic operating apparatus 1060, shown as the Measured state 1053 (comprising for example, a current end-effector pose and a current velocity associated with the robot or robot operating apparatus 1060). The velocity component of the Measured state 1053 is processed by the Filter 1025 to produce or generate the Inertial State 1054, a fdtered (smoothed) velocity estimate. The end-effector pose component in the Measured state 1053 is not modified by the Filter 1025.

[0184] As noted, in this example, Filter 1025 is applied to the velocity component ofthe current state shown as Measured State 1053, and not the pose component. This is because the end-effector pose as reported by the robot or robotic operating apparatus is considered to be sufficiently accurate and therefore does not require filtering. However, the velocity component in this case as reported by the robot or robotic operating apparatus is noisy, requiring Filter 1025 for smoothing to produce a reliable velocity7estimate that more accurately represents the current state of velocity associated with the robot or robotic operating apparatus or the end-effector. Accordingly, at least in this case, the Inertial State 1054 comprises the unfiltered pose and the filtered (smoothed) velocity.

[0185] In other embodiments or in cases where the current state comprises other components, such components may be used directly or may be pre-processed or filtered accordingly depending on the nature of the component. The disclosed techniques are not limited to the specific example as described in this case for the components in Measured State 1053, which is shown here for illustrative purposes. The method 1090 (comprising steps 1091-1094) can be performed using other metrics to determine or represent the current state associated with a robot or robotic operating apparatus without impacting or limiting the scope of what is claimed herein.

[0186] In the example shown, the 3D Tracking Module 1021 receives both the visual information and the kinematic state data and fuses them to produce or generate an estimated pose (shown as Relative Pose 1022) of the target object relative to the robotic operating apparatus 1060. The Relative Pose 1022 is then provided to the Motion Controller 1041, along with the Waypoints 1033 from the Trajectory Generator 1030 and the Inertial State 1054 (comprising the fdtered velocity' for feedforward), to generate control signals for the robotic operating apparatus 1060.

[0187] Two exemplary methods for performing kinematic pose fusion within the 3D Tracking Module 1021 are described for a constraint-based pose fusion method and a subtraction-based pose fusion method respectively with reference to FIGs. 10B and 10C below. These methods can be employed to isolate the target object’s true motion from the motion of the robot or robotic operating apparatus.

[0188] FIG. 10B shows an exemplary method 1080 for kinematic pose fusion that can be performed as part of the Visual Tracking Phase 1020 of FIG. 10A. In some embodiments, the method can be used for correcting a visual pose estimation of an object being tracked by a robot or robotic operating apparatus. This method uses the robot’s or robotic operating apparatus’ motion as a constraint in the pose optimization (e.g.. within a bundle adjustment or optimization solver.)

[0189] At 1081 , during a pose estimation operation on a moving obj ect, kinematic state data is obtained from the robotic operating apparatus 1060. In some embodiments, the kinematic state data is the Measured state 1053 shown in FIG. 10A comprising end-effector pose and velocity as reported by the robotic operating apparatus 1060 (e.g., more specifically by a robotic operating apparatus controller, which is a sub-component of the robotic operating apparatus 1060). A velocity component of the Measured state 1053 can be processed by the Filter 1025 to produce or generate an Inertial State 1054, comprising a filtered (smoothed) velocity estimate along with the unfiltered pose estimate (provided by Measured state 1053). Both the Measured state 1053 and the Inertial State 1054 can be provided as inputs to the 3D Tracking Module 1021.

[0190] At 1082, the method comprises computing a contribution of a motion of the robotic operating apparatus to a total apparent motion observed in image data generated by capturing one or more images of the object with an imaging device. In some embodiments, a tracking module such as 3D Tracking Module 1021 (described as a position determination module 220 with respect to FIGS. 2A-2B) computes a contribution of a motion of the robotic operating apparatus 1060 to a total apparent motion observed in image data generated by capturing one or more images of the object with an imaging device as described herein for example with respect to systems 200, 201 and 400. If the imaging device moves with the robotic operating apparatus, motion of the robotic operating apparatus causes apparent motion in the captured images that is independent of any motion of the target object.

[0191] At 1083, the method comprises using the computed contribution as a constraint in the pose estimation operation to isolate a motion of the object from the motion of the robotic operating apparatus. In some embodiments, a tracking module such as 3D Tracking Module 1021 (described as a position determination module 220 with respect to FIGS. 2A-2B) uses the computed contribution as a constraint in the pose estimation operation to isolate a motion of the object from the motion of the robotic operating apparatus 1060.

[0192] Method 1080 can be used to provide an estimated pose of the target object relative to the robotic operating apparatus. In the example shown, the output of method 1080 is the Relative Pose 1022 shown in FIG. 10A, which provides the estimated pose of the target object relative to the robotic operating apparatus 1060. The Inertial State 1054 (filtered velocity) can also be provided to the Motion Controller 1041 in the Control Phase 1050 for use in the feedforward component of the control signal 1059.

[0193] FIG. 10C shows an alternative exemplary method 1090 for kinematic pose fusion that can be performed as part of the Visual Tracking Phase 1020 of FIG. 10A. In someembodiments, the method can be used for correcting a visual pose estimation of an object being tracked by a robot or robotic operating apparatus. This method explicitly subtracts the robot's or robotic operating apparatus’ pose change from the total apparent visual motion.

[0194] In the example shown, method 1090 (comprising steps 1091-1094) is performed at each image capture time step — that is, each time the imaging module or imaging device (e.g. a camera) captures a new image of the target object.

[0195] In the example shown, and as used herein, a time step refers to an image capture time step. In other words, a time step as used herein is each time the imaging module or imaging device (e.g.. a camera) captures a new color and / or depth image. The robot or robotic operating apparatus 1060 reports or is configured to report its state (e.g. a state associated with the robot or robotic operating apparatus 1060, such as its end-effector pose and velocity) at each such time step.

[0196] The method 1090 comprises, at each time step, reporting a measured state comprising a current pose and current velocity of a robotic operating apparatus at 1091. In some embodiments, the robotic operating apparatus 1060 reports a measured state (e.g., Measured state 1053, shown in FIG. 10 A) comprising a current pose and current velocity of the robot or robot operating apparatus (e.g. robot operating apparatus 1060). In some embodiments, the measured state is associated with a moving component of the robot or robot operating apparatus, such as a robot arm or end-effector. For example, in the case where the robot operating apparatus comprises or uses an end-effector, or where the imaging module is mounted or attached to an end-effector, the measured state comprises a current end-effector pose and current end-effector velocity.

[0197] In some embodiments, the robot or robotic operating apparatus 1060 (e.g., the robot arm) itself reports or is configured to report its current state, in this case. Measured State 1053, via its internal joint encoders. The Measured State 1053 can comprise a current endeffector pose (e.g., the location of the imaging module or the camera in space) and a current velocity associated with the robot or robot operating apparatus 1060.

[0198] The method 1090 further comprises, computing between consecutive time steps, a change in pose based on the reported measured state at 1092. In some embodiments, a tracking module such as 3D Tracking Module 1021 (described as a position determination module 220 with respect to FIGS. 2A-2B) computes a change in pose (e.g. robot pose, robotic operating apparatus pose, or end-effector pose) between consecutive time steps based on the reported measured state (e.g. current pose) from the measured state.

[0199] In some embodiments, the method can include computing or determining a change in the current state associated with the robot or robotic operating apparatus between consecutive time steps, which includes computing or determining a change in an end-effector pose (i.e., how much the robot or robotic operating apparatus has moved the imaging module or the camera between consecutive time steps) and / or a change in a velocity' associated with movement of the robot or robotic operating apparatus. In the example of FIG. 10A, the 3D Visual Tracking Module 1021 (described as a position determination module 220 with respect to FIGS. 2A-2B) can be used or configured to compute or determine a change in the current state associated with the robot or robot operating apparatus (e.g., a change in the end-effector pose and / or velocity) between consecutive time steps.

[0200] In some embodiments, to compute or determine a change in the current state associated with the robot or robot operating apparatus, 3D Visual Tracking Module 1021 receives Inertial State 1054 (comprising a measured end-effector pose reported by the robot or by the robotic operating apparatus and a filtered velocity associated with the end-effector, robot, or robotic operating apparatus) from the robot or robotic operating apparatus 1060 at each image capture time step. 3D Visual Tracking Module 1021 then computes a difference between the current and previously reported poses (as reported by the robot or robotic operating apparatus 1060) to determine how much the robot or robot operating apparatus 1060 has moved the imaging module (e.g., the camera) between frames (i.e., images taken at consecutive time steps). 3D Visual Tracking Module 1021 also computes a difference between the current and previously reported filtered velocities associated with the movement of the end-effector, robot, or robotic operating apparatus. In this manner, 3D Visual Tracking Module 1021 is able to compute or determine a change in the current state associated with the robot or robotic operating apparatus.

[0201] Returning to FIG. 10C, the method 1090 further comprises, estimating a total apparent motion from image data generated by capturing one or more images of an object with an imaging device at 1093. In the example shown, the total apparent motion includes motion attributed to the robot or robotic operating apparatus moving the imaging module or the camera and motion of the object (e.g. an object moving on a conveyor belt). In some embodiments, estimating the total apparent motion from the image data is performed simultaneously or at about the same time as step 1092 (i.e., computing or determining a change in the current state associated with the robot or robotic operating apparatus 1060).

[0202] In some embodiments, a tracking module such as 3D Tracking Module 1021 (described as a position determination module 220 with respect to FIGS. 2A-2B) estimates atotal apparent motion from processed image data (e.g. shown at 1052 in FIG. 10A). In some embodiments, image data (e.g. Color and Depth Image 1051) provided by an imaging module as described herein is processed (e.g. by Image Pre-process 1023) to generate processed image data 1052. The 3D Tracking Module 1021 then receives the processed image data 1052 (i.e., Color and Depth Image 1051 processed by the Image Pre-process 1023) and estimates the total apparent motion by matching keypoints against the landmark map and across consecutive frames, wherein the total apparent motion includes motion attributed to the robot or robotic operating apparatus moving the imaging device and motion of the object. The total apparent motion is thus an intermediate result generated within the 3D Tracking Module 1021 — it is estimated from the image data. A contribution of motion attributed to the robot or robotic operating apparatus is subtracted in a next step, and the result representing the true motion of the object is output as the Relative Pose 1022.

[0203] Returning to FIG. 10C, the method 1090 further comprises, subtracting the computed change in pose from the estimated total apparent motion to provide a difference representing the true motion of the object at 1094.

[0204] In some embodiments, a tracking module such as 3D Tracking Module 1021 (described as aposition determination module 220 with respect to FIGS. 2A-2B) subtracts the computed change in pose (e.g. robot pose, robotic operating apparatus pose, or end-effector pose) from the estimated total apparent motion to provide or generate a difference representing the true motion of the object.

[0205] The result is the output shown in FIG. 10A as the Relative Pose 1022, which can be provided to the Motion Controller 1041 for generating control signals (e.g. control signal 1059) to control motion of the robot or robot operating apparatus 1060. The Inertial State 1054 (fdtered velocity component) is also provided to the Motion Controller 1041 for use in the feedforward component of the control signal 1059.

[0206] In some embodiments, between image capture time steps, the Motion Controller 1041 continues to update the motion of the robot or robotic operating apparatus at a faster robot control rate (e.g., 100-1000 Hz) using a most recent estimate of Relative Pose 1022 and Inertial State 1054. This corresponds to the interpolation / prediction processes as described herein (i.e. calculating an expected position and orientation between successive image frames, as shown for example in FIG. 1 A at operation 120 and as described with respect to the second computer system 222 of FIGS. 2A-2B).

[0207] In some embodiments, a system executing the method 1090 (and performing steps 1091-1094) comprises a robot or a robotic operating apparatus configured to report acurrent state or a measured state (comprising, for example, current end-effector pose and current velocity’)- The current state or measured state can be provided by the robot’s or robotic operating apparatus’ joint encoders. The system further comprises a 3D tracking module or position determination module configured to: compute the change in the current state (e.g., the end-effector pose) between consecutive time steps; estimate a total apparent motion from image data; and subtract the change in the current state (e.g. the change in end-effector pose) from the total apparent motion to isolate the target object’s true motion, outputting a relative pose. The image data (e.g. Color and Depth Image 1051) can be provided by an imaging module as described herein and processed (e.g. by Image Pre-process 1023) to generate processed image data that may be used by the system to estimate a total apparent motion. The system further comprises a calculation device or a motion controller configured to receive the relative pose and to generate a control signal for the robot or robotic operating apparatus. The system may also use waypoints and an inertial state along with the relative pose to generate the control signal.EXAMPLES

[0208] FIG. 11 shows an exemplary proportional integral controller 1100 with lead filter and feedforward components for use with the methods and systems described herein. In some embodiments, the desired trajectory path may be tracked using proportional-integral (PI)-lead-feedforward controller 1100 as shown in FIG. 11. The current state estimate of the object may be provided by a visual tracker (e.g. position determination module 220 or 3D Visual Tracking Module 1021) as the 3D relative pose of the object (e.g., Relative Pose 1022) with respect to the robot or robotic operating apparatus. In this case, the desired object pose can be provided by the trajectory' waypoints (e.g. Waypoints 1033). The feedforward component may permit consistent tracking of the trajectory even under dynamic motion of the object in relation to the robot or robotic operating apparatus.

[0209] FIG. 12 shows an exemplary' plot of the pose estimate obtained with a system using the proportional integral controller of FIG. 11 compared against a desired trajectory'. Briefly, the systems described herein were evaluated in various lighting conditions and with different 3D vision sensors to track moving parts at various speeds from 0 millimeters per second (mm / s) to 30 mm / s. The accuracy and robustness of the dynamic part tracking were assessed by comparing the estimated positions with ground truth data. As shown in FIG. 12, the estimated poses from the system closely follow the desired trajectory with high accuracy.

[0210] FIG. 13 shows another exemplary plot of the pose estimate obtained with a system using the proportional integral controller of FIG. 11 compared against a desiredtrajectory . As in the example shown in FIG. 12, the systems described herein were evaluated in various lighting conditions and with different 3D vision sensors to track moving parts at various speeds from 0 millimeters per second (mm / s) to 30 mm / s. The accuracy and robustness of the dynamic part tracking were assessed by comparing the estimated positions with ground truth data. As show n in a magnified version of a portion of the plot in FIG. 13, the estimated poses from the system closely follow the desired trajectory with high accuracy.Conclusion and Terminology

[0211] Conditional language used herein, such as, among others, “can,” “could,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements or states. Thus, such conditional language is not generally intended to imply that features, elements or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list. Further, the term “each,” as used herein, in addition to having its ordinary' meaning, can mean any subset of a set of elements to which the term “each” is applied.

[0212] Disjunctive language such as the phrase “at least one of X, Y and Z,” unless specifically stated otherwise, is to be understood with the context as used in general to convey that an item, term, etc. may be either X, Y, or Z, or a combination thereof. Thus, such conjunctive language is not generally intended to imply that certain embodiments require at least one of X, at least one of Y and at least one of Z to each be present.

[0213] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items. Accordingly, phrases such as “a device configured to” are intended to include one or more recited devices. Such one or more recited devices can also be collectively configured to carry out the stated recitations. For example, “a processor configured to cany’ out recitations A, B and C” can include a first processorconfigured to carry out recitation A working in conjunction with a second processor configured to carry out recitations B and C.

[0214] Depending on the embodiment, certain acts, events, or functions of any of the methods described herein can be performed in a different sequence, can be added, merged, or left out altogether (for example, not all described acts or events are necessary7for the practice of the method). Moreover, in certain embodiments, acts or events can be performed concurrently, for example, through multi -threaded processing, interrupt processing, or multiple processors or processor cores, rather than sequentially.

[0215] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure.

[0216] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality7of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the processor described herein encompasses circuitry, including one or more circuits configured to execute the corresponding operations.

[0217] The blocks of the methods and algorithms described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory7, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, or any other form of computer-readable storage medium known in theart. An exemplary storage medium is coupled to a processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC.

[0218] While the above detailed description has shown, described, and pointed out novel features as applied to various embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others.

Claims

1. CLAIMS1. A method comprising:(a) obtaining a landmark map associated with an object:(b) determining, based on the landmark map, a plurality of positions and orientations assumed by the object over time; and(c) determining, based on the plurality of positions and orientations assumed by the object, a trajectory of operations to be performed on the object by a robotic operating apparatus.

2. The method of claim 1 , wherein (a) comprises:(i) obtaining a plurality of images of the object from a plurality of different angles; (ii) forming a three-dimensional (3D) model of the object based on the plurality of images;(iii) identifying a plurality’ of landmarks on the object based on the 3D model; and (iv) collating the plurality' of landmarks to form the landmark map.

3. The method of claim 1, wherein (a) comprises:(i) obtaining a plurality’ of images of the object from a plurality of different angles; (ii) extracting texture-based landmark points from each image of the plurality’ of images, thereby forming a set of extracted landmark points;(iii) performing a segmentation procedure on the set of extracted landmark points to retain only landmark points on the object, thereby forming a set of object landmark points;(iv) extracting landmark points from consecutive frames of the set of object landmark points, thereby forming a set of consecutive object landmark points; and(v) collating the plurality of consecutive object landmark points to form the landmark map.

4. The method of claim 2 or 3, wherein the plurality’ of images of the object are obtained while the object is stationary.

5. The method of any one of claims 2-4. wherein each image of the plurality of images comprises a two-dimensional (2D) image or a 3D image.

6. The method of any one of claims 1-5, wherein (b) comprises:(i) obtaining a plurality’ of images of the object at a plurality’ of points in time; and (ii) determining, based on the plurality of images, a location and orientation of theobject at each point in time.

7. The method of claim 6, wherein the plurality’ of images is obtained at a frame rate of about 30 hertz (Hz).

8. The method of claim 6 or 7, wherein (b) further comprises: between successive points in time, calculating an expected position and orientation to be assumed by the object based on the location and orientation of the object at a most recent point in time.

9. The method of claim 8, wherein the expected position and onentation are calculated at a refresh rate of about 100 Hz.

10. The method of any one of claims 6-9, wherein the plurality of images of the object are obtained while the object is moving.

11. The method of any one of claims 1-10. wherein the trajectory of operations comprises a constant-speed trajectory of operations.

12. The method of any one of claims 1-11, wherein the object moves relative to the robotic operating apparatus.

13. The method of any one of claims 1-12, wherein the object moves along a conveyor belt.

14. The method of any one of claims 1-13, further comprising performing the trajectory' of operations on the object.

15. The method of any one of claims 1-14, wherein the trajectory of operations comprises a plurality of robotic manufacturing operations.

16. A system comprising:a landmark map module configured to obtain a landmark map associated yvith an object;a position determination module configured to determine, based on the landmark map, a plurality of positions and orientations assumed by the object over time; anda trajectory generation module configured to determine, based on the plurality of positions and orientations assumed by the object, a trajectory' of operations to be performed on the object by a robotic operating apparatus.

17. The system of claim 16, wherein the landmark map module comprises:a first imaging module configured to obtain a plurality of images of the object from a plurality of different angles;a first computer system configured to form a three-dimensional (3D) model of the object based on the plurality of images, to identify a plurality of landmarks on the object based on the 3D model, and to collate the plurality of landmarks to form thelandmark map.

18. The system of claim 16, wherein the landmark map module comprises:a first imaging module configured to obtain a plurality of images of the object from a lurality of different angles; anda first computer sy stem configured to extract texture-based landmark points from each image of the plurality of images, thereby forming a set of extracted landmark points, to perform a segmentation procedure on the set of extracted landmark points to retain only landmark points on the object, thereby forming a set of object landmark points, to extract landmark points from consecutive frames of the set of object landmark points, thereby forming a set of consecutive object landmark points, and to collate the plurality of consecutive object landmark points to form the landmark map.

19. The system of claim 17 or 18, wherein the plurality of images of the object are obtained while the object is stationary.

20. The system of any one of claims 17-19, wherein each image of the plurality of images comprises a two-dimensional (2D) image or a 3D image.

21. The system of any one of claims 16-20, wherein the position determination module comprises:a second imaging module configured to obtain a plurality of images of the object at a plurality of points in time; anda second computer system configured to determine, based on the lurality of images, a location and orientation of the object at each point in time.

22. The system of claim 21, wherein the plurality of images is obtained at a frame rate of about 30 hertz (Hz).

23. The system of claim 21 or 22. wherein the second computer system is further configured to: between successive points in time, calculating an expected position and orientation to be assumed by the object based on the location and orientation of the object at a most recent point in time.

24. The system of claim 23, wherein the expected position and orientation are calculated at a refresh rate of about 100 Hz.

25. The system of any one of claims 21-24, wherein the plurality of images of the object are obtained while the object is moving.

26. The system of any one of claims 16-25. wherein the trajectory of operations comprises a constant-speed trajectory of operations.

27. The system of any one of claims 16-26, wherein the object moves relative to the roboticoperating apparatus.

28. The system of any one of claims 16-27. further comprising a conveyor belt configured to move the object.

29. The system of any one of claims 16-28, further comprising an operations module configured to perform the trajectory of operations on the object.

30. The system of claim 29, wherein the operations module comprises a robotic manufacturing module and the trajectory of operations comprises a trajectory of robotic manufacturing operations.

31. A method comprising:(a) generating a landmark map associated with an object; and(b) transmitting the landmark map for subsequent use comprising:determining, based on the landmark map, a plurality of positions and orientations assumed by the object over time; anddetermining, based on the plurality of positions and orientations assumed by the object, a trajectory of operations to be performed on the object by a robotic operating apparatus.

32. The method of claim 31, wherein (a) comprises:(i) obtaining a plurality of images of the object from a plurality of different angles; (ii) forming a three-dimensional (3D) model of the object based on the plurality of images;(iii) identifying a plurality' of landmarks on the object based on the 3D model; and (iv) collating the plurality of landmarks to form the landmark map.

33. The method of claim 31, wherein (a) comprises:(i) obtaining a plurality of images of the object from a plurality of different angles; (ii) extracting texture-based landmark points from each image of the plurality of images, thereby forming a set of extracted landmark points;(iii) performing a segmentation procedure on the set of extracted landmark points to retain only landmark points on the object, thereby forming a set of object landmark points;(iv) extracting landmark points from consecutive frames of the set of object landmark points, thereby forming a set of consecutive object landmark points; and(v) collating the plurality of consecutive object landmark points to form the landmark map.

34. The method of claim 32 or 33, wherein the plurality of images of the object are obtained while the object is stationary.

35. The method of any one of claims 32-34, wherein each image of the plurality of images comprises a two-dimensional (2D) image or a 3D image.

36. The method of any one of claims 31-35, wherein the object moves relative to the robotic operating apparatus.

37. The method of any one of claims 31-36, wherein the object moves along a conveyor belt.

38. A system comprising:a landmark map module configured to generate a landmark map associated with an object; anda transmission module configured to transmit the landmark map for subsequent use comprising:determining, based on the landmark map, a plurality of positions and orientations assumed by the object over time; anddetermining, based on the plurality of positions and orientations assumed by the object, a trajectory of operations to be performed on the object by a robotic operating apparatus.

39. The system of claim 38, wherein the landmark map module comprises:an imaging module configured to obtain a plurality of images of the object from a plurality of different angles; anda computing module configured to form a three-dimensional (3D) model of the object based on the plurality of images, identity7a plurality7of landmarks on the object based on the 3D model, and collate the plurality of landmarks to form the landmark map.

40. The system of claim 38, wherein the landmark map module comprises:an imaging module configured to obtain a plurality7of images of the object from a plurality7of different angles; anda computing module configured to extract texture-based landmark points from each image of the plurality of images, thereby forming a set of extracted landmark points, to perform a segmentation procedure on the set of extracted landmark points to retain only landmark points on the object, thereby forming a set of object landmark points, to extract landmark points from consecutive frames of the set of object landmark points, thereby forming a set of consecutive object landmark points, and to collate the plurality of consecutive object landmark points to form the landmark map.

41. The system of claim 39 or 40. wherein the plurality of images of the object are obtained while the object is stationary.

42. The system of any one of claims 39-41, wherein each image of the plurality of images comprises a two-dimensional (2D) image or a 3D image.

43. The system of any one of claims 38-42, wherein the object moves relative to the robotic operating apparatus.

44. The system of any one of claims 38-43, wherein the object moves along a conveyor belt.

45. A system for processing a target object with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device, the system comprising:a calculation device configured to:generate a landmark map of the target object based on two-dimensional (2D) information associated with the target object and three-dimensional (3D) information associated with the target object; andgenerate control information for controlling the robot to process the target object with the processing device based on the landmark map and the image data; andan output device configured to output the control information.

46. The system of claim 45, wherein the calculation device is configured to generate the landmark map using the 3D information and landmarks extracted from the 2D information.

47. The system of claim 46, wherein the calculation device is configured to generate the landmark map by deleting or invalidating landmarks of objects around the target object from the 2D information using the 3D information.

48. The system of claim 45, wherein the 2D information comprises 2D image data.

49. The system of claim 48, wherein the 2D image data is captured by the imaging device.

50. The system of claim 45, wherein the 3D information comprises measurement data.

51. The system of claim 50, wherein the measurement data comprises at least one member selected from the group consisting of: distance information, depth information, three- dimensional coordinates, and combinations thereof.

52. The system of claim 50, wherein the measurement data comprises stereo image data.

53. The system of claim 52, wherein the stereo image data is captured by the imaging device.

54. The system of claim 52, wherein the calculation device is configured to generate the landmark map using the stereo image data and landmarks extracted from the 2D information.

55. The system of claim 45, wherein each of the 2D information and the 3D information comprises a plurality of image data repeatedly captured by the imaging device while the imaging device is moving relative to the target object.

56. The system of claim 55, wherein the plurality of image data included in the 2D information comprises a plurality of monocular image data and the plurality of image data included in the 3D information comprises a plurality of stereo image data.

57. The system of claim 45, wherein:each of the 2D information and the 3D information comprises:a first set of image data repeatedly captured by the imaging device while the imaging device is moving relative to the target object in a first direction; and a second set of image data repeatedly captured by the imaging device while the imaging device is moving relative to the target object in a second direction different from the first direction; andthe calculation device is configured to:generate the landmark map based on the first set of image data and the second set of image data.

58. The system of claim 46, wherein the calculation device is configured to extract the landmarks from the 2D information based on the 3D information and depth information.

59. The system of claim 45, wherein the calculation device is configured to generate an output signal for displaying the generated landmark map of target object on a display device.

60. The system of claim 45, wherein the calculation device is configured to generate an output signal that indicates position information of the target object from which non-extracted feature points are removed.

61. The system of claim 45, wherein the calculation device is configured to generate the control information based on landmarks of the target object extracted from the image data and the landmark map.

62. The system of claim 56, wherein the calculation device is configured to use the stereo image data to identify the landmarks of the target object from among landmarks extracted from the monocular image data.

63. The system of claim 62, wherein the calculation device uses the stereo image data and the set depth information to identify the landmarks of the target object from among the landmarks extracted from the monocular image data.

64. The system of claim 45, wherein the calculation device is configured to generate the control information to sequentially process different positions on the target object with theprocessing device based on each of the plurality of image data and the landmark map, and wherein the plurality of image data is generated by repeatedly capturing images of the target object with the imaging device while at least one of the processing device and the target object is moving.

65. A system for processing a target object with a processing device attached to a robot based on image data generated by capturing one or more images of the target object with an imaging device, the system comprising:a calculation device configured to:generate an operating path along which the processing device passes based on a path condition for generating the operating path and period information based on at least one of a capturing period of the imaging device and a control period of the robot, wherein the processing device is configured to use the operating path to process the target object; andgenerate control information for controlling the robot, such that the processing device moves along the operating path based on the image data; and an output device configured to output the control information.

66. The system of claim 65, wherein the calculation device is configured to generate the operating path based on the period information and a motion result of the robot operating, such that the processing device moves based on the path condition.

67. The system of claim 66, wherein the operating path comprises at least one of position information and posture information of at least one of the robot and the processing device while the processing device is moving based on the path condition.

68. The system of claim 67, wherein the path condition comprises a path which the processing device passes over the target object, and wherein the path is set by a user.

69. The system of claim 68, wherein the path condition comprises at least one of position information and posture information of at least one of the processing device and the robot at each position of the target object set by user.

70. The system of claim 65, wherein the operating path has way points, and wherein each of the waypoints has at least one of position information and posture information.

71. The system of claim 70, wherein the intervals between the plurality of waypoints are calculated such that a moving speed of the processing device over the target object is constant.

72. The system of claim 71, wherein the intervals between the waypoints are constant.

73. The system of claim 65, wherein, while the target object is moving, the processing deviceperforms the processing on the target object while moving along the operating path based on the control information.

74. A system for performing a process on an object while at least one of the target object and a robot having a processing device for performing a process on the target object is moving, the system comprising:a computing device configured to:generate an operating path along which the processing device passes based on a condition for generating the operating path;generate position information of feature points associated with the target object by extracting the feature points of the target object; and calculate at least one of a position of the processing device and a position of the robot relative to the target object based on image data captured by an imaging device that captures at least one of the target object and the processing device during processing of the target object by the processing device and the position information;a calculation device configured to generate a control signal for controlling at least one of the robot and the processing device based on the calculated position information regarding at least one of the position of the robot and the position of the processing device and the movement path; andan output device configured to output the control signal.

75. The system of claim 74, wherein the imaging device is attached to the robot.

76. The system of claim 74, wherein the robot is configured to move on a base within a space in which the processing device performs the processing on the target object.

77. The system of claim 74, wherein the calculating device is further configured to:calculate a moving speed of at least one of the target object, the processing device, and the robot based on the image data; andgenerate the control signal based on the posture information, the operating path, and the moving speed.

78. The system of claim 74, wherein the calculating device is further configured to:at each time step, report a current end-effector pose;between consecutive time steps, compute a change in end-effector pose; and subtract the change in end-effector pose from a total apparent motion to provide a difference representing the true motion of the object.

79. The system of claim 74, wherein the robot is configured to report a current end-effector pose at each image capture time step; and wherein the computing device is further configured to:compute a change in end-effector pose between consecutive image capture time steps based on reported end-effector poses;estimate, from the image data, a total apparent motion comprising motion attributed to the robot moving the imaging device and motion of the target object; and estimate a pose of the target object by performing a pose optimization in which the change in end-effector pose is used as a constraint to decouple motion attributed to the robot moving the imaging device from the motion of the target object.

80. A method for estimating a pose of a target object relative to a robotic operating apparatus, comprising:obtaining, from the robotic operating apparatus, kinematic state data comprising an end-effector pose and velocity reported by the robotic operating apparatus;computing a contribution attributed to a motion of the robotic operating apparatus to an observed motion of an imaging module;performing a pose optimization that estimates the pose of the target object based on landmarks extracted from image data captured by an imaging device and matched against a landmark map associated with the target object, wherein the pose optimization uses the computed contribution of the motion of the robotic operating apparatus as a constraint by accounting for a known motion of the imaging device between consecutive image capture time steps to decouple a motion of the robotic operating apparatus from a total apparent motion, such that a resulting pose estimate represents a true motion of the target object.