Photogrammetry system
A network of nodes with a hub controller processes stereoscopic images to enhance photogrammetry accuracy and enable real-time monitoring in complex environments by synchronizing local 3D position determination and reducing line-of-sight issues.
Patent Information
- Application Number
- JP2023546570
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-12
- Filing Date
- 2021-09-27
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-09-27
Smart Images

Figure 0007716776000001 
Figure 0007716776000002 
Figure 0007716776000003
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of photogrammetry, and more particularly to photogrammetry systems and methods for monitoring the position of a target of interest within a three-dimensional (3D) space.
Background Art
[0002] There are numerous systems for monitoring the position of a target object within a 3D space, including so-called coordinate measuring machines, laser trackers, photogrammetry, and motion capture systems.
[0003] Some photogrammetry systems include a stereo camera assembly in which two cameras are attached at a known distance from each other by a rigid structure. The two cameras each comprise a respective lens assembly associated with an image sensor operable to capture a pair of stereo images of a scene. The pair of stereo images are then analyzed using conventional photogrammetry techniques to determine the 3D position of a target within the scene, for example, based on the relative displacement of the target within the pair of stereo images.
[0004] Such photogrammetry systems were originally designed to operate in a well-controlled laboratory environment. However, the value of performing 3D measurements in more challenging environments such as those encountered in a manufacturing setting on a shop floor has been recognized, and some systems have been adapted with shop floor applications in mind. For example, it is known to direct a single stereo camera assembly towards a manufacturing shop floor scene including a target of interest, such as a robot end effector, and to determine the relative position of the target from the stereo camera assembly based on the analysis of stereo images captured by the stereo camera assembly.
[0005] Despite conventional efforts in the art, the above systems have several drawbacks that make it difficult to obtain position data accurate enough to be used in more challenging or complex settings.
[0006] First, the accuracy of a single stereo camera assembly is limited by its "baseline length", which is the separation distance between the lens assemblies. For a given baseline length, the measurement error tends to increase in relation to the square of the range, i.e., the distance to the target being monitored by the camera. Doubling the measurement range results in approximately four times the error. By increasing the baseline length, it may be possible to improve the accuracy of the camera over long ranges (such as those commonly encountered in manufacturing settings). However, this is often unrealistic as the rigid structure between the cameras is more susceptible to the effects of greater heat and vibration, which in itself introduces errors in the measurements (e.g., when used in a manufacturing setting).
[0007] Also, in complex environments, it is known that the line of sight of the camera towards the target is sometimes blocked, thereby making it difficult to capture an image of the target. For example, a target in the form of a robot end effector may face the camera at the start of a manufacturing operation (and is thus visible from the camera), but then turns to face outwards from the camera during the operation, thereby interrupting the capture of relevant data. Manufacturing processes, etc., often require continuous and real-time monitoring of the target position within the scene, and as such, conventional systems are not suitable for such applications.
[0008] It is desirable to provide a photogrammetry system that enables accurate and continuous position monitoring of a target object, particularly suitable for use in complex settings. SUMMARY OF THE INVENTION
[0009] According to one aspect, there is provided a photogrammetry system for determining the position of a target, comprising a network of two or more nodes arranged around a scene including the target, and a hub controller communicating with the network of two or more nodes. Each node is configured to capture a set of stereoscopic images of the scene in synchronization with one or more other nodes in the network, and when the target is visible in the set of stereoscopic images, to process the set of stereoscopic images to determine the three-dimensional position of the target and cause the hub controller to transmit target data including the three-dimensional position of the target to the node. The hub controller is configured to receive target data from the network of two or more nodes and determine the global position of the target in the first coordinate space based on one or more three-dimensional positions included in the target data received from the network of two or more nodes.
[0010] By determining the global position of the target based on the three-dimensional positions reported (received) by multiple nodes, the accuracy of the global positioning determination can be increased compared to what can be achieved by conventional systems using a single node or a stereo camera assembly.
[0011] Furthermore, by locally determining the three-dimensional position of the target at each node and transmitting the three-dimensional position (e.g., in the form of coordinate values) to the hub controller for further processing, the system can avoid data transfer bottlenecks that would otherwise occur if the image data itself were transferred from each node to the hub controller for central image processing. This can be advantageous in terms of increasing the processing efficiency of the system, thus shortening the overall processing time and thereby enabling real-time position monitoring.
[0012] The two or more nodes may be located at different positions around the scene to capture images of the scene from unique viewpoints. This can reduce or even avoid the risk of completely losing the line of sight to the target.
[0013] The system may be extensible in that the number of nodes in the network can be adjusted to include any number of nodes. In this regard, the number of nodes can be adjusted without substantially affecting the operation of the nodes or the hub controller. For example, the hub controller can be configured to determine the global position of a target based on all three-dimensional positions received from the network of nodes (for a given timestamp) for that target, regardless of the number of nodes.
[0014] The network can include at least four nodes. This can achieve a minimum desired level of accuracy and versatility for monitoring a target within a complex setting.
[0015] The three-dimensional positions included in the target data may be the relative positions of the target from the nodes. The three-dimensional positions may be represented in the form of three-dimensional coordinate values.
[0016] The hub controller can be configured to convert the relative position of the target to an absolute position within a first coordinate space by determining the positional relationship between the target and a reference point. The position of the reference point within the first coordinate space may be known to the hub controller. The reference point may be a reference point on the node in question, or a reference point corresponding to a reference target located within the scene.
[0017] The reference target can be located at a reference point within the scene. The processor can be configured to process a set of stereoscopic images to determine the three-dimensional position of the reference target relative to the node, and cause the hub controller to transmit target data including the three-dimensional position of the reference target relative to the node. The hub controller can be configured to determine the positional relationship between the target and the reference point by comparing their relative positions from the nodes.
[0018] The hub controller may be configured to receive reference data indicating the position of the reference point within the first coordinate space.
[0019] Each node, if visible, can be configured to process a set of stereo images to identify a target and generate metadata indicating the determined identification information of the target. The target data transmitted to the hub controller can further include the metadata. The metadata indicating the determined identification information of the target can be advantageous in enabling the hub controller to group 3D positions corresponding to the same target. Thereby, the positions of multiple targets within the scene can be monitored.
[0020] The hub controller may be configured to issue a trigger command to a network of two or more nodes to control the synchronous capture of each set of stereo images.
[0021] The hub controller can be configured to determine the global position of a target by performing an optimal operation on a plurality of 3D positions determined for the target.
[0022] The system can further include a flash unit for illuminating the scene with infrared radiation. The digital camera may respond to the infrared radiation such that the stereo images represent the intensity of the infrared radiation reflected by the scene.
[0023] The target can include a marker attached to an object within the scene.
[0024] The target can include a platform and a marker disposed at a first predetermined position on the platform. The processor can be further configured to process the set of stereo images to determine the position of the marker on the platform and identify the target based on the determined position.
[0025] The platform can have a plurality of attachment points arranged in a predetermined position array on the platform. The marker can be suitable for being removably attached to the platform at any one of the attachment points at a time.
[0026] The target can include a plurality of markers arranged in a predetermined pattern. A processor or a hub controller may be configured to determine the orientation of the target based on the three-dimensional positions of the plurality of markers.
[0027] Each marker may include a retroreflective material and / or may have a spherical shape.
[0028] According to another aspect of the present disclosure, a method of operating a photogrammetry system according to any one of the foregoing descriptions is provided. The method includes each node using a digital camera to capture a set of stereoscopic images of a scene in synchronization with one or more other nodes in the network, and when a target is visible in the stereoscopic images, processing the set of stereoscopic images to determine the three-dimensional position of the target and transmitting target data including the three-dimensional position of the target to a hub controller. The method further includes the hub controller receiving target data from a network of two or more nodes and determining a global position of the target in a first coordinate space based on one or more three-dimensional positions included in the target data received from the network of two or more nodes.
[0029] When the three-dimensional position included in the target data is the relative position of the target from the node, the method can include the hub controller converting the relative position of the target to an absolute position in the first coordinate space by determining the positional relationship between the target and a reference point, and the position of the reference point in the first coordinate space is known to the hub controller.
[0030] A reference target can be located at a reference point in the scene. In such an embodiment, the method can include a processor processing a set of stereoscopic images to determine the three-dimensional position of the reference target relative to the node and causing the node to transmit target data including the three-dimensional position of the reference target to the hub controller. The method can include the hub controller determining the positional relationship between the target and the reference point by comparing their relative positions from the node.
[0031] The method can include receiving, at a hub controller, reference data indicating a position of a reference point within a first coordinate space.
[0032] The method can include each node processing a set of stereo images to identify a target within a scene and generating metadata indicating determined identification information of the target, where the target data further includes the metadata.
[0033] The method can include the hub controller issuing a trigger command instructing nodes within a network to perform synchronized image capture. The method can include each node capturing a set of stereo images in synchronization with other nodes in response to receiving the trigger command from the hub controller.
[0034] Determining the global position of the target can include performing an optimal operation on a plurality of determined 3D positions for the target.
[0035] The optimal operation can be a weighted fit of 3D positions, where the weighting for each 3D position can be based on the relative position of the target from the node corresponding to the 3D position.
[0036] The method can include two or more nodes processing a set of stereo images in parallel.
[0037] According to one aspect of the technology described herein, a target for use in a photogrammetry system such as the aforementioned system in the foregoing description is provided. The target can have any one or more of the features described herein. Thus, the target can include a platform and a marker disposed at a first predetermined position on the platform. The platform can have a plurality of attachment points disposed in a predetermined position arrangement on the platform. The marker can be suitable for being removably attached to the platform at any one of the attachment points.
[0038] The processor(s) and controller(s) (and various related elements) described herein can comprise any suitable circuitry for causing the execution of the methods described herein and shown in the figures. The processor or controller can include at least one application specific integrated circuit (ASIC), and / or at least one field programmable gate array (FPGA), and / or a single or multi-processor architecture, and / or a sequential (Von Neumann) / parallel architecture, and / or at least one programmable logic controller (PLC), and / or at least one microprocessor, and / or at least one microcontroller, and / or a central processing unit (CPU) for executing the present method.
[0039] The processor or controller may include at least one microprocessor, may include a single core processor, may include multiple processor cores (such as a dual core processor or a quad core processor), or may include multiple processors (at least one of which may include multiple processor cores).
[0040] The processor or controller may be part of a system that includes an electronic display, which may be any suitable device for communicating information, such as location data, to a user.
[0041] The processor or controller can comprise one or more memories for storing the data described herein and / or for storing software for executing the processes described herein, and / or can communicate with one or more memories.
[0042] The memory may be any suitable non-transitory computer-readable storage medium, one or more data storage devices, and may include a hard disk and / or solid state memory (such as flash memory). The memory may be a permanent non-removable memory, or a removable memory (such as a Universal Serial Bus (USB) flash drive).
[0043] When read by a processor or a controller, the memory can store a computer program including computer-readable instructions that cause the execution of the methods described herein and shown in the figures. The computer program may be software or firmware, or a combination of software and firmware.
[0044] The computer-readable storage medium may be, for example, a USB flash drive, a compact disc (CD), a digital versatile disc (DVD), or a Blu-ray disc. In some examples, the computer-readable instructions may be transferred to the memory via a wireless signal or a wired signal.
[0045] Those skilled in the art will understand that, unless mutually exclusive, the features or parameters described in connection with any one of the above aspects can be applied to any other aspect. Further, unless mutually exclusive, any feature or parameter described herein may be applied to any aspect, and / or combined with any other feature or parameter described herein.
Brief Description of the Drawings
[0046] Here, with reference to the drawings, embodiments will be described as mere examples.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0047] It will be understood that like reference numerals are used in the drawings to label like features of the technologies described herein.
[0048] FIGS. 1 and 2 show an example of a photogrammetry system 100 suitable for determining and monitoring in real time the position of a target object (hereinafter referred to as “target”) according to the technologies described herein. In the examples of FIGS. 1 and 2, the system is used to monitor a target within a scene corresponding to a manufacturing setting, but it will be understood that system 100 is more generally applicable to any type of setting. For example, the system can be used in a medical operating room to monitor the position of surgical instruments.
[0049] System 100 comprises an interconnected network of node devices 101, 102, 103, 104, 105 that communicate wired with hub controller 106 via one or more communication lines 107 (although wireless communication is equally possible). Hub controller 106 is operable to use a processing circuit to transmit control signals to nodes 101, 102, 103, 104, 105 and to control the operation of the nodes according to the methods described herein.
[0050] As best shown in FIG. 3 which schematically shows one such node 101 of system 100, each node includes a housing 108 that surrounds a first digital camera 301 and a second digital camera 302 mounted apart on a rigid support structure 303. The cameras 301, 302 and the associated support structure 303 form a stereo camera assembly.
[0051] Each camera 301, 302 has an image sensor (not shown), such as a charge coupled device (CCD), and associated lens assemblies 110, 111. The lens assemblies 110, 111 are exposed on a common outer surface 109 of the housing 108 to direct external light (electromagnetic radiation) incident on the lens assemblies 110, 111 towards their associated image sensors. The node 101 can also include at least one flash unit (not shown) operable to illuminate the scene being imaged, for example one flash unit per camera 301, 302.
[0052] The first and second cameras 301, 302, and thus their corresponding lens assemblies 110, 111, are separated by a known distance herein referred to as the baseline length 112 of the stereo camera assembly. In the systems of FIGS. 1 and 2, the camera assemblies within each node 101, 102, 103, 104, 105 have the same baseline length 112, for example 400 mm, but in other arrangements the baseline length 112 of the camera assemblies can vary between nodes.
[0053] Node 101 further includes an on-board processor 304 operable to cause cameras 301, 302 to simultaneously capture images representative of light incident on the sensors (e.g., under the instruction of hub controller 106). Each stereo camera assembly generates a set (a pair in this example) of stereo images representative of a scene within the field of view of lens assemblies 110, 111. In this regard, lens assemblies 110, 111 are oriented such that their fields of view overlap at least partially. In this way, any target located within the overlapping portion is represented in both stereo images, despite being at an offset position within each image due to the baseline length spacing between lens assemblies 110, 111.
[0054] The on-board processor 304 is configured to perform edge processing at the node, including processing the stereo images generated by cameras 301, 302 within node 101. As will be described in more detail below, processor 304 is configured to process a pair of stereo images to determine the 3D position of targets within the scene. To facilitate this, processor 304 may receive or otherwise be provided with all information necessary to determine the position of a target based on the image data, such as the epipolar geometry of the stereo camera assembly within the node. The epipolar geometry refers to the relative separation and orientation of cameras 301, 302 within the stereo camera assembly and is accurately determined by a conventional calibration process not described herein.
[0055] After determining the 3D position of a target within the scene, processor 304 causes the node to report (i.e., transmit) the determined 3D position of the target to hub controller 106. Hub controller 106 then determines, for each target, a single global position of the target within a first coordinate space 120 based on the 3D position reported by the node of that target.
[0056] Referring to FIG. 1, nodes 101, 102, 103, 104, 105 are located at different positions around the scene to capture images of the scene from their respective perspectives. Further, the nodes are arranged and oriented such that at least some of their fields of view overlap at points within the scene that are expected to be visible from the nodes when there is one or more targets.
[0057] It will be appreciated that the target can be any object of interest (or actually a part of the object) that can be identified within the image using conventional image processing techniques. The target may be, for example, a tool such as a robotic end effector that is movable within the scene with respect to the node. Alternatively, the target may be in the form of a marker or a group of markers that are fixedly attached, for example movably, to the tool so that the position of the tool can be determined based on the position of the marker.
[0058] In the example of FIGS. 1 and 2, the scene includes a first target 113 that is attached to the arm 114 at a position fixed with respect to the end effector 115 of the movable robotic arm 114. The second target 116 is attached to a first ground reference point having a proximity fixed to the robotic arm 114, the third target 117 is attached to a second ground reference point of proximity fixed to the workpiece 118 or other component that is manipulated or otherwise used in the manufacturing process, and the fourth target 119 is attached to a third reference point of proximity fixed to the workpiece 118.
[0059] The positions of the ground reference points, and thus the second, third, and fourth targets 116, 117, 119 within the first coordinate space 120, are known to the nodes 101, 102, 103, 104, 105 and the hub controller 106 such that the position of the movable first target 113 within the first coordinate space 120 can be monitored relative to the positions of the second, third, and / or fourth targets 116, 117, 119 based on the analysis of the stereoscopic images of the targets. The positions of the ground reference points are pre-determined by conventional pre-surveying techniques not described herein.
[0060] By using a system having a network of multiple nodes as compared to a single node or a stereo camera assembly, it will be appreciated that, for example, during a manufacturing operation, the risk that the line of sight to a target is completely lost at any point is reduced. In fact, the end user can strategically place the nodes around the scene to ensure that the line of sight to the target(s) is maintained for at least one, preferably two or more, or all of the nodes within the network. Thereby, it is possible to avoid the determination of the target position being interrupted and the target can be continuously monitored.
[0061] Although Figures 1 and 2 show five nodes, this is for illustrative purposes only. The system can comprise a network with any number of two or more nodes. In fact, the advantage of the system is that the number of nodes within the network can be easily adjusted to suit a particular scene without the need to modify the positioning process. In fact, an entire production line or even an entire factory can be monitored as a single network.
[0062] Furthermore, the nodes, particularly their cameras 301, 302, may be configured to detect light in the visible or infrared range of the electromagnetic spectrum. Correspondingly, the flash unit(s) of each node can be configured to illuminate the scene with radiation in the corresponding range of the electromagnetic spectrum.
[0063] Using a digital camera and a flash unit that operate in the infrared range of the spectrum can be particularly advantageous over those that operate in the visible range. In this regard, it will be appreciated that ambient visible light is typically a weak variable light source that limits the visibility of the target over a long range. Further, a flash unit that illuminates the target by transmitting visible light is often uncomfortable or not safe for a person working within the field of view of the flash unit. This is particularly true when a high-frequency strobe is used. On the other hand, infrared radiation transmits relatively strongly (with a larger amplitude) over long distances and does not pose a danger of visible light rays to people working within the range of the flash unit. Therefore, by utilizing infrared radiation, it is possible to improve the resolution of the target in the stereo image while avoiding the danger of visible light.
[0064] Furthermore, with respect to a stereo camera assembly having only two digital cameras that capture a corresponding pair of stereo images, the system has been described above, but each stereo camera assembly can have three or more cameras spaced apart by a known distance (e.g., to widen the field of view of each node). Such a camera assembly generates a corresponding set of three or more stereo images.
[0065] Furthermore, the system does not rely on maintaining rigidly attached nodes, and their movements can be calculated and compensated without affecting the accuracy of the system. Therefore, the system can be used in a mobile configuration where nodes are mounted on, for example, a moving vehicle or a drone.
[0066] FIG. 4 schematically shows an example of targets 113, 116, 117, 119 attached to an object within a scene imaged by the system of FIGS. 1 and 2.
[0067] The target 400 comprises a plate (or generally a platform) 401 having a top surface 402 and a bottom surface 403 opposite the top surface 402. The bottom surface 403 is suitable for attachment to an object within a scene, for example by adhesion or other attachment means. A first group of markers 404 is attached to the top surface 402 of the plate 401 and arranged to define a unique pattern recognizable when imaged by the processor 304. In this example, the first group of markers 404 defines a T-shaped pattern, although other patterns are possible.
[0068] The target 400 further comprises a movable marker 405 suitable for removably attaching to the plate 401 at any one of a plurality of attachment points 406 located on the top surface 402. In the illustrated example, the movable marker 405 comprises a body portion 407, a base 408, and a post 409 projecting from the side of the base 408 opposite the body portion 407. Each attachment point 406 is in the form corresponding to a hole or recess configured to receive the post 409 of the movable marker 405. The cross-sectional profiles of the post 409 and the hole 406 are shaped to fit snugly and securely receive the post 409 within the hole 406, yet still be removable by force.
[0069] The attachment points (holes) 406 are arranged in an array at predetermined positions on the top surface 402 of the plate 401 such that the position of the movable marker 405 relative to the first group of markers 404 can be selected and changed by the user for each target. Thus, each target 400 within a scene as described above with respect to FIGS. 1 and 2 can be configured with a unique pattern of markers 404, 405, and thus the processor 304 can identify the target 400 by its unique pattern. In this way, the target 400 can encode data in a visual machine-readable format by changing the position of the marker 405.
[0070] The body portions 407 of each of the markers 404, 405 are substantially spherical. Such a shape has a more consistent appearance in an image and thus facilitates quicker and easier detection of the marker over long distances and a wide range of acceptance angles. This is particularly true when the surface facing the camera is oriented within the scene at an angle to the image plane of the camera, as compared to a circular disk-shaped marker that can appear as an ellipse (rather than a circle) within the field of view of the node. In such a case, as the ellipse becomes more elongated and oblique, automatic recognition of the marker becomes problematic. Also, the amount of light reflected by a flat disk is severely attenuated when viewed obliquely, whereas a sphere avoids this problem.
[0071] The individual markers 404, 405 can include steel or, alternatively, may be coated with a retroreflective material configured to reflect electromagnetic radiation at the wavelengths to which the image sensors of the nodes described above respond. The retroreflective material may be in the form of a paint containing glass beads within a substrate.
[0072] The marker has been described as having a spherical shape, although other forms of coded and uncoded markers can be used in the systems of FIGS. 1 and 2, which may be retroreflective, white (or other colors), passive markers and / or active markers, i.e., self-illuminating markers. An example of an active marker is an LED light source.
[0073] Here, the operation of the system will be described with respect to FIGS. 5 and 6.
[0074] FIG. 5 is a flowchart generally showing the processing method executed by each of the nodes 101, 102, 103, 104, 105 within the systems of FIGS. 1 and 2.
[0075] The method starts at block 501 where the nodes are triggered to capture a pair of stereo images of the scene. The network of nodes receives a trigger command from the hub controller 106 (or other host device) to cause the cameras 301, 302 of each node to capture the first and second stereo images of the scene simultaneously and in synchronization with the other nodes of the network.
[0076] There may be an initialization step before block 501, and the nodes receive reference data that stores the information necessary to perform the image processing operations and corresponding positioning. For example, the nodes receive data regarding a target within the scene (such as a predetermined pattern of markers forming the target), the position of a reference point within the first coordinate space 120, and / or the epipolar geometry of the stereo camera assembly.
[0077] At block 502, the processor 304 operates to process the set of stereo images to detect and identify visible targets within the images. The processor 304 can use conventional image processing techniques for object recognition to determine which one or more of the targets 113, 116, 117, 119 are visible in both the first and second stereo images of the pair. For the specific examples of FIGS. 1 - 4 where each target has a unique pattern of markers, the processor 304 can identify the targets (if visible) by their unique marker patterns. If at least one of the first target 113, i.e., the moving target of interest, and the second, third, and fourth targets 116, 117, 119 located at known reference points within the scene are visible in the first and second stereo images, the processor 304 proceeds to block 503.
[0078] At block 503, the positions of the targets identified within the first and second stereo images are analyzed by on-board processor 304 to determine their 3D positions relative to the reference points on the nodes in question. It will be appreciated that any conventional photogrammetric technique suitable for determining the 3D position of an object within a stereo image may be used for this purpose. For example, corresponding targets within the first and second stereo images can be the subject of point triangulation, whereby the processor, along with knowledge of the epipolar geometry of the stereo camera assembly, uses the offset positions of each target (or individual marker) within the first and second stereo images to determine the 3D position of the target relative to the node.
[0079] At this stage, the orientation of the targets can be determined based on the relative positions of their individual markers. That is, the node can determine not only the position of the targets but also their orientation. In this way, the system can determine the position of the targets in six degrees of freedom.
[0080] Processor 304 generates target data that includes information indicating the determined 3D positions of targets 113, 116, 117, 119. This can include the 3D positions of the plurality of markers that form the target, or the 3D position of the target based on the 3D positions of the markers themselves. The target data can also include the determined orientation of the targets. The 3D position(s) are in the form of coordinate values and are thus easily transmissible across the system. The target data can also include metadata such as a unique point (or target) identifier (i.e., name), a data quality indicator such as the measured circularity of the target, and one or more of the camera settings used for the data set, all of which can be used to enhance the calculation of the global position data, as will be described later with reference to FIG. 6. The metadata can also include timestamp data associated with the 3D position data, which indicates the time at which the set of stereo images was captured and thus the time at which the target was at the determined position.
[0081] In block 504, the target data including the position data and metadata of each target identified within the stereoscopic image is transmitted to the hub controller 106 via the communication line 107. The target data may be transmitted to the hub controller substantially simultaneously with other nodes within the network. However, if a node is unable to identify the targets within a pair of stereoscopic images, the node may be configured not to transmit the target data to the hub controller, to conserve network bandwidth, or otherwise to indicate to the hub controller that it is unable to identify the targets.
[0082] In an embodiment, the nodes are configured to process each pair of stereoscopic images in parallel such that the overall processing time is reduced. The nodes may operate in parallel by one (or more or all) of the other nodes within the network performing any one or more of the functions in blocks 502 and 503 simultaneously while the nodes perform any one or more of the functions in blocks 502 and 503.
[0083] Although it was described above that target identification is performed as part of block 502 before determining the relative position of the target from the node in block 503, it will be appreciated that the step of identifying the target may be performed after block 503. For example, in block 502, the on-board processor 304 may simply detect the presence of markers within the image and proceed to determine the relative position of the markers from the node in block 503 before the markers are classified and identified as belonging to a particular target prior to block 504.
[0084] Furthermore, although the node has been described above as being configured to report the relative position of the target from the node to the hub controller, this is not necessarily the case. For example, in some arrangements, each node may be configured to convert the relative position (coordinate value) of the target from the node to an absolute position (coordinate value) within a first coordinate space 120 and report the absolute position to the hub controller in addition to, or instead of, the relative position. Thus, the target data may include information indicating the 3D positions of a first target 113 within the first coordinate space 120 and optionally second, third, and fourth targets 116, 117, 119.
[0085] As described above, the positions of the second, third, and fourth targets 116, 117, 119 within the first coordinate space 120 may already be known to the node. In such a case, the absolute position of the first target 113 may be determined based on the positional relationship between the first target 113 and the first, second, and / or third targets 116, 117, 119. The positional relationship may be determined by a processor by comparing the relative position of the first target 113 (from the node) with the relative positions of the second, third, and fourth targets 116, 117, 119.
[0086] In other embodiments, the node position within the first coordinate space 120 may be predetermined and known to the node such that the absolute position of the first target 113 can be determined based on its relative position from the known position of the node.
[0087] FIG. 6 is a flowchart schematically showing the processing steps executed by the hub controller 106.
[0088] At block 601, the hub controller 106 receives target data from the nodes that have identified the targets within the pair of stereo images. Of course, not all nodes necessarily always send target data to the hub controller (e.g., due to an unclear line of sight to the target). However, missing or incomplete target data from individual nodes does not prevent the hub controller 106 from determining the 3D positions represented by the target data received from the nodes that have identified the targets within the pair of stereo images.
[0089] The data stream is substantially continuous as long as one node can always see the target (in that it is transmitted at a high, uninterrupted rate). When the target moves out of the field of view of all nodes, its position is not determined and is sent to the hub controller, but even in that case, the target position is re - determined and sent to the hub controller when it comes back into the line of sight of any one of the nodes.
[0090] At block 602, the global position of each target represented by the received target data is determined based on the 3D position of the target determined locally at the nodes (the 3D position of the target can be for the whole target or for the individual markers that form the target). The reported position of the target can vary as a result of the inherent inaccuracies of the nodes. Thus, the target data from the nodes can collectively define a set of 3D points of the target, i.e., a set of multiple 3D positions of the first target 113. Thus, the hub controller 106 is configured to determine a single global position representing the set of 3D points, particularly the global position that best represents the determined 3D position(s) reported by the nodes. In an embodiment, the hub controller performs an optimization operation on the set of 3D points of each target to determine the position of the target.
[0091] If the target data received by the hub controller from each node includes the relative 3D positions of the target(s) from the node, the hub controller 106 performs an initial step of converting the relative positions from the node to absolute positions within the first coordinate space 120. This can be done based on the positional relationship determined between the first target 113 and the second, third, and / or fourth targets 116, 117, 119 located at the fixed reference points, substantially as described above. The hub controller proceeds to determine the optimality of the converted 3D positions of at least the first target 113 within the first coordinate space 120.
[0092] If the target data received by the hub controller from the node network includes the absolute 3D position of the first target within the first coordinate space, the hub controller 106 determines the optimality of the target position reported by the node.
[0093] In both cases, the optimal operation can determine the weighted fit of the set of 3D positions of the targets. The weighting used for each 3D position can be selected based on several factors such as the data quality indicators described above. In some arrangements, the weighting used for each 3D position is selected based on the distance between the target and the node that reported the position. The weighting may be inversely proportional to the distance. That is, more weight is given to the 3D positions corresponding to nodes closer to the target in question. For example, a 3D position locally derived at a node closer to the target is given a greater weight than a 3D position locally derived at a node farther from the same target.
[0094] The orientation of the target(s) may also be determined by the hub controller 106 based on the target data, e.g., the reported 3D positions of individual markers, so that the system can determine the position of the target in six degrees of freedom. The orientation of the target can be determined based on the relative absolute positions of the markers (which can be arranged in a predetermined pattern known to the hub controller).
[0095] By determining the global position of the target based on the 3D positions reported by multiple nodes, the accuracy of the global positioning determination is understood to be improved compared to what can be achieved by conventional systems that use a single node or a stereo camera assembly.
[0096] In block 603, the global position and / or orientation of the target in the first coordinate space is output. The global position and / or orientation can be output and used in various ways depending on the specific application of the system. In the examples of FIGS. 1 and 2, the system operates as a robot teaching or control system, and at least the measured global position of the first target 113 is used to adjust or guide the position of the robot arm 114 and thus the end effector 115. The position of the end effector 115 can be calculated based on the known positional relationship between the end effector and the target 113.
[0097] The processing steps described above with respect to FIGS. 5 and 6 can be continuously repeated, for example, at a set frequency so that the system can track the global position and / or orientation of the target over time. For example, each node can capture and process a pair of stereo images, and the hub controller can process the 3D positions reported by the nodes to determine the global target position and / or orientation at a frequency greater than 3 Hz, such as up to 50 Hz. This is particularly suitable for manufacturing environments where real-time continuous monitoring of moving objects is required. Further, by monitoring the global position and / or orientation of the target over time, the system can track, for example, the speed and acceleration of the target based on, for example, timestamp data associated with the 3D position data.
[0098] Above, a technique has been described regarding tracking the position of a target by attaching a marker to an object in a scene, but marker-based tracking is optional. As described above, the target to be tracked may be any object in the scene that can be recognized in a stereo image.
[0099] The system determines the target position locally at the node and centrally determines the global position and / or orientation of the target at the hub controller based on the determined positions reported by one or more of any of the nodes, such that the system is scalable in that it can adjust the number of nodes in the network without adversely affecting the functionality and processing speed of the system. Accordingly, the techniques described herein provide a photogrammetry system for monitoring target positions and / or orientations that not only has better accuracy, but also improved flexibility and versatility.
[0100] Furthermore, by locally determining the 3D position(s) of the target(s) at each node and transferring the 3D position(s) (e.g., in the form of target data) to the hub controller 106 for further processing, the system can avoid data transfer bottlenecks that would otherwise occur if, for example, the image data itself were transferred from each of the nodes 101, 102, 103, 104, 105 to the hub controller 106 for central image processing. This can be advantageous in that it increases the processing efficiency of the system and thus shortens the overall processing time, thereby enabling real-time position monitoring.
Claims
1. A photogrammetry system (100) for determining the position of a target (113), comprising: a network of two or more nodes (101, 102, 103, 104, 105) arranged around a scene including the target (113); a hub controller (106) communicating with the network of the two or more nodes (101, 102, 103, 104, 105); each node (101, 102, 103, 104, 105) comprising: at least two digital cameras (301, 302) configured to capture a set of stereo images of the scene in synchronization with one or more other nodes in the network; a processor (304) configured to process the set of stereo images to determine a three-dimensional position of the target (113) if the target is visible within the set of stereo images, and to cause the node to transmit target data including the three-dimensional position of the target (113) to the hub controller (106); the hub controller (106) being configured to: receive the target data from the network of the two or more nodes (101, 102, 103, 104, 105); determine a global position of the target (113) in a first coordinate space (120) based on a plurality of the three-dimensional positions included in the target data received from the network of the two or more nodes (101, 102, 103, 104, 105); the target (113) comprising: a platform (401); a movable marker (405) disposed at a first predetermined position on the platform (401); and a plurality of markers (404) disposed in a predetermined pattern; the processor (304) further configured to process the set of stereo images to determine the position of the movable marker (405) on the platform (401) and to identify the target (113) based on the determined position; the photogrammetry system (100), wherein the processor (304) or the hub controller (106) is configured to determine an orientation of the target (113) based on the three-dimensional positions of the plurality of markers (404).
2. The photogrammetry system (100) according to claim 1, wherein the two or more nodes (101, 102, 103, 104, 105) are located at different positions around the scene to capture an image of the scene from a unique perspective.
3. The photogrammetry system (100) according to claim 1 or 2, wherein the system is extensible in that the number of nodes in the network can be adjusted to include any number of nodes.
4. The photogrammetry system (100) according to claim 1, 2, or 3, wherein the network includes at least four nodes (101, 102, 103, 104, 105).
5. The 3D position included in the target data is the relative position of the target from the node, and the hub controller is configured to convert the relative position of the target into an absolute position in the first coordinate space (120) by determining the positional relationship between the target and a reference point, and the position of the reference point in the first coordinate space is known to the hub controller. The photogrammetry system (100) according to any one of claims 1 to 4.
6. A reference target is located at the reference point in the scene, The processor (304) is configured to process the set of stereo images to determine the 3D position of the reference target with respect to the node, and cause the hub controller (106) to transmit target data including the 3D position of the reference target to the node. The hub controller (106) is configured to determine the positional relationship between the target and the reference point by comparing the relative positions of the target and the reference point from the node. The photogrammetry system (100) according to claim 5.
7. The photogrammetry system (100) according to claim 5 or 6, wherein the hub controller is configured to receive reference data indicating the position of the reference point in the first coordinate space.
8. When visible, each node is configured to process the set of stereo images to identify the target and generate metadata indicating the determined identification information of the target. The target data transmitted to the hub controller (106) further includes the metadata. The photogrammetry system (100) according to any one of claims 1 to 7.
9. The photogrammetry system (100) according to any one of claims 1 to 8, wherein the hub controller is configured to issue a trigger command to the network of the two or more nodes to control synchronous capture of each set of stereoscopic images.
10. The photogrammetry system (100) according to any one of claims 1 to 9, wherein the hub controller (106) is configured to determine the global position of the target (113) by performing an optimal operation on a plurality of three-dimensional positions determined for the target.
11. The system (100) further comprises a flash unit for illuminating the scene with infrared radiation, The digital cameras (301, 302) respond to the infrared radiation such that the intensity of the infrared radiation reflected by the scene is represented by the stereoscopic images, the photogrammetry system (100) according to any one of claims 1 to 10.
12. The target (113) is attached to an object (114) within the scene or is attached to an object (114) within the scene, the photogrammetry system (100) according to any one of claims 1 to 11.
13. The processor (304) is further configured to generate metadata indicating the determined identification information of the target (113), The target data transmitted to the hub controller (106) includes the metadata, the photogrammetry system (100) according to any one of claims 1 to 12.
14. The metadata includes one or more unique object identifiers, the photogrammetry system (100) according to claim 13.
15. The platform (401) has a plurality of attachment points (406) arranged in a predetermined position arrangement on the platform (401), The movable marker (405) is suitable for being removably attached to the platform (401) at any one of the attachment points (406) at a time, the photogrammetry system (100) according to any one of claims 1 to 14.
16. Each marker (404, 405) includes a retroreflective material and / or has a spherical shape, the photogrammetry system (100) according to any one of claims 12 to 15.
17. Each of the nodes Using the digital camera (301, 302) to capture a set of stereoscopic images of the scene in synchronization with the other one or more nodes in the network, When the target is visible in the stereoscopic image, determining the three-dimensional position of the target (113) and processing the set of stereoscopic images to transmit target data including the three-dimensional position of the target (113) to the hub controller (106). The hub controller (106) Receiving the target data from the network of the two or more nodes (101, 102, 103, 104, 105), Determining the global position of the target (113) in the first coordinate space (120) based on the plurality of three-dimensional positions included in the target data received from the network of the two or more nodes (101, 102, 103, 104, 105). A method of operating the photogrammetry system (100) according to claim 1, comprising:
18. Determining the global position of the target (113) includes performing an optimal operation on a plurality of three-dimensional positions determined for the target, the optimal operation being a weighted fit of the three-dimensional positions, and the weighting resulting from each three-dimensional position being based on the relative position of the target (113) from the node (101, 102, 103, 104, 105) to which the three-dimensional position corresponds. The method according to claim 17.
19. A method of operating a photogrammetry system (100) to determine the position of a target (113), The photogrammetry system (100) A network of two or more nodes (101, 102, 103, 104, 105) arranged around a scene including the target (113), A hub controller (106) communicating with the network of the two or more nodes (101, 102, 103, 104, 105), and The method includes Each node (101, 102, 103, 104, 105) Using at least two digital cameras (301, 302) to capture a set of stereoscopic images of the scene in synchronization with the other one or more nodes in the network, When the target is visible within the stereoscopic image, the processor (304) processes the set of stereoscopic images to determine the three-dimensional position of the target (113), and transmits target data including the three-dimensional position of the target (113) to the hub controller (106), wherein the hub controller (106) receives the target data from the network of the two or more nodes (101, 102, 103, 104, 105), and determines the global position of the target (113) in the first coordinate space (120) based on the plurality of three-dimensional positions included in the target data received from the network of the two or more nodes (101, 102, 103, 104, 105). Determining the global position of the target (113) includes performing an optimal operation on the plurality of three-dimensional positions determined for the target, and the optimal operation is a weighted fit of the three-dimensional positions, and the weighting resulting from each three-dimensional position is based on the relative position of the target (113) from the node (101, 102, 103, 104, 105) corresponding to the three-dimensional position.
20. The three-dimensional position included in the target data is the relative position of the target from the node, and the method further includes the hub controller converting the relative position of the target into an absolute position within the first coordinate space (120) by determining the positional relationship between the target and a reference point, The method according to any one of claims 17 to 19, wherein the position of the reference point within the first coordinate space is known to the hub controller.
21. A reference target is located at the reference point within the scene, and the method includes the processor (304) processing the set of stereoscopic images to determine the three-dimensional position of the reference target relative to the node, and causing the hub controller (106) to transmit target data including the three-dimensional position of the reference target relative to the node, and the hub controller (106) further determining the positional relationship between the target and the reference point by comparing the relative positions of the target and the reference point from the node. The method according to claim 20.
22. The method according to claim 20 or 21, further comprising receiving, by the hub controller, reference data indicating the position of the reference point in the first coordinate space.
23. The method according to any one of claims 17 to 22, further comprising each node processing the set of stereoscopic images to identify the target in the scene and generating metadata indicating the determined identification information of the target, wherein the target data further includes the metadata.
24. The hub controller issuing a trigger command for instructing the nodes in the network to perform synchronized image capture; The method according to any one of claims 17 to 23, further comprising each node capturing the set of stereoscopic images in synchronization with the other nodes in response to receiving the trigger command from the hub controller.
25. The method according to any one of claims 17 to 24, further comprising the two or more nodes processing the set of stereoscopic images in parallel.
Citation Information
Patent Citations
Method and system for determining the centroid of an object
JP2001508209A
Object information detection system, information detection method, and program
JP2011185897A
System and method for three-dimentional measurement
JP2013079854A
Attitude estimating system, attitude estimating device, and distance image camera
JP2019066238A
Positioning device for mobile body and method for calibration
JP2019086390A