Method, device and system for estimating uncertainty in vision-based tracking system
Through multiple key point detector neural networks, predict key points in 2D images and calculate their 3D poses, derive the change metrics to generate uncertain values, solving the problems of inaccurate pose estimation and the impact of environmental variables in the prior art, and achieving more accurate and reliable pose estimation.
Patent Information
- Application Number
- CN202411717021.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-30
AI Technical Summary
In existing vision-based tracking systems, the inaccuracy of pose estimation mainly comes from inaccurate prediction of key points, and the existing methods fail to effectively consider occlusion and environmental variables, resulting in limited effectiveness in real-world scenarios.
The key points in the 2D image are predicted through multiple key point detector neural networks, and their corresponding 3D poses are calculated, the change metrics between each 3D pose are derived, and their Euclidean norms are calculated to generate uncertainty values, thereby controlling the processing between objects.
This method can more comprehensively and accurately evaluate the reliability of pose estimation, reduce errors due to inaccurate predictions, and improve the effectiveness of the system in real-world scenarios.
Smart Images

Figure CN120070561A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to vision-based tracking systems, and more particularly to estimating uncertainty in vision-based tracking systems. Background Art
[0002] Pose estimation is often used in vision-based tracking systems to accurately locate and track the position and orientation of an object within a physical environment. This process typically relies on key points predicted within a two-dimensional (2D) image or video frame. However, inaccurate prediction of key points can lead to imprecise pose estimation, thus undermining the overall performance of the system. To mitigate the risk of inaccurate pose estimation, conventional methods employ a three-dimensional (3D) model of the object, thus allowing comparison with the predicted key points in the 2D image. However, this method does not account for occlusion and environmental variables, limiting its effectiveness in real-world scenarios. Summary of the Invention
[0003] The subject matter of the present application has been developed in response to the state of the art, and in particular, in response to problems and needs that are generated by or not fully solved by typical uncertainty estimation. Generally, the subject matter of the present application has been developed to provide a method for estimating uncertainty in a vision-based tracking system that overcomes at least some of the above disadvantages of the prior art.
[0004] Disclosed herein is a method for estimating uncertainty in a vision-based tracking system. The method includes receiving a two-dimensional (2D) image of at least a portion of a first object via a camera on a second object. The method further includes predicting a set of key points on the first object in the 2D image by each of a plurality of key point detectors. The method further includes calculating a three-dimensional (3D) pose of each of the plurality of key point detectors from the corresponding set of key points. The method additionally includes deriving a measure of variation between each of the 3D poses of the plurality of key point detectors, and calculating the Euclidean norm of the measure of variation to produce an uncertainty value. The method further includes controlling the processing between the first object and the second object in response to the uncertainty value. The foregoing subject matter of this paragraph characterizes Example 1 of the present disclosure.
[0005] Calculating the 3D pose is based on a 2D-to-3D correspondence model. The foregoing subject matter of this paragraph characterizes Example 2 of the present disclosure, where Example 2 further includes the subject matter of Example 1 above.
[0006] Each 3D pose includes a vector of six values. Each of the six values represents a corresponding one of the six degrees of freedom of the 3D pose. Deriving a measure of the change between each of the 3D poses includes deriving a measure of the change in each vector for each of the six degrees of freedom. The foregoing subject matter of this paragraph represents Example 3 of the present disclosure, where Example 3 also includes the subject matter according to any one of Examples 1 to 2 above.
[0007] The plurality of key point detectors includes at least three key point detectors. The foregoing subject matter of this paragraph represents Example 4 of the present disclosure, where Example 4 also includes the subject matter according to any one of Examples 1 to 3 above.
[0008] The method includes training each of the plurality of key point detectors individually before predicting a corresponding set of key points. Each of the plurality of key point detectors has the same architecture, training duration, and training data. Each of the plurality of key point detectors is initialized with random value weights. The foregoing subject matter of this paragraph represents Example 5 of the present disclosure, where Example 5 also includes the subject matter according to any one of Examples 1 - 4 above.
[0009] The measure of the change between each of the 3D poses of the plurality of key point detectors is the standard deviation between each of the 3D poses. The foregoing subject matter of this paragraph represents Example 6 of the present disclosure, where Example 6 also includes the subject matter according to Example 5 above.
[0010] The handling between the first object and the second object includes a coupling handling between the first object and the second object. Controlling the coupling handling includes: coupling the first object and the second object when the uncertainty value is at or below a predetermined threshold, and preventing the coupling between the first object and the second object when the uncertainty value is higher than the predetermined threshold. The foregoing subject matter of this paragraph represents Example 7 of the present disclosure, where Example 7 also includes the subject matter according to any one of Examples 1 to 6 above.
[0011] Automatically control the handling between the first object and the second object. The foregoing subject matter of this paragraph represents Example 8 of the present disclosure, where Example 8 also includes the subject matter according to any one of Examples 1 - 7 above.
[0012] Manually control the handling between the first object and the second object such that the handling is initiated by an operator. The foregoing subject matter of this paragraph represents Example 9 of the present disclosure, where Example 9 also includes the subject matter according to any one of Examples 1 - 8 above.
[0013] The 2D image includes a portion of the second object. Predicting a set of key points includes predicting additional key points on the second object in the 2D image by each of the plurality of key point detectors. The foregoing subject matter of this paragraph represents Example 10 of the present disclosure, where Example 10 also includes the subject matter according to any one of Examples 1 - 9 above.
[0014] The first object is the receiver aircraft. The second object is the tanker aircraft. The process is a refueling operation between the receiver aircraft and the tanker aircraft. The foregoing subject matter of this paragraph represents Example 11 of the present disclosure, where Example 11 also includes the subject matter according to any one of the above Examples 1-10.
[0015] The present disclosure also discloses a vision-based tracking device, which includes a processor and a non-transitory computer-readable storage medium storing code. The code can be executed by the processor to perform operations including receiving a two-dimensional (2D) image of at least a part of the first object via a camera on the second object. The code can also be executed by the processor to perform operations including predicting a set of key points on the first object in the 2D image by each of a plurality of key point detectors. The code can further be executed by the processor to perform operations including calculating a three-dimensional (3D) pose of each of the plurality of key point detectors from the corresponding set of key points. The code can additionally be executed by the processor to perform operations including deriving a measure of the variation between each of the 3D poses of the plurality of key point detectors and calculating the Euclidean norm of the measure of the variation to generate an uncertainty value. The code can also be executed by the processor to perform operations including controlling the process between the first object and the second object in response to the uncertainty value. The foregoing subject matter of this paragraph represents Example 12 of the present disclosure.
[0016] Calculating the 3D pose is based on a 2D-to-3D correspondence model. The foregoing subject matter of this paragraph represents Example 13 of the present disclosure, where Example 13 also includes the subject matter according to Example 12 above.
[0017] The code can be executed by the processor to train each of the plurality of key point detectors individually before predicting a set of key points. Each of the plurality of key point detectors has the same architecture, training duration, and training data. Each of the plurality of key point detectors is initialized with random value weights. The foregoing subject matter of this paragraph represents Example 14 of the present disclosure, where Example 14 also includes the subject matter according to Example 13 above.
[0018] The process between the first object and the second object includes a coupling process between the first object and the second object. Controlling the coupling process includes: coupling between the first object and the second object when the uncertainty value is at or below a predetermined threshold, and preventing the coupling between the first object and the second object when the uncertainty value is higher than the predetermined threshold. The foregoing subject matter of this paragraph represents Example 15 of the present disclosure, where Example 15 also includes the subject matter according to Example 14 above.
[0019] The present disclosure further discloses a vision-based tracking system. The vision-based tracking system includes a camera configured to generate a two-dimensional (2D) image of at least a portion of a first object, wherein the camera is located on a second object. The vision-based tracking system further includes a processor and a non-transitory computer-readable storage medium storing code. The code is executable by the processor to perform operations including predicting a set of key points on the first object in the 2D image by each of a plurality of key point detectors. The code is further executable by the processor to perform operations including calculating a three-dimensional (3D) pose of each of the plurality of key point detectors from the corresponding set of key points. The code is further executable by the processor to perform operations including deriving a measure of the variation between each of the 3D poses of the plurality of key point detectors and calculating the Euclidean norm of the measure of the variation to produce an uncertainty value. The code is additionally executable by the processor to perform operations including controlling the handling between the first object and the second object in response to the uncertainty value. The foregoing subject matter of this paragraph characterizes Example 16 of the present disclosure.
[0020] Calculating the 3D pose is based on a 2D-to-3D correspondence model. The foregoing subject matter of this paragraph characterizes Example 17 of the present disclosure, wherein Example 17 also includes the subject matter according to Example 16 above.
[0021] The code is executable by the processor to train each of the plurality of key point detectors individually before predicting a set of key points. Each of the plurality of key point detectors has the same architecture, training duration, and training data. Each of the plurality of key point detectors is initialized with random value weights. The foregoing subject matter of this paragraph characterizes Example 18 of the present disclosure, wherein Example 18 further includes the subject matter according to any one of Examples 16-17 above.
[0022] The measure of the variation between each of the 3D poses of the plurality of key point detectors is the standard deviation between each of the 3D poses. The foregoing subject matter of this paragraph characterizes Example 19 of the present disclosure, wherein Example 19 further includes the subject matter according to Example 18 above.
[0023] The handling between the first object and the second object includes a coupling handling between the first object and the second object. Controlling the coupling handling includes: coupling the first object and the second object when the uncertainty value is at or below a predetermined threshold, and preventing the coupling between the first object and the second object when the uncertainty value is higher than the predetermined threshold. The foregoing subject matter of this paragraph characterizes Example 20 of the present disclosure, wherein Example 20 further includes the subject matter according to any one of Examples 16-19 above.
[0024] The described features, structures, advantages, and / or characteristics of the subject matter of the present disclosure may be combined in any suitable manner in one or more examples including embodiments and / or implementations. In the following description, numerous specific details are provided to gain a thorough understanding of examples of the subject matter of the present disclosure. Those skilled in the relevant art will recognize that the subject matter of the present disclosure may be practiced without one or more of the specific features, details, components, materials, and / or methods of a particular example, embodiment, or implementation. In other instances, additional features and advantages may be recognized in certain examples, embodiments, and / or implementations that may not be present in all examples, embodiments, or implementations. Further, in some instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the subject matter of the present disclosure. The features and advantages of the subject matter of the present disclosure will become more apparent from the following description and the appended claims, or may be learned by the practice of the subject matter as set forth below. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] To better understand the advantages of the subject matter, a more specific description of the subject matter briefly described above will be presented by reference to specific examples shown in the accompanying drawings. It should be understood that these drawings depict only typical examples of the subject matter and are not considered to limit its scope. The subject matter will be described and explained with additional features and details by using the drawings, in which:
[0026] Figure 1 is a schematic block diagram showing an embodiment of a vision-based tracking system according to one or more examples of the present disclosure;
[0027] Figure 2 is a schematic perspective view of a two-dimensional image of an aircraft involving a vision-based tracking system according to one or more examples of the present disclosure;
[0028] Figure 3 is a schematic block diagram of a method for estimating uncertainty in a vision-based tracking system according to one or more examples of the present disclosure;
[0029] Figure 4 is a schematic side view of an embodiment of a vision-based tracking system involving an aircraft refueling operation according to one or more examples of the present disclosure;
[0030] Figure 5 is a schematic perspective view of two-dimensional key points projected into a three-dimensional space according to one or more examples of the present disclosure; and
[0031] Figure 6 is a schematic flowchart of a method for estimating uncertainty in a vision-based tracking system according to one or more examples of the present disclosure. DETAILED DESCRIPTION
[0032] References to "an example", "one example", or similar language throughout this specification mean that a particular feature, structure, or characteristic described in connection with the example is included in at least one example of the subject matter of the present disclosure. The appearances of the phrases "in one example", "in an example", and similar language throughout this specification may, but do not necessarily, all refer to the same example. Similarly, the use of the term "implementation" means an implementation having a particular feature, structure, or characteristic described in connection with one or more examples of the subject matter of the present disclosure, however, in the absence of an express correlation to indicate otherwise, an implementation may be associated with one or more examples. Additionally, the features, advantages, and characteristics of the described embodiments may be combined in any suitable manner. Those skilled in the relevant art will recognize that embodiments may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments.
[0033] These features and advantages of these examples will become more apparent from the following description and the appended claims, or may be learned by practice of the examples as set forth hereinafter. As will be understood by those skilled in the art, aspects of the examples of the present disclosure may be embodied as a system, method, and / or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may all be collectively referred to herein as "circuitry", "module", or "system". Additionally, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having program code embodied thereon.
[0034] Examples of methods, systems, and devices for estimating uncertainty within a vision-based tracking system are disclosed herein. Some features of at least some examples of the uncertainty estimation methods, systems, and devices, as well as the accompanying program products, are provided below. In a vision-based tracking system, uncertainty estimation is used to evaluate the accuracy and confidence of the predicted key points within a two-dimensional (2D) to three-dimensional (3D) pose estimation pipeline. Converting a 2D image to 3D pose estimation is susceptible to errors influenced by factors such as specific learning dynamics, unexpected conditions, occlusions, environmental variables, or operator-related errors. By way of non-limiting example, in some examples, environmental variables can include occlusions, self-occlusions, lighting changes, or external disruptions (e.g., adverse weather) captured in the 2D image, which may introduce inaccuracies into the pose estimation. In many cases, including vision-based tracking systems that use neural networks, there is no reliable built-in mechanism to determine the accuracy of the system output (i.e., pose estimation). Therefore, uncertainty estimation can be used to quantify the potential errors in pose estimation. Uncertainty estimation can be used in systems involving manual control, which allows human operators to adjust their trust in the pose estimation. Additionally, uncertainty estimation can also be used for decision-making in automated systems that rely on pose estimation for tasks involving robots or controllers. For example, whether using a manual or an automated system, operations can be initiated or paused in situations where the pose estimation exhibits relatively high or low reliability, respectively.
[0035] Previous methods for estimating uncertainty focused on using a set of randomly perturbed key points to construct a distribution of statistically plausible poses (i.e., Monte Carlo sampling). While this method captures the uncertainty in the key point to 3D pose correspondence, it does not directly evaluate the accuracy of the underlying predicted 2D key points themselves. Therefore, this method is short in dealing with inaccurately predicted key points as it does not directly address the key point quality.
[0036] Thus, the method described herein forms a neural network ensemble by leveraging multiple keypoint detector neural networks to estimate 3D poses and the uncertainty of the underlying predicted 2D keypoints. This neural network ensemble is used to obtain multiple 3D pose estimates, where each of these neural networks predicts a set of keypoints in a 2D image and calculates a corresponding 3D pose estimate based on that set of keypoints. That is, given a predicted set of keypoints, the neural network employs a randomized algorithm that solves for the 3D pose by minimizing the keypoint reprojection error. However, the reprojection error cannot be used as an uncertainty measure because it is a 2D space measurement while the 3D pose exists in 3D space. Thus, employing different sets of keypoint detectors allows for measuring the uncertainty associated with multiple 3D pose estimates. This ensures a more comprehensive and accurate assessment of the reliability of the pose estimates. This method of estimating uncertainty can also be combined with other uncertainty measures (such as Monte Carlo sampling) to identify and mitigate potential errors in other aspects of complex pose estimation scenarios.
[0037] Reference Figure 1 , which is a vision-based tracking system 102 located on a second object 101. As used herein, a vision-based tracking system is a system that employs visual information (typically captured by a camera or other imaging device) to monitor and track an object or subject within a given environment. The vision-based tracking system evaluates the movement, position, and orientation of these objects in real time. For example, vision-based tracking systems can be used in aviation, autonomous vehicles, medical imaging, industrial automation, etc. Vision-based tracking systems enable applications such as in-air refueling operations, object tracking, robot navigation, etc., where understanding the precise position and orientation of objects in a scene is crucial for real-world interactions and decision-making.
[0038] The second object 101 is configured to be positioned relative to the first object, as described below with reference to Figure 2 . The vision-based tracking system 102 includes a processor 104, a memory 106, a camera system 108, and a processing system 110. In some examples, non-transitory computer-readable instructions (i.e., code) stored in the memory 106 (i.e., a storage medium) cause the processor 104 to perform operations, such as operations in a vision-based tracking device (see, for example, Figure 3 ). That is, the code stored in the memory 106 can be executed by the processor 104 to perform operations. The second object 101 can include any object equipped with the camera system 108 or can be the camera system 108 itself. In some examples, the second object 101 is a tanker aircraft, as Figure 4As shown. In such a case, the vision-based tracking system 102 is a refueling system that provides aerial refueling information that can be supplied to a receiver aircraft pilot, boom operators, and / or automated aerial refueling components during the aircraft's approach to a refueling coupling position relative to a tanker. In other examples, the second object 101 is a component in a factory automation process used to determine whether a part or support tool is properly configured and oriented for subsequent manual or automated robotic assembly. In yet other examples, the second object 101 can be any object employed in a vision-based tracking system in which key points are predicted. Although reference is made throughout to an aerial refueling system, it is used merely as an example of the various applications of the vision-based tracking system 102 disclosed herein.
[0039] In various examples, the camera system 108 includes a camera 112, a video image processor 114, and an image creator 116. The camera 112 is any device capable of capturing images or videos and includes one or more lenses, which in some examples may have remotely operated focus and zoom capabilities. The video image processor 114 is configured to process the images or videos captured by the camera 112 and may adjust the focus, zoom, and / or perform other operations to improve image quality. The image creator 116 assembles the processed information from the video image processor 114 into a final image or representation that can be used for various purposes. The camera 112 is mounted to the second object 101 so that the camera 112 is fixed relative to the first object. The camera 112 is fixed at a position where the field of view of the camera 112 includes at least a portion of the first object. In the example of an aerial refueling system, the camera 112 is located behind and below the tanker to capture images of the area below and behind the tanker, such as mounted to a fixed platform within a housing attached to the lower rear fuselage of the tanker.
[0040] In various examples, the processing system 110 includes a sensor 118 and a controller 120. The processing system 110 is configured to control and / or direct a process between a second object 101 and a first object. The process can be a coupling process between the first object and the second object 101, where the first object and the second object 101 are at least temporarily coupled together (e.g., linked). In other examples, the process can be a decoupling process between the first object and the second object 101. For example, the process can be an inspection or quality control process, an assembly process without physical coupling of the first object and the second object 101, an object recognition process, a monitoring or tracking process, etc. The sensor 118 is configured to sense the position of the second object 101 and / or other components such as the first object and rely on the processor 104 for the sensed data. The controller 120 is configured to control the processing system 110 based on signals received from the processor 104, which can include sensed data from the sensor 118. Additionally, the processor 104 generates instructions, such as cues, alerts, or warnings, based on the estimated uncertainty values generated by a plurality of keypoint detectors using data captured by the camera 112. The process of estimating the uncertainty values is described in more detail below with reference to Figure 3 Based on the instructions, the controller 120 notifies, initiates, or pauses the process of the processing system 110.
[0041] Reference Figure 2 , the camera system 108 is configured to generate a two-dimensional (2D) image 200 of a three-dimensional (3D) space. The camera system 108 can produce the 2D image 200 from a single image captured by the camera 112 or extract the 2D image 200 from a video feed of the camera 112. The 2D image 200 can be provided as an RGB image, i.e., an image represented in color using red, green, and blue channels. The 2D image 200 is configured to include at least a portion of the first object 100. In some examples, the 2D image 200 also includes at least a portion of the second object 101. As shown, the 3D space represented in the 2D image 200 includes both a portion of the second object 101 and a portion of the first object 100. For example, as shown, the first object 100 is a tanker aircraft 202 that includes a boom nozzle receiver 208 (i.e., the coupling location). The second object 101 is a refueling aircraft 201 with a deployed refueling boom 204 that can be coupled to the first object 100 at the boom nozzle receiver 208. Thus, during the coupling process, in order to complete the fuel transfer from the refueling aircraft 201, the refueling boom 204 is coupled to the tanker aircraft 202.
[0042] Using a keypoint detector (i.e., a neural network), the positions of a set of keypoints 210 or salient features on a second object 101 are predicted within a 2D image 200. As used herein, a keypoint detector is a particular type of neural network configured to predict the positions of keypoints on an object within a 2D image 200. Each keypoint in the set of keypoints 210 refers to a salient feature, such as distinct and relevant visual elements or points of interest in the 2D image 200. For example, a keypoint can refer to a particular corner, edge, protrusion, or other distinctive feature that aids in identifying and tracking an object within an image. The keypoint detector can also be configured to predict a set of keypoints 220 on a first object 100, thereby allowing the spatial relationship between the first object 100 and the second object 101 to be depicted by the set of keypoints 210 and the set of keypoints 220, respectively. Thus, the set of keypoints can include only the set of keypoints 210 on the first object 100, or the set of keypoints 210 on the first object 100 and the set of keypoints 220 on the second object 101.
[0043] The 2D image 200 captured during real-time conditions includes a background 206 and different environmental variables. The background 206 and environmental variables can include, but are not limited to, occlusion, self-occlusion, lighting changes, or external disruptions (e.g., weather). The background 206 and / or environmental variables can pose challenges to predicting the set of keypoints 210 and the set of keypoints 220 because visual elements in the first object 100 and / or the second object 101 can be washed out, distorted, blurred, etc. Thus, the set of keypoints 210 and / or 220 predicted by the keypoint detector may be inaccurate due to semantic factors such as weather conditions or the presence of interfering objects, which may not be accounted for in the predicted keypoint positions.
[0044] Reference Figure 3 , a vision-based tracking device 300 is shown. The vision-based tracking device 300 is configured by a processor 104 to receive a 2D image 200 of at least a portion of a first object 100 via a camera 112 on a second object 101, predict a set of keypoints 308 - 312 on the first object 100 in the 2D image 200 by each of a plurality of keypoint detectors 302 - 306, calculate a three-dimensional (3D) pose 314 - 318 associated with a corresponding one of the plurality of keypoint detectors 302 - 306 from a corresponding one of the keypoints in the set of keypoints 308 - 312, derive a measure 320 of the change between each of the 3D poses 314 - 318 of the corresponding ones of the plurality of keypoint detectors 302 - 306, and calculate the Euclidean norm of the measure 320 of the change to produce an uncertainty value 322, and control a process 324 between the first object 100 and the second object 101 in response to the uncertainty value 322.
[0045] The vision-based tracking device 300 can be part of a larger management system that can be located on the second object 101, on a remote control system, and / or some combination of both. For example, in the case of in-air refueling operations, the vision-based tracking device 300 can be part of a flight management system located on the tanker 201 and / or on a ground control system.
[0046] The vision-based tracking device 300 is configured to receive a 2D image 200 from a camera 112 located on the second object 101. As described above, the 2D image 200 includes at least a portion of the first object 100 and, in some examples, can include at least a portion of the second object 101.
[0047] The vision-based tracking device 300 includes a plurality of keypoint detectors, such as a first keypoint detector 302, a second keypoint detector 304, and a keypoint detector N 306. Although three keypoint detectors 302 - 306 are shown, any number of keypoint detectors (up to N keypoint detectors) can be utilized as needed. In some examples, the plurality of keypoint detectors can include more or less than the three keypoint detectors 302 - 306 shown. In other examples, the vision-based tracking device 300 includes at least three keypoint detectors. Organizing the plurality of keypoint detectors into a keypoint detector ensemble means that the plurality of keypoint detectors 302 - 306 are configured to jointly make predictions on the visual data captured by the 2D image 200. An ensemble refers to a grouping of multiple keypoint detectors that collaborate to improve the accuracy and performance of keypoint prediction by comparing the results (e.g., uncertainty values) from the multiple keypoint detectors.
[0048] To prepare multiple keypoint detectors to predict keypoints, each of the multiple keypoint detectors 302-306 must undergo a training process. During training, each of the multiple keypoint detectors 302-306 operates independently and does not affect each other in any way. Each of the multiple keypoint detectors 302-306 shares the same network architecture and loss function, and is trained using the same training data for the same training duration. Additionally, each of the multiple keypoint detectors 302-306 is initialized with weights randomly assigned from the same distribution. That is, each of the multiple keypoint detectors 302-306 is diversified by training with randomly valued weights. Notably, each of the multiple keypoint detectors 302-306 uses a different random seed to select these randomly valued weights, ensuring that each keypoint detector starts with a unique initial condition, which helps with diversification within the set of keypoint detectors. After training each of the multiple keypoint detectors 302-306, the multiple keypoint detectors 302-306 are organized into an ensemble. A key feature of the multiple keypoint detectors 302-306 is their ability to produce consistent predictions in the final output (e.g., 3D pose) when the input data (e.g., 2D keypoints) falls within an expected range or is within the domain. Conversely, when the multiple keypoint detectors 302-306 are outside the expected range or are particularly challenging (i.e., out-of-domain), the multiple keypoint detectors 302-306 produce different predictions in the final output.
[0049] Training data is used to train the multiple keypoint detectors 302-306. In some examples, the training data can include multiple training images or a training video feed. The training data includes data representing possible conditions under which a process between a first object and a second object can be performed. For example, the training data can include at least one of the following: a training image including nominal conditions, a training image including at least one occlusion, a training image including at least one self-occlusion, a training image including lighting brighter than nominal lighting, a training image including lighting darker than nominal lighting, and / or a training image including a background mixed with the first object.
[0050] Each of a plurality of key point detectors 302-306 of the set predicts a corresponding one of a set of key points 308-312 in the 2D image 200. In other words, N individual key point detectors are used to extract N independent sets of key point predictions. For example, the first key point detector 302 predicts the first set of key points 308, the second key point detector 304 predicts the second set of key points 310, and the key point detector N 306 predicts a set of key points N 312. The set of key points 308-312 are predictions of the positions of the ground truth key points 230 indicated on the 2D image 200, which ground truth key points 230 include the ground truth key points 230 on the first object 100. In some examples, the ground truth key points 230 include the ground truth key points 230 on the first object 100 and the second object 101. It should be noted that the 2D image 200 does not include the ground truth key points 230, but the ground truth key points 230 are shown on the 2D image for illustrative purposes only, thereby indicating that each of the plurality of key point detectors 302-306 is trained to predict the key point positions. Each of the plurality of key point detectors 302-306 calculates a corresponding one of the 3D poses 314-318 from a corresponding one of the set of key points 308-312. For example, the first key point detector 302 calculates the first 3D pose 314 based on the set of key points 308, the second key point detector 304 calculates the second 3D pose 316 based on the set of key points 310, and the key point detector N 306 calculates the 3D pose N 318 based on the set of key points N 312.
[0051] In some examples, a 2D-to-3D correspondence model is used to generate a corresponding one of the 3D poses 314-318 from a corresponding one of the set of key points 308-312. That is, the correspondence model establishes the correspondence between the 2D image 200 and the 3D real-world object. It enables the vision-based tracking device 300 to determine how each of the predicted set of key points 308-312 in the 2D image 200 relates to a specific point or feature on the 3D object. By establishing these correspondences, each of the plurality of key point detectors 302-306 can estimate the 3D pose of the object based on the predicted set of key points 308-312. As Figure 5 shown, in the 2D space 502 (such as the 2D space in the 2D image 200) with reference to the 2D key points X' 1 、X' 2 、X' i 、X' nA set of key points is represented. The corresponding key point detector then performs 2D-to-3D correspondence by projecting the set of key points into the 3D space 504. That is, perspective-n-point (PnP) pose calculation is used to project each key point in the predicted set of key points from the 2D space 502 into the 3D space 504. Specifically, the 2D key points X' 1 、X' 2 、X' i 、X' n are respectively converted into 3D key points X 1 、X 2 、X i 、and X n .
[0052] Using iterative optimization processing, the PnP pose calculation can be solved. The PnP pose calculation serves as a mathematical framework for calculating the corresponding one of the 3D poses 314-318 of the 2D image 200. Therefore, this calculation ensures that the 3D object view of the camera is closely aligned with the predicted corresponding key points in the set of key points 308-312 in the 2D image 200. In some examples, outlier 2D key points can be removed from the set of key points 308-312 before generating the corresponding one of the 3D poses 314-318. The outlier 2D key points can be manually removed from the set of key points 308-312. In contrast, the PnP solver can handle outliers by removing outliers (such as PnP RANSAC) from a set of key points 308-312, which can be used as a mathematical framework for calculating the corresponding one of the 3D poses 314-318. When the 2D image 200 has poor detections for at least several 2D key points (i.e., poorly detected key points), the algorithm that uses the PnP solver to remove outliers can be useful. The poorly detected key points can be automatically excluded from consideration, thus ensuring that they do not introduce biases into the predicted 3D pose.
[0053] Due to the iterative nature of the optimization process, there will always be a degree of error associated with the final output (i.e., the 3D pose). This error (usually quantified as the Euclidean distance (measured in inches)) represents the difference between the predicted corresponding pose in the 3D pose 314 - 318 and the true 3D pose. Although this error is usually small, typically on the order of about one inch, the error can become significantly larger when the predicted set of key points 302 - 306 is incorrect. In a real - world scenario, the vision - based tracking device 300 lacks prior knowledge of this error during the processing between the first object 100 and the second object 101. Thus, the vision - based tracking device 300 relies on an uncertainty value 322, which serves as a quantifiable measure of the likelihood and degree of error in the 3D pose 314 - 318. The uncertainty value 322 is calculated by deriving a measure 320 of the variation between each of the 3D poses 314 - 318 of the multiple key - point detectors 302 - 306 and computing the Euclidean norm (i.e., the 2 - norm) of the measure 320 of the variation. The measure 320 of the variation can be various measures, including standard deviation or variance. The measure 320 of the variation is derived over the 3D poses 314 - 318 of the multiple key - point detectors 302 - 306. In some examples, each of the 3D poses 314 - 318 has six degrees of freedom, and thus the 3D pose is a six - dimensional vector. That is, the 3D pose is a vector with six values, each of the six values representing the corresponding one of the six degrees of freedom of the 3D pose. Thus, the uncertainty value 322, being a single value, is the Euclidean norm of the six - dimensional measure 320 of the variation over the multiple 3D poses 314 - 318. In other words, in some examples, the uncertainty value is derived as follows: U = Norm(STD([p 1 ,..., p N ), where U is the uncertainty value 322, Norm is the Euclidean norm, STD is the standard deviation, p 1 is the first 3D pose, and p N is the 3D pose N. Additionally, the measure 320 of the variation can be calculated independently for each dimension within the 3D poses 314 - 318. In other words, in some examples, it involves calculating the measure 320 of the variation separately for each of the six dimensions.
[0054] In response to the uncertainty value 322, control the process 324 between the first object 100 and the second object 101. In some examples, the process may be initiated when the uncertainty value 322 is at or below a predetermined threshold. In other words, the uncertainty value 322 represents the likelihood that the error does not exceed the predetermined threshold and thus the process can be safely performed. The initiation of the process 324 may involve allowing the automated process to proceed or signaling to the operator that the conditions are suitable for the manual process to proceed. In contrast, in some examples, when the uncertainty value 322 is above the predetermined threshold, the start of the process 324 is blocked. In other words, the uncertainty value 322 represents the likelihood that the error exceeds the predetermined threshold, and it is not advisable to continue the process. Such blocking may include preventing the start of the process, or pausing or stopping an ongoing process to avoid potential errors. In other examples, the uncertainty value 322 may be provided to the operator, and the decision on whether to continue the process 324 may depend on the operator's judgment.
[0055] As Figure 4 shown, in some examples, the process 324 is a coupling process between a receiver 202 of an in-flight refueling system and a tanker 201. Although an aircraft is shown, it should be understood that the coupling process or close-quarter operations can occur between any first object 100 and second object 101. For example, refueling or close-quarter operations between other vehicles (not just the depicted aircraft 202 and 201). The vehicle can be any vehicle that moves in space (in water, on land, in air, or in space). The vehicle can also be manned or unmanned. In various examples, the vehicle can be a motor vehicle driven by wheels and / or tracks, such as a car, truck, van, etc. The vehicle can also include vessels such as ships, boats, submarines, submersibles, autonomous underwater vehicles (AUVs), etc. In still some other examples, the vehicle can include other manned or unmanned aircraft, such as fixed-wing aircraft, rotary-wing aircraft, and lighter-than-air (LTA) aircraft.
[0056] In some examples, the fuel tanker 201 includes a light array 400 located on the lower front fuselage. The light array 400 is positioned to be clearly visible to the pilot of the receiver aircraft 202, as shown by line 402. The light array 400 includes various lights for providing direction information to the pilot of the receiver aircraft 202. That is, the light array 400 can be used to guide the pilot to position the receiver aircraft 202 relative to the fuel tanker 201 such that the field of view 404 of the camera 112 is aligned for vision-based tracking of the receiver aircraft 202. Thus, as shown, the field of view 404 of the camera 112 includes a view of a portion of the receiver aircraft 202 (including the boom nozzle receiver 208 (i.e., the coupling position)) and a view of a portion of the deployed refueling boom 204 of the fuel tanker 201.
[0057] Based on the uncertainty value, the refueling process between the receiver aircraft 202 and the fuel tanker 201 can be started or stopped. The refueling process can be automatically started or stopped, or the operator can manually control the process based on the uncertainty value.
[0058] Refer to Figure 6 , a method 600 for estimating uncertainty in a vision-based tracking system is shown. The method 600 includes (block 602) receiving a two-dimensional (2D) image 200 of at least a portion of a first object 100 via a camera on a second object 101. In some examples, the 2D image 200 also includes at least a portion of the second object 101. In the case of a coupling process, the 2D image 200 typically includes a coupling position within the 2D image 200. When the 2D image 200 is captured in real time, the 2D image 200 also includes background information and other semantic information.
[0059] The method 600 also includes (block 604) predicting a set of key points 308 - 312 on the first object 100 in the 2D image 200 by each of a plurality of key point detectors 302 - 306. In some examples, the set of key points 308 - 312 also includes key points predicted on the second object 101. Before predicting the key points, the plurality of key point detectors 302 - 306 are trained with the same training data to predict key points. The plurality of key point detectors 302 - 306 are diversified by initializing each of the key point detectors with random value weights via different random seeds.
[0060] Method 600 further includes (block 606) calculating a corresponding one of the three-dimensional (3D) poses 314-318 of each of the plurality of keypoint detectors 302-306 from a corresponding set of keypoints 308-312. In some examples, a 2D-to-3D correspondence model is used to generate a corresponding one of the 3D poses 314-318 from the corresponding set of keypoints 308-312. That is, the correspondence model enables the vision-based tracking system to determine how each of the predicted set of keypoints 308-312 in the 2D image 200 relates to a particular point or feature on the 3D object.
[0061] Method 600 additionally includes (block 608): deriving a measure 320 of the variation between each of the 3D poses 314-318 of the plurality of keypoint detectors 302-306, and calculating the Euclidean norm of the measure 320 of the variation to produce an uncertainty value 322. In some examples, the measure 320 of the variation is the standard deviation calculated over the 3D poses 314-318 of the plurality of keypoint detectors 302-306.
[0062] Method 600 further includes (block 610) controlling the processing 324 between the first object 100 and the second object 101 in response to the uncertainty value 322. In some examples, the processing can be initiated when the uncertainty value 322 is at or below a predetermined threshold. Initiation of the processing 324 can involve allowing the automated processing to proceed or signaling to the operator that the conditions are suitable for manual processing to proceed. In contrast, in some examples, when the uncertainty value 322 is above the predetermined threshold, the processing 324 is prevented from starting. Such prevention can include preventing the start of the processing, or pausing or stopping an ongoing processing to avoid potential errors. In other examples, the uncertainty value 322 can be provided to the operator, and the decision on whether to continue the processing 324 can depend on the operator's judgment.
[0063] As used herein, a computer-readable storage medium may be a tangible device that can store and retain instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory ("RAM"), a read-only memory ("ROM"), an erasable programmable read-only memory ("EPROM" or Flash memory), a static random access memory ("SRAM"), a portable compact disk read-only memory ("CD-ROM"), a digital versatile disk ("DVD"), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable) or an electrical signal transmitted through a wire.
[0064] The computer-readable program instructions described herein may be downloaded to a respective computing / processing device from a computer-readable storage medium via a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network, or to an external computer or external storage device. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0065] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture ("ISA") instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language, etc.). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array ("FPGA"), or a programmable logic array ("PLA"), may execute the computer-readable program instructions by using the state information of the computer-readable program instructions to personalize the electronic circuit so as to perform aspects of the present disclosure.
[0066] The present invention will be described below with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0067] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which instructions cause a computer, a programmable data processing device, and / or other devices to work in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes a manufacture containing instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0068] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0069] The schematic flowcharts and / or schematic block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of devices, systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the schematic flowchart and / or schematic block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function.
[0070] It should also be noted that in some alternative implementations, the functions noted in the boxes may not occur in the order noted in the figures. For example, depending on the functions involved, two boxes shown in succession may in fact be executed substantially simultaneously, or the boxes may sometimes be executed in the reverse order. Other steps and methods may be conceived that are equivalent in function, logic, or effect to one or more boxes or portions thereof of the illustrated figures.
[0071] Although various arrow types and line types may be employed in the flowchart and / or block diagram, they are understood not to limit the scope of the corresponding embodiments. In fact, some arrows or other connectors may be used to indicate only the logical flow of the depicted embodiment. For example, an arrow may indicate an unspecified duration of waiting or monitoring between enumerated steps of the depicted embodiment. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and program code.
[0072] As used herein, a list with an “and / or” conjunction includes any single item in the list or a combination of items in the list. For example, the list of A, B, and / or C includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C. As used herein, a list using the term “one or more of...” includes any single item in the list or a combination of items in the list. For example, one or more of A, B, and C includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C. As used herein, a list using the term “one of...” includes one and only one of any single item in the list. For example, “one of A, B, and C” includes only A, only B, or only C, and does not include the combination of A, B, and C. As used herein, “a member selected from the group consisting of A, B, and C” includes one or only one of A, B, or C, and excludes the combination of A, B, and C. As used herein, “a member selected from the group consisting of A, B, and C and their combinations” includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C.
[0073] In the above description, certain terms may be used, such as “upward”, “downward”, “upper”, “lower”, “horizontal”, “vertical”, “left”, “right”, “above...”, “below...”, etc. Where applicable, these terms are used to provide some clarity in the description when dealing with relative relationships. However, these terms are not intended to imply absolute relationships, positions, and / or orientations. For example, with respect to an object, the “upper” surface can simply be changed to the “lower” surface by turning the object over. Nevertheless, it is still the same object. In addition, unless otherwise expressly specified, the terms “including”, “comprising”, “having” and their variants mean “including but not limited to”. Unless otherwise expressly specified, an enumerated list of items does not imply that any or all of the items are mutually exclusive and / or mutually inclusive. Unless otherwise expressly specified, the terms “a”, “an” and “the” also refer to “one or more”. In addition, the term “plural” can be defined as “at least two”.
[0074] As used herein, when used with a list of items, the phrase "at least one of..." means that different combinations of one or more of the listed items can be used and only one item in the list may be required. The items can be specific objects, things, or categories. In other words, "at least one of..." means that any combination or multiple of the items in the list can be used, but all items in the list may not be required. For example, "at least one of item A, item B, and item C" can mean item A; item A and item B; item B; item A, item B, and item C; and item B and item C. In some cases, "at least one of item A, item B, and item C" can, for example, mean but is not limited to two of item A, one of item B, and ten of item C; four of item B and seven of item C; or some other suitable combination.
[0075] Unless otherwise indicated, the terms "first", "second", etc. are used herein only as labels and are not intended to impose ordinal, positional, or hierarchical requirements on the items to which these terms refer. Further, a reference to, for example, a "second" item does not require or preclude the presence of, for example, a "first" or lower-numbered item and / or a "third" or higher-numbered item.
[0076] As used herein, a system, device, structure, article, element, component, or hardware that is "configured to" perform a specified function is capable of performing the specified function without any change, rather than just having the possibility of performing the specified function after further modification. In other words, a system, device, structure, article, element, component, or hardware that is "configured to" perform a specified function is specifically selected, created, implemented, utilized, programmed, and / or designed for the purpose of performing the specified function. As used herein, "configured to" represents an existing characteristic of a system, device, structure, article, element, component, or hardware that enables the system, device, structure, article, element, component, or hardware to perform the specified function without further modification. For purposes of this disclosure, a system, device, structure, article, element, component, or hardware described as "configured to" perform a particular function may alternatively or additionally be described as "adapted to" and / or "operable to" perform that function.
[0077] The illustrative flowcharts included in this document are generally presented as logical flowcharts. As such, the depicted order and labeled steps indicate an example of the presented method. Other steps and methods may be conceived that are equivalent in function, logic, or effect to one or more steps or portions thereof of the illustrated method. In addition, the format and symbols employed are provided to explain the logical steps of the method and are understood not to limit the scope of the method. Although various arrow types and line types may be employed in the flowchart, they are understood not to limit the scope of the corresponding method. In fact, some arrows or other connectors may be used to indicate only the logical flow of the method. For example, an arrow may indicate a waiting or monitoring period of unspecified duration between enumerated steps of the depicted method. In addition, the order in which a particular method occurs may or may not strictly adhere to the order of the corresponding steps shown.
[0078] Without departing from the spirit or essential characteristics of the subject matter, the subject matter may be embodied in other specific forms. The described examples are considered illustrative rather than restrictive in all respects. All changes that fall within the meaning and scope of equivalents of the examples in this document will be included within their scope.
Claims
1. A method (600) for estimating uncertainty in a vision-based tracking system (102), the method (600) comprising: Receiving (602) a two-dimensional image (200) of at least a portion of the first object (100) via a camera (112) on the second object (101); predicting (604) a set of key points (308-312) on the first object (100) in the two-dimensional image (200) by each of a plurality of key point detectors (302-306); computing (606) a three-dimensional pose (314-318) of each of the plurality of keypoint detectors (302-306) from a corresponding one of the set of keypoints (308-312); deriving (608) a measure of variation (320) between each of the three-dimensional poses of the plurality of keypoint detectors (302-306), and calculating a Euclidean norm of the measure of variation (320) to produce an uncertainty value (322); and A process (324) between the first object (100) and the second object (101) is controlled (610) in response to the uncertainty value (322).
2. The method (600) according to claim 1, wherein: Calculating the 3D pose (314-318) is based on a 2D to 3D correspondence model.
3. The method (600) of claim 1, wherein: Each of the three-dimensional poses (314-318) comprises a vector of six values; Each of the six values represents a corresponding one of the six degrees of freedom of the three-dimensional pose; and Deriving the measure (320) of the variation between each of the three-dimensional poses (314-318) includes deriving the measure (320) of the variation in each vector for each of the six degrees of freedom.
4. The method (600) of claim 1, wherein: The plurality of keypoint detectors (302-306) includes at least three keypoint detectors.
5. The method (600) of claim 1, further comprising: Each of the plurality of keypoint detectors (302-306) is trained individually before predicting the corresponding set of keypoints (308-312), wherein: Each of the plurality of keypoint detectors (302-306) has the same architecture, training duration, and training data; and Each of the plurality of keypoint detectors (302-306) is initialized with a random value weight.
6. The method (600) according to claim 5, wherein: The measure (320) of variation between each of the three-dimensional poses (314-318) of the plurality of keypoint detectors (302-306) is a standard deviation between each of the three-dimensional poses (314-318).
7. The method (600) of claim 1, wherein: The processing (324) between the first object (100) and the second object (101) includes a coupling processing between the first object (100) and the second object (101); and Controlling the coupling process includes: When the uncertainty value (322) is at or below a predetermined threshold, coupling between the first object (100) and the second object (101) is performed; and When the uncertainty value is above the predetermined threshold, the coupling between the first object (100) and the second object (101) is prevented.
8. The method (600) of claim 1, wherein: The processing (324) between the first object (100) and the second object (101) is automatically controlled.
9. The method (600) of claim 1, wherein: The process (324) between the first object (100) and the second object (101) is manually controlled such that the process (324) is initiated by an operator.
10. The method (600) of claim 1, wherein: The two-dimensional image (200) also includes a portion of the second object (101); and Predicting the set of keypoints (308-312) further includes predicting, by each of the plurality of keypoint detectors (308-312), additional keypoints (220) on the second object (101) in the two-dimensional image (200).
11. The method (600) of claim 1, wherein: The first object (101) is a receiving aircraft (202); The second object (100) is a fuel dispenser (201); and The process (324) is a refueling operation between the receiving aircraft (202) and the fueling machine (201).
12. A vision-based tracking device (300), comprising: Processor (104); as well as A non-transitory computer-readable storage medium storing code executable by the processor to perform operations comprising: Receiving a two-dimensional image (200) of at least a portion of the first object (100) via a camera (112) on the second object (101); predicting a set of key points (308-312) on the first object (100) in the two-dimensional image (200) by each of a plurality of key point detectors (302-306); computing a three-dimensional pose (314-318) of each of the plurality of keypoint detectors (302-306) from a corresponding one of the set of keypoints (308-312); deriving a measure (320) of variation between each of the three-dimensional poses (314-318) of the plurality of keypoint detectors (302-306), and calculating a Euclidean norm of the measure (320) of variation to produce an uncertainty value (322); and A process (324) between the first object (100) and the second object (101) is controlled in response to the uncertainty value (322).
13. The vision-based tracking device (300) of claim 12, wherein: Calculating the 3D pose (314-318) is based on a 2D to 3D correspondence model.
14. The vision-based tracking device (300) of claim 13, wherein: The code is executable by the processor to individually train each of the plurality of keypoint detectors (302-306) prior to predicting the set of keypoints (308-312), wherein: Each of the plurality of keypoint detectors (302-306) has the same architecture, training duration, and training data; and Each of the plurality of keypoint detectors (302-306) is initialized with a random value weight.
15. The vision-based tracking device (300) of claim 14, wherein: The processing (324) between the first object (100) and the second object (101) includes a coupling processing between the first object (100) and the second object (101); and Controlling the coupling process includes: When the uncertainty value (322) is at or below a predetermined threshold, coupling between the first object (100) and the second object (101) is performed; and When the uncertainty value is above the predetermined threshold, the coupling between the first object (100) and the second object (101) is prevented.
16. A vision-based tracking system (102), comprising: A camera (112) configured to generate a two-dimensional image (200) of at least a portion of a first object (100), wherein the camera (112) is located on a second object (101); a processor (104); and A non-transitory computer-readable storage medium storing code executable by the processor to perform operations comprising: Predicting a set of key points (308-312) on the first object (100) in the two-dimensional image (200) by each of a plurality of key point detectors (302-306); computing a three-dimensional pose (314-318) of each of the plurality of keypoint detectors (302-306) from the corresponding set of keypoints (308-312); deriving a measure (320) of variation between each of the poses (314-318) of the plurality of keypoint detectors (302-306), and calculating a Euclidean norm of the measure (320) of variation to produce an uncertainty value (322); and A process (324) between the first object (100) and the second object (101) is controlled in response to the uncertainty value (322).
17. The vision-based tracking system (102) of claim 16, wherein: Calculating the 3D pose (314-318) is based on a 2D to 3D correspondence model.
18. The vision-based tracking system (102) of claim 16, wherein: The code is executable by the processor to individually train each of the plurality of keypoint detectors (302-306) prior to predicting the set of keypoints (308-312), wherein: Each of the plurality of keypoint detectors (302-306) has the same architecture, training duration, and training data; and Each of the plurality of keypoint detectors (302-306) is initialized with a random value weight.
19. The vision-based tracking system (102) of claim 18, wherein: The measure (320) of variation between each of the poses (314-318) of the plurality of keypoint detectors (302-306) is a standard deviation between each of the three-dimensional poses (314-318).
20. The vision-based tracking system (102) of claim 16, wherein: The processing (324) between the first object (100) and the second object (101) includes a coupling processing between the first object (100) and the second object (101); and Controlling the coupling process includes: When the uncertainty value (322) is at or below a predetermined threshold, coupling between the first object (100) and the second object (101) is performed; and When the uncertainty value is above the predetermined threshold, the coupling between the first object (100) and the second object (101) is prevented.