Method, device and system for training neural network for vision-based tracking
By using a key point detector in a vision-based tracking system to predict 2D key points and project them into 3D space, combining 3D models and rotation translation information, calculating the loss value and training the optimizer, the problem of neural networks in the prior art poor performance in occlusion and complex environments is solved, and high-precision and adaptable 3D key point detection is achieved.
Patent Information
- Application Number
- CN202411128041.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-08-16
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art faces the problem of poor performance in scenarios with occlusion or environmental variables when training neural networks for vision-based tracking, especially in which 2D key point detectors are difficult to accurately predict key points in these scenarios.
By receiving a two-dimensional image of an object, predicting 2D key points using a key point detector, and projecting it into a 3D space, adding depth information to generate 3D key points. Then, using the object's 3D model and known rotation and translation information, the 3D model keys are adjusted to compare with the predicted 3D keys, the loss value is calculated, and the key point detector is trained by the optimizer to minimize the loss value.
This method improves the accuracy and adaptability of neural networks when dealing with occlusion and complex environment variables, and realizes the ability to effectively process 3D geometric features while high accuracy in 2D key point detection.
Smart Images

Figure CN120070559A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to training neural networks and, more particularly, to training neural networks for vision-based tracking. Background Art
[0002] Neural networks have become important in various applications, including vision-based tracking. Before being used in a vision-based tracking system, a neural network must be trained to perform tasks such as keypoint detection. Generally, neural networks for vision-based tracking systems are divided into two categories: two-dimensional (2D) keypoint detectors, which are trained to predict 2D keypoints in an image; and three-dimensional (3D) keypoint detectors, which are trained to predict 3D keypoints in 3D space. While 2D keypoint detectors, which are typically trained with a single image, exhibit high accuracy in predicting 2D keypoints, they struggle in scenarios with occlusion or environmental variables. On the other hand, 3D keypoint detectors, which are trained differently from 2D keypoint detectors, exhibit inaccuracies in 2D keypoint detection and require multiple object views during training due to the limited depth information available in a single image, thus requiring a more extensive dataset for effective training. Summary of the Invention
[0003] The subject matter of the present application has been developed in response to the prior art and, in particular, in response to problems and needs that have arisen from the typical training of neural networks or that have not been fully addressed. Generally, the subject matter of the present application has been developed to provide a method of training a neural network for prediction that overcomes at least some of the above disadvantages of the prior art.
[0004] A computer-implemented method of training a neural network for use in a vision-based tracking system is disclosed herein. The method includes receiving a two-dimensional (2D) image of at least a portion of an object via a camera fixed relative to the object. The method further includes: predicting, by a keypoint detector, a set of keypoints on the object in the 2D image to generate predicted 2D keypoints. The method further includes: projecting the predicted 2D keypoints into three-dimensional (3D) space and adding keypoint depth information to generate predicted 3D keypoints. The method further includes: using a 3D model of the object and adding known rotation and translation information of the object in the 2D image to produce transformed 3D model keypoints. The method further includes: comparing the predicted 3D keypoints with the transformed 3D model keypoints to calculate a loss value. The method further includes: using an optimizer to train the keypoint detector to minimize the loss value during a training period. The foregoing subject matter of this paragraph represents Example 1 of the present disclosure.
[0005] The camera is configured to be fixed on the second object. The foregoing subject matter of this paragraph represents Example 2 of the present disclosure, where Example 2 further includes the subject matter according to Example 1 as described above.
[0006] The camera includes camera intrinsics. The step of projecting the predicted 2D key points into 3D space depends on knowing at least one of the camera intrinsics. The foregoing subject matter of this paragraph represents Example 3 of the present disclosure, where Example 3 further includes the subject matter according to any one of Examples 1 to 2 as described above.
[0007] The key point depth information is determined at least in part based on the camera intrinsics known to the key point detector. The foregoing subject matter of this paragraph represents Example 4 of the present disclosure, where Example 4 further includes the subject matter according to Example 3 as described above.
[0008] The 3D model of the object is in a fixed position without any rotation or translation. The foregoing subject matter of this paragraph represents Example 5 of the present disclosure, where Example 5 further includes the subject matter according to any one of Examples 1 to 4 as described above.
[0009] The known rotation and translation information of the object is obtained from the real data measured by at least one physical sensor connected to the object. The foregoing subject matter of this paragraph represents Example 6 of the present disclosure, where Example 6 further includes the subject matter according to any one of Examples 1 to 5 as described above.
[0010] The known rotation and translation information of the object is obtained from the simulated data. The known rotation and translation information is predefined in the simulated data. The foregoing subject matter of this paragraph represents Example 7 of the present disclosure, where Example 7 further includes the subject matter according to any one of Examples 1 to 6 as described above.
[0011] The object is a receiver aircraft. The foregoing subject matter of this paragraph represents Example 8 of the present disclosure, where Example 8 further includes the subject matter according to any one of Examples 1 to 7 as described above.
[0012] After training the key point detector, the key point detector is configured to calculate the predicted 3D pose based on the predicted 2D key points. The foregoing subject matter of this paragraph represents Example 9 of the present disclosure, where Example 9 further includes the subject matter according to any one of Examples 1 to 8 as described above.
[0013] The predicted 3D pose is used to guide the refueling operation between the receiver aircraft and the tanker aircraft. The foregoing subject matter of this paragraph represents Example 10 of the present disclosure, where Example 10 further includes the subject matter according to Example 9 as described above.
[0014] The loss value includes a translation loss value and a rotation loss value. The foregoing subject matter of this paragraph represents Example 11 of the present disclosure, where Example 11 further includes the subject matter according to any one of Examples 1 to 10 as described above.
[0015] The loss value includes a combination of at least one of the first loss values and at least one of the second loss values. The first loss values include a translation loss value, a rotation loss value, and a 3D keypoint loss value. The second loss values include a translation loss value, a rotation loss value, a 3D keypoint loss value, and a 2D keypoint loss value. The foregoing subject matter of this paragraph represents Example 12 of the present disclosure, where Example 12 further includes the subject matter according to any one of Examples 1 to 11 above.
[0016] The translation loss value is calculated by computing the absolute difference between the predicted translation vector of the object and the ground truth translation vector of the object. The rotation loss value is calculated by the following operations: by computing the matrix product of the ground truth rotation matrix and the transpose of the predicted rotation matrix, then subtracting the identity matrix to obtain a result, and taking the absolute value of the result. The foregoing subject matter of this paragraph represents Example 13 of the present disclosure, where Example 13 further includes the subject matter according to Example 11 above.
[0017] The present disclosure also discloses a vision-based tracking training device. The device includes at least one processor and a memory device storing instructions. The memory device, when executed by the at least one processor, causes the at least one processor to receive at least a two-dimensional (2D) image of at least a portion of the object via a camera fixed relative to the object. The memory device, when executed by the at least one processor, further causes the at least one processor to predict, by a keypoint detector, a set of keypoints on the object in the 2D image to generate predicted 2D keypoints. The memory device, when executed by the at least one processor, further causes the at least one processor to project the predicted 2D keypoints into a three-dimensional (3D) space and add keypoint depth information to generate predicted 3D keypoints. The memory device, when executed by the at least one processor, additionally causes the at least one processor to generate transformed 3D model keypoints using a 3D model of the object and known rotation and translation information of the object in the 2D image. The memory device, when executed by the at least one processor, further causes the at least one processor to compare the predicted 3D keypoints with the transformed 3D model keypoints to compute a loss value. The memory device, when executed by the at least one processor, further causes the at least one processor to train the keypoint detector using an optimizer to minimize the loss value during a training period. The foregoing subject matter of this paragraph represents Example 14 of the present disclosure.
[0018] The camera includes camera intrinsics. Projecting the predicted 2D keypoints into 3D space depends on knowing at least one of the camera intrinsics. The foregoing subject matter of this paragraph represents Example 15 of the present disclosure, where Example 15 further includes the subject matter according to Example 14 above.
[0019] The loss value includes a translation loss value and a rotation loss value. The foregoing subject matter of this paragraph represents Example 16 of the present disclosure, where Example 16 further includes the subject matter according to any one of Examples 14 to 15 above.
[0020] The translation loss value is calculated by: calculating the absolute difference between the predicted translation vector of the object and the ground truth translation vector of the object. The rotation loss value is calculated by: calculating the matrix product of the ground truth rotation matrix and the transpose of the predicted rotation matrix, then subtracting the identity matrix to obtain a result, and taking the absolute value of the result. The foregoing subject matter of this paragraph represents Example 17 of the present disclosure, where Example 17 further includes the subject matter according to Example 16 above.
[0021] This disclosure also provides a vision-based tracking training system. The vision-based tracking training system includes a camera fixed relative to the object. The vision-based tracking training system further includes at least one processor and a memory device storing instructions. The memory device, when executed by the at least one processor, causes the at least one processor to receive a two-dimensional (2D) image of at least a portion of the object via the camera. The memory device, when executed by the at least one processor, further causes the at least one processor to predict a set of key points on the object in the 2D image by a key point detector to generate predicted 2D key points. The memory device, when executed by the at least one processor, further causes the at least one processor to project the predicted 2D key points into a three-dimensional (3D) space and add key point depth information to generate predicted 3D key points. The memory device, when executed by the at least one processor, further causes the at least one processor to generate transformed 3D model key points using the 3D model of the object and the known rotation and translation information of the object in the 2D image. The memory device, when executed by the at least one processor, further causes the at least one processor to compare the predicted 3D key points with the transformed 3D model key points to calculate a loss value. The memory device, when executed by the at least one processor, further causes the at least one processor to train the key point detector using an optimizer to minimize the loss value during a training period. The foregoing subject matter of this paragraph represents Example 18 of the present disclosure.
[0022] The camera includes camera internal parameters. Projecting the predicted 2D key points into the 3D space depends on knowing at least one of the camera internal parameters. The foregoing subject matter of this paragraph represents Example 19 of the present disclosure, where Example 19 further includes the subject matter according to Example 18 above.
[0023] The loss value includes a translation loss value and a rotation loss value. The foregoing subject matter of this paragraph represents Example 20 of the present disclosure, where Example 20 further includes the subject matter according to any one of Examples 18 to 19 above.
[0024] The described features, structures, advantages, and / or characteristics of the subject matter of the present disclosure may be combined in any suitable manner in one or more examples, including embodiments and / or implementations. In the following description, numerous specific details are provided to afford a thorough understanding of examples of the subject matter of the present disclosure. Those skilled in the relevant art will recognize that the subject matter of the present disclosure may be practiced without one or more of the specific features, details, components, materials, and / or methods of a particular example, embodiment, or implementation. In other instances, additional features and advantages may be recognized in certain examples, embodiments, and / or implementations that may not be present in all examples, embodiments, or implementations. Further, in some instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the subject matter of the present disclosure. The features and advantages of the subject matter of the present disclosure will become more fully apparent from the following description and the appended claims, or may be learned by the practice of the subject matter as set forth below. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] To more readily understand the advantages of the subject matter, a more specific description of the subject matter briefly described above will be presented by reference to specific examples shown in the accompanying drawings. It should be understood that these drawings depict only typical examples of the subject matter and are not considered to limit its scope. The subject matter will be described and explained with additional features and details by using the drawings, in which:
[0026] Figure 1 is a schematic block diagram showing a neural network training system according to one or more examples of the present disclosure;
[0027] Figure 2 is a schematic perspective view of a two-dimensional image used within a neural network training system according to one or more examples of the present disclosure;
[0028] Figure 3 is a schematic side view of a vision-based tracking system using a neural network (trained using a neural network training system) according to one or more examples of the present disclosure;
[0029] Figure 4A is a schematic perspective view of two-dimensional (2D) key points projected onto an object in three-dimensional (3D) space according to one or more examples of the present disclosure;
[0030] Figure 4B is according to one or more examples of the present disclosure having 3D model key points Figure 4B of a 3D model of an object;
[0031] Figure 4C is according to one or more examples of the present disclosure in connection with Figure 4BThe key points of the 3D model compared to the 2D key points projected onto Figure 4A A schematic perspective view of the 2D key points in the 3D space of
[0032] Figure 5 Is a schematic flow chart of a method for training a neural network for use in a vision-based tracking system according to one or more examples of the present disclosure. Detailed Description
[0033] References throughout this specification to "one example", "an example", or similar language mean that a particular feature, structure, or characteristic described in connection with the example is included in at least one example of the subject matter of the present disclosure. The phrases "in one example", "in an example", and similar language throughout this specification may, but do not necessarily, all refer to the same example. Similarly, the use of the term "embodiment" means an embodiment having a particular feature, structure, or characteristic described in connection with one or more examples of the subject matter of the present disclosure, however, without express correlation to indicate otherwise, an embodiment may be associated with one or more examples. Additionally, the features, advantages, and characteristics of the described embodiments may be combined in any suitable manner. Those skilled in the relevant art will recognize that embodiments may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments.
[0034] These features and advantages of the embodiments will become more apparent from the following description and the appended claims, or may be learned by practicing the embodiments described below. As will be understood by those skilled in the art, aspects of the present invention may be embodied as a system, method, and / or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may generally be referred to herein as a "circuit", "module", or "system". Additionally, aspects of the present invention may take the form of a computer program product implemented in one or more computer-readable media having program code embodied thereon.
[0035] Examples of methods, systems, and apparatuses for training neural networks used in vision-based tracking systems are disclosed herein. Some features of at least some examples of the training methods, systems, and apparatuses are provided below. Although neural networks have become important in predictive applications, their ability to adapt and perform optimally across all input data scenarios is constrained by the data inputs provided during their training process. A keypoint detector is trained during a training cycle, which is a specific type of neural network configured to predict the positions of keypoints on an object. Specifically, the methods, systems, and apparatuses leverage the strengths of both conventionally trained 2D and 3D keypoint detectors, providing a training process that achieves high accuracy in 2D keypoint detection and effectively handles occlusions and other challenges by explicitly learning geometric features. Thus, the keypoint detectors disclosed herein are trained using 3D models to improve their proficiency in handling various input data scenarios, including occlusions and other complex variables.
[0036] Training images of an object from a single viewpoint are used to train the keypoint detector. The keypoint detector is trained to predict 2D keypoints on the object within the training images. A loss value is calculated by comparing the predicted 2D keypoints with the ground truth keypoints (i.e., the actual known keypoints), and an optimizer is used to minimize the loss value. Additionally, to improve the performance of the keypoint detector across various input data scenarios, the loss value includes geometric information. In other words, the loss incorporates 3D geometric information alongside the 2D keypoint information. Thus, during training, additional loss terms for rotation and translation are integrated into the keypoint detector by incorporating the 3D model of the object present in the training images.
[0037] Reference Figure 1 , Figure 1 is a neural network training system 102. As used herein, a neural network training system is a system designed to facilitate the training and optimization of a neural network configured to be used within a vision-based tracking system. A vision-based tracking system is a system that employs visual information typically captured by a camera or other imaging device to monitor and track an object or subject within a given environment. A vision-based tracking system evaluates the movement, position, and orientation of these objects in real time. For example, vision-based tracking systems can be used in aviation, autonomous vehicles, medical imaging, industrial automation, etc. A vision-based tracking system is configured to be located on a primary object and evaluate the object being tracked. Thus, before a neural network can be used within a vision-based tracking system, it must undergo training, such as by neural network training system 102.
[0038] The neural network training system 102 includes a processor 104, a memory 106, a camera system 108, and a neural network training device 110. In some examples, non-transitory computer-readable instructions (i.e., code) stored in the memory 106 (i.e., storage medium) cause the processor 104 to perform operations, such as operations in the neural network training device 110.
[0039] In various examples, the camera system 108 includes a camera 112, a video image processor 114, and an image creator 116. The camera 112 is any device capable of capturing an image or video and includes one or more lenses that can have remotely operable focusing and zoom capabilities in some examples. The video image processor 114 is configured to process the image or video captured by the camera 112 and can adjust focus, zoom, and / or perform other operations to improve image quality. The image creator 116 assembles the processed information from the video image processor 114 into a final image or representation that can be used for various purposes. The camera 112 is mounted to the main object such that the camera 112 is fixed relative to the object being tracked. The camera 112 is fixed at a position where the field of view of the camera 112 includes at least a portion of the object being tracked.
[0040] In various examples, the neural network training device 110 is configured by the processor 104 to train a neural network using input data from the camera system 108. The input data can include a 2D image or multiple 2D images, and in some examples, the input data can include a video feed from which 2D images can be generated. The camera system 108 can provide the input data in real time, or the camera system 108 can have previously provided input data stored in the memory 106 for later use by the processor 104. In some examples, the neural network training device 110 includes an image capture module 118, a keypoint detection module 120, a 3D projection module 122, a 3D model integration module 124, a loss calculation module 126, and an optimization module 128.
[0041] The image capture module 118 is configured to receive a 2D image of at least a portion of an object (i.e., the object being tracked). In some examples, the image capture module 118 may transform input data into a 2D image. The image capture module 118 is configured to receive at least one 2D image. In other examples, the image capture module 118 may receive multiple 2D images all from the same viewpoint captured at different times. In other examples, a single 2D image may be enhanced to produce multiple 2D images. The 2D image is received directly or indirectly from the camera system 108 via input data previously stored in the memory 106. Alternatively, the 2D image is a synthetic or simulated 2D image that is artificially generated or enhanced to represent a 2D image that would be obtained from the camera system 108. Synthetic or simulated 2D images may be used in cases where it is difficult or expensive to obtain real-world data generation.
[0042] The keypoint detection module 120 is configured to predict, by a keypoint detector (i.e., a neural network), a set of keypoints on the object being tracked in the 2D image to generate predicted 2D keypoints. That is, using the keypoint detector, the positions of a set of keypoints or salient features on the object being tracked in the 2D image are predicted to generate predicted 2D keypoints. Each keypoint in the predicted 2D keypoints refers to a salient feature, such as a unique and relevant visual element or point of interest in the 2D image. By way of example, a keypoint may refer to a specific corner, edge, protrusion, or other unique feature that aids in identifying and tracking an object within the image. Thus, the keypoint detection module 120 analyzes the visual data in the 2D image and predicts the positions of a set of keypoints based on its analysis. In some examples, the positions of one or more keypoints may be hidden or otherwise occluded in the 2D image, such as by occlusion (including self-occlusion) or environmental variables that affect the image quality of the 2D image. For example, in the case of occlusion (i.e., obscuration), the occlusion may block at least one keypoint from the field of view of the camera 112 such that the position of at least one keypoint is not visible in the 2D image. Additionally, environmental variables such as lighting and weather may result in 2D images where it is difficult to predict keypoint positions. A keypoint detector trained only on 2D input data may not be able to accurately determine the positions of occluded or hidden keypoints within the 2D image. However, in such cases, a 3D model of the object being tracked may be used to improve the accuracy of the keypoint detection module 120. Using the 3D model, the keypoint positions may influence the prediction of other keypoint positions and result in a more accurate prediction of occluded or hidden keypoints.
[0043] The 3D projection module 122 is configured to project the predicted 2D key points from the key point detection module 120 into 3D space, converting the predicted 2D key points from the 2D image into a 3D space context. That is, the predicted 2D key points are converted from a flat 2D image into 3D space. In addition, the 3D projection module 122 adds key point depth information to the predicted 2D key points to generate predicted 3D key points. In other words, the 3D projection module 122 incorporates information about how far each key point is from the viewpoint (i.e., the camera). This additional depth information enables the predicted 2D key points to be configured as predicted 3D key points with rotation and translation information. In some examples, at least one known camera intrinsic parameter can be used to determine the depth information of the predicted 2D key points. Camera intrinsic parameters refer to the parameters of the camera 112 that define its geometry and imaging characteristics. Common camera intrinsic parameters include, but are not limited to, focal length, principal point, lens distortion parameters, image sensor format, camera size, pixel aspect ratio, skew, etc. Using at least one known camera intrinsic parameter, the depth information can be obtained and added to the predicted 2D key points to generate predicted 3D key points. The 3D projection module 122 is equipped with knowledge of at least one camera intrinsic parameter. In other words, within the training pipeline of which the 3D projection module 122 is a part, there is awareness of at least one camera intrinsic parameter. In other examples, the key point depth information can be determined by at least one known camera intrinsic parameter and other depth information (such as depth information provided by a physical sensor).
[0044] The 3D model integration module 124 is configured to provide a 3D model of the tracked object in the 2D image. The 3D model will be provided to the training process of the key point detector based on the geometric parameters of the model of the tracked object. Thus, using the 3D model enables the key point detector to understand the volumetric representation of the tracked object. This understanding allows the key point detector to infer the positions of the key points based on the 3D model, thereby providing insights into where specific key points should be located.
[0045] Known 3D model key points are initially established on a 3D model in model space, lacking translation and rotation information. However, recognizing that the tracked object in the 2D image has some rotation and translation, in order to compare the 3D model and the 2D image, it is necessary to adjust the 3D model and / or the known 3D model key points to include the rotation and translation of the tracked object in the 2D image. Therefore, the known rotation and translation information of the tracked object is added to the 3D model to generate transformed 3D model key points. That is, the transformed 3D model key points are transformed from the initial known 3D model key points into a representation aligned with the changes introduced by the rotation and translation of the tracked object when the 2D image is captured. In some examples, the known rotation and translation information of the tracked object is obtained from real data measured by at least one physical sensor coupled to the tracked object. In other examples, such as when the 2D image is a synthetic or simulated 2D image, the known rotation and translation information is obtained from simulated data. For example, the known rotation and translation information can be predefined in the simulated data.
[0046] The loss calculation module 126 is configured to compare the predicted 3D key points with the transformed 3D model key points to calculate a loss value. Generally, the loss value of the 2D key point detector is calculated by comparing the 2D prediction of the 2D key point detector with the ground truth (i.e., known) 2D key points, which is referred to as the 2D key point loss value. The loss calculation module 126 of the neural network training device 110 modifies this method by adding geometric information in addition to the 2D information to the loss value. That is, during training, a loss term based on geometric information is added to the loss calculation to calculate the loss value of the key point detector. Therefore, the loss value can include any of the following: 1) a translation loss value and a rotation loss value, 2) a translation loss value, a rotation loss value, and at least one other loss term, 3) a 3D key point loss value and at least one other loss term, 4) a translation loss value and at least one other loss term, or 5) a rotation loss value and at least one other loss term. Other loss terms can include but are not limited to translation loss values, rotation loss values, 3D key point loss values, 2D key point loss values, and other standard loss terms.
[0047] The 3D key point loss value is found by minimizing L in the equation L = ||WZ - (RM + T)||, where L is the loss value, W is the projection of the predicted 2D key points into the 3D camera space, Z is the depth information of the predicted 2D key points, R is the rotation of the tracked object, M is the known 3D model key points of the 3D model, and T is the translation of the tracked object. More specifically, the term W only represents the 3D key points in terms of direction as it does not include depth information. The depth information is added to the 3D key points via the term Z, and at least one camera intrinsic parameter is required to determine Z. The term W is represented as a 3xk matrix, where k is the number of key points predicted by the key point detector in the 2D image. The term Z is represented as a k×k diagonal matrix. The term R is represented as a 3x3 rotation matrix, which represents the roll, pitch, and yaw information of the tracked object projected in the 3D space. The term T is represented as a 3x1 translation vector representing the x, y, and z positions of the tracked object projected in the 3D space. In some cases, the term T can also be represented as a 3xN matrix with all columns equal. When no roll, pitch, yaw, or translation (x, y, z) is applied to the 3D model and is represented as a 3xk matrix, the term M represents the known 3D model key points in the model space. Additionally, ||WZ - (RM + T)|| represents the vector or matrix norm of WZ - (RM + T). The norm representation can be any norm function, such as the L2 norm or the frobenius norm.
[0048] For a single 2D image used to train the key point detector, the terms Z, R, M, and T are all ground truth information, i.e., known values (against which other predictions can be evaluated). Specifically, the term M is the ground truth information of the 3D model of the tracked object, while the terms R and T are the ground truth information based on the tracked object projected in the 3D space. The term Z can be calculated using the ground truth 2D key points (i.e., the known 2D key points) before the training of the key point detector. Thus, given a training 2D image, the key point detector is employed during training to calculate the term W by projecting the predicted 2D key points into the 3D camera space. After determining the term W and using the known values of the terms Z, R, M, and T, the loss value of the key point detector trained using the single 2D image and the 3D model can be calculated, thereby incorporating both translation and rotation into the overall calculation.
[0049] Using the predicted 3D key points of the key point detector, the translation loss value predicted by the key point detector can also be calculated, representing the difference between the predicted translation of the 3D key points of the tracked object and the transformed 3D model key points. Specifically, the translation loss value predicted by the key point detector is calculated by computing the absolute difference between the predicted translation vector of the tracked object and the ground truth translation vector of the tracked object. That is, L t = ||T - T gt |, where Lt is the translation loss value, T is the translation of the predicted 3D keypoints (i.e., the predicted translation vector), and T gt is the ground truth translation of the object being tracked. This calculation quantifies how different the predicted 3D keypoints are from the transformed 3D model keypoints in terms of translation. The predicted translation vector T is obtained from the predicted 3D keypoints (WZ), the ground truth rotation (R gt ), and the model keypoints (M). Thus, the predicted translation vector T is not explicitly predicted, but rather the 3D keypoints are predicted and other values are used to determine the predicted translation vector.
[0050] Similarly, using the predicted 3D keypoints of the keypoint detector, the rotation loss value predicted by the keypoint detector can be additionally calculated to represent the difference between the predicted rotation of the 3D keypoints of the object being tracked and the transformed 3D model keypoints. Specifically, the rotation loss value is calculated by the following operations: calculating the matrix product of the ground truth rotation matrix and the transpose of the predicted rotation matrix, then subtracting the identity matrix to obtain the result, and taking the absolute value of the result. That is, L r = ||R gt R T - I||, where L t is the rotation loss value, R gt is the ground truth rotation of the object being tracked, R T is the transpose of the rotation of the predicted 3D keypoints (i.e., the predicted rotation matrix), and I is the identity matrix. This calculation quantifies how different the predicted 3D keypoints are from the transformed 3D model keypoints in terms of rotation. The predicted rotation matrix R is obtained from the predicted 3D keypoints (WZ), the ground truth transformation (T gt ), and the model keypoints (M). Thus, the predicted rotation matrix R is not explicitly predicted, but rather the 3D keypoints are predicted and other values are used to determine the predicted rotation matrix.
[0051] Integrating both the translation and rotation loss values ensures that the keypoint detector predictions are adjusted to the translation and rotation of the 3D model. Thus, in some examples, the keypoint detector is trained to understand the position of the keypoints based on the translation loss value and the rotation loss value. Thus, the keypoint detector uses these values to more accurately predict the position of any keypoints in the 2D image that may be occluded or obscured. That is, the keypoint detector is trained to understand how a particular keypoint should be located in an ideal 3D pose from the 2D keypoints. This training process allows the keypoint detector to adjust future predictions to ensure the best translation and rotation for each keypoint.
[0052] The optimization module 128 is configured to train the keypoint detector using an optimizer to minimize the loss value during a training period. That is, the optimization module 128 refines the predictions of the keypoint detector to minimize the loss value. In some examples, the loss value can include a translation loss value and a rotation loss value. In other examples, the loss value can include other combinations of loss values. By combining both the predicted 2D keypoints and the predicted 3D keypoints, the optimization module 128 ensures that the keypoint detector not only achieves high accuracy in 2D keypoint detection but also adapts to the 3D geometry of the object to more accurately handle various input data scenarios. The optimization module 128 can employ any of a variety of optimizers to minimize the loss value. In some examples, the optimizer utilizes machine learning algorithms to optimize the keypoint detector during the training period.
[0053] After training the keypoint detector, using the neural network training device 110, the keypoint detector can be used within a vision-based tracking system. After being trained using the neural network training device 110, the keypoint detector is better equipped to handle 2D keypoint predictions with increased accuracy and adaptability in different input data scenarios within the vision-based tracking system. Specifically, during the tracking process, the keypoint detector will receive the actual 2D image of the object involved in the vision-based tracking system and predict a set of 2D keypoints on the object in the 2D image. In some examples, the keypoint detector calculates the predicted 3D pose based on the predicted 2D keypoints. The process between the object and a second object, such as a coupling process, will be controlled based on the predicted 3D pose. For example, as Figure 3 shown, the predicted 3D pose can be used to guide the refueling operation between a receiver aircraft and a tanker aircraft.
[0054] As Figure 2As shown, a two-dimensional (2D) image 200 of a three-dimensional (3D) space is generated by a camera system 108, and the resulting image is received by an image capture module 118. In some examples, the 2D image 200 is an actual 2D image captured by the camera system 108. The camera system 108 can generate the 2D image 200 from a single image captured by a camera 112, or extract the 2D image 200 from a video feed of the camera 112. In other examples, the 2D image 200 is an artificially generated or enhanced synthetic or simulated 2D image to represent a 2D image that would be obtained from the camera system 108. The camera system 108 knows at least one camera intrinsic parameter, where the 2D image is an actual 2D image or a simulated 2D image. The 2D image 200 can be provided as an RGB image, i.e., an image represented in color using red, green, and blue channels. The 2D image 200 is configured to include at least a portion of a tracked object 201. That is, it is not necessary to capture the entire tracked object 201 in the 2D image 200. In some examples, the 2D image 200 also includes at least a portion of a primary object 203. As shown, the 3D space represented in the 2D image 200 includes a portion of the primary object 203 and a portion of the tracked object 201. For example, as shown, the tracked object 201 is a tanker aircraft 202 that includes a boom nozzle receiver 208 (i.e., a connection location). The primary object 203 is a refueling tanker 204 with an extended refueling boom 205 that can be connected to the tracked object 201 at the boom nozzle receiver 208. Thus, during a vision-based tracking process, in order to complete fuel delivery from the refueling tanker 204, the refueling boom 205 is connected to the tanker aircraft 202 (see, i.e., Figure 3 ).
[0055] A key point detector is used to predict a set of key points within the 2D image 200 to generate predicted 2D key points. That is, the key point detector is configured to predict a set of key points on the tracked object to generate predicted 2D key points 210. In some examples, a neural network training device 110 can be used to train a second key point detector to predict a set of key points on the primary object 203 using a 3D model of the primary object to generate predicted host 2D key points 220, thereby allowing the spatial relationship between the tracked object 201 and the primary object 203 to be depicted by the predicted 2D key points 210 and the predicted host 2D key points 220, respectively.
[0056] The 2D image 200 includes a background 206 and various environmental variables when captured during real-time conditions. Additionally, the 2D image 200 may also include a simulated background and / or various environmental variables when generated from simulated data. The background 206 and environmental variables may include, but are not limited to, occlusion, self-occlusion, lighting changes, lens flare, or external interruptions (e.g., weather). The background 206 and / or environmental variables may pose challenges to predicting the predicted 2D key points 210 because visual elements in the tracked object 201 may be cleared, distorted, blurred, etc. Thus, the predicted 2D key points 210 predicted by the key point detector may be inaccurate due to semantic factors such as weather conditions or the presence of interfering objects, which may not be considered in the predicted key point positions.
[0057] Figure 3 Shown is an example of a vision-based tracking system configured to control Figure 2 the refueling operation (e.g., coupling process) between the receiver aircraft 202 and the extended refueling boom 205 of the tanker aircraft 204 shown in. After training the vision-based tracking system to predict key points using the neural network training device 110, the vision-based tracking system is used to perform a process such as the coupling process. Thus, the 2D image 200 is generated by or simulated based on the camera 112 for use during the training process of the key point detector. The camera 112 is fixed to the tanker aircraft 204 because the field of view 304 of the camera 112 includes a view of a portion of the receiver aircraft 202, including the boom nozzle receiver 208 or the coupling location, and a view of a portion of the extended refueling boom 205 of the tanker aircraft 204. The camera 112 is fixed relative to the receiver aircraft 202 (i.e., the tracked object) such that the camera 112 internal parameters can be used to determine the depth information of the predicted 2D key points projected into 3D space.
[0058] Although shown with aircraft, it should be understood that predicting key points using a key point detector can be used with any tracked object 201 and primary object 203. Specifically, a coupling process or close operation such as shown with aircraft (receiver aircraft 202 and tanker aircraft 204) can occur between any tracked object 201 and primary object 203 (such as other vehicles). The vehicle can be any vehicle that moves in space (in water, on land, in the air, or in space). In other examples, the vision-based tracking system can be used in any system that uses visual information typically captured by a camera or other imaging device to monitor and track objects or subjects within a given environment. For example, the vision-based tracking system can be used in aviation, autonomous vehicles, medical imaging, industrial automation, etc., and for applications such as in-air refueling processes, object tracking, robot navigation, etc., where understanding the precise position and orientation of objects in the scene is crucial for real-world interactions and decision-making.
[0059] As Figure 4A shown, by way of example only, the predicted 2D key points 210 are represented by predicted 2D key points X’ 1 , X’ 2 , X’ 3 , X’ n within the 2D image 200. Although four predicted key points are shown, any number of key points, up to N key points, can be predicted as needed. The predicted 2D key points 210 are projected from the camera 112 into the 3D camera space 404 as initial 3D key points 407. The initial 3D key points 407 are the projection of the predicted 2D key points 210 into the 3D camera space 404 that has direction information but no depth information. Using the loss value equation L = ||WZ - RM + T)|, the term W represents the initial 3D key points 407. In other words, the initial 3D key points 407 are depthless 3D vectors. Therefore, the initial 3D key points 407 are modified with depth information to generate the predicted 3D key points 408. The term WZ represents the predicted 3D key points 408, which are the initial 3D key points modified by the term Z or depth information. The predicted 3D key points 408 are represented by predicted 3D key points X 1 , X 2 , X 3 , X n , where the predicted 2D key point X’ 1 is projected and modified to generate the predicted 3D key point X 1 , the predicted 2D key point X’ 2 is projected and modified to generate the predicted 3D key point X 2 , the predicted 2D key point X’ 3 is projected and modified to generate the predicted 3D key point X 3 , and the predicted 2D key point X’ n is projected and modified to generate the predicted 3D key point X n .
[0060] Refer to Figure 4B, shows a 3D model 410 represented by a box shape. The 3D model 410 is a 3D model of the tracked object in the 2D image 200. The geometry of the 3D model 410 is known to the neural network training device 110 (i.e., the training pipeline), enabling the training pipeline to understand the volume representation of the tracked object in the 2D image 200. In other words, the 3D model of the tracked object is incorporated into the learning process of the key point detector, such that in addition to learning the key points based on the 2D positions of the key points, the key point detector also learns the key point positions based on the 3D structure of the model. Any 3D modeling method such as a CAD model can be used to generate the 3D model 410. The known 3D model key points 411 are established on the 3D model 410 and lack translation and rotation information. The known 3D model key points 411 are represented by the known 3D model key points M’ 1 , M’ 2 , M’ 3 , M’ n . The term M represents the known 3D model key points 411 in the model space.
[0061] Reference Figure 4C , the known 3D model key points 411 are transformed into transformed 3D model key points 412 by adding the known rotation and translation information of the tracked object in the 2D image 200. The transformed 3D model key points 412 are represented by the transformed 3D model key points M 1 , M 2 , M 3 , M n , where the known 3D model key point M’ 1 is transformed to generate the transformed 3D model key point M 1 , the known 3D model key point M’ 2 is transformed to generate the transformed 3D model key point M 2 , the known 3D model key point M’ 3 is transformed to generate the transformed 3D model key point M 3 , and the known 3D model key point M’ n is transformed to generate the transformed 3D model key point M n . That is, the transformed 3D model key points are obtained by transforming the initial known 3D model key points into a representation aligned with the changes introduced by the rotation and translation of the tracked object when capturing the 2D image. The terms R and T represent the known rotation and translation information respectively.
[0062] During the training of the key point detector, the goal is to minimize the loss value (term L). In other words, the key point detector aims to align term WZ as closely as possible with term RM + T. Therefore, the key point detector attempts to achieve a 3D key point 408 that is approximately equal to the prediction of the 3D model key point 412 of the transformation. For example, such that the predicted 3D key point X 1 is approximately equal to the transformed 3D model key point M 1 , the predicted 3D key point X 2 is approximately equal to the transformed 3D model key point M 2 , the predicted 3D key point X 3 is approximately equal to the transformed 3D model key point M 3 , and the predicted 3D key point X n is approximately equal to the transformed 3D model key point M n .
[0063] Reference Figure 5 , shows a method 500 for training a neural network (i.e., a key point detector) used in a vision-based tracking system. Method 500 includes (block 502) receiving a two-dimensional (2D) image of at least a portion of an object via a camera fixed relative to the object. The camera includes the camera intrinsics known to the key point detector. Method 500 also includes (block 504) predicting, by the key point detector, a set of key points on the object in the 2D image to generate predicted 2D key points. Method 500 also includes (block 506) projecting the predicted 2D key points into a three-dimensional (3D) space and adding key point depth information to generate predicted 3D key points. Initially, the projected 2D key points lack depth information and are thus projected into the camera space as 3D vectors. Therefore, the depth information determined using the known camera intrinsics is added to the projected 2D key points.
[0064] Method 500 also includes (block 508) using a 3D model of the object and adding known rotation and translation information of the object in the 2D image to generate transformed 3D model key points. The known rotation and translation information can be known and predefined in the simulated data or can be measured by physical sensors in the actual data. Method 500 also includes (block 510) comparing the predicted 3D key points with the transformed 3D model key points to calculate a loss value. Method 500 additionally includes (block 512) using an optimizer to train the key point detector to minimize the loss value during a training period.
[0065] This application relates to the following clauses:
[0066] Clause 1. A computer-implemented method (500) for training a neural network used in a vision-based tracking system (102), the method comprising:
[0067] Receiving (502) a two-dimensional image (200) of at least a portion of the object (201) via a camera (112) fixed relative to the object (201);
[0068] Predicting (504), by a key point detector, a set of key points on the object (201) in the two-dimensional image (200) to generate predicted two-dimensional key points (210);
[0069] Projecting (506) the predicted two-dimensional key points (210) into a three-dimensional space (404) and adding key point depth information to generate predicted three-dimensional key points (408);
[0070] Using (508) a three-dimensional model (410) of the object (201) and adding known rotation and translation information of the object (201) in the two-dimensional image (200) to generate transformed three-dimensional model key points (412);
[0071] Comparing (510) the predicted three-dimensional key points (408) with the transformed three-dimensional model key points (412) to calculate a loss value; and
[0072] Using (512) an optimizer to train the key point detector to minimize the loss value during a training period.
[0073] Clause 2. The computer-implemented method (500) according to Clause 1, wherein the camera (112) is configured to be fixed on a main object (203).
[0074] Clause 3. The computer-implemented method (500) according to Clause 1, wherein:
[0075] The camera (112) includes camera internal parameters; and
[0076] The step of projecting the predicted two-dimensional key points into the three-dimensional space depends on knowing at least one of the camera internal parameters to project the predicted two-dimensional key points into the three-dimensional space.
[0077] Clause 4. The computer-implemented method (500) according to Clause 3, wherein the key point depth information is determined at least in part based on the camera internal parameters known to the key point detector.
[0078] Clause 5. The computer-implemented method (500) according to Clause 1, wherein the three-dimensional model (410) of the object (201) is in a fixed position without any rotation or translation.
[0079] Clause 6. The computer-implemented method (500) according to Clause 1, wherein the known rotation and translation information of the object (201) is obtained from real data measured by at least one physical sensor coupled to the object (201).
[0080] Clause 7. The computer-implemented method (500) according to Clause 1, wherein the known rotation and translation information of the object (201) is obtained from simulation data; and
[0081] the known rotation and translation information is predefined in the simulation data.
[0082] Clause 8. The computer-implemented method (500) according to Clause 1, wherein the object (201) is a receiver aircraft.
[0083] Clause 9. The computer-implemented method (500) according to Clause 1, wherein after training the key point detector, the key point detector is configured to calculate a predicted three-dimensional pose based on the predicted two-dimensional key points.
[0084] Clause 10. The computer-implemented method (500) according to Clause 9, wherein the predicted three-dimensional pose is used to guide a refueling operation between a receiver aircraft and a tanker aircraft.
[0085] Clause 11. The computer-implemented method (500) according to Clause 1, wherein the loss value includes a translation loss value and a rotation loss value.
[0086] Clause 12. The computer-implemented method (500) according to Clause 1, wherein:
[0087] the loss value includes a combination of at least one of a first loss value and at least one of a second loss value, wherein:
[0088] the first loss value includes:
[0089] a translation loss value;
[0090] a rotation loss value; and
[0091] a three-dimensional key point loss value; and
[0092] the second loss value includes:
[0093] a translation loss value;
[0094] a rotation loss value;
[0095] a three-dimensional key point loss value; and
[0096] a two-dimensional key point loss value.
[0097] Clause 13. The computer-implemented method (500) according to Clause 11, wherein:
[0098] The translational loss value is calculated by computing the absolute difference between the predicted translational vector of the object and the ground truth translational vector of the object; and
[0099] The rotational loss value is calculated by: computing the matrix product of the ground truth rotation matrix and the transpose of the predicted rotation matrix, then subtracting the identity matrix to obtain a result, and taking the absolute value of the result.
[0100] Clause 14. A vision-based tracking training device, the vision-based tracking training device comprising:
[0101] At least one processor (104); and
[0102] A memory device (106) storing instructions which, when executed by the at least one processor (104), cause the at least one processor (104) to at least:
[0103] Receive (502) a two-dimensional image (200) of at least a portion of the object (201) via a camera (112) fixed relative to the object (201);
[0104] Predict, by a key point detector, a set of key points on the object (201) in the two-dimensional image (200) to generate predicted two-dimensional key points (210);
[0105] Project the predicted two-dimensional key points (210) into a three-dimensional space (404) and add key point depth information to generate predicted three-dimensional key points (408);
[0106] Utilize a three-dimensional model (410) of the object (201) and add known rotation and translation information of the object (201) in the two-dimensional image (200) to generate transformed three-dimensional model key points (412); compare the predicted three-dimensional key points (408) with the transformed three-dimensional model key points (412) to compute a loss value; and
[0107] Utilize an optimizer to train the key point detector to minimize the loss value during a training period.
[0108] Clause 15. The vision-based tracking training device according to Clause 14, wherein:
[0109] The camera (112) includes camera intrinsic parameters; and
[0110] Projecting the predicted 2D key points (210) into a three-dimensional space depends on knowing at least one of the camera internal parameters.
[0111] Clause 16. The device according to clause 14, wherein the loss value includes a translation loss value and a rotation loss value.
[0112] Clause 17. The vision-based tracking training device according to clause 16, wherein:
[0113] The translation loss value is calculated by computing the absolute difference between the predicted translation vector of the object and the ground truth translation vector of the object; and
[0114] The rotation loss value is calculated by the following operations: computing the matrix product of the ground truth rotation matrix and the transpose of the predicted rotation matrix, then subtracting the identity matrix to obtain a result, and taking the absolute value of the result.
[0115] Clause 18. A vision-based tracking training system (102), the vision-based tracking training system comprising:
[0116] A camera (112) that is fixed relative to an object (201);
[0117] At least one processor (104); and
[0118] A memory device (106) that stores instructions which, when executed by the at least one processor (104), cause the at least one processor (104) to at least:
[0119] Receive a two-dimensional image (200) of at least a portion of the object (201) via the camera (112);
[0120] Predict, by a key point detector, a set of key points on the object (201) in the two-dimensional image (200) to generate predicted 2D key points (210);
[0121] Project the predicted 2D key points (210) into a three-dimensional space (404) and add key point depth information to generate predicted 3D key points (408);
[0122] Utilize a three-dimensional model (410) of the object (201) and add known rotation and translation information of the object (201) in the two-dimensional image (200) to generate transformed three-dimensional model key points (412);
[0123] Compare the predicted 3D key points (408) with the transformed three-dimensional model key points (412) to calculate a loss value; and
[0124] An optimizer is used to train the key point detector to minimize the loss value during a training period.
[0125] Clause 19. The vision-based tracking training system (102) according to Clause 18, wherein:
[0126] The camera (112) includes camera internal parameters; and
[0127] Projecting the predicted two-dimensional key points (210) into a three-dimensional space depends on knowing at least one of the camera internal parameters.
[0128] Clause 20. The vision-based tracking training system (102) according to Clause 18, wherein the loss value includes a translation loss value and a rotation loss value.
[0129] As cited herein, a computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory ("RAM"), a read-only memory ("ROM"), an erasable programmable read-only memory ("EPROM" or flash memory), a static random access memory ("SRAM"), a portable compact disc read-only memory ("CD-ROM"), a digital versatile disc ("DVD"), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium as used herein should not be construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable) or an electrical signal transmitted through a wire.
[0130] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0131] The computer-readable program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture ("ISA") instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on a user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network connection, including a local area network ("LAN") or a wide area network ("WAN"), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array ("FPGA"), or a programmable logic array ("PLA"), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit so as to perform aspects of the present invention.
[0132] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0133] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0134] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0135] The schematic flowcharts and / or schematic block diagrams in the figures illustrate the possible architectures, functions, and operations of apparatuses, systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the schematic flowchart and / or schematic block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function.
[0136] It should also be noted that in some alternative embodiments, the functions noted in the boxes may not occur in the order noted in the figures. For example, depending on the functions involved, two boxes shown in succession may in fact be executed substantially simultaneously, or the boxes may sometimes be executed in the reverse order. Other steps and methods may be envisioned that are equivalent in function, logic, or effect to one or more boxes or portions thereof shown in the figures.
[0137] Although various arrow types and line types may be employed in the flowchart and / or block diagram, they are understood not to limit the scope of the corresponding embodiments. In fact, some arrows or other connectors may be used to indicate only the logical flow of the depicted embodiments. For example, an arrow may indicate a waiting or monitoring period of unspecified duration between the enumerated steps of the depicted embodiment. It should also be noted that each box of the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, may be implemented by a system based on dedicated hardware that performs the specified function or act, or by a combination of dedicated hardware and program code.
[0138] As used herein, a list with a conjunction of "and / or" includes any single item in the list or a combination of items in the list. For example, the list of A, B, and / or C includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C. As used herein, a list using the term "one or more of" includes any single item in the list or a combination of items in the list. For example, one or more of A, B, and C includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C. As used herein, a list using the term "one" includes one and only one of any single item in the list. For example, "one of A, B, and C" includes only A, only B, or only C, and does not include the combination of A, B, and C. As used herein, "a member selected from the group consisting of A, B, and C" includes one and only one of A, B, or C, and excludes the combination of A, B, and C. As used herein, "a member selected from the group consisting of A, B, and C and combinations thereof" includes only A, only B, only C, the combination of A and B, the combination of B and C, the combination of A and C, or the combination of A, B, and C.
[0139] In the above description, certain terms may be used, such as "upper", "lower", "above", "below", "horizontal", "vertical", "left", "right", "up", "down", etc. These terms are used, where applicable, to provide some clarity in dealing with relative relationships. However, these terms are not intended to imply absolute relationships, positions, and / or orientations. For example, with respect to an object, the "upper" surface can simply become the "lower" surface by flipping the object. However, it is still the same object. Additionally, unless otherwise expressly stated, the terms "comprise", "include", "have", and their variants mean "include but not limited to". Unless otherwise expressly stated, a list of enumerated items does not mean that any or all of the items are mutually exclusive and / or mutually inclusive. Unless otherwise expressly stated, the terms "a", "an", and "the" also refer to "one or more". Additionally, the term "plural" may be defined as "at least two".
[0140] As used herein, when used in conjunction with a list of items, the phrase "at least one" means that different combinations of one or more of the listed items can be used and that only one of the items in the list may be required. The items can be specific objects, things, or categories. In other words, "at least one of" means that any combination of the items in the list or number of items can be used, but not all of the items in the list may be required. For example, "at least one of item A, item B, and item C" can mean item A; item A and item B; item B; item A, item B, and item C; or item B and item C. In some cases, "at least one of item A, item B, and item C" can mean, for example but not limited to, two of item A, one of item B, and ten of item C; four of item B and seven of item C; or some other suitable combination.
[0141] Unless otherwise specified, the terms "first", "second", etc. are used herein only as labels and are not intended to impose an order, position, or hierarchical requirement on the items to which these terms refer. Additionally, a reference to, for example, a "second" item does not require or preclude the presence of, for example, a "first" or lower-numbered item and / or a "third" or higher-numbered item.
[0142] As used herein, a system, apparatus, structure, article, element, component, or hardware "configured to" perform a specified function is capable of performing the specified function without any change, rather than merely having the potential to perform the specified function after further modification. In other words, for the purpose of performing the specified function, a system, apparatus, structure, article, element, component, or hardware "configured to" perform the specified function is specifically selected, created, implemented, utilized, programmed, and / or designed. As used herein, "configured to" represents an existing characteristic of a system, apparatus, structure, article, element, component, or hardware that enables the system, apparatus, structure, article, element, component, or hardware to perform the specified function without further modification. For the purposes of this disclosure, a system, apparatus, structure, article, element, component, or hardware described as "configured to" perform a particular function may additionally or alternatively be described as "adapted to" and / or "operable to" perform that function.
[0143] The schematic flowcharts included in this document are generally presented as logical flowcharts. Accordingly, the depicted order and labeled steps represent an example of the presented method. Other steps and methods can be envisioned that are equivalent in function, logic, or effect to one or more steps or portions thereof of the illustrated method. Additionally, the format and symbols employed are provided to explain the logical steps of the method and are understood not to limit the scope of the method. Although various arrow types and line types can be used in the flowchart, they are understood not to limit the scope of the corresponding method. In fact, some arrows or other connectors can be used to indicate only the logical flow of the method. For example, an arrow can indicate a waiting or monitoring period of unspecified duration between enumerated steps of the depicted method. Additionally, the order in which a particular method occurs can or can not strictly adhere to the order of the corresponding steps shown.
[0144] Without departing from the spirit or essential characteristics of the subject matter, the subject matter can be embodied in other specific forms. The described examples are to be considered in all respects only illustrative and not restrictive. All changes that come within the meaning and range of equivalents of the examples in this document are to be embraced within its scope.
Claims
1. A computer-implemented method (500) of training a neural network for use in a vision-based tracking system (102), the method comprising: Receiving (502) a two-dimensional image (200) of at least a portion of an object (201) via a camera (112) fixed relative to the object (201); predicting (504) a set of key points on the object (201) in the two-dimensional image (200) by a key point detector to generate predicted two-dimensional key points (210); Projecting (506) the predicted two-dimensional keypoints (210) into three-dimensional space (404) and adding keypoint depth information to generate predicted three-dimensional keypoints (408); Using (508) the three-dimensional model (410) of the object (201) and adding known rotation and translation information of the object (201) in the two-dimensional image (200) to generate transformed three-dimensional model key points (412); comparing (510) the predicted 3D keypoints (408) with the transformed 3D model keypoints (412) to calculate a loss value; as well as The keypoint detector is trained using (512) an optimizer to minimize the loss value during a training period.
2. The computer-implemented method (500) of claim 1, wherein: The camera (112) is configured to be fixed on a main object (203).
3. The computer-implemented method (500) of claim 1, wherein: The camera (112) includes camera intrinsic parameters; and The step of projecting the predicted two-dimensional keypoints into three-dimensional space relies on knowing at least one of the camera intrinsic parameters to project the predicted two-dimensional keypoints into three-dimensional space.
4. The computer-implemented method (500) of claim 3, wherein: The keypoint depth information is determined based at least in part on the camera intrinsic parameters known to the keypoint detector.
5. The computer-implemented method (500) of claim 1, wherein: The three-dimensional model (410) of the object (201) is in a fixed position without any rotation or translation.
6. The computer-implemented method (500) of claim 1, wherein: The known rotation and translation information of the object (201) is obtained from real data measured by at least one physical sensor coupled to the object (201).
7. The computer-implemented method (500) of claim 1, wherein: The known rotation and translation information of the object (201) is obtained from simulation data; and The known rotation and translation information is predefined in the simulation data.
8. The computer-implemented method (500) of claim 1, wherein: The object (201) is a receiving aircraft.
9. A vision-based tracking training device, the vision-based tracking training device comprising: at least one processor (104); as well as A memory device (106) storing instructions that, when executed by the at least one processor (104), cause the at least one processor (104) to at least: Receiving (502) a two-dimensional image (200) of at least a portion of an object (201) via a camera (112) fixed relative to the object (201); A key point detector is used to predict a set of key points on the object (201) in the two-dimensional image (200) to generate predicted two-dimensional key points (210); Projecting the predicted two-dimensional key points (210) into a three-dimensional space (404) and adding key point depth information to generate predicted three-dimensional key points (408); Utilizing a three-dimensional model (410) of the object (201) and adding known rotation and translation information of the object (201) in the two-dimensional image (200) to generate transformed three-dimensional model key points (412); Comparing the predicted 3D keypoints (408) with the transformed 3D model keypoints (412) to calculate a loss value; as well as The keypoint detector is trained using an optimizer to minimize the loss value during a training period.
10. A vision-based tracking training system (102), the vision-based tracking training system comprising: a camera (112) fixed relative to the object (201); at least one processor (104); as well as A memory device (106) storing instructions that, when executed by the at least one processor (104), cause the at least one processor (104) to at least: receiving, via the camera (112), a two-dimensional image (200) of at least a portion of the object (201); A key point detector is used to predict a set of key points on the object (201) in the two-dimensional image (200) to generate predicted two-dimensional key points (210); Projecting the predicted two-dimensional key points (210) into a three-dimensional space (404) and adding key point depth information to generate predicted three-dimensional key points (408); Utilizing a three-dimensional model (410) of the object (201) and adding known rotation and translation information of the object (201) in the two-dimensional image (200) to generate transformed three-dimensional model key points (412); Comparing the predicted 3D keypoints (408) with the transformed 3D model keypoints (412) to calculate a loss value; as well as The keypoint detector is trained using an optimizer to minimize the loss value during a training period.