Model-free three-dimensional object tracking method based on differential rendering
Through the combination of differentiable rendering module and tracking module, the three-dimensional model of the object is inversely rendered, which solves the efficient tracking problem in model-free scenarios, and achieves the improvement of high precision and real-time performance. It is suitable for a variety of complex scenarios without the need for specific data set training, reducing costs.
Patent Information
- Application Number
- CN202510204145.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-25
AI Technical Summary
The dependence of existing 3D object tracking technology on specific 3D models or training data limits its application in scenarios where detailed models are lacking, and model-free methods have insufficient generalization capabilities, speed and accuracy, and require specific data set training, which increases costs and limits flexibility.
The model-free three-dimensional object tracking method based on differentiable rendering is adopted. By specifying the target object in the RGB image and using the differentiable rendering module to inversely render the initial three-dimensional model, combined with the tracking module for continuous tracking and iterative optimization, the available three-dimensional model is generated, suitable for multiple scenarios without the need for specific data set training.
High-precision real-time three-dimensional object tracking under limited or no prior information conditions is realized, which improves tracking accuracy and real-time performance, and reduces application costs. The generated three-dimensional model can be reused by other methods, has strong adaptability and is suitable for a variety of complex scenarios.
Smart Images

Figure CN120374670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a model-free three-dimensional object tracking method based on differentiable rendering, belonging to the technical field of object tracking. Background Art
[0002] Three-dimensional object tracking based on monocular RGB is a key technology that can continuously estimate the 6-degree-of-freedom pose of an object relative to a camera in a dynamic scene. Although significant progress has been made in three-dimensional object tracking technology, its dependence on specific three-dimensional models or training data remains a significant limitation. Many model-based tracking methods have achieved high-precision tracking results, but these methods usually require an accurate three-dimensional CAD model as a prior condition [Wang L, Yan S, Zhen J, Liu Y, Zhang M, Zhang G, Zhou X. Deep Active Contours for Real-time 6-DoF Object Tracking. Proceedings of the IEEE International Conference on Computer Vision, 2023: 13988-13998], [Tian X, Lin X, Zhong F, Qin X. Large-Displacement 3D Object Tracking with Hybrid Non-local Optimization. European Conference on Computer Vision, 2022, 13682: 627-643], [Stoiber M, Sundermeyer M, Triebel R. Iterative Corresponding Geometry: Fusing Region and Depth for Highly Efficient 3D Tracking of Textureless Objects. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2022: 6855-6865]. This dependence limits their application in scenarios lacking detailed models.
[0003] To address this issue, some researchers have developed model-free methods based on deep learning, which attempt to reduce the dependence on 3D models by using 2D priors, such as [Lee J, Cabon Y, Brégier R, Yoo S, Revaud J. MFOS: Model-Free&One-Shot Object Pose Estimation. Proceedings of theAAAI Conference on Artificial Intelligence, 2024: 2911-2919], [He X, Sun J, WangY, Huang D, Bao H, Zhou X. OnePose++: Keypoint-Free One-Shot Object PoseEstimation without CAD Models. Advances in Neural Information ProcessingSystems, 2022.], [Liu Y, Wen Y, Peng S, Lin C, Long X, Komura T, Wang W. Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGBImages. Proceedings of the European Conference on Computer Vision, LectureNotes in Computer Science, 2022, 13692: 298-315]. However, these model-free methods face numerous challenges in practical applications. First, they have deficiencies in generalization ability and are difficult to adapt to diverse environments and different types of objects. Second, in terms of speed and accuracy, these methods often cannot compare with model-based methods. In addition, these methods usually require a large amount of training on specific datasets, which not only increases the development and deployment costs but also limits their flexibility in different application scenarios. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a model-free 3D object tracking method based on differentiable rendering, which solves the problem of high-precision real-time tracking of 3D objects in different scenarios with only limited prior information or no prior information, and moreover, does not require training on a specific dataset, can quickly generate available 3D models in various application scenarios, and achieve efficient six-degree-of-freedom pose tracking.
[0005] The technical solution of the present invention is as follows:
[0006] A model-free 3D object tracking method based on differentiable rendering, comprising the following steps:
[0007] S1: Specify the target object in the RGB image of the first frame using a 2D bounding box, and provide its initial pose and camera intrinsics as the input to the differentiable rendering module, where the camera intrinsics can be calibrated in advance. When there is a reference frame, the reference frame and its corresponding pose and 2D bounding box are also used as inputs;
[0008] S2: Segment the target object in the input frame to obtain its 2D mask;
[0009] S3: Start the differentiable rendering module, inverse render an initial 3D model, and transfer the 3D model to the tracking module;
[0010] S4: The tracking module uses the initial 3D model for continuous tracking and generates key frames to be sent back to the differentiable rendering module to further optimize the 3D model through multiple iterations.
[0011] Preferably according to the present invention, in step S2, specifically, the target object in the input frame {I i} is segmented by the large image segmentation model SAM. During this process, the 2D bounding box {b i} of the reference frame is used as a prompt for the large image segmentation model SAM to specify the target object to be segmented. When the target object occupies a small area in the image, the segmented 2D mask is scaled to a fixed size and centered in the image to obtain the final 2D mask {M i}.
[0012] Further preferably according to the present invention, the scaling and centering of the 2D mask are achieved by adjusting the camera intrinsics. Assuming that the model coordinate system of the 3D model G coincides with the world coordinate system and its corresponding pose is T, G is transformed to the camera coordinate system through T to obtain the model G C :
[0013] G C = TG
[0014] The imaging process of the 3D object in the camera is regarded as the process of projecting the 3D model G C onto the 2D image, and the internal parameter matrix of the camera is K:
[0015]
[0016] For a 3D point C on G its 2D projection point x is:
[0017]
[0018] Among them, π(.) represents de - homogenizing three - dimensional coordinates. By adjusting K, the projection mask is scaled and centered without changing the pose parameter T. The width of the two - dimensional mask image M segmented by the large - scale image segmentation model SAM is w, the height is h, and the upper - left corner coordinates of the two - dimensional bounding box are The width is w b , and the height is h b . Calculate the magnification factor s:
[0019]
[0020] where s min is the predefined minimum scaling coefficient, and λ is the adjustment coefficient;
[0021] After obtaining the magnification factor s, calculate the offset Δc to move the mask to the center of the image. To this end, calculate the center c of the two - dimensional mask image and the center c b :
[0022]
[0023] The offset Δc is calculated from s, c, and c b as follows:
[0024]
[0025] Apply s and Δc to f x , f y , c x and c y to perform a transformation to update K:
[0026] f x ← s×f x
[0027] f y ← s×f y
[0028] c x ← s×c x + Δw
[0029] c y ← s×c y + Δh
[0030] Under the new camera internal parameter K, project the three-dimensional model to obtain a new two-dimensional bounding box, or directly transform the original two-dimensional bounding box through s and Δc to obtain a new two-dimensional bounding box. In this case, only need to map the original two-dimensional bounding box area to the new two-dimensional bounding box area to obtain the magnified and centered two-dimensional mask. Before inputting into the differentiable rendering module, pad the two-dimensional mask image with zeros along the x-axis or y-axis direction to form a square without changing the camera internal parameter K.
[0031] According to a preferred embodiment of the present invention, in step S3, given an initial three-dimensional sphere or the three-dimensional mesh model G0 obtained from the previous inverse rendering, the differentiable rendering module uses the differentiable renderer SoftRas for inverse rendering, and the loss function is:
[0032]
[0033] where, represents rendering G0 as a two-dimensional mask under the reference frame pose T i and the camera internal parameter K i to obtain a two-dimensional mask, is the Laplacian matrix, m is the number of vertices of the model G0, X G ∈R m×3 is the matrix composed of all vertices in G0, and the subscript (v, d) represents the value of the element in the v-th row and d-th column of the matrix, is the total loss function, is the term associated with IoU in the loss function, is the term associated with the Laplacian, and μ is the weight of the term associated with the Laplacian.
[0034] According to a preferred embodiment of the present invention, in step S3, the edge of the two-dimensional mask segmented by the large image segmentation model SAM is not smooth, resulting in the size of the inverse-rendered model being slightly smaller than the actual size. The inverse-rendered three-dimensional model is magnified by 1.05 times to match the actual size.
[0035] According to a preferred embodiment of the present invention, in step S4, the tracking module uses the three-dimensional tracking method SLOT to provide a tracking function, receives the three-dimensional model obtained by the differentiable rendering module through inverse rendering, and uses this model to track the target object and provide the three-dimensional pose of the target object. The tracking module generates a certain number of key frames in each iteration for the differentiable rendering module to use. Specifically,
[0036] The tracking module internally establishes a view pool V to store the views corresponding to the key frames:
[0037] V = {v0, v1, v2,..., v n-1}
[0038] where, v i ∈R3 is the i-th perspective, which is calculated from the pose T i where the rotation matrix R representing rotation in T i and the translation vector t representing translation in T i are calculated as follows:
[0039]
[0040] where represents the normalized displacement t i ;
[0041] When a new frame I n arrives, the target object in it is tracked, and its perspective v n is calculated, and then the angular difference between v n and all perspectives in the perspective pool V is calculated:
[0042]
[0043] If all Δa i exceed a pre-given threshold a min , then I n is considered a candidate for a key frame, and then the 3D model is projected onto a 2D plane to obtain its 2D bounding box b n . If the distance between b n and the image boundary exceeds 10 pixels, then I n is considered a key frame, and v i is added to the perspective pool V. If the number of key frames reaches a certain amount, all key frames and their corresponding poses, camera intrinsics, and 2D bounding boxes are transmitted to the differentiable rendering module for inverse rendering.
[0044] The present invention takes an RGB video stream as input, combines a differentiable rendering module with a tracking module, and realizes fast inverse rendering of the 3D model of an object while tracking the object, transforming the model-free tracking problem into a model-based tracking problem. Through inverse rendering of the object contour, the generated 3D model can be reused by other model-based tracking methods without additional processing. This method is applicable under a small number of given prior conditions or without additional prior conditions and is suitable for various scenarios. Compared with the prior art, the present invention does not rely on the training of a specific dataset, avoids high application costs, and improves the model accuracy in an iterative manner to achieve efficient 3D object tracking.
[0045] The beneficial effects of the present invention are as follows:
[0046] 1. The present invention proposes an innovative three-dimensional object tracking method. By combining a differentiable rendering module with a tracking module, it achieves efficient tracking in a model-free scenario. This method can quickly inverse-render the three-dimensional model of an object during the tracking process, transforming the model-free tracking problem into a model-based tracking problem, thereby significantly improving the tracking accuracy and real-time performance. It can work under the condition of only a small amount of prior information, and can even be applicable to a variety of complex scenarios without additional priors. This flexibility makes the present invention have wide applicability in practical applications, especially in environments with limited or lacking prior information.
[0047] 2. The three-dimensional model obtained by inverse rendering through the present invention can be directly reused by other model-based tracking methods without additional processing, and can achieve the effect of an approximate accurate model. This not only improves the generality of the model but also reduces the application cost. The method of the present invention does not require any training, avoiding the high cost and generalization problems of training with specific datasets. This feature makes the method have higher adaptability and practicality in different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic flow chart of the present invention;
[0049] Figure 2 is a schematic flow chart of the differentiable rendering module of the present invention;
[0050] Figure 3 is a schematic flow chart of the scaling and centering process of the mask segmented by the present invention;
[0051] Figure 4 is a schematic flow chart of the tracking module of the present invention;
[0052] Figure 5 is a test result graph of the tracking and inverse rendering of different categories of objects by the present invention in a real scenario; among them, Figure 5 (a) is a graph of the tracking result and the model result obtained by inverse rendering without a reference frame; Figure 5 (b) is a graph of the tracking result and the model result obtained by inverse rendering with 3 given reference frames; Figure 5 (c) is a graph of the tracking result and the model result obtained by inverse rendering with 6 given reference frames. DETAILED DESCRIPTION OF THE INVENTION
[0053] The following further illustrates the present invention through examples in conjunction with the drawings, but is not limited thereto.
[0054] Example 1:
[0055] A model-free three-dimensional object tracking method based on differentiable rendering includes the following steps:
[0056] S1: Specify the target object in the RGB image I0 of the first frame using the two-dimensional bounding box b0, and provide its initial pose T0 and the camera intrinsics as the input to the differentiable rendering module. The camera intrinsics can be calibrated in advance. When there is a reference frame, the reference frame, its corresponding pose, and the two-dimensional bounding box are also used as inputs.
[0057] S2: Segment the target object in the input frame to obtain its two-dimensional mask. Specifically, segment the target object in the input frame {I i} using the large image segmentation model SAM. During this process, the two-dimensional bounding box {b i} of the reference frame is used as a prompt for the large image segmentation model SAM to specify the target object to be segmented. When the target object occupies a small area in the image, scale the segmented two-dimensional mask to a fixed size and center it in the image to obtain the final two-dimensional mask {M i}.
[0058] According to a further preference of the present invention, the scaling and centering of the two-dimensional mask are achieved by adjusting the camera intrinsics. Assume that the model coordinate system of the three-dimensional model G coincides with the world coordinate system, and its corresponding pose is T. Transform G to the camera coordinate system through T to obtain the model G C :
[0059] G C = TG
[0060] Regard the imaging process of the three-dimensional object in the camera as the process of projecting the three-dimensional model G C onto the two-dimensional image. The internal parameter matrix of the camera is K:
[0061]
[0062] For a three-dimensional point C on G its two-dimensional projection point x is:
[0063]
[0064] where π(.) represents dehomogenizing the three-dimensional coordinates. By adjusting K, scale and center the projection mask without changing the pose parameter T. The width of the two-dimensional mask image M segmented by the large image segmentation model SAM is w, and the height is h. The upper left corner coordinates of the two-dimensional bounding box are the width is w b , the height is h b , calculate the scaling factor s:
[0065]
[0066] where smin is the predefined minimum scaling factor, and λ is the adjustment factor;
[0067] After obtaining the magnification factor s, calculate the offset Δc to move the mask to the center of the image. To this end, calculate the center c of the two-dimensional mask image and the center c of the two-dimensional bounding box respectively b :
[0068]
[0069] The offset Δc is calculated from s, c, and c b as follows:
[0070]
[0071] Apply s and Δc to f x , f y , c x and c y for transformation to update K:
[0072] f x ← s × f x
[0073] f y ← s × f y
[0074] c x ← s × c x + Δw
[0075] c y ← s × c y + Δh
[0076] Under the new camera internal parameter K, project the three-dimensional model to obtain a new two-dimensional bounding box, or directly transform the original two-dimensional bounding box through s and Δc to obtain a new two-dimensional bounding box. In this case, only map the original two-dimensional bounding box area to the new two-dimensional bounding box area to obtain the magnified and centered two-dimensional mask. Before inputting into the differentiable rendering module, pad the two-dimensional mask image with zeros along the x-axis or y-axis direction to form a square without changing the camera internal parameter K.
[0077] S3: Start the differentiable rendering module, inverse render an initial three-dimensional model, and transfer the three-dimensional model to the tracking module;
[0078] Given the initial three-dimensional sphere or the three-dimensional mesh model G0 obtained from the previous inverse rendering, the differentiable rendering module uses the differentiable renderer SoftRas for inverse rendering, and the loss function is:
[0079]
[0080] where, Denote the pose of \(G_0\) in the reference frame as \(T\) i and the camera intrinsic parameter matrix as \(K\) i under which \(G_0\) is rendered as a 2D mask is the Laplacian matrix, \(m\) is the number of vertices of the model \(G_0\), \(X\) G \(\in\mathbb{R}\) m×3 is the matrix composed of all vertices in \(G_0\), and the subscript \((v, d)\) represents the value of the element in the \(v\)-th row and \(d\)-th column of the matrix is the total loss function is the term in the loss function related to IoU is the term related to the Laplacian, and \(\mu\) is the weight of the term related to the Laplacian
[0081] The edges of the 2D mask obtained by segmenting the large image segmentation model SAM are not smooth, resulting in the size of the inverse-rendered model being slightly smaller than the actual size. The inverse-rendered 3D model is magnified by 1.05 times to match the actual size.
[0082] S4: The tracking module uses the initial 3D model for continuous tracking and generates key frames to be passed back to the differentiable rendering module, and further optimizes the 3D model through multiple iterations.
[0083] The tracking module uses the 3D tracking method SLOT to provide the tracking function. It receives the 3D model obtained by inverse rendering from the differentiable rendering module and uses this model to track the target object, providing the 3D pose of the target object. The tracking module generates a certain number of key frames in each iteration for the differentiable rendering module to use. Specifically,
[0084] The tracking module internally establishes a view pool \(V\) to store the views corresponding to the key frames:
[0085] \(V=\{v_0, v_1, v_2, \cdots, v\) n-1 \}\)
[0086] where \(v\) i \(\in\mathbb{R}\) 3 is the \(i\)-th view, which is calculated from the rotation matrix \(R\) i representing rotation and the translation vector \(t\) i representing translation in the pose \(T\): i Calculated as:
[0087]
[0088] where represents the normalized displacement \(t\) i ;
[0089] When a new frame \(I\) n arrives, track the target object in it, calculate its view \(v\) n , and then calculate \(v\)n The angular difference from all viewpoints in the viewpoint pool V:
[0090]
[0091] If all Δa i exceed a pre-given threshold a min , then I n is considered a candidate for a key frame, and then the 3D model is projected onto a 2D plane to obtain its 2D bounding box b n . If the distance between b n and the image boundary exceeds 10 pixels, then I n is considered a key frame, and v i is added to the viewpoint pool V. If the number of key frames reaches a certain amount, then all key frames and their corresponding poses, camera intrinsics, and 2D bounding boxes are transmitted to the differentiable rendering module for inverse rendering.
[0092] The MOPED dataset is a dataset captured in a real scene, containing 11 objects. Each object has several RGB-D video sequences with a resolution of 640×480, divided into a reference set and a test set. Due to the limitation of the annotation method accuracy, the annotation accuracy error of some sequences in this dataset is relatively large. In this embodiment, following the Gen6D method, 5 objects with higher annotation accuracy are selected for testing. Each object contains 100 - 600 reference frames and 100 - 300 test frames.
[0093] This embodiment uses the ADD-AUC (ADD - Area Under the Curve, Average Distance - Area Under the Curve) metric to evaluate the tracking accuracy. The distance threshold is from 0 to 10 cm, and the step size is 0.1 cm. This metric measures the tracking accuracy by evaluating the alignment degree of the tracked pose with the ground truth pose within a series of distance thresholds. A higher ADD-AUC value indicates better tracking performance, where the ADD value is calculated as follows:
[0094]
[0095] where m is the number of 3D points on the real 3D model Ggt, R and t represent the rotation matrix and translation vector respectively, and Rgt and tgt represent the real rotation and translation.
[0096] The proposed method was tested under two conditions, which are: (1) R0-3: Only the information of the first frame is given, that is, the initial pose, camera intrinsics, and the 2D bounding box of the target object. The tracking module generates 3 key frames in each iteration; (2) R3-6: Except for the first frame, 3 reference frames are given. The tracking module generates 6 key frames in each iteration. For each model, at most 3 iterations are performed.
[0097] Without a reference frame and using only RGB data without any training, the method R0-3 in this embodiment still outperforms other methods in terms of average precision; when using 3 reference frames, the method R3-6 in this embodiment achieves the best result and has a significant advantage compared to other methods. In addition, compared to other methods, the method in this embodiment shows more stable tracking performance on each object, and there is no significant difference in precision, indicating that the method in this embodiment does not rely on data and demonstrates obvious characteristics of being training-free and good generalization.
[0098] The method in this embodiment is compared with the PoseRBPF method, two versions of the LatentFusion method, the PVNet method, and the Gen6D method. The test results are shown in Table 1, where bold indicates the best result and underlined indicates the second-best result:
[0099] Table 1: Comparison of precision results between the model-free 3D object tracking method based on differentiable rendering proposed in this paper and existing methods in the real-scene MOPED dataset
[0100]
[0101] This embodiment also conducts experiments on the RBOT dataset and tests various advanced tracking methods, including the RBOT method, the SLOT method, the RBGT method, and the SRT3D method. By comparing the performance of these methods on the accurate CAD model and the model obtained by inverse rendering of the method in this chapter, it shows the effectiveness of the method in this embodiment in transforming the model-free tracking problem into a model-based tracking problem.
[0102] Since there are large frame differences in the RBOT dataset, this embodiment uses 3 to 6 reference frames as priors and configures 3 versions, namely, (1) R3-6: using 3 reference frames, and the tracking module generates 6 key frames; (2) R6-12: using 6 reference frames, and the tracking module generates 12 key frames; (3) R12-12: using 12 reference frames, and the tracking module generates 12 key frames.
[0103] The method in this embodiment tracks the first 100 frames of the regular sequence in the RBOT dataset. For each object, 3 iterations are performed, and then the inverse-rendered 3D model is obtained. These 3D models are used as inputs for other tracking methods to test the performance of these methods in all scenarios of the RBOT dataset. In the test, the evaluation metric of 5cm / 5° is used to evaluate the tracking precision, and the normalized Hausdorff distance is used to evaluate the precision of the inverse-rendered model.
[0104] Table 2 shows the tracking accuracies of the four methods participating in the test under different models. The bold font indicates the best results except for the exact model, and the "Effect" column represents the relative tracking accuracy compared with the exact model. When given 6 frames (R6-12) and 12 frames (R12-12) of reference frames, the model obtained by inverse rendering using the method of this embodiment achieved a tracking effect very close to that of the exact model. It reached 75.2% - 83.9% of the effect of the exact model under 6 frames of reference frames and 81.2% - 88.2% of the effect of the exact model under 12 frames of reference frames, indicating the effectiveness of the method of this embodiment in transforming the model-free tracking problem into a model-based tracking problem. Even under the condition of only given 3 frames of reference frames, it can still reach 41.3% - 53.2% of the effect of the exact model.
[0105] Table 2: Tracking results of each method on the exact model and the model obtained by inverse rendering on the RBOT dataset
[0106]
[0107]
[0108] As can be seen from Table 1 and Table 2, the method of the present invention is a 3D object tracking method that does not require training. Under the condition of weak prior in the model-free scenario, through the proposed differentiable rendering module and tracking module, it can quickly inverse render the 3D model of the object while tracking the object. The 3D model obtained by inverse rendering can be reused by other tracking methods without additional processing. The method of this embodiment does not need to be trained on specific data and has achieved advanced results in various test scenarios. The above characteristics make this method have good generalization and usability.
[0109] Figure 5 It shows that in real scenarios with different prior conditions, the method of this embodiment has achieved stable tracking of objects and can inverse render the models of objects, thus proving the versatility of this method.
[0110] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A model-free three-dimensional object tracking method based on differentiable rendering, characterized in that, The steps are as follows: S1: Specify the target object in the RGB image of the first frame using a two-dimensional bounding box, and provide its initial pose and camera intrinsics as the input to the differentiable rendering module. When there is a reference frame, the reference frame and its corresponding pose and two-dimensional bounding box are also used as inputs; S2: Segment the target object in the input frame to obtain its two-dimensional mask; S3: Start the differentiable rendering module to inverse-render an initial three-dimensional model, and transfer the three-dimensional model to the tracking module; S4: The tracking module uses the initial three-dimensional model for continuous tracking, generates key frames and sends them back to the differentiable rendering module, and further optimizes the three-dimensional model through multiple iterations.
2. The model-free three-dimensional object tracking method based on differentiable rendering according to claim 1, wherein In step S2, specifically, the target object in the input frame {I i} is segmented by the large image segmentation model SAM. During this process, the two-dimensional bounding box {b i} of the reference frame is used as a prompt for the large image segmentation model SAM to specify the target object to be segmented. The resulting two-dimensional mask is scaled to a fixed size and centered in the image to obtain the final two-dimensional mask {M i}.
3. The model-free three-dimensional object tracking method based on differentiable rendering according to claim 2, wherein The scaling and centering of the two-dimensional mask are achieved by adjusting the camera internal parameters. Assuming that the model coordinate system of the three-dimensional model G coincides with the world coordinate system, and its corresponding pose is T, G is transformed into the camera coordinate system through T, and the model G is obtained C : G c = TG The imaging process of a three-dimensional object in a camera is regarded as a three-dimensional model G C projected onto a two-dimensional image. The internal parameter matrix of the camera is K: For G C a three-dimensional point X = [X, Y, Z] on T , its two-dimensional projection point x is: Among them, π(.) represents dehomogenizing three-dimensional coordinates. By adjusting K, the projection mask is scaled and centered without changing the pose parameter T. The width of the two-dimensional mask image M segmented by the large image segmentation model SAM is w, and the height is h. The upper left corner coordinates of the two-dimensional bounding box are (x b , y b ). T , the width is w b , and the height is h b , calculate the magnification factor s: where s min is a predefined minimum scaling factor, and λ is a regulation factor; After obtaining the magnification factor s, calculate the offset Δc to move the mask to the center of the image. To this end, calculate the center c of the two-dimensional mask image and the center c of the two-dimensional bounding box respectively b : The offset Δc is calculated from s, c, and c b as follows: Apply s and Δc to f x 、f y 、c x and c y for transformation to update K: f x ←s × f x f y ←s×f y c x ← s × c x + Δw c y ← s × c y + Δh Under the new camera intrinsics K, project the three-dimensional model to obtain a new two-dimensional bounding box, or directly transform the original two-dimensional bounding box through s and Δc to obtain a new two-dimensional bounding box. In this case, only map the original two-dimensional bounding box area to the new two-dimensional bounding box area to obtain the enlarged and centered two-dimensional mask. Before inputting it into the differentiable rendering module, pad the two-dimensional mask image with zeros along the x-axis or y-axis direction to form a square without changing the camera intrinsics K.
4. The model-free three-dimensional object tracking method based on differentiable rendering according to claim 3, wherein In step S3, given an initial three-dimensional sphere or the three-dimensional mesh model G0 obtained from the previous inverse rendering, the differentiable rendering module uses the differentiable renderer SoftRas for inverse rendering, and the loss function is: Among them, means rendering G0 as a two-dimensional mask under the reference frame pose T i and the camera internal parameter K i and rendering it as a two-dimensional mask, is the Laplacian matrix, m is the number of vertices of the model G0, X G ∈R m×3 is the matrix composed of all vertices in G0, and the subscript (v, d) represents the value of the element in the v-th row and d-th column of the matrix, is the total loss function, is the term in the loss function associated with IoU, is the term associated with the Laplacian, and μ is the weight of the term associated with the Laplacian.
5. The model-free three-dimensional object tracking method based on differentiable rendering according to claim 4, wherein In step S3, the inverse-rendered three-dimensional model is enlarged by 1.05 times.
6. The model-free three-dimensional object tracking method based on differentiable rendering according to claim 5, characterized in that, In step S4, the tracking module uses the three-dimensional tracking method SLOT to provide tracking functionality, receives the three-dimensional model obtained by inverse rendering from the differentiable rendering module, and uses this model to track the target object, providing the three-dimensional pose of the target object. The tracking module generates a certain number of key frames in each iteration for the differentiable rendering module to use. Specifically, The tracking module internally establishes a view pool V to store the views corresponding to the key frames: V = {v0, v1, v2,..., v n-1} where, v i ∈R 3 is the i-th view, which is calculated from the rotation matrix R i representing rotation and the translation vector t i representing translation in the pose T i as follows: Among them, represents the normalized displacement t i ; When the new frame I n arrives, track the target object in it and calculate its viewing angle v n , and then calculate the n angular differences between v and all the viewing angles in the viewing angle pool V: If all Δa i exceed a pre-given threshold a min , then I n is considered a candidate for a key frame. Then, the 3D model is projected onto a 2D plane to obtain its 2D bounding box b n . If the distance between b n and the image boundary exceeds 10 pixels, then I n is considered a key frame, and v i is added to the view pool V. If the number of key frames reaches a certain amount, then all key frames and their corresponding poses, camera intrinsics, and 2D bounding boxes are transmitted to the differentiable rendering module for inverse rendering.