Unknown object 6D pose estimation method and system based on single reference view
Through a deep learning method based on a single reference view, the spatial and image features of the object are extracted and point cloud iteratively aligned, the time and performance consumption problems of estimating 6D poses of unknown objects in the prior art are solved, and efficient and accurate poses of estimation are achieved.
Patent Information
- Application Number
- CN202510076822.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-01
AI Technical Summary
The existing 6D pose estimation method of unknown objects relies on CAD models or intensive reference views, which is time-consuming and labor-intensive, and is not very scalable, making it difficult to meet real-time requirements.
Using a single reference view-based method, the spatial features and image features of the object are extracted through the deep learning model, and iteratively aligned point clouds to solve the position pose, realizing accurate 6D position pose estimation of unknown objects.
No CAD models or intensive reference views are required, which improves scalability and real-time performance, and enables accurate pose estimation of unknown objects.
Smart Images

Figure CN120235933A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot environmental perception, and particularly relates to a method and system for estimating the 6D pose of an unknown object based on a single reference view. Background Art
[0002] With the rapid development of the field of computer vision and the continuous progress of three-dimensional scene understanding technology, it is promoting the transformation from two-dimensional to three-dimensional vision, enabling machines to have higher cognitive and decision-making abilities in the real environment. The three-dimensional space scene understanding technology has gradually become a research hotspot. These technologies aim to enable machines to perceive and understand the complex three-dimensional scene layout, the spatial relationships of objects, and the pose information of each object. Object 6D pose estimation is a key task among them, which determines the position (3D translation) and orientation (3D rotation) of an object in space through RGB or RGB-D image data, and plays a crucial role in fields such as robot control, augmented reality (AR), and virtual reality (VR).
[0003] In early research, traditional object 6D pose estimation methods relied on manually annotated feature correspondences and solved the pose through feature matching. Such methods were overly dependent on manual annotation. With the development of deep learning and neural networks, many subsequent works applied them to the object 6D pose estimation task. Compared with traditional methods, the object 6D pose estimation method based on deep learning has shown advantages in terms of accuracy, generalization, and robustness, and has become the mainstream of current research.
[0004] In the object 6D pose estimation method based on deep learning, previous research was divided into instance-level and category-level object 6D pose estimation methods according to whether the algorithm uses a specific object or objects of the same category for training. However, these methods are only applicable to specified objects or categories and cannot be generalized to new objects or categories that were not seen during training. To solve this problem, researchers proposed the concept of unknown object 6D pose estimation. However, current unknown object 6D pose estimation methods usually rely on CAD models or dense reference views. Accurate texture and high-precision CAD models require professional equipment for acquisition or professional knowledge for design. Such methods are time-consuming and laborious and have poor scalability. In contrast, the acquisition of reference views is relatively easy, so its scalability is also stronger. However, methods based on reference views usually solve the pose by first reconstructing a 3D model or directly obtain a rough pose through template matching. These methods all rely on dense reference views and consume a huge amount of time and performance, and it is difficult to meet the real-time requirements in practical applications. The device performance also limits the scalability of the algorithm.
[0005] In summary, there are still many challenges in the generalization and scalability of current pose estimation algorithms. Summary of the Invention
[0006] An object of an embodiment of the present invention is to provide a method and system for 6D pose estimation of an unknown object based on a single reference view, which does not rely on the CAD model of the object or dense reference views, but only uses a single reference view, extracts the spatial features and image features of the object through a deep learning model, and calculates the pose after iterative alignment of the point cloud, so as to achieve accurate 6D pose estimation of the unknown object, thereby solving at least one technical problem involved in the background art.
[0007] In order to solve the above technical problems, the present invention is implemented as follows:
[0008] An embodiment of the present invention provides a method for 6D pose estimation of an unknown object based on a single reference view, including the following steps:
[0009] Step S1, use an RGB-D camera to obtain a reference view and a query view of the object, and obtain an RGB image and a point cloud through segmentation and back-projection;
[0010] Step S2, use the object pose of the reference view and the rough object pose of the query view to focus the point cloud on the same coordinate system;
[0011] Step S3, input the focused point cloud into a point cloud state space model to extract spatial features, input the RGB image into a visual state space model to extract image features, then fuse the extracted image features and spatial features of the reference view to obtain reference features, and fuse the image features and spatial features of the query view to obtain query features;
[0012] Step S4, input the reference features and the query features into a GeoTransformer model to obtain the point-by-point features of the reference view and the query view respectively, then establish a correspondence matrix of the point-by-point features of the reference view and the query view and solve the pose of the object through weighted singular value decomposition;
[0013] Step S5, use the solved pose as the rough object pose of the query view, and repeat steps S2 to S4 for iterative alignment of the point cloud to obtain the 6D pose estimation of the object in the query view.
[0014] Optionally, step S1 specifically includes:
[0015] Use an RGB-D camera to obtain a reference view and a query view of the object, use a CNOS or MASK R-CNN segmentation model to segment the reference view and the query view to obtain an RGB image, and back-project the RGB image to obtain a point cloud.
[0016] Optionally, the RGB image includes a reference view RGB image and a query view RGB image; the point cloud includes a reference point cloud and a query point cloud.
[0017] Optionally, step S2 specifically includes:
[0018] Performing coordinate transformation on the original reference point cloud P using the object pose of the reference view r to obtain the reference point cloud in the camera origin coordinate system which is represented by the following formula:
[0019]
[0020] In the formula, R r and t r respectively represent the translation and rotation of the object pose of the reference view; T represents matrix transpose; c represents the camera origin coordinate system; r represents the reference view;
[0021] Performing coordinate transformation on the original query point cloud P using the rough object pose of the query view q to obtain the query point cloud in the camera origin coordinate system, which is represented by the following formula:
[0022]
[0023] In the formula, t i and R i respectively represent the translation and rotation of the object solved in the i-th iteration; represents the query point cloud in the camera origin coordinate system after the i-th iteration.
[0024] Optionally, step S3 specifically includes:
[0025] Step S31: The point cloud state space model uses two scanning methods, the three-dimensional Hilbert curve and the transposed variant, to sequentially traverse all points of the point cloud in three-dimensional space and extract the spatial features of each point;
[0026] Step S32: The visual state space model uses a cross-scanning module to cross-scan the RGB image in four different directions. Each pixel point of the RGB image will integrate information from all other pixels in different directions, retain the global feature information while maintaining linear complexity, and extract the image features;
[0027] Step S33: Fusing the extracted image features and spatial features of the reference view to obtain the reference feature F r , and fusing the image features and spatial features of the query view to obtain the query feature which is represented by the following formula:
[0028]
[0029] In the formula, represents matrix addition; and respectively represent the reference features and query features extracted in the i-th iteration; N r represents the number of reference features; N q represents the number of query features; C represents the dimension of the feature channel; PointMamba() represents the point cloud state space model; VMamba() represents the visual state space model; I r represents the reference image; I q represents the query image.
[0030] Optionally, step S4 specifically includes:
[0031] Step S41, input the reference feature F r and the query feature into the GeoTransformer model to obtain the point-wise features of the reference view and the query view and Then establish the correspondence matrix A of the point-wise features of the reference view and the query view i , where represents matrix multiplication; select the point pairs with the highest similarity for alignment, and use the cross-entropy loss function to supervise the alignment relationship, which is expressed by the following formula:
[0032]
[0033] In the formula, Loss represents the overall loss function; CE(·, ·) represents the cross-entropy loss function; and respectively represent the ground truth of the query point cloud and the reference point cloud ;
[0034] Step S42: Perform weighted singular value decomposition on the point-wise correspondence matrix to calculate the object pose pose i+1 , which is expressed by the following formula:
[0035]
[0036] where pose i+1 represents the object pose solved in the i-th iteration; R i+1 represents rotation; t i+1 represents translation; WSVD() represents weighted singular value decomposition; A i represents the correspondence matrix of the point-wise features of the reference view and the query view.
[0037] The present invention also provides a single-reference-view-based unknown object 6D pose estimation system for implementing the above method, including:
[0038] The RGB image and point cloud acquisition module is used to obtain the reference view and query view of an object using an RGB-D camera, and obtain the RGB image and point cloud through segmentation and back-projection;
[0039] The point cloud focusing module is used to focus the point cloud into the same coordinate system using the object pose of the reference view and the rough object pose of the query view;
[0040] The reference feature and query feature acquisition module is used to input the focused point cloud into the point cloud state space model to extract spatial features, input the RGB image into the visual state space model to extract image features, and then fuse the extracted image features and spatial features of the reference view to obtain reference features, and fuse the image features and spatial features of the query view to obtain query features;
[0041] The pose calculation module is used to input the reference feature and query feature into the GeoTransformer model to obtain the point-by-point features of the reference view and query view respectively, and then establish a correspondence matrix of the point-by-point features of the reference view and query view and solve the pose of the object through weighted singular value decomposition;
[0042] The 6D pose acquisition module is used to use the calculated pose as the rough object pose of the query view, perform point cloud iterative alignment, and obtain the 6D pose of the object in the query view.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] 1. The present invention proposes a pose estimation framework without a CAD model, which can realize the 6D pose estimation of an unknown object only relying on a single reference view, eliminating the dependence on textured CAD models and dense reference views, thereby further improving scalability.
[0045] 2. The present invention proposes to learn a generalizable point cloud alignment model to perform point-by-point alignment of the reference view and query view in the camera coordinate system to realize object pose estimation, and further improve the accuracy and robustness through the iterative alignment process.
[0046] 3. The present invention explores the potential of the state space model. Since the establishment of the point-by-point alignment relationship depends on rich spatial information, a point cloud and visual state space model are designed to extract the point-by-point features of the reference and query views, which can reduce the computational complexity and the number of model parameters while improving the accuracy of the alignment relationship compared with the Transformer architecture. Description of the Drawings
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings, where:
[0048] Figure 1 It is a flow framework diagram of a method for estimating the 6D pose of an unknown object based on a single reference view provided by the present invention;
[0049] Figure 2 It is a network structure diagram of the visual state space model used in the present invention;
[0050] Figure 3 It is a framework diagram of a system for estimating the 6D pose of an unknown object based on a single reference view provided by the present invention. Specific embodiments
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0052] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0053] Please refer to Figure 1 As shown, the embodiments of the present invention provide a method for estimating the 6D pose of an unknown object based on a single reference view, including the following steps:
[0054] Step S1, use an RGB-D camera to obtain the reference view and query view of the object, and obtain the RGB image and point cloud through segmentation and back-projection;
[0055] Step S2, use the object pose of the reference view and the rough object pose of the query view to focus the point cloud on the same coordinate system;
[0056] Step S3: Input the focused point cloud into the point cloud state space model to extract spatial features, and input the RGB image into the visual state space model to extract image features. Then, fuse the extracted image features and spatial features of the reference view to obtain reference features, and fuse the image features and spatial features of the query view to obtain query features;
[0057] It should be noted that the point cloud state space model is a new framework for processing and analyzing three-dimensional point cloud data. It has recently shown linear complexity and long-range context learning ability in sequential data, sorts irregular point cloud data, and retains the spatial proximity of the original point cloud. It is used in the present invention to extract the spatial features of point cloud data.
[0058] The visual state space model is a model for processing and analyzing visual data, especially image sequences. It models by mapping the input data to a latent state space, and can capture the dynamic changes and time series information of the data. It is used in the present invention to extract the image features of RGB. For the specific network design, reference can be made to Figure 2 as shown.
[0059] Step S4: Input the reference features and query features into the GeoTransformer model to obtain the point-by-point features of the reference view and the query view respectively. Then, establish a correspondence matrix of the point-by-point features of the reference view and the query view and solve the pose of the object through weighted singular value decomposition;
[0060] It should be noted that the GeoTransformer model is a deep learning model based on the Transformer architecture, which is specifically used to process spatial data or inputs with spatial features. This model is mainly applied to tasks that require point-by-point alignment, feature extraction, and spatial relationship modeling of objects or point cloud data. The GeoTransformer model combines the powerful self-attention mechanism of the Transformer and the special requirements of spatial data, and can effectively capture global and local information when processing spatial data, and then perform high-precision matching, registration, and pose estimation. In the present invention, it is used for point-by-point alignment of the reference view and the query view to generate a correspondence matrix.
[0061] Step S5: Use the solved pose as the rough pose of the object in the query view, and repeat Steps S2 to S4 for iterative alignment of the point cloud to obtain the 6D pose estimation of the object in the query view.
[0062] Step S1 specifically includes:
[0063] Use an RGB-D camera to obtain the reference view and query view of the object. Adopt the CNOS or MASK R-CNN segmentation model to segment the reference view and query view to obtain RGB images, and back-project the RGB images to obtain point clouds. In this embodiment, the number of points in the point cloud is 2048 points.
[0064] Specifically, the RGB images include the reference view RGB image and the query view RGB image; the point clouds include the reference point cloud and the query point cloud.
[0065] Step S2 specifically includes:
[0066] Use the object pose of the reference view to perform coordinate transformation on the original reference point cloud P r to obtain the reference point cloud in the camera origin coordinate system which is expressed by the following formula:
[0067]
[0068] In the formula, R r and t r respectively represent the translation and rotation of the object pose of the reference view; T represents matrix transpose; c represents the camera origin coordinate system; r represents the reference view;
[0069] Use the rough object pose of the query view to perform coordinate transformation on the original query point cloud P q to obtain the query point cloud in the camera origin coordinate system, which is expressed by the following formula:
[0070]
[0071] In the formula, t i and R i respectively represent the translation and rotation of the object solved in the i-th iteration; represents the query point cloud in the camera origin coordinate system after the i-th iteration.
[0072] It should be noted that since the object pose in the query view is initially unknown, no rotation transformation is performed during the first point cloud focusing, but the translation t1 is set to the average coordinates of the object. In subsequent iterations, the object pose Pose i+1 solved in the previous round is used for coordinate transformation.
[0073] Step S3 specifically includes:
[0074] Step S31, the point cloud state space model uses two scanning methods, the three-dimensional Hilbert curve and the transposed variant, to sequentially traverse all points of the point cloud in three-dimensional space and extract the spatial features of each point;
[0075] Step S32: The visual state space model uses a cross-scanning module to cross-scan the RGB image in four different directions. Each pixel point of the RGB image will integrate information from all other pixels in different directions, retaining global feature information while maintaining linear complexity and extracting image features;
[0076] Step S33: Fuse the image features and spatial features of the extracted reference view to obtain the reference feature F r , and fuse the image features and spatial features of the query view to obtain the query feature It is expressed by the following formula:
[0077]
[0078] In the formula, represents matrix addition; and respectively represent the reference feature and the query feature extracted in the i-th iteration; N r represents the number of reference features; N q represents the number of query features; C represents the dimension of the feature channel; PointMamba() represents the point cloud state space model; VMamba() represents the visual state space model; I r represents the reference image; I q represents the query image.
[0079] Step S4 specifically includes:
[0080] Step S41, input the reference feature F r and the query feature into the GeoTransformer model to obtain the point-by-point features and of the reference view and the query view respectively. Then establish the correspondence matrix A i of the point-by-point features of the reference view and the query view, where represents matrix multiplication; select the point pair with the highest similarity for alignment, and use the cross-entropy loss function to supervise the alignment relationship, which is expressed by the following formula:
[0081]
[0082] In the formula, Loss represents the overall loss function; CE(·,·) represents the cross-entropy loss function; and respectively represent the true values of the query point cloud and the reference point cloud ;
[0083] It should be noted that, each element p inq , corresponding to the points in the query point cloud The reference point cloud can be found by using the given ground truth rotation R gt and translation t gt to obtain the reference point cloud the index of the point closest to the query point cloud in is obtained. If the distance between the closest point pairs exceeds a specified threshold (set to 15 cm), the point pairs are discarded. Additionally, the elements in are also obtained in the same way.
[0084] It should be noted that by training and learning two GeoTransformer models with non-shared weights, one for establishing the initial point-by-point alignment and the other for the subsequent iterative process, the point-by-point correspondence matrix is output;
[0085] Step S42: Perform weighted singular value decomposition on the point-by-point correspondence matrix to calculate the object pose pose i+1 , which is expressed as:
[0086]
[0087] where pose i+1 represents the object pose solved in the i-th iteration; R i+1 represents rotation; t i+1 represents translation; WSVD() represents weighted singular value decomposition; A i represents the correspondence matrix of the point-by-point features of the reference view and the query view.
[0088] Regarding step S5, it should be further noted that considering the significant pose differences that may exist between the initial query point cloud and the reference point cloud, especially in terms of rotation, this may lead to unaligned point pairs during the alignment process, thereby affecting the accuracy of pose estimation. By feeding back the estimated 6D pose of the object obtained in step S404 to step S202 and iteratively executing steps S2 to S4, the pose transformation of the query point cloud is refined. This process promotes the gradual convergence between the query point cloud and the reference point cloud, ultimately achieving a more accurate 6D pose estimation of the unknown object.
[0089] The present invention also provides a single-reference-view-based unknown object 6D pose estimation system for implementing the described method, including an RGB image and point cloud acquisition module, a point cloud focusing module, a reference feature and query feature acquisition module, a pose solution module, and a 6D pose acquisition module.
[0090] The RGB image and point cloud acquisition module is used to acquire the reference view and query view of the object using an RGB-D camera, and obtain the RGB image and point cloud through segmentation and back-projection.
[0091] The point cloud focusing module is used to focus the point cloud into the same coordinate system by using the object pose of the reference view and the rough object pose of the query view.
[0092] The reference feature and query feature acquisition module is used to input the focused point cloud into the point cloud state space model to extract spatial features, input the RGB image into the visual state space model to extract image features, then fuse the extracted image features and spatial features of the reference view to obtain reference features, and fuse the image features and spatial features of the query view to obtain query features.
[0093] The pose calculation module is used to input the reference features and query features into the GeoTransformer model to obtain the point-by-point features of the reference view and the query view respectively, then establish a corresponding relationship matrix of the point-by-point features of the reference view and the query view and solve the pose of the object through weighted singular value decomposition.
[0094] The 6D pose acquisition module is used to use the calculated pose as the rough object pose of the query view, perform point cloud iterative alignment, and obtain the 6D pose of the object in the query view.
[0095] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0096] In addition, it should be pointed out that the scope of the methods and systems in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0097] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims, and all of them belong to the protection scope of the present invention.
Claims
1. A method for estimating 6D pose of an unknown object based on a single reference view, characterized in that: The steps include: Step S1, using an RGB-D camera to obtain a reference view and a query view of the object, and obtaining an RGB image and a point cloud through segmentation and back-projection; Step S2, focusing the point cloud into the same coordinate system using the object pose of the reference view and the rough object pose of the query view; Step S3, inputting the focused point cloud into the point cloud state space model to extract spatial features, inputting the RGB image into the visual state space model to extract image features, then fusing the extracted image features and spatial features of the reference view to obtain reference features, and fusing the image features and spatial features of the query view to obtain query features; Step S4, inputting the reference features and the query features into the GeoTransformer model, obtaining the point-by-point features of the reference view and the query view respectively, then establishing a correspondence matrix between the point-by-point features of the reference view and the query view and calculating the pose of the object through weighted singular value decomposition; Step S5, using the solved pose as the rough pose of the object in the query view, repeating steps S2 to S4 to perform iterative point cloud alignment to obtain a 6D pose estimate of the object in the query view.
2. The method according to claim 1, characterized in that: Step S1 specifically includes: Use an RGB-D camera to obtain the reference view and query view of the object, use the CNOS or MASK R-CNN segmentation model to segment the reference view and query view to obtain RGB images, and back-project the RGB images to obtain point clouds.
3. The method according to claim 1 or 2, characterized in that: The RGB image includes a reference view RGB image and a query view RGB image; the point cloud includes a reference point cloud and a query point cloud.
4. The method according to claim 3, characterized in that Step S2 specifically includes: The original reference point cloud P is rectified using the object pose of the reference view. r Perform coordinate transformation to obtain the reference point cloud in the camera origin coordinate system It is expressed by the following formula: In the formula, R r and t r Respectively represent the translation and rotation of the object pose of the reference view; represents matrix transpose; c represents the camera origin coordinate system; r represents the reference view; The original query point cloud P is rectified using the rough pose of the object in the query view. q Perform coordinate transformation to obtain the query point cloud in the camera origin coordinate system, which is expressed as follows: Where, t i and R i denote the translation and rotation of the object solved in the i-th iteration respectively; represents the query point cloud in the camera origin coordinate system after the i-th iteration.
5. The method according to claim 4, characterized in that Step S3 specifically includes: Step S31, the point cloud state space model uses two scanning methods, the three-dimensional Hilbert curve and the transposed variant, to sequentially traverse all points of the point cloud in the three-dimensional space to extract point-by-point spatial features; Step S32: the visual state space model uses a cross-scan module to cross-scan the RGB image along four different directions. Each pixel of the RGB image integrates information from all other pixels in different directions, retains global feature information and extracts image features while maintaining linear complexity. Step S33: Fusing the extracted image features and spatial features of the reference view to obtain the reference feature F r , the image features and spatial features of the query view are fused to obtain the query features It is expressed by the following formula: In the formula, represents matrix addition; and Respectively represent the reference features and query features extracted in the i-th iteration; N r Represents the number of reference features; N q represents the number of query features; C represents the dimension of the feature channel; PointMamba() represents the point cloud state space model; VMamba() represents the visual state space model; I r I represents the reference image; q represents the query image.
6. The method according to claim 5, characterized in that Step S4 specifically includes: Step S41: reference feature F r and query features Input to the GeoTransformer model to obtain the point-by-point features of the reference view and the query view respectively and Then, the correspondence matrix A between the point-by-point features of the reference view and the query view is established i ,in, represents matrix multiplication; select the point pair with the highest similarity for alignment, and use the cross entropy loss function to supervise the alignment relationship, which is expressed as follows: Where Loss represents the overall loss function; CE(·,·) represents the cross entropy loss function; and Represent the query point cloud and reference point cloud The true value of Step S42: Perform weighted singular value decomposition on the point-by-point correspondence matrix to calculate the object pose i+1 , expressed by the following formula: Among them, pose i+1 represents the object pose solved in the i-th iteration; R i+1 Indicates rotation; t i+1 represents translation; WSVD() represents weighted singular value decomposition; A i A correspondence matrix representing the point-wise features of the reference view and the query view.
7. A system for estimating 6D pose of an unknown object based on a single reference view for implementing the method according to any one of claims 1 to 6, characterized in that: include: RGB image and point cloud acquisition module, which is used to obtain the reference view and query view of the object using the RGB-D camera, and obtain the RGB image and point cloud through segmentation and back projection; A point cloud focusing module, which is used to focus the point cloud to the same coordinate system using the object pose of the reference view and the rough pose of the object in the query view; A reference feature and query feature acquisition module, which is used to input the focused point cloud into the point cloud state space model to extract spatial features, input the RGB image into the visual state space model to extract image features, and then fuse the extracted image features and spatial features of the reference view to obtain the reference features, and fuse the image features and spatial features of the query view to obtain the query features; The pose solving module is used to input the reference features and query features into the GeoTransformer model, obtain the point-by-point features of the reference view and the query view respectively, then establish the correspondence matrix of the point-by-point features of the reference view and the query view and solve the pose of the object through weighted singular value decomposition; The 6D pose acquisition module is used to use the solved pose as the rough pose of the object in the query view, perform iterative point cloud alignment, and obtain the 6D pose of the object in the query view.