Object six-dimensional posture recognition method, system and equipment and medium

By performing semantic segmentation and feature fusion on RGB-D images, combining attention mechanism and multi-layer perceptron, the problem of low accuracy in six-dimensional pose recognition of objects in the prior art is solved, achieving higher recognition accuracy and robustness to low-textured objects.

CN120107362APending Publication Date: 2025-06-06NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510187135.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing six-dimensional pose recognition methods for objects have low accuracy when dealing with low-texture or non-textured objects, and cannot effectively utilize the fusion strategy between color and geometric information.

Method used

By semantic segmentation of RGB-D images, the color features and geometric features in the target color image and the target depth image are extracted, and the attention mechanism is used to gradually fusion of multiple features to obtain the fusion features. The six-dimensional pose of the object is then determined using a pre-trained multi-layer perceptron.

Benefits of technology

Improve the recognition accuracy of the six-dimensional pose of objects, especially when dealing with low-texture or non-textured objects, the accuracy is improved by more than 5%, and the fusion strategy between color and geometric information is effectively utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107362A_ABST
    Figure CN120107362A_ABST
Patent Text Reader

Abstract

The invention discloses an object six-dimensional posture recognition method, system and device and a medium, and relates to the field of posture recognition, and the method comprises the steps: obtaining an RGB-D image of an object; performing semantic segmentation on the RGB-D image to obtain a target color image and a target depth image; extracting color features from the target color image, and extracting geometric features from the target depth image; carrying out multi-feature progressive fusion on the color features and the geometric features by adopting an attention mechanism to obtain fusion features; and determining the six-dimensional attitude of the object by adopting a pre-trained first multi-layer perceptron according to the fusion features. The method fully explores the relationship between the color and the geometric information, can extract more representative features, and improves the recognition precision of the six-dimensional attitude of the object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of posture recognition, and in particular to a method, system, device and medium for six-dimensional posture recognition of an object based on multi-feature fusion. Background Art

[0002] The six-dimensional pose recognition of objects plays an important role in the fields of virtual reality, robot grasping and autonomous navigation. However, due to the different three-dimensional shapes of objects, noisy captured images, changing lighting conditions and occlusion between objects, the six-dimensional pose recognition of objects is full of challenges. To solve this problem, many pose recognition methods have emerged:

[0003] 1. Pose recognition based on RGB images: By extracting and matching local features or predicting the 2D projection of predefined 3D key points, a 2D-3D correspondence between 2D key points and 3D models is established. Based on these correspondences, the 6D pose of the object is estimated by solving the perspective-n-point (PnP) problem. Although these algorithms are effective and fast for objects with rich textures, they have difficulty dealing with low-texture or no-texture objects.

[0004] 2. Posture recognition based on RGB-D images: Extract three-dimensional features from color and depth image pairs, and then perform a corresponding matching process to predict the six-dimensional pose of the object. For example, the Ipose method uses an encoder-decoder architecture to extract features from color images, and then obtains the two-dimensional-three-dimensional correspondence between the color image and the 3D model. The PnP problem is solved by using the obtained correspondence and depth information to estimate the six-dimensional pose of the object. However, these methods cannot effectively utilize the fusion strategy between color and geometric information.

[0005] In summary, the accuracy of related posture recognition methods is low. Summary of the invention

[0006] The purpose of this application is to provide a method, system, device and medium for six-dimensional posture recognition of an object, which can improve the recognition accuracy of the six-dimensional posture of an object.

[0007] To achieve the above objectives, this application provides the following solutions:

[0008] In a first aspect, the present application provides a method for six-dimensional posture recognition of an object, comprising:

[0009] Get the RGB-D image of the object;

[0010] Performing semantic segmentation on the RGB-D image to obtain a target color image and a target depth image;

[0011] Extracting color features from the target color image and extracting geometric features from the target depth image;

[0012] Using an attention mechanism to perform multi-feature progressive fusion on the color feature and the geometric feature to obtain a fusion feature;

[0013] According to the fusion features, a pre-trained first multi-layer perceptron is used to determine the six-dimensional posture of the object.

[0014] In a second aspect, the present application provides a six-dimensional posture recognition system for an object, comprising:

[0015] An image acquisition module, used to acquire an RGB-D image of an object;

[0016] A semantic segmentation module, used to perform semantic segmentation on the RGB-D image to obtain a target color image and a target depth image;

[0017] A feature extraction module, used to extract color features from the target color image and extract geometric features from the target depth image;

[0018] A feature fusion module, used for performing multi-feature progressive fusion of the color feature and the geometric feature by using an attention mechanism to obtain a fusion feature;

[0019] The posture determination module is used to determine the six-dimensional posture of the object according to the fusion features by using a pre-trained first multi-layer perceptron.

[0020] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned six-dimensional posture recognition method of an object.

[0021] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for six-dimensional posture recognition of an object.

[0022] According to the specific embodiments provided in this application, this application has the following technical effects:

[0023] The present application provides a method, system, device and medium for six-dimensional posture recognition of an object. By performing semantic segmentation on an RGB-D image, a target color image and a target depth image are obtained; color features are extracted from the target color image, and geometric features are extracted from the target depth image. The color features and geometric features are integrated to make up for the misestimation caused by the lack of geometric information. An attention mechanism is used to perform multi-feature progressive fusion of color features and geometric features, and the relationship between color and geometric information is fully explored. More representative features can be extracted. Finally, a multi-layer perceptron is used to determine the six-dimensional posture of the object, thereby improving the recognition accuracy of the six-dimensional posture of the object. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0025] Figure 1 This is a diagram of an application environment of a six-dimensional posture recognition method for an object in an embodiment of the present application;

[0026] Figure 2 A schematic diagram of a flow chart of a method for six-dimensional posture recognition of an object provided in one embodiment of the present application;

[0027] Figure 3 This is a schematic diagram of a six-dimensional posture recognition process of an object in an embodiment of the present application;

[0028] Figure 4 A schematic diagram of a convolutional layer and a maximum pooling layer in a semantic segmentation framework in an embodiment of the present application;

[0029] Figure 5 This is a schematic diagram of the object posture recognition result of a banana in this application;

[0030] Figure 6 This is a schematic diagram of the object posture recognition result of a biscuit box in this application;

[0031] Figure 7 This is a schematic diagram of the object posture recognition result of the thermos cup in this application;

[0032] Figure 8 This is a schematic diagram of the object posture recognition result of a milk carton in this application;

[0033] Fig. 9 This application presents the AUC curve of the water cup with and without fusion features;

[0034] Fig.10This application presents the AUC curve of the candy box with and without fusion features;

[0035] Fig.11 A schematic diagram of functional modules of a six-dimensional posture recognition system for an object provided in one embodiment of the present application;

[0036] Fig.12 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0037] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0038] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0039] The object six-dimensional posture recognition method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, or it can be integrated on the server 104, or it can be placed on the cloud or other servers. The terminal 102 can send an RGB-D image of the object to be processed to the server 104. After the server 104 receives the RGB-D image of the object, it processes the RGB-D image to determine the six-dimensional posture of the object. The server 104 can feedback the obtained six-dimensional posture of the object to the terminal 102. In addition, in some embodiments, the six-dimensional posture recognition method of the object can also be implemented by the server 104 or the terminal 102 alone.

[0040] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.

[0041] In an exemplary embodiment, Figure 2 and Figure 3As shown, a method for six-dimensional posture recognition of an object is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used for explanation, and the steps include the following steps 201 to 205.

[0042] Step 201, obtaining an RGB-D image of an object.

[0043] Step 202: semantically segment the RGB-D image to obtain a target color image and a target depth image.

[0044] Specifically, the semantic segmentation framework in PoseCNN is used to perform semantic segmentation on the RGB-D image to obtain a target color image and a target depth image.

[0045] In an exemplary embodiment, step 202 includes steps 21 to 27 below.

[0046] Step 21, using 13 convolutional layers and 4 maximum pooling layers to extract the features of the RGB-D image to obtain the first feature map and the second feature map.

[0047] The number of channels of the first feature map and the second feature map are both 512, the resolution of the first feature map is 1 / 8 of the RGB-D image, and the resolution of the second feature map is 1 / 16 of the RGB-D image.

[0048] like Figure 4 As shown, the connection between the 13 convolutional layers (blue cubes) and the 4 maximum pooling layers (green cubes) is: the 1st convolutional layer - the 2nd convolutional layer - the 1st maximum pooling layer - the 3rd convolutional layer - the 4th convolutional layer - the 2nd maximum pooling layer - the 5th convolutional layer - the 6th convolutional layer - the 7th convolutional layer - the 3rd maximum pooling layer - the 8th convolutional layer - the 9th convolutional layer - the 10th convolutional layer - the 4th maximum pooling layer - the 11th convolutional layer - the 12th convolutional layer - the 13th convolutional layer.

[0049] The number of channels of the feature map output by the first maximum pooling layer is 64, the number of channels of the feature map output by the second maximum pooling layer is 128, the number of channels of the feature map output by the third maximum pooling layer is 256, the number of channels of the feature map output by the fourth maximum pooling layer is 512, and the number of channels of the feature maps output by the 11th to 13th convolutional layers are all 512.

[0050] In step 22, the number of channels of the first feature map and the second feature map is reduced to 64 through two convolutional layers (the kernel size of the first convolutional layer is 1×1 convolution, the kernel size of the second convolutional layer is 3×3, and the step size is 1), and the third feature map and the fourth feature map are obtained.

[0051] In step 23, the resolution of the fourth feature map is doubled through a deconvolution layer (with a kernel size of 4×4, a stride of 2, and no padding) to obtain the fifth feature map.

[0052] In step 24, the third feature map is added to the fifth feature map, and the resolution is increased by 8 times through a deconvolution layer (kernel size is 16×16, stride is 8, and no padding is used) to obtain the sixth feature map. The resolution of the sixth feature map is the same as the resolution of the RGB-D image.

[0053] Step 25, generate the semantic label score of each pixel in the sixth feature map through a convolutional layer (with a sum size of 1×1).

[0054] Step 26, the semantic label score of each pixel is normalized using the softmax function so that the sum of the semantic label scores of each pixel is 1, thereby determining the probability distribution of each pixel belonging to each semantic category. For each pixel, the semantic category with the highest probability is selected as the semantic category of the pixel to generate a semantic segmentation map.

[0055] Step 27, determine the bounding box of the target object according to the semantic segmentation map. Use the bounding box to crop the color depth image pair from the RGB-D image to obtain a target color image and a target depth image.

[0056] Step 203: extract color features from the target color image, and extract geometric features from the target depth image.

[0057] In an exemplary embodiment, extracting color features from the target color image specifically includes: encoding the target color image using a pre-trained encoder to obtain an encoded image. Decoding the encoded image using a pre-trained decoder to obtain color features. The encoder is Resnet-18. The decoder includes 4 upsampling layers.

[0058] The encoder-decoder architecture can generate a target color image of size H×W×3 into a color image of size H×W×d rgb The characteristic diagram of rgb is the dimension of each pixel in the feature map.

[0059] In an exemplary embodiment, extracting geometric features from the target depth image specifically includes: projecting the target depth image onto a point cloud. Using a k-nearest neighbors algorithm (kNN) to construct a local map for each point on the point cloud. Based on the local map of each point, using a pre-trained second multi-layer perceptron to determine the geometric features.

[0060] The target depth image is projected onto the point cloud according to the camera intrinsic matrix. The number of layer neurons of the second multilayer perceptron is {3, 64, 64, 64}.

[0061] Step 204: Use an attention mechanism to perform multi-feature progressive fusion of the color feature and the geometric feature to obtain a fused feature.

[0062] In an exemplary embodiment, step 204 includes the following steps 41 to 43 .

[0063] Step 41, position encoding the color feature to obtain more high-frequency information. Specifically, the color feature is position encoded using the following formula:

[0064] H(x)=(sin(2 0 πx),cos(2 0 πx),…,sin(2 L-1 πx),cos(2 L-1 πx));

[0065] Among them, x is the color feature, H(x) is the color feature after position encoding, and L is the encoding dimension.

[0066] Step 42, using the self-attention layer to learn the self-attention features of the color features after position encoding and the self-attention features of the geometric features, and obtain the color self-attention feature e self1 and geometric self-attention feature e self2 .

[0067] Step 43, using a cross attention layer to fuse the color self-attention feature and the geometric self-attention feature to obtain a fused feature e cross .

[0068] Specifically, the process features F learned by the self-attention layer and the cross-attention layer are e for:

[0069] F e =∑ j:(i,j)∈e a ij V j ;

[0070] Among them, e is the attention feature set, e∈{e self1 ,e self2 ,e cross}, a ij is the attention weight, Q i According to the weight relationship K j With feature V j More representative features are obtained.

[0071] Assumptions is the fusion feature of element i at layer l, then the feature transfer update process can be defined as:

[0072]

[0073] Among them, || represents a cascade operation, and MLP is a multi-layer perceptron.

[0074] Step 205: Determine the six-dimensional pose of the object using a pre-trained first multi-layer perceptron based on the fused features. The present application uses a first multi-layer perceptron to fuse features to perform pixel-level six-dimensional pose estimation of the object.

[0075] To train the first multi-layer perceptron, the present application defines a loss function based on the Euclidean distance between the points after the real pose transformation and the points after the predicted pose transformation:

[0076]

[0077] Among them, l m is the loss function value, N is the number of selected points, R represents the rotation matrix, t represents the translation matrix, [R′ m |t′ m ], m∈N is the estimated six-dimensional pose of the object, [R|t] is the real six-dimensional pose of the object, x n is a point of the three-dimensional model.

[0078] In order to further improve the accuracy of the six-dimensional posture, the object six-dimensional posture recognition method further includes the following step 206.

[0079] Step 206, optimizing the six-dimensional posture of the object.

[0080] This application trains a pose residual estimation network, regards the previously predicted pose as the standard frame estimate of the target object, and converts the target depth image into point cloud data according to the current six-dimensional pose of the object. The transformed point cloud implies the six-dimensional pose of the object. Then, the converted point cloud is input into the feature fusion network, new features are re-extracted for fusion, and then the six-dimensional pose of the object is estimated based on the fused features. This process can be applied iteratively and produces potentially more refined poses in each iteration. After multiple iterations, the final six-dimensional pose of the object is obtained. In order to reduce training time, the network is refined after the main network converges.

[0081] This application is more robust to occluded and low-texture objects, and uses depth images to compensate for misestimation caused by missing geometric information. In order to effectively extract geometric information, this application uses the k-nearest neighbor algorithm to process point clouds, avoiding the basic limitations of local feature loss caused by previous methods. The use of multi-feature fusion to fully explore the relationship between color and geometric information can extract more representative features and achieve effective expression of features, thereby improving the accuracy of object posture recognition. Compared with the method without multi-feature fusion, the accuracy rate is improved by more than 5%. In addition, this application further optimizes the recognized six-dimensional posture through an end-to-end iterative optimization method, achieving effective estimation of the posture of low-texture and heavily occluded objects.

[0082] This application performs performance evaluation in the LineMOD dataset and the YCB-Video dataset, using the accuracy threshold curve (Area under curve, AUC) and the average closest point distance (ADD(-S)) as performance evaluation indicators for pose estimation. The implementation results show that the ADD(-S) error of this application is less than 3 cm, and the AUC error is less than 30 degrees. As shown in Table 1, the average AUC value of this application on all objects in the YCB-Video dataset is 90.8%, and the average ADD(-S) value is 93.3%. Therefore, this application has a high accuracy and success rate as a whole. At the same time, the ADD(-S) values ​​in estimating the poses of the four objects 004_sugar_box, 009_gelatin_box, 019_pitcher_base and 061_foam_brick reached an accuracy rate of more than 99%, indicating that this application can accurately estimate the pose. As Figures 5 to 8 As shown, the object is highly accurately superimposed on the original image in the form of a red model through the posture recognition of the present application. The present application is more accurate in the six-dimensional posture recognition of the object and has higher robustness for objects lacking low texture. Fig. 9 and Fig.10As shown in the figure, the posture recognition method with fusion features has greatly improved the posture recognition accuracy compared with the posture recognition method without fusion features, which can meet the requirements of robot object grasping and provide technical support for robot object grasping.

[0083] Table 1 Quantitative evaluation results of this application on the YCB-Video dataset

[0084]

[0085]

[0086] Based on the same inventive concept, the embodiment of the present application also provides a six-dimensional posture recognition system for implementing the above-mentioned six-dimensional posture recognition method for objects. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in one or more six-dimensional posture recognition system embodiments provided below can refer to the limitations of the six-dimensional posture recognition method for objects above, and will not be repeated here.

[0087] In an exemplary embodiment, Fig.11 As shown, a six-dimensional posture recognition system for an object is provided, including: an image acquisition module 301, a semantic segmentation module 302, a feature extraction module 303, a feature fusion module 304, a posture determination module 305 and a posture optimization module 306.

[0088] The image acquisition module 301 is used to acquire an RGB-D image of an object.

[0089] The semantic segmentation module 302 is used to perform semantic segmentation on the RGB-D image to obtain a target color image and a target depth image.

[0090] The feature extraction module 303 is used to extract color features from the target color image and extract geometric features from the target depth image.

[0091] The feature fusion module 304 is used to use the attention mechanism to perform multi-feature progressive fusion on the color feature and the geometric feature to obtain a fused feature.

[0092] The posture determination module 305 is used to determine the six-dimensional posture of the object according to the fusion features by using a pre-trained first multi-layer perceptron.

[0093] The posture optimization module 306 is used to optimize the six-dimensional posture of the object.

[0094] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Fig.12As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store RGB-D images of objects. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a six-dimensional posture recognition method of an object is implemented.

[0095] Those skilled in the art will understand that Fig.12 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0096] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0097] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0098] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0099] In this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0100] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0101] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0102] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0103] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application; at the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for six-dimensional posture recognition of an object, characterized in that: The object six-dimensional posture recognition method comprises: Get the RGB-D image of the object; Performing semantic segmentation on the RGB-D image to obtain a target color image and a target depth image; Extracting color features from the target color image and extracting geometric features from the target depth image; Using an attention mechanism to perform multi-feature progressive fusion on the color feature and the geometric feature to obtain a fusion feature; According to the fusion features, a pre-trained first multi-layer perceptron is used to determine the six-dimensional posture of the object.

2. The object six-dimensional posture recognition method according to claim 1, characterized in that: Performing semantic segmentation on the RGB-D image to obtain a target color image and a target depth image specifically includes: The semantic segmentation framework in PoseCNN is used to perform semantic segmentation on the RGB-D image to obtain a target color image and a target depth image.

3. The object six-dimensional posture recognition method according to claim 1, characterized in that: Extracting color features from the target color image specifically includes: Encode the target color image using a pre-trained encoder to obtain an encoded image; the encoder is Resnet-18; The encoded image is decoded using a pre-trained decoder to obtain color features; the decoder includes 4 upsampling layers.

4. The object six-dimensional posture recognition method according to claim 1, characterized in that: Extracting geometric features from the target depth image specifically includes: Projecting the target depth image onto a point cloud; Using a k-nearest neighbor algorithm to construct a local graph for each point on the point cloud; According to the local map of each point, a pre-trained second multi-layer perceptron is used to determine the geometric features.

5. The object six-dimensional posture recognition method according to claim 1, characterized in that: The attention mechanism is used to perform multi-feature progressive fusion of the color feature and the geometric feature to obtain fusion features, which specifically include: Position encoding the color feature; Using a self-attention layer to learn the self-attention features of the color features after position encoding and the self-attention features of the geometric features, respectively, to obtain color self-attention features and geometric self-attention features; The color self-attention feature and the geometric self-attention feature are fused using a cross attention layer to obtain a fused feature.

6. The object six-dimensional posture recognition method according to claim 5, characterized in that: The color feature is positionally encoded using the following formula: H(x)=(sin(2 0 πx),cos(2 0 πx),…,sin(2 L-1 πx),cos(2 L-1 πx)); Among them, x is the color feature, H(x) is the color feature after position encoding, and L is the encoding dimension.

7. The object six-dimensional posture recognition method according to claim 1, characterized in that: The object six-dimensional posture recognition method also includes: The six-dimensional posture of the object is optimized.

8. A six-dimensional posture recognition system for an object, applied to the six-dimensional posture recognition method for an object according to any one of claims 1 to 7, characterized in that: The object six-dimensional posture recognition system comprises: An image acquisition module, used to acquire an RGB-D image of an object; A semantic segmentation module, used to perform semantic segmentation on the RGB-D image to obtain a target color image and a target depth image; A feature extraction module, used to extract color features from the target color image and extract geometric features from the target depth image; A feature fusion module, used for performing multi-feature progressive fusion of the color feature and the geometric feature by using an attention mechanism to obtain a fusion feature; The posture determination module is used to determine the six-dimensional posture of the object according to the fusion features by using a pre-trained first multi-layer perceptron.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the six-dimensional posture recognition method of an object according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the six-dimensional posture recognition method of an object described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • 6D attitude estimation method based on deep learning

    CN114742888A

Cited By

  • Pallet goods identification method and device and storage medium

    CN121640040A