In-cabin grabbing target characterization method based on visual touch fusion

Through the three-dimensional model reconstruction of the targets that capture in the space station cabin, the method of fusing visual and tactile information is solved, and the reliability of the robot's grasping task is improved.

CN119991944APending Publication Date: 2025-05-13江淮前沿技术协同创新中心 +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510040405.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the space station cabin operation, the three-dimensional model reconstruction method of the 3D model of the grab target based on robot vision has blind spots in the field of view, resulting in global information loss, low reconstruction accuracy, and unable to support complete three-dimensional target reconstruction, which in turn affects the reliability of the robot's grab task.

Method used

A method of in-cabin grab target representation based on visual-tactile fusion is proposed. By acquiring single-view visual point clouds, blind-spot tactile point clouds and real point clouds, a three-dimensional model is trained to reconstruct the network, and fuse visual and tactile information to improve reconstruction accuracy.

Benefits of technology

It effectively improves the gripping target perception ability and three-dimensional model reconstruction accuracy of the robot in the cabin, makes up for the shortcomings of the visual reconstruction method, and improves the reliability of the gripping task in the robot in the cabin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991944A_ABST
    Figure CN119991944A_ABST
Patent Text Reader

Abstract

The invention discloses an in-cabin grabbing target characterization method based on visual touch fusion. The in-cabin grabbing target characterization method comprises the steps that an in-cabin grabbing target three-dimensional model reconstruction data set based on visual touch fusion information is acquired; training an initial three-dimensional model reconstruction network by using the data set to obtain an in-cabin target grabbing three-dimensional model reconstruction network based on visual touch fusion information; and the in-cabin target single-visual-angle visual point cloud and the visual field blind area touch point cloud of which the three-dimensional model needs to be reconstructed are input into the in-cabin grabbing target three-dimensional model reconstruction network based on the visual touch fusion information, and the in-cabin grabbing target three-dimensional model is obtained. According to the method, the in-cabin grabbing target three-dimensional model is reconstructed with high precision in a visual touch fusion mode, the defects of a robot vision-based grabbing target three-dimensional model reconstruction method are overcome, the precision of robot grabbing target reconstruction is improved, and then the reliability of a robot in-cabin grabbing task in a space station is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for representing a target to be captured in a cabin based on visual-touch fusion, and belongs to the fields of multimodal perception technology and artificial intelligence technology. [Background technology]

[0002] With the continuous development of multimodal perception technology and the improvement of performance level, the demand for target reconstruction is also moving towards complex and specialized scenes, as well as in the direction of multi-sensor fusion.

[0003] Three-dimensional reconstruction is an important means of characterizing the grasping target in the cabin. However, in the scenario of working in the space station cabin, the visual information is often from a single perspective, with a large blind spot, which easily causes global information loss, making it insufficient to describe the geometric shape information of the object, unable to support relatively complete three-dimensional target reconstruction, and the reconstruction accuracy is low, which brings challenges to the task of reconstructing the three-dimensional model of the grasping target in the cabin, and also brings difficulties to the subsequent robot grasping and task operation in the space cabin. Therefore, it is considered to make up for the shortcomings of the three-dimensional model reconstruction method of the grasping target based on robot vision through tactile information, improve the accuracy of the robot grasping target reconstruction, and thus ensure the reliability of the robot's grasping task in the cabin. [Summary of the invention]

[0004] In view of this, the present invention proposes a method for characterizing in-cabin targets based on vision-touch fusion for reconstructing a three-dimensional model of in-cabin targets, which can effectively improve the model accuracy.

[0005] The present invention provides a method for representing a target captured in a cabin based on visual-touch fusion, comprising:

[0006] (1) Obtaining a dataset of three-dimensional model reconstruction of a grasping target in a cabin based on visual-tactile fusion information, wherein the dataset includes a single-view visual point cloud, a tactile point cloud in a blind area of ​​vision, and a real point cloud;

[0007] (2) Using the single-view visual point cloud, blind-spot tactile point cloud, and real point cloud in the in-cabin grasping target 3D model reconstruction dataset based on visual-tactile fusion information to train the initialized 3D model reconstruction network, and obtain the in-cabin grasping target 3D model reconstruction network based on visual-tactile fusion information;

[0008] (3) The single-view visual point cloud of the in-cabin target whose three-dimensional model needs to be reconstructed and the tactile point cloud of the blind spot of the field of vision are input into the in-cabin grasping target three-dimensional model reconstruction network based on visual and tactile fusion information to obtain the in-cabin grasping target three-dimensional model.

[0009] In the above method, the step (1) specifically comprises:

[0010] (1.1) Selecting a number of objects in the cabin with different surface shapes, and using a depth camera in the cabin to obtain the single-view visual point cloud;

[0011] (1.2) using a robot hand equipped with a tactile sensor on its fingertips to actively touch one side of the blind area of ​​the depth camera's field of view, collect a pressure array, and convert the pressure array into a tactile point cloud of the blind area of ​​the field of view;

[0012] (1.3) Scanning the objects in the cabin using a point cloud scanner to obtain the real point cloud;

[0013] (1.4) Execute steps (1.1) to (1.3) M times to obtain a data set of M groups of data samples for reconstructing a three-dimensional model of the in-cabin grasping target based on visual-tactile fusion information.

[0014] In the above method, the step (2) specifically comprises:

[0015] (2.1) inputting the single-view visual point cloud and the blind-spot tactile point cloud into a center point extraction module to obtain a visual center point and a tactile center point;

[0016] (2.2) inputting the visual center point and the tactile center point into a geometric perception encoder module respectively to obtain visual features and tactile features;

[0017] (2.3) inputting the visual features and the tactile features into a cross fusion module to obtain visual-tactile spatial fusion features;

[0018] (2.4) inputting the visual-tactile spatial fusion feature into a geometric perception decoder module to obtain a predicted point cloud;

[0019] (2.5) Based on the predicted point cloud and the real point cloud, determine whether the three-dimensional model reconstruction network has reached the preset convergence condition of the loss function value. If so, use the three-dimensional model reconstruction network as a new three-dimensional model reconstruction network for in-cabin grasping targets based on visual-touch fusion information; if not, execute steps (2.1) to (2.4) again until the three-dimensional model reconstruction network reaches the preset convergence condition.

[0020] Wherein, the loss function includes:

[0021] set up Contains R Points, is the real point cloud, including n G points, the loss function J is expressed as:

[0022] J=J0+J1

[0023] Among them, J0 and J1 can be expressed as:

[0024]

[0025] Among them, J0 represents the one-way distance from the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information to the real point cloud; J1 represents the one-way distance from the real point cloud to the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information. The two are combined to form a bidirectional distance, ensuring that the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information matches the geometric structure of the real point cloud.

[0026] In the above method, converting the pressure array into a tactile point cloud comprises:

[0027] The pressure array is represented as F={F(i,j)|i=1,...,m; j=1,...,n}, where m×n is the number of pressure measurement points of the tactile sensor, and F(i,j) is the force value measured by the measurement point with coordinates (i,j);

[0028] According to the nonlinear mechanical behavior of the surface material of the tactile sensor, each pressure value F(i,j) is mapped to the corresponding deformation height h(i,j):

[0029] h(i,j)=C·(F(i,j)) k

[0030] Wherein: C is the proportionality coefficient, which is related to the elastic coefficient of the material; k is the nonlinear index, which can be selected in the range of 0.5 to 0.8.

[0031] In the above method, the center point extraction module includes:

[0032] (3.1) Given a point cloud Where N0 is the number of points in the input point cloud;

[0033] (3.2) Specify the number of center point cloud samples N1;

[0034] (3.3) Randomly select a point s1∈S in the point cloud as the first center point and initialize the center point set S t ={s1}, and delete S in S t The points included.

[0035] (3.4) Select p that satisfies the following conditions as the next center point, that is, S t The next element of S and delete p from S:

[0036]

[0037] Among them, the variable q represents a point in the point cloud S.

[0038] (3.5) Repeat step (3.4) until the center point cloud S t The number of elements reaches the specified center point cloud sampling number N1.

[0039] In the above method, the geometry-aware encoder comprises:

[0040] Using the fully connected layer Mapping to point cloud features

[0041] S f =φj f (W f S t +b f )

[0042] Among them, B is the number of data samples in a batch processed by the network, C is the number of features, and W f is the weight matrix of the fully connected layer, b f is the bias, φ f It is the nonlinear activation function Relu;

[0043] Then, according to the characteristic number, S f Divide into blocks and obtain point cloud block features Where C2 is the specified feature dimension of the point cloud block feature, and Ns is the number of point cloud blocks;

[0044] Then the Transformer network is used to encode the block point cloud to obtain the feature

[0045] In the above method, the cross-fusion module includes:

[0046] For the encoded visual and tactile features Build a cross-fusion module based on the dot product attention mechanism:

[0047]

[0048] The visual-tactile spatial fusion feature is the output of the attention head, and Softmax is the normalized exponential function.

[0049] In the above method, the geometry-aware decoder comprises:

[0050] The Transformer network is used to fuse the visual and tactile spatial features. Decode and reconstruct the visual and tactile spatial features obtained using the Transformer network Decode and obtain the reconstructed proxy features

[0051] For N1 intra-block features Make it weighted summation within the block to the global feature

[0052] The global features are transformed into Mapping

[0053] Y3=φ g (W g Y2+b g )

[0054] Among them, W g is the weight matrix of the fully connected layer, b g is the bias, φ g It is the nonlinear activation function Relu;

[0055] Using a multi-layer perceptron, the complete global features are used as input to restore the complete global three-dimensional spatial form, and a three-dimensional model of the in-cabin grasping target based on visual-tactile fusion information is obtained. Where N is the number of model points required for reconstructing the three-dimensional model of the grasping target in the cabin based on the visual-touch fusion information:

[0056]

[0057] Among them, W1, W2, W3 are the weight matrices of the multi-layer perceptron, b1, b2, b3 are biases, and φ is the nonlinear activation function Relu.

[0058] It can be seen from the above technical solutions that the embodiments of the present invention have the following beneficial effects:

[0059] The present invention adopts a method for representing in-cabin grasping targets based on vision-touch fusion, uses devices such as depth cameras, tactile sensors and point cloud scanners to obtain single-view visual point clouds, blind-area tactile point clouds and real point clouds, and trains an initial 3D model reconstruction network to obtain a three-dimensional model reconstruction network for in-cabin grasping targets based on vision-touch fusion information, and then uses the network and the given single-view visual point cloud and blind-area tactile point cloud input to infer the three-dimensional model of the in-cabin grasping target, which can effectively improve the in-cabin robot's grasping target perception ability and the accuracy of the robot's grasping target reconstruction, thereby ensuring the reliability of the robot's in-cabin grasping task.

Brief Description of the Drawings

[0060] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It is obvious that the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creativity and labor.

[0061] Figure 1 It is a flow chart of a method for representing a target captured in a cabin based on visual-touch fusion provided by an embodiment of the present invention;

[0062] Figure 2 It is a schematic diagram of a three-dimensional model reconstruction network for grabbing a target in a cabin based on visual-tactile fusion information in an embodiment of the present invention. [Specific implementation method]

[0063] In order to better understand the technical solution of the present invention, the embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0064] It should be clear that the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0065] The embodiment of the present invention provides a method for representing a target captured in a cabin based on visual-touch fusion. Figure 1 , which is a flow chart of a method for representing a target captured in a cabin based on visual-touch fusion provided by an embodiment of the present invention, such as Figure 1 As shown, the method comprises the following steps:

[0066] Step 101, obtaining a dataset of reconstructing a three-dimensional model of a grasping target in a cabin based on visual-tactile fusion information, wherein the dataset includes a single-view visual point cloud, a tactile point cloud in a blind area of ​​vision, and a real point cloud.

[0067] (1.1) Selecting a number of objects in the cabin with different surface shapes, and using a depth camera in the cabin to obtain the single-view visual point cloud;

[0068] (1.2) Using a robot hand equipped with a tactile sensor on its fingertips to actively touch one side of the blind area of ​​the depth camera, collect a pressure array, and convert the pressure array into a tactile point cloud of the blind area of ​​the field of view. The specific method is as follows:

[0069] The pressure array is represented as F={F(i,j)|i=1,...,m; j=1,...,n}, where m×n is the number of pressure measurement points of the tactile sensor, and F(i,j) is the force value measured by the measurement point with coordinates (i,j);

[0070] According to the nonlinear mechanical behavior of the surface material of the tactile sensor, each pressure value F(i,j) is mapped to the corresponding deformation height h(i,j):

[0071] h(i,j)=C·(F(i,j)) k

[0072] Wherein: C is the proportionality coefficient, which is related to the elastic coefficient of the material; k is the nonlinear index, which can be selected in the range of 0.5 to 0.8.

[0073] (1.3) Scanning the objects in the cabin using a point cloud scanner to obtain the real point cloud;

[0074] (1.4) Execute steps (1.1) to (1.3) M times to obtain a data set of M groups of data samples for reconstructing a three-dimensional model of the in-cabin grasping target based on visual-tactile fusion information.

[0075] Step 102, use the single-view visual point cloud, the blind spot tactile point cloud and the real point cloud in the three-dimensional model reconstruction data set of the in-cabin grasping target based on vision-touch fusion information to train the initial three-dimensional model reconstruction network to obtain the in-cabin grasping target three-dimensional model reconstruction network based on vision-touch fusion information.

[0076] Please refer to Figure 2 , which is a schematic diagram of a three-dimensional model reconstruction network for grabbing targets in a cabin based on visual-touch fusion information in an embodiment of the present invention.

[0077] First, the single-view visual point cloud and the blind spot tactile point cloud are respectively input into the center point extraction module to obtain the visual center point and the tactile center point. The specific method is as follows:

[0078] (3.1) Given a point cloud Where N0 is the number of points in the input point cloud;

[0079] (3.2) Specify the number of center point cloud samples N1;

[0080] (3.3) Randomly select a point s1∈S in the point cloud as the first center point and initialize the center point set S t = {s1}, and delete S in S t The points included.

[0081] (3.4) Select p that satisfies the following conditions as the next center point, that is, S t The next element of S and delete p from S:

[0082]

[0083] Among them, the variable q represents a point in the point cloud S.

[0084] (3.5) Repeat step (3.4) until the center point cloud S t The number of elements reaches the specified center point cloud sampling number N1.

[0085] Secondly, the visual center point and the tactile center point are respectively input into the geometric perception encoder module to obtain the visual features and the tactile features. The specific method is as follows:

[0086] Using the fully connected layer Mapping to point cloud features

[0087] S f =φ f (W f S t +b f )

[0088] Among them, B is the number of data samples in a batch processed by the network, C is the number of features, and W f is the weight matrix of the fully connected layer, b f is the bias, φ f It is the nonlinear activation function Relu;

[0089] Then, according to the characteristic number, S f Divide into blocks and obtain point cloud block features Where C2 is the specified feature dimension of the point cloud block feature, and Ns is the number of point cloud blocks;

[0090] Then the Transformer network is used to encode the block point cloud to obtain the feature

[0091] Subsequently, the visual features and tactile features are input into a cross fusion module to obtain visual-tactile spatial fusion features. The specific method is as follows:

[0092] For the encoded visual and tactile features Build a cross-fusion module based on the dot product attention mechanism:

[0093]

[0094] The visual-tactile spatial fusion feature is the output of the attention head, and Softmax is the normalized exponential function.

[0095] Then, the visual-tactile spatial fusion feature is input into the geometric perception decoder module to obtain the predicted point cloud. The specific method is as follows:

[0096] The Transformer network is used to fuse the visual and tactile spatial features. Decode and obtain the reconstructed proxy features

[0097] For N1 intra-block features Make it weighted summation within the block to the global feature

[0098] The global features are transformed into Mapping

[0099] Y3=φ g (W g Y2+b g )

[0100] Among them, W g is the weight matrix of the fully connected layer, b g is the bias, φ g It is the nonlinear activation function Relu;

[0101] Using a multi-layer perceptron, the complete global features are used as input to restore the complete global three-dimensional spatial form, and a three-dimensional model of the in-cabin grasping target based on visual-tactile fusion information is obtained. Where N is the number of model points required for reconstructing the three-dimensional model of the grasping target in the cabin based on the visual-touch fusion information:

[0102]

[0103] Among them, W1, W2, W3 are the weight matrices of the multi-layer perceptron, b1, b2, b3 are biases, and φ is the nonlinear activation function Relu.

[0104] Finally, based on the predicted point cloud and the real point cloud, it is determined whether the current 3D model reconstruction network has reached the preset convergence condition of the loss function value. If so, the current 3D model reconstruction network is used as the in-cabin grasping target 3D model reconstruction network based on visual-touch fusion information; if not, the training is repeated until the 3D model reconstruction network reaches the preset convergence condition.

[0105] Wherein, the loss function includes:

[0106] set up Contains R Points, is the real point cloud, including n G points, the loss function J is expressed as:

[0107] J=J0+J1

[0108] Among them, J0 and J1 can be expressed as:

[0109]

[0110] Among them, J0 represents the one-way distance from the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information to the real point cloud; J1 represents the one-way distance from the real point cloud to the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information. The two are combined to form a bidirectional distance, ensuring that the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information matches the geometric structure of the real point cloud.

[0111] Step 103 , input the single-view visual point cloud and the blind-spot tactile point cloud of the in-cabin target to be reconstructed into the in-cabin target grasping three-dimensional model reconstruction network trained based on visual-tactile fusion information to obtain the in-cabin target grasping three-dimensional model.

[0112] The single-view visual point cloud of the in-cabin target to be reconstructed and the tactile point cloud of the blind area of ​​the field of vision are input into the trained in-cabin grasping target 3D model reconstruction network based on visual-tactile fusion information to obtain the in-cabin grasping target 3D model result.

[0113] According to the above method provided by an embodiment of the present invention, the initial terrain classification model is trained using a portion of the acquired training data set, and the accuracy of the initial in-cabin target grasping three-dimensional model reconstruction network based on visual-touch fusion information is tested using a portion of the acquired training data set and compared with the existing method. The evaluation method adopts the widely used average chamfer distance, which is calculated as follows.

[0114] The average chamfer distance is an indicator that has nothing to do with the point arrangement. It is invariant to the disorder of the point cloud and is suitable for measuring the similarity between point clouds:

[0115]

[0116] Among them, A is the predicted point cloud and B is the original point cloud.

[0117] The results are shown in Table 1:

[0118] Table 1 Comparison between the present invention and the existing method

[0119]

[0120] It can be seen from Table 1 that the in-cabin grasping target representation method based on vision-touch fusion proposed in the present invention can effectively realize the three-dimensional reconstruction and representation of the grasped target, and the average chamfer distance is only 1.371, which is better than the existing method.

[0121] The technical solution of the embodiment of the present invention has the following beneficial effects:

[0122] In the technical solution of the embodiment of the present invention, a method for representing the in-cabin grasping target based on visual-touch fusion is used to realize the reconstruction of the three-dimensional model of the in-cabin grasping target. The initial three-dimensional model reconstruction network is trained by using the single-view visual point cloud, the tactile point cloud in the blind spot of the field of vision and the real point cloud in the in-cabin grasping target three-dimensional model reconstruction data set based on visual-touch fusion information to obtain the in-cabin grasping target three-dimensional model reconstruction network based on visual-touch fusion information. The single-view visual point cloud and the tactile point cloud in the blind spot of the field of vision of the in-cabin target to be reconstructed are input into the initial in-cabin grasping target three-dimensional model reconstruction network based on visual-touch fusion information obtained by training, and the in-cabin grasping target three-dimensional model is obtained. The present invention effectively improves the grasping perception ability of the robot in the space station cabin, can reconstruct the three-dimensional model of the in-cabin grasping target with higher precision, makes up for the shortcomings of the grasping target three-dimensional model reconstruction method based on robot vision, improves the accuracy of the robot grasping target reconstruction, and thus ensures the reliability of the robot's grasping task in the space station cabin.

[0123] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

[0124] The contents not described in detail in the specification of the present invention belong to the common knowledge of those skilled in the art.

Claims

1. A method for representing in-cabin grasping targets based on visual-touch fusion, characterized in that: The method comprises the following steps: (1) Obtaining a dataset of three-dimensional model reconstruction of a grasping target in a cabin based on visual-tactile fusion information, wherein the dataset includes a single-view visual point cloud, a tactile point cloud in a blind area of ​​vision, and a real point cloud; (2) Using the single-view visual point cloud, blind-spot tactile point cloud, and real point cloud in the in-cabin grasping target 3D model reconstruction dataset based on visual-tactile fusion information to train the initialized 3D model reconstruction network, and obtain the in-cabin grasping target 3D model reconstruction network based on visual-tactile fusion information; (3) The single-view visual point cloud of the in-cabin target whose three-dimensional model needs to be reconstructed and the tactile point cloud of the blind spot of the field of vision are input into the in-cabin grasping target three-dimensional model reconstruction network based on visual and tactile fusion information to obtain the in-cabin grasping target three-dimensional model.

2. The method for characterizing in-cabin grasping targets based on visual-touch fusion according to claim 1, characterized in that: The step (1) specifically comprises: (1.1) Selecting a number of objects in the cabin with different surface shapes, and using a depth camera in the cabin to obtain the single-view visual point cloud; (1.2) using a robot hand equipped with a tactile sensor on its fingertips to actively touch one side of the blind area of ​​the depth camera's field of view, collect a pressure array, and convert the pressure array into a tactile point cloud of the blind area of ​​the field of view; (1.3) Scanning the objects in the cabin using a point cloud scanner to obtain the real point cloud; (1.4) Execute steps (1.1) to (1.3) M times to obtain a data set of M groups of data samples for reconstructing a three-dimensional model of the in-cabin grasping target based on visual-tactile fusion information.

3. The method according to claim 1, characterized in that The step (2) specifically comprises: (2.1) inputting the single-view visual point cloud and the blind-spot tactile point cloud into a center point extraction module to obtain a visual center point and a tactile center point; (2.2) inputting the visual center point and the tactile center point into a geometric perception encoder module respectively to obtain visual features and tactile features; (2.3) inputting the visual features and the tactile features into a cross fusion module to obtain visual-tactile spatial fusion features; (2.4) inputting the visual-tactile spatial fusion feature into a geometric perception decoder module to obtain a predicted point cloud; (2.5) Based on the predicted point cloud and the real point cloud, determine whether the three-dimensional model reconstruction network has reached the preset convergence condition of the loss function value. If so, use the three-dimensional model reconstruction network as a new three-dimensional model reconstruction network for in-cabin grasping targets based on visual-touch fusion information; if not, execute steps (2.1) to (2.4) again until the three-dimensional model reconstruction network reaches the preset convergence condition.

4. The method according to claim 2, characterized in that: The method of converting the pressure array into a tactile point cloud of the blind area of ​​the visual field comprises: The pressure array is represented as F={F(i,j)|i=1,…,m; j=1,…,n}, where m×n is the number of pressure measurement points of the tactile sensor, and F(i,j) is the force value measured by the measurement point with coordinates (i,j); According to the nonlinear mechanical behavior of the surface material of the tactile sensor, each pressure value F(i,j) is mapped to the corresponding deformation height h(i,j): h(i,j)=C·(F(i,j)) k Wherein: C is the proportionality coefficient, which is related to the elastic coefficient of the material; k is the nonlinear index, which can be selected in the range of 0.5 to 0.

8.

5. The method according to claim 3, characterized in that: The center point extraction module comprises: (3.1) Given a point cloud Where N0 is the number of points in the input point cloud; (3.2) Specify the number of center point cloud samples N1; (3.3) Randomly select a point s1∈S in the point cloud as the first center point and initialize the center point set S t = {s1}, and delete S in S t The points included; (3.4) Select p that satisfies the following conditions as the next center point, that is, S t The next element of S and delete p from S: Among them, the variable q represents a point in the point cloud S; (3.5) Repeat step (3.4) until the center point cloud S t The number of elements reaches the specified center point cloud sampling number N1.

6. The method according to claim 3, characterized in that The geometry-aware encoder comprises: Using the fully connected layer Mapping to point cloud features S f =φ f (W f S t +b f ) Among them, B is the number of data samples in a batch processed by the network, C is the number of features, and W f is the weight matrix of the fully connected layer, b f is the bias, φ f It is the nonlinear activation function Relu; Then, according to the characteristic number, S f Divide into blocks and obtain point cloud block features Where C2 is the specified feature dimension of the point cloud block feature, and Ns is the number of point cloud blocks; Then the Transformer network is used to encode the block point cloud to obtain the feature 7. The method according to claim 3, characterized in that The cross-fusion module specifically includes: For the encoded visual and tactile features Build a cross-fusion module based on the dot product attention mechanism: The visual-tactile spatial fusion feature is the output of the attention head, and Softmax is the normalized exponential function.

8. The method according to claim 3, characterized in that The geometry-aware decoder module comprises: The Transformer network is used to fuse the visual and tactile spatial features. Decode and obtain the reconstructed proxy features For N1 intra-block features Make it weighted summation within the block to the global feature The global features are transformed into Mapping Y3=φ g (W g Y2+b g ) Among them, W g is the weight matrix of the fully connected layer, b g is the bias, φ g It is the nonlinear activation function Relu; Using a multi-layer perceptron, the complete global features are used as input to restore the complete global three-dimensional spatial form, and a three-dimensional model of the in-cabin grasping target based on visual-tactile fusion information is obtained. Where N is the number of model points required for reconstructing the three-dimensional model of the grasping target in the cabin based on the visual-touch fusion information: Among them, W1, W2, W3 are the weight matrices of the multi-layer perceptron, b1, b2, b3 are biases, and φ is the nonlinear activation function Relu.

9. The method according to claim 3, characterized in that: The loss function includes: set up Contains R Points, is the real point cloud, including n G points, the loss function J is expressed as: J=J0+J1 Among them, J0 and J1 can be expressed as: Among them, J0 represents the one-way distance from the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information to the real point cloud; J1 represents the one-way distance from the real point cloud to the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information. The two are combined to form a bidirectional distance, ensuring that the three-dimensional model of the in-cabin target grasping based on vision-touch fusion information matches the geometric structure of the real point cloud.

Citation Information

Cited By

  • A point cloud-based cross-modal alignment visual-haptic fusion object recognition method

    CN122821300A