Soft object clamping state evaluation method based on vision-curvature fusion

Through the visual-curve fusion method, combined with image acquisition and curvature data processing, feature extraction and modal fusion are used using Kalman filtering and YOLOv8 algorithm, which solves the problem of low accuracy of soft object clamping state evaluation and achieves efficient grasping state evaluation.

CN120244967APending Publication Date: 2025-07-04TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510492828.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has low accuracy in the evaluation of soft object clamping state, and cannot effectively deal with the unstructured characteristics of soft objects, resulting in inaccurate grasping evaluation.

Method used

The visual-curvature fusion method is adopted to obtain the image data and curvature data of the object through the image acquisition device and the curvature data acquisition module, and feature extraction and modal fusion are combined with Kalman filtering and YOLOv8 algorithm, and classification is used to realize the evaluation of the clamping state of soft objects.

Benefits of technology

It improves the accuracy of the evaluation of the clamping state of soft objects, enhances the practicality and adaptability of the clamping state of robotic arm, optimizes the signal processing effect, and improves the evaluation accuracy of the clamping state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244967A_ABST
    Figure CN120244967A_ABST
Patent Text Reader

Abstract

The invention discloses a soft object clamping state evaluation method based on vision-curvature fusion, and belongs to the technical field of mechanical arm clamping state evaluation. In order to solve the problem that in an existing method, the accuracy rate of flexible object judgment is low, image data and curvature data of an evaluated object when the evaluated object is grabbed are collected, and the compression ratio of the evaluated object is defined to obtain a clamping state to serve as a label; performing corresponding data preprocessing on the collected image data and curvature data to obtain a data structure type required by the model; dividing the preprocessed image data and curvature data into a training set and a test set, and inputting the training set and the test set into a corresponding feature extraction module for training to obtain feature vectors of the image data and the curvature data; inputting the obtained feature vectors into a modal fusion module, and fusing two original uncorrelated modals to obtain attribute-correlated fusion features; and the obtained fusion features are connected to a classifier for classification, and evaluation of the clamping state of the soft object is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robotic arm grasping state evaluation, and particularly relates to a method for evaluating the grasping state of soft objects based on vision-curvature fusion. Background Art

[0002] With the progress of the times and the development of technology, the demand for robot control and dexterous hand operation in the service and production fields is increasing day by day, and the requirements are getting higher and higher. The grasping ability of robotic dexterous hands has received wide attention. Efficiently and accurately evaluating the grasping state is a very crucial part of improving the overall grasping quality of robotic dexterous hands. The traditional focus of evaluating the grasping quality of robotic arm dexterous hands mainly lies in the visual feedback and force feedback generated after the object slides and undergoes deformation due to over-grasping during the grasping process. Many scholars have conducted research on grasping stability evaluation, slip detection, and deformation detection or prediction. However, the methods adopted are mostly visual-tactile fusion decision-making. The advantage of this method is that it has a relatively accurate detection effect for objects with obvious changes in elastic force after deformation. However, for soft objects, due to their material properties, the change in elastic force feedback during the deformation process is not obvious. Therefore, the visual-tactile fusion method has poor judgment effect on the grasping state. With the continuous in-depth research, the demand for soft object grasping is increasing, and curvature perception has gradually entered various fields. Currently, researchers tend to use curvature perception to conduct research on grasping state evaluation.

[0003] Grasping state evaluation has a wide range of application spaces. In the field of embodied intelligence, through grasping state evaluation, robots can judge in real time whether the grasping action is successful, thereby adjusting the grasping strategy and improving the stability and safety of the operation. In pipeline operations, grasping state evaluation ensures the grasping accuracy of the robotic arm for workpieces, avoiding production interruption or losses caused by grasping failures. In research on exploring the optimal grasping strategy, grasping force control, etc., grasping state evaluation is an important indicator for verifying the performance of algorithms, and its evaluation results can be used to improve the robot control model.

[0004] Chinese invention patent CN107953329B proposes an object recognition and pose estimation method. By using a multi-level deep learning model to perform object category and pose learning on image samples, a set of feature descriptors is established, and object recognition and pose estimation are performed in combination with the set of feature descriptors. The pose estimation is relatively accurate, but it is limited to image acquisition, and the application scenario is relatively limited. US patent application 20240091951A1 proposes a collaborative task-aware grasping estimation method between picking and placing. By extracting the features of the picking action and the placing action for grasping estimation, due to the large dynamic changes in the process, the accuracy is questionable. Summary of the Invention

[0005] Aiming at the problem of low accuracy of grasping evaluation caused by the lack of handling of the unstructured characteristics of soft objects in the existing grasping evaluation methods, the present invention provides a method for evaluating the clamping state of soft objects based on vision-curvature fusion.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for evaluating the clamping state of soft objects based on vision-curvature fusion, the method comprising the following steps:

[0008] Step 1: Use the image acquisition device and the curvature data acquisition module installed on the robot to collect the image data and curvature data of the object to be evaluated when it is being grasped, obtain the clamping state by defining the compression ratio of the object to be evaluated, and use the clamping state as a label;

[0009] Further, the specific operation of step 1 is as follows:

[0010] Step 1.1: Fix the image acquisition device by installing a support on the wrist of the robotic arm, drive the image acquisition device to take pictures at different angles by rotating the support, install a curvature data acquisition module on the gripper, and when the gripper makes an opening and closing movement, the state of the curvature data acquisition module changes;

[0011] Step 1.2: Use the image acquisition device and the curvature data acquisition module to collect the image data and curvature data of the object to be evaluated when it is being grasped, obtain the clamping state by defining the compression ratio of the object to be evaluated, and use the clamping state as a label; The labels include sliding, excessive, appropriate, and extreme; where:

[0012] Sliding: There is relative sliding between the object and the gripper during the grasping process;

[0013] Appropriate: The compression rate of the object during the grasping process does not exceed 10%;

[0014] Excessive: The compression rate of the object during the grasping process does not exceed 70%;

[0015] Extreme: The compression rate of the object during the grasping process exceeds 70%.

[0016] Step 2: Perform corresponding data preprocessing on the collected image data and curvature data respectively to obtain the required data structure types;

[0017] Further, the specific operation of step 2 is as follows:

[0018] Step 2.1: Perform two filtering processes on the curvature data. The first filtering uses the mean filtering method in the front-end acquisition device processor (STM32F103C8T6) of the robot to perform preliminary smoothing processing on the original signal. The formula is:

[0019]

[0020] Wherein, y[n] is the output sequence, the window size is 5; i represents the summation index variable, representing the index position within the current window, n represents the index at the current moment, that is, the index of the data after calculating the mean filter;

[0021] Step 2.2: The signal after preliminary smoothing is sent to the computer through the serial bus on the robot, and the required curvature data is obtained by performing secondary filtering using Kalman filtering;

[0022] The resolution of the collected image is 1280×960, and the input size of the image data required by the model is 640×640. In order to reduce the image size, the image is downsampled and compressed, and the resolution of the compressed image is 640×480. Pixel supplementation operation is performed on the image to obtain an image with a resolution of 640×640.

[0023] The specific process of the Kalman filtering is as follows:

[0024] The input value of the Kalman filter is the observed data sequence x n = z n + v n , where x n , z n and v n represent the observed value, the true position, and the observation noise respectively. The observation noise follows the normal distribution v n ~(0, R), where R = 5;

[0025] In the prediction stage, the Kalman filter predicts the current state based on the state at the previous time, and the formula is:

[0026]

[0027] Wherein, and z n-1 represent the state at the previous time and the state at the current time respectively;

[0028] Covariance matrix prediction:

[0029]

[0030] Wherein, Q = 10 -5 ; and P n-1 represent the prior covariance and the posterior covariance respectively; T represents the transpose;

[0031] In the update stage, the Kalman filter is based on the current observed value x nand the predicted value update status; at this stage, the Kalman gain K n is calculated using the following formula:

[0032]

[0033] The relationship between the input and output is derived and expressed as:

[0034]

[0035] where is the data after Kalman filtering, represents the predicted value based on the state estimate at the previous moment n - 1 and the system model at moment n, that is, the state estimate before measurement update; x n is the input observation value, and the initial state is The initial covariance is

[0036] Step 3: Divide the preprocessed image data X v and the curvature data (x c1 , x c2 , x c3 ,......, x cn ) into training sets and test sets respectively, and input them into the corresponding feature extraction modules for training to obtain the feature vectors of the image data and the curvature data;

[0037] Furthermore, the specific operation of step 3 is as follows:

[0038] The image feature extraction module consists of YOLOv8 and performs the image feature extraction task; the visual image feature extraction module consists of the following three main parts: Backbone (main network), Neck (neck network), and Head (head network). Among them, the main network is the foundation and is responsible for extracting features from the input image. The head network is the decision-making part and is responsible for generating the final detection results. The neck network is located between the main network and the head network and functions to perform feature fusion and enhancement.

[0039] The curvature feature extraction module consists of a front-end network and a back-end network. The front-end network is based on the random forest algorithm composed of three decision trees and undergoes three classifications in total. Its formulaic expression is:

[0040]

[0041] In the above formula, c is the parameter of the decision tree, h1(x), h2(x), and h3(x) represent the three decision trees respectively. Each decision tree has 4 classification results. The result of the random forest algorithm is input into the back-end network composed of the linear regression algorithm to complete the extraction of the curvature features.

[0042] Step 4: Input the obtained feature vectors into the modality fusion module to fuse two originally uncorrelated feature vectors, and obtain the fused features with attribute correlation;

[0043] Further, the specific operation of Step 4 is as follows:

[0044] Given visual image data X v and curvature data (x c1 , x c2 , x c3 ,......, x cn ), first use the image feature extraction module Y v to extract visual feature information F v , and the formula is as follows:

[0045] F v = Y v (X v )

[0046] Use the curvature feature extraction module Y c to obtain curvature feature information F c :

[0047] F c = Y c (X c1 , X c2 , X c3 ,......, X cn )

[0048] Input the two sets of features into the feature fusion module F v,c to obtain the fused features with attribute correlation.

[0049] Step 5: Connect the obtained fused features to the classifier for classification to realize the evaluation of the grasping state of the soft object;

[0050] Further, the specific operation of Step 5 is as follows:

[0051] Input the fused features into the classifier C to predict the current grasping state g, and the formula is:

[0052]

[0053] g = C(F v,c )

[0054] In the above formula, g ∈ {0, 1, 2, 3}, where 0, 1, 2, and 3 respectively represent four grasping states: sliding, appropriate, excessive, and extreme.

[0055] Compared with the prior art, the present invention has the following advantages:

[0056] First, the present invention uses a signal acquisition device installed on the gripper of the robotic arm to collect curvature information, which has stronger practicability and adaptability;

[0057] Second, the present invention uses a signal processing method based on the Kalman filtering principle to reduce the noise of the collected original signal and optimize the effect of subsequent processing.

[0058] Third, the present invention uses a curvature feature extraction module composed of an image feature extraction module based on the YOLOv8 algorithm and a random forest algorithm to improve the evaluation accuracy of the grasping state. Description of the Drawings

[0059] Figure 1 is the overall flowchart of the present invention;

[0060] Figure 2 is a schematic diagram with the grasping state of the present invention as a label;

[0061] Figure 3 is the image processing flowchart;

[0062] Figure 4 is a schematic diagram of the Kalman filtering process.

[0063] Figure 5 is the structural diagram of the YOLOv8 algorithm.

[0064] Figure 6 is a schematic diagram of the random forest algorithm.

[0065] Figure 7 is the structural diagram of the front-end acquisition device processor. Detailed Description of the Invention

[0066] To understand the present invention in depth, we will describe it comprehensively and meticulously. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples is intended to deepen the comprehensive understanding of the disclosed content of the present invention.

[0067] A method for evaluating the grasping state of a soft object based on vision-curvature fusion, the method comprising the following steps:

[0068] Step 1: Use an image acquisition device and a curvature data acquisition module to collect image data and curvature data of the object to be evaluated during grasping, obtain the grasping state by defining the compression ratio of the object to be evaluated, and use the grasping state as a label;

[0069] Furthermore, the specific operation of Step 1 is:

[0070] Step 1.1: Fix the image acquisition device by installing a support on the wrist of the robotic arm. Drive the image acquisition device to take pictures at different angles by rotating the support. Install a curvature data acquisition module on the gripper. When the gripper makes an opening and closing action, the state of the curvature data acquisition module changes;

[0071] Step 1.2: Use the image acquisition device and the curvature data acquisition module to collect the image data and curvature data of the object to be evaluated when it is being grasped. Obtain the grasping state by defining the compression ratio of the object to be evaluated, and use the grasping state as a label; The labels include sliding, excessive, appropriate, and extreme; Among them:

[0072] Sliding: There is relative sliding between the object and the gripper during the grasping process;

[0073] Appropriate: The compression rate of the object during the grasping process does not exceed 10%;

[0074] Excessive: The compression rate of the object during the grasping process does not exceed 70%;

[0075] Extreme: The compression rate of the object during the grasping process exceeds 70%.

[0076] Step 2: Perform corresponding data preprocessing on the collected image data and curvature data respectively to obtain the required data structure types;

[0077] Furthermore, the specific operation of Step 2 is as follows:

[0078] Step 2.1: Perform two filtering processes on the curvature data. The first filtering is performed in the front-end acquisition device processor (STM32F103C8T6) of the robot, and the mean filtering method is used to perform preliminary smoothing on the original signal. The structure is as Figure 7 shown, and the formula is:

[0079]

[0080] In the formula, y[n] is the output sequence, the window size is 5; i represents the summation index variable, represents the index position within the current window, n represents the index of the current moment, that is, the index for calculating the data after mean filtering;

[0081] Step 2.2: The signal after preliminary smoothing is sent to the computer through the serial bus on the robot, and Kalman filtering is used for secondary filtering to obtain the required ideal curvature data;

[0082] Step 2.3: The resolution of the captured image is 1280×960, and the input size of the image data required by the model is 640×640. To reduce the image size, the image is downsampled and compressed, and the resolution of the compressed image is 640×480. Pixel supplementation operation is performed on the image to obtain an image with a resolution of 640×640.

[0083] The specific process of the Kalman filter is as follows:

[0084] The input value of the Kalman filter is the observation data sequence x n = z n + v n , where x n , z n and v n represent the observed value, the true position, and the observation noise respectively. The observation noise follows a normal distribution v n ~(0, R), where R = 5;

[0085] In the prediction stage, the Kalman filter predicts the current state based on the state at the previous time. The formula is:

[0086]

[0087] In the formula, and z n-1 represent the state at the previous time and the state at the current time respectively;

[0088] Covariance matrix prediction:

[0089]

[0090] In the formula, where Q = 10 -5 ; and P n-1 represent the prior covariance and the posterior covariance respectively; T represents the transpose;

[0091] In the update stage, the Kalman filter updates the state based on the current observed value x n and the predicted value; in this stage, the Kalman gain K n is calculated using the following formula:

[0092]

[0093] The relationship between the input and the output is derived and expressed as:

[0094]

[0095] In the formula, is the data after Kalman filtering, denotes the predicted value at time n based on the state estimate at the previous time n-1 and the system model, that is, the state estimate before measurement update; x n is the input observation value, and the initial state is The initial covariance is

[0096] Step 3: Divide the preprocessed image data X v and the curvature data (x c1 , x c2 , x c3 ,......, x cn ) into training sets and test sets respectively, and input them into the corresponding feature extraction modules for training to obtain the feature vectors of the image data and the curvature data;

[0097] Furthermore, the specific operation of Step 3 is as follows:

[0098] The image feature extraction module consists of YOLOv8 and performs the image feature extraction task; The visual image feature extraction module consists of the following three main parts, Backbone (main network), Neck (neck network) and Head (head network). Among them, the main network is the basis and is responsible for extracting features from the input image. The head network is the decision-making part and is responsible for generating the final detection result. The neck network is located between the main network and the head network, and its role is to perform feature fusion and enhancement. Its structure is as Figure 5 shown.

[0099] The curvature feature extraction module consists of a front-end network and a back-end network. The front-end network is based on the random forest algorithm composed of three decision trees and undergoes three classifications in total. Its formulaic expression is:

[0100]

[0101] In the above formula, c is the parameter of the decision tree, h1(x), h2(x), and h3(x) respectively represent the three decision trees, and each decision tree has 4 classification results. Its structural schematic diagram is as Figure 6 shown. Input the result of the random forest algorithm into the back-end network composed of the linear regression algorithm to complete the extraction of the curvature feature;

[0102] Step 4: Input the obtained feature vectors into the modal fusion module to fuse the two originally uncorrelated feature vectors to obtain the fusion feature with attribute association;

[0103] Furthermore, the specific operation of Step 4 is as follows:

[0104] Given the visual image data X v and the curvature data (x c1 , x c2,x c3 ,......,x cn ), first use the image feature extraction module Y v to extract the visual feature information F v , the formula is as follows:

[0105] F v = Y v (X v )

[0106] Use the curvature feature extraction module Y c to obtain the curvature feature information F c :

[0107] F c = Y c (X c1 , X c2 , X c3 ,......, X cn )

[0108] Input the two sets of features into the feature fusion module F v,c to obtain the fusion feature with attribute association.

[0109] Step 5: Connect the obtained fusion feature to the classifier for classification to realize the evaluation of the soft object grasping state;

[0110] Furthermore, the specific operation of the said Step 5 is:

[0111] Input the fused feature into the classifier C to predict the current grasping state g, and the formula is:

[0112]

[0113] g = C(F v,c )

[0114] In the above formula, g ∈ {0, 1, 2, 3}, where 0, 1, 2, and 3 respectively represent four grasping states: sliding, appropriate, excessive, and extreme.

[0115] Table 1 Scores of several evaluation indexes of the present invention

[0116] Accuracy Precision Recall F1 Score (×100) 99.51% 99.46% 99.51% 99.48

[0117] According to the experimental results in Table 1, the effect of visual curvature fusion perception is very ideal. Theoretically, the visual image provides geometric information about the contact conditions, which helps to better distinguish extreme grasping states. The curvature modality obtains more geometric detail information during the structural deformation of the soft object. The combination of the two modalities enables the present method to show good performance in the evaluation of the soft object grasping state.

[0118] Regarding the specific structure of the present invention, it should be noted that the connection relationships between the various component modules adopted by the present invention are definite and achievable. Except for the special descriptions in the embodiments, the specific connection relationships can bring corresponding technical effects and, on the premise of not relying on the execution of corresponding software programs, solve the technical problems proposed by the present invention. The models of the components, modules, and specific components, the connection methods between them, as well as the conventional usage methods and predictable technical effects brought by the above technical features, except for the specific descriptions, all belong to the publicly disclosed content in patents, journal papers, technical manuals, technical dictionaries, and textbooks that those skilled in the art could obtain before the filing date, or belong to the prior art such as the conventional technologies and common general knowledge in the art, and there is no need to elaborate. This makes the technical solution provided in this case clear, complete, and achievable, and the corresponding physical product can be reproduced or obtained based on this technical means.

[0119] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention have been described above for the convenience of those skilled in the art of the present technology to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. A method for evaluating the clamping state of a soft object based on vision-curvature fusion, characterized in that, The method includes the following steps: Step 1: Use the image acquisition device and the curvature data acquisition module installed on the robot to collect the image data and curvature data of the object to be evaluated when it is being grasped. Obtain the clamping state by defining the compression ratio of the object to be evaluated, and use the clamping state as a label; Step 2: Perform corresponding data preprocessing on the collected image data and curvature data respectively to obtain the required data structure types; Step 3: Divide the preprocessed image data X v and the curvature data (x c1 , x c2 , x c3 ,......, x cn ) into training sets and test sets respectively, and input them into the corresponding feature extraction modules for training to obtain the feature vectors of the image data and the curvature data; Step 4: Input the obtained feature vectors into the modal fusion module to fuse two originally uncorrelated feature vectors to obtain a fusion feature with attribute correlation; Step 5: Connect the obtained fusion feature to a classifier for classification to realize the evaluation of the clamping state of the soft object.

2. The method for evaluating the clamping state of a soft object based on vision-curvature fusion according to claim 1, wherein: The said Step 1: Use the image acquisition device and the curvature data acquisition module to collect the image data and curvature data of the object to be evaluated when it is being grasped. Obtain the clamping state by defining the compression ratio of the object to be evaluated, and use the clamping state as a label. The specific process is as follows: Step 1.1: Fix the image acquisition device by installing a support on the wrist of the robotic arm. Drive the image acquisition device to take pictures at different angles by rotating the support. Install the curvature data acquisition module on the gripper. When the gripper makes an opening and closing movement, the state of the curvature data acquisition module changes; Step 1.2: Use the image acquisition device and the curvature data acquisition module to collect the image data and curvature data of the object to be evaluated when it is being grasped. Obtain the clamping state by defining the compression ratio of the object to be evaluated, and use the clamping state as a label; The labels include sliding, excessive, appropriate, and extreme; where: Sliding: There is relative sliding between the object and the gripper during the grasping process; Appropriate: The compression rate of the object during the grasping process does not exceed 10%; Excessive: The compression rate of the object during the grasping process does not exceed 70%; Extreme: The compression rate of the object during the grasping process exceeds 70%.

3. The method for evaluating the clamping state of a soft object based on vision-curvature fusion according to claim 2, wherein: The said Step 2: Perform corresponding data preprocessing on the collected image data and curvature data respectively to obtain the required data structure types. The specific process is as follows: Step 2.1: Perform two filtering processes on the curvature data. In the first filtering, use the mean filtering method to perform preliminary smoothing on the original signal in the processor of the front-end acquisition device of the robot. The formula is: In the formula, y[n] is the output sequence, the window size is 5; i represents the summation index variable, represents the index position within the current window, n represents the index at the current moment, that is, the index for calculating the data after mean filtering; Step 2.2: Send the signal after preliminary smoothing to the computer through the serial bus on the robot, and use Kalman filtering for secondary filtering to obtain the required curvature data; Step 2.3: Downsample and compress the image. The resolution of the compressed image is 640×480, and perform a pixel filling operation on the image to obtain an image with a resolution of 640×640.

4. A method for evaluating the clamping state of a soft object based on vision-curvature fusion according to claim 3, characterized in that: The specific process of Kalman filtering in Step 2.2 is as follows: The input value of the Kalman filter is the observed data sequence x n = z n + v n , where x n , z n and v n represent the observed value, the true position, and the observation noise respectively. The observation noise follows a normal distribution v n ~(0, R), where R = 5; In the prediction stage, the Kalman filter predicts the current state based on the state at the previous time. The formula is: wherein, and z n-1 represent the state at the previous time and the state at the current time, respectively; Covariance matrix prediction: In the formula, where Q = 10 -5 ; and P n-1 represent the prior covariance and the posterior covariance respectively; T represents the transpose; In the update phase, the Kalman filter updates the state based on the current observation x n and the predicted value; in this phase, the Kalman gain K n is calculated using the following formula: The relationship between input and output is derived and expressed as: Wherein, is the data after Kalman filtering, represents the state estimation and prediction value at time n based on the state at the previous time n - 1, that is, the state estimation before measurement update, x n is the input observation value, and the initial state is The initial covariance is 5. A method for evaluating the clamping state of a soft object based on vision-curvature fusion according to claim 4, characterized in that: Step 3: The preprocessed image data X v and the curvature data (x c1 , x c2 , x c3 ,......, x cn ) are divided into a training set and a test set and input into the corresponding feature extraction module for training to obtain the feature vectors of the image data and the curvature data. The specific process is as follows: The image feature extraction module consists of YOLOv8 and performs the task of image feature extraction. The visual image feature extraction module consists of the following three main parts: Backbone, Neck, and Head. Among them, the backbone network is the foundation, responsible for extracting features from the input image. The head network is the decision-making part, responsible for generating the final detection results. The neck network is located between the backbone network and the head network and functions to perform feature fusion and enhancement. The curvature feature extraction module consists of a front-end network and a back-end network. The front-end network is based on the random forest algorithm composed of three decision trees and undergoes three classifications in total. Its formulaic expression is: In the above formula, c is the parameter of the decision tree, h1(x), h2(x), and h3(x) represent the three decision trees respectively. Each decision tree has 4 classification results. The result of the random forest algorithm is input into the back-end network composed of the linear regression algorithm to complete the extraction of the curvature feature. The formulaic expression of the back-end network composed of the linear regression algorithm is: In the above formula, ω1 = ω2 = ω3 = ω4 = 1 are the weight coefficients.

6. A method for evaluating the clamping state of a soft object based on vision-curvature fusion according to claim 5, characterized in that: Step 4: Input the obtained feature vector into the modality fusion module to fuse two originally uncorrelated modalities to obtain the fusion feature with attribute association. The specific process is as follows: Given visual image data X v and curvature data (x c1 , x c2 , x c3 ,......, x cn ), first use the image feature extraction module Y v to extract visual feature information F v , the formula is as follows: F v = Y v (X v ) Use the curvature feature extraction module Y c to obtain curvature feature information F c : F c = Y c (X c1 , X c2 , X c3 ,......, X cn ) Input two sets of features into the feature fusion module F v,c to obtain the fused features with attribute associations.

7. The method for evaluating the clamping state of a soft object based on vision-curvature fusion according to claim 6, characterized in that: Step 5: Connect the obtained fusion feature to the classifier for classification to realize the evaluation of the soft object grasping state. The specific process is as follows: Input the fused feature into classifier C to predict the current grasping state g. The formula is: g = C(F v,c ) In the above formula, g ∈ {0, 1, 2, 3}, where 0, 1, 2, and 3 represent four grasping states: sliding, appropriate, excessive, and extreme.

Citation Information

Patent Citations

  • Mechanical arm autonomous grabbing method based on visual touch fusion under weak rigidity characteristic condition

    CN115431279A

  • Grasped object classification method based on visual touch fusion

    CN117611919A

  • Low-voltage distribution network operation state detection method based on fusion terminal

    CN119150210A

  • Gesture pose estimation method based on kalman filtering and deep learning

    WO2024094227A1