Manipulator high-precision pressure estimation method based on visual perception

Through a high-precision pressure estimation method based on visual perception, the dual-channel feature enhancement module and deep learning network are used to solve the reliability and cost of pressure perception of robots in complex environments, and high-precision pressure estimation and safe interaction are achieved.

CN120259212AActive Publication Date: 2025-07-04JIANGTAI INTELLIGENT TECHNOLOGY (SUZHOU) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510316946.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing robot pressure sensing technology is insufficient in complex environments. Traditional sensors have problems with blind spots, friction and wear and high cost. Vision-based pressure estimation methods have deteriorated performance under light and surface texture changes.

Method used

A high-precision pressure estimation method based on visual perception is adopted to build a deep learning network through a dual-channel feature enhancement module, and a RGB image is used to extract and map pressure features, combining the encoder-decoder architecture and feature pyramid network to achieve the generation of pressure distribution maps.

Benefits of technology

It realizes high-precision pressure estimation of robots in complex environments, avoids the physical limitations of traditional contact sensors, improves the practicality and reliability of the system, and reduces the complexity and cost of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259212A_ABST
    Figure CN120259212A_ABST
Patent Text Reader

Abstract

The invention discloses a manipulator high-precision pressure estimation method based on visual perception, and belongs to the field of visual perception, and the method comprises the following steps: obtaining an RGB image, and preprocessing the RGB image to obtain a standardized image; constructing a deep learning network based on a dual-channel feature enhancement module, and inputting the standardized image into the deep learning network to extract pressure features; and mapping the pressure characteristics into a pressure distribution diagram to obtain a pressure estimation value. According to the method, high-precision pressure estimation of the manipulator based on visual perception is realized, and key support is provided for precise operation and safe interaction of the manipulator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual perception, and particularly relates to a high-precision pressure estimation method for a manipulator based on visual perception. Background Art

[0002] The international academic community has shown multi-dimensional innovative explorations in the field of manipulator contact surface pressure perception, mainly making significant progress in three major directions: high-performance tactile sensing skin, intelligent pressure distribution analysis algorithms, and bionic tactile systems. In the field of high-performance tactile sensing skin, research teams such as Ntagios have pioneered the combination of 3D printing technology and soft material sensors, successfully developing a manipulator system with intrinsic tactile perception capabilities. This research has broken through the limitations of traditional tactile sensors in terms of structural complexity and integration efficiency, providing a paradigm-shifting technical path for the manufacturing of a new generation of intelligent mechanical structures. The breakthrough in the research field of pressure distribution analysis algorithms is reflected in the innovative classification and segmentation method proposed by Albini and Cannata. A dedicated pressure mapping algorithm is designed for the contact force pattern in the human-robot interaction scenario. This method converts the complex contact pressure distribution into an interpretable interaction pattern through an advanced computational model, significantly enhancing the depth and accuracy of the robot system's understanding of the human-robot contact state. This progress not only improves the precision of human-robot interaction but also lays an algorithmic foundation for the development of safe collaborative robots. In the research on bionic tactile systems, Roberts et al. have constructed a theoretical framework for the robot's soft tactile perception skin through systematic summarization and analysis. Their research deeply explores various technical paths and material selection strategies for simulating biological tactile systems, and comparatively analyzes the performance characteristics and application limitations of different bionic methods, providing comprehensive theoretical guidance and design ideas for the development of a new type of manipulator tactile system with high sensitivity and versatility.

[0003] The research in the field of contact pressure perception of robotic manipulators in the prior art also shows a diversified development trend, mainly focusing on three directions: the development of new sensing materials, intelligent algorithm-assisted perception, and multi-modal system integration. In the research of flexible sensors, Zhang et al. developed a multi-functional sensing unit integrating pressure measurement and ultrasonic detection functions. By adopting an innovative composite material structure, this sensor achieved high-precision contact force detection, providing a new perception dimension for robotic grasping operations. Such new flexible sensors not only improved the perception accuracy of robotic manipulators but also enhanced their adaptability on irregular surfaces. Significant progress has also been made in the field of artificial intelligence-assisted perception. Yang et al. combined machine learning algorithms with pressure perception mechanisms to develop an intelligent fingertip tactile system for humanoid robots. This system can accurately identify the contact pressure distribution on complex surfaces and classify contact types and object materials based on pressure characteristics. Through feature extraction and pattern recognition using deep neural networks, the research team effectively solved the non-linear problems in contact force analysis and improved the intelligent level of the interaction between robotic manipulators and the environment. In terms of the integration of multi-functional sensing systems, Chen et al. designed a flexible dual-mode sensor, which organically combined the functions of precise contact pressure measurement and non-contact distance detection. Using a multi-layer composite structure and micro-nano manufacturing processes, this sensor achieved the perception conversion between contact and non-contact states. This multi-modal fusion perception method provides more comprehensive perception information for the fine operations of robots in complex environments and expands the adaptability of robotic manipulators in uncertain environments.

[0004] Despite the significant progress made in the prior art, the current solutions for the pressure perception research of robotic manipulators still have obvious limitations. Traditional contact sensors face many engineering challenges: flexible sensors are difficult to perfectly fit complex curved surfaces and often produce perception blind spots; rigid sensors, on the other hand, will interfere with fine operations. More critically, the reliability of these systems in actual application environments is insufficient, and friction, wear, pollution, and environmental changes will all cause a significant decline in performance. The complex wiring and signal processing of sensing units not only increase the volume and weight of the system but also raise the manufacturing and maintenance costs.

[0005] Vision-based pressure estimation faces reliability challenges in practical applications. Due to factors such as posture, illumination, occlusion, and surface texture, the appearance of soft grippers under the same pressure varies significantly. The performance of data-driven methods highly depends on the coverage of training data. Models trained in controlled environments (fixed perspective, smooth surface) are difficult to generalize to complex scenarios (textured surface, multi-angle, occlusion, etc.). The direct installation of sensors will affect surface characteristics and contact mechanical characteristics. Although physical simulation provides ideas for solving this problem, accurately simulating gripper deformation and visual phenomena is still challenging. To address the above challenges, the present invention breaks through the limitations of traditional sensing methods and provides a high-precision pressure estimation method for robotic manipulators based on visual perception. Summary of the Invention

[0006] To solve the above technical problems, the present invention proposes a high-precision pressure estimation method for a manipulator based on visual perception to solve the problems existing in the above prior art.

[0007] To achieve the above object, the present invention provides a high-precision pressure estimation method for a manipulator based on visual perception, including:

[0008] Obtain an RGB image, and preprocess the RGB image to obtain a normalized image;

[0009] Construct a deep learning network based on a dual-channel feature enhancement module, and input the normalized image into the deep learning network to extract pressure features;

[0010] Map the pressure features to a pressure distribution map to obtain a pressure estimation value.

[0011] Optionally, the deep learning network constructed based on the dual-channel feature enhancement module includes: an encoder, a feature pyramid network, and a decoder;

[0012] Among them, the process of extracting pressure features based on the deep learning network includes:

[0013] Input the normalized image into the encoder, and the encoder processes the normalized image based on a plurality of cascaded dual-channel feature enhancement modules to obtain multi-scale features;

[0014] The multi-scale features are subjected to cross-scale feature fusion through the feature pyramid network to obtain fused features;

[0015] Obtain pressure features based on the fused features;

[0016] The pressure features are subjected to upsampling and convolution operations of the decoder to obtain predicted pressure categories.

[0017] Optionally, the dual-channel feature enhancement module includes a spatial feature modulation module and a VP-Mamba module;

[0018] Among them, the calculation expression for processing the normalized image based on the dual-channel feature enhancement module to obtain multi-scale features is:

[0019] F out =α·F SFM +(1 - α)·F Mamba

[0020] In the formula, F SFM is the output of the SFM module, F Mamba is the output of the VP-Mamba module, and α is an adaptive fusion weight.

[0021] Optionally, the process of calculating the output of the SFM module includes:

[0022] Obtain input features based on the input normalized image;

[0023] Group the input features to obtain a number of feature groups;

[0024] Apply a multi-scale feature generation unit to each feature group to obtain multi-scale features;

[0025] Perform feature aggregation and attention generation on the multi-scale features to obtain a multi-scale feature map;

[0026] Perform feature modulation on the multi-scale feature map and the input features to obtain the output of the SFM module.

[0027] Optionally, the process of calculating the output of the VP-Mamba module includes:

[0028] Divide the input features into blocks and serialize them to obtain a number of feature blocks;

[0029] Use SSM to perform state space processing on a number of the feature blocks to obtain a state output;

[0030] Scan the state output to obtain a scan output;

[0031] Perform feature recombination on the scan output to obtain the output of the VP-Mamba module.

[0032] Optionally, the calculation expression for the multi-scale features to be cross-scale feature fused by the feature pyramid network to obtain the fused features is:

[0033] M n =Conv 1×1 (F n )

[0034] M i =Conv 1×1 (F i ) + Upsample(M i+1 ), i ∈ n - 1, n - 2,..., 1

[0035] In the formula, Conv 1×1 represents a 1×1 convolution operation, Upsample represents an upsampling operation, M i is the fused feature of the i-th layer, and F i represents the feature representation of the i-th layer.

[0036] Optionally, the expression for the pressure estimate value is:

[0037]

[0038] In the formula, N = 9 represents the number of pressure intervals, and P pred (x, y) is the predicted pressure value of the pixel (x, y), O(x, y, i) represents the probability that the pixel belongs to the i-th pressure interval, and P i is the representative pressure value of the i-th interval.

[0039] Optionally, after constructing the deep learning network, a network training process is further included, wherein a composite loss function is used to guide the training of the deep learning network;

[0040] The expression of the composite loss function is: L total = L CE + λ1·L SSIM + λ2·L smooth

[0041] In the formula, L total represents the composite loss value, L CE represents the classification cross-entropy loss, L SSIM represents the structural similarity loss, λ1 represents the weight of the structural similarity loss, and L smooth represents the smoothing loss, and λ2 represents the weight of the smoothing loss.

[0042] Compared with the prior art, the present invention has the following advantages and technical effects:

[0043] The present invention realizes high-precision pressure estimation of the manipulator based on visual perception, provides key support for the precise operation and safe interaction of the manipulator, and promotes the development of the manipulator pressure perception technology towards a more practical, reliable and intelligent direction. Especially in application scenarios with extremely high requirements for hygiene and safety such as medical surgery, precision assembly, and food processing, the non-contact pressure perception method provided by the present invention has significant advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0045] Figure 1 is the overall architecture diagram of the high-precision pressure estimation model based on visual perception according to the embodiment of the present invention;

[0046] Figure 2 is the DCFE module diagram according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The following will describe the present application in detail with reference to the drawings and in combination with the embodiments.

[0048] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0049] Embodiment 1

[0050] The present invention proposes a high-precision visual pressure estimation method, which realizes high-precision estimation of contact pressure through a single RGB image. This method is based on the robotic hand contact image segmentation technology, avoiding the process of pressure estimation pixel by pixel. By introducing a Dual-Channel Feature Enhancement (DCFE) module, effective extraction of the detailed features and regional relevance of the robotic hand is achieved, so as to be able to effectively estimate the contact pressure of the robotic hand in a complex environment.

[0051] As Figure 1 shown, it is a deep learning framework for robotic gripper pressure estimation. Accurate estimation of contact pressure is realized through a single RGB image. The network is designed based on an encoder-decoder architecture, and spatial information is retained through skip connections. The encoder integrates multiple DCFE modules pre-trained on ImageNet and has strong feature extraction capabilities. The decoder uses a Feature Pyramid Network (FPN) to achieve multi-scale feature fusion, enhancing the recognition ability for pressure distributions at different scales. The model outputs a pressure map with the same dimension as the input image, realizing pixel-level contact position and pressure estimation.

[0052] The model proposed by the present invention transforms pressure estimation into a classification problem, and uses 8 logarithmically equally spaced intervals and 1 zero-pressure interval for pressure value discretization. The model classifies each pixel of the input image into a predefined pressure interval.

[0053] This model uses synchronously collected RGB images and pressure sensing data for training and verification. Through homography transformation, the high-resolution pressure sensor data is mapped into the RGB image space, realizing the precise alignment of pressure data and images, which is convenient for quantitative comparison between the model prediction results and the measured values.

[0054] This model includes: an encoder: using a DCFE module to extract multi-scale features; a Feature Pyramid Network (FPN): realizing cross-scale feature fusion; a decoder: generating a pressure map with the same dimension as the input image.

[0055] Among them

[0056] The present invention provides a high-precision pressure estimation method for a manipulator based on visual perception, which can achieve high-precision estimation of contact pressure through a single RGB image, avoiding various challenges faced by traditional contact sensors in applications.

[0057] In this embodiment, a high-precision pressure estimation method for a manipulator based on visual perception is provided, including the following steps:

[0058] S1. Data acquisition and preprocessing: Obtain images of the manipulator and the contact surface through an external RGB camera, and perform standardization processing on the images.

[0059] As a specific implementation manner of this embodiment, the process of data acquisition and preprocessing includes:

[0060] S1.1 Image acquisition: The present invention uses a standard RGB camera to collect interaction images of the manipulator and the contact surface from an external perspective, and the image resolution is 1920×1080 pixels. The images collected by the camera contain deformation information of the manipulator, shadow changes, and characteristics of the contact area.

[0061] S1.2 Training data construction:

[0062] (1) Synchronously collect RGB images of the manipulator operation and high-precision pressure sensor data;

[0063] (2) Map the pressure sensor data to the RGB image space through a homography transformation:

[0064] P image = H·P sensor

[0065] where P image is the pressure value in the image space, P sensor is the sensor measurement value, and H is the homography transformation matrix;

[0066] (3) Pressure value discretization design: Divide the continuous pressure value into 8 logarithmically equally spaced intervals and 1 zero-pressure interval to form a classification problem. The specific interval division formula is:

[0067]

[0068] where P min and P max respectively represent the minimum and maximum pressure values, N is the total number of pressure intervals (in the present invention, N = 9), and P threshold is the threshold for determining contact.

[0069] S2. Pressure Feature Extraction: The designed deep learning network architecture based on the DCFE module extracts the deformation and contact features of the manipulator from RGB images. The process of extracting pressure features based on the deep learning network includes: inputting the normalized image into the encoder, and the encoder processes the normalized image based on several cascaded dual-channel feature enhancement modules to obtain multi-scale features; the multi-scale features are subjected to cross-scale feature fusion through the feature pyramid network to obtain the fused features; the pressure features are obtained based on the fused features; and the pressure features are subjected to upsampling and convolution operations of the decoder to obtain the predicted pressure categories.

[0070] The network architecture designed in this embodiment is as Figure 1 shown. Based on the encoder-decoder structure, spatial information is retained through skip connections. The encoder integrates multiple DCFE modules pre-trained on ImageNet and has strong feature extraction capabilities. The decoder uses the feature pyramid network (FPN) to achieve multi-scale feature fusion, enhancing the recognition ability for pressure distributions at different scales. The model outputs a pressure map with the same dimension as the input image, realizing pixel-level contact position and pressure estimation.

[0071] The model includes: Encoder: Uses the DCFE module to extract multi-scale features; Feature Pyramid Network (FPN): Achieves cross-scale feature fusion; Decoder: Generates a pressure map with the same dimension as the input image.

[0072] S2.1 Encoder

[0073] The encoder is composed of multiple cascaded DCFE modules, and the input is the RGB image The output is the multi-scale feature representation F = F1, F2,..., F n , where H i 、W i are the spatial dimensions of the i-th layer feature map, and C i is the number of channels.

[0074] The encoding process can be expressed as:

[0075] F i = DCFE i (F i-1 ), i ∈ 1, 2,..., n

[0076] where F0 = I is the input image, and DCFE i is the i-th dual-channel feature enhancement module.

[0077] As Figure 2 shown, the dual-channel feature enhancement (DCFE) module: includes two sub-modules, spatial feature modulation (SFM) and VP-Mamba, which are used to enhance the model's perception ability of details and the modeling of long-range dependencies.

[0078] Among them, the SFM module realizes multi-scale feature extraction through independent calculation and dynamic aggregation, enhancing the model's ability to perceive details. The SFM module can dynamically adjust the feature processing strategy according to the image content, improving the model's adaptability to diverse low-resolution images. The VP-Mamba module enhances the global feature modeling ability by establishing long-range dependencies, which is of great significance for improving the performance of tasks such as image classification, segmentation, and object detection.

[0079] The overall calculation process of the Dual-channel Feature Enhancement (DCFE) module is as follows:

[0080] F out = α·F SFM +(1 - α)·F Mamba

[0081] Among them, F SFM is the output of the SFM sub-module, F Mmba is the output of the VP-Mamba sub-module, and α is the adaptive fusion weight.

[0082] S2.1.1 Spatial Feature Modulation (SFM) Module

[0083] The SFM module realizes multi-scale feature extraction through independent calculation and dynamic aggregation, and is dedicated to the image segmentation task. The SFM first divides the input features into four groups, which are processed by the Multi-scale Feature Generation Unit (MFGU). The MFGU uses 3×3 depth convolution and self-adjusting max pooling to extract local and global features. After the extracted features are concatenated in the channel dimension and aggregated by 1×1 convolution, an attention map is generated through the GELU activation function. Finally, the SFM adaptively modulates the input features based on the attention map. The SFM module has three core features: multi-scale feature representation, adaptive feature modulation, and non-local feature interaction. The module realizes multi-scale information processing through adaptive pooling and the GELU activation function, can capture local details and global semantics simultaneously, and dynamically adapts to different image features. Compared with the traditional attention mechanism, the SFM uses 1×1 convolution to realize multi-scale feature aggregation, significantly improving the feature representation ability.

[0084] The specific calculation steps of the SFM module include: obtaining input features based on the input normalized image; performing feature grouping on the input features to obtain several feature groups; applying the multi-scale feature generation unit to each feature group to obtain multi-scale features; performing feature aggregation and attention generation on the multi-scale features to obtain a multi-scale feature map; and performing feature modulation on the multi-scale feature map and the input features to obtain the output of the SFM module.

[0085] (1) Feature grouping: Divide the input features into four groups:

[0086] F1,F2,F3,F4 = Split(F in )

[0087] where B is the batch size, C is the number of channels, and H and W are the spatial dimensions.

[0088] (2) Multi-scale feature generation: Apply a multi-scale feature generation unit (MFGU) to each feature group:

[0089] F i ' = MFGU(F i ), i ∈ 1, 2, 3, 4

[0090] The specific operations of MFGU include:

[0091] 3×3 depth convolution to extract local features: F local = DWConv 3×3 (F i )

[0092] Adaptive max pooling to extract global features: F global = AdaptiveMaxPool(F i )

[0093] Feature combination: F i ' = Concat(F local , F global )

[0094] (3) Feature aggregation and attention generation to obtain a multi-scale feature map:

[0095] F cat = Concat(F1', F2', F3', F4')

[0096] F agg = Conv 1×1 (F cat )

[0097] A map = GELU(F agg )

[0098] where F cat is a multi-scale feature formed by concatenating four feature groups processed by MFGU in the channel dimension, F agg is the aggregated feature obtained by applying a 1×1 convolution operation to F cat , and A map is the attention map generated by applying the GELU activation function to F agg . GELU is the Gaussian error linear unit activation function:

[0099]

[0100] (4) Feature modulation:

[0101] F SFM = F in ⊙A map

[0102] where ⊙ represents element-wise multiplication to achieve adaptive modulation of features.

[0103] S2.1.2 VP-Mamba module

[0104] VP-Mamba is designed using the state space model (SSM) architecture to effectively capture the long-range dependent features of images through SSM. The module consists of an input layer and an SSM layer: the input layer is responsible for image chunking, and the SSM layer processes the chunked data and establishes inter-region dependencies. The module also integrates an adaptive selection mechanism, which, through dynamic filtering and output layer processing, achieves efficient feature extraction and classification. The main advantage of VP-Mamba lies in its long-range dependent modeling ability, which performs excellently in image classification and understanding tasks. Based on the dynamic characteristics and selection mechanism of SSM, VP-Mamba can adaptively process different image contents and effectively extract classification-related features. This design achieves excellent performance while maintaining a low number of parameters and retains the ability to perceive details in high-resolution image processing, meeting the requirements of high-precision application scenarios.

[0105] VP-Mamba introduces a selection mechanism for dynamically filtering information and generating classification results. This selection mechanism reduces the computational complexity while focusing the model on key features by dynamically processing and filtering spatial information. It extends the one-dimensional sequence model to two-dimensional image processing, achieving efficient processing and stable performance for inputs of different complexities. The specific steps are as follows:

[0106] First, perform a scan expansion operation to divide the input image into multiple blocks, and then scan these blocks in sequence so that the model can gradually obtain the global information of the image.

[0107] Second, perform a selective scan on each token sequence to obtain a context token sequence, and use SSM to implement long-range dependent modeling to capture the spatial correlation of image features. Finally, fuse the local and global information and output the fused features.

[0108] The VP-Mamba module is designed based on the state space model (SSM) to capture long-range dependencies in images. The specific process includes: The process of calculating the output of the VP-Mamba module includes: dividing the input features into blocks and serializing them to obtain a number of feature blocks; performing state space processing on the number of feature blocks using SSM to obtain a state output; scanning the state output to obtain a scan output; and performing feature recombination on the scan output to obtain the output of the VP-Mamba module.

[0109] The processing flow is as follows:

[0110] (1) Image chunking and serialization to obtain a number of feature blocks:

[0111] X tokens = Tokenize(F in )

[0112] where the Tokenize operation divides the feature map F in into blocks of size n×n and arranges them in scan order;

[0113] (2) State space processing to obtain a state output: Applying SSM to obtain long-range dependencies:

[0114] X ssm = SSM(X tokens )

[0115] X ssm represents the state output, and X tokens represents the feature block.

[0116] The core calculation process of SSM is:

[0117] h t = A·h t-1 + B·x t

[0118] y t = C·h t

[0119] where, is the hidden state vector, x t is the input sequence, y t is the output sequence, is the state transition matrix, and are the input mapping and output mapping matrices respectively.

[0120] (3) Selective scanning to obtain a scan output:

[0121] X scanned = SelectiveScan(X ssm )

[0122] Selective scanning processes the sequence from four directions, capturing omnidirectional spatial dependencies:

[0123] X out = ScanH(X ssm ) + ScanV(X ssm ) + ScanHR(X ssm ) + ScanVR(X ssm )

[0124] where ScanH, ScanV, ScanHR, and ScanVR represent horizontal, vertical, horizontal reverse, and vertical reverse scans respectively; X out represents the scan output.

[0125] (4) Feature recombination:

[0126] F MAmba = Reshape(X out , [B, C, H, W])

[0127] Recombines the processed sequence into the shape of the original feature map.

[0128] S2.2 Feature Pyramid Network (FPN)

[0129] The FPN module implements cross-scale feature fusion, and the specific calculation process is as follows:

[0130] M n = Conv 1×1 (F n )

[0131] M i = Conv 1×1 (F i ) + Upsample(M i+1 ), i ∈ n - 1, n - 2,..., 1

[0132] where Conv 1×1 represents the 1×1 convolution operation, Upsample represents the upsampling operation, and M i is the feature after fusion at the i-th layer.

[0133] S2.3 Decoder

[0134] The decoder restores the feature map to the original image size through upsampling and convolution operations and outputs the predicted pressure category:

[0135] O = Softmax(Conv 3×3 (P1))

[0136] where, Represents the probability distribution of each pixel belonging to N pressure intervals, where N = 9 represents the number of pressure intervals. The final pressure value can be obtained by weighted summation:

[0137]

[0138] Among them, P pred (x, y) is the predicted pressure value of pixel (x, y), O(x, y, i) represents the probability that this pixel belongs to the i-th pressure interval, and P i is the representative pressure value of the i-th interval.

[0139] S2.4 Model Training and Optimization

[0140] S2.4.1 Loss Function Design

[0141] The present invention uses a composite loss function to guide network training:

[0142] L total = L CE + λ1·L SSIM + λ2·L smooth

[0143] Categorical cross-entropy loss:

[0144] Among them, y x,y,c is the true label of pixel (x, y) in category c, and p x,y,c is the predicted probability;

[0145] Structural similarity loss: L SSIM = 1 - SSIM(P pred , P gt )

[0146] Among them, SSIM represents the structural similarity index, which is used to ensure the spatial coherence of the predicted pressure map;

[0147] Smoothing loss:

[0148]

[0149] Among them, and represent the gradient operators in the horizontal and vertical directions. This loss prompts the pressure map to remain sharp at the image edges and smooth in the smooth regions.

[0150] S2.4.2 Training Strategy

[0151] The AdamW optimizer is adopted, and the initial learning rate is set to 1×10 -4 ;

[0152] Cosine annealing learning rate scheduling with a minimum learning rate of 1×10 -6 , and the calculation formula is:

[0153]

[0154] where η t is the learning rate for the t-th round, and T is the total number of training rounds;

[0155] The batch size is set to 8, and the number of training rounds is 100;

[0156] During the training process, data augmentation techniques are applied, including random cropping, rotation, horizontal flipping, and brightness variation.

[0157] S3. Pressure value estimation: Map the extracted features to a pressure distribution map to achieve pixel-level pressure estimation;

[0158] S4. Result output and application: Generate a pressure map with the same dimension as the input image for subsequent precise control of the manipulator.

[0159] As a specific implementation manner of this embodiment, the online prediction process includes:

[0160] Obtain a single RGB image I from the camera; preprocess the image, including size adjustment and pixel value normalization; input the processed image into the trained model to obtain the pressure distribution map P pred ; Generate a visual pressure distribution map based on the prediction result and overlay it on the original image.

[0161] The present invention can be applied to the precise operation control of the manipulator, and the specific implementation method is as follows:

[0162] (1) Based on the visual servo principle, establish a closed-loop system for pressure estimation and manipulator position control;

[0163] (2) The system control formula is:

[0164]

[0165] where u t is the control command, P target is the target pressure distribution, P pred is the current estimated pressure, K p and K d are the proportional and differential control coefficients respectively;

[0166] (3) Based on the above control strategy, the manipulator can achieve: precisely controlling the magnitude and distribution of the contact pressure; maintaining a stable pressure when sliding along the surface; adapting to objects with different stiffnesses and shapes.

[0167] The present invention has the following technical advantages:

[0168] Contactless sensing: Avoids the physical limitations and reliability issues of traditional contact sensors in applications;

[0169] High-precision estimation: Through the dual-channel design of the DCFE module, fine-grained pressure estimation is achieved, with the mean absolute error lower than traditional methods;

[0170] Robustness: The model maintains stable performance under different lighting, viewing angles, and object material conditions;

[0171] Real-time performance: The optimized network structure can achieve a real-time inference speed of 30fps on a standard GPU;

[0172] Easy integration: Only requires an external RGB camera, without the need to install additional sensors on the robotic arm, significantly reducing system complexity and cost.

[0173] Through the above technical solutions, the present invention realizes high-precision pressure estimation of a robotic arm based on visual perception, provides key support for the precise operation and safe interaction of the robotic arm, and promotes the development of robotic arm pressure sensing technology towards a more practical, reliable, and intelligent direction. Especially in application scenarios with extremely high requirements for hygiene and safety, such as medical surgeries, precision assembly, and food processing, the contactless pressure sensing method provided by the present invention has significant advantages.

[0174] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A high-precision pressure estimation method for a manipulator based on visual perception, characterized in that It includes the following steps: Obtain an RGB image, and preprocess the RGB image to obtain a normalized image; Construct a deep learning network based on a dual-channel feature enhancement module, and input the normalized image into the deep learning network to extract pressure features; Map the pressure features to a pressure distribution map to obtain a pressure estimation value.

2. The high-precision pressure estimation method for a manipulator based on visual perception according to claim 1, wherein The deep learning network constructed based on the dual-channel feature enhancement module includes: an encoder, a feature pyramid network, and a decoder; Among them, the process of extracting pressure features based on the deep learning network includes: Input the normalized image into the encoder, and the encoder processes the normalized image based on a plurality of cascaded dual-channel feature enhancement modules to obtain multi-scale features; The multi-scale features are subjected to cross-scale feature fusion through the feature pyramid network to obtain fused features; Obtain pressure features based on the fused features; The pressure features are subjected to upsampling and convolution operations of the decoder to obtain predicted pressure categories.

3. The high-precision pressure estimation method for a manipulator based on visual perception according to claim 2, characterized in that The dual-channel feature enhancement module includes a spatial feature modulation module and a VP-Mamba module; Among them, the calculation expression for processing the normalized image based on the dual-channel feature enhancement module to obtain multi-scale features is: F out = α·F SFM +(1 - α)·F Mamba Where F SFM is the output of the SFM module, F Mamba is the output of the VP-Mamba module, and α is the adaptive fusion weight.

4. The method for high-precision pressure estimation of a manipulator based on visual perception according to claim 3, wherein The process of calculating the output of the SFM module includes: Obtain input features based on the input normalized image; Perform feature grouping on the input features to obtain a plurality of feature groups; Apply a multi-scale feature generation unit to each feature group to obtain multi-scale features; Perform feature aggregation and attention generation on the multi-scale features to obtain a multi-scale feature map; Perform feature modulation on the multi-scale feature map and the input features to obtain the output of the SFM module.

5. The high-precision pressure estimation method for a manipulator based on visual perception according to claim 4, wherein The process of calculating the output of the VP-Mamba module includes: Divide the input features into blocks and serialize them to obtain a plurality of feature blocks; Perform state space processing on the plurality of feature blocks using SSM to obtain a state output; Perform scanning on the state output to obtain a scanning output; Perform feature recombination on the scanning output to obtain the output of the VP-Mamba module.

6. The high-precision pressure estimation method for a manipulator based on visual perception according to claim 2, characterized in that The calculation expression for the multi-scale features to be subjected to cross-scale feature fusion through the feature pyramid network to obtain fused features is: M n = Conv 1×1 (F n ) M i = Conv 1×1 (F i ) + Upsample(M i+1 ), i ∈ n - 1, n - 2,..., 1 where Conv 1×1 represents a 1×1 convolution operation, Upsample represents an upsampling operation, and M i is the feature after fusion at the i-th layer, and F i represents the feature representation at the i-th layer.

7. The high-precision pressure estimation method for a manipulator based on visual perception according to claim 1, wherein The expression for the pressure estimation value is: where N = 9 represents the number of pressure intervals, P pred (x, y) is the predicted pressure value of the pixel (x, y), and O(x, y, i) represents the probability that the pixel belongs to the i-th pressure interval, P i is the representative pressure value of the i-th interval.

8. The high-precision pressure estimation method for a manipulator based on visual perception according to claim 1, wherein After constructing the deep learning network, it further includes a network training process, in which a composite loss function is used to guide the training of the deep learning network; The expression of the composite loss function is: L total = L CE + λ1·L SSIM + λ2·L smooth where L total represents the composite loss value, L CE represents the categorical cross-entropy loss, L SSIM represents the structural similarity loss, λ1 represents the structural similarity loss weight, L smooth represents the smoothing loss, λ2 represents the smoothing loss weight.

Citation Information

Patent Citations

  • Method for predicting grip strength of hand grabbing object based on single RGB image

    CN111897436A

  • Deep learning-based violent behavior recognition method, storage device and server

    CN112906516A

  • Foot pressure detection algorithm based on multi-information image

    CN117017268A

  • Method for building light-weight monocular endoscope depth estimation model

    CN118470081A

  • Monocular depth estimation method based on CNN-Transform hybrid architecture

    CN119478000A