End-to-end visual-haptic perception method and system based on topography-force field analysis model, terminal and storage medium
By combining shape reconstruction and force field analysis models into an end-to-end visual-tactile perception method, the problems of insufficient spatial resolution, mechanical measurement range and real-time performance of visual-tactile sensing in existing technologies are solved. It achieves high-precision simultaneous estimation of shape and multi-axis force, and is suitable for scenarios such as precision robot operation, flexible manufacturing inspection, medical rehabilitation and virtual interaction.
Patent Information
- Application Number
- CN202511715651.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-11-21
AI Technical Summary
Existing visual-tactile sensing methods have shortcomings in spatial resolution, mechanical measurement range and accuracy, multi-axis force detection and real-time performance, making it difficult to maintain robustness in complex scenarios.
An end-to-end visual-tactile perception method based on a shape-force field analytical model is adopted. By combining shape reconstruction and force field analysis, a deep structure engine model is used to perform high-precision collaborative reasoning of shape and force field. Combined with physical constraints for optimization training, the method can achieve simultaneous estimation of shape, normal force and multiple shear forces.
It achieves micron-level topography analysis and multi-axis force synchronous estimation, possesses real-time robust performance under high-frequency dynamic tasks, reduces the cost of tactile perception, and improves accuracy.
Smart Images

Figure CN121170541B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image analysis technology, and in particular to an end-to-end visual-tactile perception method, system, terminal, and computer-readable storage medium based on a shape-force field analytical model. Background Technology
[0002] With the rapid development of emerging applications such as robot operation, intelligent manufacturing, medical rehabilitation, and virtual reality or augmented reality, the demand for high-resolution tactile perception between end effectors and the environment is becoming increasingly urgent.
[0003] Existing visual-tactile sensing methods still have core limitations. One type of method relies on surface markers or speckle patterns for displacement field tracking. While this can obtain the deformation distribution of the contact area, it suffers from poor real-time performance, high dependence on image quality, and difficulty in stable operation in dynamic tasks. Another type of method can reconstruct micron-level contact morphology, but force estimation heavily relies on material parameter calibration, resulting in limited shear force detection accuracy, high computational complexity, and insufficient real-time performance.
[0004] Although end-to-end deep learning methods that have emerged in recent years have improved computation speed, their results often lack physical consistency constraints, and predictions are prone to drift, making it difficult to maintain robustness in complex scenarios.
[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0006] The main objective of this invention is to provide an end-to-end visual-tactile perception method, system, terminal, and computer-readable storage medium based on a shape-force field analytical model, aiming to solve the problems of insufficient spatial resolution, mechanical measurement range and accuracy, multi-axis force detection, and real-time performance in existing technologies for tactile analysis.
[0007] To achieve the above objectives, the present invention provides an end-to-end visual-tactile perception method based on a shape-force field analytical model, wherein the end-to-end visual-tactile perception method based on the shape-force field analytical model includes the following steps:
[0008] Obtain a deformation image of the test object, preprocess the deformation image to obtain an enhanced image, and obtain the deflection field of the test object based on the enhanced image;
[0009] The deflection field is input into the constructed force field analytical model, which generates the load, normal, and multiple shear force fields of the test object based on the deflection field.
[0010] The deformation image, the deflection field, the load, the normal force, and all the shear force fields are input into the deep model to establish a mapping relationship. Based on the mapping relationship, the topography, normal force, and multiple shear forces are constructed and output.
[0011] Consistency constraints are applied to all shear forces, and partial differential equation residuals are constructed using the morphology and the normal force. The deep engine model is then optimized and trained using the partial differential equation residuals.
[0012] The target object is analyzed using a trained deep engine model, and the target shape, target normal force, and multiple target shear forces of the target object are output.
[0013] Optionally, the end-to-end visual-tactile perception method based on the topography-force field analytical model, wherein acquiring a deformation image of the test object, preprocessing the deformation image to obtain an enhanced image, and obtaining the deflection field of the test object based on the enhanced image specifically includes:
[0014] Obtain the deformation image of the test object and input the deformation image into the data domain shaper in the topography reconstruction model;
[0015] The data domain shaper uses multiple parameters to perform AND transformation on the deformed image to obtain the transformed image:
[0016] ;
[0017] in, Represents a transformed image. This indicates an implicit modulation family. Represents a deformed image. Represents a parameter vector;
[0018] The transformed image is converted into a binary image, and the binary image is segmented to obtain the ROI regions with deformation. The enhanced image is then obtained by integrating all pixels within each ROI region.
[0019] ;
[0020] ;
[0021] in, Indicates a binary mask. Represents the height of a binary image. Indicates the ROI region. express ROI region in This indicates the bias correction term. Indicates an enhanced image. This indicates that the ROI region is being integrated. Represents the gradient operator. express The square of the modulus;
[0022] Based on the enhanced image, the contact surface of the test object in the world coordinate system is determined, and the horizontal and vertical deformations of the test object in the world coordinate system are determined based on the contact surface.
[0023] The deflection field of the test object is determined based on the horizontal axis deformation and the vertical axis deformation.
[0024] Optionally, the end-to-end visual-tactile perception method based on the shape-force field analytical model includes, wherein the shape reconstruction model comprises: a flexible sensing layer, a multi-wavelength light source, and an image acquisition model.
[0025] The acquisition of the deformation image of the test object specifically includes:
[0026] The light emitted by the multi-wavelength light source is scattered onto the flexible sensing layer through a light-blocking plate, a reflective lens, and a scattering plate;
[0027] The image acquisition model captures light scattered onto the flexible sensing layer to collect the deformation of the flexible sensing layer after it comes into contact with the test object, thereby obtaining a deformation image.
[0028] Optionally, in the end-to-end visual-tactile perception method based on the topography-force field analytical model, the shear force field includes: a transverse shear force field and a longitudinal shear force field.
[0029] The step of inputting the deflection field into the constructed force field analytical model, wherein the force field analytical model generates the load, normal, and multiple shear force fields of the test object based on the deflection field, specifically includes:
[0030] The deflection field is input into the constructed force field analytical model, which determines the load on the test object based on the deflection field.
[0031] ;
[0032] in, Indicates load, Denotes the divergence operator, Indicates the membrane tension coefficient. Represents the gradient operator. Represents the deflection field. Indicates the subgrade reaction coefficient. Indicates the shear bed parameters. Represents the Laplace operator. This represents the deformation of the flexible sensing layer on the horizontal axis after it comes into contact with the test object. This indicates the deformation of the flexible sensing layer along the vertical axis after it comes into contact with the test object.
[0033] The load is decomposed to obtain and output the normal, transverse shear force field and longitudinal shear force field of the test object.
[0034] Optionally, the end-to-end visual-tactile perception method based on the topography-force field analytical model, wherein the step of establishing a mapping relationship between the deformation image, the deflection field, the load, the normal force, and all the shear force fields input into the deep structure engine model, and constructing and outputting the topography, normal force, and multiple shear forces according to the mapping relationship, specifically includes:
[0035] The deformation image, the deflection field, the load, the normal force, the transverse shear force field, and the longitudinal shear force field are input into the deep structure engine model. The deep structure engine model converts the deflection field, the load, the normal force, the transverse shear force field, and the longitudinal shear force field into their corresponding tensor forms, and constructs the mapping relationship between all the tensor forms and the deformation image to obtain the final mapping features.
[0036] ;
[0037] in, An index representing the number of levels, Indicates the first Each mapping feature Indicates the first Layered convolutional-nonlinear composite operator, Indicates the first Each mapping feature Indicates the number of levels;
[0038] The deep engine model constructs the shape, normal force, transverse shear force, and longitudinal shear force of the test object based on the final mapping features:
[0039] ;
[0040] ;
[0041] ;
[0042] ;
[0043] in, Describing appearance, Represents normal force, Represents the shear force field along the transverse axis. Represents the shear force field along the longitudinal axis. The width of the mapping feature representing the shape. The width of the mapping feature representing the normal force. The width of the mapping characteristic representing the transverse shear force field. The width of the mapping characteristic of the longitudinal shear force field. This indicates that the volume satisfies the physical prior non-negative activation. This represents the final mapping feature.
[0044] Optionally, the end-to-end visual-tactile perception method based on a shape-force field analytical model, wherein the step of applying consistency constraints to all shear forces, constructing partial differential equation residuals using the shape and the normal force, and optimizing and training the deep engine model using the partial differential equation residuals, specifically includes:
[0045] The horizontal and vertical shear force fields are input into the physical law field model, which performs consistent characterization of the horizontal and vertical shear force fields and outputs the horizontal and vertical partial derivatives respectively.
[0046] ;
[0047] ;
[0048] in, Indicates the lateral partial derivative. Indicates the longitudinal partial derivative, and Both represent taking partial derivatives. This represents the combination of second-order partial derivatives of the abscissa in space. Represents the Laplace operator. This represents the displacement constraint field of the test object;
[0049] Based on the topography, normal force, lateral partial derivative, and longitudinal partial derivative at different resolutions, construct the partial differential equation residuals:
[0050] ;
[0051] in, Let represent the residual of the partial differential equation, s represent the index of different resolutions, and S represent the scale set. This represents the weight at resolution s. Represents the residual functional. This represents the morphology at resolution s. This represents the updated normal force at resolution s. This represents the lateral partial derivative at resolution s. This represents the longitudinal partial derivative at resolution s. express The square of the modulus;
[0052] The deep engine model is optimized and trained using the residuals of the partial differential equations to obtain the target deep engine model.
[0053] Optionally, the end-to-end visual-tactile perception method based on a shape-force field analytical model, wherein the step of analyzing the target object using a trained deep engine model and outputting the target object's shape, target normal force, and multiple target shear forces specifically includes:
[0054] The target image of the target object is acquired, and the target deflection field of the target object is determined based on the target image;
[0055] Input the target deflection field into the force field analytical model, and output the target normal, the target transverse shear force field and the target longitudinal shear force field of the target object;
[0056] A mapping relationship is constructed between the target deflection field, target load, target normal, target transverse shear force field, target longitudinal shear force field, and target image to obtain the target mapping image;
[0057] The target mapping image is input into the target deep structure engine model for analysis, and the target shape, target normal force, target transverse shear force and target longitudinal shear force of the target object are output.
[0058] Furthermore, to achieve the above objectives, the present invention also provides an end-to-end visual-tactile perception system based on a shape-force field analytical model, wherein the end-to-end visual-tactile perception system based on the shape-force field analytical model includes:
[0059] The data augmentation module is used to acquire the deformation image of the test object, preprocess the deformation image to obtain an augmented image, and obtain the deflection field of the test object based on the augmented image;
[0060] A force field analytical model is used to input the deflection field into the constructed force field analytical model, which generates the load, normal, and multiple shear force fields of the test object based on the deflection field.
[0061] The synchronous calculation module is used to input the deformation image, the deflection field, the load, the normal and all the shear force fields into the deep structure engine model to establish a mapping relationship, and construct the morphology, normal force and multiple shear forces according to the mapping relationship and output them.
[0062] The model training module is used to apply consistency constraints to all the shear forces, construct partial differential equation residuals using the morphology and the normal force, and optimize and train the deep engine model using the partial differential equation residuals.
[0063] The tactile analysis module is used to analyze the target object using a trained deep engine model, and output the target object's shape, target normal force, and multiple target shear forces.
[0064] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and an end-to-end visual-tactile perception program based on a shape-force field analytical model stored in the memory and executable on the processor, wherein when the end-to-end visual-tactile perception program based on the shape-force field analytical model is executed by the processor, it implements the steps of the end-to-end visual-tactile perception method based on the shape-force field analytical model as described above.
[0065] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an end-to-end visual-tactile perception program based on a shape-force field analytical model, wherein when the end-to-end visual-tactile perception program based on the shape-force field analytical model is executed by a processor, it implements the steps of the end-to-end visual-tactile perception method based on the shape-force field analytical model as described above.
[0066] In this invention, a deformation image of a test object is acquired, preprocessed to obtain an enhanced image, and the deflection field of the test object is obtained based on the enhanced image. The deflection field is input into a pre-constructed force field analytical model, which generates the load, normal force, and multiple shear force fields of the test object based on the deflection field. The deformation image, the deflection field, the load, the normal force, and all the shear force fields are input into a deep structure engine model to establish a mapping relationship. Based on the mapping relationship, the shape, normal force, and multiple shear forces are constructed and output. Consistency constraints are applied to all the shear forces. Simultaneously, the shape and normal force are used to construct partial differential equation residuals, and the deep structure engine model is optimized and trained using these residuals. The trained deep structure engine model is used to analyze the target object, outputting the target object's target shape, target normal force, and multiple target shear forces. This invention achieves micron-level shape analysis and simultaneous multi-axis force estimation, and possesses real-time robustness under high-frequency dynamic tasks, reducing the cost of tactile perception while improving the accuracy of tactile perception results. Attached Figure Description
[0067] Figure 1 This is a flowchart of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analytical model of the present invention;
[0068] Figure 2 This is a hardware schematic diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analytical model of the present invention.
[0069] Figure 3 This is a data production flowchart of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analytical model of the present invention.
[0070] Figure 4 This is a flowchart of the force calculation of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analytical model of the present invention.
[0071] Figure 5 This is an end-to-end result-morphology 2D comparison diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the morphology-force field analytical model of the present invention;
[0072] Figure 6 This is an end-to-end result-morphology 3D comparison diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the morphology-force field analytical model of the present invention.
[0073] Figure 7 This is a 2D comparison diagram of end-to-end results and force fields of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analytical model of the present invention.
[0074] Figure 8 This is an end-to-end result-force field 3D comparison diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analytical model of the present invention;
[0075] Figure 9 This is a structural diagram of a preferred embodiment of the end-to-end visual-tactile perception system based on the topography-force field analytical model of the present invention;
[0076] Figure 10 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation
[0077] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0078] Visor-tactile sensors achieve sub-millimeter to tens of micrometer-level topographic resolution by optically imaging the deformation of flexible membranes or gel surfaces. They possess the potential to simultaneously acquire contact morphology and mechanical information, making them a crucial development direction for intelligent tactile sensing. However, existing vision-tactile sensing methods still have core limitations. For example, marker-based or speckle-tracking-based methods suffer from poor real-time performance, with lengthy image registration and optical flow calculations, making them unsuitable for high-frame-rate applications. They also lack stability, being highly dependent on marker density and imaging conditions, and susceptible to distortion due to occlusion or noise. Topographic reconstruction methods based on multi-source coding have limited mechanical estimation accuracy, with results highly dependent on material parameters and boundary conditions, making them sensitive to biases. Multi-axial force estimation is insufficient, resulting in limited accuracy in shear force detection, with errors easily occurring in direction and amplitude. Furthermore, they lack real-time performance, have high computational complexity, and are unsuitable for high-speed or dynamic operations. The lack of physical constraints in deep learning-based end-to-end prediction methods makes the results prone to drift and distortion, and lacks mechanical consistency; the measurement accuracy is limited, and it is difficult to distinguish minute force changes of less than 0.1N under low load conditions; the functionality is limited, and most of them can only predict a single mode (such as morphology or normal force), making it difficult to output multi-axis mechanical information at the same time.
[0079] For the reasons mentioned above, this invention proposes a high-resolution visual-tactile perception system based on an end-to-end shape-force analytical neural network. Its core idea is to combine "high-precision dataset creation—shape reconstruction and force field calculation" with "an end-to-end neural network for physical constraint fusion (GeoMech-Net, i.e., the deep engine model proposed in this invention)" to achieve collaborative high-precision reasoning of shape and force field.
[0080] The end-to-end visual-tactile perception method based on a shape-force field analytical model described in the preferred embodiment of the present invention, such as... Figure 1 As shown, the end-to-end visual-tactile perception method based on the shape-force field analytical model includes the following steps:
[0081] Step S10: Obtain the deformation image of the test object, preprocess the deformation image to obtain an enhanced image, and obtain the deflection field of the test object based on the enhanced image.
[0082] First, in the hardware section, a shape reconstruction module is used to acquire the deformation of the test object in the form of images; among which, such as Figure 2 As shown, the shape reconstruction model includes: a flexible sensing layer, a multi-wavelength light source, and an image acquisition model; where F represents the force applied when the test object touches the flexible sensing layer.
[0083] The multi-wavelength light source scatters its light onto the flexible sensing layer through a light-blocking plate, a reflecting mirror, and a scattering plate. The image acquisition model captures the light scattered onto the flexible sensing layer to collect the deformation of the flexible sensing layer after it comes into contact with the test object, thus obtaining a deformation image. When the test object comes into contact with the flexible sensing layer, the flexible sensing layer is subjected to a force. By irradiating it with the multi-wavelength light source, a deformation image of the flexible sensing layer is captured, simulating the deformation image of the test image.
[0084] In one embodiment of this invention, the flexible sensing layer and the sensor housing form a sealed space, effectively shielding the influence of ambient light and reducing background errors in the image. The multi-wavelength light source and image acquisition module are integrated onto the same PCB board, reducing size and improving system portability. The flexible sensing layer is made by uniformly mixing silicone, aluminum alloy powder, and fumed silica in a specific ratio and then curing it to achieve strong internal diffuse reflection and external dryness and wear resistance, providing a foundation for sensing external deformation and force. During illumination, a light-blocking plate isolates the multi-wavelength light source, preventing it from being directly captured by the image acquisition module without reflection. The multi-wavelength light source is scattered onto the flexible sensing layer through a light guide plate, reflecting mirror, and scattering acrylic plate, making the deformation of the flexible sensing layer easier to identify and providing a basis for subsequent processing.
[0085] Furthermore, the size of the film can be adjusted to achieve precise sensing of different areas, and the thickness of the film can be adjusted to achieve relatively precise sensing of different ranges. For the equipment used for shooting, mobile phones can be used directly to take pictures, which significantly reduces costs.
[0086] For the software portion, a pre-trained topography reconstruction network is used to reconstruct the depth of the thin film, ultimately obtaining high-precision depth data. Specifically, a deformation image of the test object is acquired and input into the data domain shaper in the topography reconstruction model; the data domain shaper performs AND transformations on the deformation image using various parameters to obtain a transformed image.
[0087] ;
[0088] in, Represents a transformed image. This indicates an implicit modulation family. Represents a deformed image. The parameter vector is represented; the transformed image is converted into a binary image, and the binary image is segmented to obtain the ROI regions with deformation. The enhanced image is obtained by integrating all pixels within the ROI regions.
[0089] ;
[0090] ;
[0091] in, Indicates a binary mask. Represents the height of a binary image. Indicates the ROI region. express ROI region in This indicates the bias correction term. Indicates an enhanced image. This indicates that the ROI region is being integrated. Represents the gradient operator. express The square of the modulus; based on the enhanced image, determine the contact surface of the test object in the world coordinate system, and based on the contact surface, determine the transverse and longitudinal deformations of the test object in the world coordinate system; based on the transverse and longitudinal deformations, determine the deflection field of the test object.
[0092] In image processing and computer vision, the Region of Interest (ROI) is a core concept, referring to a specific sub-region selected from an image to focus on key targets and improve processing efficiency and accuracy. To maintain high image resolution, in addition to acquiring contact surface details through a micrometer-level topography reconstruction module, a data domain shaper is needed to implicitly restructure the input image. Spatial domain modulation is achieved through implicit modulation families, employing non-explicit parameterized operators to perform domain transformations on the input image, ensuring the model remains robust and adaptable under different experimental conditions. Implicit modulation families are techniques that achieve signal modulation through implicit functions or non-explicit parameterization. Their core characteristic is that the modulation process does not directly change the explicit parameters of the carrier (such as amplitude and phase), but indirectly affects signal characteristics through implicit mapping or nonlinear transformations.
[0093] Step S20: Input the deflection field into the constructed force field analytical model. The force field analytical model generates the load, normal, and multiple shear force fields of the test object based on the deflection field.
[0094] The shear force field includes: the transverse shear force field and the longitudinal shear force field; such as Figure 3 As shown, based on the enhanced image (i.e., the processed image), the deformation of the flexible sensing layer along the horizontal and vertical axes of the world coordinate system (i.e., Figure 3The relative depth of the membrane is used to determine the deflection field of the test object. Then, a force field analytical model is used to calculate the force field of the test object by combining the deflection field with the physical mechanics model (i.e., the governing equations between the deflection field and the load). This enhances multi-dimensional tactile perception and yields the corresponding absolute depth, relevant physical parameter calibration results, and the true force field. The force field analytical model achieves accurate inference of the normal force and shear force in the contact area through the coupling of the membrane deflection field and the physical mechanics model. This module eliminates the uncertainty in multi-axis force estimation using traditional empirical fitting methods, ensuring that the output results have strict physical consistency.
[0095] Specifically, the deflection field is input into the constructed force field analytical model, which determines the load on the test object based on the deflection field.
[0096] ;
[0097] in, Indicates load, Denotes the divergence operator, Indicates the membrane tension coefficient. Represents the gradient operator. Represents the deflection field. Indicates the subgrade reaction coefficient. Indicates the shear bed parameters. Represents the Laplace operator. This represents the deformation of the flexible sensing layer on the horizontal axis after it comes into contact with the test object. This represents the deformation along the longitudinal axis after the flexible sensing layer comes into contact with the test object; the load is decomposed to obtain and output the normal, the transverse shear force field, and the longitudinal shear force field of the test object.
[0098] Within the ROI region, at the thin film, the load on the test object (referring to the external force acting on an object or structure to cause internal forces and deformation) can be generated based on the generated deflection field according to the governing equations. The deflection field output by the topography reconstruction model serves as the input to the force field analytical model. After numerical calculation using the governing equations, the total force on the test object (i.e., the total force on the flexible sensing layer when the test object comes into contact with it) is obtained. After decomposition, the normal and shear components (i.e., the shear force fields along the horizontal and vertical axes) can be determined. Based on these two shear force fields, the in-plane friction and tangential action laws can be determined.
[0099] Specifically, by employing a force field analytical model, a high-precision mapping from topography to force field is achieved while ensuring physical constraints, providing reliable supervision signals and calibration benchmarks for neural network training and practical applications. Furthermore, by acquiring contact surface details through a micrometer-level topography reconstruction module and combining it with a force field analytical model based on physical equations, high-precision load inference within the range of 0–20 N is achieved, with a normal force error of less than 5%, capable of resolving minute force changes of less than 0.1 N.
[0100] Step S30: Input the deformation image, the deflection field, the load, the normal and all the shear force fields into the deep structure engine model to establish a mapping relationship, and construct the morphology, normal force and multiple shear forces according to the mapping relationship and output them.
[0101] Based on the above-mentioned total force decomposition, the GeoMech-Net neural network (i.e., the deep engine model described in this invention) combines data-driven and physical consistency constraints, avoiding the prediction drift and numerical distortion problems of pure deep learning methods, and achieving stable and reliable end-to-end solution.
[0102] Specifically, the deformation image, the deflection field, the load, the normal force, the transverse shear force field, and the longitudinal shear force field are input into the deep structure engine model. The deep structure engine model converts the deflection field, the load, the normal force, the transverse shear force field, and the longitudinal shear force field into their corresponding tensor forms, and constructs a mapping relationship between all the tensor forms and the deformation image to obtain the final mapping features.
[0103] ;
[0104] in, An index representing the number of levels, Indicates the first Each mapping feature Indicates the first Layered convolutional-nonlinear composite operator, Indicates the first Each mapping feature The depth engine model represents the number of levels; based on the final mapping features, it constructs the shape, normal force, transverse shear force, and longitudinal shear force of the test object.
[0105] ;
[0106] ;
[0107] ;
[0108] ;
[0109] in, Describing appearance, Represents normal force, Represents the shear force field along the transverse axis. Represents the shear force field along the longitudinal axis. The width of the mapping feature representing the shape. The width of the mapping feature representing the normal force. The width of the mapping characteristic representing the transverse shear force field. The width of the mapping characteristic of the longitudinal shear force field. This indicates that the volume satisfies the physical prior non-negative activation. This represents the final mapping feature.
[0110] In the embodiments disclosed in this invention, the deep engine model undertakes the task of hierarchical analysis and calculation of high-dimensional features. The deep engine model constructs a feature pyramid through layer-by-layer convolution and upsampling, thereby realizing the mapping from the input image to a multi-scale representation. First, various quantity fields (i.e., deflection field, load, normal force, transverse shear force field, and longitudinal shear force field) need to be input into the already constructed initial deep engine model for preliminary optimization, so that the deep engine model can identify these quantity fields in the image, thereby constructing a mapping relationship between the image and these quantity fields. Then, the deformed image is input into the pre-optimized deep engine model, which converts these quantity fields in the deformed image into tensor form (its dimension is B×C×H×W, i.e., batch size × number of feature channels × feature map height × feature map width). After feature mapping at each layer, the mapping relationship between these feature mappings and the original deformed image is established, resulting in a deformed image with the final mapped features. Then, branch decoupling is performed based on the final mapped features. This process is implemented by constructing four branches on the shared backbone features, such as... Figure 4 As shown, the final output includes the shape, normal force, transverse shear force, and longitudinal shear force of the test object.
[0111] The deep-structure engine model, through the construction of a multi-scale feature pyramid, can fully capture both the minute local texture information contained in the contact area and the large-scale changes in the overall morphology, thus providing a comprehensive analysis of complex contact scenes. Based on this, the deep-structure engine model adopts a four-head decoupling design, mapping the shared feature space to morphology, normal force, and shear force components (horizontal and vertical shear forces), achieving end-to-end complete mechanical field output without additional intermediate processes, avoiding the cascading error accumulation problem of the traditional "morphology first, then mechanics" approach. Simultaneously, through the subsequent physical law field module, the model introduces non-negative constraints and physical priors conforming to physical laws at the output end, making the prediction results not only more numerically accurate but also exhibiting strict consistency and rationality in mechanical interpretation, thus achieving both high accuracy and high interpretability.
[0112] Step S40: Apply consistency constraints to all shear forces, construct partial differential equation residuals using the morphology and the normal force, and optimize and train the deep engine model using the partial differential equation residuals.
[0113] In contrast to learning approaches that rely solely on data supervision, the embodiments disclosed in this invention embed physical constraints directly into the training process, constructing differentiable PDE residuals (Partial Differential Equation Residuals) and soft boundary constraints. These constraints are aggregated across multiple scales and can be exchanged and combined with stable boundary treatments (such as mirroring, DCT equivalence, finite difference, etc.), ensuring both numerical stability and boundary physicality while maintaining adaptability to complex contact geometries.
[0114] Specifically, the horizontal and vertical shear force fields are input into the physical law field model, which performs consistent characterization of the horizontal and vertical shear force fields, respectively, and outputs the horizontal and vertical partial derivatives:
[0115] ;
[0116] ;
[0117] in, Indicates the lateral partial derivative. Indicates the longitudinal partial derivative, and Both represent taking partial derivatives. This represents the combination of second-order partial derivatives of the abscissa in space. Represents the Laplace operator. Represent the displacement constraint field of the test object; construct the partial differential equation residuals based on the topography, normal force, lateral partial derivative, and longitudinal partial derivative at different resolutions:
[0118] ;
[0119] in, Let represent the residual of the partial differential equation, s represent the index of different resolutions, and S represent the scale set. This represents the weight at resolution s. Represents the residual functional. This represents the morphology at resolution s. This represents the updated normal force at resolution s. This represents the lateral partial derivative at resolution s. This represents the longitudinal partial derivative at resolution s. express The square of the modulus; the deep engine model is optimized and trained using the residual of the partial differential equation to obtain the target deep engine model.
[0120] Among them, the physical law field model is based on multiple quantity fields output by the deep engine model, constructs their consistency, and performs a consistent characterization of the tangential force component to obtain the lateral and longitudinal partial derivatives. Then, based on backgrounds of different resolutions, multi-scale constraints are applied, and the final calculated PDE residuals are used as the loss term of the deep engine model for physical constraints.
[0121] By transforming physical priors into a unified optimizable objective and integrating them with representation learning, the model can maintain overall consistent convergence of the morphology-force field in weakly labeled or domain-transfer scenarios, improving boundary stability, significantly suppressing edge artifacts and leakage, and taking into account tension anisotropy, bed parameters and higher-order regularization terms. It can also absorb material, thickness or pixel size differences into physical terms and scaling strategies, reducing the recalibration burden.
[0122] Step S50: Analyze the target object using the trained deep engine model and output the target shape, target normal force, and multiple target shear forces of the target object.
[0123] Specifically, the target image of the target object is acquired, and the target deflection field of the target object is determined based on the target image; the target deflection field is input into the force field analytical model, and the target normal, target transverse shear force field, and target longitudinal shear force field of the target object are output; the mapping relationship between the target deflection field, target load, target normal, target transverse shear force field, and target longitudinal shear force field and the target image is constructed to obtain a target mapping image; the target mapping image is input into the target deep structure engine model for analysis, and the target morphology, target normal force, target transverse shear force, and target longitudinal shear force of the target object are output.
[0124] Specifically, based on the trained deep engine model, the target image is processed through the same steps described above using a morphology reconstruction model and a force field analysis model to obtain a target mapping image with mapping relationships. This image is then input into the trained deep engine model to perform contact surface analysis on the target object, obtaining the target morphology, target normal force, target transverse shear force, and target longitudinal shear force. Figure 5 and Figure 6 As shown, Figure 5 (a) and Figure 6 In the examples (a), each image represents the input target image. Figure 5 In the figure, (b) represents the predicted two-dimensional depth of the target object. Figure 6 In the figure, (b) represents the predicted three-dimensional depth of the target object; Figure 5In the figure, (c) represents the two-dimensional depth of the actual target object. Figure 6 In the figure, (c) represents the actual three-dimensional depth of the target object; Figure 5 In the figure, (d) represents the error between the true 2D depth and the predicted 2D depth. Figure 6 In the figure, (d) represents the error between the actual 3D depth and the predicted 3D depth.
[0125] Furthermore, such as Figure 7 and Figure 8 As shown, Figure 7 (a) and Figure 8 In the examples (a), each image represents the input target image. Figure 7 In (b), the predicted two-dimensional load on the target object is shown. Figure 8 (b) in the figure represents the predicted three-dimensional load of the target object; Figure 7 In the diagram, (c) represents the two-dimensional load of the actual target object. Figure 8 (c) in the figure represents the three-dimensional load of the actual target object; Figure 7 In the figure, (d) represents the error between the actual two-dimensional load and the predicted two-dimensional load. Figure 8 In the equation (d), the error between the actual three-dimensional load and the predicted three-dimensional load is represented.
[0126] The force field analytical model disclosed in this invention differs from traditional force field analysis, which only provides an estimate of the normal force. This model can simultaneously output the distribution of the normal force and the in-plane shear force, accurately providing their direction and amplitude, thus enhancing multi-dimensional tactile perception capabilities. The proposed deep engine model combines data-driven approaches with physical consistency constraints, avoiding the prediction drift and numerical distortion problems inherent in purely deep learning methods, achieving stable and reliable end-to-end computation. Furthermore, the image-enhanced preprocessing ensures that the results maintain high resolution and accuracy while possessing real-time inference capabilities, adapting to different contact objects and complex operating conditions, and maintaining stable output even in dynamic tasks. The system architecture is compatible with various sensors and execution platforms, suitable for scenarios such as precision robot operation, flexible manufacturing inspection, medical rehabilitation, and virtual interaction, possessing broad industrial application prospects.
[0127] This invention achieves micron-level morphology analysis and simultaneous multi-axis force estimation, and possesses real-time robust performance under high-frequency dynamic tasks, reducing the cost of tactile perception while improving the accuracy of tactile perception results.
[0128] Furthermore, such as Figure 9 As shown, based on the above-mentioned end-to-end visual-tactile perception method based on the shape-force field analytical model, the present invention also provides an end-to-end visual-tactile perception system based on the shape-force field analytical model, wherein the end-to-end visual-tactile perception system based on the shape-force field analytical model includes:
[0129] The data augmentation module 51 is used to acquire the deformation image of the test object, preprocess the deformation image to obtain an augmented image, and obtain the deflection field of the test object based on the augmented image.
[0130] Force field analytical model 52 is used to input the deflection field into the constructed force field analytical model, which generates the load, normal and multiple shear force fields of the test object based on the deflection field;
[0131] The synchronous calculation module 53 is used to input the deformation image, the deflection field, the load, the normal and all the shear force fields into the deep structure engine model to establish a mapping relationship, and construct the morphology, normal force and multiple shear forces according to the mapping relationship and output them.
[0132] The model training module 54 is used to apply consistency constraints to all the shear forces, construct partial differential equation residuals using the morphology and the normal force, and optimize and train the deep engine model using the partial differential equation residuals.
[0133] The tactile analysis module 55 is used to analyze the target object using a trained deep engine model and output the target object's shape, target normal force, and multiple target shear forces.
[0134] Furthermore, such as Figure 10 As shown, based on the above-mentioned end-to-end visual-tactile perception method and system based on the shape-force field analytical model, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 10 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0135] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores an end-to-end visual-tactile perception program 40 based on a shape-force field analytical model. This end-to-end visual-tactile perception program 40 based on the shape-force field analytical model can be executed by the processor 10, thereby implementing the end-to-end visual-tactile perception method based on the shape-force field analytical model in this application.
[0136] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the end-to-end visual-tactile perception method based on the topography-force field analytical model.
[0137] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0138] In one embodiment, when the processor 10 executes the end-to-end visual-tactile perception program 40 based on the shape-force field analytical model in the memory 20, it implements the steps of the end-to-end visual-tactile perception method based on the shape-force field analytical model as described above.
[0139] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an end-to-end visual-tactile perception program based on a shape-force field analytical model, wherein when the end-to-end visual-tactile perception program based on the shape-force field analytical model is executed by a processor, it implements the steps of the end-to-end visual-tactile perception method based on the shape-force field analytical model as described above.
[0140] In summary, this invention provides an end-to-end visual-tactile perception method and related equipment based on a shape-force field analytical model. The method includes: acquiring a deformation image of a test object; preprocessing the deformation image to obtain an enhanced image; and obtaining the deflection field of the test object based on the enhanced image; inputting the deflection field into a pre-constructed force field analytical model, which generates the load, normal force, and multiple shear force fields of the test object based on the deflection field; inputting the deformation image, the deflection field, the load, the normal force, and all the shear force fields into a deep structure engine model to establish a mapping relationship; constructing and outputting the shape, normal force, and multiple shear forces based on the mapping relationship; applying consistency constraints to all the shear forces; simultaneously constructing partial differential equation residuals using the shape and the normal force; and optimizing and training the deep structure engine model using the partial differential equation residuals; and analyzing the target object using the trained deep structure engine model, outputting the target shape, target normal force, and multiple target shear forces of the target object. This invention achieves micron-level morphology analysis and simultaneous multi-axis force estimation, and possesses real-time robust performance under high-frequency dynamic tasks, reducing the cost of tactile perception while improving the accuracy of tactile perception results.
[0141] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0142] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.
[0143] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. An end-to-end haptic perception method based on a topography-force field analytical model, characterized in that, The end-to-end visual-haptic perception method based on the topography-force field analytical model comprises: obtaining a deformation image of a test object, pre-processing the deformation image to obtain an enhanced image, and obtaining a deflection field of the test object according to the enhanced image; The obtaining of the deformation image of the test object, the pre-processing of the deformation image to obtain the enhanced image, and the obtaining of the deflection field of the test object according to the enhanced image specifically comprises: obtaining the deformation image of the test object, and inputting the deformation image into a data domain shaper in a topography reconstruction model; The data domain shaper utilizes multiple parameters to perform transformation on the deformation image to obtain a transformed image: ; wherein, denotes a transformed image, denotes an implicit modulation family, denotes a morphed image, denotes a parameter vector; convert the transformed image into a binary image, segment the binary image to obtain an ROI region with deformation, and integrate all pixels in the ROI region to obtain an enhanced enhanced image: ; ; wherein, denotes a binarization mask, denotes the height of the binarized image, denotes the ROI region, denotes the ROI region in, denotes a bias correction term, denotes the enhanced image, denotes the integration of the ROI region, denotes a gradient operator, denotes the square of the modulus of According to the enhanced image, the contact surface of the test object in the world coordinate system is determined, and the lateral axis deformation and the longitudinal axis deformation of the test object in the world coordinate system are determined according to the contact surface; According to the lateral axis deformation and the longitudinal axis deformation, the deflection field of the test object is determined; inputting the deflection field into the constructed force field analytical model, the force field analytical model generating the load, the normal and multiple shear force fields of the test object according to the deflection field; inputting the deformation image, the deflection field, the load, the normal and all the shear force fields into a deep construction engine model to establish a mapping relationship, constructing topography, normal force and multiple shear forces according to the mapping relationship and outputting; consistently constrain all the shear forces, construct partial differential equation residuals by using the topography and the normal force, and optimize and train the deep construction engine model by using the partial differential equation residuals; using the trained deep construction engine model to analyze a target object and output target topography, target normal force and multiple target shear forces of the target object.
2. The end-to-end haptic perception method based on a topography-force field analytical model according to claim 1, wherein, The topography reconstruction model comprises a flexible perception layer, a multi-wavelength light source and an image acquisition model; The obtaining of the deformation image of the test object specifically comprises: the light source emitted by the multi-wavelength light source is scattered onto the flexible perception layer through a light blocking plate, a reflecting lens and a scattering plate; The image acquisition model captures the light source scattered onto the flexible perception layer to acquire the deformation of the flexible perception layer after contacting with the test object to obtain a deformation image.
3. The end-to-end haptic perception method based on a topography-force field analytical model according to claim 1, wherein, The shear force field comprises a lateral axis shear force field and a longitudinal axis shear force field; The inputting of the deflection field into the constructed force field analytical model, the force field analytical model generating the load, the normal and multiple shear force fields of the test object according to the deflection field, specifically comprises: inputting the deflection field into the constructed force field analytical model, the force field analytical model determining the load of the test object according to the deflection field; ; wherein, denotes the load, denotes the divergence operator, denotes the membrane tension coefficient, denotes the gradient operator, denotes the deflection field, denotes the base reaction coefficient, denotes the shear base parameter, denotes the Laplace operator, denotes the deformation of the flexible sensing layer in the lateral axis after the flexible sensing layer contacts with the test object, denotes the deformation of the flexible sensing layer in the longitudinal axis after the flexible sensing layer contacts with the test object. decomposing the load to obtain the normal of the test object, the lateral axis shear force field and the longitudinal axis shear force field and outputting.
4. The haptic perception method based on the analytical model of topography-force field according to claim 3, wherein, The deformation image, the deflection field, the load, the normal force and all the shear force fields are input into a deep construction engine model to establish a mapping relationship, and the topography, normal force and multiple shear forces are constructed and output according to the mapping relationship, specifically including: The deformation image, the deflection field, the load, the normal force, the transverse shear force field and the longitudinal shear force field are input into a deep construction engine model, and the deflection field, the load, the normal force, the transverse shear force field and the longitudinal shear force field are respectively converted into corresponding tensor forms by the deep construction engine model, and a mapping relationship between all the tensor forms and the deformation image is constructed to obtain final mapping features: ; wherein, an index representing the number of levels, an index representing the th mapped feature, an index representing the th convolution-nonlinear composite operator, an index representing the th mapped feature, an index representing the number of levels; The deep construction engine model constructs the topography, normal force, transverse shear force and longitudinal shear force of the test object according to the final mapping features: ; ; ; ; wherein, represents a topography, represents a normal force, represents a transverse shear force field, represents a longitudinal shear force field, represents a width of a mapped feature of the topography, represents a width of a mapped feature of the normal force, represents a width of a mapped feature of the transverse shear force field, represents a width of a mapped feature of the longitudinal shear force field, represents a non-negative activation that the volume satisfies a physical prior, represents a final mapped feature.
5. The end-to-end haptic perception method based on a topography-force field analytical model according to claim 1, wherein, The consistency of all the shear forces is constrained, and a partial differential equation residual is constructed using the topography and the normal force, and the deep construction engine model is optimized and trained using the partial differential equation residual, specifically including: The transverse shear force field and the longitudinal shear force field are input into a physical law field model, and the transverse shear force field and the longitudinal shear force field are respectively described by the physical law field model, and the transverse partial derivative and the longitudinal partial derivative are output: ; ; wherein denotes the lateral deflection, denotes the longitudinal deflection, and each denote the partial derivative, denotes the second order partial derivative combination of the spatial abscissa, denotes the Laplace operator, denotes the displacement constraint field of the test object; A partial differential equation residual is constructed according to the topography, the normal force, the transverse partial derivative and the longitudinal partial derivative under different resolutions: ; wherein, represents a partial differential equation residual, s represents an index of different resolutions, S represents a set of scales, represents a weight at resolution s, represents a residual functional, represents a topography at resolution s, represents an updated normal force at resolution s, represents a lateral partial derivative at resolution s, represents a longitudinal partial derivative at resolution s, represents a square of a modulus of The deep construction engine model is optimized and trained using the partial differential equation residual to obtain a target deep construction engine model.
6. The end-to-end haptic perception method based on a topography-force field analytical model according to claim 5, wherein, The trained deep construction engine model is used to analyze the target object, and the target topography, target normal force and multiple target shear forces of the target object are output, specifically including: A target image of the target object is obtained, and a target deflection field of the target object is determined according to the target image; The target deflection field is input into the force field analysis model to output the target normal force, target transverse shear force field and target longitudinal shear force field of the target object; A mapping relationship between the target deflection field, target load, target normal force, target transverse shear force field and target longitudinal shear force field and the target image is constructed to obtain a target mapping image; The target mapping image is input into the target deep construction engine model for analysis, and the target topography, target normal force, target transverse shear force and target longitudinal shear force of the target object are output.
7. A topography-force field analytical model based end-to-end visuo-haptic perception system for implementing the topography-force field analytical model based end-to-end visuo-haptic perception method of any one of claims 1-6, characterized in that, The end-to-end visual-haptic perception system based on the topography-force field analysis model includes: A data enhancement module is configured to obtain a deformation image of a test object, pre-process the deformation image to obtain an enhanced image, and obtain a deflection field of the test object according to the enhanced image; A force field analysis model is configured to input the deflection field into a constructed force field analysis model, and the force field analysis model generates a load, a normal force and multiple shear force fields of the test object according to the deflection field; A synchronization solving module is configured to input the deformation image, the deflection field, the load, the normal and all the shear force fields into a deep construction engine model to establish a mapping relationship, construct a topography, a normal force and multiple shear forces according to the mapping relationship and output the topography, the normal force and the multiple shear forces; A model training module is configured to perform consistency constraint on all the shear forces, construct a partial differential equation residual by using the topography and the normal force, and perform optimization training on the deep construction engine model by using the partial differential equation residual; A haptic analysis module is configured to analyze a target object by using the trained deep construction engine model and output a target topography, a target normal force and multiple target shear forces of the target object.
8. A terminal, characterized by comprising: The terminal comprises a memory, a processor and an end-to-end visual-haptic perception program based on a topography-force field analysis model stored on the memory and executable on the processor, and the end-to-end visual-haptic perception program based on the topography-force field analysis model implements the steps of the end-to-end visual-haptic perception method based on the topography-force field analysis model according to any one of claims 1-6 when executed by the processor.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores an end-to-end visual-haptic perception program based on a topography-force field analysis model, and the end-to-end visual-haptic perception program based on the topography-force field analysis model implements the steps of the end-to-end visual-haptic perception method based on the topography-force field analysis model according to any one of claims 1-6 when executed by the processor.
Citation Information
Patent Citations
Intelligent video monitoring system and method based on deep learning
CN120263940A
End-to-end automatic driving control method and device based on multi-camera fusion
CN120411902A