End-to-end visual tactile perception method and system based on morphology-force field analytical model, terminal and storage medium
By adopting an end-to-end visual-tactile sensing method based on a topography-force field analytical model, the problems of insufficient resolution, accuracy and real-time performance of visual-tactile sensing in existing technologies are solved. It achieves micron-level topography analysis and multi-axis force synchronous estimation, and has a highly efficient technical sensing effect.
Patent Information
- Application Number
- CN202511715651.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2045-11-21
AI Technical Summary
Existing visual-tactile sensing methods have shortcomings in spatial resolution, mechanical measurement range and accuracy, multi-axis force detection and real-time performance, making it difficult to maintain robustness in complex scenarios.
An end-to-end visual-tactile perception method based on a topography-force field analytical model is adopted. The deflection field is obtained through a topography reconstruction model, and the load, normal and shear force fields are generated by combining the force field analytical model. The mapping relationship is established using a deep structure engine model to construct topography, normal force and multiple shear forces. Consistency constraints and optimization training are performed to achieve high-precision collaborative reasoning of topography and force field.
It achieves micron-level topography analysis and multi-axis force synchronous estimation, possesses real-time robust performance under high-frequency dynamic tasks, reduces the cost of tactile perception, and improves accuracy.
Smart Images

Figure CN121170541A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image analysis, in particular to an end-to-end visual-haptic perception method and system based on a topography-force field analytical model, a terminal and a computer readable storage medium. BACKGROUND
[0002] With the rapid development of emerging applications such as robot operation, intelligent manufacturing, medical rehabilitation, and virtual reality or augmented reality, the demand for high-resolution haptic perception between the end effector and the environment is increasingly urgent.
[0003] The existing visual-haptic sensing method still has core limitations. One method relies on surface marker points or speckles for displacement field tracking, which can obtain the deformation distribution of the contact area, but has poor real-time performance, is highly dependent on imaging quality, and is difficult to work stably in dynamic tasks. Another method can reconstruct the micron-level contact topography, but the force estimation is heavily dependent on material parameter calibration, the shear force detection precision is limited, and the computational complexity is high, which is not real-time enough.
[0004] Although the end-to-end deep learning method appeared in recent years has improved the speed of solution, the results often lack physical consistency constraints, and the prediction has drift, which is difficult to maintain robustness in complex scenarios.
[0005] Therefore, the prior art still needs to be improved and developed. SUMMARY
[0006] The main purpose of the present application is to provide an end-to-end visual-haptic perception method and system based on a topography-force field analytical model, a terminal and a computer readable storage medium, which aims to solve the problems of insufficient spatial resolution, mechanical measurement range and precision, multi-axis force detection, and real-time performance in the prior art when performing haptic analysis.
[0007] To achieve the above purpose, the present application provides an end-to-end visual-haptic perception method based on a topography-force field analytical model, which comprises the following steps: Obtain the deformation image of the test object, pre-process the deformation image to obtain an enhanced image, and obtain the deflection field of the test object according to the enhanced image; Input the deflection field into the force field analytical model, which generates the load, normal and multiple shear force fields of the test object according to the deflection field; Input the deformation image, deflection field, load, normal and all shear force fields into a deep learning engine model to establish a mapping relationship, and construct the topography, normal force and multiple shear forces according to the mapping relationship and output; The consistency constraints are performed on all the shear forces, while a partial differential equation residual is constructed using the topography and the normal force, and the deep structure engine model is optimized and trained using the partial differential equation residual; A target object is analyzed using the trained deep structure engine model, and target topography, target normal force and multiple target shear forces of the target object are output.
[0008] Optionally, the end-to-end visual-haptic perception method based on the topography-force field analysis model, wherein the deformed image of the test object is obtained, the deformed image is preprocessed to obtain an enhanced image, and the deflection field of the test object is obtained according to the enhanced image, specifically comprising: The deformed image of the test object is obtained, and the deformed image is input into the data domain shaper in the topography reconstruction model; The data domain shaper uses multiple parameters to transform the deformed image to obtain a transformed image: ; Wherein, represents the transformed image, represents an implicit modulation family, represents the deformed image, represents a parameter vector; The transformed image is converted into a binary image, and the binary image is segmented to obtain an ROI region with deformation, and all pixels in the ROI region are integrated to obtain an enhanced enhanced image: ; ; Wherein, represents a binary mask, represents the height of the binary image, represents the ROI region, represents the ROI region in represents a bias correction term, represents the enhanced image, represents the integration processing of the ROI region, represents a gradient operator, represents the square of the modulus of ; According to the enhanced image, the contact surface of the test object in the world coordinate system is determined, and the lateral axis deformation and the longitudinal axis deformation of the test object in the world coordinate system are determined according to the contact surface; According to the lateral axis deformation and the longitudinal axis deformation, the deflection field of the test object is determined.
[0009] Optionally, the end-to-end visual-haptic perception method based on the topography-force field analytical model, wherein the topography reconstruction model comprises a flexible perception layer, a multi-wavelength light source, and an image acquisition model. The deformed image of the test object is obtained, specifically comprising: The light source emitted by the multi-wavelength light source is scattered onto the flexible perception layer through a light blocking plate, a reflecting lens, and a scattering plate. The image acquisition model captures the light source scattered onto the flexible perception layer to acquire the deformation of the flexible perception layer after contacting the test object, thereby obtaining a deformed image.
[0010] Optionally, the end-to-end visual-haptic perception method based on the topography-force field analytical model, wherein the shear force field comprises a horizontal-axis shear force field and a vertical-axis shear force field. The deflection field is input into the constructed force field analytical model, and the force field analytical model generates a load, a normal force, and a plurality of shear force fields of the test object according to the deflection field, specifically comprising: The deflection field is input into the constructed force field analytical model, and the force field analytical model determines the load of the test object according to the deflection field. ; wherein, represents the load, represents the divergence operator, represents the membrane tension coefficient, represents the gradient operator, represents the deflection field, represents the base bed reaction coefficient, represents the shear base bed parameter, represents the Laplace operator, represents the deformation of the flexible perception layer in the horizontal axis after contacting the test object, represents the deformation of the flexible perception layer in the vertical axis after contacting the test object. The load is decomposed to obtain the normal force, the horizontal-axis shear force field, and the vertical-axis shear force field of the test object and output.
[0011] Optionally, the end-to-end visual-haptic perception method based on the topography-force field analytical model, wherein the deformed image, the deflection field, the load, the normal force, and all the shear force fields are input into a deep learning engine model to establish a mapping relationship, and a topography, a normal force, and a plurality of shear forces are constructed and output according to the mapping relationship, specifically comprising: inputting the deformation image, the deflection field, the load, the normal force, the transverse shear force field and the longitudinal shear force field into a deep structure engine model, the deep structure engine model converting the deflection field, the load, the normal force, the transverse shear force field and the longitudinal shear force field into corresponding tensor forms respectively, and constructing a mapping relationship between all the tensor forms and the deformation image to obtain final mapping features: ; wherein, represents an index of a layer number, represents the th mapping feature, represents a convolution-nonlinear composite operator of the th layer, represents the th mapping feature, represents a layer number; the deep structure engine model constructing a topography, a normal force, a transverse shear force and a longitudinal shear force of the test object according to the final mapping features: ; ; ; ; wherein, represents a topography, represents a normal force, represents a transverse shear force field, represents a longitudinal shear force field, represents a width of a mapping feature of the topography, represents a width of a mapping feature of the normal force, represents a width of a mapping feature of the transverse shear force field, represents a width of a mapping feature of the longitudinal shear force field, represents a non-negative activation satisfying a physical prior, represents a final mapping feature.
[0012] Optionally, the end-to-end visual-haptic perception method based on a topography-force field analysis model, wherein the consistency constraint is performed on all the shear forces, a partial differential equation residual is constructed by using the topography and the normal force, and the deep structure engine model is optimized and trained by using the partial differential equation residual, and specifically includes: inputting the transverse shear force field and the longitudinal shear force field into a physical law field model, the physical law field model respectively performing consistency description on the transverse shear force field and the longitudinal shear force field, and outputting a transverse partial derivative and a longitudinal partial derivative: ; ; wherein, denotes the lateral derivative, denotes the longitudinal derivative, and both denote the partial derivative, denotes the second order partial derivative combination of the spatial horizontal coordinate, denotes the Laplace operator, denotes the displacement constraint field of the test object; constructing a partial differential equation residual according to the topography, the normal force, the lateral derivative and the longitudinal derivative under different resolutions: ; wherein, denotes the partial differential equation residual, s denotes the index of different resolutions, S denotes the scale set, denotes the weight under the resolution s, denotes the residual functional, denotes the topography under the resolution s, denotes the updated normal force under the resolution s, denotes the lateral derivative under the resolution s, denotes the longitudinal derivative under the resolution s, denotes the square of the modulus of ; optimizing and training the deep structure engine model using the partial differential equation residual to obtain a target deep structure engine model.
[0013] Optionally, the end-to-end visiotactile perception method based on the topography-force field analytical model, wherein the trained deep structure engine model is used to analyze the target object, and the target topography, the target normal force and a plurality of target shear forces of the target object are output, specifically comprising: obtaining a target image of the target object, and determining a target deflection field of the target object according to the target image; inputting the target deflection field into the force field analytical model to output a target normal force, a target lateral shear force field and a target longitudinal shear force field of the target object; constructing a mapping relationship between the target deflection field, the target load, the target normal force, the target lateral shear force field and the target longitudinal shear force field and the target image to obtain a target mapping image; inputting the target mapping image into the target deep structure engine model for analysis to output the target topography, the target normal force, the target lateral shear force and the target longitudinal shear force of the target object.
[0014] In addition, to achieve the above object, the present application also provides an end-to-end visual-haptic perception system based on a topography-force field analytical model, wherein the end-to-end visual-haptic perception system based on the topography-force field analytical model comprises: a data enhancement module configured to obtain a deformation image of a test object, pre-process the deformation image to obtain an enhanced image, and obtain a deflection field of the test object according to the enhanced image; a force field analytical model configured to input the deflection field into the constructed force field analytical model, and generate a load, a normal force and a plurality of shear force fields of the test object according to the deflection field; a synchronous calculation module configured to input the deformation image, the deflection field, the load, the normal force and all the shear force fields into a deep construction engine model to establish a mapping relationship, and construct topography, normal force and a plurality of shear forces and output according to the mapping relationship; a model training module configured to perform consistency constraint on all the shear forces, construct partial differential equation residuals by using the topography and the normal force, and perform optimization training on the deep construction engine model by using the partial differential equation residuals; a haptic analysis module configured to analyze a target object by using the trained deep construction engine model, and output target topography, target normal force and a plurality of target shear forces of the target object.
[0015] In addition, to achieve the above object, the present application also provides a terminal, wherein the terminal comprises a memory, a processor, and a topography-force field analytical model based end-to-end visual-haptic perception program stored in the memory and executable on the processor, and the topography-force field analytical model based end-to-end visual-haptic perception program implements the steps of the topography-force field analytical model based end-to-end visual-haptic perception method when executed by the processor.
[0016] In addition, to achieve the above object, the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a topography-force field analytical model based end-to-end visual-haptic perception program, and the topography-force field analytical model based end-to-end visual-haptic perception program implements the steps of the topography-force field analytical model based end-to-end visual-haptic perception method when executed by a processor.
[0017] In the present application, the deformation image of the test object is obtained, the deformation image is preprocessed to obtain an enhanced image, and the deflection field of the test object is obtained according to the enhanced image; the deflection field is input into the constructed force field analysis model, and the force field analysis model generates the load, normal and multiple shear force fields of the test object according to the deflection field; the deformation image, the deflection field, the load, the normal and all the shear force fields are input into the deep structure engine model to establish a mapping relationship, and the topography, normal force and multiple shear force are constructed and output according to the mapping relationship; all the shear forces are subjected to consistency constraint, and the partial differential equation residual is constructed by using the topography and the normal force, and the deep structure engine model is optimized and trained by using the partial differential equation residual; the trained deep structure engine model is used to analyze the target object, and the target topography, target normal force and multiple target shear force of the target object are output. The present application realizes micron-level topography analysis and multi-axis force synchronous estimation, has real-time robust performance under high-frequency dynamic tasks, reduces the cost of tactile perception, and improves the accuracy of tactile perception results. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 is a flowchart of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analysis model of the present application; Figure 2 is a hardware schematic diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analysis model of the present application; Figure 3 is a data making flowchart of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analysis model of the present application; Figure 4 is a stress calculation flowchart of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analysis model of the present application; Figure 5 is an end-to-end result-topography 2D comparison diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analysis model of the present application; Figure 6 is an end-to-end result-topography 3D comparison diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analysis model of the present application; Figure 7 is an end-to-end result-force field 2D comparison diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analysis model of the present application; Figure 8 is an end-to-end result-force field 3D comparison diagram of a preferred embodiment of the end-to-end visual-tactile perception method based on the topography-force field analysis model of the present application; Figure 9is a structural diagram of a preferred embodiment of an end-to-end visual-haptic perception system based on a topography-force field analysis model of the present application; Figure 10 is a structural diagram of a preferred embodiment of a terminal of the present application. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions and advantages of the present application clearer and more explicit, the present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0020] Visual-haptic sensors achieve sub-millimeter or even tens of micrometer level topography resolution by optically imaging the deformation of flexible films or gel surfaces, and have the potential to simultaneously acquire contact morphology and mechanical information, and have become an important development direction of intelligent haptic perception. However, the existing visual-haptic sensing methods still have core limitations. For example: the visual-haptic sensing method based on tracking of marker points or speckles has poor real-time performance, and the image registration and optical flow calculation are time-consuming, which makes it difficult to support high frame rate applications; the stability is insufficient, and the marker point density and imaging conditions are highly dependent, and occlusion or noise can easily cause distortion of the results. The mechanical estimation accuracy of the topography reconstruction method based on multi-light source coding is limited, and the results are highly dependent on material parameters and boundary conditions, and are sensitive to deviations; multi-axis force estimation is insufficient, and the shear force detection accuracy is limited, and the direction and amplitude are prone to errors; the real-time performance is insufficient, and the computational complexity is high, which is not conducive to high-speed or dynamic operation tasks. The physical constraints are missing in the end-to-end prediction method based on deep learning, and the results are prone to drift and distortion, and lack of mechanical consistency; the measurement accuracy is limited, and it is difficult to distinguish subtle force changes less than 0.1N under low load conditions; the function is limited, and most of them can only predict a single modality (such as topography or normal force), and it is difficult to simultaneously output multi-axis mechanical information.
[0021] Based on the above reasons, the present application proposes a high-resolution visual-haptic perception system based on an end-to-end topography-force analysis neural network. The core idea is to combine "high-precision dataset making-topography reconstruction and force field calculation" and "end-to-end neural network with physical constraints fusion (GeoMech-Net, the deep construction engine model proposed in the present application)", to realize collaborative high-precision inference of topography and force field.
[0022] The end-to-end visual-haptic perception method based on a topography-force field analysis model according to the preferred embodiment of the present application, as shown in Figure 1 The end-to-end visual-haptic perception method based on a topography-force field analysis model includes the following steps: Step S10, obtain the deformation image of the test object, pre-process the deformation image to obtain an enhanced image, and obtain the deflection field of the test object according to the enhanced image.
[0023] First, in the hardware part, the morphological reconstruction module is used to collect the deformation of the test object in the form of an image. Figure 2 As shown in the figure, the morphological reconstruction model includes a flexible sensing layer, a multi-wavelength light source and an image acquisition model; wherein F represents the force applied by the test object when touching the flexible sensing layer.
[0024] The multi-wavelength light source scatters the light source onto the flexible sensing layer through a light barrier, a reflecting lens and a scattering plate; the image acquisition model captures the light source scattered onto the flexible sensing layer to collect the deformation of the flexible sensing layer after being contacted by the test object, and obtains a deformation image. When the test object contacts the flexible sensing layer, the flexible sensing layer will be subjected to an acting force, and the multi-wavelength light source is irradiated to shoot the deformation image of the flexible sensing layer, simulating the deformation image of the test image.
[0025] In one embodiment of the present application, the flexible sensing layer and the sensor shell form a sealed space, which can effectively shield the influence of environmental light and reduce the background error in the image; the multi-wavelength light source and the image acquisition module are integrated on the same PCB board, reducing the volume and improving the portability of the system; and the flexible sensing layer is mixed uniformly according to a certain proportion of silica gel, aluminum alloy powder and fumed silica, and is solidified to realize internal strong diffuse reflection and external dry wear resistance, providing a basis for sensing external deformation and force. When irradiation is performed, a light barrier is used to isolate the multi-wavelength light source to prevent the multi-wavelength light source from being directly captured by the image acquisition module without reflection; the multi-wavelength light source is scattered onto the flexible sensing layer through the light guide plate, the reflecting lens and the scattering acrylic plate, so that the deformation of the flexible sensing layer is more easily identified, providing a basis for subsequent processing.
[0026] Further, the size of the film can be adjusted to realize accurate perception of different areas, and by adjusting the thickness of the film, relative accurate perception of different ranges can be realized. For the equipment used for shooting, a mobile phone can be directly used for shooting, which significantly reduces the cost.
[0027] For the software part, a pre-trained morphological reconstruction network is used to perform deep reconstruction on the film to finally obtain high-precision depth data. Specifically, the deformation image of the test object is obtained, and the deformation image is input into a data domain shaper in the morphological reconstruction model; the data domain shaper uses multiple parameters to perform transformation on the deformation image to obtain a transformed image: wherein, represents the transformed image, represents an implicit modulation family, represents the deformation image, represents a parameter vector; converting the transformed image into a binary image, and segmenting the binary image to obtain an ROI region with deformation, and integrating all pixels in the ROI region to obtain an enhanced image: ; ; represents a binary mask, represents the height of the binary image, represents the ROI region, represents the ROI region in represents a bias correction term, represents the enhanced image, represents the integration processing of the ROI region, represents a gradient operator, represents the square of the modulus of According to the enhanced image, the contact surface of the test object in the world coordinate system is determined, and the lateral axis deformation and the longitudinal axis deformation of the test object in the world coordinate system are determined according to the contact surface; according to the lateral axis deformation and the longitudinal axis deformation, the deflection field of the test object is determined.
[0028] Wherein, the ROI region (Region of Interest) is a core concept in image processing and computer vision, which refers to a specific sub-region selected from an image, used to focus on key targets to improve processing efficiency and accuracy; in order to maintain the high resolution of the image, in addition to obtaining the details of the contact surface through the micron level topography reconstruction module, the input image also needs to be implicitly restructured and shaped by the data domain shaper. Through the implicit modulation family to realize the spatial domain modulation, the input image is domain transformed by using the non-display parameterized operator, so as to ensure that the model still has robustness and adaptability under different experimental conditions. Among them, the implicit modulation family is a kind of technology that realizes signal modulation through implicit function or non-explicit parameterization, and its core feature is that the modulation process does not directly change the explicit parameters (such as amplitude, phase) of the carrier, but indirectly affects the signal characteristics through implicit mapping or nonlinear transformation.
[0029] Step S20, input the deflection field into the constructed force field analysis model, and the force field analysis model generates the load, normal and multiple shear force fields of the test object according to the deflection field.
[0030] Wherein, the shear force field includes: lateral axis shear force field and longitudinal axis shear force field; as shown in Figure 3 Based on the enhanced image (i.e. processed image), the deformation of the flexible perception layer in the horizontal and vertical axes of the world coordinate system (i.e.Figure 3 determining a deflection field of the test object, and then performing force field calculation on the test object by a force field analytical model based on the deflection field and a physical mechanics model (i.e., a control equation between the deflection field and the load), improving the multi-dimensional tactile perception capability, and obtaining corresponding absolute depth, related physical parameter calibration results and real force field. The force field analytical model realizes accurate reasoning of the normal force and shear force in the contact area through coupling of the membrane deflection field and the physical mechanics model. This module eliminates the uncertainty of the traditional empirical fitting method in multi-axis force estimation, and ensures that the output result has strict physical consistency.
[0031] Specifically, the deflection field is input into the force field analytical model, and the force field analytical model determines the load of the test object according to the deflection field: ; wherein, represents the load, represents a divergence operator, represents a membrane tension coefficient, represents a gradient operator, represents a deflection field, represents a base bed reaction coefficient, represents a shear base bed parameter, represents a Laplace operator, represents a deformation of the flexible perception layer in the transverse axis after contacting the test object, represents a deformation of the flexible perception layer in the longitudinal axis after contacting the test object; and the load is decomposed to obtain the normal force, the transverse shear force field and the longitudinal shear force field of the test object and output.
[0032] wherein, based on the generated deflection field, the load of the test object can be generated according to the control equation (i.e., the external force acting on the object or structure to generate internal force and deformation). The deflection field output by the topography reconstruction model is input into the force field analytical model, and after numerical calculation by the control equation, the total force acting on the test object (i.e., the total force acting on the flexible perception layer when the test object contacts the flexible perception layer) is obtained. After decomposition, the normal force and shear component (i.e., the transverse and longitudinal shear force fields) can be determined, and according to the two shear force fields, the in-plane friction and tangential action law can be determined.
[0033] The high-precision mapping from the topography to the force field is realized under the premise of guaranteeing the physical constraints through the force field analytical model, and reliable supervision signals and calibration benchmarks are provided for neural network training and actual application. Through the micron-level topography reconstruction module, the details of the contact surface are obtained, and the high-precision load inference in the range of 0-20 N is realized by combining the force field analytical model based on the physical equation, the normal force error is less than 5%, and the subtle force change less than 0.1 N can be distinguished.
[0034] In step S30, the deformation image, the deflection field, the load, the normal force and all the shear force fields are input into a deep construction engine model to establish a mapping relationship, and the topography, the normal force and the plurality of shear forces are constructed according to the mapping relationship and output.
[0035] Based on the decomposition of the total force, the data-driven and physical consistency constraints are combined through the GeoMech-Net neural network (i.e. the deep construction engine model described in the application) to avoid the problems of prediction drift and numerical distortion of the pure deep learning method, and stable and reliable end-to-end calculation is realized.
[0036] Specifically, the deformation image, the deflection field, the load, the normal force, the transverse axis shear force field and the longitudinal axis shear force field are input into a deep construction engine model, the deep construction engine model converts the deflection field, the load, the normal force, the transverse axis shear force field and the longitudinal axis shear force field into corresponding tensor forms respectively, and establishes a mapping relationship between all the tensor forms and the deformation image to obtain final mapping features: ; wherein, represents the index of the number of layers, represents the first mapping feature, represents the convolution-nonlinear composite operator of the first layer, represents the first mapping feature, represents the number of layers; the deep construction engine model constructs the topography, the normal force, the transverse axis shear force and the longitudinal axis shear force of the test object according to the final mapping features: ; ; ; ; wherein, represents the topography, represents the normal force, represents the transverse axis shear force field, a width of a mapped feature representing a longitudinal shear force field, a width of a mapped feature representing a normal force, a width of a mapped feature representing a normal force, a width of a mapped feature representing a transverse shear force field, a width of a mapped feature representing a longitudinal shear force field, a non-negative activation representing a winding satisfying a physical prior, a final mapped feature.
[0037] In the embodiments disclosed in the present application, the deep structure engine model undertakes the task of hierarchical analysis and calculation of high-dimensional features. The deep structure engine model constructs a feature pyramid through layer-by-layer convolution and upsampling, thereby realizing the mapping from the input image to the multi-scale representation. First, the various quantity fields (i.e. the deflection field, the load, the normal force, the transverse shear force field and the longitudinal shear force field) need to be input into the initial deep structure engine model that has been constructed for preliminary optimization, so that the deep structure engine model can identify these quantity fields in the image, thereby constructing the mapping relationship between the image and these quantity fields. Then, the deformed image is input into the deep structure engine model that has been preliminarily optimized, and the deep structure engine model converts these quantity fields in the deformed image into the form of a tensor (with dimensions BxCxHxW, i.e. batch size x feature channel number x feature map height x feature map width). After each layer of feature mapping, the mapping relationship between these feature mappings and the original deformed image is established, and a deformed image with a final mapped feature is obtained. Then, based on the final mapped feature, branch decoupling is performed. This process is implemented on the shared backbone features, as shown in FIG. 1, and finally outputs the topography, the normal force, the transverse shear force and the longitudinal shear force of the test object. Figure 4
[0038] The deep structure engine model can fully capture both the small local texture information contained in the contact area and the large-scale changes of the overall topography through the construction of a multi-scale feature pyramid, thereby forming a comprehensive analysis of complex contact scenarios. On this basis, the deep structure engine model adopts a four-head decoupling design idea, mapping the shared feature space into the topography, the normal force, and the shear force components (the transverse shear force and the longitudinal shear force), respectively, thereby realizing end-to-end complete mechanical field output without the need for additional intermediate processes, avoiding the cascading error accumulation problem of the traditional method of "topography first, then mechanics". At the same time, the model introduces non-negative constraints and physical priors that conform to physical laws through subsequent physical law modules at the output end, so that the prediction results are not only more accurate in numerical value, but also have strict consistency and reasonableness in mechanical interpretation, thereby combining high precision and high interpretability.
[0039] Step S40, consistent constraints are performed on all the shear forces, a partial differential equation residual is constructed by using the topography and the normal force, and the deep structure engine model is optimized and trained by using the partial differential equation residual.
[0040] In the embodiments of the present disclosure, unlike the learning route that simply relies on data supervision, a physical constraint term is directly embedded in the training process to construct a differentiable PDE residual and a boundary soft constraint, which are aggregated under multiple scales and are interchangeably combined with stable boundary processing (such as mirror or DCT equivalent, finite difference, etc.), thereby ensuring numerical stability and boundary physics, and maintaining the ability to adapt to complex contact geometry.
[0041] Specifically, the horizontal shear force field and the vertical shear force field are input into a physical law field model, the physical law field model performs consistent description on the horizontal shear force field and the updated vertical shear force field respectively, and outputs horizontal partial derivatives and vertical partial derivatives: ; ; wherein, represents the horizontal partial derivative, represents the vertical partial derivative, and both represent partial derivatives, represents a combination of second-order partial derivatives of spatial horizontal coordinates, represents a Laplacian operator, represents a displacement constraint field of a test object; a partial differential equation residual is constructed according to the topography, the normal force, the horizontal partial derivative and the vertical partial derivative under different resolutions: ; wherein, represents the partial differential equation residual, s represents an index of different resolutions, and S represents a scale set, represents a weight under resolution s, represents a residual functional, represents the topography under resolution s, represents the updated normal force under resolution s, represents the horizontal partial derivative under resolution s, represents the vertical partial derivative under resolution s, represents a square of a modulus of ; the deep structure engine model is optimized and trained by using the partial differential equation residual to obtain a target deep structure engine model.
[0042] The physical law field model is based on multiple quantity fields output by the deep construction engine model, constructs consistency thereof, and uniformly depicts tangential force components to obtain transverse partial derivatives and longitudinal partial derivatives. Then, based on different resolutions of the background, multi-scale constraints are performed. Finally, the calculated PDE residual is taken as a loss term of the deep construction engine model for physical constraint.
[0043] By converting the physical prior into a unified optimizable target and fusing it with the representation learning, the model can still maintain the overall consistent convergence of the topography-force field in a weakly labeled or domain transfer scenario, improve the boundary stability, significantly suppress edge artifacts and leakage, and balance the tension anisotropy, base parameters and high-order regular terms. The material, thickness or pixel size difference can also be absorbed into the physical term and scale strategy, reducing the burden of re-calibration.
[0044] In step S50, the trained deep construction engine model is used to analyze the target object, and the target topography, target normal force and multiple target shear forces of the target object are output.
[0045] Specifically, a target image of the target object is obtained, and a target deflection field of the target object is determined according to the target image; the target deflection field is input into the force field analysis model to output a target normal force, a target transverse shear force field and a target longitudinal shear force field of the target object; a mapping relationship between the target deflection field, the target load, the target normal force, the target transverse shear force field and the target longitudinal shear force field and the target image is constructed to obtain a target mapping image; the target mapping image is input into the target deep construction engine model for analysis to output the target topography, the target normal force, the target transverse shear force and the target longitudinal shear force of the target object.
[0046] Wherein, based on the trained deep construction engine model, the target image is input into the topography reconstruction model and the force field analysis model in the same way as the above steps to obtain a target mapping image with a mapping relationship, and then the trained deep construction engine model is input to realize the analysis of the contact surface of the target object to obtain the target topography, the target normal force, the target transverse shear force and the target longitudinal shear force. As shown in Figure 5 and Figure 6 , Figure 5 (a) in (a) and Figure 6 in (a) are input target images, Figure 5 (b) in (b) is the predicted two-dimensional depth of the target object, Figure 6 (b) in (b) is the predicted three-dimensional depth of the target object; Figure 5 (c) in (c) is the two-dimensional depth of the real target object, Figure 6 (c) in (c) is the three-dimensional depth of the real target object; Figure 5In the figure, (d) represents the error between the true 2D depth and the predicted 2D depth. Figure 6 In the figure, (d) represents the error between the actual 3D depth and the predicted 3D depth.
[0047] Furthermore, such as Figure 7 and Figure 8 As shown, Figure 7 (a) and Figure 8 In the examples (a), each image represents the input target image. Figure 7 In (b), the predicted two-dimensional load on the target object is shown. Figure 8 (b) in the figure represents the predicted three-dimensional load of the target object; Figure 7 In the diagram, (c) represents the two-dimensional load of the actual target object. Figure 8 (c) in the figure represents the three-dimensional load of the actual target object; Figure 7 In the figure, (d) represents the error between the actual two-dimensional load and the predicted two-dimensional load. Figure 8 In the equation (d), the error between the actual three-dimensional load and the predicted three-dimensional load is represented.
[0048] The force field analytical model disclosed in this invention differs from traditional force field analysis, which only provides an estimate of the normal force. This model can simultaneously output the distribution of the normal force and the in-plane shear force, accurately providing their direction and amplitude, thus enhancing multi-dimensional tactile perception capabilities. The proposed deep engine model combines data-driven approaches with physical consistency constraints, avoiding the prediction drift and numerical distortion problems of purely deep learning methods, achieving stable and reliable end-to-end solutions. Furthermore, the image-enhanced preprocessing ensures that the results maintain high resolution and accuracy while possessing real-time inference capabilities, adapting to different contact objects and complex operating conditions, and maintaining stable output even in dynamic tasks. The system architecture is compatible with various sensors and execution platforms, suitable for scenarios such as precision robot operation, flexible manufacturing inspection, medical rehabilitation, and virtual interaction, possessing broad industrial application prospects.
[0049] This invention achieves micron-level morphology analysis and simultaneous multi-axis force estimation, and possesses real-time robust performance under high-frequency dynamic tasks, reducing the cost of tactile perception while improving the accuracy of tactile perception results.
[0050] Furthermore, such as Figure 9 As shown, based on the above-mentioned end-to-end visual-tactile perception method based on the shape-force field analytical model, the present invention also provides an end-to-end visual-tactile perception system based on the shape-force field analytical model, wherein the end-to-end visual-tactile perception system based on the shape-force field analytical model includes: The data augmentation module 51 is used to acquire the deformation image of the test object, preprocess the deformation image to obtain an augmented image, and obtain the deflection field of the test object based on the augmented image. a force field analytical model 52, configured to input the deflection field into a constructed force field analytical model, the force field analytical model generating a load, a normal and a plurality of shear force fields of the test object according to the deflection field; a synchronization solving module 53, configured to input the deformation image, the deflection field, the load, the normal and all the shear force fields into a deep construction engine model to establish a mapping relationship, and construct a topography, a normal force and a plurality of shear forces according to the mapping relationship and output; a model training module 54, configured to perform consistency constraint on all the shear forces, construct a partial differential equation residual by using the topography and the normal force, and optimize and train the deep construction engine model by using the partial differential equation residual; a haptic analysis module 55, configured to analyze a target object by using the trained deep construction engine model, and output a target topography, a target normal force and a plurality of target shear forces of the target object.
[0051] Further, as shown in Figure 10 Based on the above-mentioned end-to-end visual-haptic perception method and system based on a topography-force field analytical model, the application further provides a terminal, which comprises a processor 10, a memory 20 and a display 30. Figure 10 Only some components of the terminal are shown, but it should be understood that all the shown components are not required, and more or less components can be alternatively implemented.
[0052] The memory 20 can be an internal storage unit of the terminal in some embodiments, for example, a hard disk or a memory of the terminal. The memory 20 can also be an external storage device of the terminal in other embodiments, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 20 can include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software and various data installed on the terminal, for example, program codes of the terminal, etc. The memory 20 can also be used to temporarily store data that has been output or will be output. In an embodiment, the memory 20 stores an end-to-end visual-haptic perception program 40 based on a topography-force field analytical model, which can be executed by the processor 10, so as to realize the end-to-end visual-haptic perception method based on a topography-force field analytical model in the application.
[0053] The processor 10 can be a Central Processing Unit (CPU), a microprocessor or other data processing chip in some embodiments, for running program codes stored in the memory 20 or processing data, such as executing the end-to-end visual-haptic perception method based on the topography-force field analytical model.
[0054] The display 30 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, etc. in some embodiments. The display 30 is used to display information of the terminal and to display a visualized user interface. The components of the terminal communicate with each other through a system bus.
[0055] In an embodiment, when the processor 10 executes the end-to-end visual-haptic perception program 40 based on the topography-force field analytical model in the memory 20, the steps of the end-to-end visual-haptic perception method based on the topography-force field analytical model as described above are implemented.
[0056] The application also provides a computer readable storage medium, wherein the computer readable storage medium stores an end-to-end visual-haptic perception program based on the topography-force field analytical model, and the end-to-end visual-haptic perception program based on the topography-force field analytical model implements the steps of the end-to-end visual-haptic perception method based on the topography-force field analytical model as described above when executed by a processor.
[0057] In summary, the application provides an end-to-end visual-haptic perception method based on a topography-force field analytical model and related equipment, the method comprising: obtaining a deformation image of a test object, pre-processing the deformation image to obtain an enhanced image, and obtaining a deflection field of the test object according to the enhanced image; inputting the deflection field into a constructed force field analytical model, the force field analytical model generating a load, a normal and a plurality of shear force fields of the test object according to the deflection field; inputting the deformation image, the deflection field, the load, the normal and all the shear force fields into a deep structure engine model to establish a mapping relationship, constructing topography, normal force and a plurality of shear forces according to the mapping relationship and outputting; performing consistency constraint on all the shear forces, constructing a partial differential equation residual using the topography and the normal force, and optimizing and training the deep structure engine model using the partial differential equation residual; analyzing a target object using the trained deep structure engine model, and outputting a target topography, a target normal force and a plurality of target shear forces of the target object. The application realizes micron-level topography analysis and synchronous multi-axis force estimation, has real-time robust performance under high-frequency dynamic tasks, reduces the cost of haptic perception, and improves the accuracy of haptic perception results.
[0058] It should be noted that, in the present document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that processes, methods, articles, or terminals that comprise a list of elements are not limited to those elements, but can also include other elements not expressly listed, or inherent to such processes, methods, articles, or terminals. Without further limitation, an element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or terminal that includes the element.
[0059] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer readable computer readable storage medium, and the program can include the processes of the above-mentioned method embodiments when executed. The computer readable storage medium can be a memory, a magnetic disc, an optical disc, etc.
[0060] It should be understood that the application is not limited to the above examples, and those skilled in the art can make improvements or changes according to the above description, and all these improvements and changes shall fall within the protection scope of the claims of the present application.
Claims
1. An end-to-end haptic perception method based on a topography-force field analytical model, characterized in that, The end-to-end visual-haptic perception method based on the topography-force field analytical model comprises: Obtaining a deformation image of a test object, pre-processing the deformation image to obtain an enhanced image, and obtaining a deflection field of the test object according to the enhanced image; Inputting the deflection field into a constructed force field analytical model, and the force field analytical model generates a load, a normal force and a plurality of shear force fields of the test object according to the deflection field; Inputting the deformation image, the deflection field, the load, the normal force and all the shear force fields into a deep construction engine model to establish a mapping relationship, constructing topography, normal force and a plurality of shear forces according to the mapping relationship and outputting; Conducting consistency constraint on all the shear forces, constructing partial differential equation residuals by using the topography and the normal force, and optimizing and training the deep construction engine model by using the partial differential equation residuals; Analyzing a target object by using the trained deep construction engine model, and outputting a target topography, a target normal force and a plurality of target shear forces of the target object.
2. The end-to-end haptic perception method based on a topography-force field analytical model according to claim 1, wherein, The obtaining of the deformation image of the test object, the pre-processing of the deformation image to obtain the enhanced image, and the obtaining of the deflection field of the test object according to the enhanced image specifically comprises: Obtaining the deformation image of the test object, and inputting the deformation image into a data domain shaper in the topography reconstruction model; The data domain shaper uses a plurality of parameters to transform the deformation image to obtain a transformed image: ; wherein, denotes a transformed image, denotes an implicit modulation family, denotes a morphed image, denotes a parameter vector; Converting the transformed image into a binary image, segmenting the binary image to obtain an ROI region where deformation exists, and integrating all pixels in the ROI region to obtain an enhanced enhanced image: ; ; wherein, denotes a binarization mask, denotes the height of the binarized image, denotes the ROI region, denotes the ROI region in the image, denotes a bias correction term, denotes the enhanced image, denotes the integration of the ROI region, denotes a gradient operator, denotes the square of the modulus of According to the enhanced image, determining a contact surface of the test object in a world coordinate system, and determining a horizontal axis deformation and a vertical axis deformation of the test object in the world coordinate system according to the contact surface; According to the horizontal axis deformation and the vertical axis deformation, determining the deflection field of the test object.
3. The end-to-end haptic perception method based on a topography-force field analytical model according to claim 2, wherein, The topography reconstruction model comprises a flexible perception layer, a multi-wavelength light source and an image acquisition model; The obtaining of the deformation image of the test object specifically comprises: The light source emitted by the multi-wavelength light source is scattered to the flexible perception layer through a light blocking plate, a reflecting lens and a scattering plate; The image acquisition model captures the light source scattered to the flexible perception layer to acquire the deformation of the flexible perception layer after contacting with the test object, and obtains a deformation image.
4. The haptic perception method based on a topography-force field analytical model according to claim 1, wherein, The shear force field comprises a horizontal axis shear force field and a vertical axis shear force field; The inputting of the deflection field into the constructed force field analytical model, and the force field analytical model generating a load, a normal force and a plurality of shear force fields of the test object according to the deflection field specifically comprises: Inputting the deflection field into the constructed force field analytical model, and the force field analytical model determining the load of the test object according to the deflection field; ; wherein, denotes the load, denotes the divergence operator, denotes the membrane tension coefficient, denotes the gradient operator, denotes the deflection field, denotes the base reaction coefficient, denotes the shear base parameter, denotes the Laplace operator, denotes the deformation of the flexible sensing layer in the lateral axis after the flexible sensing layer contacts with the test object, denotes the deformation of the flexible sensing layer in the longitudinal axis after the flexible sensing layer contacts with the test object. Decomposing the load to obtain the normal force, the horizontal axis shear force field and the vertical axis shear force field of the test object and outputting.
5. The haptic perception method based on the analytical model of topography-force field according to claim 4, wherein, The deformation image, the deflection field, the load, the normal force and all the shear force fields are input into a deep construction engine model to establish a mapping relationship, and the topography, normal force and multiple shear forces are constructed and output according to the mapping relationship, specifically including: The deformation image, the deflection field, the load, the normal force, the transverse shear force field and the longitudinal shear force field are input into a deep construction engine model, and the deflection field, the load, the normal force, the transverse shear force field and the longitudinal shear force field are respectively converted into corresponding tensor forms by the deep construction engine model, and a mapping relationship between all the tensor forms and the deformation image is constructed to obtain final mapping features: ; wherein, an index representing the number of levels, an index representing the th mapped feature, an index representing the th convolution-nonlinear composite operator, an index representing the th mapped feature, an index representing the number of levels; The deep construction engine model constructs the topography, normal force, transverse shear force and longitudinal shear force of the test object according to the final mapping features: ; ; ; ; wherein, represents a topography, represents a normal force, represents a transverse shear force field, represents a longitudinal shear force field, represents a width of a mapped feature of the topography, represents a width of a mapped feature of the normal force, represents a width of a mapped feature of the transverse shear force field, represents a width of a mapped feature of the longitudinal shear force field, represents a non-negative activation that the volume satisfies a physical prior, represents a final mapped feature.
6. The haptic-aware end-to-end rendering method based on a haptic-force field analytical model according to claim 1, wherein, The consistency of all the shear forces is constrained, and a partial differential equation residual is constructed using the topography and the normal force, and the deep construction engine model is optimized and trained using the partial differential equation residual, specifically including: The transverse shear force field and the longitudinal shear force field are input into a physical law field model, and the transverse shear force field and the longitudinal shear force field are respectively described by the physical law field model, and the transverse partial derivative and the longitudinal partial derivative are output: ; ; wherein denotes the lateral deflection, denotes the longitudinal deflection, and each denote the partial derivative, denotes the second order partial derivative combination of the spatial abscissa, denotes the Laplace operator, denotes the displacement constraint field of the test object; A partial differential equation residual is constructed according to the topography, the normal force, the transverse partial derivative and the longitudinal partial derivative under different resolutions: ; wherein, represents a partial differential equation residual, s represents an index of different resolutions, S represents a set of scales, represents a weight at resolution s, represents a residual functional, represents a topography at resolution s, represents an updated normal force at resolution s, represents a lateral partial derivative at resolution s, represents a longitudinal partial derivative at resolution s, represents a square of a modulus of The deep construction engine model is optimized and trained using the partial differential equation residual to obtain a target deep construction engine model.
7. The haptic perception method based on the analytical model of topography-force field according to claim 6, wherein, The trained deep construction engine model is used to analyze the target object, and the target topography, target normal force and multiple target shear forces of the target object are output, specifically including: A target image of the target object is obtained, and a target deflection field of the target object is determined according to the target image; The target deflection field is input into the force field analysis model to output the target normal force, target transverse shear force field and target longitudinal shear force field of the target object; A mapping relationship between the target deflection field, target load, target normal force, target transverse shear force field and target longitudinal shear force field and the target image is constructed to obtain a target mapping image; The target mapping image is input into the target deep construction engine model for analysis, and the target topography, target normal force, target transverse shear force and target longitudinal shear force of the target object are output.
8. A topography-force field analytical model based end-to-end visuo-haptic perception system for implementing the topography-force field analytical model based end-to-end visuo-haptic perception method of any one of claims 1-7, characterized in that, The end-to-end visual-haptic perception system based on the topography-force field analysis model includes: A data enhancement module is configured to obtain a deformation image of a test object, pre-process the deformation image to obtain an enhanced image, and obtain a deflection field of the test object according to the enhanced image; A force field analysis model is configured to input the deflection field into a constructed force field analysis model, and the force field analysis model generates a load, a normal force and multiple shear force fields of the test object according to the deflection field; A synchronization solving module is configured to input the deformation image, the deflection field, the load, the normal and all the shear force fields into a deep construction engine model to establish a mapping relationship, construct a topography, a normal force and multiple shear forces according to the mapping relationship and output the topography, the normal force and the multiple shear forces; A model training module is configured to perform consistency constraint on all the shear forces, construct a partial differential equation residual by using the topography and the normal force, and perform optimization training on the deep construction engine model by using the partial differential equation residual; A haptic analysis module is configured to analyze a target object by using the trained deep construction engine model and output a target topography, a target normal force and multiple target shear forces of the target object.
9. A terminal, characterized by comprising: The terminal comprises a memory, a processor and an end-to-end visual-haptic perception program based on a topography-force field analysis model stored on the memory and executable on the processor, and the end-to-end visual-haptic perception program based on the topography-force field analysis model implements the steps of the end-to-end visual-haptic perception method based on the topography-force field analysis model according to any one of claims 1-7 when executed by the processor.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores an end-to-end visual-haptic perception program based on a topography-force field analysis model, and the end-to-end visual-haptic perception program based on the topography-force field analysis model implements the steps of the end-to-end visual-haptic perception method based on the topography-force field analysis model according to any one of claims 1-7 when executed by the processor.
Citation Information
Patent Citations
Visual tactile sensor three-dimensional contact force field measurement method based on inverse finite element analysis
CN118114524A
Intelligent video monitoring system and method based on deep learning
CN120263940A
End-to-end automatic driving control method and device based on multi-camera fusion
CN120411902A
Cited By
Tactile perception system and method, electronic device, storage medium, and program product
CN121614037A
Haptic perception system and method, electronic device, storage medium, program product
CN121614037B
Intensive normal distribution force field reconstruction method fusing force and visual tactile information
CN122154490A