Visual tactile sensor gradient calibration method and device based on deep learning and medium

Through deep learning segmentation model and gradient recognition model, the gradient information of visual haptic sensors is automatically calculated, which solves the inefficient light intensity to gradient correction problem in the prior art, and achieves efficient and accurate gradient calibration.

CN120255734AActive Publication Date: 2025-07-04NANJING YIMU INTELLIGENT TECHNOLOGY CO LTD +1

Patent Information

Application Number
CN202510704794.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-04
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing visual haptic sensors are inefficient and error-prone in the correction process of light intensity to gradients, requiring professionals to manually mark the sphere position and record data under experimental conditions.

Method used

The deep learning segmentation model is used to identify the spherical contact area, and the gradient information is automatically calculated by training the gradient recognition model, establishing the mapping relationship between RGB light intensity and gradient, and using the pre-trained deep learning segmentation model and gradient recognition model for automatic calibration.

Benefits of technology

The efficiency and accuracy of visual haptic sensor gradient calibration are improved, and an automated gradient calibration process is realized, reducing human error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255734A_ABST
    Figure CN120255734A_ABST
Patent Text Reader

Abstract

The invention discloses a visual tactile sensor gradient calibration method and device based on deep learning and a medium, and the method comprises the steps: obtaining sensor image information of a preset ball pressed on a target touch film of a visual tactile sensor to be calibrated, inputting the sensor image information into a pre-trained deep learning segmentation model, obtaining a ball contact area of the target touch film; coordinate values and RGB values of all points in the ball contact area are determined, and gradient information of all the points in the ball contact area is calculated according to the radius of a preset ball; constructing a training data set according to the coordinate values, the RGB values and the gradient information, and training according to the training data set to obtain a gradient recognition model; and calibrating a mapping relationship between the RGB light intensity of each point on the target touch film and the gradient according to the gradient recognition model. The visual tactile sensor gradient calibration efficiency and accuracy are improved, and the method can be widely applied to the technical field of artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device and medium for calibrating the gradient of a visual tactile sensor based on deep learning. Background Art

[0002] The basic principle of a visual tactile sensor represented by Gelsight is to calculate the height gradient generated by a point in two directions on a plane according to the amount of reflected light of the light of different colors in the RGB three directions at this point. The depth data of the surface can be obtained by integrating the gradient (solving the Poisson equation). The basic principle of the design of this type of sensor is that light is diffusely reflected on the surface of the touch film. When the film undulates, the incident angle of the light during diffuse reflection on the film surface changes. The surface normal vector of the film at this point, that is, the gradient direction, can be inferred according to the change in the intensity of the light captured at this point. Therefore, the calibration of the non-linear relationship from light intensity to gradient is crucial for the visual tactile sensor.

[0003] Since there are assembly errors in the positions of the three color light sources and the position of the film during the production and manufacturing of sensors such as Gelsight, and the reflectivity parameters of the films used in each sensor cannot be completely consistent, therefore, before leaving the factory or before the user uses each sensor, the calibration process from light intensity to gradient is essential.

[0004] Currently, for the method of the calibration process from light intensity to gradient, generally, the film is pressed on experimental molds with different gradients under a single experimental condition, and the light intensity situation is observed to establish the mapping relationship from light intensity to gradient. This method requires professionals to complete it under good experimental conditions, and the completion process is relatively complex. A method provided by Gelsight simplifies this experimental mold, that is, a sphere with a known radius is used as the experimental mold. According to the radius of the sphere, the gradient of the points on the surface of the sphere can be solved. However, this method still requires professionals to mark the position of the sphere and record data, with low efficiency and easy to make mistakes. Summary of the Invention

[0005] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0006] To this end, an object of an embodiment of the present invention is to provide a method for calibrating the gradient of a visual tactile sensor based on deep learning, which improves the efficiency and accuracy of calibrating the gradient of the visual tactile sensor.

[0007] Another object of an embodiment of the present invention is to provide a device for calibrating the gradient of a visual tactile sensor based on deep learning.

[0008] In order to achieve the above technical objectives, the technical solutions adopted in the embodiments of the present invention include: On the one hand, an embodiment of the present invention provides a method for calibrating the gradient of a visual tactile sensor based on deep learning, including the following steps: Obtain sensor image information of a preset spherical ball pressing on a target touch film of the tactile sensor to be calibrated, and input the sensor image information into a pre-trained deep learning segmentation model to obtain the spherical ball contact area of the target touch film; Determine the coordinate values and RGB values of each point in the spherical ball contact area, and calculate the gradient information of each point in the spherical ball contact area according to the radius of the preset spherical ball; Construct a training data set according to the coordinate values, the RGB values, and the gradient information, and train a gradient recognition model according to the training data set; Calibrate the mapping relationship between the RGB light intensity and the gradient of each point on the target touch film according to the gradient recognition model.

[0009] Further, in an embodiment of the present invention, the deep learning segmentation model is trained through the following steps: Obtain a plurality of sensor sample images of the preset spherical ball pressing on different positions of the sample touch film with different forces; Perform segmentation annotation on the sensor sample images to determine the contact area labels corresponding to the sensor sample images; Input the sensor sample images into a pre-constructed deep learning neural network to obtain a predicted contact area; Determine a first loss value according to the predicted contact area and the contact area label; Update the parameters of the deep learning neural network according to the first loss value to obtain the trained deep learning segmentation model.

[0010] Further, in an embodiment of the present invention, the calculation of the gradient information of each point in the spherical ball contact area according to the radius of the preset spherical ball specifically includes: Determine the actual bending radius according to the radius of the preset spherical ball and the thickness of the target touch film; Fit the spherical ball contact area to obtain a fitted circular area, and determine the center position of the fitted circular area; Calculate the gradient information of each point in the spherical ball contact area according to the center position and the actual bending radius.

[0011] Further, in an embodiment of the present invention, the gradient information is calculated by the following formula:

[0012]

[0013] Among them, and respectively represent the gradient values of the point (x, y) in the x - direction and y - direction, (x0, y0) represents the position of the center of the circle, and R represents the actual bending radius.

[0014] Furthermore, in an embodiment of the present invention, constructing a training data set according to the coordinate values, the RGB values, and the gradient information, and training a gradient recognition model according to the training data set specifically includes: Determining the x - coordinate and y - coordinate according to the coordinate values, determining the R value, G value, and B value according to the RGB values, and determining a training sample according to the R value, the G value, the B value, the x - coordinate, and the y - coordinate; Determining the gradient value in the x - direction and the gradient value in the y - direction according to the gradient information, and determining the gradient label corresponding to the training sample according to the gradient value in the x - direction and the gradient value in the y - direction; Inputting the training sample into a pre - constructed ANN neural network to obtain a gradient recognition result; Determining a second loss value according to the gradient recognition result and the gradient label; Updating the parameters of the ANN neural network according to the second loss value to obtain the trained gradient recognition model.

[0015] Furthermore, in an embodiment of the present invention, the ANN neural network includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, a fourth hidden layer, and an output layer. The input layer is used to input the training sample, the output layer is used to output the gradient recognition result, and the first hidden layer, the second hidden layer, the third hidden layer, and the fourth hidden layer respectively include 64, 128, 128, and 64 neurons.

[0016] Furthermore, in an embodiment of the present invention, calibrating the mapping relationship between the RGB light intensity and the gradient of each point on the target touch film according to the gradient recognition model specifically includes: Dividing the target touch film into multiple grid points according to a preset spacing; Determining the gradient of each grid point under different RGB light intensities according to the gradient recognition model, and establishing a mapping relationship according to the coordinate position, RGB light intensity, and corresponding gradient of the grid point; Generating a three - dimensional look - up table containing coordinate values, RGB values, and gradient values according to the mapping relationship.

[0017] On the other hand, an embodiment of the present invention provides a gradient calibration device for a visual - tactile sensor based on deep learning, including: A contact area recognition module, configured to obtain sensor image information of a preset spherical ball pressing on a target touch film of a to-be-calibrated visual tactile sensor, input the sensor image information into a pre-trained deep learning segmentation model, and obtain a spherical ball contact area of the target touch film; A gradient information calculation module, configured to determine coordinate values and RGB values of each point within the spherical ball contact area, and calculate gradient information of each point within the spherical ball contact area according to the radius of the preset spherical ball; A model training module, configured to construct a training data set according to the coordinate values, the RGB values, and the gradient information, and train a gradient recognition model according to the training data set; A calibration module, configured to calibrate a mapping relationship between the RGB light intensity and the gradient of each point on the target touch film according to the gradient recognition model.

[0018] On the other hand, an embodiment of the present invention provides an electronic device, which includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the method for calibrating the gradient of a visual tactile sensor based on deep learning as described above.

[0019] On the other hand, an embodiment of the present invention further provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method for calibrating the gradient of a visual tactile sensor based on deep learning as described above.

[0020] The advantages and beneficial effects of the present invention will be partially given in the following description, partially will become obvious from the following description, or can be understood through the practice of the present invention: In an embodiment of the present invention, sensor image information of a preset spherical ball pressed on a target touch film of a tactile sensor to be calibrated is obtained, and the sensor image information is input into a pre-trained deep learning segmentation model to obtain a spherical contact area of the target touch film. Coordinate values and RGB values of each point within the spherical contact area are determined, and gradient information of each point within the spherical contact area is calculated according to the radius of the preset spherical ball. A training data set is constructed based on the coordinate values, RGB values, and gradient information, and a gradient recognition model is trained according to the training data set. The mapping relationship between the RGB light intensity and the gradient of each point on the target touch film is calibrated according to the gradient recognition model. In the embodiment of the present invention, a pre-trained deep learning segmentation model is used to identify the spherical contact area, a training data set is constructed based on the coordinate values, RGB values, and gradient information of each point within the spherical contact area, a gradient recognition model is trained, and the gradient calibration of each point on the target touch film is completed by using the gradient recognition model, improving the efficiency and accuracy of the gradient calibration of the tactile sensor. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention, and those skilled in the art can obtain other drawings according to these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of steps of a method for calibrating the gradient of a tactile sensor based on deep learning provided by an embodiment of the present invention; Figure 2 It is a flowchart of steps of training a deep learning segmentation model provided by an embodiment of the present invention; Figure 3 It is a flowchart of steps of step S102 provided by an embodiment of the present invention; Figure 4 It is a flowchart of steps of step S103 provided by an embodiment of the present invention; Figure 5 It is a flowchart of steps of step S104 provided by an embodiment of the present invention; Figure 6 It is a schematic structural diagram of a device for calibrating the gradient of a tactile sensor based on deep learning provided by an embodiment of the present invention; Figure 7 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention; Figure 8 It is a schematic structural diagram of a storage medium provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as limiting the present application. It should be noted that although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division from that in the system schematic diagram or a different order from that in the flowchart. For the step numbers in the following embodiments, they are only set for the convenience of explanation and do not limit the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0024] In the description of the present invention, "a plurality" means two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0025] The method for calibrating the gradient of a visual tactile sensor based on deep learning provided by the embodiments of the present application can be applied to a terminal, or to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a set-top box, etc.; the server can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the method for calibrating the gradient of a visual tactile sensor based on deep learning, etc., but is not limited to the above forms.

[0026] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0027] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of relevant countries and regions. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.

[0028] The process of establishing the mapping relationship of the visual-tactile sensor from RGB light intensity to gradient, that is, the calibration process of RGB light intensity to gradient. Currently, the mainstream method is to collect experimental data of RGB light intensity under different gradients in the laboratory and establish the mapping relationship of RGB light intensity to gradient. A method provided by Gelsight uses a simple sphere to simplify the experimental tool, which improves the calibration efficiency to a certain extent, but still requires professionals to mark the position of the sphere and record data, and cannot achieve automatic calibration.

[0029] This invention uses deep learning segmentation to identify the specific position where the film presses on the sphere, eliminating the need for experimenters to manually mark the range of the sphere. At the same time, by training a gradient recognition model, the gradient calibration of each point on the touch film can be completed, and automatic gradient calibration of the visual-tactile sensor can be achieved.

[0030] As Figure 1 shown is a step flow chart of a method for calibrating the gradient of a visual-tactile sensor based on deep learning provided by an embodiment of this invention. Referring to Figure 1, an embodiment of the present invention provides a method for calibrating the gradient of a visual tactile sensor based on deep learning, which specifically includes the following steps: S101. Obtain the sensor image information of a preset spherical ball pressing on the target touch film of the tactile sensor to be calibrated, and input the sensor image information into a pre-trained deep learning segmentation model to obtain the spherical ball contact area of the target touch film; S102. Determine the coordinate values and RGB values of each point within the spherical ball contact area, and calculate the gradient information of each point within the spherical ball contact area according to the radius of the preset spherical ball; S103. Construct a training data set based on the coordinate values, RGB values, and gradient information, and train a gradient recognition model according to the training data set; S104. Calibrate the mapping relationship between the RGB light intensity and the gradient of each point on the target touch film according to the gradient recognition model.

[0031] The present invention uses deep learning segmentation to detect the position where the spherical ball presses on the film, obtains the spherical ball contact area, automatically extracts the coordinate values and RGB values of each point in this area, calculates the gradient information of each point in this area, trains a gradient recognition model based on the coordinate values, RGB values, and gradient information, and can complete the calibration of the mapping relationship between the RGB light intensity and the gradient of each point on the touch film through this gradient recognition model. The specific process is as follows: 1) Press the spherical ball on a certain position on the film; 2) Collect the image of the current sensor; 3) For the current image, use the pre-trained deep learning segmentation model to identify the position where the spherical ball contacts, and record this area as the valid area; 4) Within this valid area, calculate the gradients of the height of each point in the x and y directions according to the radius of the spherical ball; 5) Record the gradient information (Gx, Gy) of each point within this valid area, and extract the coordinate values (x, y) and RGB values (r, g, b) of each point based on the image; 6) Change the position where the spherical ball presses on the film, and repeat steps 1) to 5) so that the spherical ball is squeezed at multiple positions on the touch film with multiple different forces; 7) Organize the data collected in the above steps, and record each group of data as (r, g, b, x, y) → (Gx, Gy); 8) Construct a neural network, input five nodes, and output two nodes; 9) Train a neural network model using the data in 7). The trained gradient recognition model can be directly deployed on this visual-tactile sensor or the corresponding host computer, so as to perform the mapping transformation from RGB light intensity to gradient information when this visual-tactile sensor is applied. It is also possible to directly establish the mapping relationship between the RGB light intensity and the gradient of each point on the touch film according to this gradient recognition model, form a mapping table and store it in this visual-tactile sensor or the host computer to complete the gradient calibration of the visual-tactile sensor.

[0032] In the embodiment of the present invention, a pre-trained deep learning segmentation model is used to identify the spherical contact area. A training data set is constructed based on the coordinate values, RGB values, and gradient information of each point within the spherical contact area, and a gradient recognition model is trained. The gradient calibration of each point on the target touch film is completed using this gradient recognition model, improving the efficiency and accuracy of the gradient calibration of the visual-tactile sensor.

[0033] The following further describes the specific implementation process of the embodiment of the present invention with reference to the accompanying drawings.

[0034] As Figure 2 shown, it is a step flow chart for training a deep learning segmentation model provided by an embodiment of the present invention. Referring to Figure 2 , further as an optional implementation manner, the deep learning segmentation model is trained through the following steps: S201. Obtain a plurality of sensor sample images of a preset sphere pressing on different positions of a sample touch film with different forces; S202. Perform segmentation annotation on the sensor sample images to determine the contact area labels corresponding to the respective sensor sample images; S203. Input the sensor sample images into a pre-constructed deep learning neural network to obtain a predicted contact area; S204. Determine a first loss value according to the predicted contact area and the contact area label; S205. Update the parameters of the deep learning neural network according to the first loss value to obtain a trained deep learning segmentation model.

[0035] Specifically, press the sphere on different positions of the sample touch film with different forces, and collect the image data of the sensor; for the collected sample images, use the segmentation annotation method to mark the positions where the sphere is squeezed to obtain the contact area labels; construct a deep learning segmentation model, and use the sample images and the contact area labels to train this segmentation model. This model can use common deep learning segmentation models, such as yolov8-seg.

[0036] After inputting the sample image into the initialized deep learning neural network, the predicted result output by the model, that is, the predicted contact area, can be obtained. The accuracy of the model prediction can be evaluated based on the predicted contact area and the aforementioned contact area label, so as to update the parameters of the model. For a deep learning segmentation model, the accuracy of the model prediction result can be measured by a loss function. The loss function is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of a training data is determined by the label of the single training data and the prediction result of the model for this training data. During actual training, a training data set has many training data. Therefore, generally, a cost function is used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction errors of all training data, which can better measure the prediction effect of the model. For a general machine learning model, based on the aforementioned cost function, plus a regularization term that measures the model complexity, it can be used as the objective function for training. Based on this objective function, the loss value of the entire training data set can be obtained. There are many types of commonly used loss functions. For example, the 0-1 loss function, square loss function, absolute loss function, logarithmic loss function, cross-entropy loss function, etc. can all be used as the loss function of the machine learning model, which will not be elaborated one by one here. In the embodiments of the present invention, any one of the loss functions can be selected to determine the loss value of training. Based on the loss value of training, the parameters of the model are updated using the backpropagation algorithm, and after several rounds of iteration, a trained deep learning segmentation model can be obtained. The specific number of iteration rounds can be preset in advance, or it is considered that the training is completed when the accuracy requirement is met in the test set.

[0037] As Figure 3 shown is a step flowchart of step S102 provided by an embodiment of the present invention. Referring to Figure 3 , further as an optional implementation manner, the gradient information of each point within the spherical contact area is calculated according to the radius of the preset sphere, which specifically includes: S1021. Determine the actual bending radius according to the radius of the preset sphere and the thickness of the target touch film; S1022. Fit the spherical contact area to obtain a fitted circular area, and determine the center position of the fitted circular area; S1023. Calculate the gradient information of each point within the spherical contact area according to the center position and the actual bending radius.

[0038] Further as an optional implementation manner, the gradient information is calculated by the following formula:

[0039]

[0040] Among them, and respectively represent the gradient values of the point (x, y) in the x-direction and y-direction. (x0, y0) represents the center position of the circle, and R represents the actual bending radius.

[0041] Specifically, assuming that the radius of the spherical ball is r and the thickness of the touch film is d, the actual bending radius can be calculated as R = r + d.

[0042] The depth pressed out by the spherical ball is expressed as z = f(x, y). When the center of the contact area of the spherical ball is used as the coordinate origin, then there is:

[0043] The gradients of z in the x- and y-directions are respectively solved as:

[0044]

[0045] It should be noted that the x- and y-coordinates in the above formula are calculated with the center of the contact area of the spherical ball as the coordinate origin. When converting to the image coordinates of the sensor, just substitute x = x - x0 and y = y - y0, where (x0, y0) is the center position of the contact area of the spherical ball. After the contact area of the spherical ball is segmented, the center position can be calculated through circle fitting and will not be elaborated here.

[0046] Such as Figure 4 shown is a step flow chart of step S103 provided by an embodiment of the present invention. Referring to Figure 4 , further as an optional implementation manner, a training data set is constructed according to the coordinate values, RGB values, and gradient information, and a gradient recognition model is trained according to the training data set, which specifically includes: S1031. Determine the x-coordinate and y-coordinate according to the coordinate values, determine the R value, G value, and B value according to the RGB values, and determine the training samples according to the R value, G value, B value, x-coordinate, and y-coordinate; S1032. Determine the gradient value in the x-direction and the gradient value in the y-direction according to the gradient information, and determine the gradient label corresponding to the training sample according to the gradient value in the x-direction and the gradient value in the y-direction; S1033. Input the training samples into a pre-constructed ANN neural network to obtain a gradient recognition result; S1034. Determine the second loss value according to the gradient recognition result and the gradient label; S1035. Update the parameters of the ANN neural network according to the second loss value to obtain a trained gradient recognition model.

[0047] As a further optional implementation, the ANN neural network includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, a fourth hidden layer, and an output layer. The input layer is used to input training samples, and the output layer is used to output gradient recognition results. The first hidden layer, the second hidden layer, the third hidden layer, and the fourth hidden layer include 64, 128, 128, and 64 neurons respectively.

[0048] Specifically, when the positions of the light sources are fixed, that is, when the positions of the red (R), green (G), and blue (B) color light sources are fixed, and after the position of the touch film is fixed, the relationship between the gradient at each point on the film and the RGB light intensity is fixed. This relationship is a fixed non-linear relationship f, that is, G(x, y) = f(r, g, b, x, y). The calculation of the gradient is related to 5 input parameters, that is, the R value, G value, B value of this point under the illumination of three color light sources, and the x coordinate value and y coordinate value of this point. The gradient of this point is the gradient in the x and y directions.

[0049] In the embodiment of the present invention, the input layer of the ANN neural network has 5 nodes, which respectively input the R value, G value, B value, x coordinate, and y coordinate, that is, (r, g, b, x, y). The output layer has 2 nodes, which are the gradient Gx in the x direction and the gradient Gy in the y direction respectively. There are four hidden layers in the middle, and the number of nodes in the four hidden layers is 64, 128, 128, and 64 respectively. During the training process, the loss function adopts the MSE loss function, and the parameters of the ANN neural network are updated through the backpropagation algorithm until the model converges.

[0050] As Figure 5 shown is another step flow chart of step S104 provided by the embodiment of the present invention. Referring to Figure 5 , as a further optional implementation, according to the gradient recognition model, calibrate the mapping relationship between the RGB light intensity and the gradient of each point on the target touch film, which specifically includes: S1041. Divide the target touch film into multiple grid points according to a preset spacing; S1042. Determine the gradients of each grid point under different RGB light intensities according to the gradient recognition model, and establish a mapping relationship based on the coordinate position, RGB light intensity, and corresponding gradient of the grid point; S1043. Generate a three-dimensional look-up table containing coordinate values, RGB values, and gradient values according to the mapping relationship.

[0051] Specifically, divide the touch film into grid points with a spacing of 0.1 mm. For each grid point, determine its gradient value under different RGB light intensities through the gradient recognition model, so as to establish a mapping relationship between the coordinate position, RGB light intensity, and gradient, and finally generate a three-dimensional look-up table of coordinate values, RGB values, and gradient values, then the gradient calibration of the visual tactile sensor can be completed.

[0052] After the tactile sensor acquires the touch image, the gradient information can be recognized through the gradient recognition model, or the gradient information can be found through the three-dimensional look-up table generated by calibration, and then the integral method is used to obtain the depth data. The depth data can be used for other applications, such as calculating the pressure according to the elastic characteristics of the membrane, which will not be elaborated in the embodiments of the present invention.

[0053] The method steps of the embodiments of the present invention are described above. It can be realized that the embodiments of the present invention use a pre-trained deep learning segmentation model to identify the spherical contact area, construct a training data set based on the coordinate values, RGB values, and gradient information of each point in the spherical contact area, train to obtain a gradient recognition model, and use the gradient recognition model to complete the gradient calibration of each point on the target touch membrane, improving the efficiency and accuracy of the gradient calibration of the tactile sensor.

[0054] As Figure 6 shown is a schematic structural diagram of a gradient calibration device for a tactile sensor based on deep learning provided by an embodiment of the present invention. Referring to Figure 6 , an embodiment of the present invention provides a gradient calibration device for a tactile sensor based on deep learning, including: A contact area recognition module, configured to obtain sensor image information of a preset sphere pressing on a target touch membrane of a tactile sensor to be calibrated, and input the sensor image information into a pre-trained deep learning segmentation model to obtain the spherical contact area of the target touch membrane; A gradient information calculation module, configured to determine the coordinate values and RGB values of each point in the spherical contact area, and calculate the gradient information of each point in the spherical contact area according to the radius of the preset sphere; A model training module, configured to construct a training data set according to the coordinate values, RGB values, and gradient information, and train to obtain a gradient recognition model according to the training data set; A calibration module, configured to calibrate the mapping relationship between the RGB light intensity and the gradient of each point on the target touch membrane according to the gradient recognition model.

[0055] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0056] An embodiment of the present invention also provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the above-mentioned gradient calibration method for a tactile sensor based on deep learning. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0057] As shown Figure 7 in the following figure is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Referring to Figure 7 , the embodiment of the present invention provides an electronic device, including: A processor 701, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention; A memory 702, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702 and are called by the processor 701 to execute the method for calibrating the gradient of the visual tactile sensor based on deep learning provided by the embodiments of the present invention; An input / output interface 703, which is used to implement information input and output; A communication interface 704, which is used to implement communication and interaction between this device and other devices. It can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); A bus 705, which transmits information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704); Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 are communicatively connected to each other inside the device through the bus 705.

[0058] As shown Figure 8 in the following figure is a schematic diagram of the structure of the storage medium provided by the embodiment of the present invention. Referring to Figure 8 , the embodiment of the present invention also provides a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs 801, and the one or more programs 801 can be executed by one or more processors to implement the above-mentioned method for calibrating the gradient of the visual tactile sensor based on deep learning.

[0059] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0060] Embodiments of the present invention also disclose a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.

[0061] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the above blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.

[0062] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0063] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes of various kinds.

[0064] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0065] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which the above program can be printed, because the above program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways when necessary, and then storing it in a computer memory.

[0066] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0067] In the above description of this specification, the description referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0068] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0069] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A gradient calibration method for a visual-tactile sensor based on deep learning, characterized in that, It includes the following steps: Obtain the sensor image information of a preset spherical ball pressing on the target touch film of the tactile sensor to be calibrated, and input the sensor image information into a pre-trained deep learning segmentation model to obtain the spherical ball contact area of the target touch film; Determine the coordinate values and RGB values of each point within the spherical ball contact area, and calculate the gradient information of each point within the spherical ball contact area according to the radius of the preset spherical ball; Construct a training data set based on the coordinate values, the RGB values, and the gradient information, and train a gradient recognition model according to the training data set; Calibrate the mapping relationship between the RGB light intensity and the gradient of each point on the target touch film according to the gradient recognition model.

2. The method for gradient calibration of a visual tactile sensor based on deep learning according to claim 1, wherein, The deep learning segmentation model is trained through the following steps: Obtain multiple sensor sample images of the preset spherical ball pressing on different positions of a sample touch film with different forces; Perform segmentation annotation on the sensor sample images to determine the contact area labels corresponding to the sensor sample images; Input the sensor sample images into a pre-constructed deep learning neural network to obtain a predicted contact area; Determine a first loss value according to the predicted contact area and the contact area label; Update the parameters of the deep learning neural network according to the first loss value to obtain the trained deep learning segmentation model.

3. A method for calibrating the gradient of a visual-tactile sensor based on deep learning according to claim 1, characterized in that, The calculation of the gradient information of each point within the spherical ball contact area according to the radius of the preset spherical ball specifically includes: Determine the actual bending radius according to the radius of the preset spherical ball and the thickness of the target touch film; Fit the spherical ball contact area to obtain a fitted circular area, and determine the center position of the fitted circular area; Calculate the gradient information of each point within the spherical ball contact area according to the center position and the actual bending radius.

4. A method for gradient calibration of a visual-tactile sensor based on deep learning according to claim 3, characterized in that, The gradient information is calculated by the following formula: Among them, and respectively represent the gradient values of the point (x, y) in the x - direction and y - direction, (x0, y0) represents the position of the center of the circle, and R represents the actual bending radius.

5. A method for gradient calibration of a visual-tactile sensor based on deep learning according to claim 1, characterized in that, The construction of the training data set based on the coordinate values, the RGB values, and the gradient information, and the training of the gradient recognition model according to the training data set specifically includes: Determine the x coordinate and the y coordinate according to the coordinate values, determine the R value, the G value, and the B value according to the RGB values, and determine a training sample according to the R value, the G value, the B value, the x coordinate, and the y coordinate; Determine the x-direction gradient value and the y-direction gradient value according to the gradient information, and determine the gradient label corresponding to the training sample according to the x-direction gradient value and the y-direction gradient value; Input the training sample into a pre-constructed ANN neural network to obtain a gradient recognition result; Determine a second loss value according to the gradient recognition result and the gradient label; Update the parameters of the ANN neural network according to the second loss value to obtain the trained gradient recognition model.

6. A method for calibrating the gradient of a visual-tactile sensor based on deep learning according to claim 5, characterized in that: The ANN neural network includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, a fourth hidden layer, and an output layer. The input layer is used to input the training samples, and the output layer is used to output the gradient recognition result. The first hidden layer, the second hidden layer, the third hidden layer, and the fourth hidden layer respectively include 64, 128, 128, and 64 neurons.

7. A method for calibrating the gradient of a visual tactile sensor based on deep learning according to any one of claims 1 to 6, characterized in that, Calibrating the mapping relationship between the RGB light intensity and the gradient of each point on the target touch film according to the gradient recognition model specifically includes: Dividing the target touch film into multiple grid points according to a preset spacing; Determining the gradient of each grid point under different RGB light intensities according to the gradient recognition model, and establishing a mapping relationship based on the coordinate position, RGB light intensity, and corresponding gradient of the grid point; Generating a three-dimensional lookup table containing coordinate values, RGB values, and gradient values according to the mapping relationship.

8. A visual-tactile sensor gradient calibration device based on deep learning, characterized in that, Including: A contact area recognition module, configured to obtain sensor image information of a preset sphere pressing on the target touch film of the to-be-calibrated visual tactile sensor, input the sensor image information into a pre-trained deep learning segmentation model, and obtain the sphere contact area of the target touch film; A gradient information calculation module, configured to determine the coordinate values and RGB values of each point within the sphere contact area, and calculate the gradient information of each point within the sphere contact area according to the radius of the preset sphere; A model training module, configured to construct a training data set according to the coordinate values, the RGB values, and the gradient information, and train a gradient recognition model according to the training data set; A calibration module, configured to calibrate the mapping relationship between the RGB light intensity and the gradient of each point on the target touch film according to the gradient recognition model.

9. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, the steps of the method for calibrating the gradient of a visual tactile sensor based on deep learning according to any one of claims 1 to 7 are implemented.

10. A storage medium, the storage medium being a computer-readable storage medium for computer-readable storage, characterized in that The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the method for calibrating the gradient of a visual tactile sensor based on deep learning according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Photometric stereoscopic vision system calibration method based on visual tactile sensor

    CN115615342A

  • Visual tactile sensor based on lensless imaging and measuring method thereof

    CN118482656A

  • Six-axis distributed force touch sensing method, system and terminal based on binocular visual touch

    CN119625073A

Cited By

  • Noise processing method and device based on visual tactile sensor, equipment and medium

    CN121640071A