Construction method and three-dimensional reconstruction method of normal vector prediction model based on contact surface rendering

By constructing a normal vector prediction model and boundary depth prior based on contact surface rendering, the problems of difficulty in obtaining normal vectors and large integral errors in three-dimensional reconstruction of complex surfaces by visual-tactile sensors are solved, and the rapid and accurate acquisition of normal vectors and improved accuracy of three-dimensional reconstruction are achieved.

CN120655798APending Publication Date: 2025-09-16HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510663635.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing visual-tactile sensors have difficulty obtaining normal vectors in three-dimensional reconstruction of complex perception surfaces and have limited accuracy, and there are cumulative errors in the normal vector integration process.

Method used

By constructing a normal vector prediction model based on contact surface rendering, the normal vector prediction model is trained using the image information when the calibration sphere contacts the sensor, and 3D reconstruction is performed in combination with the boundary depth prior, which reduces the difficulty of normal vector prediction and improves the accuracy.

Benefits of technology

It achieves fast and accurate acquisition of normal vectors, reduces the cumulative error of normal vector integrals, and improves the accuracy of three-dimensional reconstruction of complex surfaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655798A_ABST
    Figure CN120655798A_ABST
Patent Text Reader

Abstract

The invention belongs to the related technical field of computer vision and three-dimensional reconstruction, and discloses a construction method of a normal vector prediction model based on contact surface rendering and a three-dimensional reconstruction method. The method comprises the following steps that a real image generated when a calibration ball in an actual scene makes contact with a sensor serves as input, normal vectors of all points in a contact area serve as output to train a normal vector prediction model, and the trained normal vector prediction model is the needed normal vector prediction model. Acquiring a real contact image of the object and the sensor in an actual scene, predicting a normal vector of each point on the surface of the contact object by using the normal vector prediction model, and calculating the depth of each point on the surface of the contact object by using the normal vector and the depth prior by taking the image boundary depth as the depth prior. And three-dimensional coordinates of all points on the surface of the contact object are obtained through depth calculation, so that three-dimensional reconstruction of the contact object is realized. According to the invention, the problems of difficult acquisition of a real normal vector and large integral error of the normal vector during calibration are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to computer vision and three-dimensional reconstruction, and more specifically, relates to a method for constructing a normal vector prediction model based on contact surface rendering and a three-dimensional reconstruction method. Background Art

[0002] Visual tactile sensors use camera systems to capture contact surfaces as tactile images. Combined with computer vision technology, these images are then analyzed into higher-level tactile features, such as the contact surface's three-dimensional geometry, slip, force / torque, and roughness. Contact surface geometry reconstruction uses visual information (tactile images captured by camera systems) to build a three-dimensional geometric model of the contact surface. This typically involves identifying and measuring the shape, structure, texture, and other geometric properties of an object's surface. This is crucial for understanding the physical properties of objects and how we interact with them, and has widespread applications in robotic grasping and manipulation, product quality control and inspection, and medical imaging.

[0003] However, current visual-tactile sensors are mostly flat or simply curved, significantly limiting their applicability in scenarios such as confined spaces and dexterous manipulation. The main challenges in 3D reconstruction of visual-tactile sensors with complex sensing surfaces (such as finger shapes) are obtaining the true surface normal, limited normal reconstruction accuracy, and cumulative errors in the normal integration process. Therefore, a method for quickly and accurately obtaining the true normal value is needed. Summary of the Invention

[0004] In response to the above defects or improvement needs of the existing technology, the present invention provides a method for constructing a normal vector prediction model based on contact surface rendering and a three-dimensional reconstruction method to solve the problems of difficulty in obtaining normal vectors and large normal vector integral errors of visual sensors.

[0005] To achieve the above object, according to one aspect of the present invention, a method for constructing a normal vector prediction model based on contact surface rendering is provided, the method comprising the following steps:

[0006] Obtain the real image when the calibration ball contacts the sensor in the actual scene;

[0007] Establishing a node of the image acquisition device using the parameters of the calibrated image acquisition device in the rendering scene, establishing a three-dimensional geometric model of the calibration sphere and the sensor surface, and acquiring a virtual image of the calibration sphere in contact with the sensor from the node of the image acquisition device;

[0008] Adjusting the three-dimensional coordinates of the calibration sphere so that the virtual image and the real image coincide with each other, obtaining the current depth of each point in the currently rendered scene, subtracting the current depth of each point from the background depth to obtain the depth difference between the calibration sphere and the contact area of ​​the sensor, and using the depth difference to calculate the normal vector of each point in the contact area;

[0009] The real image of the calibration ball in contact with the sensor in the actual scene is used as input, and the normal vectors of each point in the contact area are used as output to train the normal vector prediction model. The trained normal vector prediction model is the required normal vector prediction model.

[0010] Further preferably, the parameters of the calibrated image acquisition device are obtained according to the following steps:

[0011] Acquire an actual image of the calibration object captured by an image acquisition device in an actual scene;

[0012] A three-dimensional geometric model of a virtual image acquisition device node and a calibration object is established in a rendering scene, and a rendering image of the calibration object is captured by the virtual image acquisition device node; the relative positions of the image acquisition device and the calibration object are the same in the rendering scene and the actual scene;

[0013] Adjust the parameters of the image acquisition device in the virtual scene until the rendered image of the calibration object coincides with the actual image. At this time, the parameters of the image acquisition device are the required parameters.

[0014] Further preferably, the background depth is the depth of each point in the rendered image when no calibration sphere exists in the rendered scene.

[0015] Further preferably, the normal vector prediction model is a multi-layer perceptron network structure.

[0016] Further preferably, the formula for calculating the normal vector of each point in the contact area using the depth difference is as follows:

[0017]

[0018] n=(p,q,-1)

[0019] Where z is the depth difference function, p and q are the gradients of z in the x and y directions, and n is the normal vector at (x, y).

[0020] Further preferably, the parameters of the image acquisition device include a rotation matrix, a translation vector, focal lengths in the X-axis and Y-axis directions, and principal point coordinates.

[0021] According to another aspect of the present invention, a 3D reconstruction method based on boundary depth prior is provided, the method comprising the following steps:

[0022] Establishing a virtual shooting node using the parameters of the image acquisition device calibrated in the rendering scene, capturing an image of the contact object from the virtual shooting node, and using the pixel width within a preset range of the edge of the rendered image as the image boundary depth;

[0023] Acquire a real image of the contact object in the actual scene, use the normal vector prediction model described above to predict the normal vector of each point on the surface of the contact object, use the image boundary depth as the depth prior, use the normal vector and the depth prior to calculate the depth of each point on the surface of the contact object, and use the parameters of the image acquisition device and the depth calculation to obtain the three-dimensional coordinates of each point on the surface of the contact object, thereby realizing three-dimensional reconstruction of the contact object.

[0024] Further preferably, the deformation depth is calculated according to the following relationship:

[0025]

[0026] Among them, A is the sparse matrix of the encoding depth coefficient, λ is the depth prior weight, I prior is a diagonal matrix whose diagonal elements are 1 at the positions corresponding to pixels with valid depth priors, b is the Poisson equation divergence term, z prior is the depth prior and z is the deformation depth.

[0027] Further preferably, when the image size is m×n pixels, the sparse matrix A is an mn×mn square matrix, each row has a maximum of 5 non-zero elements, the main diagonal element value is -4, and the remaining non-zero elements are 1.

[0028] The calculation formula for the three-dimensional coordinates of each point on the surface of the contact object is as follows:

[0029]

[0030] Z i,j =z i,j

[0031] Among them, z i,j is the depth of pixel (i, j), f x and f y is the focal length of the camera in the x and y directions, (c x ,c y ) are the main point coordinates.

[0032] In general, the above technical solutions conceived by the present invention have the following beneficial effects compared with the prior art:

[0033] 1. The present invention constructs a virtual scene similar to the actual scene and creates a virtual image. Then, the normal vector is calculated using depth information. This method quickly obtains the true value of the normal vector of a complex surface by rendering the contact scene between the calibration sphere and the sensor. The corresponding relationship between the actual image and the contact surface normal vector is then established using the contact surface rendering method. A normal vector prediction model is then established based on this correspondence between the actual image and the contact surface normal vector. The normal vector prediction model is then trained to obtain the desired normal vector prediction model. This model reduces the difficulty of normal vector prediction and improves the accuracy and speed of normal vector prediction.

[0034] 2. The present invention introduces the boundary depth prior from the sensor CAD model to force the depth at the boundary to be zero, providing a global reference benchmark for the normal vector integration process, ensuring that all integration paths converge to the same reference plane, avoiding the global error divergence caused by local normal vector errors, and reducing the cumulative error of normal vector integration. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A three-dimensional reconstruction method based on boundary depth prior constructed according to a preferred embodiment of the present invention;

[0036] Figure 2 is a flow chart of a sensor calibration method constructed according to a preferred embodiment of the present invention;

[0037] Figure 3 Schematic diagram of a normal vector true value calibration device constructed according to a preferred embodiment of the present invention;

[0038] Figure 4 Schematic diagram of the structure of a multispectral curved surface visual-tactile sensor constructed according to a preferred embodiment of the present invention;

[0039] Figure 5 This is a diagram of three-dimensional reconstruction experimental results constructed according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0041] like Figure 1 As shown, a method for constructing a normal vector prediction model based on contact surface rendering includes the following steps:

[0042] (1) Establishing a training dataset

[0043] In one embodiment of the present invention, the image acquisition device includes RGB and NIR cameras, and the contact object is a sensor.

[0044] S1 calibrates the parameters of the image acquisition device in the rendering scene

[0045] 1) Obtain RGB and NIR camera intrinsic parameters.

[0046] By taking multi-angle images of known geometric patterns (such as chessboard), the Zhang Zhengyou calibration method is used to calculate the internal parameters of the two cameras, including the focal length (f x ,f y ), principal point (c x ,c y ).

[0047] 2) Obtain the affine transformation parameters between RGB and NIR images through feature point matching.

[0048] A calibration object with characteristic protrusions that fits the sensor surface is placed on the sensor surface and pressed down to ensure full contact between the calibration object and the sensor. A dual-camera system is used to synchronously capture RGB and NIR images.

[0049] Feature point matching is used to obtain the affine transformation parameters of the RGB image relative to the NIR image, including the scaling factor, rotation angle, and translation. Feature point matching can be performed using corner detection and brute force matching.

[0050] 3) Obtain the camera intrinsic and extrinsic parameters in the rendered scene through rendering.

[0051] like Figure 2 As shown in the figure, the camera internal and external parameter calibration process specifically includes:

[0052] The calibration object CAD model was read and a rendering scene was created using the OpenGL-based 3D rendering engine Pyrender. The trimesh mesh was converted into a Pyrender mesh and added to the rendering scene. A renderer with a size of 640 × 480 pixels (the same as the captured image) was created. The renderer was set up with the 3D model of the calibration object.

[0053] In the renderer, initialize the camera's extrinsic parameters. The 3D coordinates are initialized to [0, 0, 10], and the pitch, yaw, and roll angles are initialized to 0. Calculate the camera's initial extrinsic matrix based on the initialization parameters. The specific steps are: convert the angles to radians, construct the X, Y, and Z axis rotation matrices, combine the rotation matrix R in the order of Z, Y, and X, add the translation component T to construct the complete extrinsic matrix, and convert the coordinate system from OpenCV to OpenGL.

[0054] By aligning the data, the RGB camera internal parameters (f x1 ,f y1 ,cx1 ,c y1 ) and the NIR camera internal reference (f x2 ,f y2 ,c x2 ,c y2 ) is unified as (f x ',f y ',c x ',c y '), initialize the camera internal parameters (f x ',f y ',c x ',c y '), create a camera node with external parameters (R, T) and add it to the rendering scene, add light sources to the scene, and import the actual calibration object image taken.

[0055] Build a virtual node from the rendering scene, obtain the rendered image of the calibration object through this node, overlay the rendered image of the calibration object with the actual image taken by the camera, adjust the camera's extrinsic and intrinsic parameters, update the camera's intrinsic and extrinsic parameter matrices in real time and synchronize them to the rendering scene. When the characteristic patterns of the rendered image and the actual image are observed to coincide, record the camera parameters at this time.

[0056] S2 obtains the virtual image when the calibration ball contacts the sensor

[0057] 1) Calibrate the sensor and obtain the true normal value by rendering an image of a small ball with a known radius.

[0058] Normal vector true value calibration device such as Figure 3 As shown. The steps to obtain the true normal value include:

[0059] The sensor is fixed on the base, and computer numerical control technology is used to control the probe with a small ball of known radius at the end to contact the sensor at different positions, so as to evenly collect calibration images of different areas of the surface.

[0060] The CAD model of the sensor surface is read and a rendering scene is created. The rendering scene includes the camera node, the ball node, and the 3D model of the sensor surface. In the current rendering scene, the ball is in contact with the sensor. A camera node is created using the calibrated camera parameters and the background depth is read from the rendering scene.

[0061] Enter a known radius to create a sphere mesh and add a material. Combine the mesh and material into a renderable object and add it to the scene.

[0062] 2) Read the rendered scene pixels and depth, and obtain the surface normal by calculating the depth gradient.

[0063] Use central difference to calculate the gradient of the depth map in the x and y directions

[0064]

[0065] Then the true value of the surface normal can be calculated as follows

[0066]

[0067] The image of the rendered scene is superimposed with the calibrated image of the actual ball, the 3D coordinates of the ball node are adjusted and the rendered scene is updated synchronously so that the rendered ball coincides with the actual ball.

[0068] In the rendered scene, a virtual image of the calibration sphere in contact with the sensor is captured from the camera's node.

[0069] The depth difference between the ball and the sensor contact area is obtained by subtracting the background depth from the depth of each point in the image of the current rendered scene, that is, the contact area mask. The normal of each point in the mask area is saved as the true normal value required for MLP training.

[0070] Construct a scene that does not contain a ball node, so that the depth of each point on the obtained rendered image is the background depth.

[0071] (2) Training the normal vector prediction MLP model.

[0072] The normal vector prediction model takes as input the image of the ball captured in the actual scene, the pixel values ​​and pixel coordinates of each point in the image, and outputs the normal vector of each pixel in the image corresponding to the spatial point.

[0073] In one embodiment of the present invention, an MLP model is used. The model structure is as follows: the input pixel value feature uses the difference between the foreground and the background to eliminate background noise and highlight dynamic tactile deformation; the initial number of input channels is 3, and you can choose whether to enable NIR / position encoding to expand the number of channels to 6 / 9; the core network structure is convolution layer (1×1 convolution kernel) → batch normalization layer → ReLU activation layer → DropOut layer → output layer (1×1 convolution layer). Among them, the 1×1 convolution layer does not change the image size and is suitable for pixel-level normal vector prediction; the batch normalization layer accelerates convergence, stabilizes training, and allows a higher learning rate; the ReLU activation layer introduces nonlinearity to fit complex mapping relationships; the DropOut layer randomly masks 40% of neurons to prevent overfitting of tactile noise; the output layer compresses the features to 3 channels, corresponding to the three components of the normal vector.

[0074] In one embodiment of the present invention, the training configuration is as follows:

[0075] GPU-accelerated training was performed, with FP32 precision, an initial learning rate of 0.01, a batch size of 16, 300 epochs, and a weight decay coefficient of 0.0001. Weight decay was applied to standard convolutional layers to prevent overfitting. To avoid interfering with batch normalization statistics and to minimize the impact of the bias term on model complexity, weight decay was not applied to batch normalization layers or the bias term.

[0076] Since the adaptive learning rate feature is suitable for scenarios with large gradient changes in tactile data, the Adam optimizer is used.

[0077] The learning rate is reduced to 0.1 every 100 epochs to achieve rapid convergence in the early stage and fine-tune in the later stage.

[0078] num_workers is set to 8 to fully utilize multi-core CPUs; persistent_workers is set to True to avoid the overhead of repeated creation and destruction.

[0079] Select the L1 loss function and monitor the gradient error calculated from the normal vector.

[0080] The specific process of making the dataset is as follows:

[0081] The affine transformation parameters obtained in step S1 are applied to process the RGB image so as to align it with the NIR image data.

[0082] Normalize the X-axis coordinate to [-1, 1], the Y-axis coordinate to H / W (H is the image height, W is the image height), and add the Z-axis coordinate to 0.

[0083] The ball contact area mask processed image obtained in step S3 retains the pixel values ​​and pixel coordinates of the contact area.

[0084] By generating random numbers, we load contact data with a 50% probability and background data with a 50% probability, so that the model can learn to distinguish between contact and non-contact states.

[0085] If NIR training is enabled, concatenate RGB and NIR data and expand the input channels to 6.

[0086] Output each sample as a dictionary, including the current frame RGB / NIR data, background frame RGB / NIR data, normal vector label, and normalized position encoding.

[0087] The input to the MLP network is the difference between the foreground and background pixels and the position encoding. If only RGB training is used, the number of input channels is 3; if NIR training is enabled, the number of input channels is 6; if position encoding is enabled, the number of input channels is 6; if both NIR and position encoding are enabled, the number of input channels is 9.

[0088] The specific structure of the network is such as the vector prediction model, which obtains the contact normal vector map.

[0089] Table 1 shows the data. The 1×1 convolutional layer is used for cross-channel feature interaction and multimodal information integration; the batch normalization layer standardizes activation values ​​to accelerate training convergence; the ReLU activation function layer introduces nonlinear factors to improve the model's expressiveness; the DropOut layer randomly blocks 40% of neurons to prevent overfitting of tactile noise; and the output layer compresses the features into three channels, representing the three components of the normal vector.

[0090] A three-dimensional reconstruction method based on boundary depth prior comprises the following steps:

[0091] 1) Obtaining depth boundary prior

[0092] Import the CAD model of the sensor surface into the rendering scene, create a camera node using the camera intrinsic and extrinsic parameters obtained by the above steps, and capture an image of the sensor surface from the virtual node, i.e., the rendered image of the sensor surface. Read the 10-pixel width of the edge in the rendered image as the boundary of the rendered image, which serves as the depth boundary prior for the normal vector integration.

[0093] The depth prior is integrated to solve the depth field through normal vector integration. Depth refers to the height of each point on the sensor surface.

[0094] Given the surface normal n(x,y)=(n x ,n y ,n z ), the gradient of depth z(x,y) can be calculated as follows

[0095]

[0096] Furthermore, the 3D reconstruction problem can be modeled as solving the following Poisson equation:

[0097]

[0098] The above equation is discretized using central difference, and for pixel (i, j) we have

[0099]

[0100] For all valid pixels in the image, the Poisson equation can be combined into the following discrete linear system

[0101] Az=b

[0102] Where A is the sparse matrix of the encoded depth coefficients, z is the vector containing all pixel depths, and b is the divergence term of the Poisson equation.

[0103] The sparse matrix A is in the following form: if the image size is m×n pixels, then A is an mn×mn square matrix with a maximum of 5 non-zero elements in each row (one for the current pixel and its four neighbors), the main diagonal element value is -4, and the remaining non-zero elements are 1. Taking a 3×3 image as an example, the elements in Az=b are:

[0104]

[0105] The actual depth of the image boundary is used as the prior z prior , and integrate the weight λ into the above formula to obtain the augmented sparse linear system

[0106] A total z=b total

[0107] in, I prior is a diagonal matrix whose diagonal elements are 1 at positions corresponding to pixels with valid depth priors.

[0108] Solving overdetermined systems of equations using the least squares method

[0109] A total T A total z=A total T b total

[0110] The sparse matrix solvers sparseqr or spsolve can be used for fast solutions.

[0111] 2) 3D reconstruction

[0112] Assume that the camera internal parameter is focal length (f x ,f y ), principal point (c x ,c y ), depth z i,j The corresponding three-dimensional point is

[0113]

[0114] Z i,j =z i,j

[0115] This can complete the reconstruction of the three-dimensional point cloud of the object surface.

[0116] The present invention is further described below with reference to specific embodiments.

[0117] The image of the object to be 3D reconstructed in contact with the sensor is input into the obtained normal vector prediction model to obtain a contact normal vector map.

[0118] Table 1 Normal vector prediction network structure

[0119]

[0120]

[0121] Normalize the normal vector value from [0,255] to [-1,1] and unitize it. Establish the gradient equation for the four neighboring vertices of each pixel, such as the gradient of the upper left vertex All gradient equations are combined to obtain a sparse linear system.

[0122] A 10-pixel width depth boundary prior is introduced, and the influence of the depth prior is controlled by weights to obtain an augmented sparse linear system.

[0123] The coefficient matrix condition number is improved by matrix mean subtraction to avoid ill-conditioned problems.

[0124] Select sparseqr or spsolve solver for sparse matrix solving. sparseqr is based on the QR decomposition of Householder reflection and has high stability and is suitable for large-scale problems. spsolve is based on the LU decomposition method and has low memory usage but is slower, making it suitable for small-scale applications.

[0125] The vertex depth is converted to pixel depth through 2×2 mean filtering to eliminate meshing artifacts.

[0126] The three-dimensional points are back-projected according to the camera intrinsic parameters, as shown in the above formula, and the surface mesh is generated based on Open3D.

[0127] like Figure 4 The figure shows an embodiment of the multispectral curved surface visual and tactile sensing application of the present invention. Based on the RGB light source configuration, a NIR light source is added, a beam splitter prism is used to enable RGB / NIR cameras to capture images simultaneously, and an RGB / NIR filter is used to filter the required wavelengths of light.

[0128] like Figure 5 The figure shows an example of a 3D reconstruction application of the present invention, using four objects: a screwdriver, a fingertip, a grid pen holder, and a color-coded resistor. Observing the normal vectors and 3D point cloud reconstruction results, the addition of NIR imaging reveals subtle texture variations in the test objects more clearly than using only the RGB channel, qualitatively demonstrating the effectiveness of the proposed method in improving the accuracy of visual and tactile 3D reconstruction of curved surfaces.

[0129] To quantitatively analyze the effectiveness of the method of the present invention, a comparative experiment on the normal reconstruction error of different methods was conducted. The x-axis gradient error, y-axis gradient error, and error sum were compared. The results are shown in Table 2. It can be observed that for both the lookup table method and the neural network method used in the present invention, adding NIR training can reduce the normal reconstruction error; compared to the lookup table method, the MLP modeling method used in the present invention can further reduce the normal reconstruction error. Compared with the lookup table method using only RGB training, the MLP method proposed in the present invention using RGB-NIR training reduced the error from 0.0706 to 0.0130 (a reduction of 81.59%), greatly improving the normal reconstruction accuracy.

[0130] Furthermore, a comparison experiment on the normal vector integral error of different methods was conducted, with the unit being mm. The experimental results are shown in Table 3. It can be observed that the sparse matrix solution method with depth prior introduced in this invention greatly improves the accuracy of the normal vector integral compared to the fast Poisson method without depth prior, reducing the error from 0.579 mm to 0.0406 mm (a reduction of 92.99%).

[0131] Table 2 Normal vector reconstruction error comparison experiment

[0132]

[0133] Table 3 Normal vector integral error comparison experiment

[0134]

[0135] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for constructing a normal vector prediction model based on contact surface rendering, characterized in that: The method comprises the following steps: Obtain the real image when the calibration ball contacts the sensor in the actual scene; Establishing a node of the image acquisition device using the parameters of the calibrated image acquisition device in the rendering scene, establishing a three-dimensional geometric model of the calibration sphere and the sensor surface, and acquiring a virtual image of the calibration sphere in contact with the sensor from the node of the image acquisition device; Adjusting the three-dimensional coordinates of the calibration sphere so that the virtual image and the real image coincide with each other, obtaining the current depth of each point in the currently rendered scene, subtracting the current depth of each point from the background depth to obtain the depth difference between the calibration sphere and the contact area of ​​the sensor, and using the depth difference to calculate the normal vector of each point in the contact area; The real image of the calibration ball in contact with the sensor in the actual scene is used as input, and the normal vectors of each point in the contact area are used as output to train the normal vector prediction model. The trained normal vector prediction model is the required normal vector prediction model.

2. The method for constructing a normal vector prediction model based on contact surface rendering according to claim 1, wherein: The parameters of the calibrated image acquisition device are obtained according to the following steps: Acquire an actual image of the calibration object captured by an image acquisition device in an actual scene; A three-dimensional geometric model of a virtual image acquisition device node and a calibration object is established in a rendering scene, and a rendering image of the calibration object is captured at the virtual image acquisition device node; the relative positions of the image acquisition device and the calibration object are the same in the rendering scene and the actual scene; Adjust the parameters of the image acquisition device in the virtual scene until the rendered image of the calibration object coincides with the actual image. At this time, the parameters of the image acquisition device are the required parameters.

3. The method for constructing a normal vector prediction model based on contact surface rendering according to claim 1 or 2, wherein: The background depth is the depth of each point in the rendered image when there is no calibration sphere in the rendered scene.

4. The method for constructing a normal vector prediction model based on contact surface rendering according to claim 1 or 2, wherein: The normal vector prediction model is a multi-layer perceptron network structure.

5. The method for constructing a normal vector prediction model based on contact surface rendering according to claim 1 or 2, wherein: The formula for calculating the normal vector of each point in the contact area using the depth difference is as follows: n=(p,q,-1) Where z is the depth difference function, p and q are the gradients of z in the x and y directions, and n is the normal vector at (x, y).

6. The method for constructing a normal vector prediction model based on contact surface rendering according to claim 1 or 2, wherein: The parameters of the image acquisition device include a rotation matrix, a translation vector, focal lengths in the X-axis and Y-axis directions, and principal point coordinates.

7. A 3D reconstruction method based on boundary depth prior, characterized in that: The method comprises the following steps: Establishing a virtual shooting node based on the parameters of the image acquisition device calibrated in the rendering scene, capturing an image of the contact object from the virtual shooting node, and using the pixel width within a preset range of the edge of the rendered image as the image boundary depth; Obtain a real image of the contact object in the actual scene, use the normal vector prediction model described in any one of claims 1-6 to predict the normal vector of each point on the surface of the contact object, use the image boundary depth as the depth prior, use the normal vector and the depth prior to calculate the depth of each point on the surface of the contact object, use the parameters of the image acquisition device and the depth calculation to obtain the three-dimensional coordinates of each point on the surface of the contact object, thereby realizing three-dimensional reconstruction of the contact object.

8. The three-dimensional reconstruction method based on boundary depth prior according to claim 7, characterized in that: The deformation depth is calculated according to the following relationship: Among them, A is the sparse matrix of the encoding depth coefficient, λ is the depth prior weight, I prior is a diagonal matrix whose diagonal elements are 1 at the positions corresponding to pixels with valid depth priors, b is the Poisson equation divergence term, z prior is the depth prior and z is the deformation depth.

9. The three-dimensional reconstruction method based on boundary depth prior according to claim 8, characterized in that: When the image size is m×n pixels, the sparse matrix A is an mn×mn square matrix, each row has a maximum of 5 non-zero elements, the main diagonal element value is -4, and the remaining non-zero elements are 1.

10. The three-dimensional reconstruction method based on boundary depth prior according to claim 7 or 8, characterized in that: The calculation formula for the three-dimensional coordinates of each point on the surface of the contact object is as follows: WITH i,j =z i,j Among them, z i,j is the depth of pixel (i, j), f x and f y is the focal length of the camera in the x and y directions, (c x ,c y ) are the main point coordinates.

Citation Information

Cited By

  • Finger type visual tactile sensor calibration method and device and server

    CN121527200A