A method for depth image repair of transparent objects based on a multi-stage neural network

Through the method based on multi-stage neural network, the problem of the existing technology dealing with transparent objects in different scenario environments is solved, and efficient depth image repair and fast inference are achieved, which is suitable for real-time detection and capture of robots.

CN114943654BActive Publication Date: 2025-06-10ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210546060.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-06-10
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle transparent objects in different scenario environments, resulting in problems such as robots being unable to grasp transparent objects, and navigation obstacle avoidance cannot identify glass obstacles.

Method used

The deep image repair method of transparent objects based on multi-stage neural network is adopted, and the segmentation is performed through the Deeplab v3+ network model, the dynamic convolution layer and codec architecture are extracted, the spatial attention map is generated, and the multi-stage fusion is performed, and the repaired depth image is finally obtained.

Benefits of technology

It improves the generalization of the method, improves the repair accuracy, fast model inference speed, can achieve good results in real-time detection and crawling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943654B_ABST
    Figure CN114943654B_ABST
Patent Text Reader

Abstract

A method for repairing depth images of transparent objects based on a multi-stage neural network, comprising: reading an input image and performing scaling; predicting a segmentation image; removing noise in the depth image; initially extracting features using dynamic convolution; extracting first-stage features; generating spatial attention maps for different branches in the first stage; generating repaired images for different branches in the first stage; fusing the prediction results of the first stage; extracting second-stage features; generating a repaired image in the second stage; extracting third-stage features; generating a repaired image in the third stage; fusing the prediction results of the three stages to obtain a repaired depth image. The present invention improves the inference speed, enhances the model accuracy, gets rid of the dependence on the external environment, and improves the generalization of the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for repairing the depth image of a transparent object. Background Art

[0002] A depth image is a special image obtained by a depth camera. The depth image records the distance information between an object and the camera, and is a very important source of distance information in vision-based robotics. With the popularization of depth cameras and the in-depth research of computer vision, the increasingly developed computer vision technology can be introduced into the research field of robots, such as vision-based robot grasping, navigation and obstacle avoidance. Due to the unique optical properties of transparent objects, when light passes through a transparent object, refraction or specular reflection occurs, making the current depth camera unable to capture the depth information of the transparent object. Therefore, in vision-based robotics, it is impossible to handle transparent objects well, which leads to problems such as robots being unable to grasp transparent objects and navigation and obstacle avoidance being unable to recognize glass obstacles.

[0003] Traditional transparent object recognition and depth repair methods rely too much on external devices, such as a fixed background, fixed devices, etc., which limits their generalization ability in different scene environments and cannot be applied on a large scale. And some neural network-based repair methods have too slow inference speed and too low efficiency, which is unacceptable for robots pursuing real-time detection and grasping. Summary of the Invention

[0004] The present invention overcomes the above-mentioned disadvantages in the prior art and proposes a method for repairing transparent objects based on a multi-stage neural network.

[0005] A method for repairing the depth image of a transparent object based on a multi-stage neural network according to the present invention includes the following steps:

[0006] S1. Read the input color RGB image I rgb , and the corresponding depth image I dep .

[0007] S2. Scale the read images, and use bilinear interpolation to adjust the image resolution to the resolution required by the network, height×width. height represents the vertical resolution of the image, and width represents the horizontal resolution of the image.

[0008] S3. Use the Deeplab v3+ network model to obtain the segmentation image Iseg, and the formula is as follows:

[0009] I seg = Deeplab(I rgb ) (1)

[0010] Among them, Deeplab() represents the Deeplab v3+ network model, as Figure 2 shown.

[0011] S4: Remove the noise in the depth image. Use the predicted segmentation image I seg , and modify the depth values of the pixels in the depth image that belong to the transparent object pixels in I dep to 0, so as to remove the noise in the depth image. The formula is as follows:

[0012]

[0013] where represents the pixel at the i-th row and j-th column in the segmentation image, represents the pixel at the i-th row and j-th column in the depth image.

[0014] S5. Use the dynamic convolution layer to initially extract image features. The formula is as follows:

[0015]

[0016] where Conv dynamic () represents the dynamic convolution layer, F rgb0 represents the initial features extracted from I rgb , and F dep0 represents the initial features extracted from I dep .

[0017] S6. Use the encoder-decoder architecture to extract the features in the first stage. The formula is as follows:

[0018]

[0019] where UNet() represents the UNet sub-network model of the encoder-decoder architecture, as Figure 3 shown. F rgb1 represents the features obtained by extracting features from F rgb0 , and F dep1 represents the features obtained by extracting features from F dep0 .

[0020] S7. Generate the spatial attention map. The formula is as follows:

[0021]

[0022] where SAB() represents the spatial attention module, SA rgb represents the spatial attention map obtained by decoding F rgb1 , and SA dep represents the spatial attention map obtained by decoding F dep1 .

[0023] S8. Generate the first-stage repaired image Out rgb With Out dep . The formula is as follows:

[0024]

[0025] Where Dec() represents the decoding module, which is used to output the repaired result. Out rgb Represents the result obtained by decoding F rgb1 , and Out dep Represents the result obtained by decoding F dep1 .

[0026] S9. Fuse the prediction results to obtain the output result Out of the first stage 1 . The formula is as follows:

[0027] Out1 = SA rgb ·Out rgb + SA dep ·Out dep (7)

[0028] S10. Add F rgb1 and F dep1 to obtain the fused feature F 1 , and input it into the sub-network of the second stage to extract features. The formula is as follows:

[0029]

[0030] Where F 2 Represents the feature extracted by the second-stage network.

[0031] S11. Use the decoding module to decode the feature F 2 to obtain the output result Out of the second stage 2 .

[0032] S12. Input the feature F 2 into the network of the third stage to obtain the feature F 3 . The formula is as follows:

[0033] F 3 = ORNet(F 2 ) (9)

[0034] Where ORNet() represents the scale-invariant network with spatial attention, as Figure 4 shown.

[0035] S13. Use the decoding module to decode the feature F 3 to obtain the output result Out of the third stage 3 .

[0036] S14. Predict the output result Out using the spatial attention module 1 of the spatial attention map SA 1 and the output result Out 2 of the spatial attention map SA 2 and the output result Out 3 of the spatial attention map SA 3 .

[0037] S15. Fuse the output results of each stage to obtain the restored depth image I out . The formula is as follows:

[0038] I out = SA 1 ·Out 1 + SA 2 ·Out 2 + SA 3 ·Out 3 (10)

[0039] Compared with the prior art, the present invention has the following advantages:

[0040] (1) Solve the difference problem between the synthetic dataset and the real environment dataset, enabling the method to achieve better results in the real environment dataset and improving the generalization of the method.

[0041] (2) Improve the repair accuracy. The multi-stage network proposed by the invention, the dynamic convolution layer can dynamically generate convolution kernel parameters suitable for the current picture for global information; the scale-invariant network can repair fine details.

[0042] (3) The model has a fast inference speed. The network proposed by the invention has an average inference speed of 0.02 s on the GPU, and the inference speed is increased by 100 times compared with the previous method. Description of the Drawings

[0043] Figure 1 is the overall flow chart of the method of the present invention.

[0044] Figure 2 is the data processing network structure diagram in the method of the present invention.

[0045] Figure 3 is the encoder-decoder network structure diagram in the method of the present invention.

[0046] Figure 4 is the scale-invariant network structure diagram in the method of the present invention. Detailed Embodiments

[0047] To make the process of the present invention easier to understand, the present invention will be described in detail with reference to examples.

[0048] S1. Read the input color RGB image I rgb , and the corresponding depth image I dep .

[0049] S2. Resize the read images. Use bilinear interpolation to adjust the image resolution to the resolution required by the network, height×width. height represents the vertical resolution of the image, and width represents the horizontal resolution of the image.

[0050] S3. Use the Deeplab v3+ network model to obtain the segmentation image I seg , and the formula is as follows:

[0051] I seg = Deeplab(I rgb ) (11)

[0052] where Deeplab() represents the Deeplab v3+ network model, as Figure 2 shown.

[0053] S4: Design to remove the part of incorrect depth information. Use the predicted segmentation image I seg to modify the depth information of the pixels in the depth image I dep that belong to the transparent object pixels to 0, thereby removing the incorrect depth information. The formula is as follows:

[0054]

[0055] where represents the pixel at the i-th row and j-th column in the segmentation image, and represents the pixel at the i-th row and j-th column in the depth image.

[0056] S5. Use the dynamic convolution layer for preliminary extraction of image features. The formula is as follows:

[0057]

[0058] where Conv dynamic () represents the dynamic convolution layer, F rgb0 represents the initial features extracted from I rgb , and F dep0 represents the initial features extracted from I dep .

[0059] S6. Use the encoder-decoder architecture to extract the features of the first stage. The formula is as follows:

[0060]

[0061] Among them, UNet() represents the UNet sub-network model of the encoder-decoder architecture, as Figure 3 shown. F rgb1 represents the feature obtained by performing feature extraction on F rgb0 , and F dep1 represents the feature obtained by performing feature extraction on F dep0 .

[0062] S7. Generate a spatial attention map. The formula is as follows:

[0063]

[0064] Among them, SAB() represents the spatial attention module, and SA rgb represents the spatial attention map obtained by decoding F rgb1 , and SA dep represents the spatial attention map obtained by decoding F dep1 .

[0065] S8. Generate the first-stage repaired image Out rgb and Out dep . The formula is as follows:

[0066]

[0067] Among them, Dec() represents the decoding module, which is used to output the repaired result. Out rgb represents the result obtained by decoding F rgb1 , and Out dep represents the result obtained by decoding F dep1 .

[0068] S9. Fuse the prediction results to obtain the output result Out 1 of the first stage. The formula is as follows:

[0069] Out1 = SA rgb ·Out rgb + SA dep ·Out dep (17)

[0070] S10. Add F rgb1 and F dep1 to obtain the fused feature F 1 , and input it into the sub-network of the second stage for feature extraction. The formula is as follows:

[0071]

[0072] Among them, F 2 represents the feature extracted by the second-stage network.

[0073] S11. Use the decoding module to decode the feature F 2 to obtain the output result Out of the second stage 2 .

[0074] S12. Input the feature F 2 into the network of the third stage to obtain the feature F 3 . The formula is as follows:

[0075] F 3 = ORNet(F 2 ) (19)

[0076] where ORNet() represents a scale-invariant network with spatial attention, as Figure 4 shown.

[0077] S13. Use the decoding module to decode the feature F 3 to obtain the output result Out of the third stage 3 .

[0078] S14. Use the spatial attention module to predict the spatial attention map SA 1 of the output result Out 1 , the spatial attention map SA 2 of the output result Out 2 , and the spatial attention map SA 3 of the output result Out 3 .

[0079] S15. Fuse the output results of each stage to obtain the restored depth image I out . The formula is as follows:

[0080] I out = SA 1 ·Out 1 + SA 2 ·Out 2 + SA 3 ·Out 3 (20)

[0081] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the examples, and the protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the inventive concept of the present invention.

Claims

1. A method for repairing the depth image of a transparent object based on a multi-stage neural network, comprising the following steps: S1. Read the input color RGB image I rgb , and the corresponding depth image I dep ; S2. Scale the read picture, and use bilinear interpolation to adjust the image resolution to the resolution height×width required by the network; hright represents the vertical resolution of the image, and width represents the horizontal resolution of the image; S3. Obtain the segmented image I using the Deeplab v3+ network model seg , the formula is as follows: I seg = Deeplab(I rgb ) (1) wherein, Deeplab() represents the Deeplab v3+ network model; S4: Remove the noise in the depth image; use the predicted segmentation image I seg , and modify the depth values of the pixels in the depth image that belong to the transparent object to 0 in I dep , so as to remove the noise in the depth image; the formula is as follows: Among them represents the pixel at the i-th row and j-th column in the segmented image, represents the pixel at the i-th row and j-th column in the depth image; S5. Use a dynamic convolutional layer to initially extract image features, and the formula is as follows: Among them, Conv dynamic ( ) represents the dynamic convolution layer, F rgb0 represents the initial features extracted from I rgb , and F dep0 represents the initial features extracted from I dep ; S6. Use an encoder-decoder architecture to extract the features of the first stage; the formula is as follows: Among them, UNet() represents the UNet sub-network model of the encoder-decoder architecture; F rgb1 represents the feature obtained by performing feature extraction on F rgb0 , and F dep1 represents the feature obtained by performing feature extraction on F dep0 ; S7. Generate a spatial attention map; the formula is as follows: Among them, SAB( ) represents the spatial attention module, and SA rgb represents the spatial attention map obtained by decoding F rgb1 , and SA dep represents the spatial attention map obtained by decoding F dep1 ; S8. Generate the first-stage repaired image Out rgb and Out dep ; The formula is as follows: Among them, Dec( ) represents the decoding module, which is used to output the repaired result; Out rgb represents the result obtained by decoding F rgb 1, and Out dep represents the result obtained by decoding F dep1 ; S9. The prediction results are fused to obtain the output result Out of the first stage 1 ; The formula is as follows: Out 1 = SA rgb ·Out rgb + SA dep ·Out dep (7) S10. Add F rgb1 to F dep1 to obtain the fused feature F 1 , and input it into the sub-network of the second stage to extract features; the formula is as follows: Among which F 2 represents the features extracted by the second-stage network; S11. Use the decoding module to decode the feature F 2 to obtain the output result Out of the second stage 2 ; S12. Input feature F 2 into the network in the third stage to obtain feature F 3 ; The formula is as follows: P 3 = ORNet(F 2 ) (9) where ORNet() represents a scale-invariant network with spatial attention; S13. Use the decoding module to decode the feature F 3 to obtain the output result Out of the third stage 3 ; S14. Use the spatial attention module to predict the output result Out 1 of the spatial attention map SA 1 , and the output result Out 2 of the spatial attention map SA 2 , and the output result Out 3 of the spatial attention map SA 3 ; S15. The output results of each stage are fused to obtain the repaired depth image I out ; The formula is as follows: I out = SA 1 ·Out 1 + SA 2 ·Out 2 + SA 3 ·OUt 3 (10).

Citation Information

Patent Citations

  • Real-time semantic segmentation method based on context attention mechanism and information fusion

    CN112541503A

  • Single-image rain removal method and system based on deep convolutional neural network, and storage medium

    CN113450288A