A method for depth image repair of transparent objects based on a multi-stage neural network
Through the method based on multi-stage neural network, the problem of the existing technology dealing with transparent objects in different scenario environments is solved, and efficient depth image repair and fast inference are achieved, which is suitable for real-time detection and capture of robots.
Patent Information
- Application Number
- CN202210546060.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-17
AI Technical Summary
The prior art is difficult to effectively handle transparent objects in different scenario environments, resulting in problems such as robots being unable to grasp transparent objects, and navigation obstacle avoidance cannot identify glass obstacles.
The deep image repair method of transparent objects based on multi-stage neural network is adopted, and the segmentation is performed through the Deeplab v3+ network model, the dynamic convolution layer and codec architecture are extracted, the spatial attention map is generated, and the multi-stage fusion is performed, and the repaired depth image is finally obtained.
It improves the generalization of the method, improves the repair accuracy, fast model inference speed, can achieve good results in real-time detection and crawling.
Smart Images

Figure CN114943654B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for repairing the depth image of a transparent object. Background Art
[0002] A depth image is a special image obtained by a depth camera. The depth image records the distance information between an object and the camera, and is a very important source of distance information in vision-based robotics. With the popularization of depth cameras and the in-depth research of computer vision, the increasingly developed computer vision technology can be introduced into the research field of robots, such as vision-based robot grasping, navigation and obstacle avoidance. Due to the unique optical properties of transparent objects, when light passes through a transparent object, refraction or specular reflection occurs, making the current depth camera unable to capture the depth information of the transparent object. Therefore, in vision-based robotics, it is impossible to handle transparent objects well, which leads to problems such as robots being unable to grasp transparent objects and navigation and obstacle avoidance being unable to recognize glass obstacles.
[0003] Traditional transparent object recognition and depth repair methods rely too much on external devices, such as a fixed background, fixed devices, etc., which limits their generalization ability in different scene environments and cannot be applied on a large scale. And some neural network-based repair methods have too slow inference speed and too low efficiency, which is unacceptable for robots pursuing real-time detection and grasping. Summary of the Invention
[0004] The present invention overcomes the above-mentioned disadvantages in the prior art and proposes a method for repairing transparent objects based on a multi-stage neural network.
[0005] A method for repairing the depth image of a transparent object based on a multi-stage neural network according to the present invention includes the following steps:
[0006] S1. Read the input color RGB image I rgb , and the corresponding depth image I dep .
[0007] S2. Scale the read images, and use bilinear interpolation to adjust the image resolution to the resolution required by the network, height×width. height represents the vertical resolution of the image, and width represents the horizontal resolution of the image.
[0008] S3. Use the Deeplab v3+ network model to obtain the segmentation image Iseg, and the formula is as follows:
[0009] I seg = Deeplab(I rgb ) (1)
[0010] Among them, Deeplab() represents the Deeplab v3+ network model, as Figure 2 shown.
[0011] S4: Remove the noise in the depth image. Use the predicted segmentation image I seg , and modify the depth values of the pixels in the depth image that belong to the transparent object pixels in I dep to 0, so as to remove the noise in the depth image. The formula is as follows:
[0012]
[0013] where represents the pixel at the i-th row and j-th column in the segmentation image, represents the pixel at the i-th row and j-th column in the depth image.
[0014] S5. Use the dynamic convolution layer to initially extract image features. The formula is as follows:
[0015]
[0016] where Conv dynamic () represents the dynamic convolution layer, F rgb0 represents the initial features extracted from I rgb , and F dep0 represents the initial features extracted from I dep .
[0017] S6. Use the encoder-decoder architecture to extract the features in the first stage. The formula is as follows:
[0018]
[0019] where UNet() represents the UNet sub-network model of the encoder-decoder architecture, as Figure 3 shown. F rgb1 represents the features obtained by extracting features from F rgb0 , and F dep1 represents the features obtained by extracting features from F dep0 .
[0020] S7. Generate the spatial attention map. The formula is as follows:
[0021]
[0022] where SAB() represents the spatial attention module, SA rgb represents the spatial attention map obtained by decoding F rgb1 , and SA dep represents the spatial attention map obtained by decoding F dep1 .
[0023] S8. Generate the first-stage repaired image Out rgb With Out dep . The formula is as follows:
[0024]
[0025] Where Dec() represents the decoding module, which is used to output the repaired result. Out rgb Represents the result obtained by decoding F rgb1 , and Out dep Represents the result obtained by decoding F dep1 .
[0026] S9. Fuse the prediction results to obtain the output result Out of the first stage 1 . The formula is as follows:
[0027] Out1 = SA rgb ·Out rgb + SA dep ·Out dep (7)
[0028] S10. Add F rgb1 and F dep1 to obtain the fused feature F 1 , and input it into the sub-network of the second stage to extract features. The formula is as follows:
[0029]
[0030] Where F 2 Represents the feature extracted by the second-stage network.
[0031] S11. Use the decoding module to decode the feature F 2 to obtain the output result Out of the second stage 2 .
[0032] S12. Input the feature F 2 into the network of the third stage to obtain the feature F 3 . The formula is as follows:
[0033] F 3 = ORNet(F 2 ) (9)
[0034] Where ORNet() represents the scale-invariant network with spatial attention, as Figure 4 shown.
[0035] S13. Use the decoding module to decode the feature F 3 to obtain the output result Out of the third stage 3 .
[0036] S14. Predict the output result Out using the spatial attention module 1 of the spatial attention map SA 1 and the output result Out 2 of the spatial attention map SA 2 and the output result Out 3 of the spatial attention map SA 3 .
[0037] S15. Fuse the output results of each stage to obtain the restored depth image I out . The formula is as follows:
[0038] I out = SA 1 ·Out 1 + SA 2 ·Out 2 + SA 3 ·Out 3 (10)
[0039] Compared with the prior art, the present invention has the following advantages:
[0040] (1) Solve the difference problem between the synthetic dataset and the real environment dataset, enabling the method to achieve better results in the real environment dataset and improving the generalization of the method.
[0041] (2) Improve the repair accuracy. The multi-stage network proposed by the invention, the dynamic convolution layer can dynamically generate convolution kernel parameters suitable for the current picture for global information; the scale-invariant network can repair fine details.
[0042] (3) The model has a fast inference speed. The network proposed by the invention has an average inference speed of 0.02 s on the GPU, and the inference speed is increased by 100 times compared with the previous method. Description of the Drawings
[0043] Figure 1 is the overall flow chart of the method of the present invention.
[0044] Figure 2 is the data processing network structure diagram in the method of the present invention.
[0045] Figure 3 is the encoder-decoder network structure diagram in the method of the present invention.
[0046] Figure 4 is the scale-invariant network structure diagram in the method of the present invention. Detailed Embodiments
[0047] To make the process of the present invention easier to understand, the present invention will be described in detail with reference to examples.
[0048] S1. Read the input color RGB image I rgb , and the corresponding depth image I dep .
[0049] S2. Resize the read images. Use bilinear interpolation to adjust the image resolution to the resolution required by the network, height×width. height represents the vertical resolution of the image, and width represents the horizontal resolution of the image.
[0050] S3. Use the Deeplab v3+ network model to obtain the segmentation image I seg , and the formula is as follows:
[0051] I seg = Deeplab(I rgb ) (11)
[0052] where Deeplab() represents the Deeplab v3+ network model, as Figure 2 shown.
[0053] S4: Design to remove the part of incorrect depth information. Use the predicted segmentation image I seg to modify the depth information of the pixels in the depth image I dep that belong to the transparent object pixels to 0, thereby removing the incorrect depth information. The formula is as follows:
[0054]
[0055] where represents the pixel at the i-th row and j-th column in the segmentation image, and represents the pixel at the i-th row and j-th column in the depth image.
[0056] S5. Use the dynamic convolution layer for preliminary extraction of image features. The formula is as follows:
[0057]
[0058] where Conv dynamic () represents the dynamic convolution layer, F rgb0 represents the initial features extracted from I rgb , and F dep0 represents the initial features extracted from I dep .
[0059] S6. Use the encoder-decoder architecture to extract the features of the first stage. The formula is as follows:
[0060]
[0061] Among them, UNet() represents the UNet sub-network model of the encoder-decoder architecture, as Figure 3 shown. F rgb1 represents the feature obtained by performing feature extraction on F rgb0 , and F dep1 represents the feature obtained by performing feature extraction on F dep0 .
[0062] S7. Generate a spatial attention map. The formula is as follows:
[0063]
[0064] Among them, SAB() represents the spatial attention module, and SA rgb represents the spatial attention map obtained by decoding F rgb1 , and SA dep represents the spatial attention map obtained by decoding F dep1 .
[0065] S8. Generate the first-stage repaired image Out rgb and Out dep . The formula is as follows:
[0066]
[0067] Among them, Dec() represents the decoding module, which is used to output the repaired result. Out rgb represents the result obtained by decoding F rgb1 , and Out dep represents the result obtained by decoding F dep1 .
[0068] S9. Fuse the prediction results to obtain the output result Out 1 of the first stage. The formula is as follows:
[0069] Out1 = SA rgb ·Out rgb + SA dep ·Out dep (17)
[0070] S10. Add F rgb1 and F dep1 to obtain the fused feature F 1 , and input it into the sub-network of the second stage for feature extraction. The formula is as follows:
[0071]
[0072] Among them, F 2 represents the feature extracted by the second-stage network.
[0073] S11. Use the decoding module to decode the feature F 2 to obtain the output result Out of the second stage 2 .
[0074] S12. Input the feature F 2 into the network of the third stage to obtain the feature F 3 . The formula is as follows:
[0075] F 3 = ORNet(F 2 ) (19)
[0076] where ORNet() represents a scale-invariant network with spatial attention, as Figure 4 shown.
[0077] S13. Use the decoding module to decode the feature F 3 to obtain the output result Out of the third stage 3 .
[0078] S14. Use the spatial attention module to predict the spatial attention map SA 1 of the output result Out 1 , the spatial attention map SA 2 of the output result Out 2 , and the spatial attention map SA 3 of the output result Out 3 .
[0079] S15. Fuse the output results of each stage to obtain the restored depth image I out . The formula is as follows:
[0080] I out = SA 1 ·Out 1 + SA 2 ·Out 2 + SA 3 ·Out 3 (20)
[0081] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the examples, and the protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the inventive concept of the present invention.
Claims
1. A method for repairing the depth image of a transparent object based on a multi-stage neural network, comprising the following steps: S1. Read the input color RGB image I rgb , and the corresponding depth image I dep ; S2. Scale the read picture, and use bilinear interpolation to adjust the image resolution to the resolution height×width required by the network; hright represents the vertical resolution of the image, and width represents the horizontal resolution of the image; S3. Obtain the segmented image I using the Deeplab v3+ network model seg , the formula is as follows: I seg = Deeplab(I rgb ) (1) wherein, Deeplab() represents the Deeplab v3+ network model; S4: Remove the noise in the depth image; use the predicted segmentation image I seg , and modify the depth values of the pixels in the depth image that belong to the transparent object to 0 in I dep , so as to remove the noise in the depth image; the formula is as follows: Among them represents the pixel at the i-th row and j-th column in the segmented image, represents the pixel at the i-th row and j-th column in the depth image; S5. Use a dynamic convolutional layer to initially extract image features, and the formula is as follows: Among them, Conv dynamic ( ) represents the dynamic convolution layer, F rgb0 represents the initial features extracted from I rgb , and F dep0 represents the initial features extracted from I dep ; S6. Use an encoder-decoder architecture to extract the features of the first stage; the formula is as follows: Among them, UNet() represents the UNet sub-network model of the encoder-decoder architecture; F rgb1 represents the feature obtained by performing feature extraction on F rgb0 , and F dep1 represents the feature obtained by performing feature extraction on F dep0 ; S7. Generate a spatial attention map; the formula is as follows: Among them, SAB( ) represents the spatial attention module, and SA rgb represents the spatial attention map obtained by decoding F rgb1 , and SA dep represents the spatial attention map obtained by decoding F dep1 ; S8. Generate the first-stage repaired image Out rgb and Out dep ; The formula is as follows: Among them, Dec( ) represents the decoding module, which is used to output the repaired result; Out rgb represents the result obtained by decoding F rgb 1, and Out dep represents the result obtained by decoding F dep1 ; S9. The prediction results are fused to obtain the output result Out of the first stage 1 ; The formula is as follows: Out 1 = SA rgb ·Out rgb + SA dep ·Out dep (7) S10. Add F rgb1 to F dep1 to obtain the fused feature F 1 , and input it into the sub-network of the second stage to extract features; the formula is as follows: Among which F 2 represents the features extracted by the second-stage network; S11. Use the decoding module to decode the feature F 2 to obtain the output result Out of the second stage 2 ; S12. Input feature F 2 into the network in the third stage to obtain feature F 3 ; The formula is as follows: P 3 = ORNet(F 2 ) (9) where ORNet() represents a scale-invariant network with spatial attention; S13. Use the decoding module to decode the feature F 3 to obtain the output result Out of the third stage 3 ; S14. Use the spatial attention module to predict the output result Out 1 of the spatial attention map SA 1 , and the output result Out 2 of the spatial attention map SA 2 , and the output result Out 3 of the spatial attention map SA 3 ; S15. The output results of each stage are fused to obtain the repaired depth image I out ; The formula is as follows: I out = SA 1 ·Out 1 + SA 2 ·Out 2 + SA 3 ·OUt 3 (10).
Citation Information
Patent Citations
Real-time semantic segmentation method based on context attention mechanism and information fusion
CN112541503A
Single-image rain removal method and system based on deep convolutional neural network, and storage medium
CN113450288A