A Method for Infrared Image Depth Estimation Supervised by Multi-Spectrum Images
By building a spectrum conversion module and a depth estimation module, using multi-spectral images for spectrum conversion and depth estimation, and iteratively optimized through loss function, the existing infrared image depth estimation method requires depth labels and multi-spectral image matching problem is solved, and the effect of accurately estimating infrared image depth without increasing cost and difficulty is achieved.
Patent Information
- Application Number
- CN202111531301.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-14
AI Technical Summary
Existing infrared image depth estimation methods require depth tags as supervision, which increases the cost and difficulty of data set acquisition, and the spectrum differences between multi-spectral images are large and difficult to match directly.
A multi-spectral image supervised infrared image depth estimation method is proposed. By building a spectrum conversion module and a depth estimation module, using multi-spectral images for spectrum conversion and depth estimation, and iteratively optimized through spectrum conversion loss, depth estimation loss and auxiliary loss to achieve depth information estimation.
Without the need for depth tags, the depth information of a single infrared image can be estimated more accurately, which reduces the difficulty and cost of obtaining the training data set, and solves the problem of appearance differences between multi-spectral images through the spectrum conversion network, and improves the accuracy of depth estimation.
Smart Images

Figure CN114494386B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for infrared image depth estimation, belonging to the technical field of computer graphics, and specifically to a method for infrared image depth estimation supervised by multi-spectrum images. Background Art
[0002] For many large and complex projects, there is an urgent need for a low-cost solution that can monitor the project quality for a long time and regularly detect project defects to ensure project safety and meet the needs of daily maintenance. Therefore, inspection robots have attracted the attention of many people. Inspection robots obtain multi-modal information through various sensors to complete a series of two-dimensional or three-dimensional tasks, such as defect detection in two-dimensional tasks and three-dimensional reconstruction in three-dimensional tasks. And depth information plays an important role in these tasks.
[0003] Considering the characteristic that infrared cameras are insensitive to the environment, whether there is an external light source or not, infrared cameras can directly measure the infrared radiation of objects and the environment. Therefore, obtaining depth information from infrared images obtained by monocular infrared cameras has more prominent advantages than other methods. For example, compared with active sensors such as lidar and depth cameras with structured light, which are expensive and have different defects in the face of complex scenes, the passive sensor based on standard imaging technology, namely the infrared camera, is cheaper, lighter in weight, and more adaptable. It can be more flexibly deployed on inspection robots and can adapt to different complex environments; compared with the method of obtaining depth information by combining RGB images obtained by monocular RGB cameras with depth estimation technology, which cannot perform well in night environments and low-light or even zero-light environments, such as in Document 1: Godard C, Aodha O M, Firman M, et al. Digging Into Self-Supervised Monocular Depth Estimation. International Conference on Computer Vision, 2019, 3827 - 3837., the infrared camera with the characteristic of environmental insensitivity can better obtain the geometric information of objects in the image and better utilize depth estimation technology for image depth estimation, and can overcome the disadvantage that the visible light spectrum cannot perform well in depth estimation in low-light environments.
[0004] At present, the infrared image depth estimation method for obtaining depth information from a single infrared image adopts a supervised approach. For example, in Document 2: Wang, Q., Zhao, H., Hu, Z. et al. Discrete convolutional CRF networks for depth estimation from monocular infrared images. Int. J. Mach. Learn. & Cyber. 12, 2021, 187–200., however, this method that requires depth labels increases the acquisition cost of the dataset. Without using depth labels, it is difficult to obtain depth information from a single infrared image. At this time, it is hoped to use a cheaper RGB camera to generate a supervision signal. However, there is a problem in multi-spectrum supervised infrared image depth estimation that there are large appearance differences between spectra and they cannot be directly matched. Therefore, this paper proposes a multi-spectrum image supervised infrared image depth estimation method. This method can use multi-modal information as a supervision signal without depth label supervision to obtain depth information from a single infrared image. Summary of the Invention
[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a multi-spectrum image supervised infrared image depth estimation method for the deficiencies of the prior art, which can more accurately estimate the depth information of a single infrared image.
[0006] To solve the above technical problem, the present invention discloses a multi-spectrum image supervised infrared image depth method, including the following steps:
[0007] Step 1, construct a spectrum conversion module: build a spectrum conversion network, input the multi-spectrum image into the spectrum conversion network model, and obtain the spectrum conversion image of the multi-spectrum image while ignoring the parallax, that is, convert the infrared right image and the RGB left image into the RGB right image and the infrared left image; the right image is the image obtained by the right camera in the binocular camera, the left image is the image obtained by the left camera in the binocular camera, the infrared indicates that the spectrum of the image is the infrared spectrum, and the RGB indicates that the spectrum of the image is the visible light spectrum, that is, an image composed of three channels of red, green, and blue.
[0008] Step 2, construct a depth estimation module: build a depth estimation network, and input the infrared right image into the depth estimation network model to obtain the parallax;
[0009] Step 3, construct a spectrum conversion loss module: use the spectrum conversion network model and the spectrum conversion image obtained in Step 1 to obtain the cyclic conversion image and the consistent reconstruction image, calculate the spectrum conversion loss, and use this loss to iteratively optimize the spectrum conversion network model;
[0010] Step 4, construct the depth estimation loss module: Use the disparity obtained in Step 2 to warp the image, calculate the depth estimation loss, and use this loss to iteratively optimize the depth estimation network model;
[0011] Step 5, construct the auxiliary loss module: Use the spectral conversion image and disparity obtained in Step 1 and Step 2 to calculate the auxiliary loss through image warping, and use this loss to iteratively optimize the spectral conversion network model.
[0012] Step 6, overall framework training: Unify the multi-spectral image dataset to a consistent number of channels through channel expansion, input the processed data into the spectral conversion module to obtain the spectral conversion image, preheat the spectral conversion module through the spectral conversion loss module, input it into the depth estimation module to obtain the disparity, and sequentially iteratively optimize the spectral conversion module and the depth estimation module through the depth estimation loss module, the spectral conversion loss module, and the auxiliary loss module, and use the trained depth estimation module and post-processing to achieve the depth estimation of a single infrared image.
[0013] Step 1 includes the following steps:
[0014] Step 1-1: Build the spectral conversion network including constructing the spectral conversion generator G and the spectral conversion discriminator D
[0015] Step 1-2: Input the multi-spectral image to obtain the spectral conversion image while ignoring the disparity, there is and That is, convert the infrared right image I A (p) to the RGB right image Convert the RGB left image I B (p) to the infrared left image Among them, the superscript fake indicates that this image is the output item obtained by the spectral conversion module.
[0016] Step 2 includes the following steps:
[0017] Step 2-1: Build the depth estimation network including constructing the depth estimation network M;
[0018] Step 2-2: Generate the left and right disparities d A (p) from the input single infrared right image I l and d r . Among them, the left disparity d l corresponds to the disparity of the RGB left image I B (p) in the multi-spectral image, and the right disparity d r corresponds to the disparity of the input infrared right image I A (p). l and r represent the left image and the right image respectively;
[0019] Step 3 includes the following steps:
[0020] Step 3-1: Obtain the cyclic transformation image and the consistent reconstruction image;
[0021] Step 3-2: Design the spectral transformation loss. When optimizing the spectral transformation network for the first iteration, the spectral transformation loss is utilized. The spectral transformation loss consists of L G and L D and is in the form of:
[0022]
[0023]
[0024] where λ cyc , λ rec , λ g and λ d are the weights of the loss terms and respectively, and are set to 10, 5, 1, 1 in the present invention. and are the adversarial losses of the generator G and the discriminator D of the cyclic generative adversarial network F-CycleGAN of the shared encoder, and this loss completes the task of image transformation. is the cyclic consistency loss, is the consistent reconstruction loss, and these two losses together complete the task of ignoring the parallax during spectral transformation. The subscripts cyc indicate that the variable is related to the cyclic consistency loss, rec indicates that the variable is related to the consistent reconstruction loss, adv indicates that the variable is related to the adversarial loss, and g and d respectively indicate that the variable is related to the adversarial loss of the generator G and the discriminator D.
[0025] Step 3-1 includes the following steps:
[0026] Step 3-1-1: Input the spectral transformation images and into the generator of the spectral transformation network to obtain the cyclic transformation images and That is, there are
[0027] and
[0028] Step 3-1-2: Input the multi-spectral images I A (p) and I B (p) into the generator of the spectral transformation network, but use the generator opposite to the one used to obtain the multi-spectral images before to obtain the consistent reconstruction images and That is, there are and
[0029] Step 4 includes the following steps:
[0030] Step 4-1: Image warping.
[0031] Step 4-2: Design the depth estimation loss. The depth estimation loss L is used when iteratively optimizing the depth estimation network MEN . The depth estimation loss L MEN is in the following form:
[0032]
[0033] where α ap , α ds and α lr are the corresponding loss weights, which are set to 1.0, 0.2, and 0.1 respectively in the present invention. They are the appearance reconstruction loss, the disparity smoothness loss, and the left-right disparity consistency loss respectively. The subscripts ap, ds, and lr in the upper and lower indices respectively indicate that the variable is related to the appearance reconstruction loss, the disparity smoothness loss, and the left-right disparity consistency loss.
[0034] Step 4-1 includes the following steps:
[0035] Step 4-1-1: Construct a warping module and define the warping operation ω, which performs the following operations. For p = (x, y), there is:
[0036]
[0037] where I l and I r represent the left image and the right image respectively, and represent the warped pseudo-left image and pseudo-right image respectively.
[0038] Step 4-1-2: Warp the infrared right image I r (p) and the infrared left image I l (p) respectively using the left disparity d l and the right disparity d r to perform the warping operation, obtaining the pseudo-infrared left image and the pseudo-infrared right image
[0039] Step 5 includes the following steps:
[0040] Step 5-1: Image warping of the auxiliary loss module. Using the warping operation ω constructed in Step 4-1, warp the infrared right image I A (p) and the RGB left image I B(p) respectively use the left parallax d l and right parallax d r Perform warping operation, Recalculate the infrared right composite image And RGB left composite image
[0041] Step 5-2: Design auxiliary loss. The auxiliary loss is used in the second iteration to optimize the spectrum conversion network. The auxiliary loss uses the original infrared right image I A (p) and RGB left image I B (p) and the image converted by the spectrum conversion network and the image obtained by warping the disparity obtained by the depth estimation network. Auxiliary loss The design is as follows:
[0042]
[0043] in, and It is a hyperparameter, which is set to 20 here. and are the infrared right composite image and the RGB left composite image, respectively, and N is the number of pixels in the image. The superscript aux indicates that the loss is related to the auxiliary loss.
[0044] Step 6 includes the following steps:
[0045] Step 6-1: In the process of data preprocessing, infrared image channel expansion is performed, that is, the infrared right image I a (p) Expanded to three-channel image I A (p). Specifically, a single-channel matrix I is replicated once on each of the three channels. a (p) , achieve channel consistency between infrared images and RGB images.
[0046] Step 6-2: Model framework warm-up. The model warm-up process only uses the spectrum conversion loss L in one iteration. G and L D The spectrum conversion network is trained. After the model warm-up process, the spectrum conversion network can obtain the spectrum conversion image from the multi-spectrum image while ignoring the parallax.
[0047] Step 6-3: Model framework training. The complete model training process is performed in one iteration. First, the spectrum conversion loss L is used. G and L D The spectrum conversion network is trained, and then the depth estimation loss L is used MEN Train the depth estimation and finally use the auxiliary loss Retrain the spectral conversion network.
[0048] Step 6-4: Model framework testing.
[0049] Step 6-4 includes the following steps:
[0050] Step 6-4-1: Expand the channels of the target infrared image, and then input it into the trained depth estimation network to obtain the right disparity.
[0051] Step 6-4-2: Post-process the right disparity, and use the formula to convert the disparity into depth, where is the depth, d is the disparity, B is the baseline length, and f is the camera focal length, to obtain the depth estimation result of the target infrared image.
[0052] Beneficial effects: The present invention has the following advantages: First, through the depth estimation method of the present invention, the depth estimation of a single infrared image can be completed, and no depth label is required as supervision during training, but only multi-spectral images are needed, which reduces the difficulty and cost of obtaining the training data set. Second, the present invention uses a spectral conversion network to solve the problem of large appearance differences between spectra, enabling multi-spectral images to be used for depth estimation. Finally, the auxiliary loss designed in the present invention enables the spectral conversion network to obtain clearer images and improves the accuracy of the depth estimation network for estimating depth. Description of the Drawings
[0053] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0054] Figure 1 is the schematic diagram of the processing flow of the present invention.
[0055] Figure 2a is an example of the RGB left image of the embodiment.
[0056] Figure 2b is an example of the infrared right image of the embodiment.
[0057] Figure 2c is an example of the right disparity estimation result of the embodiment. Detailed Embodiments
[0058] Embodiment
[0059] As Figure 1 shown, a method for depth of an infrared image supervised by multi-spectral images according to the present invention, the specific implementation process includes the following steps:
[0060] 1. Construct a spectral conversion module
[0061] Input: The multi-spectral image includes an infrared right image and an RGB left image.
[0062] Output: The spectral conversion image includes an RGB right image and an infrared left image.
[0063] 1.1 Building the spectral conversion network includes constructing a spectral conversion generator G and a spectral conversion discriminator D
[0064] The spectral conversion network can obtain the spectral conversion image from the spectral image to be converted while ignoring the parallax, that is, converting the infrared right image into an RGB right image and converting the RGB left image into an infrared left image, which solves the problem that there are large appearance differences between spectra and they cannot be directly matched in the depth estimation network. The spectral conversion network in this example is constructed based on F-CycleGAN in Document 3: Liang M, Guo X, Li H, et al. Unsupervised cross-spectral stereo matching by learning to synthesize. Proceedings of the AAAI Conference on Artificial Intelligence. 2019, 33(01): 8706-8713. The spectral conversion generator G includes an encoder F, two decoders G A and G B , and the spectral conversion discriminator D consists of two discriminators D A and D B . Among them, the encoder F includes 2 convolutional layers for downsampling and 4 residual blocks. The two decoders G A and G B include 2 convolutional layers for upsampling and 4 residual blocks. The discriminators D A and D B are composed of 5 convolutional layers.
[0065] 1.2 Inputting the multi-spectral image to obtain the spectral conversion image while ignoring the parallax, there are and That is, converting the infrared right image I A (p) into an RGB right image Converting the RGB left image I B (p) into an infrared left image Among them, the superscript fake indicates that the image is an output item obtained by the spectral conversion module.
[0066] 2. Constructing the depth estimation module
[0067] Input: Infrared right image.
[0068] Output: Disparity, including left disparity and right disparity.
[0069] 2.1 Building the depth estimation network includes constructing the depth estimation network M
[0070] The depth estimation network takes a single infrared right image as input and outputs the left and right disparities corresponding to this image. The converted infrared left image after being transformed by the spectral conversion network and the input single infrared right image are used together as the supervision signal. The depth estimation network in this example is constructed based on monodepth in reference 4: C. Godard, O. Mac Aodha, and G. J. Brostow. Unsupervised monocular depth estimation with left - right consistency. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017. The depth estimation network M includes an encoder and a decoder. Resnet18 is used as the encoder, and the decoder consists of 4 convolutional layers.
[0071] 2.2 For the input single infrared right image I A (p) generate the left and right disparities d l and d r . Among them, the left disparity d l corresponds to the disparity of the RGB left image I B (p) in the multi - spectral image, and the right disparity d r corresponds to the disparity of the input infrared right image I A (p). l and r represent the left image and the right image respectively.
[0072] 3. Construct the spectral conversion loss module
[0073] Input: Spectral conversion network, multi - spectral image, and spectral conversion image.
[0074] Output: Spectral conversion loss.
[0075] 3.1 Obtain the cyclic conversion image and the consistent reconstruction image
[0076] Step 1, input the spectral conversion image and into the generator of the spectral conversion network to obtain the cyclic conversion image and That is, there are and
[0077] Step 2, for the multi - spectral image I A(p) and I B (p) is input into the generator of the spectrum conversion network, but the generator opposite to the previously obtained multi-spectrum image is used to obtain a consistent reconstructed image and That is and
[0078] 3.2 Design of spectrum conversion loss
[0079] As Figure 1 shown in stage 1 of, the spectrum conversion loss is utilized when optimizing the spectrum conversion network for the first time in the iteration. The spectrum conversion loss consists of L G and L D and is in the form of:
[0080]
[0081]
[0082] Among them, λ cyc , λ rec , λ g and λ d are the weights of the loss terms and respectively, and are set to 10, 5, 1, 1 in the present invention. and are the adversarial losses of the generator G and the discriminator D of the cycle generative adversarial network F-CycleGAN of the shared encoder respectively, and this loss completes the task of image conversion. is the cycle consistency loss, is the consistent reconstruction loss. These two losses jointly complete the task of ignoring parallax during spectrum conversion. The subscripts cyc indicate that the variable is related to the cycle consistency loss, rec indicates that the variable is related to the consistent reconstruction loss, adv indicates that the variable is related to the adversarial loss, and g and d respectively indicate that the variable is related to the adversarial loss of the generator G and the discriminator D.
[0083] The design of the cycle consistency loss is:
[0084]
[0085] Among them, N is the number of picture pixels, and are the cycle conversion images.
[0086] The design of the consistent reconstruction loss is:
[0087]
[0088] where N is the number of pixels in the image. and is the consistent reconstructed image.
[0089] 4. Construct a depth estimation loss module
[0090] Input: Infrared right image, infrared left image in the spectral conversion image, left disparity, and right disparity.
[0091] Output: Depth estimation loss.
[0092] 4.1 Image warping
[0093] In this paper, a single infrared right image I A (p) is relabeled as I r (p), and the converted infrared left image is relabeled as I l (p).
[0094] Step 1: Construct a warping module and define a warping operation ω. For p = (x, y), the operation is as follows:
[0095]
[0096] where I l and I r represent the left image and the right image respectively, and represent the warped pseudo-left image and pseudo-right image respectively.
[0097] Step 2: Warp the infrared right image I r (p) and the infrared left image I l (p) using the left disparity d l and the right disparity d r respectively to obtain the pseudo-infrared left image and the pseudo-infrared right image
[0098] 4.2 Design the depth estimation loss
[0099] As shown in Stage 2 of Figure 1 , the depth estimation loss L MEN is used when iteratively optimizing the depth estimation network. The depth estimation loss L MEN has the following form:
[0100]
[0101] where α ap , α ds and α lris the corresponding loss weight, which is set to 1.0, 0.2, and 0.1 in this example. They are the appearance reconstruction loss, the disparity smoothness loss, and the left-right disparity consistency loss respectively. The superscripts and subscripts ap, ds, and lr in the above indicate that the variable is related to the appearance reconstruction loss, the disparity smoothness loss, and the left-right disparity consistency loss respectively. and represent the left term and the right term respectively. In this paper, only the left term will be explained because the right term can be derived by swapping the labels l and r. is the appearance reconstruction loss, which is designed as:
[0102]
[0103] where α is a hyperparameter, which is set to 0.9 in this example, SSIM is the structural similarity function, and N is the number of pixels in the image.
[0104] is the disparity smoothness loss, which is designed as:
[0105]
[0106] where and are the gradients of d l and I l respectively, and N is the number of pixels in the image.
[0107] is the left-right disparity consistency loss, which is designed as:
[0108]
[0109] where N is the number of pixels in the image.
[0110] 5. Construct the auxiliary loss module
[0111] Input: Infrared right image, RGB left image, left disparity, and right disparity.
[0112] Output: Auxiliary loss.
[0113] 5.1 Image warping in the auxiliary loss module
[0114] Using the warping operation ω constructed in step 4.1, warp the infrared right image I A (p) and the RGB left image I B (p) respectively using the left disparity d l and the right disparity d r for warping operations, and we have Recalculate to obtain the infrared right synthesized image and and the RGB left composite image
[0115] 5.2 Design auxiliary loss
[0116] As Figure 1 shown in stage 3 of [], the auxiliary loss is utilized when optimizing the spectral conversion network in the second iteration. The auxiliary loss makes use of the original infrared right image I A (p) and the RGB left image I B (p), as well as the image obtained after warping the image converted by the spectral conversion network and the disparity obtained by the depth estimation network. The auxiliary loss is designed as follows:
[0117]
[0118] wherein, and are hyperparameters, both set to 20 here, and are the infrared right composite image and the RGB left composite image respectively, and N is the number of pixels of the picture. The superscript aux indicates that the variable is related to the auxiliary loss.
[0119] 6. Overall framework training
[0120] The input multi-spectral image dataset is unified to a consistent number of channels through channel expansion, and the processed data is input into the spectral conversion module to obtain a spectral conversion image. The spectral conversion module is preheated through the spectral conversion loss module, and the disparity is obtained by inputting it into the depth estimation module. The spectral conversion module and the depth estimation module are sequentially iteratively optimized through the spectral conversion loss module, the depth estimation loss module, and the auxiliary loss module, and the depth estimation of a single infrared image is realized by using the trained depth estimation module and post-processing.
[0121] 6.1 Data preprocessing
[0122] Input: Multi-spectral image.
[0123] Output: Multi-spectral image with consistent channels.
[0124] During data preprocessing, infrared image channel expansion is performed, that is, the infrared right image I a (p) is expanded into a three-channel image I A (p). Specifically, the single-channel matrix I a (p) is copied once for each of the three channels to achieve channel consistency between the infrared image and the RGB image. And the sizes of all images are unified to 256×256.
[0125] 6.2 Model framework preheating
[0126] Input: Multispectral images with consistent channels.
[0127] Output: Spectral conversion image
[0128] In the model framework warm-up process, only the spectral conversion losses L G and L D are used to train the spectral conversion network in one iteration round. After the model warm-up process, the spectral conversion network can obtain the spectral conversion image from the multispectral image while ignoring the disparity. That is, only the Figure 1 stage 1 is executed.
[0129] 6.3 Model framework training
[0130] Input: Multispectral images with consistent channels
[0131] Output: Disparity, including left disparity and right disparity
[0132] In the model framework training process, in one iteration round, first the spectral conversion losses L G and L D are used to train the spectral conversion network, then the depth estimation loss L MEN is used to train the depth estimation, and finally the auxiliary loss is used to retrain the spectral conversion network. That is, the Figure 1 stage 1, stage 2, and stage 3 are executed.
[0133] 6.4 Model framework testing
[0134] Input: Target infrared image
[0135] Output: Depth estimation result
[0136] Step 1, expand the channels of the target infrared image, and then input it into the trained depth estimation network to obtain the right disparity.
[0137] Step 2, post-process the right disparity, and use the formula to convert the disparity to depth, where is the depth, d is the disparity, B is the baseline length, and f is the camera focal length, to obtain the depth estimation result of the target infrared image.
[0138] In this embodiment, Figure 2a is an example of the RGB left image, Figure 2b is an example of the infrared right image, Figure 2c is an example of the right disparity estimation result. Through the depth estimation method of this embodiment, the depth estimation of a single infrared image is completed, as shown in Figure 2c , and no depth label is required as supervision during training, but only multispectral images, such as Figure 2a andFigure 2b As shown, this reduces the difficulty and cost of obtaining the training dataset.
[0139] The present invention provides a method for depth estimation of infrared images with multi-spectrum supervision. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.
Claims
1. A method for infrared image depth estimation supervised by multi - spectral images, characterized in that: It includes the following steps, Step 1, construct a spectral conversion module: Build a spectral conversion network model, input the multi - spectral image into the spectral conversion network model, and obtain a spectral conversion image of the multi - spectral image ignoring parallax; Step 2, construct a depth estimation module: Build a depth estimation network model, input the infrared right image into the depth estimation network model to obtain parallax; Step 3, construct a spectral conversion loss module: Use the spectral conversion network model and spectral conversion image obtained in Step 1 to obtain a cyclic conversion image and a consistent reconstruction image, calculate the spectral conversion loss, and use the spectral conversion loss to iteratively optimize the spectral conversion network model; Step 4, construct a depth estimation loss module: Use the parallax obtained in Step 2 for image warping, calculate the depth estimation loss, and use the depth estimation loss to iteratively optimize the depth estimation network model; Step 5, construct an auxiliary loss module: Use the spectral conversion image and parallax obtained in Step 1 and Step 2 to calculate the auxiliary loss through image warping, and use the auxiliary loss to iteratively optimize the spectral conversion network model; Step 6, overall framework training: Unify the multi - spectral image dataset to a consistent channel through channel expansion, input the processed data into the spectral conversion module to obtain a spectral conversion image ignoring parallax; Pre - heat the spectral conversion module through the spectral conversion loss module, input the infrared right image into the depth estimation module to obtain parallax; Sequentially iteratively optimize the spectral conversion module and the depth estimation module through the spectral conversion loss module, the depth estimation loss module, and the auxiliary loss module, and use the trained depth estimation module and post - processing to achieve the depth estimation of a single infrared image; Step 4 includes the following steps, Step 4 - 1: Image warping; Step 4-2: Design the depth estimation loss; utilize the depth estimation loss L when iteratively optimizing the depth estimation network MEN ; Depth estimation loss L MEN The form is as follows: Among them, α ap , α ds and α lr represent the appearance reconstruction loss weight, the disparity smoothness loss weight, and the left - right disparity consistency loss weight; and represent the appearance reconstruction loss, the disparity smoothness loss, and the left - right disparity consistency loss; ap, ds, and lr respectively indicate that the variables are related to the appearance reconstruction loss, the disparity smoothness loss, and the left - right disparity consistency loss.
2. A method for infrared image depth estimation supervised by multi - spectral images according to claim 1, characterized in that: Step 1 includes the following steps, Step 1-1: Build a spectrum conversion network, including constructing a spectrum conversion generator G and a spectrum conversion discriminator D; the spectrum conversion generator G includes an encoder F and two decoders G A and G B , and the spectrum conversion discriminator D consists of two discriminators D A and D B ; Step 1 - 2: Input the multi - spectral image to obtain a spectral conversion image ignoring parallax, and Convert the infrared right image I A (p) into an RGB right image Convert the RGB left image I B (p) into an infrared left image where p represents the pixel coordinates of any point within the image, the subscript A represents the infrared spectrum, the subscript B represents the RGB spectrum, the superscript fake indicates that the image is an output item obtained by the spectrum conversion module, and RGB indicates that the image consists of three channels: red, green, and blue.
3. A method for infrared image depth estimation supervised by multi - spectral images according to claim 2, characterized in that: Step 2 includes the following steps, Step 2 - 1: Build a depth estimation network, including constructing a depth estimation network M; Step 2-2: The single input infrared right image I A (p) generates the left and right disparities d l and d r ; among them, the left disparity d l corresponds to the disparity of the RGB left image I B (p) in the multi-spectral image, and the right disparity d r corresponds to the disparity of the input infrared right image I A (p); l and r represent the left image and the right image respectively.
4. A method for infrared image depth estimation supervised by multi - spectral images according to claim 3, characterized in that: Step 3 includes the following steps, Step 3 - 1: Obtain a cyclic conversion image and a consistent reconstruction image; Step 3-2: Design the spectrum conversion loss; the spectrum conversion loss consists of L G and L D and is expressed as: Among them, and are the adversarial losses of the generator G and the discriminator D respectively; represents the cycle consistency loss, represents the consistent reconstruction loss; λ cyc 、λ rec 、λ g and λ d are the weights of the loss terms and respectively; λ indicates that the variable is a hyperparameter, cyc indicates that the variable is related to the cycle consistency loss, rec indicates that the variable is related to the consistent reconstruction loss, adv indicates that the variable is related to the adversarial loss, and g and d respectively indicate that the variable is related to the adversarial loss of the generator G and the discriminator D.
5. A method for infrared image depth estimation supervised by multi - spectral images according to claim 4, characterized in that: Step 3 - 1 includes the following steps: Step 3-1-1: Input the spectrum conversion image and into the generator of the spectrum conversion network to obtain the cyclic conversion image and That is: and Step 3-1-2: Input the infrared right image I A (p) and the RGB left image I B (p) into the generator of the spectral conversion network, and use the generator opposite to the one for obtaining the multi-spectral image to obtain a consistent reconstructed image and That is: and 6. A method for infrared image depth estimation supervised by multi - spectral images according to claim 5, characterized in that: Step 4 - 1 includes, Step 4 - 1 - 1: Construct a warping module, define a warping operation ω, for p=(x,y), Among them, I l and I r respectively represent the left image and the right image, and respectively represent the warped pseudo-left image and pseudo-right image. The integer x ∈ [1, H] represents the horizontal coordinate of the pixel, and the integer y ∈ [1, W] represents the vertical coordinate of the pixel. W and H are the width and length of the image respectively; Step 4-1-2: Warp the infrared right image I r (p) and the infrared left image I l (p) respectively using the left parallax d l and the right parallax d r to perform a warping operation, obtaining a pseudo-infrared left image and a pseudo-infrared right image 7. A method for infrared image depth estimation supervised by multi - spectral images according to claim 6, characterized in that: Step 5 includes, Step 5 - 1: Image warping of the auxiliary loss module Using the warping operation ω constructed in step 4-1, warp the infrared right image I A (p) and the RGB left image I B (p) respectively using the left disparity d l and the right disparity d r to perform the warping operation, that is Recalculate to obtain the infrared right composite image and the RGB left composite image Step 5-2: Design the auxiliary loss; when the spectral conversion network is iteratively optimized for the second time, the auxiliary loss is utilized. The auxiliary loss uses the original infrared right image I A (p) and the RGB left image I B (p), as well as the image obtained after warping the image converted by the spectral conversion network and the disparity obtained by the depth estimation network; the auxiliary loss is designed as follows: Among them, and represent the loss term weights, α represents the variable which is a hyperparameter, and aux represents the variable related to the auxiliary loss; and represent the infrared right composite image and the RGB left composite image respectively, and N is the number of picture pixels.
8. A method for infrared image depth estimation supervised by multi - spectral images according to claim 7, characterized in that: Step 6 includes, Step 6 - 1: During data pre - processing, infrared image channel expansion is performed. The infrared right image is expanded into a three - channel image; the single - channel image matrix is copied once for each of the three channels to achieve consistent channels for infrared images and RGB images; Step 6-2: Model framework preheating: In an iteration round, use the spectral conversion losses L G and L D to train the spectral conversion network, and obtain a spectral conversion image ignoring parallax from the multi-spectrum image; Step 6-3: Model framework training: In one iteration, use the spectral conversion losses L G and L D to train the spectral conversion network, use the depth estimation loss L MEN to train the depth estimation network, and use the auxiliary loss to retrain the spectral conversion network; Step 6 - 4: Model framework testing.
9. A method for infrared image depth estimation supervised by multi - spectral images according to claim 8, characterized in that: Step 6 - 4 includes the following steps, Step 6 - 4 - 1, expand the target infrared image in channels and input it into the trained depth estimation network to obtain the right disparity; Step 6-4-2, perform post-processing on the right parallax. Through the formula convert the parallax to depth, where is the depth, d is the parallax, B is the baseline length, and f is the camera focal length, to obtain the depth estimation result of the target infrared image.
Citation Information
Patent Citations
Monocular image depth estimation method and system
CN107204010A
Unsupervised monocular depth estimation method based on generative adversarial network
CN110443843A