Single infrared image-based monocular depth estimation method and apparatus
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- IND ACAD COOP GRP OF SEJONG UNIV
- Filing Date
- 2022-09-21
- Publication Date
- 2026-08-03
Smart Images

Figure 112022099259495-PAT00047_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a single thermal image-based depth estimation method and apparatus. Background Technology
[0002] In recent years, monocular depth estimation has become an important component in the fields of computer vision and robotics, along with applications such as autonomous driving and augmented reality, due to its potential to replace LiDAR.
[0003] Many studies have shown that monocular depth estimation performance can be improved through supervised learning using LiDAR as ground truth.
[0004] However, in supervised learning, there is a problem in that data collection is difficult due to the high cost of LiDAR, and accurate and dense depth information cannot be estimated due to the sparse depth information from LiDAR.
[0005] To address this problem, researchers proposed a self-supervised learning method that does not require actual depth information throughout the learning process.
[0006] Many researchers are proposing various self-supervised learning-based monocular depth estimation methods, and the performance gap between supervised and self-supervised learning is narrowing.
[0007] However, since this methodology uses color images (RGB images) as input, it does not guarantee performance in low-light situations such as at night, and also has the problem of being sensitive to changes in the external environment, such as rain or cloudiness, due to the inherent limitations of RGB sensors.
[0008] As a realistic alternative to these problems, long-wavelength thermal imaging cameras, which are robust against various environmental changes, can be used. Unlike color imaging, thermal imaging cameras record the radiant energy of objects as an image, which is unaffected by changes in the external environment.
[0009] However, replacing the input of RGB-based depth estimation with thermal images remains a difficult problem due to spectral differences.
[0010] Figure 1 is a diagram illustrating a thermal image-based depth estimation process according to the prior art.
[0011] Referring to Fig. 1, conventionally, appearance matching loss is calculated using a depth map predicted from a color image and a thermal image.
[0012] However, in conventional technology, there are limitations to depth estimation due to the domain gap between thermal images and color images. Additionally, problems arise where learning is not effective due to the disadvantages of blurred edges in thermal images and low contrast within the image. Prior art literature
[0013] KR Registered Patent Publication 10-1947782 The problem to be solved
[0014] To solve the problems of the aforementioned prior art, the present invention proposes a single thermal image-based monocular depth estimation method and apparatus capable of predicting more accurate depth information during the day and night by applying various learning methods to compare depth information obtained from a color image and depth information predicted from a thermal image across both local and global regions. means of solving the problem
[0015] To achieve the above-mentioned purpose, according to one embodiment of the present invention, a single thermal image-based monocular depth estimation device is provided, comprising: a processor; and a memory connected to the processor, wherein the memory stores program instructions executed by the processor to generate a first depth map by inputting one input color image among stereo color images into a deep learning-based first depth network, generate a second depth map by inputting a thermal image corresponding to the input color image into a deep learning-based second depth network, repeat the learning of the first depth network by comparing the estimated color image estimated using the first depth map with the input color image, repeat the learning of the second depth network by comparing the first depth map and the second depth map according to the repetition of the learning of the first depth network, and generate a second depth map from a newly input thermal image using the learned second depth network.
[0016] The above program instructions can repeat the learning of the second depth network by calculating the self-guided loss of the first depth map and the second depth map.
[0017] The above self-map loss may include a first loss based on a combination of the structural simulation index measure (SSIM) and L1 distance of the first depth map and the second depth map, a second loss calculated using features obtained by inputting the first depth map and the second depth map into a VGG network, and a third loss calculated using global descriptors and local descriptors extracted from the first depth map and the second depth map, respectively.
[0018] The above third loss can be defined as Patch-NetVLAD loss.
[0019] A preset weight may be applied to each of the first loss and the third loss.
[0020] The first depth map above can be defined as a pseudo-label that is updated as the learning of the first depth network is repeated.
[0021] The above program instructions can repeat the training of the first depth network by calculating the appearance matching loss and image matching loss between the estimated color image and the input color image.
[0022] The above image matching loss may include a perceptual loss calculated using features obtained by inputting the estimated color image and the input color image into a VGG network, and a Patch-NetVLAD loss calculated using global descriptors and local descriptors extracted from the estimated color image and the input color image, respectively.
[0023] According to another aspect of the present invention, a method for estimating monocular depth based on a single thermal image in a device comprising a processor and a memory is provided, comprising the steps of: inputting one input color image among stereo color images into a deep learning-based first depth network to generate a first depth map; inputting a thermal image corresponding to the input color image into a deep learning-based second depth network to generate a second depth map; repeating the learning of the first depth network by comparing the estimated color image estimated using the first depth map with the input color image; repeating the learning of the second depth network by comparing the first depth map and the second depth map according to the repetition of the learning of the first depth network; and generating a second depth map from a newly input thermal image using the learned second depth network.
[0024] According to another aspect of the present invention, a computer program stored in a computer-readable recording medium for performing the above-described method is provided. Effects of the invention
[0025] According to the present embodiment, the depth network is trained by comparing depth information obtained from a high-precision color image with depth information predicted from a thermal image, which has the advantage of increasing the accuracy of depth estimation. Brief explanation of the drawing
[0026] Figure 1 is a diagram illustrating a thermal image-based depth estimation process according to the prior art. FIG. 2 is a diagram illustrating a single thermal image-based monocular depth estimation framework according to a preferred embodiment of the present invention. FIG. 3 is a diagram illustrating the Patch-NetVLAD loss according to the present embodiment. FIG. 4 is a diagram showing the accuracy of depth estimation according to the present embodiment. FIG. 5 is a diagram illustrating the configuration of a depth estimation device according to the present embodiment. Specific details for implementing the invention
[0027] The present invention is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention.
[0028] The terms used herein are merely for describing specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0029] Furthermore, the components of the embodiments described with reference to each drawing are not limited to the respective embodiments and may be implemented to be included in other embodiments within the scope of maintaining the technical spirit of the present invention. It is also obvious that multiple embodiments may be re-implemented as a single embodiment that integrates multiple embodiments, even if a separate description is omitted.
[0030] Furthermore, in the description referring to the attached drawings, identical components are assigned the same or related reference numerals regardless of drawing symbols, and redundant descriptions thereof are omitted. In describing the present invention, if it is determined that a detailed description of related prior art could unnecessarily obscure the essence of the present invention, such detailed description is omitted.
[0032] This embodiment proposes a Self-Guided Framework that uses depth information generated from a color image as a pseudo-label.
[0033] FIG. 2 is a diagram illustrating a single thermal image-based monocular depth estimation framework according to a preferred embodiment of the present invention.
[0034] Referring to FIG. 2, one input color image among stereo color images (for example, the left color image among the left and right color images, ) is a deep learning-based depth network N R Color image-based depth map input into (1st depth network) Generates (1st depth map).
[0035] According to the present embodiment, an estimated color image (estimated using a first depth map) The first depth network is trained repeatedly by comparing the input color image, and through this, a precise color image-based depth map It generates, This is used as a doctor's label.
[0036] In addition, according to the present embodiment, a thermal image corresponding to an input color image 2nd depth network N T Input into to generate a thermal image-based depth map (second depth map).
[0037] According to the present embodiment, a color image depth map for training a second depth network Self-Guided Loss of the first depth map and second depth map using pseudo-labels Calculate )
[0038] According to the present embodiment As shown in the formula below, is a combination of the structural simulation index measure (SSIM) commonly used in image translation and the L1 distance. and perceptual loss ( ) and Patch-NetVLAD loss( It consists of ).
[0039]
[0040] Here, is a static value that readjusts the loss and can be set to 10. M is the contour match loss. It is an automatic mask calculated from.
[0041] SSIM and L1 distance loss, the most traditional losses in image transformation tasks Using the first depth map and pseudo-label Wow, the second depth map It compares, and is equal to the formula below.
[0042]
[0043] Here, Is It is the value of, and Is It is the variance of, and , , am.
[0044] SSIM compares two values of luminance (l) and contrast and structure (cs) to compare the similarity between two images, and L1 distance directly compares the two image values.
[0045] In addition, according to the present embodiment, pseudo-labels using perceptual loss and predicted depth map from thermal image Measures high-level perceptual and semantic differences between the liver.
[0046] Perceptual loss uses features obtained by inputting two depth maps into a VGG network trained for image classification, and the formula is as follows.
[0047]
[0048] In addition, to further enhance local and global information, the Patch-NetVLAD loss is applied by estimating and comparing the image's local and global descriptors.
[0049] FIG. 3 is a diagram illustrating the Patch-NetVLAD loss according to the present embodiment.
[0050] Referring to FIG. 3, depth information predicted from a thermal image ( Local descriptors () from the depth information (Pseudo-Label) predicted from ) and color images ) and global technician( ) is estimated using the Patch-NetvLAD model and then compared using the difference operation.
[0051] Through this, depth information predicted from the thermal image ( ) becomes similar to the global / local depth information of the depth information (Pseudo-Label) predicted from the color image.
[0052] Patch-NetVLAD loss is NetVLAD network Global descriptors and local descriptors estimated in the first depth map using ) and estimated global and local descriptors from the second depth map( It is calculated from ), and the formula is as follows.
[0053]
[0054] is the size of multiple patches, and is the total number of patches.
[0055] Similar to perceptual loss, the comparison of global technologists is Improves image recognition and semantic information.
[0056] In addition, through a comparison of local technicians obtained using the patch The regional meaning information and details of are improved.
[0057] In the case of a Self-Guided Framework using pseudo-labels as in this embodiment, the performance of the pseudo-labels Determines the performance of.
[0058] According to the present embodiment, to improve the performance of pseudo-labels, image matching loss Suggests.
[0059] Since SSIM relies only on low-level differences between pixels, the image matching loss consists of a perceptual loss and a Patch-NetVLAD loss that calculates high-level similarity.
[0060] The semantic information of pseudo-labels is further enhanced through perceptual loss and comparison using global descriptors of Patch-NetVLAD loss. Additionally, local descriptors of Patch-NetVLAD loss improve the accuracy of detailed regions.
[0061] Perceptual loss of image registration loss and Patch-NetVLAD loss It proceeds similarly to what was explained in self-supervised learning, with the input changed only from a depth map to a color image. The above equation is formulated as follows.
[0062]
[0063] In conventional self-supervised learning, a new left image is synthesized using the predicted depth image and the right color image as shown in the conventional Fig. 1, and then compared with the actual left image using a shape matching loss. However, this formula has the disadvantage of using only SSIM and L1 distance when comparing the two images.
[0064] Since the two equations only look at actual values, it is difficult to compare semantic information and there were limitations in comparing local information. To overcome these shortcomings, an image matching loss consisting of VGG loss (perceptual loss) and Patch-NetVLAD loss is added.
[0065] In this way, adding image matching loss improves the performance of pseudo-labels, which in turn improves the performance of depth estimation models using thermal images as input.
[0066] The proposed Self-Guided Framework demonstrates high performance improvement when applied to any depth estimation model, and as shown in Figure 4, it can be seen that high performance improvements were achieved in both local and global aspects even at night.
[0067] FIG. 5 is a diagram illustrating the configuration of a depth estimation device according to the present embodiment.
[0068] As illustrated in FIG. 5, the device according to the present embodiment may include a processor (500) and a memory (502).
[0069] The processor (500) may include a CPU (central processing unit) capable of executing computer programs or other virtual machines.
[0070] The memory (502) may include a non-volatile storage device such as a fixed hard drive or a removable storage device. The removable storage device may include a compact flash unit, a USB memory stick, etc. The memory (502) may also include a volatile memory such as various random access memory and may be defined as a computer-readable recording medium.
[0071] In the memory (502) according to the present embodiment, program instructions for single thermal image-based monocular depth estimation are stored.
[0072] The program instructions according to the present embodiment input one of the stereo color images into a deep learning-based first depth network to generate a first depth map, input a thermal image corresponding to the input color image into a deep learning-based second depth network to generate a second depth map, repeat the learning of the first depth network by comparing the estimated color image estimated using the first depth map with the input color image, repeat the learning of the second depth network by comparing the first depth map and the second depth map according to the repetition of the learning of the first depth network, and generate a second depth map from a newly input thermal image using the learned second depth network.
[0073] The program instructions according to the present embodiment calculate the self-guided loss of the first depth map and the second depth map and repeat the learning of the second depth network.
[0074] Here, the self-map loss may include a first loss based on a combination of the SSIM (structural simulation index measure) and L1 distance of the first depth map and the second depth map, a second loss calculated using features obtained by inputting the first depth map and the second depth map into a VGG network, and a third loss calculated using global descriptors and local descriptors extracted from the first depth map and the second depth map, respectively, and the third loss is defined as the Patch-NetVLAD loss.
[0075] The embodiments of the present invention described above are disclosed for illustrative purposes only, and those skilled in the art with ordinary knowledge of the present invention may make various modifications, changes, and additions within the spirit and scope of the present invention, and such modifications, changes, and additions should be considered to fall within the scope of the following claims.
Claims
Claim 1 As a single thermal imaging-based monocular depth estimation device, a processor; and includes a memory connected to the processor, wherein the memory inputs one input color image among stereo color images into a deep learning-based first depth network to generate a first depth map, inputs a thermal image corresponding to the input color image into a deep learning-based second depth network to generate a second depth map, calculates an appearance matching loss and an image matching loss between an estimated color image estimated using the first depth map and the input color image, and repeats the learning of the first depth network for a predetermined number of iterations or until the sum of the appearance matching loss and the image matching loss converges, calculates a self-guided loss of the first depth map and the second depth according to the repetition of the learning of the first depth network, and repeats the learning of the second depth network for a predetermined number of iterations or until the self-guided loss converges, and the repetition of the learning of the first and second depth networks includes a stereo color image and a thermal image corresponding to each of the stereo color images. A single thermal image-based monocular depth estimation device that stores program instructions executed by the processor to generate a second depth map from a newly input thermal image using a trained second depth network, which is performed using training data. Claim 2 delete Claim 3 A single thermal image-based monocular depth estimation device according to claim 1, wherein the self-map loss comprises a first loss based on a combination of the SSIM (structural simulation index measure) and L1 distance of the first depth map and the second depth map, a second loss calculated using features obtained by inputting the first depth map and the second depth map into a VGG network, and a third loss calculated using global descriptors and local descriptors extracted from the first depth map and the second depth map, respectively. Claim 4 In paragraph 3, the third loss is a single thermal image-based monocular depth estimation device defined as Patch-NetVLAD loss. Claim 5 A single thermal image-based monocular depth estimation device in which a preset weight is applied to each of the first loss and the third loss in paragraph 3. Claim 6 A single thermal image-based monocular depth estimation device, wherein the first depth map is defined by pseudo-labels that are updated as the learning of the first depth network is repeated. Claim 7 delete Claim 8 A single thermal image-based monocular depth estimation device according to claim 1, wherein the image matching loss comprises a perceptual loss calculated using features obtained by inputting the estimated color image and the input color image into a VGG network, and a Patch-NetVLAD loss calculated using global descriptors and local descriptors extracted from the estimated color image and the input color image, respectively. Claim 9 A method for estimating monocular depth based on a single thermal image in a device including a processor and memory, comprising: a step of generating a first depth map by inputting one input color image among stereo color images into a deep learning-based first depth network; a step of generating a second depth map by inputting a thermal image corresponding to the input color image into a deep learning-based second depth network; a step of calculating an appearance matching loss and an image matching loss between the estimated color image and the input color image using the first depth map, and repeating the training of the first depth network for a predetermined number of iterations or until the sum of the appearance matching loss and the image matching loss converges; and a step of calculating a self-guided loss between the first depth map and the second depth based on the iteration of the training of the first depth network, and repeating the training of the second depth network for a predetermined number of iterations or until the self-guided loss converges. A single thermal image-based monocular depth estimation method comprising the step of generating a second depth map from a newly input thermal image using a learned second depth network, wherein the iteration of learning of the first and second depth networks is performed using training data including a stereo color image and a thermal image corresponding to each of the stereo color images. Claim 10 A computer program stored on a computer-readable recording medium that performs the method according to paragraph 9.