Thermal image depth estimation method based on manual feature camouflage
By using manual feature camouflage and thermal image enhancement techniques in the monocular image depth estimation method, the existing methods have solved the problem of low accuracy and time-consuming depth estimation in harsh environments, and more efficient and real-time depth map estimation is achieved.
Patent Information
- Application Number
- CN202410426737.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-04-10
AI Technical Summary
The existing self-supervised monocular image depth estimation method uses low accuracy, long time, and high performance requirements for hardware equipment, making it difficult to meet the real-time and efficient requirements.
The thermal image depth estimation method based on manual feature camouflage is adopted to enhance the contrast and detail extraction of thermal images, and long-distance information is mined using manual feature camouflage structures, and edge-aware smoothing loss and thermal image reconstruction loss are introduced to improve the depth estimation ability of the model under harsh lighting conditions.
The accuracy of depth map estimation under harsh lighting conditions is improved, the calculation burden is reduced, the real-time and efficiency of the model are enhanced, and the accurate depth map in the poor lighting environment can be estimated more quickly.
Smart Images

Figure CN119559230B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to computer vision and image processing technology, and in particular to a thermal image depth estimation method based on manual feature camouflage. Background Art
[0002] Depth estimation is a method of estimating the vertical distance from each pixel in the scene to the camera plane from image data. According to the number of images, it can be divided into monocular image depth estimation and stereo image depth estimation. However, due to the inconsistency between the camera size and the vehicle platform used to shoot the stereo image, stereo image depth estimation cannot be effectively applied in reality. Existing monocular image depth estimation methods are mainly divided into the following categories: (1) monocular depth estimation based on parameter learning; (2) monocular depth estimation methods based on parameter-free learning; (3) monocular depth estimation methods based on supervised learning; (4) monocular depth estimation methods based on semi-supervised learning; (5) monocular depth estimation methods based on unsupervised learning. Since the monocular depth estimation method based on self-supervised learning only needs to input a monocular image to accurately estimate the depth map when there is a true depth value, this type of method has attracted widespread attention from academia and industry, and a large number of novel methods have been proposed. The main focus of these methods is how to efficiently use the information of monocular images and true depth values to estimate an accurate depth map.
[0003] At present, there are seven main types of monocular depth estimation methods based on self-supervision: (1) methods of improving network structure; (2) methods of introducing auxiliary information; (3) methods of improving loss functions; (4) methods based on classification; (5) methods using conditional random fields; (6) methods using generative adversarial networks; and (7) methods based on partial depth information. The fourth type of method uses a lightweight network model to speed up the algorithm while ensuring the reliability of the results. The other five types of methods prioritize accuracy. Existing depth estimation methods have made great progress in estimating accurate depth maps from monocular indoor images. However, for scenes with poor lighting conditions, existing monocular image depth estimation methods based on self-supervision need to be further optimized.
[0004] In addition, with the popularity and widespread use of mobile devices such as cameras and vehicle-mounted shooting platforms, it has become very convenient for people to obtain image data of different scenes. These large amounts of image data of different scenes have posed a series of new challenges to the self-supervised monocular image depth estimation method: (a) Although the accuracy of the existing self-supervised monocular image depth estimation method has been greatly improved, the network training time and video memory usage have also gradually increased. Therefore, it cannot meet the real-time requirements in high-level computer vision practical applications such as autonomous driving and 3D reconstruction; (b) The existing self-supervised monocular image depth estimation method is limited by the training data set and will lose its original accuracy in different environments. For example, some models are only applicable to image data collected indoors, but it is difficult to obtain accurate depth maps in outdoor scenes. (c) The existing self-supervised monocular image depth estimation method is very time-consuming, especially when processing large-scale outdoor image data in harsh weather environments. Therefore, high-performance computing equipment is required as a prerequisite for training the network and estimating the depth, otherwise research work cannot be carried out.
[0005] In summary, the above problems seriously hinder the theoretical development and application promotion of self-supervised monocular image depth estimation methods, and reveal the shortcomings of existing self-supervised monocular image depth estimation methods. Therefore, people urgently need a fast and effective monocular image depth estimation algorithm to quickly estimate accurate depth maps from outdoor image data in harsh weather environments.
[0006] In recent years, deep learning technology has achieved good application results in image classification, target detection, target segmentation, text classification, natural language processing and speech recognition. Inspired by this, some researchers at home and abroad have tried to apply deep learning technology to the problem of monocular image depth estimation, and have made good research progress. The following papers have been published: "Maximizing Self-supervision from Thermal Image for Effective Self-supervised Learning of Depth and Ego-motion" discloses a monocular thermal image depth estimation method that improves the self-supervised signal of thermal images; "AdaBins: Depth Estimation Using Adaptive Bins" discloses a monocular image depth estimation method based on Transformer architecture; "Self-Supervised Lightweight DepthEstimation in Endoscopy Combining CNN and Transformer" discloses a monocular image depth estimation method that combines a lightweight network of CNN and Transformer; the above monocular image depth estimation methods focus on improving the accuracy of the estimated depth map.
[0007] However, the above methods still face the following challenges when applied to outdoor image data in harsh weather environments: (1) When processing outdoor image data in harsh weather environments, the depth maps estimated by these methods have low accuracy and are difficult to meet the requirements of some high-level application systems, such as autonomous driving, 3D reconstruction, virtual reality, navigation and positioning, human posture detection, obstacle detection, visual SLAM, etc.; (2) Existing monocular image depth estimation methods based on self-supervision are very time-consuming and are difficult to meet the real-time requirements of high-level real-world applications; (3) As the amount of image data increases and the network becomes more complex, the performance requirements of hardware devices for existing monocular image depth estimation methods based on self-supervision are also gradually increasing. Summary of the invention
[0008] Purpose of the invention: The purpose of the present invention is to address the deficiencies in the prior art and to provide a thermal image depth estimation method based on manual camouflage features, which utilizes the powerful self-supervision capability of the enhanced thermal image itself under harsh environmental conditions to further make it possible to estimate an accurate depth map.
[0009] Technical solution: A thermal image depth estimation method based on manual feature camouflage of the present invention comprises the following steps:
[0010] Step S1: Input the thermal image and the internal parameters of the camera. There are N thermal images with a width of W and a height of H in each image. Each thermal image has internal and external parameters of the corresponding camera. Where i represents the serial number of the thermal image, and the internal parameters of the camera of the i-th thermal image are K i , the external parameters of the camera are T i ;
[0011] Step S2, calculating and enhancing the thermal image;
[0012] First, for each source thermal image The histogram h composed of its thermal radiation values is extracted θ ;
[0013] h θ =s θ ,θ=t0,t1,…,t max ;
[0014] [t θ ,t θ+1 ] is the temperature interval of the histogram of the source thermal image, s θ Yes θ ,t θ+1 ]The number of samples in the temperature interval, t max It refers to the maximum temperature of the source thermal image histogram;
[0015] Then all histograms h θ According to the thermal radiation value v ra Rearrange to get the thermal image after the thermal radiation values are rearranged The formula for reordering is as follows:
[0016]
[0017] Among them, γ θ is the scaling factor, t′ θ The new offset for scaling each subhistogram, It refers to the proportional factor proportional to the number of sub-histograms. If no samples are observed in a certain temperature interval, the invalid area is directly removed by the re-ordering formula. After re-processing by calculation, different temperature intervals are extended or shortened according to the number of samples observed to improve the image contrast. In order to maintain the temporal consistency between adjacent images, the thermal images of each scene are calculated in groups of at least 4.
[0018] Finally, the thermal image after rearranging the thermal radiation values Input the thermal image processing structure, first divide the image into non-overlapping local images internally, then use the predefined limit threshold to perform histogram equalization processing on each local image to obtain a thermal image with enhanced details The resolution is still H×W; the thermal image processing structure here mainly divides the image into non-overlapping local images, and uses a predefined clipping threshold to perform histogram equalization processing on each local image;
[0019] Step S3, constructing a feature extraction encoder and calculating a feature vector set;
[0020] Based on thermal image Construct the feature map of the first stage For each feature map Five feature maps of different scales are extracted And the scales of the feature maps are: H×W, and Finally, the five feature maps are combined into a feature set
[0021] Step S4: construct a feature camouflage structure to hide and filter the information in the feature set; first, Extracting hidden feature maps Then, Input into the feature camouflage structure to calculate the uncertain feature map D t , the specific method is:
[0022] Step S4.1. Use |sinx|, |cosx|, |2sin x cos x|, and sin 2 x constitutes a manual feature disguiser, and the fifth-size feature map of step S3 is Input into the manual feature disguiser to obtain the disguised feature map
[0023] Step S4.2: Disguise feature map Linear combination to obtain the manual disguise feature map F mc , the calculation formula is as follows:
[0024]
[0025] Among them, α and β are hyperparameters and are set to 0.2 and 0.4, and the feature map F is manually disguised. mc Combine into feature sets;
[0026] Step S5, construct a feature vector decoder (including a convolution layer, a normalization layer and a Relu activation function), the feature vector decoder includes a decoding residual module and an activation function; the obtained uncertain feature map Dt Perform linear combination and input the result of linear combination into the feature vector decoder to obtain the depth map I d ;
[0027] Step S6: Calculate the mixed loss Loss, Loss = γLoss gc +δLoss eds +Loss trl ;
[0028] Among them, Loss gc is the geometric loss between the true depth value and the predicted depth value, Loss eds For the depth estimation diagram I d The introduced edge-aware smoothing loss, Loss trl For the depth estimation diagram I d The introduced thermal image reconstruction loss, γ and δ are hyper parameters;
[0029] Step S7: Repeat steps S3 to S6 until the algorithm converges, and a depth map estimation model can be obtained.
[0030] Furthermore, the specific steps of constructing a feature extraction encoder and calculating a feature vector set in step S3 are as follows:
[0031] Step S3.1: Enhance the thermal image Input 1×1 convolutional network to extract thermal image features and generate preliminary feature maps
[0032] Step S3.2: Use three 2×2 convolutional layers, a fully connected layer, and an ELU activation function to form a feature extraction encoder; convert the preliminary feature map obtained in step S3.1 into Input to the feature extraction encoder to finally generate the starting feature map
[0033] Step S3.3: The starting feature map obtained in step S3.2 Repeat steps S3.1 to S3.2 to obtain a feature map of the second size.
[0034] Step S3.4: The second size feature map obtained in step S3.3 Repeat step S3.2 to obtain the feature map of the third size
[0035] Step S3.5: The third dimension feature map obtained in step S3.4 Repeat step S3.2 to get the feature map of the fourth size
[0036] Step S3.6: The fourth dimension feature map obtained in step S3.5 Repeat step S3.2 to obtain the feature map of the fifth size
[0037] Step S3.7: feature map and Combined into feature sets
[0038] Furthermore, in step S5, the depth map I is obtained by the feature vector decoder. d The specific steps are as follows:
[0039] Step S5.1, use a 3×3 convolution layer, a normalization layer and a ReLU activation function to form a feature vector decoder; manually disguise the feature map F mc Input to the feature vector decoder to finally generate an estimated depth map
[0040] Step S5.2: Estimated depth map Repeat step S5.1 to get the second size The estimated depth map The estimated depth map Repeat step S5.1 to get the third size Estimated Depth Map The estimated depth map Repeat step S5.1 to get the fourth size Estimated Depth Map The estimated depth map Repeat step S5.1 to obtain the fifth size H×W estimated depth map
[0041] Step S5.3: Estimating the depth map of the fifth size H×W Depth estimation is calculated using the Sigmoid activation function. d .
[0042] Furthermore, the hybrid loss Loss in step S6 includes three parts: geometric loss of the depth map, edge-aware smoothing loss and thermal image reconstruction loss. The specific calculation process of the three parts is as follows:
[0043] Step S6.1, the calculation method of the geometric loss between the true depth value and the predicted depth value is as follows:
[0044]
[0045] Among them, D′ S is the deformation source depth map D s and relative posture Pt→s The synthesized depth map, D′ t is the target image depth map D t and D′ S The depth map obtained by sharing the same information, V is I s Projection to I t A collection of points.
[0046] Step S6.2: The low-frequency texture structure cannot provide substantial self-supervision. For the generated depth map, the edge-aware smoothing loss is calculated:
[0047]
[0048] in, It is differentiated along the spatial direction to prevent the estimated depth from stretching.
[0049] Step S6.3, calculate thermal image I t and I s The estimated depth map D t and D s The thermal image reconstruction loss and filtering of invalid pixels in the depth map are calculated as follows:
[0050]
[0051] Among them, I′ ent isI ent It is obtained by inverse warping, δ is the scale factor, SSIM is the structural similarity index map, Q gm Invalid pixels caused by occlusion and movement of objects in thermal images are filtered out. sm Filters out invalid pixels caused by camera motion during shooting.
[0052] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0053] (1) The present invention rearranges the thermal radiation value histogram of the original thermal image, thereby enhancing the contrast of the original thermal image and highlighting the local area details hidden in the original thermal image, making the self-supervision signal of the original thermal image stronger.
[0054] (2) The present invention uses a manual feature camouflage structure to mine the long-distance information that is ignored when the model estimates the depth map of the original thermal image. This introduces edge-aware smoothing loss and thermal image reconstruction loss, which strengthens the model's learning ability to estimate the depth of long-distance areas and improves the model's depth estimation ability in outdoor scenes at night.
[0055] (3) The present invention can not only solve the problem of insufficient long-distance information perception caused by the existing self-supervised monocular image depth estimation method when processing large-scale image data in harsh outdoor environments, but also improve the efficiency of depth estimation, laying an important foundation for the application of large-scale image data in harsh outdoor environments in the field of depth estimation and the development of three-dimensional reconstruction and visual navigation technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0057] Figure 2 is a sample schematic diagram of thermal image data in an embodiment;
[0058] Figure 3 is a sample schematic diagram of an enhanced thermal image in an embodiment;
[0059] Figure 4 Schematic diagram of a depth map output in an embodiment. DETAILED DESCRIPTION
[0060] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.
[0061] like Figure 1 As shown, a thermal image depth estimation method based on manual feature camouflage of the present invention first inputs the original thermal image and the internal parameters of the camera; then, the enhanced thermal image is calculated; secondly, a feature extraction encoder is constructed, and the thermal image feature map and its set are calculated; again, a feature camouflage structure is constructed, and the uncertain feature map is calculated; then, a feature vector decoder is constructed, and the estimated depth map is calculated; and the mixed loss is calculated with the ground truth value; finally, the mixed loss is calculated according to the estimated depth map and the ground truth value to obtain an accurate depth map.
[0062] The thermal image depth estimation method based on manual feature camouflage in this embodiment specifically includes the following steps:
[0063] Step S1: Input the thermal image and the internal parameters of the camera. There are N thermal images with a width of W and a height of H in each image. Each thermal image has internal and external parameters of the corresponding camera. Where i represents the serial number of the thermal image, and the internal parameters of the camera of the i-th thermal image are K i , the external parameters of the camera are T i ;
[0064] Step S2, calculating and enhancing the thermal image;
[0065] First, for each source thermal image The histogram h composed of its thermal radiation values is extractedθ ;
[0066] h θ =s θ ,θ=t0,t1,…,t max ;
[0067] [t θ ,t θ+1 ] is the temperature interval of the histogram of the source thermal image, s θ Yes θ ,t θ+1 ]The number of samples in the temperature interval, t max It refers to the maximum temperature of the source thermal image histogram;
[0068] Then all histograms h i According to the thermal radiation value v ra Rearrange to get the thermal image after the thermal radiation values are rearranged The formula for reordering is as follows:
[0069]
[0070] Among them, γ θ is the scaling factor, t′ θ The new offset for scaling each subhistogram, It refers to the proportional factor proportional to the number of sub-histograms. If no samples are observed in a certain temperature interval, the invalid area is directly removed by the re-ordering formula. After re-processing by calculation, different temperature intervals are extended or shortened according to the number of samples observed to improve the image contrast. In order to maintain the temporal consistency between adjacent images, the thermal images of each scene are calculated in groups of at least 4.
[0071] Finally, the thermal image after rearranging the thermal radiation values Input the thermal image processing structure, first divide the image into non-overlapping local images internally, then use the predefined limit threshold to perform histogram equalization processing on each local image to obtain a thermal image with enhanced details The resolution is still H×W;
[0072] Step S3, constructing a feature extraction encoder and calculating a feature vector set;
[0073] First based on thermal image Construct the first stage feature map For each feature map Five feature maps of different scales are extracted And the scales of the feature maps are: H×W, and Finally, the five feature maps are combined into a feature set
[0074] Step S4: construct a feature camouflage structure to hide and filter the information in the feature set; first, Extracting hidden feature maps Then, Input into the feature camouflage structure to calculate the uncertain feature map D t ;
[0075] Step S4.1. Use |sinx|, |cosx|, |2sin x cos x|, and sin 2 x constitutes a manual feature disguiser, and the fifth-size feature map of step S3 is Input into the manual feature disguiser to obtain the disguised feature map
[0076] Step S4.2: Disguise feature map Linear combination to obtain the manual disguise feature map F mc , the calculation formula is as follows:
[0077]
[0078] Among them, α and β are hyperparameters and are set to 0.2 and 0.4, and the feature map F is manually disguised. mc Combine into feature sets;
[0079] Step S5, construct a feature vector decoder (including a convolution layer, a normalization layer and a Relu activation function), the feature vector decoder includes a decoding residual module and an activation function; the obtained uncertain feature map D t Perform linear combination and input the result of linear combination into the feature vector decoder to obtain the depth map I d ;
[0080] Step S6: Calculate the mixed loss Loss, Loss = γLoss gc +δLoss eds +Loss trl ;
[0081] Among them, Loss gc is the geometric loss between the true depth value and the predicted depth value, Loss eds For the depth estimation diagram I d The introduced edge-aware smoothing loss, Loss trl For the depth estimation diagram I d The introduced thermal image reconstruction loss, γ and δ are hyper parameters;
[0082] Step S7: Repeat steps S3 to S6 until the algorithm converges, and a depth map estimation model can be obtained.
[0083] The specific steps of constructing a feature extraction encoder and calculating a feature vector set in step S3 of this embodiment are as follows:
[0084] Step S3.1: Enhance the thermal image Input 1×1 convolutional network to extract thermal image features and generate preliminary feature maps
[0085] Step S3.2: Use three 2×2 convolutional layers, a fully connected layer, and an ELU activation function to form a feature extraction encoder; convert the preliminary feature map obtained in step S3.1 into Input to the feature extraction encoder to finally generate the starting feature map
[0086] Step S3.3: The starting feature map obtained in step S3.2 Repeat steps S3.1 to S3.2 to obtain a feature map of the second size.
[0087] Step S3.4: The second size feature map obtained in step S3.3 Repeat step S3.2 to obtain the feature map of the third size
[0088] Step S3.5: The third dimension feature map obtained in step S3.4 Repeat step S3.2 to get the feature map of the fourth size
[0089] Step S3.6: The fourth dimension feature map obtained in step S3.5 Repeat step S3.2 to obtain the feature map of the fifth size
[0090] Step S3.7: feature map and Combined into feature sets
[0091] In step S5 of this embodiment, the depth map I is obtained by the feature vector decoder. d The specific steps are as follows:
[0092] Step S5.1, use a 3×3 convolution layer, a normalization layer and a ReLU activation function to form a feature vector decoder; manually disguise the feature map F mc Input to the feature vector decoder to finally generate an estimated depth map
[0093] Step S5.2: Estimated depth map Repeat step S5.1 to get the second size The estimated depth map The estimated depth map Repeat step S5.1 to get the third size Estimated Depth Map The estimated depth map Repeat step S5.1 to get the fourth size Estimated Depth Map The estimated depth map Repeat step S5.1 to obtain the fifth size H×W estimated depth map
[0094] Step S5.3: Estimating the depth map of the fifth size H×W Depth estimation is calculated using the Sigmoid activation function. d .
[0095] Furthermore, the hybrid loss Loss in step S6 includes three parts: geometric loss of the depth map, edge-aware smoothing loss and thermal image reconstruction loss. The specific calculation process of the three parts is as follows:
[0096] Step S6.1, the calculation method of the geometric loss between the true depth value and the predicted depth value is as follows:
[0097]
[0098] Among them, D′ S is the deformation source depth map D s and relative posture P t→s The synthesized depth map, D′ t is the target image depth map D t and D′ S The depth map obtained by sharing the same information, V is I s Projection to I t A collection of points.
[0099] Step S6.2: The low-frequency texture structure cannot provide substantial self-supervision. For the generated depth map, the edge-aware smoothing loss is calculated:
[0100]
[0101] in, It is differentiated along the spatial direction to prevent the estimated depth from stretching.
[0102] Step S6.3, calculate thermal image I t and I sThe estimated depth map D t and D s The thermal image reconstruction loss and filtering of invalid pixels in the depth map are calculated as follows:
[0103]
[0104] Among them, I′ ent isI ent It is obtained by inverse warping, δ is the scale factor, SSIM is the structural similarity index map, Q hm Invalid pixels caused by occlusion and movement of objects in thermal images are filtered out. sm Filters out invalid pixels caused by camera motion during shooting.
[0105] Example 1
[0106] In this embodiment, the original thermal image input is as follows: Figure 2 As shown, Figure 2 The two thermal images in the first row belong to outdoor scenes, and the two thermal images are taken at different locations. The two thermal images in the second row belong to indoor scenes, and the two thermal images are taken at different lighting conditions.
[0107] The enhanced thermal image output after calculation is as follows Figure 3 As shown, for Figure 2 The final output depth map of the two scenes is as follows: Figure 4 As shown, it can be seen that the contrast and details of the original thermal image are greatly enhanced after the thermal radiation value histogram is rearranged by the present invention, and the estimated depth map has a high degree of match with the real scene in local area details and long-distance information.
[0108] It can be seen from the above embodiments that the present invention uses a thermal image enhancement method and embeds a manual feature camouflage structure in the thermal image depth estimation method based on the Transformer structure, thereby enhancing the ability of the Transformer structure to capture long-distance information and obtaining a multi-scale feature map set. Thus, a more accurate depth map is calculated. According to the final experimental results ( Figure 4 ) It can be seen that the present invention can not only enhance the accuracy of the model in estimating the depth map when the lighting conditions are poor, but also improve the time efficiency of the depth estimation, making it possible to quickly estimate the accurate depth maps of different scenes when the lighting environment is poor.
[0109] Example 2
[0110] The technology of the present invention is compared with the evaluation results of existing thermal image depth estimation methods on the VIVID++ dataset in recent years. The experimental results are shown in Table 1. Table 1 is the quantitative results of different image depth estimation methods on the VIVID++ dataset. The evaluation indicators are the mean Error value (the lower the better) and the mean Accuracy value (the higher the better).
[0111] Table 1 Comparison of evaluation results between the present invention and the prior art
[0112]
[0113]
[0114] The prior art in Table 1 calculates feature maps from images of different scenes, and then uses the feature maps and a mixed loss function to calculate a high-quality estimated depth map. Although the prior art can generate a relatively accurate depth map, it increases the computational burden and makes it difficult to handle scenes under harsh lighting conditions. The present invention can better handle low-light and long-distance areas and obtain a more accurate depth map.
Claims
1. A thermal image depth estimation method based on manual feature camouflage, characterized in that: The following steps are involved: Step S1: Input the thermal image and the internal parameters of the camera. There are N thermal images with a width of W and a height of H in each image. Each thermal image has internal and external parameters of the corresponding camera. Where i represents the serial number of the thermal image, and the internal parameters of the camera of the i-th thermal image are K i , the external parameters of the camera are T i ; Step S2, calculating and enhancing the thermal image; First, for each source thermal image The histogram h composed of its thermal radiation values is extracted θ ; h θ =s θ ,θ=t0,t1,…,t max ; [t θ ,t θ+1 ] is the temperature interval of the histogram of the source thermal image, s θ Yes θ ,t θ+1 ] The number of samples in the temperature interval; t max It refers to the maximum temperature of the source thermal image histogram; Then all histograms h θ According to the thermal radiation value v ra Rearrange to get the thermal image after the thermal radiation values are rearranged The formula for reordering is as follows: Among them, γ θ is the scaling factor, t′ θ The new offset for scaling each subhistogram, It refers to the scaling factor proportional to the number of subhistograms; Finally, the thermal image after rearranging the thermal radiation values Input the thermal image processing structure, first divide the image into non-overlapping local images internally, then use the predefined limit threshold to perform histogram equalization processing on each local image to obtain a thermal image with enhanced details Step S3, constructing a feature extraction encoder and calculating a feature vector set; Based on thermal image Construct the feature map of the first stage For each feature map Five feature maps of different scales are extracted And the scales of the feature maps are: H×W, and Finally, the five feature maps are combined into a feature set Step S4: construct a feature camouflage structure to hide and filter the information in the feature set; first, Extracting hidden feature maps Then, Input into the feature camouflage structure to calculate the uncertain feature map D t , the specific method is: Step S4.
1. Use |sin x|, |cos x|, |2sin x cos x|, and sin 2 x constitutes a manual feature disguiser, and the fifth-size feature map of step S3 is Input into the manual feature disguiser to obtain the disguised feature map Step S4.2: Disguise feature map Linear combination to obtain the manual disguise feature map F mc , the calculation formula is as follows: Among them, α and β are hyperparameters and are set to 0.2 and 0.4, and the feature map F is manually disguised. mc Combine into feature sets; Step S5: construct a feature vector decoder, which includes a decoding residual module and an activation function; convert the obtained uncertain feature map D t Perform linear combination and input the result of linear combination into the feature vector decoder to obtain the depth map I d ; Step S6: Calculate the mixed loss Loss, Loss = γLoss gc +δLoss eds +Loss trl ; Among them, Loss gc is the geometric loss between the true depth value and the predicted depth value, Loss eds For the depth estimation diagram I d The introduced edge-aware smoothing loss, Loss trl For the depth estimation diagram I d The introduced thermal image reconstruction loss, γ and δ are hyper parameters; Step S7: Repeat steps S3 to S6 until the algorithm converges, and a depth map estimation model can be obtained.
2. The thermal image depth estimation method based on manual feature camouflage according to claim 1 is characterized in that: The specific steps of constructing a feature extraction encoder and calculating a feature vector set in step S3 are as follows: Step S3.1: Enhance the thermal image Input 1×1 convolutional network to extract thermal image features and generate preliminary feature maps Step S3.2, use three 2×2 convolutional layers, a fully connected layer and an ELU activation function to form a feature extraction encoder; The preliminary feature map obtained in step S3.1 Input to the feature extraction encoder to finally generate the starting feature map Step S3.3: The starting feature map obtained in step S3.2 Repeat steps S3.1 to S3.2 to obtain a feature map of the second size. Step S3.4: The second size feature map obtained in step S3.3 Repeat step S3.2 to obtain the feature map of the third size Step S3.5: The third dimension feature map obtained in step S3.4 Repeat step S3.2 to get the feature map of the fourth size Step S3.6: The fourth dimension feature map obtained in step S3.5 Repeat step S3.2 to obtain the feature map of the fifth size Step S3.7: feature map and Combined into feature sets 3. The thermal image depth estimation method based on manual feature camouflage according to claim 1 is characterized in that: In step S5, the depth map I is obtained by the feature vector decoder d The specific steps are as follows: Step S5.1, use a 3×3 convolution layer, a normalization layer and a ReLU activation function to form a feature vector decoder; manually disguise the feature map F mc Input to the feature vector decoder to finally generate an estimated depth map Step S5.2: Estimated depth map Repeat step S5.1 to get the second size The estimated depth map The estimated depth map Repeat step S5.1 to get the third size Estimated Depth Map The estimated depth map Repeat step S5.1 to get the fourth size Estimated Depth Map The estimated depth map Repeat step S5.1 to obtain the fifth size H×W estimated depth map Step S5.3: Estimating the depth map of the fifth size H×W Depth estimation is calculated using the Sigmoid activation function. d .
4. The thermal image depth estimation method based on manual feature camouflage according to claim 1 is characterized in that: The mixed loss Loss in step S6 includes the geometric loss of the depth map, the edge-aware smoothing loss and the thermal image reconstruction loss. The specific calculation process of the three parts is as follows: The geometric loss between the true depth value and the predicted depth value is calculated as follows: Among them, D S ′ is the deformation source depth map D s and relative posture P t→s The synthesized depth map, D t ′ is the target image depth map D t and D S ′ The depth map obtained by sharing the same information, V is I s Projection to I t The set of points of The low-frequency texture structure cannot provide substantial self-supervision, so for the generated depth map, the edge-aware smoothing loss is calculated: in, It is differentiated along the spatial direction to prevent the expansion and contraction of the estimated depth; Calculate thermal image I t and I s The estimated depth map D t and D s The thermal image reconstruction loss and filtering of invalid pixels in the depth map are calculated as follows: Among them, I e ′ nt isI ent It is obtained by inverse warping, δ is the scale factor, SSIM is the structural similarity index map, Q gm Invalid pixels caused by occlusion and movement of objects in thermal images are filtered out. sm Filters out invalid pixels caused by camera motion during shooting.
Citation Information
Patent Citations
Monocular depth estimation system and method for enhancing feature fusion in three-dimensional scene reconstruction
CN115294282A
Systems and methods for patient structure estimation during medical imaging
US20210201476A1