Depth Map Prediction Method and Device Based on Phase-Guided Transformer-CNN Dual-Path Fusion

By introducing a phase-guided Fourier-Transformer-CNN network in deep learning, the problem that the prior art is difficult to capture complex surface features in phase expansion and depth map prediction is solved, and the depth map prediction with high accuracy and robustness is achieved.

CN119445310BActive Publication Date: 2025-05-27SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510020085.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-27
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The prior art has limitations in phase expansion and depth map prediction, and it is difficult to effectively capture global and local features on complex surfaces, resulting in phase expansion errors and low depth map quality.

Method used

The phase-guided Transformer-CNN dual-path fusion method is adopted to extract phase frequency features through the Fourier model, and the global and local features are extracted respectively by combining Transformer and CNN models, and cross-fusion and multi-scale feature enhancement are performed to generate a depth map.

Benefits of technology

Effectively predict phase information on complex surfaces, significantly improve the accuracy and robustness of depth map prediction, and better deal with phase jumps and error propagation problems on complex surfaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445310B_ABST
    Figure CN119445310B_ABST
Patent Text Reader

Abstract

The present invention discloses a depth map prediction method and device based on phase-guided Transformer-CNN dual-path fusion, including: obtaining a fringe image to be processed; extracting frequency features regarding phase information from the fringe image through a Fourier model, and obtaining long-distance and global frequency features by using a Transformer model; extracting detailed features from the fringe image by using a CNN model; performing cross-fusion on the long-distance and global frequency features and the detailed features, and performing multi-scale feature enhancement to obtain the depth map of the fringe image. PG-FTCNet can effectively predict phase information on complex surfaces while significantly improving the accuracy of depth map prediction. This method takes a single fringe image in fringe projection profilometry as the input, simultaneously realizes phase unwrapping and depth map prediction, while improving the interpretability of the results, retains an efficient end-to-end calculation mode, and provides a reliable and practical solution for the 3D reconstruction task of FPP.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and in particular, to a depth map prediction method and device based on phase-guided Transformer-CNN dual-path fusion. Background Art

[0002] Fringe Projection Profilometry (FPP) is a non-contact three-dimensional reconstruction technology widely used in the field of computer vision. Its basic principle is to project a fringe pattern onto the surface of the object to be measured, calculate the wrapped phase through the deformation of the fringes, then perform phase unwrapping, and finally map the unwrapped phase to a depth value to generate the depth map required for three-dimensional reconstruction. In this process, the accuracy of phase unwrapping plays a crucial role in the quality of the final depth map. Therefore, phase unwrapping is one of the key steps in FPP.

[0003] In recent years, due to its excellent performance in automatically learning complex information, deep learning has been widely applied in many fields and achieved remarkable progress. In FPP-related research, many scholars have begun to explore how to introduce deep learning techniques into phase unwrapping and depth map prediction tasks. For example, Convolutional Neural Network (CNN) has been proven to perform well in phase unwrapping, capable of effectively extracting local phase features and handling medium-level noise. However, the receptive field of CNN is limited, and it has deficiencies in capturing global features. It may wrongly assume the continuity of local phases, making it difficult to handle discontinuities or large phase jumps on complex surfaces, thus leading to phase unwrapping errors.

[0004] In addition, Generative Adversarial Network (GAN) has also been applied to the phase unwrapping task. Through the adversarial learning between the generator and the discriminator, GAN performs outstandingly in synthesizing data and suppressing high-frequency noise. However, its results often tend to smooth out details, especially when the generator is overfitted, artifacts may appear, limiting the credibility of the results. On the other hand, due to its ability to capture long-range dependencies and rich context information, the Transformer-based model has recently become a hot topic in phase unwrapping research. However, the Transformer still faces challenges in balancing global and local features. It may ignore local details due to excessive focus on the global background, thus limiting its performance in solving phase ambiguity and noise interference problems.

[0005] Although deep learning technology has shown potential in related research on FPP, most of the current mainstream deep prediction neural networks adopt an end-to-end manner from the fringe pattern directly to the depth map, completely skipping the traditional phase unwrapping step. Although this method has improved computational efficiency, due to the lack of explicit modeling of the phase unwrapping process, the interpretability of its results is insufficient, and it cannot fully utilize the phase information to optimize the depth map prediction. This makes it difficult for existing methods to handle problems such as phase jumps and error propagation when dealing with complex surfaces, resulting in limited quality of the final depth map. Summary of the Invention

[0006] To solve the above technical problems, an embodiment of the present invention provides a depth map prediction method based on phase-guided Transformer-CNN dual-path fusion, including:

[0007] Obtain the fringe image to be processed;

[0008] Extract the frequency features regarding phase information from the fringe image through the Fourier model, and use the Transformer model to obtain long-distance and global frequency features;

[0009] Use the CNN model to extract detailed features from the fringe image;

[0010] Perform cross-fusion on the long-distance and global frequency features and the detailed features, and perform multi-scale feature enhancement to obtain the depth map of the fringe image.

[0011] Further, the extracting the frequency features regarding phase information from the fringe image through the Fourier model includes:

[0012] Convert the fringe image from the spatial domain image to the frequency domain image through Fourier transform;

[0013] Extract the phase information encoded in the ±1 level spectra from the frequency domain image and remove the zero-frequency features;

[0014] Process the image after removing the zero-frequency features through a feedforward neural network to obtain the frequency features containing phase information.

[0015] Further, the Transformer model includes: the MHSA model and the FFN model; the using the Transformer model to obtain long-distance frequency features includes:

[0016] Input the frequency features into the MHSA model and the FFN model in sequence to obtain the features with long-distance dependence and global frequency.

[0017] Further, the using the CNN model to extract detailed features from the fringe image includes:

[0018] Input the stripe image into a plurality of convolutional blocks, and use a convolutional layer with an activation function of LeakyReLU and batch normalization to extract the detailed features of the stripe image.

[0019] Further, the cross-fusion of the long-distance and global frequency features and the detailed features includes:

[0020] In the encoder, the Transformer branch and the CNN branch perform feature cross-fusion at each layer;

[0021] In the decoder, the feature image after feature cross-fusion is upsampled to restore the resolution of the original stripe image, and the feature image after feature cross-fusion in the encoder is merged with the upsampled features in the decoder using skip connections to generate the unwrapped phase.

[0022] Further, in the encoder, the Transformer branch and the CNN branch perform feature cross-fusion at each layer, including:

[0023] Input the stripe image containing the detailed features into a convolutional embedding model to convert it into the data shape features of the Transformer branch;

[0024] Fuse the data shape features converted into the Transformer branch with the long-distance and global frequency features, and after restoring them to a standard feature map through an image reconstruction model, optimize them using a CNN model.

[0025] Further, the multi-scale feature enhancement to obtain the depth map of the stripe image includes:

[0026] Use the unwrapped phase as a physical information reference, and fuse it with the upsampled features and the input image in a multi-scale global-local fusion model MSEF to obtain high-resolution depth features.

[0027] The present invention discloses a depth map prediction device based on phase-guided Transformer-CNN dual-path fusion, including:

[0028] An acquisition model for acquiring the stripe image to be processed;

[0029] A processing model for extracting the frequency features regarding phase information from the stripe image through a Fourier model, and using a Transformer model to obtain long-distance and global frequency features;

[0030] The processing model is also used to extract detailed features from the stripe image using a CNN model;

[0031] An execution model for cross - fusing the long - distance and global frequency features and the detail features, and performing multi - scale feature enhancement to obtain the depth map of the fringe image.

[0032] Further, the acquisition model includes:

[0033] A first acquisition sub - model for converting the fringe image from the spatial - domain image to the frequency - domain image through Fourier transform;

[0034] A first processing sub - model for extracting the phase information encoded in the ±1 - level spectra from the frequency - domain image and removing the zero - frequency feature;

[0035] A first execution sub - model for processing the image after removing the zero - frequency feature through a feed - forward neural network to obtain the frequency features containing phase information.

[0036] Further, the Transformer model includes: an MHSA model and an FFN model; the processing model includes:

[0037] A second processing sub - model for sequentially inputting the frequency features into the MHSA model and the FFN model to obtain the features with long - distance dependence and global frequency.

[0038] Further, the processing sub - model includes:

[0039] A third processing sub - model for inputting the fringe image into a plurality of convolutional blocks, and using convolutional layers with the activation function LeakyReLU and batch normalization to extract the detail features of the fringe image.

[0040] Further, the execution model includes:

[0041] A second execution sub - model for, in the encoder, the Transformer branch and the CNN branch performing feature cross - fusion at each layer through the multi - scale global - local fusion model MGLF;

[0042] A third execution sub - model for, in the decoder, the feature image after completing feature cross - fusion being restored to the resolution of the original fringe image through upsampling, and using skip connections to merge the feature image after feature cross - fusion in the encoder with the upsampled features in the decoder to generate the unwrapped phase.

[0043] Further, the second execution sub - model includes:

[0044] A fourth execution sub - model for inputting the fringe image containing the detail features into a convolutional embedding model to convert it into the data - shape features of the Transformer branch;

[0045] The fifth execution sub-model is used to fuse the data shape features converted into the Transformer branch with the long-distance and global frequency features, and after being restored to the standard feature map by the image reconstruction model, it is optimized by the CNN model.

[0046] Further, the execution model includes:

[0047] The sixth execution sub-model is used to use the unwrapped phase as a physical information reference, and fuse it with the upsampled features and the input image in the multi-scale global-local fusion model MSEF to obtain high-resolution depth features.

[0048] A computer device includes a memory and a processor. When the computer-readable instructions stored in the memory are executed by the processor, the processor executes the steps of the depth map prediction method based on phase-guided Transformer-CNN dual-path fusion as described above.

[0049] A storage medium storing computer-readable instructions, when the computer-readable instructions are executed by one or more processors, enables the one or more processors to execute the steps of the depth map prediction method based on phase-guided Transformer-CNN dual-path fusion as described above.

[0050] The present invention is based on a deep learning network architecture - the Phase-Guided Fourier-Transformer-CNN Network (PG-FTCNet). Different from existing end-to-end depth prediction schemes, PG-FTCNet explicitly introduces a phase unwrapping process. In a phase-guided manner, phase unwrapping and depth prediction are integrated into a unified network architecture. The network adopts a dual-branch architecture of Transformer and CNN to extract global and local features respectively, and a Fourier transform model is introduced into the Transformer branch to extract frequency features. The introduction of the Fourier transform enhances the network's ability to capture phase jumps and long-distance dependence features on complex surfaces, while the dual-branch architecture performs excellently in the balanced integration of global features and local details. PG-FTCNet can effectively predict the phase information on complex surfaces and significantly improve the accuracy of depth map prediction. This method takes a single fringe image in fringe projection profilometry as the input, and simultaneously realizes phase unwrapping and depth map prediction. While improving the interpretability of the results, it retains an efficient end-to-end calculation mode, providing a reliable and practical solution for the 3D reconstruction task of FPP. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0052] Figure 1 It is a schematic flowchart of a depth map prediction method based on phase-guided Transformer-CNN dual-path fusion provided by an embodiment of the present invention;

[0053] Figure 2 It is a schematic diagram of model processing of another depth map prediction method based on phase-guided Transformer-CNN dual-path fusion provided by an embodiment of the present invention. Detailed implementation manners

[0054] To enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention.

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0056] As Figure 1 shown, the present invention provides a depth map prediction method based on phase-guided Transformer-CNN dual-path fusion, including:

[0057] S1. Obtain a fringe image to be processed;

[0058] S2. Extract frequency features regarding phase information from the fringe image through a Fourier model, and use a Transformer model to obtain long-distance and global frequency features;

[0059] S3. Use a CNN model to extract detailed features from the fringe image;

[0060] S4. Cross-fuse the long-distance and global frequency features and the detailed features, and perform multi-scale feature enhancement to obtain the depth map of the fringe image.

[0061] In an embodiment of the present invention, in step S2, extracting frequency features regarding phase information from the fringe image through a Fourier model includes:

[0062] Step 1: Convert the fringe image from the spatial domain image to the frequency domain image through Fourier transform;

[0063] Step 2: Extract the phase information encoded in the ±1-level spectra from the frequency domain image and remove the zero-frequency feature;

[0064] Step 3: Process the image after removing the zero-frequency feature through a feedforward neural network to obtain the frequency feature containing phase information.

[0065] It should be noted that the Transformer model includes: the MHSA model and the FFN model; obtaining long-distance frequency features using the Transformer model includes: sequentially inputting the frequency features into the MHSA model and the FFN model to obtain features with long-distance dependence and global frequency.

[0066] As Figure 2 shown in (a) of , in the Fourier model, the fringe image is sequentially processed through fast Fourier transform (FFT), learnable filter, inverse fast Fourier transform (IFFT), layer normalization (Layer Norm), feedforward neural network (FFN), and regularization (Dropout). The fringe image is converted from the spatial domain to the frequency domain after FFT to obtain frequency features, extract the phase information encoded in the ±1-level spectra, remove the DC component and other irrelevant frequencies through the learnable filter and filter frequencies to suppress low-frequency noise, and obtain frequency features after IFFT.

[0067] It should be noted that the frequency features after screening and optimization, that is, provide support for the Transformer model to capture global frequency-related features through FFN. It should be noted that after Fourier transform, there are zero-level and ±1-level spectra, and the phase information is contained in the ±1-level spectra.

[0068] As Figure 2 shown in (b) of , (b-1) is a schematic diagram of the Transformer model. The Transformer model is composed of a multi-head self-attention mechanism (MHSA) model and a feedforward neural network (FFN) model, captures long-distance dependence and global frequency-related features, and provides support for the global feature extraction of the encoder.

[0069] Specifically, the frequency features are sequentially processed through layer normalization (Layer Norm), self-attention mechanism (MHSA), and regularization (Dropout), and, the frequency image is sequentially processed through layer normalization (Layer Norm), feedforward neural network (FFN), and regularization (Dropout).

[0070] The present invention utilizes the phase-frequency feature extraction method in FTP and combines the feature learning ability of deep learning to enhance the adaptability and robustness of frequency information extraction. At the same time, frequency components are screened through learnable parameters, taking into account the retention of high-frequency details and the suppression of low-frequency noise to achieve more accurate frequency-domain modeling. In addition, the extracted frequency features are input into the Transformer model to further capture long-distance frequency-related features, enabling global modeling of phase information in depth prediction.

[0071] Through the optimized extraction of phase features in the frequency domain, the Fourier model retains the physical interpretability of phase information and makes up for the deficiency of traditional deep learning methods in directly predicting from the fringe pattern to the depth map end-to-end. Taking frequency information as an explicit feature and inputting it into the network reduces noise and background interference and improves the accuracy and robustness of depth prediction. Experiments show that the model can effectively alleviate the problems of phase jumps and error propagation in complex surfaces (such as severely undulating or surface discontinuous regions).

[0072] In one embodiment of the present invention, in step S3, a CNN model is used to extract detailed features from the fringe image, including: inputting the fringe image into multiple convolutional blocks, and using a convolutional layer with an activation function carrying LeakyReLU and batch normalization to extract the detailed features of the fringe image.

[0073] As Figure 2 shown in (b) of , (b-2) is a schematic diagram of the convolutional (CNN) model. In the CNN model, after the fringe image undergoes two-dimensional convolution Conv2d with a convolution kernel k = 3, a stride s = 1, and a padding value p = 1 for convolution processing, it undergoes batch normalization (Btach Norm) and extracts local features through an activation function with LeakyReLU. After multiple extractions, the detailed information on the object surface can be effectively captured.

[0074] In one embodiment of the present invention, in step S4, cross-fusion of long-distance and global frequency features and detailed features is performed, including:

[0075] Step 1: In the encoder, the Transformer branch and the CNN branch perform feature cross-fusion through the multi-scale global-local fusion model MGLF at each layer;

[0076] It should be noted that the Multi-scale Global-Local Fusion (MGLF) model is used to fuse local and global features at multiple scales. This module converts local features into a data shape consistent with global features through a Convolutional Patch Embedding (CPE), and then performs global-local feature fusion. The fused features are input into the Transformer model for further calculation. The obtained global features are transmitted forward through the global branch. At the same time, the global features are reshaped into the data shape of local features through a Patch Reconstruction (PR) model, fused with the initial local features, and then input into the CNN model for further calculation. The obtained local features are transmitted forward through the local branch. This fusion method can effectively integrate local and global information and improve the richness of feature representation.

[0077] The specific process is as follows:

[0078] Input the fringe image containing the detailed features into the convolutional embedding model to convert it into the data shape features of the Transformer branch;

[0079] Fuse the data shape features converted into the Transformer branch with the long-distance and global frequency features, and after being restored to the standard feature map through the image reconstruction model, optimize it with the CNN model.

[0080] Step 2: In the decoder, the feature image after feature cross-fusion is upsampled to restore the resolution of the original fringe image, and the feature image after feature cross-fusion in the encoder is merged with the upsampled features in the decoder using skip connections to generate the unwrapped phase.

[0081] As Figure 2 shown in (c) of , input the fringe image, pass through the Fourier model and Transformer model in the Transformer branch, and the CNN model in the CNN branch, perform feature fusion in the MGLF model, and after fusion, perform upsampling through the Multi-Scale Feature Enhancement Fusion (MSEF) model, and extract the unwrapped phase through 1x1 convolution to obtain the depth map.

[0082] It should be noted that in the MGLF model, in the encoder, the Transformer branch and the CNN branch perform feature cross-fusion through the MGLF model at each layer. As Figure 2As shown in (b-3) therein, the local features extracted by the CNN (where W is the width, H is the height, and C is the number of channels) are convolved by the CPE model through a two-dimensional convolution Conv2d with a convolution kernel k = 4, a stride s = 2, and a padding value p = 1, and then transformed into the data shape of the Transformer branch through reshaping and permutation, that is, it becomes width W / 2, height H / 2, and number of channels C / 2. After being fused with the global features processed by the Transformer model, as Figure 2 shown in (b-4) therein, it is restored to the standard feature map by the PR model, that is, through reshaping and permutation, interpolation, and further optimized by a two-dimensional convolution Conv2d with a convolution kernel k = 1. Through this process, the global frequency features and local spatial features are effectively integrated, thus enhancing the overall feature extraction ability of the network. It should be noted that, as Figure 2 shown in (c) therein, in each layer of the MGLF model, the image processed by the Transformer model is processed by the PR model and then feature-fused by the CNN model. At the same time, the image processed by the CNN model is processed by the CPE model and then feature-fused by the Transformer model.

[0083] In addition, in Figure 2 (c) of, in the MSEF model, in the decoder stage, the feature map is upsampled layer by layer to restore the resolution of the original image, and at the same time, the multi-scale fusion features of the encoder are merged with the upsampled features of the decoder layer by layer using skip connections. Finally, the unfolded phase generates a depth map through a 1×1 convolution and participates in the final depth prediction as a guidance prior.

[0084] In one embodiment of the present invention, in step S4, obtaining the depth map of the fringe image through multi-scale feature enhancement includes: using the unfolded phase as a physical information reference and fusing it with the upsampled features and the input image in the MSEF model to obtain high-resolution depth features.

[0085] It should be noted that the MSEF model is used to fuse features of different scales and generate the final depth map. This module first adjusts the features of different scales to the same number of channels, and then performs upsampling and addition operations layer by layer, so that the feature maps of all scales are fused at the highest resolution. Finally, the fused feature map is converted into a depth map through convolution. The MSEF model can effectively integrate features of different scales and improve the accuracy of depth prediction.

[0086] In the embodiments of the present invention, the unwrapped phase is used as physical guidance and is fused with the upsampled features and the input image in the MSEF model to hierarchically integrate high-resolution and depth features. In the last step, the upsampled feature map is restored to the input resolution, and through a 1×1 convolution operation, a high-precision depth map is finally generated. This network architecture ensures the accuracy and robustness of the depth map by effectively capturing spatial and frequency-related information and incorporating the unwrapped phase as physical prior knowledge to guide depth prediction. Especially in complex scenes, it can provide reliable 3D reconstruction results.

[0087] In the decoder of the present invention, a CNN model is used to generate the unwrapped phase, which combines the fringe image and the feature map extracted by the network and participates in the multi-scale feature fusion process for generating the final depth map. The multi-scale feature enhanced fusion MSEF model uses the unwrapped phase as a guiding prior and fuses it with the feature maps upsampled layer by layer and the high-resolution feature maps in the decoder to improve the accuracy and interpretability of depth prediction. This method retains the physical modeling characteristics of the fringe projection profilometry (FPP), making the depth prediction process more in line with physical laws. By guiding and optimizing the training process of the deep learning model with the phase, the problems of phase jumps and error propagation are effectively alleviated, and the robustness in complex surface scenes is improved. The physical prior of phase unwrapping is introduced into the deep learning model to explicitly model and generate the unwrapped phase, which is used as important guiding information in the depth prediction process. The performance comparison is shown in Table 1 below:

[0088] Table 1. Performance Comparison of the PG-FTCNet Model

[0089]

[0090] In the present invention, PG-FTCNet combines frequency-domain processing with the Fourier model, combines global frequency-related feature extraction with the Transformer model, and combines local spatial detail capture with the CNN model. These components are unified by the MGLF model and the MSEF model, jointly enhancing the model's ability to accurately reconstruct 3D surfaces. At the same time, the root mean square error RMSE of the unwrapped phase prediction is 0.0212 mm, the structural similarity SSIM is 0.9924, the RMSE of the depth map prediction is 0.0220 mm, and the SSIM is 0.9925. Experiments show that the lack of the unwrapped phase as a physical prior guidance (no unwrapped phase in Table 1) will lead to a significant decrease in accuracy (RMSE 0.0398 mm, SSIM 0.9728), highlighting its key role in depth prediction. In addition, experiments show that excluding key models such as the Fourier model or the MGLF model will lead to a significant decrease in performance, emphasizing the importance of these models in capturing meaningful frequency-related features and achieving cross-fusion of global and local information. This cross-fusion is crucial for accurately integrating various features, thus maximizing the prediction accuracy of the model.

[0091] These results indicate that PG-FTCNet not only outperforms traditional methods but also, by leveraging the synergistic effects of its dedicated components, makes it a robust and efficient solution for high-precision 3D reconstruction in practical applications. In the field of architecture, it can be used for precise building modeling, providing support for design, planning, and project monitoring. In the medical field, this technology can obtain the 3D structure of human organs and is widely used in surgical planning, orthodontics, facial reconstruction, and dental treatment. In the field of cultural heritage protection and digitization, it can be used for the digitization and 3D reconstruction of cultural relics, ancient buildings, and artworks for more effective protection, restoration, and research. In industrial manufacturing, this technology is applied to product quality control, shape detection, and component inspection, helping to accurately measure part dimensions, surface defects, and assembly quality. In game development and virtual reality technology, it is used to build realistic virtual scenes and characters, significantly enhancing the game interaction experience and the visual effects of virtual reality applications. With the progress of technology, some smartphones and cameras have integrated 3D reconstruction functions to support consumer-level applications such as augmented reality, face unlocking, and gesture recognition, further promoting the popularization and application of this technology in daily life.

[0092] The embodiment of the present invention also provides a depth map prediction device based on phase-guided Transformer-CNN dual-path fusion, including: The present invention discloses a depth map prediction device based on phase-guided Transformer-CNN dual-path fusion, including: an acquisition model for acquiring a fringe image to be processed; a processing model for extracting frequency features regarding phase information from the fringe image through a Fourier model and obtaining long-distance and global frequency features by using a Transformer model; the processing model is also used for extracting detailed features from the fringe image by using a CNN model; an execution model for cross-fusing the long-distance and global frequency features and the detailed features and performing multi-scale feature enhancement to obtain a depth map of the fringe image.

[0093] The network architecture based on deep learning of the present invention - Phase-Guided Fourier-Transformer-CNN Network (PG-FTCNet) is different from existing end-to-end depth prediction schemes. PG-FTCNet explicitly introduces a phase unwrapping process and integrates phase unwrapping and depth prediction into a unified network architecture in a phase-guided manner. The network adopts a dual-branch architecture of Transformer and CNN to extract global and local features respectively, and a Fourier transform model is introduced into the Transformer branch to extract frequency features. The introduction of the Fourier transform enhances the network's ability to capture phase jumps and long-distance dependence features on complex surfaces, while the dual-branch architecture performs excellently in the balanced integration of global features and local details. PG-FTCNet can effectively predict phase information on complex surfaces and significantly improve the accuracy of depth map prediction. This method takes a single fringe image in fringe projection profilometry as the input, simultaneously realizes phase unwrapping and depth map prediction, improves the interpretability of the results while retaining an efficient end-to-end computing mode, and provides a reliable and practical solution for the 3D reconstruction task of FPP.

[0094] In some embodiments, the acquisition model includes: a first acquisition sub-model for converting the fringe image from the spatial domain image to the frequency domain image through Fourier transform; a first processing sub-model for extracting the phase information encoded in the ±1 level spectra from the frequency domain image and removing the zero-frequency features; a first execution sub-model for processing the image after removing the zero-frequency features through a feedforward neural network to obtain frequency features containing phase information.

[0095] In some embodiments, the Transformer model includes: an MHSA model and an FFN model; the processing model includes: a second processing sub-model, configured to sequentially input the frequency features into the MHSA model and the FFN model to obtain features with long-range dependencies and global frequencies.

[0096] In some embodiments, the processing sub-model includes: a third processing sub-model, configured to input the fringe image into a plurality of convolutional blocks, and use convolutional layers with an activation function carrying LeakyReLU and batch normalization to extract the detailed features of the fringe image.

[0097] In some embodiments, the execution model includes: a second execution sub-model, configured to perform feature cross-fusion between the Transformer branch and the CNN branch in each layer through a multi-scale global-local fusion model MGLF in the encoder; a third execution sub-model, configured to, in the decoder, restore the feature image after feature cross-fusion to the resolution of the original fringe image through upsampling, and use skip connections to merge the feature image after feature cross-fusion in the encoder with the upsampled features in the decoder to generate the unwrapped phase.

[0098] In some embodiments, the second execution sub-model includes: a fourth execution sub-model, configured to input the fringe image containing the detailed features into a convolutional embedding model to convert it into the data shape features of the Transformer branch; a fifth execution sub-model, configured to fuse the data shape features converted into the Transformer branch with the long-range and global frequency features, and after restoring them to standard feature maps through an image reconstruction model, optimize them with a CNN model.

[0099] In some embodiments, the execution model includes: a sixth execution sub-model, configured to use the unwrapped phase as a physical information reference to fuse it with the upsampled features and the input image in a multi-scale global-local fusion model MSEF to obtain high-resolution depth features.

[0100] To solve the above technical problems, an embodiment of the present invention further provides a computer device. The computer device includes a processor, a non-volatile storage medium, a memory, and a network interface connected through a system bus. Among them, the non-volatile storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The control information sequence can be stored in the database. When the computer-readable instructions are executed by the processor, the processor can implement a depth map prediction method based on phase-guided Transformer-CNN dual-path fusion. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. Computer-readable instructions can be stored in the memory of the computer device. When the computer-readable instructions are executed by the processor, the processor can execute a method for predicting protein-ligand binding affinity based on distance and atomic contact number. The network interface of the computer device is used to connect and communicate with a terminal. In this embodiment, the processor is used to execute the specific content of obtaining and processing the model. The memory stores the program code and various types of data required to execute the above model. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program code and data required to execute all sub-models in the image processing method. The server can call the program code and data of the server to execute the functions of all sub-models.

[0101] The present invention also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the depth map prediction method based on phase-guided Transformer-CNN dual-path fusion according to any one of the above embodiments.

[0102] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0103] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this text, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0104] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

[0105] The above content is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A depth map prediction method based on phase-guided Transformer-CNN dual-path fusion, characterized in that: include: Acquire a fringe image to be processed; Extracting frequency features of phase information from the fringe image using a Fourier model, and acquiring long-range and global frequency features using a Transformer model; Extracting detail features from the stripe image using a CNN model; Cross-fusing the long-distance and global frequency features and the detail features, and performing multi-scale feature enhancement to obtain a depth map of the fringe image; The cross-fusion of the long-distance and global frequency features and the detail features includes: In the encoder, the Transformer branch and the CNN branch perform feature cross-fusion at each layer; In the decoder, the feature image after feature cross-fusion is restored to the resolution of the original stripe image through upsampling, and the feature image after feature cross-fusion in the encoder is merged with the upsampled features of the decoder using skip connections to generate the unwrapped phase; Among them, in the encoder, the Transformer branch and the CNN branch perform feature cross-fusion at each layer through the multi-scale global-local fusion model MGLF for fusing multi-scale local and global features, including: Inputting the stripe image containing the detail features into a convolutional embedding model to convert it into a data shape feature of a Transformer branch, wherein the convolutional embedding model is used to convert the local features into a data shape consistent with the global features; The data shape features converted into the Transformer branch are fused with the long-distance and global frequency features, and are restored to a standard feature map by an image reconstruction model and then optimized by a CNN model, wherein the image reconstruction model is used to re-expand the global features into the data shape of local features and fuse them with the initial local features.

2. The method according to claim 1, characterized in that The extracting frequency characteristics of phase information from the fringe image by using a Fourier model comprises: Converting the fringe image from a spatial domain image to a frequency domain image by Fourier transform; Extracting phase information encoded in the ±1-level spectrum from the frequency domain image and removing zero-frequency features; The image after removing the zero-frequency feature is processed through a feedforward neural network to obtain the frequency feature containing phase information.

3. The method according to claim 1, characterized in that The Transformer model includes: an MHSA model and an FFN model; and the method of using the Transformer model to obtain long-distance and global frequency features includes: The frequency features are sequentially input into the MHSA model and the FFN model to obtain features with long-distance dependency and global frequency.

4. The method according to claim 1, characterized in that The extracting detail features from the stripe image using a CNN model includes: The stripe image is input into a plurality of convolution blocks, and the detail features of the stripe image are extracted by using a convolution layer with an activation function of LeakyReLU and batch normalization.

5. The method according to claim 1, characterized in that: The performing multi-scale feature enhancement to obtain the depth map of the stripe image includes: The unwrapped phase is used as a physical information reference and fused with the upsampled features and input image to obtain high-resolution deep features.

6. A depth map prediction device based on phase-guided Transformer-CNN dual-path fusion, characterized in that: include: An acquisition model is used to acquire a fringe image to be processed; A processing model for extracting frequency features of phase information from the fringe image using a Fourier model, and obtaining long-range and global frequency features using a Transformer model; The processing model is also used to extract detail features from the stripe image using a CNN model; An execution model is used to cross-fuse the long-range and global frequency features and the detail features, and perform multi-scale feature enhancement to obtain a depth map of the stripe image; Wherein, the execution model includes: The second execution sub-model is used in the encoder to perform feature cross-fusion between the Transformer branch and the CNN branch at each layer through the multi-scale global-local fusion model MGLF; The third execution sub-model is used to restore the feature image after feature cross-fusion to the resolution of the original stripe image by upsampling in the decoder, and merge the feature image after feature cross-fusion in the encoder with the upsampled features of the decoder by using a jump connection to generate an unwrapped phase; Furthermore, the second execution sub-model includes: A fourth execution sub-model, used for inputting the stripe image containing the detail features into a convolution embedding model to convert it into a data shape feature of a Transformer branch, wherein the convolution embedding model is used for converting the local features into a data shape consistent with the global features; The fifth execution sub-model is used to fuse the data shape features converted into the Transformer branch with the long-distance and global frequency features, and optimize them by the CNN model after restoring them to the standard feature map through the image reconstruction model, wherein the image reconstruction model is used to re-expand the global features into the data shape of local features and fuse them with the initial local features.

7. The depth map prediction device according to claim 6, characterized in that The acquisition model includes: A first acquisition sub-model is used to convert the fringe image from a spatial domain image to a frequency domain image by Fourier transform; A first processing sub-model is used to extract the phase information encoded in the ±1 level spectrum from the frequency domain image and remove the zero frequency feature; The first execution sub-model is used to process the image after removing the zero-frequency feature through a feedforward neural network to obtain a frequency feature containing phase information.

8. The depth map prediction device according to claim 6, characterized in that: The Transformer model includes: an MHSA model and an FFN model; the processing model includes: The second processing sub-model is used to input the frequency features into the MHSA model and the FFN model in sequence to obtain features with long-distance dependency and global frequency.

Citation Information

Patent Citations

  • Fringe projection time phase unwrapping method based on deep learning

    CN109253708A

  • Transformer-based single fringe pattern depth estimation method

    CN114066959A