Intelligent Two-Dimensional Phase Unwrapping System, Application, Its Training and Dataset Construction Method
By using a semantic segmentation module composed of a symmetric encoder and decoder, semantic segmentation and package number prediction of the two-dimensional phase is solved, and the accuracy problem under the influence of noise in two-dimensional phase expansion in the prior art is achieved, and a high-precision and robust phase expansion effect is achieved.
Patent Information
- Application Number
- CN202210128483.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-02-11
AI Technical Summary
The prior art has low accuracy due to strong noise in two-dimensional phase expansion, and lacks the ability to learn relative positional relationships and non-local correlations. The data set is small and the performance is poor, and there is a problem of category imbalance.
The semantic segmentation module composed of a symmetric encoder and decoder performs semantic segmentation of the two-dimensional phase and predicts the phase wrap count of the segmented area, thereby realizing the expansion of the two-dimensional phase. The system includes a semantic segmentation module and a phase synthesis module, which is trained through weighted cross-entropy loss function to alleviate the problem of category imbalance.
High accuracy is achieved in two-dimensional phase data with strong noise, improved the processing ability of discontinuous phases, and has higher accuracy and stronger robustness. It is suitable for practical measurement technologies.
Smart Images

Figure CN114529723B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of measurement, and more specifically, relates to an intelligent two-dimensional phase unwrapping system, application, its training and dataset construction method. Background Art
[0002] The two-dimensional phase unwrapping problem often appears in various measurement techniques, such as optical interferometry, radar interferometry, magnetic resonance imaging, and fringe projection profilometry (FPP). The signals obtained by these methods are usually represented in the form of complex exponentials, and the wrapped phase is obtained using the arctangent function. The wrapped phase values obtained by the arctangent function are limited to the range of (-π, π], losing part of the information of the original phase. Two-dimensional phase unwrapping needs to infer the original phase map based on the known wrapped phase map. Without constraints on the phase, this problem is ill-posed, mainly because the stability of the solution is poor due to data variations; with the addition of continuity constraints on the original phase, the problem becomes a well-posed problem that is very easy to solve. However, in practical applications, the modalities of the phase are infinite, and factors such as noise, insufficient sampling, and phase mutations will cause the phase to lose continuity. Therefore, phase unwrapping in real-world scenarios is still a challenging ill-posed problem. To solve this problem, many methods have been proposed, mainly including path-dependent methods, path-independent methods, and deep learning-based methods.
[0003] For path-dependent methods, the general idea is to perform integration along a certain path to perform local phase unwrapping. Commonly used methods include the branch cut method (BC), quality-guided path unwrapping (QGPU), minimum discontinuity method (MD), and non-continuous path reliability ranking algorithm (SRNCP), etc. These methods adopt different integral path selection strategies respectively. Path-dependent methods usually have good unwrapping results in the case of low noise, but there is a problem of inconsistent path integration in the case of high noise, resulting in incorrect unwrapping of some points. Moreover, the errors of some points will spread along a certain path, resulting in incorrect phase unwrapping of a region. Generally speaking, this type of method is efficient, but its biggest drawback is the lack of robustness to noise.
[0004] For path-independent methods, the phase unwrapping problem is usually converted into a global optimization problem. Among them, the most commonly used are the least squares (LS) method and the method based on the transport of intensity equation (TIE), both of which are to solve the discrete Poisson equation. The LS method has a certain robustness to noise, but the unwrapped phase is too smooth and the solution efficiency is relatively low. In the case of noisy wrapped phase, the phase map obtained by using the LS method for phase unwrapping has a global error. Guo et al. improved the LS method to alleviate the problem of global error in the phase unwrapping result of the least squares method. The research by Zhao et al. shows that the iterative method based on TIE is more robust than the iterative LS method, but the problem of global error still exists.
[0005] In recent years, the development of deep learning has promoted progress in many fields. Convolutional neural networks (CNNs) have outperformed many traditional methods in solving various problems in computer vision, especially in solving various ill-posed problems in the field of computer vision, demonstrating capabilities that traditional methods do not possess. For example, in pixel-level problems such as image segmentation, image denoising, image generation, image inpainting, and super-resolution, CNN-based methods have surpassed all traditional methods. Phase unwrapping can also be regarded as an image-to-image mapping problem, so CNNs have naturally been applied to the phase unwrapping problem and shown great potential. Spoorthi et al. were the first to transform the phase unwrapping problem into an image semantic segmentation problem, that is, using the wrap-count matrix in the phase unwrapping process as the target of semantic segmentation, and proposed a relatively simple structure PhaseNet and a post-processing method based on merging complementary information. After training PhaseNet using a phase dataset simulated with a mixture of Gaussian noises, the phase unwrapping results obtained by this method are far better than those of QGPU and SRNCP. Dardikman et al. used a residual neural network (ResNet) model consisting only of convolutional layers to perform phase unwrapping on wrapped phases containing steep spatial gradients. Zhang et al. proposed a convolutional denoising segmentation network for phase unwrapping. First, a denoising network is used to denoise the wrapped phase, and then the denoised wrapped phase is fed into the segmentation network, and the segmentation structure is post-processed to obtain the original phase. Since DeepLabv3+ performs well in the field of semantic segmentation, Zhang et al. used it for phase unwrapping and achieved good results. However, if only DeepLabv3+ is used for segmentation without post-processing, the final unwrapping effect will deteriorate. Wang et al. used a network composed of U-Net and residual blocks to perform robust single-step phase unwrapping. Different from the paper, this method directly uses U-Net to learn the mapping from the wrapped phase to the unwrapped phase. Although it is superior to traditional methods in terms of anti-noise and anti-aliasing, this method requires a large dataset for training. Wu et al. proposed a residual encoder-decoder network (REDN) for phase unwrapping. REDN has an encoder-decoder branch and a high-resolution residual branch, and information fusion is gradually carried out between the two branches. It performs better than previous deep learning methods in the application of optical coherence tomography. Recently, Spoorthi et al. proposed PhaseNet2.0, designed a convolutional network based on dense blocks for phase unwrapping, and the network uses a composite loss function for training to alleviate the class imbalance problem. PhaseNet2.0 has achieved significantly better results than PhaseNet in the case of extremely high phase map noise, and performs well in phase unwrapping real 3D models.
[0006] Although current deep learning-based phase unwrapping methods have outperformed traditional methods, as data-driven methods, data, the training process, and the model are equally important. From the perspectives of model structure, dataset, and training, we believe that existing deep learning-based two-dimensional phase unwrapping methods still face the following challenges:
[0007] Lack the ability to learn relative position relationships and non-local correlations. In phase unwrapping, the wrapped number at a certain position usually depends on other regions. However, the receptive field of a CNN usually only gradually expands locally. If the receptive field of the CNN is not large enough, some non-local dependencies cannot be captured. Although downsampling and more convolutional layers can expand the receptive field, the translational invariance of convolution results in the CNN lacking the ability to model relative positions.
[0008] Poor performance when the dataset is small. Current methods using deep learning for phase unwrapping usually use data on the order of 10k for training. Although large datasets are beneficial for the model to learn more generalizable features, it also indicates the insufficient learning ability of existing methods in the few-shot case.
[0009] In practical situations, the height distribution of the phase is always uneven, which leads to the problem of class imbalance. The network trained in this case usually has a relatively low overall error rate, but a relatively high error rate for some classes with few samples. This will have a serious impact on the overall phase unwrapping accuracy. Summary of the Invention
[0010] To address the above deficiencies or improvement requirements of the prior art, the present invention provides an intelligent two-dimensional phase unwrapping system, application, its training, and dataset construction method. The purpose is to use a semantic segmentation module composed of a symmetric encoder and decoder to perform semantic segmentation on the two-dimensional phase, and predict the wrapped number of the segmented regions, thereby unwrapping the two-dimensional phase, and having good accuracy for two-dimensional phase data with strong noise. Thus, it solves the technical problem of low accuracy caused by strong data noise when applying the two-dimensional phase unwrapping method to measurement technology in the prior art.
[0011] To achieve the above object, according to one aspect of the present invention, an intelligent two-dimensional phase unwrapping system is provided, which includes a semantic segmentation module and a phase synthesis module;
[0012] The semantic segmentation module is used to input the wrapped phase, perform semantic segmentation on the wrapped phase into different regions, classify the regions to obtain the wrapped number of each pixel, and submit it to the phase synthesis module; the semantic segmentation module includes a connected encoder and decoder; the encoder and decoder have a symmetric structure;
[0013] The phase synthesis module is used to multiply the number of wraps of each pixel by 2π and synthesize it with the input wrapped phase into the unwrapping of the phase.
[0014] Preferably, in the intelligent two-dimensional phase unwrapping system, the encoder is used to shrink the feature map, and includes a feature map shrinking unit composed of a plurality of residual blocks and a max pooling layer, and the plurality of feature map shrinking units are cascaded; the decoder is used to expand the shrunk feature map, and includes a feature map expanding unit composed of a plurality of transposed convolutional layers and SRB layers, and the plurality of feature map expanding units are cascaded.
[0015] Preferably, in the intelligent two-dimensional phase unwrapping system, the feature map output by the residual block of the encoder is also merged with the feature map of the same size after downsampling output by the corresponding residual block in the decoder on the channel.
[0016] Preferably, in the intelligent two-dimensional phase unwrapping system, the input end of the encoder and the output end of the decoder are connected by an edge enhancement block.
[0017] Preferably, in the intelligent two-dimensional phase unwrapping system, a bottleneck is connected between the encoder and the decoder; the bottleneck includes a connected residual block, a multi-scale fusion block, and / or a spatial self-attention layer.
[0018] Preferably, in the intelligent two-dimensional phase unwrapping system, the spatial self-attention layer takes the feature map as the input, performs linear transformation after convolving the feature map to obtain a sharpened feature map; performs a tensor product of the sharpened feature map and its transposed map, processes it with a softmax function to obtain a spatial attention matrix, and performs a tensor product with the sharpened feature map and then performs linear transformation and adds it to the input feature map pixel by pixel and outputs.
[0019] Preferably, in the intelligent two-dimensional phase unwrapping system, the edge enhancement block includes a convolutional layer and a linear transformation layer; the linear transformation function used by the linear transformation layer is:
[0020] O = r·conv(Ψ, k c ) + b
[0021] where O is the output of the edge enhancement module; r and b are respectively learnable scale parameters and bias parameters, and conv(Ψ, k c ) represents convolving the wrapped phase Ψ using k c as the convolution kernel.
[0022] Preferably, in the intelligent two-dimensional phase unwrapping system, the residual block includes a plurality of sequentially connected convolutional layers, BN layers, and activation layers; the activation layer preferably uses the LeakyReLU activation function; preferably, after the first sequentially connected convolutional layer, BN layer, and activation layer, the output of the branch channel and the adjacent BN layer are added before each convolutional layer.
[0023] According to another aspect of the present invention, there is provided a training method for the intelligent two-dimensional phase unwrapping system, which includes the following steps:
[0024] Construct a data set, the data set includes the unwrapped phase used for output supervision, and the wrapped phase of the training input; the unwrapped phase preferably includes continuous phase and discontinuous phase; the continuous phase is generated by superimposing multiple two-dimensional Gaussian distributions and adding noise; the discontinuous phase is obtained by setting the phase values of the continuous phase and randomly sized rectangular regions with random numbers to zero to simulate phase jumps; the wrapped phase is obtained by wrapping the unwrapped phase with the arctangent function;
[0025] Use weighted cross-entropy loss As the loss function for training.
[0026] Preferably, in the training method of the intelligent two-dimensional phase unwrapping system, during training, the stochastic gradient descent algorithm with momentum is used as the optimizer.
[0027] According to another aspect of the present invention, there is provided a method for constructing a data set for training an intelligent two-dimensional phase unwrapping system, which includes the following steps:
[0028] Superimpose multiple two-dimensional Gaussian distributions and Gaussian noise to generate an unwrapped phase image as the output supervision of the continuous phase type; superimpose multiple two-dimensional Gaussian distributions and Gaussian noise and set the phase values of rectangular regions with random sizes and numbers to zero to generate an unwrapped phase image as the output supervision of the discontinuous phase type;
[0029] Perform an arctangent function operation on the supervised output to obtain a wrapped phase image as the training input;
[0030] Combine the training input and the corresponding output supervision into a data set for training the intelligent two-dimensional phase unwrapping system.
[0031] According to another aspect of the present invention, there is provided an application of the intelligent two-dimensional phase unwrapping system, which is applied to measurement and includes the following steps:
[0032] (1) Obtain measurement data recorded in two-dimensional phase; the measurement data is recorded in the form of complex exponential and the wrapped two-dimensional phase is obtained by using the arctangent function, such as optical interference, radar interference, nuclear magnetic resonance imaging, or fringe projection profilometry;
[0033] (2) Input the two-dimensional phase obtained in step (1) into the intelligent two-dimensional phase unwrapping system to obtain the unwrapping of the phase;
[0034] (3) Obtain the measured value according to the unwrapping of the phase obtained in step (2).
[0035] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects can be achieved:
[0036] The present invention provides an intelligent two-dimensional phase unwrapping system and method. The symmetric encoder-decoder structure is adopted because the symmetric structure performs better than other network structures in the phase unwrapping task. In the preferred solution, a more dense skip connection residual block is used as the basic feature extraction block of the network, which can greatly improve the feature expression ability of the network without increasing the number of parameters. For the phase unwrapping task, a multi-scale fusion block, a spatial self-attention layer, and an edge enhancement block are added to the network to enhance its performance specifically. In the experimental part, compared with two excellent traditional methods and five deep learning-based methods, our method has higher accuracy, stronger robustness, and also has good effects on discontinuous phases. The final ablation experiment shows that the multi-scale fusion block, the spatial self-attention layer, and the edge enhancement block under the condition of continuous phase can effectively improve the network performance. Description of the Drawings
[0037] Figure 1 is a schematic structural diagram of the intelligent two-dimensional phase unwrapping system provided by the present invention;
[0038] Figure 2 is a schematic structural diagram of the semantic segmentation module provided by the present invention;
[0039] Figure 3 is a detailed structural diagram of the module adopted by the present invention;
[0040] Figure 4 is a schematic diagram of the data set constructed in the embodiment of the present invention;
[0041] Figure 5 is a statistical result diagram of the data set in the embodiment of the present invention;
[0042] Figure 6 is a result diagram of the ablation experiment in the embodiment of the present invention. Detailed Embodiments
[0043] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0044] The intelligent two-dimensional phase unwrapping system provided by the present invention includes a semantic segmentation module and a phase synthesis module;
[0045] The semantic segmentation module is used to input the wrapped phase, semantically segment the wrapped phase into different regions, classify the regions to obtain the wrapped number of each pixel, and submit it to the phase synthesis module; the wrapped phase is a two-dimensional phase wrapped by the arctangent function obtained by the measurement technology; the pixels in the region are classified according to the wrapped number they have;
[0046] The semantic segmentation module includes a connected encoder and decoder, and a bottleneck is connected between the encoder and decoder; the encoder and decoder have a symmetric structure, and the symmetric encoder-decoder structure is adopted because the symmetric structure performs better than other structures in the phase unwrapping task.
[0047] The encoder is used to shrink the feature map, preferably including a feature map shrinking unit composed of a plurality of residual blocks and a max pooling layer, and the plurality of feature map shrinking units are cascaded; considering the richness of feature extraction and the model complexity, generally 1-10 feature map shrinking units are cascaded, and preferably 4 feature map shrinking units are cascaded.
[0048] The decoder is used to expand the shrunk feature map, preferably including a feature map expanding unit composed of a plurality of transposed convolutional layers and SRB (Serried Residual Blocks) layers, and the plurality of feature map expanding units are cascaded;
[0049] The bottleneck includes a serried residual block (SRB), an atrous spatial pyramid pooling (ASPP) layer, and a positional self-attention (PSA) layer connected in sequence; the ASPP layer is used to extract and fuse the feature map information of each scale, and can capture and fuse the large-scale spatial semantic features and small-scale detail features at the same time, so that the decoder can utilize richer features; the positional self-attention layer is used to enhance the useful information in the spatial dimension and weaken the weight of the useless information;
[0050] The feature map output by the residual block of the encoder is also merged with the feature map of the same size after downsampling output by the corresponding residual block in the decoder on the channel;
[0051] The input end of the encoder and the output end of the decoder are connected by an edge enhancement block; the edge enhancement block (ERB) is used to explicitly introduce edge information to enhance the boundary of each category domain and further improve the classification accuracy of each pixel in the image.
[0052] The residual block (SRB) includes a plurality of sequentially connected convolutional layers, BN layers and activation layers; preferably includes 1 to 10, preferably 4 sequentially connected convolutional layers, BN layers and activation layers; the convolutional layer preferably uses a 3×3 to 5×5 convolution; the activation layer preferably uses the LeakyReLU activation function, which avoids feature loss and to a certain extent avoids network degradation. Experiments show that compared with the tanh and Sigmoid activation functions, LeakyReLU can avoid gradient explosion and gradient disappearance. Compared with the ReLU function, LeakyReLU avoids the weights with a gradient of zero all the time during network training and increases the fitting ability of the network; after the first sequentially connected convolutional layer, BN layer and activation layer, there is a branch channel before each convolutional layer and it is added to the output of the activation layer, so that the residual block of the present invention has a denser skip connection compared with the traditional residual block, which is not only beneficial to the backpropagation of gradients but also increases the number of paths for feature flow and improves the richness of features. On the other hand, adding batch normalization makes the training more stable and faster, and also helps the network generalization ability to a certain extent.
[0053] The spatial self-attention layer takes the feature map as the input, performs linear transformation after convolving the feature map to obtain a sharpened feature map; performs a tensor product of the sharpened feature map and its transposed map, processes it with the softmax function to obtain a spatial attention matrix, performs a tensor product with the sharpened feature map, then performs linear transformation and adds it to the input feature map pixel by pixel and outputs; the input of the self-attention layer is the feature map, and the larger the value of the feature in the feature map, the more important this feature is. The spatial attention matrix reflects the importance degree of the features at each position in the space relative to other positions, and uses it as a weight to strengthen the important features, thereby improving the accuracy of the network.
[0054] The edge enhancement block includes a convolutional layer and a linear transformation layer; the linear transformation function used in the linear transformation layer is:
[0055] O=r·conv(Ψ,k c )+b
[0056] where O is the output of the edge enhancement module; r and b are learnable scale parameter and bias parameter respectively, and conv(Ψ, k c ) represents convolving the wrapped phase Ψ using k c as the convolution kernel. The edge information obtained by the edge enhancement module will be added to the feature map before the classification output layer, explicitly attaching the edge information of the wrap to the classification output layer, thereby improving the classification accuracy of the wrapped edge pixels.
[0057] The phase synthesis module is used to multiply the number of wraps of each pixel by 2π and synthesize it with the input wrapped phase into the unwrapping of the phase.
[0058] The training method of the intelligent two-dimensional phase unwrapping system provided by the present invention includes the following steps:
[0059] Construct a dataset, the dataset includes the unwrapped phase used as output supervision and the input wrapped phase for training; the unwrapped phase includes continuous phase and discontinuous phase; the continuous phase is generated by superimposing multiple two-dimensional Gaussian distributions with different means and variances and adding different degrees of noise superposition; the discontinuous phase is obtained by setting the phase values of continuous phase and random number of rectangular regions with random sizes to zero to simulate phase jumps; the wrapped phase is obtained by wrapping the unwrapped phase using the arctangent function;
[0060] Use weighted cross-entropy loss as the loss function for training to alleviate the problem that the classification accuracy of the sample categories with fewer numbers caused by class imbalance is significantly lower;
[0061] The weighted cross-entropy loss is written as:
[0062]
[0063] where W c is the weight of category c, and the reciprocal of the proportion of the number of pixels of this category is used as the weight of this category. n is the number of wraps of the pixel, and the pixel is classified according to the number of wraps it has, divided into n categories; represents whether the classification category c of each pixel of sample i conforms to its true category. When the pixel of sample i at (x, y) belongs to category c, y ic (x, y) is 1, otherwise it is 0; p ic is the predicted probability that each pixel of sample i belongs to category c, and B is the batch size during training. Assigning weights can effectively balance the loss values of all categories, thereby increasing the classification accuracy of pixels in the minority categories.
[0064] Preferably, during training, the Stochastic Gradient Descent with Momentum (SGDM) algorithm is used as the optimizer to enable the model to have stronger generalization ability.
[0065] The method for constructing a training dataset for an intelligent two-dimensional phase unwrapping system provided by the present invention includes the following steps:
[0066] Superimpose multiple two-dimensional Gaussian distributions and Gaussian noise to generate an unwrapped phase image as the output supervision for the continuous phase type; the multiple two-dimensional Gaussian distributions have different means and variances, preferably:
[0067] Superimpose multiple two-dimensional Gaussian distributions and Gaussian noise and set the phase values of rectangular regions with random sizes and numbers to zero to generate an unwrapped phase image as the output supervision for the discontinuous phase type;
[0068] Perform an arctangent function operation on the supervised output to obtain a wrapped phase image as the training input;
[0069] Combine the training input and the corresponding output supervision into a training dataset for the intelligent two-dimensional phase unwrapping system.
[0070] The measurement method provided by the present invention includes the following steps:
[0071] (1) Obtain measurement data recorded in two-dimensional phase; the measurement data is recorded in complex exponential form and the wrapped two-dimensional phase is obtained by using the arctangent function, such as optical interference, radar interference, nuclear magnetic resonance imaging, or fringe projection profilometry;
[0072] (2) Input the two-dimensional phase obtained in step (1) into the two-dimensional phase unwrapping system provided by the present invention to obtain the unwrapping of the phase;
[0073] (3) According to the unwrapping of the phase obtained in step (2), obtain the measurement value, specifically transform according to the two-dimensional phase data type to obtain the measurement value; for example, fringe projection profilometry needs to use the phase-height transformation formula to obtain the final surface height..
[0074] The following are examples:
[0075] Two-dimensional phase unwrapping refers to the process of recovering the original phase from the wrapped two-dimensional phase, which can be considered as the inverse process of wrapping the unwrapped phase. First, analyze the process of phase wrapping.
[0076] Usually, in various measurement applications, we will obtain a complex matrix P of size H×W in the form of formula (1):
[0077] P(x, y) = e jΦ(x,y) = cos(Φ(x, y)) + j sin(Φ(x, y)) (1)
[0078] where (x, y) are the coordinates of the pixel and x ∈ [1, H], y ∈ [1, W], and Φ is the unwrapped two-dimensional phase. It is worth mentioning that due to the periodicity of the trigonometric function, when the value range of Φ is greater than a certain range, the function represented by formula (1) is irreversible. Further, we generally use the arctangent function to obtain the phase Ψ of the complex field:
[0079]
[0080] where Ψ(x, y) is the wrapped phase at (x, y), arctan2(·) is the four-quadrant arctangent, and its value range is (-π, π]. Therefore, the value range of the wrapped phase Ψ is (-π, π]. We denote the wrapping operation as wrap(·). The combined effect of formula (1) and formula (2) results in the phase being wrapped, and the wrapped phase loses part of the information of the original phase, making the phase wrapping process irreversible.
[0081] Although the phase unwrapping problem is ill-posed, it can be solved under the constraints of certain assumptions. First, we formulate the phase unwrapping problem. From formula (2), it can be seen that there is a difference of an integer multiple of 2π between the unwrapped phase Φ and the wrapped phase Ψ, that is:
[0082] Φ(x, y) = Ψ(x, y) + 2πk(x, y) (3)
[0083] where k(x, y) is the wrap-count at (x, y), representing the integer value obtained by dividing the difference between the wrapped phase and the original phase by 2π. When only the wrapped phase Ψ is known, as long as the wrap-count matrix k is obtained, the unwrapped phase Φ can be obtained. The process of solving k can be explained as inferring k based on the information provided by Ψ, that is:
[0084]
[0085] where, represents the phase unwrapping method, is the set of phase unwrapping methods, is the wrap-count obtained by solving the phase unwrapping method. From formula (4), it can be seen that the result of is jointly determined by Ψ and f. When Ψ is fixed, f determines the closeness between and k. Our goal is to create an optimal phase unwrapping method f * such that is as close as possible to k, that is:
[0086]
[0087] Among them, M(·, ·) represents a certain distance metric.
[0088] In the ideal case, we assume that the phase Φ is continuous, which is a very strong prior. At this time, the wrapped phase Ψ satisfies the Itoh condition and can be directly phase-unwrapped row by row (column by column). However, in the actual situation, the prior information contained in Ψ is reduced due to the influence of noise and the discontinuity of Φ. In order to specifically analyze the role of prior information, in this paper, we define the prior information as follows:
[0089] Define as the amount of prior information contained in Ψ. I Ψ can be calculated through the formula I Ψ = g(Ψ), and the definition of the function is as follows:
[0090]
[0091] where exp(·) is the exponential function with base e, is the residual point indication matrix, and its definition is as follows:
[0092]
[0093] where the value of q(x, y) can indicate whether the position (x, y) is a residual point [9], and q(x, y) is calculated through the following formula:
[0094] q(x, y) = wrap[Ψ(x + 1, y) - Ψ(x, y)] + wrap[Ψ(x + 1, y + 1) - Ψ(x + 1, y)] - wrap[Ψ(x + 1, y + 1) - Ψ(x, y + 1)] - wrap[Ψ(x, y + 1) - Ψ(x, y)] 8)
[0095] where wrap(·) is the wrapping operation defined by formula (2).
[0096] It can be known that: I Ψ ∈[0, 1]. The fewer the residual points in Ψ, the larger the prior information I Ψ . When I Ψ = 1, all points in Ψ conform to the Itoh condition and can be directly phase-unwrapped row by row (column by column); when I Ψ is close to 0, all traditional methods cannot be used. In other words, the prior information I Ψ reflects the degree to which all phase points in Ψ violate the Itoh condition as a whole. For the phase-unwrapping problem, I Ψ is one of the important factors affecting the phase-unwrapping accuracy. Traditional phase-unwrapping methods can only utilize the limited prior information in Ψ (such as residual points or quality maps) and perform unwrapping according to fixed rules. Therefore, when the noise is large, IΨ is small, and the performance of traditional phase unwrapping methods is poor. The deep learning method learns the rules of phase unwrapping and a large amount of prior information through training. During prediction, it uses a large amount of prior knowledge learned through training to guide phase unwrapping, so it can have stronger robustness to noise.
[0097] The intelligent two-dimensional phase unwrapping system provided in this embodiment, as Figure 1 shown, its input is the wrapped phase, and the output is the unwrapped phase. Among them, the wrapped phase is used as the input of the intelligent two-dimensional phase unwrapping system. The intelligent two-dimensional phase unwrapping system realizes the function of semantic segmentation, divides the wrapped phase into different regions and classifies them. Naturally, the output of the intelligent two-dimensional phase unwrapping system is the category of each pixel, and different categories represent different numbers of wraps. After being trained with a large number of samples, the intelligent two-dimensional phase unwrapping system can be used to predict the number of wraps required for phase unwrapping of the wrapped phase. Finally, the number of wraps predicted by the intelligent two-dimensional phase unwrapping system is multiplied by 2π and then added to the wrapped phase to obtain the unwrapped phase.
[0098] The intelligent two-dimensional phase unwrapping system includes a semantic segmentation module and a phase synthesis module;
[0099] The semantic segmentation module is used to input the wrapped phase, perform semantic segmentation on the wrapped phase into different regions, classify the regions to obtain the number of wraps of each pixel, and submit it to the phase synthesis module; the wrapped phase is the two-dimensional phase wrapped by the arctangent function obtained by the measurement technology; the pixels are classified according to the number of wraps they have;
[0100] The semantic segmentation module, its structure is as Figure 2 shown, including a connected encoder and decoder, and a bottleneck is connected between the encoder and decoder; the encoder and decoder have a symmetric structure;
[0101] The encoder is used to shrink the feature map, including a feature map shrinking unit composed of multiple residual blocks and a max pooling layer, and the multiple feature map shrinking units are cascaded; Figure 2 Among them, the part where the feature map continuously shrinks is called the encoder. In this embodiment, it is composed of four residual blocks (SRBs) and a max pooling layer alternatingly.
[0102] The decoder is used to expand the shrunk feature map, including a feature map expanding unit composed of multiple transposed convolutional layers and SRB layers, and the multiple feature map expanding units are cascaded; Figure 2 Among them, the part where the feature map continuously expands is called the decoder. In this embodiment, it is composed of four transposed convolutions and SRBs alternatingly.
[0103] The bottleneck includes a sequentially connected residual block (SRB), an ASPP layer, and a spatial self-attention layer (PSA layer); the ASPP layer is used to extract and fuse feature map information of multiple scales, which can capture large-scale spatial semantic features and small-scale detailed features simultaneously and fuse them, enabling the decoder to utilize richer features; the spatial self-attention layer is used to enhance the useful information in the spatial dimension and weaken the weight of the useless information; Figure 2 In this embodiment, the part called the bottleneck between the encoder and the decoder contains an SRB, an ASPP, and a PSA. The size of the feature map in the bottleneck module is the smallest in the entire network and the features it contains are also the most abstract. To obtain the largest possible receptive field and keep the original information from being lost, we borrowed the ASPP structure in DeepLab to extract and fuse information of each scale. The PSA can effectively enhance the useful information in the spatial dimension and weaken the weight of the useless information. Therefore, we added this module after the ASPP to further screen the fused information. In addition, the semantic segmentation module designed in this embodiment has the characteristics of edge enhancement and self-attention enhancement, named the Edge-Enhanced Self-Attention Network (EESANet), and also includes an additional branch connecting the input image and the decoder output. This branch contains an edge reinforce block (ERB), and its output is finally connected to the classification output layer after being merged with the decoder output in the channel direction. The final classification output layer is a 1×1 convolutional layer with 16 channels and uses the softmax activation function.
[0104] The input end of the encoder and the output end of the decoder are connected by an edge enhancement block; for the edge enhancement block (ERB), for each feature map output by the residual block, an additional branch needs to be led out and merged with the feature map of the same size after downsampling in the decoder in the channel dimension, so that the decoder can fuse the information at different scales at the encoder end as a detail supplement.
[0105] The residual block (SRB), as Figure 3 (a) shows, respectively includes four convolutional layers, a BN layer, and a LeakyReLU layer, and the connection order is: convolution, BN, LeakyReLU. The SRB adopts a more dense skip connection. Except for the first 3×3 convolution for channel transformation, a branch is separated before each subsequent convolution layer and added to the output of the adjacent BN layer.
[0106] The spatial self-attention layer (PSA), as Figure 3 (b), the input feature map and the output feature map Have the same size (H is the height, W is the width, and N is the number of channels). The input X is divided into four branches. After three of the branches are linearly transformed through 1×1 convolutional layers with N channels, the feature maps B, C, and D are obtained, where Reshape B, C, and D into in size and perform a tensor product of B after transposing it with C. Then use softmax to process the result of the tensor product to obtain the spatial attention matrix
[0107]
[0108] where S ij is the influence of the i-th position on the j-th position in the space (i, j = 1, 2,..., HW). C i is the i-th row vector of C, is the transposed j-th column vector of B. exp(·) is the exponential function with base e. At the same time, the tensor obtained by performing a tensor product of S and the reshaped D has a size of Reshape this result into in size and multiply it by a scale parameter α. Then add it to the original feature map element-wise to obtain the output The whole process is as follows:
[0109] Z = α·reshape(S·D, {H, W, N}) + X (11)
[0110] where α is a learnable parameter, and here we set its initial value to 0.1. reshape(A, {·}) means reshaping the tensor A into the shape of {·}.
[0111] In this embodiment, the multi-scale fusion block (ASPP), its structure is as Figure 3 (c) shown. The input feature map is divided into five branches, one of which passes through a 1×1 convolutional layer. The other three branches respectively pass through 3×3 dilated convolutional layers with dilation rates of 3, 5, and 8. The last branch first performs global average pooling and then up-samples (bilinear interpolation) to the size of the original feature map after passing through a 1×1 convolutional layer. The feature maps output by the five branches are merged in the channel dimension and then passed through a 1×1 convolutional layer to compress the number of channels to the same as that of the input feature map. The ASPP structure can capture both large-scale spatial semantic features and small-scale detailed features and fuse them, enabling the decoder to have richer features to utilize.
[0112] The edge enhancement block obtains a quality map by performing a certain operation on the wrapped phase and classifies pixels under the guidance of the quality map. Its structure is as Figure 3As shown in (d), its input is a wrapped phase diagram with 1 channel. After passing through a 3×3 convolution with fixed parameters and 1 channel, a linear transformation is then performed on the entire feature map, which is expressed as follows:
[0113] O = r·conv(Ψ, k c ) + b (12)
[0114] where O is the output of the ERB module. r and b are learnable scale parameter and bias parameter respectively, and are initialized to 1 and 0 respectively before training. conv(Ψ, k c ) represents convolving the wrapped phase Ψ using k c as the convolution kernel. To obtain the edge information of the wrapped phase using the ERB module, the Laplacian commonly used in image processing is selected here as the convolution kernel. The edge information obtained by the ERB will be added to the feature map before the classification output layer, explicitly attaching the edge information of the package to the classification output layer.
[0115] The number of channels of each residual module and transposed convolution depends on the parameter C. The number of output channels of each convolution layer in the first residual block is C, and then the number of channels doubles after each max pooling layer and halves after each transposed convolution layer. At the same time, the number of channels of the first transposed convolution layer is 8C, and the number of channels of each subsequent layer is halved. To balance the running efficiency and accuracy, C is set to 48.
[0116] The phase synthesis module is used to multiply the number of wraps of each pixel by 2π and synthesize it with the input wrapped phase into the unwrapping of the phase.
[0117] The training method of the intelligent two-dimensional phase unwrapping system provided by the present invention includes the following steps:
[0118] Construct a dataset:
[0119] Generally, methods based on deep learning require a large amount of data pairs for training. It is very difficult to obtain a large amount of wrapped phase and unwrapped phase data pairs in actual measurements. Therefore, we need to construct a dataset by generating. Since phase unwrapping is the inverse process of phase wrapping, according to formula (1) and formula (2), any phase can be wrapped, which can help us generate a large number of data pairs required for training.
[0120] The dataset we constructed is divided into two categories: continuous phase and discontinuous phase. As Figure 4 (a) shows, for the continuous phase, we use the superposition of multiple two-dimensional Gaussian distributions with different means and variances to generate phase images (H×W: 256×256 pixels, and the phase is in the range of 0rad to 100rad), and add Gaussian noise with different degrees to the unwrapped phase. As Figure 4As shown in Fig. (d), the discontinuous phase sets the phase values of rectangular regions with random sizes and numbers to zero on the basis of the continuous phase to simulate phase jumps. Substituting the generated phase data into Eqs. (1) and (2) gives the wrapped phase, as shown in Fig. Figure 4 (b) and Fig. 4(e). To train the pixel classification network, we further use the generated phase and the wrapped phase to calculate the wrapped number k:
[0121]
[0122] where round(·) represents rounding to the nearest integer. Figure 4 Figs. (c) and 4(f) show the corresponding wrapped number maps, and regions of different colors represent classes corresponding to different wrapped numbers. Since the simulated phase value has a dynamic range of 0 rad to 100 rad, the value of the wrapped number is from 0 to 15, that is, there are 16 classes of wrapped numbers. Taking the wrapped phase as the input data and the wrapped number as the output data, the two are combined as a pair of samples of the dataset.
[0123] To more comprehensively verify the performance of the proposed network, we generated multiple datasets with different characteristics. First, we generated 1000 samples of the first type of data with signal-to-noise ratios of 0 dB, 3 dB, 5 dB, 10 dB, and 100 dB, and named them WPDC0, WPDC3, WPDC5, WPDC10, and WPDC100 respectively; randomly generated 5000 samples of the first type of data from 0 dB to 15 dB and named it WPDCA. Randomly generated 10000 samples of the second type of data from 0 dB to 15 dB and named it WPDCS. All datasets were divided into training sets, validation sets, and test sets according to the ratios of 0.8, 0.1, and 0.1 respectively. To more clearly describe the datasets we generated, we listed the details of each dataset in Table 1. It is worth mentioning that since the model proposed in this paper is a classification model, the wrapped number data needs to be one-hot encoded.
[0124] Table 1. List of generated datasets
[0125]
[0126] Methods such as PhaseNet that have now been publicly published use the superposition of two-dimensional Gaussian distribution functions to generate an unwrapped phase map. The present invention also adopts a similar method. In addition, on the basis of superposing multiple two-dimensional Gaussian distributions and Gaussian noise, the present invention sets the phase values of rectangular regions with random sizes and numbers to zero to generate an unwrapped phase map as the output supervision of discontinuous phases. The discontinuous phases enrich the distribution of phase data, thereby enabling the trained model to have stronger practicality. The output supervision data is subjected to a wrapping function operation to obtain a wrapped phase as the training input. The training input and the corresponding output supervision are combined into a dataset sample for training the intelligent two-dimensional phase unwrapping system.
[0127] As Figure 5 shown, we have statistically analyzed the true categories of all phases in the dataset used. It can be clearly seen that the proportions of each category are seriously unbalanced, and the number of some categories is even more than a hundred times that of the least category. To alleviate the category imbalance, the loss function in the training stage of this network uses weighted cross-entropy loss
[0128]
[0129] where W c is the weight of category c; represents the true category of each pixel of sample i. For example, when the pixel of sample i at (x, y) belongs to category c, y ic (x, y) is 1, otherwise it is 0; represents the predicted probability that each pixel of sample i belongs to category c; B is the batch size during training. The weight W c is determined according to the required dataset. In this article, we statistically analyze the category of each pixel in all images of the dataset, and use the reciprocal of the proportion of the number of pixels of a certain category as the weight of this category:
[0130]
[0131] where ω c ∈[0,1] is the proportion of all pixels of category c in the total number of pixels. When calculating the loss, the weights allocated according to formula (12) can effectively balance the loss values of all categories, thereby increasing the classification accuracy of pixels of minority categories. However, because the value of ω c depends on the dataset, using a larger and more realistic dataset during training can obtain better results.
[0132] Before the training starts, we initialize all the parameters of the convolutional layers. During the training phase, we use Stochastic Gradient Descent with Momentum (SGDM) as the optimizer because, for low-level vision tasks, networks trained with SGDM have stronger generalization performance than those trained with Adam. Here, the momentum is 0.9, the initial learning rate is 0.02, and the training batch size is 10. For datasets with different numbers of samples, we set different maximum iteration numbers. Specifically, the five datasets WPDC0 to WPDC100, each containing 1000 pairs of samples, are trained for 400 epochs, the dataset WPDCA is trained for 150 epochs, and the dataset WPDCS is trained for 75 epochs. The specific rules for dividing the training set and test set are given in Section 2.2. The input to the network is a 256×256 single-channel wrapped phase image, and the output is one-hot encoded classification data of H×W×Class: 256×256×16. To avoid excessive impact of precision loss on the results, we save the wrapped image data using 16-bit encoding instead of the commonly used 8-bit encoding for grayscale images. Due to the property of convolutional networks sharing parameters in space, the size of the input image is variable during the prediction phase. However, since the parameters of structures such as the depth of convolutional layers, up / downsampling, and ASPP are fixed, too large or too small input image sizes during the prediction phase will cause performance degradation.
[0133] The network structure we proposed is implemented in the Deep Learning Toolbox based on MATLAB 2020b. Both the training and testing of the network are carried out on a workstation with an Intel Xeno E5-2678 CPU, 128 GB of memory, and an NVIDIA GeForce GTX 1080Ti GPU.
[0134] In this embodiment, the implemented phase unwrapping system is compared and tested with existing phase unwrapping methods. The phase unwrapping methods used for comparison include: SRNCP, TIEPU, PhaseNet, DeepLabV3+, DLPU, REDN, and PhaseNet2.0. Among them, SRNCP is a path-dependent method, TIEPU is a path-independent method, and the rest are deep learning methods.
[0135] The performance of different network structures in the phase unwrapping task was compared. The results show that the encoder and decoder with symmetric structures have smaller root mean square error (RMSE) and smaller computational complexity (#FLOPs) for predicting a single sample, thus having excellent prediction performance. Experiments were carried out under different levels of Gaussian noise, and wrapped phases with different levels of Gaussian noise were used to analyze the anti-noise performance of each phase unwrapping algorithm. The results show that there will be a large number of errors in the phases unwrapped by SRNCP at high noise, but the phase unwrapping can be basically error-free at low noise. An overall deviation appears in the results of the TIEPU method, which is a characteristic of path-independent methods. For example, when SNR = 3 dB (I Ψ = 0.5713), although there are only sporadic points with errors in the unwrapped phase, there is a large deviation overall. The results of PhaseNet and REDN have similar characteristics. Their unwrapping results are better in the case of high noise and small number of phase wraps than in the case of low noise and large number of phase wraps. In addition, the results of PhaseNet 2.0 and DeepLabV3+ have strong similarities, especially at low noise (SNR = 100 dB, I Ψ = 1), where errors always occur at the edges of the wraps in the phases they unwrap. Particularly, DLPU is a method that directly learns the mapping from the wrapped phase to the true phase, so its unwrapping error is usually not a jump of 2π times, but relatively continuous, and the error distribution range is wide. In the case of high noise, only sporadic unwrapping errors occur in the present invention, and the phase unwrapping can be almost error-free in the case of low noise.
[0136] Experiment on discontinuous phases: To further verify the performance of the proposed method in unwrapping discontinuous phases, we conducted experiments on the WPDCS dataset with discontinuous phases. Looking at the phase errors of the phase unwrapping results of each method, there are longitudinal linear phase errors in the results of SRNCP; there are large-area phase errors in the results of TIEPU; Phase errors still occur at the edges of the wraps in DeepLabV3+ and PhaseNet 2.0; Phase errors occur at the edges of a part of the discontinuous phases in the results of PhaseNet and REDN; the results of DLPU remain stable and are not affected by the discontinuous edges; the errors in the results of the present invention are scattered in the phase wrap region, and no obvious errors occur at the discontinuous edges.
[0137] Phase unwrapping test using an object model: The phase of an actual object has both continuous and discontinuous parts. It is very important to test various phase unwrapping algorithms on an actual object model. For deep learning methods, we must understand the generalization ability of the trained network on new samples with different distributions in order to estimate the practicality of the network. The method of this embodiment has the highest phase unwrapping accuracy among all methods. In particular, the RMSE of the phase unwrapping result for the rabbit model reaches 0.54, which is one-fourth of that of the sub-optimal method DeepLabV3+. This shows that the phase unwrapping system and method provided by the present invention have good generalization ability.
[0138] Application test in fringe projection profilometry: Fringe projection profilometry (FPP) is a widely used active three-dimensional measurement method. Its core step is to project a grating fringe image with certain rules onto the surface of the object to be measured, and then use an image acquisition device to obtain the projected image and restore the surface depth of the object to be measured through the modulated deformed fringe information in the image. Taking Fourier transform profilometry (FTP) as an example, FTP usually includes three main steps: wrapped phase extraction, phase unwrapping, and phase-to-height mapping. Due to the possible phenomenon of spectral aliasing in the image, the extracted wrapped phase may be discontinuous and have fringe artifacts. There are many distortions in the wrapped phase obtained by FTP. The simulation test results show that the RMSE of the phase unwrapping method provided by this embodiment is 0.78, which is the best result among all tested phase unwrapping methods. The other better results are: SRNCP (RMSE = 1.22) and REDN (RMSE = 1.10).
[0139] We conducted an ablation study on the proposed method to quantitatively analyze the impact of each part of the network on the network performance. The results of each network in the ablation experiment on datasets with different noises are as Figure 6 shown. At a high noise level, the network without ERB is better than the original network with ERB. However, at a low noise level, the situation is the opposite. This is because ERB directly generates the edges of the wrapped phase and directly merges them into the feature map before the output layer. At low noise, the final output layer can make full use of the edge information. However, at high noise, the serious interference generated by the noise affects the performance of the output layer instead. The performance of the network with further removal of ASPP and PSA has a significant decline compared to the previous two networks at all noise levels. This shows that the ability of ASPP to increase the receptive field and the spatial self-attention are very effective.
[0140] The training and test results of each network in the ablation experiment on the WPDCA and WPDCS datasets are given in Table 3. Among them, the test results on the phase-continuous WPDCA dataset show that the network with ERB is better than the network without ERB, and the network containing ASPP and PSA has little difference from the network without both. The test results on the discontinuous WPDCS dataset show that the error increases by about 25% after removing ASPP and PSA. However, adding ERB on the basis of including ASPP and PSA in the network reduces the performance, indicating that ERB has a negative impact on the unwrapping of discontinuous phases.
[0141] Table 2. RMSE of each network in the ablation experiment tested on the WPDCA and WPDCS datasets
[0142]
[0143] For CNN, inference speed, resolution sensitivity, and anti-noise ability are extremely important. First, we counted the time required for phase unwrapping of a single-phase image by the deep learning-based method and listed it in Table 3. The speed of the proposed method is slower than that of DLPU, DeepLab V3+, and PhaseNet 2.0, but faster than REDN and PhaseNet. In the case of a resolution of 256×256 pixels and a GPU of RTX 1080Ti, the time consumption for phase unwrapping is 23.9 ms. It is worth mentioning that usually CNN runs on the GPU, and other traditional methods are not compared because they can only run on the CPU. Second, we tested the phase unwrapping performance of EESANet under two cases of non-uniform noise. The results show that EESANet trained only with samples of uniform noise distribution still maintains high accuracy under non-uniform noise. Third, in practical applications, the phase images obtained usually have different resolutions, so we tested the phase unwrapping performance of EESANet under different resolutions. The results show that when the resolution is not close to 256×256 pixels, the performance of the network will drop significantly. However, if the resolution of the phase image is adjusted to 256×256 pixels by upsampling / downsampling, the performance of the network still remains good, and the corresponding downsampling / upsampling of the results can restore the original resolution.
[0144] Table 3 Time for phase unwrapping of 256×256 pixel phase images by deep learning methods on the RTX 1080Ti GPU
[0145]
[0146] The accuracy and stability of existing deep learning-based methods have been greatly improved after being trained on larger datasets. However, the improvement of the present invention is relatively small, indicating that the present invention can achieve good results with a smaller number of samples without the need for a large number of labeled samples, thus being more practical in actual situations where it is difficult to obtain labeled samples. On the other hand, dynamic noise levels enhance sample diversity, and stronger sample diversity is beneficial for the convolutional neural network to learn more accurate inductive biases. Therefore, the performance of all deep learning-based methods has been improved. Generally speaking, the present invention has good effects under different noise levels and has strong robustness to noise.
[0147] The experimental results show that the method we proposed has obvious advantages over other methods in the problem of discontinuous phase unwrapping. This advantage mainly benefits from the ASPP and PSA in EESANet. PSA can enhance spatial information with a global receptive field, thereby being able to distinguish the edges of phase mutations and the edges generated by wrapping from a global perspective. This is one of the reasons why our method has good performance in unwrapping discontinuous phases. In addition, the newly designed SRB also plays a key role. On the one hand, more skip connections not only facilitate the backpropagation of gradients but also increase the number of paths for feature flow, improving the richness of features. On the other hand, adding batch normalization makes the training more stable and faster, and also helps to some extent with the generalization ability of the network. It is worth mentioning that using LeakyReLU as the activation function avoids feature loss and, to a certain extent, avoids network degradation.
[0148] The present invention has obvious advantages in the accuracy of phase unwrapping for samples with features different from the training set compared to other methods. This shows that the method we proposed has strong generalization ability and also verifies the feasibility of applying it to FPP and even other measurement techniques that require phase unwrapping.
[0149] The ablation study shows that in the network we proposed, ASPP and PSA can bring stable performance improvement to the network, and ERB can improve the performance of the network in the case of continuous phases and not serious noise. ASPP can effectively increase the receptive field of the feature map and add the global information provided by global pooling. PSA weights the spatial dimension of the feature map to reduce the interference caused by unimportant features. Therefore, ASPP and PSA in EESANet provide more important global features.
[0150] To approximate the phase distribution in actual applications as closely as possible, the dataset we designed inevitably exhibits class imbalance. Among the deep learning-based phase unwrapping methods for comparison, PhaseNet 2.0 takes into account the class imbalance problem and uses a composite loss function to alleviate it. To mitigate the class imbalance problem, we adopted weighted cross-entropy loss instead of other methods such as focal loss. This is mainly because the loss was unstable when we trained using focal loss, and the final results were also inferior to those of weighted cross-entropy loss. It should be noted that the weights were obtained by statistically analyzing the dataset before training. When the samples in the dataset are sufficiently rich and the sample size is large enough, the network trained with these weights has better generalization ability. In addition to the loss function, some training techniques for convolutional neural networks and knowledge distillation methods can also be applied to our network to further improve performance.
[0151] The system provided by the present invention is trained using WPDCA and predictions are made on the test set. The classification matrix is obtained by counting the predicted classes of all pixels and their corresponding true classes. The misclassified pixels mainly concentrate on the cases with smaller numbers of wrappings. As the number of wrappings increases, the proportion of misclassified pixels gradually decreases. However, due to the class imbalance phenomenon, as the number of wrappings gradually increases, the possibility of classification error gradually increases, and the misclassification index is usually less than the correct class index and only differs by 1. This phenomenon is still caused by the class imbalance of the dataset samples.
[0152] It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. An intelligent two-dimensional phase unwrapping system, characterized in that, it includes a semantic segmentation module and a phase synthesis module; The semantic segmentation module is used to input the wrapped phase, semantically segment the wrapped phase into different regions, classify the regions to obtain the wrapped number of each pixel, and submit it to the phase synthesis module; the semantic segmentation module includes a connected encoder and decoder; the encoder and decoder have a symmetric structure; The encoder is used to shrink the feature map, including a feature map shrinking unit composed of multiple residual blocks and a max pooling layer, and multiple feature map shrinking units are cascaded; the decoder is used to expand the shrunk feature map, including a feature map expanding unit composed of multiple transposed convolutional layers and SRB layers, and multiple feature map expanding units are cascaded; The feature map output by the residual block of the encoder is also merged with the feature map of the same size after downsampling output by the corresponding residual block in the decoder on the channel; a bottleneck is connected between the encoder and the decoder; the bottleneck includes a connected residual block, a multi-scale fusion block, and / or a spatial self-attention layer; The input end of the encoder and the output end of the decoder are connected by an edge enhancement block; The phase synthesis module is used to multiply the wrapped number of each pixel by 2π and synthesize it with the input wrapped phase into the unwrapped phase.
2. The intelligent two-dimensional phase unwrapping system according to claim 1, characterized in that, The spatial self-attention layer takes the feature map as input, convolves the feature map and then performs a linear transformation to obtain a sharpened feature map; performs a tensor product of the sharpened feature map and its transposed map, processes it with a softmax function to obtain a spatial attention matrix, performs a tensor product with the sharpened feature map, then performs a linear transformation and adds it to the input feature map pixel by pixel and outputs.
3. The intelligent two-dimensional phase unwrapping system according to claim 1, characterized in that, The edge enhancement block includes a convolutional layer and a linear transformation layer; the linear transformation function used by the linear transformation layer is: O = r·conv(Ψ, k c ) + b where O is the output of the edge enhancement module; r and b are learnable scale parameter and bias parameter respectively, and conv(Ψ,k c ) represents convolving the wrapped phase ψ using k c as the convolution kernel.
4. The intelligent two-dimensional phase unwrapping system according to claim 1, characterized in that, The residual block includes multiple sequentially connected convolutional layers, BN layers and activation layers.
5. The intelligent two-dimensional phase unwrapping system according to claim 4, characterized in that, The activation layer uses the LeakyReLU activation function.
6. The intelligent two-dimensional phase unwrapping system according to claim 4, characterized in that, After the first sequentially connected convolutional layer, BN layer and activation layer, the output of the branch channel and the adjacent BN layer are added before each convolutional layer.
7. The application of the intelligent two-dimensional phase unwrapping system according to any one of claims 1 to 6, characterized in that, for measurement, including the following steps: (1) Obtain measurement data recorded in two-dimensional phase; the measurement data is recorded in complex exponential form and the wrapped two-dimensional phase is obtained by using the arctangent function; (2) Input the two-dimensional phase obtained in step (1) into the intelligent two-dimensional phase unwrapping system to obtain the unwrapped phase; (3) Obtain the measurement value according to the unwrapped phase obtained in step (2).
8. The training method of the intelligent two-dimensional phase unwrapping system according to any one of claims 1 to 6, characterized in that, it includes the following steps: Construct a data set, the data set includes the unwrapped phase used for output supervision and the wrapped phase of the training input; the unwrapped phase includes continuous phase and discontinuous phase; the continuous phase is generated by superimposing multiple two-dimensional Gaussian distributions and adding noise superposition; the discontinuous phase is obtained by setting the phase values of continuous phase and rectangular regions with random sizes and random numbers to zero to simulate phase jumps; the wrapped phase is obtained by wrapping the unwrapped phase with the arctangent function; Use weighted cross-entropy loss as the loss function for training.
9. The training method of the intelligent two-dimensional phase unwrapping system according to claim 8, characterized in that, the random gradient descent algorithm with momentum is used as the optimizer.
10. The training method of the intelligent two-dimensional phase unwrapping system according to claim 8, characterized in that, its method for constructing the training data set includes the following steps: Superimpose multiple two-dimensional Gaussian distributions and Gaussian noise to generate an unwrapped phase image as the output supervision of the continuous phase type; superimpose multiple two-dimensional Gaussian distributions and Gaussian noise and set the phase values of rectangular regions with random sizes and numbers to zero to generate an unwrapped phase image as the output supervision of the discontinuous phase type; Perform an arctangent function operation on the supervised output to obtain a wrapped phase image as the training input; Combine the training input and the corresponding output supervision into a data set for training the intelligent two-dimensional phase unwrapping system.
Citation Information
Patent Citations
Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field
AU2020103901A4
Rapid automatic semantic image segmentation model method
CN104504725A