An artificial intelligence-based virtual phase shifting method for interferograms
Patent Information
- Application Number
- CN202610589241.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-04-30
AI Technical Summary
然而,传统相移干涉术高度依赖物理移相装置(如压电陶瓷移相器)对相位步长的精确控制,这带来了环境敏感性高(振动、气流扰动导致帧间相位误差)、测量速度受限(机械移相过程耗时,难以满足动态过程观测需求)以及系统成本与复杂度增加等问题
[0013]本发明通过构建深度学习模型作为虚拟移相器,仅需采集少量(M≥1)参考相位随机且未知的干涉图,即可实时生成一组具有标准相移量的参考干涉图序列,从而彻底摆脱对物理移相器(如压电陶瓷)的依赖,消除了环境振动与气流扰动引入的帧间相位误差,同时突破了空间载波法对条纹方向与频率的苛刻要求;相较于直接从单帧干涉图端到端预测相位的现有深度学习方法,本发明保留了成熟多步相移算法的解析解算优势,显著提升了相位恢复的精度与泛化能力,实现了单次曝光或少量随机曝光下的高精度动态测量,降低了系统成本与复杂度。
Smart Images

Figure CN122473490B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of optical precision measurement and artificial intelligence, and specifically relates to an artificial intelligence-based virtual phase shifting method for interferograms. Background Technology
[0002] Phase-shifting interferometry is a core technology in optical precision measurement, widely used in high-precision measurement scenarios such as lens surface shape inspection and micro / nano structure characterization. Its basic principle is to introduce multiple known reference phase shifts, acquire a corresponding number of interferograms, and use phase-shifting algorithms to calculate the phase distribution of the measured object. However, traditional phase-shifting interferometry heavily relies on the precise control of the phase step size by physical phase-shifting devices (such as piezoelectric ceramic phase shifters). This leads to problems such as high environmental sensitivity (vibration and airflow disturbances cause inter-frame phase errors), limited measurement speed (the mechanical phase-shifting process is time-consuming and difficult to meet the needs of dynamic process observation), and increased system cost and complexity.
[0003] To overcome the aforementioned bottlenecks, existing technologies are mainly developing in two directions: one is to use the space carrier method to achieve single-frame interferometry. This method requires the interference fringes to have a specific direction and a sufficiently high spatial frequency, which is extremely demanding on optical path adjustment and does not perform well when dealing with complex or closed fringes; the other is to use deep learning to directly predict the phase distribution from end to end in a single-frame interferogram in recent years. Although this method avoids physical phase shifting, it often suffers from weak generalization ability and limited accuracy upper limit because it skips the physically rigorous phase shift calculation process, making it difficult to meet the application requirements of metrology.
[0004] Therefore, how to generate a set of interferogram sequences with standard phase shifts from a small number of randomly acquired interferograms without relying on physical phase shifters and avoiding poor generalization of end-to-end phase prediction, so as to reuse mature multi-step phase shifting algorithms for high-precision phase calculation, has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] To address the problems existing in the background art, one aspect of the present invention provides an artificial intelligence-based interferogram virtual phase shifting method, comprising:
[0006] S1: Obtain the initial interferogram training set, which includes training samples obtained by observing and collecting data from the object under test. Each training sample contains M frames of random interferograms and N frames of reference interferogram labels with standard phase delay collected under the phase distribution to be tested.
[0007] S2: Construct a deep learning model as a virtual phase shifter, use M frames of random interferograms as input to the virtual phase shifter, and use the virtual phase shifter to predict N frames of reference interferograms with standard phase delay.
[0008] S3: Construct a loss function for training the virtual phase shifter, and train the virtual phase shifter based on the constructed loss function. The loss function includes at least one of spatial domain basic loss, high-frequency region enhancement loss, frequency domain basic loss, and frequency domain high-frequency loss.
[0009] S4: The M-frame random interferograms of the object under test are acquired by the interferometer and input into the trained virtual phase shifter for inference to predict the N-frame reference interferograms with standard phase delay.
[0010] Another aspect of the present invention provides an artificial intelligence-based interferogram virtual phase-shifting system, the system comprising a memory and a processor; the memory is used to store an application program; the processor is used to run the application program and execute the artificial intelligence-based interferogram virtual phase-shifting method.
[0011] Another aspect of the present invention provides a computer storage medium storing a computer program, which, when executed by a processor, implements the aforementioned artificial intelligence-based interferogram virtual phase shifting method.
[0012] The present invention has at least the following beneficial effects
[0013] This invention constructs a deep learning model as a virtual phase shifter. It only requires the acquisition of a small number (M≥1) random and unknown reference phase interferograms to generate a set of reference interferogram sequences with standard phase shifts in real time. This completely eliminates the dependence on physical phase shifters (such as piezoelectric ceramics) and eliminates inter-frame phase errors introduced by environmental vibrations and airflow disturbances. At the same time, it breaks through the stringent requirements of the space carrier method on fringe direction and frequency. Compared with existing deep learning methods that directly predict the phase from a single frame interferogram end-to-end, this invention retains the analytical solution advantages of mature multi-step phase shifting algorithms, significantly improves the accuracy and generalization ability of phase recovery, realizes high-precision dynamic measurement under single exposure or a small number of random exposures, and reduces system cost and complexity. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the method flow of the present invention;
[0015] Figure 2 This is a schematic diagram of the model of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0018] Please see Figure 1 and Figure 2 This invention provides an artificial intelligence-based virtual phase-shifting method for interferograms, comprising:
[0019] S1: Obtain the initial interferogram training set, which includes training samples obtained by observing and collecting data from the object under test. Each training sample contains M frames of random interferograms and N frames of reference interferogram labels with standard phase delay collected under the phase distribution to be tested.
[0020] Preferably, the M-frame random interferogram includes: an M-frame interferogram with unknown reference phase and unknown inter-frame phase difference; the N-frame standard phase delay reference interferogram includes: N equally spaced standard phase delay reference interferograms required by the selected phase resolution algorithm.
[0021] There are various phase resolution algorithms in the prior art, among which typical phase resolution algorithms include two-step phase resolution algorithms and four-step phase resolution algorithms. This embodiment mainly uses these two algorithms to specifically illustrate this application:
[0022] For the two-step phase resolution algorithm:
[0023] The two-step phase resolution algorithm requires a standard phase shift of... Two interferograms; in this scenario, first generate a random interferogram. Its true reference phase It is unknown, but can be considered as 0, with a corresponding standard phase shift of 0; refer to the interferogram. for A frame virtual interferogram is defined as a virtual interferogram in phase. Up phase shift The resulting interferogram corresponds to the standard phase shift. According to the interferogram and The phase distribution can then be calculated using a two-step phase-solving algorithm. During this process, the training set must contain all possible phase distributions of the object being measured. and various background light intensities Harmony system The distribution pattern; in this embodiment, the Zernike polynomial (circular pupil) is used to generate the phase distribution for the circular pupil. The phase distribution is generated using Legendre polynomials for square pupils. In combination with the set background light intensity Harmony system ,pass , Training samples are generated, and Gaussian noise and a small amount of speckle noise can be applied to the training samples to enhance the model's robustness. Simultaneously, in this embodiment, various standard mirrors, lenses with local defects, or samples with different reflectivities can be acquired using an interferometer (e.g., a Twyman-Green interferometer or a Fizeau interferometer equipped with a PZT phase shifter), and the phase shifting is driven by the PZT phase shifter. , generate containing and The training samples were used; real images were randomly cropped, rotated, and slightly scaled, and different levels of measured noise were added for data augmentation; all images were normalized to [value missing]. The simulation data ensures coverage of various smooth phase distributions and light intensity distributions by performing effective area masking (retaining the circular pupil area); real data provides real system errors (such as PZT phase shift nonlinearity, detector response nonlinearity, and residual environmental vibration).
[0024] For the four-step phase solution algorithm:
[0025] When M equals 1, the four-step phase solution algorithm requires a standard phase shift of: The input is a random interferogram, which consists of four frames of interferograms. Its true reference phase It is unknown, but can be considered as 0, with a corresponding standard phase shift of 0; the reference interferogram is... , and Similar to the two-step method, simulation training samples are generated using Legendre or Zernike polynomials, and real training samples are acquired using an interferometer to form the initial interferogram training set. In this scenario, the model can only accurately predict if the phase of the object being measured is smooth and of low order. The objects in the actual acquired data must also be smooth or approximately smooth. In actual industrial scenarios, if the object being measured exhibits characteristics dominated by low-order aberrations, quasi-dynamic measurement can be achieved with a single frame input.
[0026] When M equals 2, the input in this scenario is two frames of random interferograms. and Its phase difference Unknown, the output is the standard phase shift. The four interferograms are generated; similar to the two-step method, simulation training samples are generated using Legendre or Zernike polynomials, and real training samples are acquired using an interferometer (the interferometer must be equipped with a phase shifter with arbitrarily adjustable phase shift), together forming the initial interferogram training set; in this scenario, the phase difference between the two random interferograms input cannot be close to 0 or π; the training set must cover all possible reference phases. The model can be combined with a phase difference estimation module (a small fully connected network that outputs the predicted phase difference between two frames of random interferograms) that can be connected in parallel before the model output layer. This module is not directly used for output but can serve as an auxiliary task to help the network understand the relationships between input frames, thereby improving convergence speed. The loss function can also be the mean squared error loss function, based on the true phase difference between two random interferograms. and the predicted phase difference .
[0027] When M equals 3, the input in this scenario is 3 frames of random interferograms, each with a completely unknown reference phase, and the output is the standard phase shift. The four interferograms are used; similar to the two-step method, simulation training samples are generated using Legendre or Zernike polynomials, and real training samples are generated by acquiring data using an interferometer (the interferometer must be equipped with a phase shifter with arbitrarily adjustable phase shift), which together form the initial interferogram training set; in this scenario, three interferograms with different phases are sufficient to uniquely determine the phase distribution. Background light intensity Harmony system Therefore, the model can reconstruct the interferogram of any target phase shift very accurately.
[0028] S2: Construct a deep learning model as a virtual phase shifter, use M frames of random interferograms as input to the virtual phase shifter, and use the virtual phase shifter to predict N frames of reference interferograms with standard phase delay.
[0029] The deep learning model described in this embodiment may employ a deep learning model including CNN, GAN, or Transformer network architectures.
[0030] The deep learning model (virtual phase shifter) constructed in this invention can be implemented using various mainstream network architectures, and this embodiment does not impose any specific limitations. For example, the deep learning model can employ U-Net or a variant based on a convolutional neural network (CNN), achieving end-to-end interferogram generation through an encoder-decoder structure and skip connections; it can also employ a generative adversarial network (GAN), utilizing adversarial training between the generator and discriminator to improve the fringe sharpness and realism of the output interferogram; or it can employ a model based on a Transformer architecture (such as Swin-UNet), utilizing a self-attention mechanism to capture long-range fringe correlations in the interferogram. These are merely some feasible implementations of this invention; any deep learning model capable of mapping from M frames of random interferograms to N frames of standard phase-delay reference interferograms falls within the protection scope of this invention.
[0031] Furthermore, this invention provides a deep learning model specifically designed for virtual phase-shifting tasks in interferograms. This model employs a parallel multi-branch downsampling structure, capturing globally connected domain features of the interferogram through a large receptive field spatial feature extraction branch, performing partitioned texture modeling based on fringe direction and grayscale distribution through an adaptive fringe segmentation branch, simulating the random process of brightness gradation in the interferogram to extract gradation-related features through a Markov gradient learning branch, and learning phase representation in Fourier space through a frequency domain feature extraction branch. Subsequently, a feature alignment and fusion layer performs channel alignment and adaptive fusion of the above multi-source features to obtain globally effective features. In the decoding stage, the model further employs a parallel multi-branch upsampling architecture, including a large receptive field spatial reconstruction branch, a Markov gradient reconstruction branch, and a frequency domain mapping branch, reconstructing the initial virtual phase-shifting interferogram sequence from the spatial domain, probability domain, and frequency domain, respectively. Finally, a mask output layer defines the effective measurement area, outputting a reference interferogram with N frames of standard phase delay. This model, by introducing interferogram physical priors (fringe continuity, gradient consistency, and frequency domain phase characteristics) as inductive biases, can achieve higher generation accuracy and generalization ability with fewer training samples, as detailed below:
[0032] Preferably, the deep learning model includes: a parallel downsampling stage, a feature alignment and fusion layer, a parallel upsampling stage, and a mask output layer;
[0033] The parallel downsampling stage includes a large receptive field spatial feature extraction branch, an image feature downsampling branch, an adaptive stripe segmentation branch, a Markov gradient learning branch, and a frequency domain feature extraction branch, which are used to extract global spatial features, contextual semantic features, texture features, gradient correlation features, and phase representation features, respectively.
[0034] The feature alignment and fusion layer is used to align and fuse the global spatial features, contextual semantic features, texture features, gradient association features and phase representation features extracted in the parallel downsampling stage to obtain global effective features;
[0035] The parallel upsampling stage is used to reconstruct the initial virtual phase-shifting interferogram sequence based on the global effective features;
[0036] The mask output layer is used to mask the effective measurement region of the initial virtual phase-shifted interferogram sequence according to a predetermined mask of the effective measurement region, and output a reference interferogram with N frames of standard phase delay.
[0037] This embodiment details the implementation of each branch in the parallel downsampling stage. Let the input be an M-frame sequence of random interferograms. The downsampling step size is The spatial size of the feature map after downsampling is , The number of output feature channels for each branch is denoted as follows: ;
[0038] In the parallel downsampling stage, the input M-frame random interferogram sequence is... Each branch of the parallel downsampling stage is processed separately; and These represent the height and width of the interferogram, respectively.
[0039] In the large receptive field spatial feature extraction branch, large-size convolution, batch normalization, and activation functions are sequentially applied to the input feature map. Processing yields global spatial features ;
[0040] The input will be in this branch of the implementation. The data is sequentially fed into a large-size convolutional layer, a batch normalization layer, and an activation function. Specifically, a convolutional kernel with a size of [size missing] is first used. (For example A convolutional layer with a stride of ) The number of output channels is After batch normalization and ReLU activation, intermediate features are obtained; these intermediate features are then subjected to 3×3 convolution with a stride of 1, resulting in an output channel count of [missing value]. Finally, global spatial features are obtained. ;
[0041] In the image feature downsampling branch, the input feature map is first processed... Feature maps are obtained by downsampling. Then, the feature map is processed by convolution. Dimensionality reduction is performed to obtain features , , Where S represents a sampling step size greater than or equal to 2; and Indicates the height and width of the corresponding feature;
[0042] In this branch of the implementation, the input is first... Spatial downsampling can be performed using a step size of... Max pooling or average pooling is used to obtain the feature map. Then use one or more convolutional layers (e.g., 3×3 kernel, stride 1, output channels). For feature maps Dimensionality reduction and feature extraction are performed to obtain contextual semantic features. This branch directly preserves the overall outline and low-frequency information of the interferogram.
[0043] In the adaptive fringe segmentation branch, local fringe orientation is detected on the input interferogram based on image gradient or structure tensor. Combining the orientation detection information, a region growing method guided by SLIC superpixel segmentation or grayscale thresholding is used to divide the input interferogram into K spatially continuous and oriented regions. Each region undergoes dimensionality reduction through a convolutional module to extract local texture features. The convolutional outputs of all regions are then concatenated according to their spatial positions to form a complete texture feature. ;
[0044] In this branch example, local fringe direction detection is performed on the input interferogram based on image gradient or structural tensor. For example, the input feature map... This can be viewed as an M-frame interferogram, and the Sobel gradient of each frame of the interferogram is calculated. and ;
[0045] Sobel gradient for all M-frame interferograms and The fusion gradient is obtained by performing a sum of squares fusion. and ;
[0046]
[0047]
[0048] Construct the structure tensor:
[0049]
[0050] in, This represents a Gaussian weighted average, used to suppress noise and enhance directional consistency.
[0051] Calculate the principal direction angle:
[0052]
[0053] in, ; ; ; Indicates the pixel index;
[0054] The SLIC (Simple Linear Iterative Clustering) superpixel segmentation algorithm is adopted, taking into account both spatial location and directional similarity.
[0055] The feature vector of each pixel in the image is defined as:
[0056] in, Indicates the first The coordinates of one pixel;
[0057] Initialize cluster centers: Grid the image and set the grid spacing. ,in, This represents the desired number of superpixels; a cluster center is initialized at the center of each grid cell. To avoid falling on the strong gradient boundary, the center is perturbed to the position with the minimum gradient in the 3×3 neighborhood;
[0058] Iterative allocation and update:
[0059] Repeat the following process 10 times: for each cluster center In its Pixels in the neighborhood Calculate the distance:
[0060]
[0061] in:
[0062]
[0063]
[0064] It is the directional similarity weight (usually taken as 0.3 to 0.5);
[0065] Pixels Assigned to make The smallest cluster center;
[0066] For each cluster Recalculate the cluster centers:
[0067]
[0068]
[0069]
[0070]
[0071] Then normalize:
[0072] First, the original interferogram Feature map obtained by downsampling ;
[0073] For each region Constructing a mask ,if Belongs to the region ,but ;otherwise ;
[0074] feature map Multiplying them together yields the M-channel patch for that region. ;
[0075] This block Input a lightweight convolutional network with shared weights. The network structure consists of two 3×3 convolutions, followed by BatchNorm and ReLU after each convolution. Output a feature map. ;
[0076] Then all areas By adding elements one by one, we obtain the feature map. ;
[0077] Use 1×1 convolution to transform the feature map The number of channels from Transform to To obtain texture features .
[0078] In the Markov gradual learning branch, shallow fully convolutional layers are first used to process the input interferogram. Dimensionality reduction is performed; then, a Markov gradient learning mechanism inspired by Stable Diffusion is introduced. By constructing a multi-scale forward noise addition process, image information is gradually abstracted at different noise levels. A reverse recursive structure is used to recover and enhance the local smoothness and global gradient consistency between adjacent pixels from coarse to fine, and spatial gradient correlation features with Markov chain characteristics are extracted. ;
[0079] This example borrows the idea of the diffusion model, gradually transforming the input interferogram into a noise distribution through a forward noise addition process, and then learning a reverse denoising process to extract the gradual correlation features between pixels in the interferogram. The specific implementation is as follows:
[0080] To reduce computational complexity, a shallow fully convolutional network (composed of Conv, BatchNorm, and ReLU activation functions) is first used to process the input. Mapped to a low-dimensional latent space and downsampled to Size yields potential features ;
[0081] Define the total number of time steps T (usually 5~10), and implement noise scheduling. Linear scheduling is typically used. ,in, , ; Indicates the current time step;
[0082] definition , Gaussian noise is gradually added during the forward process, and the latent features after adding noise are calculated. :
[0083] ,
[0084] in, Represents a standard normal distribution;
[0085] An inverse denoising network is constructed, which adopts a lightweight UNet structure, including an encoder, a bottleneck layer and a decoder cascaded in sequence; the encoder includes multiple encoding blocks and the decoder includes multiple decoding blocks; the output feature residual of each encoding block is connected to the corresponding decoding block;
[0086] The latent features after adding noise Input the reverse denoising network for forward propagation, and set the time step... Sine position encoding:
[0087]
[0088]
[0089] Where i=0,…, The length is obtained as Embedded vector;
[0090] The length is The embedding vector is mapped to two MLP layers respectively. and ;according to and Adjusting the output of each convolutional layer in the inverse denoising network is expressed as:
[0091]
[0092] in, This represents the output feature map of the convolutional layer;
[0093] Latent features are processed using an inverse denoising network. The noise prediction results are obtained by performing forward propagation. Based on noise prediction results and actual noise The inverse denoising network is optimized by constructing a loss function, which is:
[0094]
[0095] in, Represents the square of the L2 norm;
[0096] After training is complete, fix the network parameters;
[0097] Select a fixed time step ( , Sample a noise ,calculate:
[0098]
[0099] Will and time step The embedding vector after sinusoidal position encoding is input into the inverse denoising network. The feature map output from the last decoder block is taken, and the output feature map is convolved with a 1×1 convolution to obtain the spatial gradient correlation feature. .
[0100] In the frequency domain feature extraction branch, the input feature map is first processed. Perform a Fourier transform and extract the amplitude from the complex spectrum after the Fourier transform. and phase ; regarding amplitude and phase Convolutional processing is performed separately to reduce dimensionality and extract amplitude and phase features; the amplitude and phase features are then concatenated to obtain the phase representation feature. ; This indicates the number of channels for the corresponding feature.
[0101] In this branch example, the frequency domain characteristics of the interferogram are modeled. The spatial interferogram is transformed to the frequency domain using Fourier transform, and features are extracted from the amplitude spectrum and phase spectrum respectively, thereby capturing the frequency distribution and phase structure information of the interferogram and enhancing the model's ability to perceive fringe period. The specific process includes the following:
[0102] For input Each frame of the interferogram is subjected to a Fast Fourier Transform to obtain a complex spectrum:
[0103]
[0104] in, express The Frame interferogram;
[0105] For each complex spectrum Extracting the amplitude spectrum and phase spectrum Stack the amplitude spectra of all M frames along the channel dimension to obtain the amplitude. and phase ;
[0106] Amplitude is processed separately using convolutional layers. and phase Amplitude characteristics were obtained and phase characteristics The phase characterization feature is obtained by concatenating the amplitude and phase features along the channel dimension. To enable full interaction between amplitude and phase information, a 3×3 convolutional layer can be added after splicing.
[0107] Preferably, in the feature alignment and fusion layer, the global spatial features are first... Contextual semantic features Texture features Gradual correlation features and phase characterization features Align and stitch along the channel dimension to obtain features. Then, convolution and fully connected operations are used to process the features. Perform feature cross-interaction to obtain globally effective features. ; ; Representation of features The number of channels; Representation of features The number of channels.
[0108] In this embodiment, the spliced features are first... By using multiple cascaded 1×1 convolutional layers, each pixel location is independently linearly reorganized along the channel dimension to obtain a feature map after convolutional mixing. ; for feature maps Perform global average pooling to obtain channel descriptors. , channel descriptor Input two consecutive fully connected layers, perform dimensionality reduction and then dimensionality increase, and finally apply sigmoid activation to obtain the channel weight vector. ;Will Multiply by each channel To obtain globally effective features .
[0109] The parallel upsampling stage includes: a large receptive field space reconstruction branch, a Markov gradient reconstruction branch, and a frequency domain mapping branch.
[0110] In the large receptive field space reconstruction branch, features are first processed through convolution. Features are obtained by mapping to N output channels Then, upsampling is performed using deconvolution with a large receptive field to obtain features. ; Features Spatial reconstruction features are obtained through attention processing. , ; This indicates element-wise multiplication; Indicates the activation function; Indicates attention score;
[0111] In this embodiment, a 3×3 convolutional layer is first used to... The number of channels is mapped to the number of target frames N to obtain the features. Using a large-size deconvolution layer (transposed convolution) to reduce the spatial size from Upsampling to ,right A spatial attention mechanism is applied to adaptively highlight effective stripe regions. An attention score is generated through a 1×1 convolution plus a sigmoid function. The attention score is then multiplied element-wise with the original features to obtain the spatially reconstructed features. ;
[0112] In the Markov gradient reconstruction branch, the Markov denoising idea based on Stable Diffusion is used for globally effective features. Perform multi-scale denoising state transition probability modeling to obtain the modeled features. Then, the scale is expanded by upsampling layer by layer through multi-scale deconvolution to restore the original image size. Gradual correlation reconstruction features are obtained. ;
[0113] In this embodiment, this branch draws on the Markov denoising idea in Stable Diffusion, and reconstructs an interferogram sequence with continuous gradient characteristics from the global effective features through multi-scale denoising state transition probability modeling, which specifically includes:
[0114] according to Construct conditional features for three resolutions:
[0115]
[0116]
[0117]
[0118] Randomly sample an initial noise tensor from a standard normal distribution:
[0119]
[0120] in, This indicates the total number of denoising steps;
[0121] Pre-compute a noise scheduling sequence (This process is the same as the aforementioned forward noise addition process)
[0122] Define calculation , ;
[0123] For the current time step t, perform the following sub-steps:
[0124] Current state Current time step The multi-scale conditional features are input into the conditional denoising network to predict the noise added in this step. The conditional denoising network employs a lightweight UNet; the encoder downsampling path corresponds to the scale of the conditional features, and the conditional features at the corresponding scale are injected into each residual block through concatenation; time step t is mapped to FiLM parameters through sinusoidal position encoding and MLP. and Adjust the output of each convolutional layer; calculate the state of the previous time step. The update formula using DDIM (Denoising Diffusion Implicit Models) is as follows:
[0125]
[0126] A preset upsampling timetable is defined, for example, after specific steps such as t=T, T / 2, T / 4, etc. Perform spatial upsampling; simultaneously switch the corresponding scale in the conditional features to a higher resolution (e.g., the scale previously used). Now changed to In this way, the denoising process is combined with the gradual improvement of spatial resolution. First, the global structure is restored at a large scale (low resolution), and then details are added at a small scale (high resolution). After T iterations and multiple upsampling steps, the final state is obtained. This tensor has been completely recovered from the initial noise to a latent representation with a clear structure, and its gradual consistency is guaranteed by the probabilistic model of the Markov chain.
[0127] Use a 3×3 convolutional layer to The number of channels from Converted to the target phase-shifted frame number N, the gradual correlation reconstruction features are obtained. ;
[0128] In the frequency domain mapping branch, features are first processed through convolution. Features are obtained through processing Then, for the features Perform a Fourier transform and extract the amplitude from the complex spectrum after the Fourier transform. and phase ; Amplitude was adjusted by deconvolution respectively and phase The amplitude is obtained through processing. and phase According to amplitude and phase Perform an inverse Fourier transform to obtain the frequency domain reconstructed features. ;
[0129] In this embodiment, the features are first processed by 3×3 convolution. Features are obtained by mapping to N channels ;right Each frame (N frames in total) undergoes a 2D Fourier transform to obtain the complex spectrum:
[0130]
[0131] Extracting the amplitude spectrum and phase spectrum The amplitude is obtained by stacking the amplitude and phase of all frames separately. and phase Using deconvolutional networks to analyze the amplitude and phase Perform upsampling to make its size from Expand to get and phase ; the amplitude of each frame and phase Combining into a complex spectrum ; for complex spectrum Performing an inverse Fourier transform yields the spatial domain features of the frame; concatenating the spatial domain features of all frames along the channel dimension yields the reconstructed frequency domain features. .
[0132] Reconstructing spatial features Gradual association reconstruction features and frequency domain reconstruction features The sequences are added together and then convolved to obtain the initial virtual phase-shifting interferogram sequence.
[0133] Through multiple branches in the parallel downsampling stage, including large receptive field spatial feature extraction, adaptive fringe segmentation, Markov gradient learning, and frequency domain feature extraction, the model captures the global connected components, directional texture, gradient consistency, and frequency domain phase structure of the interferogram, thereby achieving deep decoupling and representation of the multi-dimensional physical properties of the interferogram. The feature alignment and fusion layer fully cross-integrates heterogeneous features through convolution and fully connected operations to generate information-rich global effective features. In the parallel upsampling stage, the three branches of spatial reconstruction, Markov gradient reconstruction, and frequency domain mapping work together to reconstruct the interferogram from the spatial, probabilistic, and frequency domains, respectively, ensuring high consistency in fringe continuity, gradient smoothness, and spectral fidelity of the virtual phase-shifting interferogram. By introducing the inherent physical prior of the interferogram as an inductive bias, this model can achieve higher generation accuracy and stronger generalization ability with fewer training samples compared to general network architectures, and is particularly suitable for complex scenarios such as sparse fringes, high noise interference, or dynamic measurements.
[0134] S3: Construct a loss function for training the virtual phase shifter, and train the virtual phase shifter based on the constructed loss function. The loss function includes at least one of spatial domain basic loss, high-frequency region enhancement loss, frequency domain basic loss, and frequency domain high-frequency loss.
[0135] The airspace fundamental loss includes:
[0136]
[0137] in, Indicates basic airspace loss; Indicates the first Frame prediction virtual phase shift map; No. Frame reference interferogram labels; This represents the reconstruction error metric function, which includes one or more combinations of L1 loss, L2 loss, Charbonnier loss, and Wasserstein distance; Indicates the pixel value index;
[0138] The high-frequency enhancement loss includes:
[0139]
[0140]
[0141]
[0142]
[0143]
[0144] in, This indicates the enhancement loss in the high-frequency region; Represents the error map of the k-th frame. exist Error value at; Indicates a high-pass filter; Indicates the first Frame error map The mean; Represents the error map of the k-th frame. Standard deviation; and express and High-frequency components; Represents a small constant;
[0145] The fundamental frequency domain loss includes:
[0146]
[0147]
[0148]
[0149] in, Represents the fundamental loss in the frequency domain; Indicates Fourier transform; and express and Frequency domain characteristics;
[0150] The frequency domain high-frequency loss includes:
[0151]
[0152]
[0153]
[0154] in, Indicates high-frequency loss in the frequency domain; and They represent and The high-frequency components. In this embodiment, it is aimed at... and Common approaches include: using the difference in complex modulus, using the Euclidean distance of complex vectors, and calculating the loss of amplitude and phase separately.
[0155] The total loss function used when training the virtual phase shifter is as follows:
[0156]
[0157] in, Represents the total loss function; , , and This represents the weighting parameter.
[0158] In this embodiment, a spatial domain fundamental loss (such as L1, L2, or Charbonnier loss) ensures that the predicted interferogram matches the true label in overall grayscale distribution. A high-frequency region enhancement loss, based on the Laida criterion, identifies and weights high-frequency anomalous regions in the prediction error, forcing the model to focus on spatial high-frequency components such as fringe edges and texture details, significantly improving the clarity and edge sharpness of the output interferogram. The frequency domain fundamental loss and the frequency domain high-frequency loss jointly constrain the amplitude and phase distribution of the prediction results in the Fourier spectrum, particularly strengthening the supervision of high-frequency spectral components and effectively suppressing the blurring and artifact phenomena easily generated by traditional spatial domain losses. This four-component weighted loss function finely constrains the model generation process from multiple dimensions, including spatial and frequency domains, and global and local dimensions, ensuring that the standard phase-shifted interferogram output by the virtual phase shifter meets metrological requirements in terms of visual quality, spectral fidelity, and subsequent phase calculation accuracy. It is particularly suitable for complex measurement environments with noise interference or low fringe contrast.
[0159] S4: The M-frame random interferograms of the object under test are acquired by the interferometer and input into the trained virtual phase shifter for inference to predict the N-frame reference interferograms with standard phase delay.
[0160] In this embodiment, an interferometer is first used to expose the object under test once, acquiring M frames of random interferograms. These frames are normalized to the [0,1] interval and multiplied by a binary mask of the effective measurement area before being input into a trained virtual phase shifter. The model forward inference directly outputs four frames of reference interferograms with standard phase delay. Then, the pixel values of each output frame are cropped to a reasonable range and the mask is applied again, resulting in a complete interferogram sequence that can be directly used for phase resolution using the traditional four-step phase shift algorithm. This embodiment achieves real-time conversion from a single exposure to high-precision four-step phase shift data, eliminating the need for a physical phase shifter and its precise control.
[0161] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0162] In summary, this invention constructs a deep learning model as a virtual phase shifter. It can generate a set of reference interferogram sequences with standard phase shifts (such as the two-step method {0,π} or the four-step method {0,π / 2,π,3π / 2}) in real time with only a single or a small number (M≥1) reference interferograms of randomly unknown phase. This completely eliminates the dependence on physical phase shifters such as piezoelectric ceramics, eliminates inter-frame phase errors introduced by environmental vibrations and airflow disturbances, and overcomes the stringent requirements of the space carrier method on fringe direction and frequency. Compared to existing deep learning methods that directly predict phase end-to-end from a single frame interferogram, this invention retains the analytical solution advantages of mature multi-step phase shifting algorithms, significantly improving the accuracy and generalization ability of phase recovery. Furthermore, the parallel multi-branch deep learning model provided by this invention further enhances the model's ability to model complex interference fringes by fusing spatial, frequency, and Markov gradient features. The high-frequency enhancement and frequency domain loss in the composite loss function effectively improve the edge sharpness and spectral fidelity of the predicted interferogram. In summary, this invention achieves low-cost, highly robust, and near-real-time virtual phase-shifting measurement, which can be widely applied to scenarios such as mirror inspection, micro / nano structure characterization, and dynamic process observation.
[0163] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A virtual phase-shifting method for interferograms based on artificial intelligence, characterized in that, include: S1: Obtain the initial interferogram training set, which includes training samples obtained by observing and collecting data from the object under test. Each training sample contains M frames of random interferograms and N frames of reference interferogram labels with standard phase delay collected under the phase distribution to be tested. The M-frame random interferogram includes: an M-frame interferogram with unknown reference phase and unknown inter-frame phase difference; the N-frame standard phase delay reference interferogram includes: N equally spaced standard phase delay reference interferograms required by the selected phase resolution algorithm; S2: Construct a deep learning model as a virtual phase shifter, use M frames of random interferograms as input to the virtual phase shifter, and use the virtual phase shifter to predict N frames of reference interferograms with standard phase delay. The deep learning model includes: a parallel downsampling stage, a feature alignment and fusion layer, a parallel upsampling stage, and a mask output layer; The parallel downsampling stage includes a large receptive field spatial feature extraction branch, an image feature downsampling branch, an adaptive stripe segmentation branch, a Markov gradient learning branch, and a frequency domain feature extraction branch, which are used to extract global spatial features, contextual semantic features, texture features, gradient correlation features, and phase representation features, respectively. The feature alignment and fusion layer is used to align and fuse the global spatial features, contextual semantic features, texture features, gradient association features and phase representation features extracted in the parallel downsampling stage to obtain global effective features; The parallel upsampling stage is used to reconstruct the initial virtual phase-shifting interferogram sequence based on the global effective features; The mask output layer is used to mask the effective measurement region of the initial virtual phase-shifting interferogram sequence according to a predetermined mask of the effective measurement region, and output a reference interferogram with N frames of standard phase delay; S3: Construct a loss function for training the virtual phase shifter, and train the virtual phase shifter based on the constructed loss function. The loss function includes at least one of spatial domain basic loss, high-frequency region enhancement loss, frequency domain basic loss, and frequency domain high-frequency loss. S4: The M-frame random interferograms of the object under test are acquired by the interferometer and input into the trained virtual phase shifter for inference to predict the N-frame reference interferograms with standard phase delay.
2. The artificial intelligence-based virtual phase-shifting method for interferograms according to claim 1, characterized in that, In the parallel downsampling stage, the input M-frame random interferogram sequence is... Each branch of the parallel downsampling stage is processed separately; and These represent the height and width of the interferogram, respectively. In the large receptive field spatial feature extraction branch, large-size convolution, batch normalization, and activation functions are sequentially applied to the input feature map. Processing yields global spatial features ; In the image feature downsampling branch, the input feature map is first processed... Feature maps are obtained by downsampling. Then, the feature map is processed by convolution. Dimensionality reduction is performed to obtain features , , Where S represents a sampling step size greater than or equal to 2; and Indicates the height and width of the corresponding feature; In the adaptive fringe segmentation branch, local fringe orientation is detected on the input interferogram based on image gradient or structure tensor. Combining the orientation detection information, a region growing method guided by SLIC superpixel segmentation or grayscale thresholding is used to divide the input interferogram into K spatially continuous and oriented regions. Each region undergoes dimensionality reduction through a convolutional module to extract local texture features. The convolutional outputs of all regions are then concatenated according to their spatial positions to form a complete texture feature. ; In the Markov gradual learning branch, shallow fully convolutional layers are first used to process the input interferogram. Dimensionality reduction is performed; then, a Markov gradient learning mechanism inspired by Stable Diffusion is introduced. By constructing a multi-scale forward noise addition process, image information is gradually abstracted at different noise levels. A reverse recursive structure is used to recover and enhance the local smoothness and global gradient consistency between adjacent pixels from coarse to fine, and spatial gradient correlation features with Markov chain characteristics are extracted. ; In the frequency domain feature extraction branch, the input feature map is first processed. Perform a Fourier transform and extract the amplitude from the complex spectrum after the Fourier transform. and phase ; regarding amplitude and phase Convolutional processing is performed separately to reduce dimensionality and extract amplitude and phase features; the amplitude and phase features are then concatenated to obtain the phase representation feature. ; This indicates the number of channels for the corresponding feature.
3. The artificial intelligence-based virtual phase-shifting method for interferograms according to claim 2, characterized in that, In the feature alignment and fusion layer, the global spatial features are first processed. Contextual semantic features Texture features Gradual correlation features and phase characterization features Align and stitch along the channel dimension to obtain features. Then, convolution and fully connected operations are used to process the features. Perform feature cross-interaction to obtain globally effective features. ; ; Representation of features The number of channels; Representation of features The number of channels.
4. The artificial intelligence-based virtual phase-shifting method for interferograms according to claim 3, characterized in that, The parallel upsampling stage includes: a large receptive field space reconstruction branch, a Markov gradient reconstruction branch, and a frequency domain mapping branch. In the large receptive field space reconstruction branch, features are first processed through convolution. Features are obtained by mapping to N output channels Then, upsampling is performed using deconvolution with a large receptive field to obtain features. ; Features Spatial reconstruction features are obtained through attention processing. , ; This indicates element-wise multiplication; Indicates the activation function; Indicates attention score; In the Markov gradient reconstruction branch, the Markov denoising idea based on Stable Diffusion is used for globally effective features. Perform multi-scale denoising state transition probability modeling to obtain the modeled features. Then, the scale is expanded by upsampling layer by layer through multi-scale deconvolution to restore the original image size. Gradual correlation reconstruction features are obtained. ; In the frequency domain mapping branch, features are first processed through convolution. Features are obtained through processing Then, for the features Perform a Fourier transform and extract the amplitude from the complex spectrum after the Fourier transform. and phase ; Amplitude was adjusted by deconvolution respectively and phase The amplitude is obtained through processing. and phase According to amplitude and phase Perform an inverse Fourier transform to obtain the frequency domain reconstructed features. ; Reconstructing spatial features Gradual association reconstruction features and frequency domain reconstruction features The sequences are then added together and fused, followed by convolution to obtain the initial virtual phase-shifting interferogram sequence.
5. The artificial intelligence-based virtual phase-shifting method for interferograms according to claim 4, characterized in that, The airspace fundamental loss includes: in, Indicates basic airspace loss; Indicates the first Frame prediction virtual phase shift map; No. Frame reference interferogram labels; This represents the reconstruction error metric function, which includes one or more combinations of L1 loss, L2 loss, Charbonnier loss, and Wasserstein distance; Indicates the pixel value index; The high-frequency enhancement loss includes: in, This indicates the enhancement loss in the high-frequency region; Represents the error map of the k-th frame. exist Error value at; Indicates a high-pass filter; Indicates the first Frame error map The mean; Represents the error map of the k-th frame. Standard deviation; and express and High-frequency components; Represents a small constant; The fundamental frequency domain loss includes: in, Represents the fundamental loss in the frequency domain; Indicates Fourier transform; and express and Frequency domain characteristics; The frequency domain high-frequency loss includes: in, Indicates high-frequency loss in the frequency domain; and They represent and The high-frequency components.
6. The artificial intelligence-based virtual phase-shifting method for interferograms according to claim 5, characterized in that, The total loss function used when training the virtual phase shifter is as follows: in, Represents the total loss function; , , and This represents the weighting parameter.
7. An interferogram virtual phase-shifting system based on artificial intelligence, characterized in that, The system includes a memory and a processor; the memory is used to store an application program; the processor is used to run the application program and execute an artificial intelligence-based interferogram virtual phase shifting method as described in any one of claims 1 to 6.
8. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which, when executed by a processor, implements an artificial intelligence-based virtual phase-shifting method for interferograms as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Surveying and mapping geographic system based on remote sensing technology
CN121323596A
Measurement and imaging instruments and beamforming method
US20190129026A1