An interference phase disentanglement method based on cross-attention mechanism fusion network

By constructing an interference phase unwrap method based on the cross attention mechanism fusion network, the problem of insufficient unwrap accuracy of interference synthetic aperture radar under low signal-to-noise ratio conditions is solved, and high-precision and rapid convergence interference phase unwrap effect is achieved.

CN120181136BActive Publication Date: 2025-08-19NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510622248.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-19
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The existing interference synthetic aperture radar technology has insufficient phase disintegration accuracy under low signal-to-noise ratio conditions, and deep learning methods are incomplete in multi-scale feature learning, high feature redundancy, and strong data set dependence, resulting in limited effectiveness in low coherence and high gradient deformation areas.

Method used

Using a method based on the cross attention mechanism fusion network, a multi-scale pooled cross attention mechanism fusion network is constructed, including input layer, convolutional layer, maximum pooling layer, depth separable convolutional unit, residual unit, channel attention module and cross attention fusion module, training and verification of interference phase data sets are carried out to improve the detangling accuracy.

Benefits of technology

It realizes high-precision, fast convergence and strong robust interference phase unwrap under low signal-to-noise ratio conditions, improving the accuracy and real-time performance of synthetic aperture radar interference measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181136B_ABST
    Figure CN120181136B_ABST
Patent Text Reader

Abstract

The present application relates to an interferometric phase unwrapping method based on a cross-attention mechanism fusion network. The method comprises: forming an interferometric phase dataset from each interferometric entangled phase image dataset and the corresponding real interferometric phase image; constructing a cross-attention mechanism fusion network structure based on multi-scale pooling; dividing the interferometric phase dataset into a training dataset, a validation dataset, and a test dataset, and inputting the training dataset into the cross-attention mechanism fusion network structure for training, thereby obtaining a cross-attention mechanism fusion network that takes the interferometric entangled phase image dataset as input and outputs the corresponding interferometric phase unwrapped image dataset. This method can achieve interferometric phase unwrapping under low signal-to-noise ratio conditions and obtain highly accurate synthetic aperture radar interferometric unwrapping results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of synthetic aperture radar interferometry technology, and in particular to an interferometric phase unwrapping method based on a cross-attention mechanism fusion network. Background Art

[0002] Interferometric Synthetic Aperture Radar (InSAR) is a remote sensing technology that combines electromagnetic interferometry with synthetic aperture radar (SAR). Phase unwrapping is a key technique for measuring elevation or deformation in InSAR data processing. However, practical challenges such as decorrelation noise, terrain variability, and algorithmic limitations can significantly reduce phase unwrapping accuracy.

[0003] Since interferometric phase unwrapping requires timeliness and high precision in its application, exploring and developing a new phase unwrapping method is of great significance. In recent years, deep learning has become a very promising phase unwrapping method, which has improved the efficiency and accuracy of phase unwrapping to a certain extent. However, it still faces problems such as incomplete multi-scale feature learning, high feature redundancy, and strong dataset dependence. These problems undermine their effectiveness in areas with low coherence and high gradient deformation. Therefore, for interferometric phase unwrapping using neural networks, obtaining an interferometric phase unwrapping method with high timeliness, good robustness, and high unwrapping accuracy is exactly the problem that needs to be solved. Summary of the Invention

[0004] Based on this, it is necessary to provide an interferometric phase unwrapping method based on a cross-attention mechanism fusion network, which can realize interferometric phase unwrapping under low signal-to-noise ratio conditions and improve the accuracy of synthetic aperture radar interferometry measurement, in order to address the above technical problems.

[0005] The present application provides an interference phase unwrapping method based on a cross-attention mechanism fusion network, which is used to execute steps S1 to S3 to obtain a cross-attention mechanism fusion network for interference phase unwrapping, including:

[0006] Step S1, generating an interference entanglement phase image dataset based on a Gaussian function according to a real interference phase image, wherein each interference entanglement phase image dataset and the corresponding real interference phase image constitute an interference phase dataset;

[0007] Step S2, constructing a cross-attention mechanism fusion network structure based on multi-scale pooling, including an input layer, a first convolutional layer, a maximum pooling layer, a first depth-separable convolutional unit, a second depth-separable convolutional unit, a third depth-separable convolutional unit, a first batch of normalization and linear activation combination layers, a channel attention module, a first residual unit, a second residual unit, a third residual unit, a fourth residual unit, a fusion module based on a cross-attention mechanism, a second convolutional layer, a third convolutional layer, a second batch of normalization and linear activation combination layers, a fourth convolutional layer, a third batch of normalization and linear activation combination layers, an upsampling layer, a fifth convolutional layer, and an output layer;

[0008] Step S3, divide the interference phase dataset into a training dataset, a verification dataset and a test dataset, and input the training dataset into the cross-attention mechanism fusion network structure for training, and iterate by comparing the interference phase unwrapping images corresponding to each interference entangled phase map dataset output by the cross-attention mechanism fusion network structure with the real interference phase images corresponding to each interference entangled phase map dataset in the interference phase dataset, to obtain a cross-attention mechanism fusion network with the interference entangled phase map dataset as input and the corresponding interference phase unwrapping map dataset as output; and input the verification dataset into the cross-attention mechanism fusion network for verification, and input the test dataset into the cross-attention mechanism fusion network to obtain the predicted interference phase unwrapping result.

[0009] In one embodiment, in step S2, the cross-attention mechanism fusion network structure based on multi-scale pooling is connected to the first convolution layer and the maximum pooling layer in sequence from input to output, and then divided into a first branch and a second branch, wherein the first branch is connected to the first residual unit, the second residual unit, the third residual unit, and the fourth residual unit in sequence from input to output, and the output end of the fourth residual unit is connected to the first input end of the fusion module based on the cross-attention mechanism; the second branch is connected to the first depth-separable convolution unit, the second depth-separable convolution unit, the third depth-separable convolution unit, the first batch of normalization and linear activation combination layer, and the channel attention module in sequence from input to output, and the output end of the channel attention module is connected to the second input end of the fusion module based on the cross-attention mechanism, and the fusion module based on the cross-attention mechanism is connected to the second convolution layer, the third convolution layer, the second batch of normalization and linear activation combination layer, the fourth convolution layer, the third batch of normalization and linear activation combination layer, the upsampling layer, the fifth convolution layer and the output layer in sequence from input to output.

[0010] In one embodiment, in step S2, the first residual unit includes a first residual module and a second residual module in sequence from input to output, wherein the input end of the first residual module is the input end of the first residual unit, which is connected to the output end of the maximum pooling layer, and the output end of the second residual module is the output end of the first residual unit, which is connected to the input end of the second residual unit; the second residual unit includes a third residual module, a fourth residual module, and a fifth residual module in sequence from input to output, wherein the input end of the third residual module is the input end of the second residual unit, which is connected to the output end of the first residual unit, and the output end of the fifth residual module is the output end of the second residual unit, which is connected to the input end of the third residual unit; the third residual unit includes a fourth residual module and a fifth residual module in sequence from input to output, wherein the input end of the third residual module is the input end of the second residual unit, which is connected to the output end of the first residual unit, and the output end of the fifth residual module is the output end of the second residual unit, which is connected to the input end of the third residual unit; The output direction includes the sixth residual module, the seventh residual module, the eighth residual module, the ninth residual module, and the tenth residual module in sequence, wherein the input end of the sixth residual module is the input end of the third residual unit, which is connected to the output end of the second residual unit, and the output end of the tenth residual module is the output end of the third residual unit, which is connected to the input end of the fourth residual unit; the fourth residual unit includes the eleventh residual module, the twelfth residual module, and the thirteenth residual module in sequence from input to output, wherein the input end of the eleventh residual module is the input end of the fourth residual unit, which is connected to the output end of the third residual unit, and the output end of the thirteenth residual module is the output end of the fourth residual unit, which is connected to the first input end of the fusion module based on the cross-attention mechanism.

[0011] In one embodiment, each residual module in each residual unit in step S2 includes an input layer, a first residual convolution layer, a residual batch normalization and linear activation combination layer, a second residual convolution layer, a first residual batch normalization layer, a residual addition fusion layer, a residual linear activation layer, a judgment block, a third residual convolution layer, a second residual batch normalization layer, and an output layer. The input layer serves as the input end of the residual module and is connected to the input end of the judgment block and the first residual convolution layer, respectively. The first residual convolution layer, the residual batch normalization and the linear activation layer are connected to the residual convolution layer, the residual batch normalization layer, and the linear activation layer. The linear activation combination layer, the second residual convolution layer, and the first residual batch normalization layer are connected in sequence, the judgment block, the third residual convolution layer, and the second residual batch normalization layer are connected in sequence, the output ends of the first residual batch normalization layer and the second residual batch normalization layer are connected to the input end of the residual addition fusion layer, the residual addition fusion layer, the residual linear activation layer, and the output layer of the residual module are connected in sequence, and the output layer serves as the output end of the residual module; when the input step size of each residual module in each residual unit is not 1, the residual module The input layer is connected to the third residual convolution layer and the first residual convolution layer respectively. The branch of the first residual convolution layer is connected to the residual batch normalization and linear activation combination layer, the second residual convolution layer, and the first residual batch normalization layer in sequence, as the first input end of the residual addition fusion layer. The branch of the third residual convolution layer is connected to the second residual batch normalization layer, as the second input end of the residual addition fusion layer. The output end of the residual addition fusion layer is connected to the residual linear activation layer, and the residual module result is output through the output layer of the residual module. ; When the input stride of each residual module is 1, the input layer of the residual module is respectively connected to the residual addition fusion layer as the first input end of the residual addition fusion layer and the first residual convolution layer. The branch of the first residual convolution layer is sequentially connected to the residual batch normalization and linear activation combination layer, the second residual convolution layer, and the first residual batch normalization layer as the second input end of the residual addition fusion layer. The output of the residual addition fusion layer is connected to the residual linear activation layer, and the residual module result is output through the output layer of the residual module.

[0012] In one embodiment, the first depth-wise separable convolution unit in step S2 includes a first depth-wise separable convolution layer, a second depth-wise separable convolution layer, and a third depth-wise separable convolution layer in sequence from input to output, wherein the input end of the first depth-wise separable convolution layer is the input end of the first depth-wise separable convolution unit, which is connected to the output end of the maximum pooling layer, and the output end of the third depth-wise separable convolution layer is the output end of the first depth-wise separable convolution unit, which is connected to the input end of the second depth-wise separable convolution unit.

[0013] In one embodiment, in step S2, the second depth-wise separable convolution unit includes, from input to output, a sixth convolution layer, a fourth depth-wise separable convolution layer, a fifth depth-wise separable convolution layer, and a sixth depth-wise separable convolution layer, wherein the input end of the sixth convolution layer is the input end of the second depth-wise separable convolution unit, and is connected to the output end of the first depth-wise separable convolution unit. The output end of the sixth depth-wise separable convolution layer is the output end of the second depth-wise separable convolution unit, and is connected to the input end of the third depth-wise separable convolution unit; the third depth-wise separable convolution unit includes, from input to output, a seventh convolution layer, a seventh depth-wise separable convolution layer, an eighth depth-wise separable convolution layer, and a ninth depth-wise separable convolution layer, wherein the input end of the seventh convolution layer is the input end of the third depth-wise separable convolution unit, and is connected to the output end of the second depth-wise separable convolution unit. The output end of the ninth depth-wise separable convolution layer is the output end of the third depth-wise separable convolution unit, and is connected to the input end of the first batch of normalization and linear activation combination layers.

[0014] In one embodiment, the channel attention module in step S2 includes an input layer, a global average pooling layer, a first channel attention fully connected layer, a second channel attention fully connected layer, a linear activation function layer, a third channel attention fully connected layer, a channel attention sigmoid function layer, a channel attention multiplication fusion layer, a channel attention addition fusion layer and an output layer. The first branch of the input layer of the channel attention module passes through the global average pooling layer, the first channel attention fully connected layer, the first branch of the second channel attention fully connected layer, the linear activation function layer, the third channel attention fully connected layer, the channel attention sigmoid function layer, the output end of the channel attention sigmoid function layer serves as the first input end of the channel attention multiplication fusion layer, the output end of the channel attention multiplication fusion layer serves as the first input end of the channel attention addition fusion layer, the second branch of the input layer of the channel attention module serves as the second input end of the channel attention addition fusion layer, the second branch of the second channel attention fully connected layer serves as the second input end of the channel attention multiplication fusion layer, and the output end of the channel attention addition fusion layer is connected to the output layer of the channel attention module.

[0015] In one embodiment, the fusion module based on the cross-attention mechanism in step S2 includes a first fusion input layer, a second fusion input layer, a first average pooling layer, a second average pooling layer, a third average pooling layer, a first fusion convolution layer, a second fusion convolution layer, a third fusion convolution layer, a fourth fusion convolution layer, a first fusion upsampling layer, a second fusion upsampling layer, a third fusion upsampling layer, a fourth fusion upsampling layer, a cross-attention module, a first fusion batch normalization and linear activation combination layer, a second fusion batch normalization and linear activation combination layer, a third fusion batch normalization and linear activation combination layer, a fusion splicing layer, a fusion batch normalization layer, a cross-addition fusion layer, a fusion linear activation layer and a fusion output layer, wherein the first branch of the first fusion input layer serves as the first input end of the cross-attention module, the second branch of the first fusion input layer is connected to the first fusion upsampling layer, the first fusion upsampling layer is connected to four branches respectively, and the first branch of the first fusion upsampling layer is directly used as the first input end of the fusion splicing layer; the second branch of the first fusion upsampling layer is sequentially connected to the first average pooling layer, the first fusion layer, and the second branch of the first fusion layer. The first fusion convolution layer, the first fusion batch normalization and linear activation combination layer, and the second fusion upsampling layer are used as the second input of the fusion splicing layer; the third branch of the first fusion upsampling layer is connected in sequence to the second average pooling layer, the second fusion convolution layer, the second fusion batch normalization and linear activation combination layer, and the third fusion upsampling layer as the third input of the fusion splicing layer; the fourth branch of the first fusion upsampling layer is connected in sequence to the third average pooling layer, the third fusion convolution layer, the third fusion batch normalization and linear activation combination layer, and the fourth fusion upsampling layer as the fourth input of the fusion splicing layer; the output of the fusion splicing layer is connected in sequence to the fourth fusion convolution layer and the fusion batch normalization layer as the first input of the cross-addition fusion layer; the second fusion input layer is directly used as the second input of the cross-attention module, the output of the cross-attention module is used as the second input of the cross-addition fusion layer, the output of the cross-addition fusion layer is connected to the fusion linear activation layer, the output of the fusion linear activation layer is connected to the fusion output layer, and the result of the fusion module based on the cross-attention mechanism is output through the fusion output layer.

[0016] In one embodiment, the cross-attention module in step S2 includes a first cross-fully connected layer, a second cross-fully connected layer, a third cross-fully connected layer, a first cross-convolutional layer, a second cross-convolutional layer, a third cross-convolutional layer, a fourth cross-convolutional layer, a first deformation layer, a second deformation layer, a third deformation layer, a subtraction layer, a cross-sigmoid function layer, an enhancement layer, a transposition layer, a calculation layer, a cross-splitting layer, an instance normalization layer, and a cross-upsampling layer. The first branch of the first fusion input layer serves as the first input end of the cross-attention module, and the second fusion input layer directly serves as the second input end of the cross-attention module. The first input end of the cross-attention module is divided into three branches. The first branch of the first input end of the cross-attention module is connected to the first cross-fully connected layer and the first deformation layer in sequence as the first input end of the calculation layer; the second branch of the first input end of the cross-attention module is connected to the first cross-convolutional layer and serves as the first input end of the enhancement layer; the third branch of the first input end of the cross-attention module serves as the first input end of the subtraction layer; the second input end of the cross-attention module is divided into four branches. The first branch of the second input end of the cross-attention module serves as the second input end of the subtraction layer, and the second branch of the second input end of the cross-attention module is connected to the third cross-convolution layer as the second input end of the reinforcement layer; the third branch of the second input end of the cross-attention module is connected to the second cross-fully connected layer and the second deformation layer in sequence as the second input end of the calculation layer; the fourth branch of the second input end of the cross-attention module serves as the first input end of the cross-splice layer; the output end of the subtraction layer is connected to the second cross-convolution layer and the cross-sigmoid function layer in sequence, and serves as the third input end of the reinforcement layer; the output end of the reinforcement layer is connected to the third cross-fully connected layer, the third deformation layer, and the transpose layer in sequence, and serves as the third input end of the calculation layer; the output end of the calculation layer is connected to the cross-splice layer as the second input end of the cross-splice layer; the output end of the cross-splice layer is connected to the fourth cross-convolution layer, the instance normalization layer, and the cross-upsampling layer in sequence, and the output end of the cross-upsampling layer serves as the output end of the cross-attention module, and the output end of the cross-addition fusion layer serves as the second input end of the cross-attention module.

[0017] In one embodiment, step S3 further includes:

[0018] The normalized root mean square error is used to evaluate the network phase disentanglement accuracy of the cross-attention mechanism fusion network. The calculation formula is as follows:

[0019] ;

[0020] in, is the normalized root mean square error, is a sample of the training dataset, For training data set samples The difference between the true value and the estimated value of the cross-attention mechanism fusion network training, Represents the phase pixel coordinates.

[0021] The above-mentioned interferometric phase unwrapping method based on the cross-attention mechanism fusion network creates an interferometric phase data set, adopts a reasonable residual module for network construction, designs a channel attention mechanism-based enhanced depth-separable convolution module, designs a network fusion module based on the cross-attention mechanism, and constructs, trains and verifies the cross-attention mechanism fusion network, etc., to achieve interferometric phase unwrapping under low signal-to-noise ratio conditions, improve the precision of synthetic aperture radar interferometry measurement, and obtain an interferometric phase unwrapping method with fast convergence speed, strong real-time performance, strong robustness and high precision. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 1 is a flow chart of an interference phase unwrapping method based on a cross-attention mechanism fusion network in one embodiment;

[0023] Figure 2 A schematic diagram of the structure of a cross-attention mechanism fusion network in one embodiment;

[0024] Figure 3 2 is a schematic diagram of the structure of a fusion module based on a cross-attention mechanism in one embodiment;

[0025] Figure 4 This is a loss accuracy graph of a cross-attention mechanism fusion network in one embodiment;

[0026] Figure 5 An interferometric phase unwrapping method based on a cross-attention mechanism fusion network in one embodiment is used to interferometrically entangle a phase image under noise;

[0027] Figure 6 is a true interferometric phase image corresponding to an interferometric entangled phase image obtained by an interferometric phase unwrapping method based on a cross-attention mechanism fusion network under noise in one embodiment;

[0028] Figure 7 It is an interference unwrapped phase map predicted by a cross-attention mechanism fusion network in an interference phase unwrapping method based on a cross-attention mechanism fusion network in one embodiment. DETAILED DESCRIPTION

[0029] Interferometric synthetic aperture radar (ISAR) can extract surface deformation information with millimeter-level accuracy by analyzing the phase differences between SAR images acquired over the same area at different times. This technology, with its ability to penetrate the atmosphere and acquire deformation information of monitored targets around the clock and in all weather conditions, has been successfully applied in a variety of fields, including surface deformation monitoring, landslides, volcanic activity, glacial movement, and deformation information extraction of man-made buildings. In ISAR data processing, phase unwrapping is a key technique for measuring elevation or deformation. However, practical challenges such as decorrelation noise, terrain variability, and algorithmic limitations can significantly reduce phase unwrapping accuracy. These challenges are particularly prominent in areas with low coherence and steep gradients, where traditional methods often fail, limiting the reliability of InSAR-derived measurements.

[0030] Early research has primarily addressed phase unwrapping using three traditional approaches: path tracking, optimization, and filtering-based joint denoising and unwrapping. Path tracking methods perform phase unwrapping by integrating along a selected path. However, uncertainty in path selection and noise can lead to large errors during path integration. Optimization methods achieve phase unwrapping by minimizing the difference between the phase gradient and the estimated gradient. These methods are global and typically use various objective functions to optimize the unwrapping result. They offer high robustness but high computational complexity. Filtering-based joint denoising and unwrapping methods simultaneously suppress phase noise and unwrap phase through filtering, offering strong robustness but complex implementation. Phase unwrapping in the presence of high phase noise and steep phase gradients in interferometric synthetic aperture radar (InSAR) has long been a challenge, and the least squares phase unwrapping method (LS) often suffers from large unwrapping errors in such situations. Given the timeliness and high accuracy requirements of interferometric phase unwrapping, exploring and developing novel phase unwrapping methods is of great significance.

[0031] Phase unwrapping, the process of extracting a true phase image from noisy phase measurement data, plays a crucial role in scientific imaging. In particular, in the presence of significant noise, challenging nonlinear ill-posed problems must be addressed. In recent years, deep learning has emerged as a promising approach for phase unwrapping. Inspired by the successful application of convolutional neural networks (CNNs) in image restoration, neural networks have been used for phase unwrapping. However, due to the local nature of the convolution kernel, they are less than ideal in capturing global spatial dependencies. Subsequently, many researchers have used various deep convolutional neural network (DCNN) models to improve the efficiency and accuracy of phase unwrapping. Examples include U-net (U-shaped convolutional neural network), PhaseNet (phase unwrapping neural network), PUnet (phase unwrapping neural network), and SCAPU (phase unwrapping neural network). These networks have improved the efficiency and accuracy of phase unwrapping to a certain extent, but they still have certain limitations. For example, the PhaseNet network is suitable for unwrapping in situations with little noise impact and requires large data sets for training. Unet is not suitable for phase unwrapping problems in high-precision applications. Although SCAPU has good interferometric phase unwrapping performance, the network is complex and the training time is long. PUNet suffers from poor robustness. These also represent the problems encountered by common networks in dealing with interferometric phase unwrapping. To overcome these limitations, it is of great significance to explore an interferometric phase unwrapping method with high precision, high efficiency, and strong robustness.

[0032] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0033] It is understood that the terms "first," "second," and so on, used herein may be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, a first convolutional layer may be referred to as a second convolutional layer, and similarly, a second convolutional layer may be referred to as a first convolutional layer, without departing from the scope of this application. Both the first convolutional layer and the second convolutional layer are convolutional layers, but they are not the same convolutional layer.

[0034] It can be understood that the terms used in this application such as "fused convolution layer", "residual convolution layer", "fused batch normalization and linear activation combination layer", "residual batch normalization and linear activation combination layer", etc. are only used to distinguish the same network structure in different modules.

[0035] When used herein, the singular forms "a", "an", and "the" may also include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "include / comprise" or "have" and the like specify the presence of stated features, integers, steps, operations, components, parts, or combinations thereof, but do not preclude the possibility of the presence or addition of one or more other features, integers, steps, operations, components, parts, or combinations thereof.

[0036] In an exemplary embodiment, a method for interferometric phase unwrapping based on a cross-attention mechanism fusion network is provided. The method is used to execute steps S1 to S3 to obtain a cross-attention mechanism fusion network for interferometric phase unwrapping, including:

[0037] Step S1, generating an interference entanglement phase image dataset based on a Gaussian function according to a real interference phase image, wherein each interference entanglement phase image dataset and the corresponding real interference phase image constitute an interference phase dataset;

[0038] Step S2, constructing a cross-attention mechanism fusion network structure based on multi-scale pooling, including an input layer, a first convolutional layer, a maximum pooling layer, a first depth-separable convolutional unit, a second depth-separable convolutional unit, a third depth-separable convolutional unit, a first batch of normalization and linear activation combination layers, a channel attention module, a first residual unit, a second residual unit, a third residual unit, a fourth residual unit, a fusion module based on a cross-attention mechanism, a second convolutional layer, a third convolutional layer, a second batch of normalization and linear activation combination layers, a fourth convolutional layer, a third batch of normalization and linear activation combination layers, an upsampling layer, a fifth convolutional layer, and an output layer;

[0039] Step S3, divide the interference phase dataset into a training dataset, a verification dataset and a test dataset, and input the training dataset into the cross-attention mechanism fusion network structure for training, and iterate by comparing the interference phase unwrapping images corresponding to each interference entangled phase map dataset output by the cross-attention mechanism fusion network structure with the real interference phase images corresponding to each interference entangled phase map dataset in the interference phase dataset, to obtain a cross-attention mechanism fusion network with the interference entangled phase map dataset as input and the corresponding interference phase unwrapping map dataset as output; and input the verification dataset into the cross-attention mechanism fusion network for verification, and input the test dataset into the cross-attention mechanism fusion network to obtain the predicted interference phase unwrapping result.

[0040] Specifically, if Figure 1As shown, step S1 includes interference phase preprocessing, step S2 includes the structural design of the cross-attention mechanism fusion network, and step S3 includes the training of the cross-attention mechanism fusion network, the verification of the cross-attention mechanism fusion network, the output of the phase untangling result, and the evaluation of the phase untangling result.

[0041] In an exemplary embodiment, a method for interferometric phase unwrapping based on a cross-attention fusion network includes generating an interferometric phase dataset in step S1. The dataset used in this embodiment consists of a synthetic phase image containing random shapes and its corresponding wrapped phase image. These random shapes are created by adding and subtracting several Gaussian functions of different shapes and positions. This mixing of Gaussian functions ensures that the resulting irregular shapes are achieved.

[0042] Specifically, by adding and subtracting several Gaussian functions of different shapes and positions, irregular and arbitrary interferometric phase images are created. Slopes along the vertical and horizontal directions and Gaussian additive noise are added to the interferometric phase image to create an interferometric phase image dataset. The steps are as follows:

[0043] By adding and subtracting several Gaussian functions of different shapes and positions, irregular and arbitrary shaped interference phase patterns can be created;

[0044] Randomly select ramps and add them to the interference phase image in vertical and horizontal directions to form an interference phase image with ramp phase;

[0045] The interference phase image phase is pixel phase wrapped, and the wrapped phase The calculation formula is as follows:

[0046] ;

[0047] in, is an exponential function with the natural constant e as the base, is the original true phase of the interference phase image pixel, is the spatial coordinate of the pixel in the phase image, To find the angle sign.

[0048] Following the aforementioned method, a dataset of 1400 noise-affected interferometric phase images and their corresponding true interferometric phase images was created. The interferometric phase image dataset was randomly divided into training, validation, and test datasets in a 7:2:1 ratio. Each interferometric phase image had a pixel size of 256×256, with pixel values ranging from -55 to 55. The interferometric phase image data in the dataset were randomly assigned Gaussian additive noise of 0dB, 3dB, 5dB, and 10dB to construct the interferometric phase dataset.

[0049] In an exemplary embodiment, a method for interferometric phase unwrapping based on a cross-attention mechanism fusion network, step S2, includes designing a cross-attention mechanism fusion network. The cross-attention mechanism fusion network structure is shown in FIG. Figure 2 As shown, where c is the number of channels and s is the stride, it consists of an input layer, a 7x7 convolutional layer with a stride of 2 (i.e., the first convolutional layer), a maximum pooling layer with a pooling size of 3 and a stride of 2, nine depth-separable convolutional layers (i.e., the first depth-separable convolutional unit, the second depth-separable convolutional unit, and each depth-separable convolutional layer in the third depth-separable convolutional unit), two 1x1 convolutional layers with a stride of 2 (i.e., the sixth convolutional layer in the second depth-separable convolutional unit and the seventh convolutional layer in the third depth-separable convolutional unit), and a channel attention module. , thirteen residual modules (i.e., the residual modules in the first residual unit, the second residual unit, the third residual unit, and the fourth residual unit), a fusion module based on the cross attention mechanism, a 1x1 convolution layer (i.e., the second convolution layer), three 3x3 convolution layers (i.e., the third convolution layer, the fourth convolution layer, and the fifth convolution layer), three batch normalization and linear activation combination layers (i.e., the first batch normalization and linear activation combination layer, the second batch normalization and linear activation combination layer, the third batch normalization and linear activation combination layer), an upsampling layer, and an output layer, and according to Figure 2 Connected.

[0050] In an exemplary embodiment, Figure 2As shown in Figure 1, the input layer is sequentially connected to a 7x7 convolutional layer with a stride of 2 (i.e., the first convolutional layer) and a maximum pooling layer with a pooling size of 3 and a stride of 2, and then divided into two branches. The first branch is sequentially connected to thirteen residual modules (i.e., each residual module in the first residual unit, the second residual unit, the third residual unit, and the fourth residual unit) as the first input end of the fusion module based on the cross attention mechanism; the second branch is sequentially connected to three depthwise separable convolutional layers (i.e., the first depthwise separable convolutional layer, the second depthwise separable convolutional layer, and the third depthwise separable convolutional layer in the first depthwise separable convolutional unit, the first depthwise separable convolutional layer, the second depthwise separable convolutional layer, and the third depthwise separable convolutional layer in the first depthwise separable convolutional unit are all 3x3 convolutional layers with a channel number c=64), a 1x1 convolutional layer with a stride of 2 (i.e., the sixth convolutional layer in the second depthwise separable convolutional unit), and three depthwise separable convolutional layers (i.e., the second depthwise separable convolutional layer). The fourth depth-wise separable convolutional layer, the fifth depth-wise separable convolutional layer, and the sixth depth-wise separable convolutional layer in the unit, the fourth depth-wise separable convolutional layer, the fifth depth-wise separable convolutional layer, and the sixth depth-wise separable convolutional layer in the second depth-wise separable convolutional unit are all 3x3 convolutional layers with c=128 channels), one 1x1 convolutional layer with a stride of 2 (i.e., the seventh convolutional layer in the third depth-wise separable convolutional unit), three depth-wise separable convolutional layers (i.e., the seventh depth-wise separable convolutional layer, the eighth depth-wise separable convolutional layer, and the ninth depth-wise separable convolutional layer in the third depth-wise separable convolutional unit, the seventh depth-wise separable convolutional layer, the eighth depth-wise separable convolutional layer, and the ninth depth-wise separable convolutional layer in the third depth-wise separable convolutional unit are all 3x3 convolutional layers with c=256 channels), a batch normalization and linear activation combination layer (i.e., the first batch normalization and linear activation combination layer), and a channel attention module as the second input of the fusion module based on the cross-attention mechanism.

[0051] In an exemplary embodiment, Figure 2 As shown in the figure, the output end of the fusion module based on the cross attention mechanism is sequentially connected to the 1x1 convolution layer (i.e., the second convolution layer), the 3x3 convolution layer (i.e., the third convolution layer), the batch normalization and linear activation combination layer (i.e., the second batch normalization and linear activation combination layer), the 3x3 convolution layer (i.e., the fourth convolution layer), the batch normalization and linear activation combination layer (i.e., the third batch normalization and linear activation combination layer), the upsampling layer, the 3x3 convolution layer (i.e., the fifth convolution layer) and the output layer.

[0052] In an exemplary embodiment, Figure 2As shown in the figure, each residual module in each residual unit consists of an input layer, two 3x3 convolution layers (i.e., the first residual convolution layer and the second residual convolution layer), a batch normalization and linear activation combination layer (i.e., residual batch normalization and linear activation combination layer), a 1x1 convolution layer (i.e., the third residual convolution layer), a judgment block, two batch normalization layers (i.e., the first residual batch normalization layer and the second residual batch normalization layer), an additive fusion layer (i.e., residual additive fusion layer), a linear activation layer (i.e., residual linear activation layer) and an output layer (i.e., residual output layer). The residual module structure is different depending on the input step size. Specifically, when the step size is not 1, the input layer of the residual module is connected to a 1x1 convolution layer (i.e., the third residual convolution layer) with the same step size and channel as the input and a 3x3 convolution layer (i.e., the first residual convolution layer) with the same step size and channel as the input. The branch connecting the 3x3 convolution layer (i.e., the first residual convolution layer) is connected in sequence to a batch normalization and linear activation function combination layer (i.e., residual batch normalization and linear activation combination layer), a 3x3 convolution layer (i.e., the first residual convolution layer), and a linear activation function combination layer (i.e., residual batch normalization and linear activation combination layer). The first residual convolution layer (i.e., the second residual convolution layer) and a batch normalization layer (i.e., the first residual batch normalization layer) are used as the first input of the additive fusion layer (i.e., the residual additive fusion layer). The branch connected to the 1x1 convolution layer (i.e., the third residual convolution layer) is connected to a batch normalization layer (i.e., the second residual batch normalization layer) and then used as the second input of the additive fusion layer (i.e., the residual additive fusion layer). The output of the additive fusion layer (i.e., the residual additive fusion layer) is connected to a linear activation layer (i.e., the residual linear activation layer) and then the residual module result is output. When the stride is 1, the input layer of the residual module is connected to an additive fusion layer (i.e., residual additive fusion layer) and a 3x3 convolution layer (i.e., the first residual convolution layer) with the same stride and channel as the input. The branch connecting the 3x3 convolution layer (i.e., the first residual convolution layer) is sequentially connected to a batch normalization and linear activation function combination layer (i.e., residual batch normalization and linear activation combination layer), a 3x3 convolution layer (i.e., the second residual convolution layer), and a batch normalization layer (i.e., the first residual batch normalization layer) as an input of the additive fusion layer (i.e., residual additive fusion layer). The output of the additive fusion layer (i.e., residual additive fusion layer) is connected to a linear activation layer (i.e., residual linear activation layer) before the residual module result is output. Among them, Figure 2 The s of the first residual convolution layer and the third residual convolution layer in each residual module will change according to the input step size, and the s of the second residual convolution layer defaults to 1. The c of the convolution layer will change with the number of input channels.

[0053] In an exemplary embodiment, Figure 2As shown in the figure, the channel attention module consists of an input layer, a global average pooling layer, three fully connected layers (i.e., the first channel attention fully connected layer, the second channel attention fully connected layer, and the third channel attention fully connected layer), an additive fusion layer (i.e., the channel attention additive fusion layer), a sigmoid function layer (i.e., the channel attention sigmoid function layer), a linear activation function layer, a multiplicative fusion layer (i.e., the channel attention multiplicative fusion layer), and an output layer. The first branch of the channel attention module's input layer passes through a global average pooling layer, the first channel attention fully connected layer, the first branch of the second channel attention fully connected layer, a linear activation function layer, a third channel attention fully connected layer, and a channel attention sigmoid function layer. The output of the channel attention sigmoid function layer serves as the first input of the channel attention multiplication fusion layer, and the output of the channel attention multiplication fusion layer serves as the first input of the channel attention addition fusion layer. The second branch of the channel attention module's input layer serves as the second input of the channel attention addition fusion layer, and the second branch of the second channel attention fully connected layer serves as the second input of the channel attention multiplication fusion layer. The output of the channel attention addition fusion layer is connected to the output layer of the channel attention module. Convolutional layers without a specified stride have a stride of 1.

[0054] In an exemplary embodiment, the fusion module structure based on the cross attention mechanism is as follows: Figure 3As shown in Figure 1, the module consists of two input layers (i.e., the first fusion input layer, the second fusion input layer), three average pooling layers (i.e., the first average pooling layer, the second average pooling layer, the third average pooling layer), three 1x1 convolutional layers (i.e., the first fusion convolutional layer, the second fusion convolutional layer, the third fusion convolutional layer), one 3x3 convolutional layer (i.e., the fourth fusion convolutional layer), four upsampling layers (i.e., the first fusion upsampling layer, the second fusion upsampling layer, the third fusion upsampling layer, the fourth fusion upsampling layer), a cross attention module, three batch normalization and linear activation combination layers (i.e., the first fusion batch normalization and linear activation combination layer, the second fusion batch normalization and linear activation combination layer, the third fusion batch normalization and linear activation combination layer), a splicing layer (i.e., the fusion splicing layer), a batch normalization layer (i.e., the fusion batch normalization layer), an addition fusion layer (i.e., the cross addition fusion layer), a linear activation layer (i.e., the fusion linear activation layer) and an output layer (i.e., the fusion output layer). Specifically, the first branch of the first fusion input layer serves as the first input end of the cross-attention module, the second branch of the first fusion input layer is connected to the first fusion upsampling layer, the first fusion upsampling layer is connected to four branches respectively, and the first branch of the first fusion upsampling layer is directly used as the first input end of the fusion splicing layer; the second branch of the first fusion upsampling layer is sequentially connected to the first average pooling layer (P=2), the first fusion convolution layer (1x1 convolution layer), the first fusion batch normalization and linear activation combination layer, and the second fusion upsampling layer, and the output end of the second fusion upsampling layer serves as the second input end of the fusion splicing layer; the third branch of the first fusion upsampling layer is sequentially connected to the second average pooling layer (P=4), the second fusion convolution layer (1x1 convolution layer), the second fusion batch normalization and linear activation combination layer, and the third fusion upsampling layer, and the output end of the third fusion upsampling layer serves as the fusion splicing layer The third input end of the first fusion upsampling layer; the fourth branch of the first fusion upsampling layer is connected to the third average pooling layer (P=8), the third fusion convolution layer (1x1 convolution layer), the third fusion batch normalization and linear activation combination layer, and the fourth fusion upsampling layer in sequence, and the output end of the fourth fusion upsampling layer serves as the fourth input end of the fusion splicing layer; the output end of the fusion splicing layer is connected to the fourth fusion convolution layer (3x3 convolution layer) and the fusion batch normalization layer in sequence, and the output end of the fusion batch normalization layer serves as the first input end of the cross-addition fusion layer; the second fusion input layer is directly used as the second input end of the cross-attention module, and the output end of the cross-attention module serves as the second input end of the cross-addition fusion layer. The output end of the cross-addition fusion layer is connected to the fusion linear activation layer, and the output end of the fusion linear activation layer is connected to the fusion output layer, and the result of the fusion module based on the cross-attention mechanism is output through the fusion output layer.

[0055] In an exemplary embodiment, Figure 3As shown, the cross attention module includes three fully connected layers (i.e., the first cross fully connected layer, the second cross fully connected layer, and the third cross fully connected layer), four convolutional layers (i.e., the first cross convolutional layer, the second cross convolutional layer, the third cross convolutional layer, and the fourth cross convolutional layer), three deformation layers (i.e., the first deformation layer, the second deformation layer, and the third deformation layer), a subtraction layer, a sigmoid function layer (i.e., a cross sigmoid function layer), a reinforcement layer, a transposition layer, a calculation layer, a splicing layer (i.e., a cross splicing layer), an instance normalization layer, and a cross upsampling layer. The first branch of the first fusion input layer serves as the cross attention layer. The first input end of the attention module and the second fusion input layer are directly used as the second input end of the cross attention module. The first input end of the cross attention module is divided into three branches. The first branch of the first input end of the cross attention module is connected to the first cross fully connected layer and the first deformation layer in sequence as the first input end of the calculation layer; the second branch of the first input end of the cross attention module is connected to the first cross convolution layer (1x1 convolution layer) and serves as the first input end of the reinforcement layer; the third branch of the first input end of the cross attention module serves as the first input end of the subtraction layer; the third branch of the first input end of the cross attention module serves as the first input end of the subtraction layer; the third branch of the first input end of the cross attention module serves as the first input end of the subtraction layer; the third branch of the first input end of the cross attention module serves as the first input end of the subtraction layer. The two input ends are divided into four branches. The first branch of the second input end of the cross attention module serves as the second input end of the subtraction layer. The second branch of the second input end of the cross attention module is connected to the third cross convolution layer (1x1 convolution layer) as the second input end of the reinforcement layer. The third branch of the second input end of the cross attention module is connected to the second cross fully connected layer and the second deformation layer in sequence as the second input end of the calculation layer. The fourth branch of the second input end of the cross attention module serves as the first input end of the cross splicing layer. The output end of the subtraction layer is connected to the second cross convolution layer (3x3 convolution layer) in sequence. The output of the reinforcement layer is connected to the third cross fully connected layer, the third deformation layer, and the transposition layer in sequence, and the output of the transposition layer is used as the third input of the calculation layer; the output of the calculation layer is connected to the cross splicing layer as the second input of the cross splicing layer; the output of the cross splicing layer is connected to the fourth cross convolution layer (3x3 convolution layer), the instance normalization layer, and the cross upsampling layer in sequence, and the output of the cross upsampling layer is used as the output of the cross attention module, and the output of the cross attention module is used as the second input of the cross-add fusion layer.

[0056] Furthermore, the cross attention mechanism is used in the cross attention module. Assume that the input is two feature sequences: X 1 and X 2 ,in X 1 After passing through the fully connected layer, it is called query Q (Query), where X1 After passing through the 1x1 convolutional layer, it is called the query Q`, where X 2 After passing through the fully connected layer, it is called key K (Key), where X 2 After passing through the fully connected layer, it is called value V (Value). X 1 and X 2 The difference D is obtained after the subtraction and the 3x3 convolution layer and the sigmoid function layer. The reinforcement layer is defined as the following steps:

[0057] First, get strengthened , as follows:

[0058] ;

[0059] Afterwards, it is strengthened , as follows:

[0060] ;

[0061] in represents element-wise multiplication, is the output of the reinforcement layer.

[0062] The computation layer is defined as the following steps:

[0063] First, the attention score is as follows:

[0064] ;

[0065] in, is the dimension of the key vector, used to scale the dot product result to avoid gradient explosion.

[0066] Normalize the attention scores to a probability distribution.

[0067] The final step is weighted summation, as follows:

[0068] ;

[0069] Where V is the output of the fully connected layer, and the value vector is weighted summed using the attention weight to obtain the final output. Figure 3 The V, K, and Q of the input calculation layer are the feature maps after deformation, so we use Represents the feature map of the input computation layer.

[0070] In an exemplary embodiment, step S3 includes training a cross-attention mechanism fusion network. Relevant parameters are set, and the interferometric phase dataset is randomly divided into a training dataset, a validation dataset, and a test dataset in a ratio of 7:2:1. The training dataset is fed into the cross-attention mechanism fusion network structure, and the cross-attention mechanism fusion network structure is trained until the cross-attention mechanism fusion network structure converges, and the training weights of the cross-attention mechanism fusion network structure are saved. Specifically, the steps are as follows:

[0071] Step S301, setting the training starting learning rate, maximum learning rate, training batch size and number of training rounds; optionally, the starting learning rate is 0.0001, the maximum learning rate is 0.01, the training batch size is 4 and the number of training rounds is 100;

[0072] Step S302: During the training process, L2 norm regularization is used to prevent the network from overfitting.

[0073] Step S303: Use the Adam gradient optimization algorithm to optimize the cross attention mechanism fusion network structure training. During the optimization process, the following function is used as the loss function: composite loss function The definition is as follows:

[0074] ;

[0075] in,

[0076] ;

[0077] Where, As weights, set the two weights to 1 and 0.1 respectively. Experience E represents the expectation. To predict the phase and true phase The mean square error between To predict the phase and true phase Between The total average error in the direction.

[0078] Step S304, repeat steps S301 and S302 until the cross-attention mechanism fusion network structure converges, obtain the final cross-attention mechanism fusion network model and weights for interference phase unwrapping, and save the network training weights.

[0079] Furthermore, step S3 also includes verifying the cross-attention mechanism fusion network. The network training weights are loaded, and the validation dataset is fed into the trained cross-attention mechanism fusion network for verification. Then, the test dataset is fed into the trained cross-attention mechanism fusion network for testing to obtain the predicted interference phase disentanglement result.

[0080] In an exemplary embodiment, an interference phase unwrapping method based on a cross-attention mechanism fusion network further includes selecting an evaluation metric. Specifically, using the normalized root mean square error Evaluate the phase disentanglement accuracy of the cross-attention mechanism fusion network. The smaller the value, the higher the phase unwrapping accuracy and the better the network performance. The calculation formula is as follows:

[0081] ;

[0082] in, is the normalized root mean square error, is a sample of the training dataset, For training data set samples The difference between the true value of and its network training estimate, Represents the phase pixel coordinates.

[0083] The cross attention mechanism fusion network loss accuracy diagram is as follows Figure 4 As shown in the figure, the analysis shows that the loss of the cross attention mechanism fusion network decreases as the number of iterative words increases. The interference entanglement phase diagram under noise is as follows Figure 5 As shown, the real interference phase diagram corresponding to the interference entanglement phase diagram under noise is as follows Figure 6 As shown, analysis Figure 5 and Figure 6 It can be seen that in the case of noise, the phase information has changed significantly, indicating that noise has a great adverse effect on phase unwrapping. The interference unwrapped phase map predicted by the cross-attention mechanism fusion network is shown in Figure 7 As shown, analysis Figure 6 and Figure 7 It can be seen that by processing the interference entangled phase under noise through the network, the interference unwrapped phase with high similarity to the true interference phase can be predicted, indicating that the interference phase unwrapping method of the cross-attention mechanism fusion network can perform phase unwrapping with high accuracy. At the same time, the network is calculated by the root mean square error NRMSE. The present application can achieve high-precision unwrapping of the interference entangled phase under low signal-to-noise ratio. The NRMSE result reaches 1.10% and the required training time is 7 hours. Under the same other conditions, the root mean square error NRMSE obtained by phase unwrapping of other networks such as U-net, PhaseNet, PUnet and SCAPU method is 3.93%, 12.81%, 3.22% and 2.12% respectively. The root mean square error describes the gap between the predicted phase and the true phase. The smaller the value, the smaller the gap between the two, which also means the higher the accuracy of unwrapping. The experimental results show that the present invention is a highly robust and accurate interference phase unwrapping method under low signal-to-noise ratio conditions, which can effectively improve the accuracy of synthetic aperture radar interferometry.

[0084] The above-mentioned interference phase unwrapping method based on the cross-attention mechanism fusion network includes the use of a reasonable residual module for network construction, the design of a channel attention module-based enhanced deep separable convolution module, the design of a network fusion module based on the cross-attention mechanism, the construction, training and verification of the cross-attention mechanism fusion network, etc., to achieve interference phase unwrapping under low signal-to-noise ratio conditions and improve the precision of synthetic aperture radar interferometry. This application uses a deep separable convolution module based on the channel attention mechanism and a residual module as two inputs of the cross-attention mechanism fusion module respectively to strengthen the network's feature extraction of the input image, enhance the feature discrimination ability and capture the dependency between channels; at the same time, a cross-attention mechanism fusion module is used to strengthen the dependency between the output feature maps of the two branches, and the information between the two is fused and aligned to increase the expression ability of the model. This application uses a residual module and a deep separable module based on the channel attention mechanism to improve the feature extraction function in the encoding process. At the same time, the deep separable convolution module can reduce the computational complexity of the network and reduce the training time of the network.

[0085] It should be noted that in this application, the function of the residual branch is to retain multi-level feature information and alleviate the problem of gradient disappearance or explosion through thirteen residual modules. However, using multiple residual blocks alone lacks dynamic adjustment of channel importance and will cause feature redundancy and noise sensitivity. To avoid such problems, this application adds a depth-separable convolution branch based on channel attention. The main functions of this branch also include: first, the depth-separable convolution branch based on channel attention is different from the feature extraction information of the residual branch, and serves as another input end of different input sources based on the cross-fusion module; second, the context information of the input feature map is extracted while reducing the amount of calculation through depth-separable convolution, where the number of channels of the input channel attention is 256, which is lower than the number of channels at the end of the residual module, which is convenient for enhancing important channel features; third, the multi-fully connected layer structure and activation function in the channel attention increase the nonlinear modeling capability of the channel attention module, enabling the channel attention module to learn more complex relationships between channels; fourth, by using the feature map of the input layer of the channel attention module as an input end of the additive fusion layer, this residual structure is used to enhance the integrity of the feature map while enhancing feature information.

[0086] For the fusion module based on the cross-attention mechanism, the cross-attention module is used in this application to realize the information interaction and fusion tasks of different feature maps. In the fusion module based on the cross-attention mechanism, the main functions of the cross-attention fusion module also include: first, the K in the cross-attention is strengthened based on the difference between the two input feature maps. This strengthening can help the model better capture the complementary information between the modalities and enhance the difference of the features. In the interference phase map, these differences are specifically manifested in the noise or fringe phase, which is the key point for evaluating the performance of phase unwrapping. Second, the data set used in this application is relatively small, so this application uses instance normalization in the cross-attention module to better highlight local features, which can help the model better capture the local correlation between different information. Third, this application abandons the combination of the average pooling layer with P=1, the convolution layer, the batch normalization and the linear activation combination layer and the upsampling layer, simplifies the network structure, avoids redundant branches, and fuses the feature maps obtained by average pooling of different sizes to enrich the feature information at different levels. Fourth, the input end of the input layer of the second fusion layer is used as the feature map of the input information of one input end of the splicing fusion, thereby enhancing the feature fusion capability and reducing information loss, while improving the generalization capability of the model.

[0087] At least a portion of the steps in the flowcharts involved in the various embodiments described above may include multiple steps or multiple stages. These steps or stages may be executed at different times, or may be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.

[0088] Those skilled in the art will be able to make several modifications and improvements without departing from the concept of the present application, and these modifications and improvements are all within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be based on the appended claims.

Claims

1. An interference phase unwrapping method based on a cross-attention mechanism fusion network, characterized in that: The method is used to execute steps S1 to S3 to obtain a cross-attention mechanism fusion network for interference phase unwrapping, including: Step S1, generating an interference entanglement phase image dataset based on a Gaussian function according to a real interference phase image, wherein each interference entanglement phase image dataset and the corresponding real interference phase image constitute an interference phase dataset; Step S2, constructing a cross-attention mechanism fusion network structure based on multi-scale pooling, including an input layer, a first convolutional layer, a maximum pooling layer, a first depth-separable convolutional unit, a second depth-separable convolutional unit, a third depth-separable convolutional unit, a first batch of normalization and linear activation combination layers, a channel attention module, a first residual unit, a second residual unit, a third residual unit, a fourth residual unit, a fusion module based on a cross-attention mechanism, a second convolutional layer, a third convolutional layer, a second batch of normalization and linear activation combination layers, a fourth convolutional layer, a third batch of normalization and linear activation combination layers, an upsampling layer, a fifth convolutional layer and an output layer; each residual unit includes each residual module, and the number of channels of the input channel attention module is lower than that of the fourth residual unit. The number of channels of each residual module in the difference unit; wherein, the fusion module based on the cross attention mechanism includes the first fusion input layer, the second fusion input layer, the first average pooling layer, the second average pooling layer, the third average pooling layer, the first fusion convolution layer, the second fusion convolution layer, the third fusion convolution layer, the fourth fusion convolution layer, the first fusion upsampling layer, the second fusion upsampling layer, the third fusion upsampling layer, the fourth fusion upsampling layer, the cross attention module, the first fusion batch normalization and linear activation combination layer, the second fusion batch normalization and linear activation combination layer, the third fusion batch normalization and linear activation combination layer, the fusion splicing layer, the fusion batch normalization layer, the cross addition fusion layer, the fusion linear activation layer and the fusion output layer, wherein the first fusion input layer The first branch of the first fusion input layer is connected to the first fusion upsampling layer. The first fusion upsampling layer is connected to four branches respectively. The first branch of the first fusion upsampling layer is directly used as the first input of the fusion splicing layer; the second branch of the first fusion upsampling layer is connected to the first average pooling layer, the first fusion convolution layer, the first fusion batch normalization and linear activation combination layer, and the second fusion upsampling layer in sequence as the second input of the fusion splicing layer; the third branch of the first fusion upsampling layer is connected to the second average pooling layer, the second fusion convolution layer, the second fusion batch normalization and linear activation combination layer, and the third fusion upsampling layer in sequence as the third input of the fusion splicing layer; The fourth branch of the first fusion upsampling layer is connected to the third average pooling layer, the third fusion convolution layer, the third fusion batch normalization and linear activation combination layer, and the fourth fusion upsampling layer in sequence as the fourth input end of the fusion splicing layer; the output end of the fusion splicing layer is connected to the fourth fusion convolution layer and the fusion batch normalization layer in sequence as the first input end of the cross-addition fusion layer; the second fusion input layer is directly used as the second input end of the cross-attention module, the output end of the cross-attention module is used as the second input end of the cross-addition fusion layer, the output end of the cross-addition fusion layer is connected to the fusion linear activation layer, the output end of the fusion linear activation layer is connected to the fusion output layer, and the result of the fusion module based on the cross-attention mechanism is output through the fusion output layer; Step S3, divide the interference phase dataset into a training dataset, a verification dataset and a test dataset, and input the training dataset into the cross-attention mechanism fusion network structure for training, and iterate by comparing the interference phase unwrapping images corresponding to each interference entangled phase map dataset output by the cross-attention mechanism fusion network structure with the real interference phase images corresponding to each interference entangled phase map dataset in the interference phase dataset, to obtain a cross-attention mechanism fusion network with the interference entangled phase map dataset as input and the corresponding interference phase unwrapping map dataset as output; and input the verification dataset into the cross-attention mechanism fusion network for verification, and input the test dataset into the cross-attention mechanism fusion network to obtain the predicted interference phase unwrapping result.

2. The interference phase unwrapping method based on the cross-attention mechanism fusion network according to claim 1 is characterized in that: In step S2, the input layer is connected to the first convolution layer and the maximum pooling layer in sequence from input to output and is divided into a first branch and a second branch, wherein the first branch is connected to the first residual unit, the second residual unit, the third residual unit, and the fourth residual unit in sequence from input to output, and the output end of the fourth residual unit is connected to the first input end of the fusion module based on the cross-attention mechanism; the second branch is connected to the first depth-separable convolution unit, the second depth-separable convolution unit, the third depth-separable convolution unit, the first batch of normalization and linear activation combination layer, and the channel attention module in sequence from input to output, and the output end of the channel attention module is connected to the second input end of the fusion module based on the cross-attention mechanism, and the fusion module based on the cross-attention mechanism is connected to the second convolution layer, the third convolution layer, the second batch of normalization and linear activation combination layer, the fourth convolution layer, the third batch of normalization and linear activation combination layer, the upsampling layer, the fifth convolution layer, and the output layer in sequence from input to output.

3. The interference phase unwrapping method based on the cross-attention mechanism fusion network according to claim 2 is characterized in that: In the step S2, the first residual unit includes a first residual module and a second residual module in order from input to output, wherein the input end of the first residual module is the input end of the first residual unit, which is connected to the output end of the maximum pooling layer, and the output end of the second residual module is the output end of the first residual unit, which is connected to the input end of the second residual unit; the second residual unit includes a third residual module, a fourth residual module, and a fifth residual module in order from input to output, wherein the input end of the third residual module is the input end of the second residual unit, which is connected to the output end of the first residual unit, and the output end of the fifth residual module is the output end of the second residual unit, which is connected to the input end of the third residual unit; the third residual unit includes a fourth residual module and a fifth residual module in order from input to output, wherein the input end of the third residual module is the input end of the second residual unit, which is connected to the output end of the first residual unit, and the output end of the fifth residual module is the output end of the second residual unit, which is connected to the input end of the third residual unit; It includes the sixth residual module, the seventh residual module, the eighth residual module, the ninth residual module, and the tenth residual module, wherein the input end of the sixth residual module is the input end of the third residual unit, which is connected to the output end of the second residual unit, and the output end of the tenth residual module is the output end of the third residual unit, which is connected to the input end of the fourth residual unit; the fourth residual unit includes the eleventh residual module, the twelfth residual module, and the thirteenth residual module from the input to the output direction, wherein the input end of the eleventh residual module is the input end of the fourth residual unit, which is connected to the output end of the third residual unit, and the output end of the thirteenth residual module is the output end of the fourth residual unit, which is connected to the first input end of the fusion module based on the cross-attention mechanism.

4. The interference phase unwrapping method based on the cross-attention mechanism fusion network according to claim 3 is characterized in that: Each residual module in each residual unit in step S2 includes an input layer, a first residual convolution layer, a residual batch normalization and linear activation combination layer, a second residual convolution layer, a first residual batch normalization layer, a residual addition fusion layer, a residual linear activation layer, a judgment block, a third residual convolution layer, a second residual batch normalization layer, and an output layer. The input layer serves as the input end of the residual module and is connected to the input ends of the judgment block and the first residual convolution layer, respectively. The first residual convolution layer, the residual batch normalization and linear activation combination layer, the second residual convolution layer, and the first residual batch normalization layer are connected in sequence. The judgment block, the third residual convolution layer, and the second residual batch normalization layer are connected in sequence. The output ends of the first residual batch normalization layer and the second residual batch normalization layer are connected to the input end of the residual addition fusion layer. The ends are connected, the residual addition fusion layer, the residual linear activation layer, and the output layer of the residual module are connected in sequence, and the output layer serves as the output end of the residual module; when the input step size of each residual module in each residual unit is not 1, the input layer of the residual module is connected to the third residual convolution layer and the first residual convolution layer respectively, and the branch of the first residual convolution layer is connected to the residual batch normalization and linear activation combination layer, the second residual convolution layer, and the first residual batch normalization layer in sequence as the first input end of the residual addition fusion layer, the branch of the third residual convolution layer is connected to the second residual batch normalization layer as the second input end of the residual addition fusion layer, and the output end of the residual addition fusion layer is connected to the residual linear activation layer, and the residual module result is output through the output layer of the residual module; When the input stride of each residual module is 1, the input layer of the residual module is connected to the residual addition fusion layer as the first input end of the residual addition fusion layer and the first residual convolution layer respectively. The branch of the first residual convolution layer is sequentially connected to the residual batch normalization and linear activation combination layer, the second residual convolution layer, and the first residual batch normalization layer as the second input end of the residual addition fusion layer. The output of the residual addition fusion layer is connected to the residual linear activation layer, and the residual module result is output through the output layer of the residual module.

5. The interference phase unwrapping method based on the cross-attention mechanism fusion network according to claim 1 is characterized in that: In step S2, the first depth-wise separable convolution unit includes a first depth-wise separable convolution layer, a second depth-wise separable convolution layer, and a third depth-wise separable convolution layer in sequence from input to output, wherein the input end of the first depth-wise separable convolution layer is the input end of the first depth-wise separable convolution unit, and is connected to the output end of the maximum pooling layer. The output end of the third depth-wise separable convolution layer is the output end of the first depth-wise separable convolution unit, and is connected to the input end of the second depth-wise separable convolution unit.

6. The interference phase unwrapping method based on the cross-attention mechanism fusion network according to claim 5 is characterized in that: In step S2, the second depth-wise separable convolution unit includes, from input to output, a sixth convolution layer, a fourth depth-wise separable convolution layer, a fifth depth-wise separable convolution layer, and a sixth depth-wise separable convolution layer, wherein the input end of the sixth convolution layer is the input end of the second depth-wise separable convolution unit, and is connected to the output end of the first depth-wise separable convolution unit. The output end of the sixth depth-wise separable convolution layer is the output end of the second depth-wise separable convolution unit, and is connected to the input end of the third depth-wise separable convolution unit. The third depth-wise separable convolution unit includes, from input to output, a seventh convolution layer, a seventh depth-wise separable convolution layer, an eighth depth-wise separable convolution layer, and a ninth depth-wise separable convolution layer, wherein the input end of the seventh convolution layer is the input end of the third depth-wise separable convolution unit, and is connected to the output end of the second depth-wise separable convolution unit. The output end of the ninth depth-wise separable convolution layer is the output end of the third depth-wise separable convolution unit, and is connected to the input end of the first batch of normalization and linear activation combination layers.

7. The interference phase unwrapping method based on the cross-attention mechanism fusion network according to claim 1 is characterized in that: In step S2, the channel attention module includes an input layer, a global average pooling layer, a first channel attention fully connected layer, a second channel attention fully connected layer, a linear activation function layer, a third channel attention fully connected layer, a channel attention sigmoid function layer, a channel attention multiplication fusion layer, a channel attention addition fusion layer and an output layer. The first branch of the input layer of the channel attention module passes through the global average pooling layer, the first channel attention fully connected layer, the first branch of the second channel attention fully connected layer, the linear activation function layer, the third channel attention fully connected layer, the channel attention sigmoid function layer, the output end of the channel attention sigmoid function layer serves as the first input end of the channel attention multiplication fusion layer, the output end of the channel attention multiplication fusion layer serves as the first input end of the channel attention addition fusion layer, the second branch of the input layer of the channel attention module serves as the second input end of the channel attention addition fusion layer, the second branch of the second channel attention fully connected layer serves as the second input end of the channel attention multiplication fusion layer, and the output end of the channel attention addition fusion layer is connected to the output layer of the channel attention module.

8. The interference phase unwrapping method based on the cross-attention mechanism fusion network according to claim 1 is characterized in that: In step S2, the cross-attention module includes a first cross-fully connected layer, a second cross-fully connected layer, a third cross-fully connected layer, a first cross-convolutional layer, a second cross-convolutional layer, a third cross-convolutional layer, a fourth cross-convolutional layer, a first deformation layer, a second deformation layer, a third deformation layer, a subtraction layer, a cross-sigmoid function layer, an enhancement layer, a transposition layer, a calculation layer, a cross-splitting layer, an instance normalization layer, and a cross-upsampling layer. The first branch of the first fusion input layer serves as the first input end of the cross-attention module, and the second fusion input layer directly serves as the second input end of the cross-attention module. The first input end of the cross-attention module is divided into three branches. The first branch of the first input end of the cross-attention module is sequentially connected to the first cross-fully connected layer and the first deformation layer as the first input end of the calculation layer; the second branch of the first input end of the cross-attention module is connected to the first cross-convolutional layer and serves as the first input end of the enhancement layer; the third branch of the first input end of the cross-attention module serves as the first input end of the subtraction layer; the second input end of the cross-attention module is divided into four branches. The first branch of the second input end of the cross-attention module serves as the second input end of the subtraction layer, and the second branch of the second input end of the cross-attention module is connected to the third cross-convolution layer as the second input end of the reinforcement layer; the third branch of the second input end of the cross-attention module is connected to the second cross-fully connected layer and the second deformation layer in sequence as the second input end of the calculation layer; the fourth branch of the second input end of the cross-attention module serves as the first input end of the cross-splicing layer; the output end of the subtraction layer is connected to the second cross-convolution layer and the cross-sigmoid function layer in sequence, and serves as the third input end of the reinforcement layer; the output end of the reinforcement layer is connected to the third cross-fully connected layer, the third deformation layer, and the transposition layer in sequence, and serves as the third input end of the calculation layer; the output end of the calculation layer is connected to the cross-splicing layer as the second input end of the cross-splicing layer; the output end of the cross-splicing layer is connected to the fourth cross-convolution layer, the instance normalization layer, and the cross-upsampling layer in sequence, and the output end of the cross-upsampling layer serves as the output end of the cross-attention module, and the output end of the cross-addition fusion layer serves as the second input end of the cross-attention module.

9. The interference phase unwrapping method based on the cross-attention mechanism fusion network according to claim 1 is characterized in that: After step S3, the following steps are also included: The normalized root mean square error is used to evaluate the network phase disentanglement accuracy of the cross-attention mechanism fusion network. The calculation formula is as follows: Among them, NRMSE is the normalized root mean square error, N is the total number of samples in the training data set, and y′ k (i, j) is the difference between the true value of sample k in the training dataset and the estimated value of the cross-attention mechanism fusion network training, and (i, j) represents the phase pixel coordinates.

Citation Information

Patent Citations

  • Video crowd counting method based on double-branch space-time interaction network

    CN118781553A

  • Interferometric phase unwrapping method based on double-branch decoding residual network

    CN119667681A