High-fidelity hologram generation method using multi-domain transformation and multi-resolution training
By constructing the Holo-UFormer network model, combining three-level wavelet transform and Fourier transform, and adopting a dual-resolution migration training strategy, the problem of insufficient details and depth perception of hologram generation in the existing technology is solved, and the generation of high-fidelity holograms is realized.
Patent Information
- Application Number
- CN202510493530.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
When the existing deep learning model processes the generation of three-dimensional holograms with continuous depth, it is difficult to effectively process the depth information of complex scenes, resulting in insufficient details and depth perception of the generated holograms. The existing methods mainly rely on airspace information and single-resolution images, and fail to make full use of multi-scale information and frequency domain features.
The Holo-UFormer network model is built, three-level wavelet transform and Fourier transform are introduced, and the multi-domain feedforward network is improved in combination with the window attention mechanism. The two-phase migration training strategy of dual resolution is adopted. First pre-trained at low resolution, and then migrated to high resolution for fine-tuning to improve the model's ability to capture detailed features.
The generated holograms are more realistic and three-dimensional in detail performance and depth perception, improving the quality and reality of the holograms, significantly better than the reconstruction accuracy of the amplitude and phase maps of traditional methods.
Smart Images

Figure CN120353111A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of general image data processing or generation, and more particularly to a method for generating high-fidelity holograms using multi-domain transformation and multi-resolution training in the technical field of methods or devices for pattern recognition or machine learning of images or videos. The present invention can be used to generate high-fidelity holograms by capturing multi-scale features of images in complex scenes. Background Art
[0002] In the process of computer-generated holograms based on the point-source method, a three-dimensional object is discretized into a series of point light sources, and each point light source is regarded as an independent spherical wave emission source. These spherical waves propagate on the hologram plane and interfere with each other to form complex interference fringes, which are finally superimposed to generate a hologram. The prior art mainly relies on spatial domain information and single-resolution image processing methods. This method often fails to capture the multi-scale features of images when dealing with complex scenes, resulting in insufficiently fine details in the generated holograms. Especially when dealing with objects with complex textures and fine structures, the phenomenon of detail loss is more obvious. Since frequency domain features are crucial for understanding the global structure and depth information of images, the lack of this aspect in the existing methods limits the performance of holograms in depth perception. The prior art usually uses a single-resolution dataset in the training process, which limits the generalization ability of the model and its adaptability to images of different resolutions. This training method not only reduces the robustness of the model but also limits its performance when dealing with high-resolution images, resulting in the quality of the generated holograms not meeting the high-standard application requirements.
[0003] Scholars such as Shi Liang proposed a point-based occlusion trajectory tracking method for generating a large-scale Fresnel three-dimensional hologram dataset MIT-CGH-4K in their published paper "Towards real-time photorealistic 3D holography with deep neural networks" (Nature, 2021, 591(7849): 234-239.). Based on this dataset, the researchers designed a convolutional neural network model TensorHolo to achieve the efficient generation of three-dimensional holograms. However, this method still has the following deficiencies. Given that in many computer vision tasks, neural networks combining Transformer and convolution are often superior to pure convolutional networks, the combination of the two can further improve the quality of hologram generation.
[0004] Scholars such as Shi Liang proposed a method in their published paper "End-to-end learning of 3D phase-only holograms for holographic display" (Light: Science & Applications, 2022, 11(1): 247.) that uses a two-stage training strategy to achieve the generation of high-quality pure-phase holograms at the same resolution. The shortcoming of this method is that although the scheme has completed training and verification on high-resolution data, it has not fully exploited the potential information in low-resolution data.
[0005] Scholars such as Dong Zhenxing explored an application method of the UFormer network based on Vision Transformer in the hologram generation task in their published paper "Vision transformer-based, high-fidelity, computer-generated holography" (Advances in Display Technologies XIII. SPIE, 2023, 12443: 47-53.), taking into account both generation quality and efficiency. However, the shortcoming of this method is that it only uses RGB images as input, lacks the support of depth information, and has not been verified in three-dimensional scenes containing depth information.
[0006] Scholars such as Dong Zhenxing effectively improved the generation quality of holograms by introducing a network structure that combines Fourier transform and convolutional modules in their published paper "Fourier-inspired neural module for real-time and high-fidelity computer-generated holography" (Optics Letters, 2023, 48(3): 759-762.), which is better than models based only on spatial domain convolution. However, the shortcoming of this method is that it only uses Fourier transform to extract frequency domain features, and the feature representation is still limited. Multi-scale frequency domain analysis methods such as three-level discrete wavelet transform can be further introduced to enhance the representation ability. Summary of the Invention
[0007] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies by proposing a high-fidelity hologram generation method using multi-domain transformation and multi-resolution training, aiming to solve the following two deficiencies of the existing technologies:
[0008] First, when existing deep learning models process the generation of deep continuous three-dimensional holograms, they adopt pure convolutional neural network models. Such models are difficult to effectively process the depth information of complex scenes, resulting in deficiencies in the detail performance and depth perception of the generated holograms.
[0009] Second, during the image generation process, existing methods mainly rely on spatial domain information and single-resolution images, and fail to fully utilize multi-scale information and frequency domain features. These problems limit the quality and realism of holograms.
[0010] The technical idea for achieving the purpose of the present invention is that the present invention constructs a Holo-UFormer network model by improving the UFormer model. In the Transformer module of this network model, a three-level wavelet transform (Wavelet Transform) and Fourier transform (Fourier Transform) are innovatively introduced, and combined with a window-based attention mechanism to efficiently extract the high and low frequency features of the image. The feed-forward network inside the Holo-UFormer network model is improved to a multi-domain feed-forward network (Multi-domain Feed-forward Network, MDFFN). In this way, the problem that existing deep learning models are difficult to effectively process the depth information of complex scenes when generating deep continuous three-dimensional holograms, resulting in deficiencies in the detail performance and depth perception of the generated holograms, is solved. The wavelet transform used in the invention is a discrete wavelet transform (DWT, Discrete Wavelet Transform), and its core idea is to decompose the image step by step through low-pass and high-pass filters to extract feature information at different scales. The present invention designs a two-stage transfer training strategy: first, pre-train the model at a low resolution to learn the basic feature representation; then transfer the model to a high resolution for fine-tuning to further optimize the model's ability to capture detail features. This training strategy can not only accelerate the model convergence but also effectively improve the generalization performance of the model on high-resolution data. Considering that the network structures used in the two-stage training are the same, there are two differences. One is the different resolutions of the input pictures, and the other is that under the requirements of actual physical devices, it is necessary to ensure that the low-resolution * pixel pitch = high-resolution * pixel pitch. In this way, the problem that existing methods mainly rely on spatial domain information and single-resolution images, and the problem of failing to fully utilize multi-scale information and frequency domain features are solved.
[0011] To achieve the above object, the specific implementation steps of the present invention are as follows:
[0012] Step 1, construct a multi-domain Transformer module;
[0013] Step 2, construct the Holo-UFormer network model;
[0014] Step 3, obtain a dataset containing RGB images, depth images, amplitude images, and phase images, and each type of image contains two samples with high and low resolutions;
[0015] Step 4, train the Holo-UFormer network model using a two-stage transfer training strategy with double resolutions;
[0016] Step 5, use the trained network model to generate high-fidelity holograms.
[0017] Furthermore, the multi-domain Transformer module is composed of a first normalization layer, a window multi-head attention layer, a second normalization layer, and a multi-domain feed-forward network connected in series in sequence.
[0018] Furthermore, each attention head in the window multi-head attention layer performs self-attention calculation using the following formula:
[0019]
[0020] where Attention(Q, K, V) represents the calculation process of the attention mechanism used in the multi-domain Transformer module. Denote the shape of the feature map when Holo-UFormer enters this module as {batch size, height, width, number of channels}. After window partitioning and flattening, the shape becomes {total number of windows, window height * window width, number of channels}, and Q (Query), K (Key), and V (Value) are calculated using a fully connected layer. k represents the serial number of the attention head, k = 1, 2,..., K, and the dimension of k is d k = C / k, C represents the number of channels, SoftMax represents the operation of normalizing the attention weights so that the sum of the values in each row is 1, the superscript T represents the transpose operation, and B represents the bias of the relative position encoding.
[0021] Furthermore, in the multi-domain feed-forward network, the data performs boundary padding of 1 in the spatial domain, with a kernel size of 3, a stride of 1, and a depthwise separable convolution with an output number of channels 4 times the input number of channels. In addition, a three-level wavelet transform is performed in the frequency domain branch, and the Fourier transform is performed after convolution.
[0022] Furthermore, the wavelet transform is a three-level discrete wavelet transform, and its implementation steps are as follows:
[0023] The first step is to gradually decompose the feature map through a low-pass filter and a high-pass filter to extract feature information at different scales;
[0024] Step 2: Decompose the feature map through a low-pass filter to generate multiple sub-bands, including the low-frequency sub-band LL and high-frequency sub-bands {LH, HL, HH}:
[0025]
[0026] Among them, LL j+1 is the low-frequency sub-band, representing the global features of the extracted image. LH j+1 is the horizontal high-frequency component, representing the edge features in the horizontal direction. HL j+1 is the vertical high-frequency component, representing the edge features in the vertical direction. HH j+1 is the diagonal high-frequency component, representing the detailed information in the diagonal direction. I(x, y) represents the pixel value of the feature map at position (i, j). φ(x) represents the scaling function of the low-pass filter used to extract the low-frequency components of the image, and ψ(x) represents the wavelet function of the high-pass filter used to extract the edge and texture information of the image. The expressions of the two functions φ(x) and ψ(x) are as follows:
[0027]
[0028] Among them, h k is the coefficient of the low-pass filter, which determines the weighting method of the scaling function. g k is the coefficient of the high-pass filter, which determines the weighting method of the wavelet function. k is the index of the filter, representing the coefficient range of the filter;
[0029] Step 3: Continue to perform wavelet transform on the wavelet low-frequency component LL at each level, using wavelet transform three times in total. For the first-level wavelet low-frequency component LL and the wavelet high-frequency components HL, LH, and HH at each level, after using bilinear interpolation to enlarge them to the original size, they are concatenated along the channel axis.
[0030] Furthermore, the Holo-UFormer network model adopts an encoder-bottleneck-decoder architecture. The encoding layer consists of four cascaded multi-domain Transformer modules. Each module downsamples the input features and extracts multi-scale spatio-frequency domain features. The bottleneck layer consists of 1 multi-domain Transformer module for deep feature representation and information compression. The decoding layer consists of four cascaded multi-domain Transformer modules. Each module gradually restores the spatial resolution through upsampling. Among them, skip connections are provided between the corresponding levels of the encoding layer and the decoding layer for fusing low-level detail features and high-level semantic features. The dimension of the feature map accepted by the encoder is set to 16. Downsampling is implemented using convolution with a stride of 2, a kernel size of 4, a padding size of 1, and an output channel twice the number of input channels. Upsampling is implemented using transposed convolution with a stride of 2 and a kernel size of 2, and the number of output channels is half the number of input channels. The number of input feature channels accepted by the multi-domain Transformer modules in the encoder is 16, 16*2, 16*4, 16*8 respectively. The corresponding levels of the decoder are the same as those of the encoder. The number of input feature channels accepted by the bottleneck layer is 16*8. The number of multi-head attention heads used in the Transformer modules of the Holo-UFormer network model is 1, 2, 4, 8, 16, 16, 8, 4, 2 in sequence.
[0031] Furthermore, the two-stage transfer training strategy means that the RGB map and depth map in the dataset are concatenated and input into the Holo-UFormer network model. The Holo-UFormer network model is pre-trained using the low-resolution samples in the dataset, and the amplitude map and phase map are predicted through the network. The pre-trained Holo-UFormer network model is retrained using the high-resolution samples in the dataset, and the network parameters are iteratively updated until the total loss function of the network converges, obtaining a trained network model.
[0032] The total loss function of the network is as follows:
[0033] L = L hologram + L propagation
[0034]
[0035] L propagation = σL amp + εL tv
[0036] L amp = ||AMP focal (H, D, d, d in ) - AMP focal (Hgt , D, d, d in ) || 2
[0037]
[0038] L tv = || TV(AMP focal (H, D, d, d in )) - TV(AMP focal (H gt , D, d, d in )) || 1
[0039] Among them, L represents the total loss function used for training the Holo-UFormer network. L hologram represents the difference between the amplitude map and phase map in the attention dataset and the amplitude map and phase map predicted by the network. α, β, σ, and ε are all weight parameters. A, represents the amplitude map and phase map in the dataset. A gt , respectively represent the amplitude and phase of the predicted hologram. e (·) represents the exponential function with the natural constant e as the base. i represents the imaginary unit symbol. π represents the pi. PhaseDifferenceCorrected represents the correction of the phase map predicted by the model to the true phase map in the dataset during the training process, which is calculated by . The hologram H is calculated from the amplitude map and phase map in the dataset using the formula . H gt calculated using the amplitude map and phase map predicted by the Holo-UFormer network also uses for calculation. In the ASM angular spectrum method, m and n are the horizontal and vertical coordinates of the hologram in the form of a discretized matrix. h and w respectively represent the height and width of the physical hologram. ||·||1 represents the L1 norm operation. ||·||2 represents the L2 norm operation. d represents the focusing depth used for propagation. d in represents the depth image pixel matrix in the form of a two-dimensional matrix. The EXP exponential term and the parameter μ are used to adjust the attention weight between the focus and defocus regions. D is the distance between the near and far clipping planes of the frustum. The predicted amplitude map with attention weight and the true amplitude map in the focus stack are propagated to the depth d through the ASM (angular spectrum method) to predict the hologram and the true hologram.
[0040] Compared with the prior art, the present invention has the following advantages:
[0041] First, in the Transformer module of the Holo-UFormer network model proposed in the present invention, three-level wavelet transform and Fourier transform are adopted, and combined with the window attention mechanism to efficiently extract the high and low frequency features of the image. And the multi-domain feed-forward network in the Holo-UFormer network model can ensure the full learning of the network depth and feature scale. It overcomes the problem that the existing technology is difficult to effectively process the depth information of complex scenes, and the generated holograms have deficiencies in detail performance and depth perception. Compared with other methods, the Holo-UFormer network proposed in the present invention presents clearer edge contours, more accurate depth of field structures, and higher amplitude-phase consistency in the generated holograms. Through the simulation experiments of the present invention, the Holo-UFormer proposed in the present invention is significantly superior to the traditional U-Net model in terms of the reconstruction accuracy of the amplitude and phase diagrams, demonstrating stronger image restoration ability and network modeling performance.
[0042] Second, the present invention adopts a two-stage transfer training strategy with double resolutions to train the Holo-UFormer network model: first, the network model is pre-trained at a low resolution to learn the basic feature representation; then the model is transferred to a high resolution for fine-tuning to further optimize the model's ability to capture detailed features. This training strategy can not only accelerate the model convergence, but also effectively improve the generalization performance of the model on high-resolution data. It overcomes the defect that the existing technology only relies on spatial domain information and single-resolution images and cannot make full use of multi-scale information and frequency domain features, enabling the present invention to increase the diversity of the feature space, thereby making full use of multi-scale information and frequency domain features. This method can not only improve the accuracy of the hologram in detail performance, but also enhance its expressiveness in depth perception, making the generated hologram more realistic and three-dimensional. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flowchart of the method of the present invention;
[0044] Figure 2 is a schematic diagram of the Holo-UFormer network structure constructed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0045] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0046] Refer to Figure 1 to further describe the implementation steps of the embodiments of the present invention.
[0047] Step 1, construct a multi-domain Transformer module.
[0048] The constructed multi-domain Transformer module consists of a normalization layer, a window multi-head attention layer, and a multi-domain feed-forward network connected in series in sequence. Among them,
[0049] For the window multi-head attention layer part, given the input feature map where C is the number of channels, H and W are the height and width of the feature map respectively, the window-based multi-head attention mechanism divides the feature map into non-overlapping windows of size M*M. For each window, the features are flattened and transposed, and self-attention calculation is performed within the window. For multi-head attention, assume the model uses k attention heads, and the dimension of each head is d k = C / k.
[0050] Each attention head in the window multi-head attention layer performs self-attention calculation using the following formula:
[0051]
[0052] where Attention(Q, K, V) represents the calculation process of the attention mechanism used in the multi-domain Transformer module. Denote the shape of the feature map when Holo-UFormer enters this module as {batch size, height, width, number of channels}. After window partitioning and flattening, the shape is {total number of windows, window height * window width, number of channels}. Use a fully connected layer to calculate Q (Query), K (Key), and V (Value). k represents the serial number of the attention head, k = 1, 2,..., K, and the dimension of k is d k = C / k, C represents the number of channels, SoftMax represents the operation of normalizing the attention weights so that the sum of the values in each row is 1, the superscript T represents the transpose operation, and B represents the bias of the relative position encoding.
[0053] In the multi-domain feed-forward network, data performs boundary padding of 1 in the spatial domain, with a kernel size of 3 and a stride of 1. In addition to the depthwise separable convolution with an output channel number 4 times the input channel number, a three-level discrete wavelet transform is performed in the frequency domain branch, and then the Fourier transform is performed after convolution. The complete process of performing the Fourier transform on the feature map in the feed-forward network is to first perform a two-dimensional fast Fourier transform (FFT) to map the features from the real space to the complex frequency domain. At this time, each frequency domain point contains real and imaginary part information. Next, a complex convolution is applied in the frequency domain, with a convolution kernel size of 3, a stride of 1, and a boundary padding of 1, to achieve the modeling and enhancement of the frequency domain features. This complex convolution operation processes both the real and imaginary parts simultaneously and effectively captures the frequency domain context through the complex multiplication rule. Finally, the frequency domain features are restored back to the spatial domain through the inverse fast Fourier transform (IFFT) to obtain a feature representation that integrates the frequency domain enhanced information.
[0054] The wavelet transform is a discrete wavelet transform, and its implementation steps are:
[0055] In the first step, the image is decomposed step by step through a low-pass filter and a high-pass filter to extract feature information at different scales;
[0056] In the second step, the input image is decomposed by a low-pass filter to generate multiple sub-bands, including a low-frequency sub-band LL and high-frequency sub-bands {LH, HL, HH}:
[0057]
[0058] Among them, LL j+1 is the low-frequency sub-band, representing the global features of the extracted image. LH j+1 is the horizontal high-frequency component, representing the edge features in the horizontal direction. HL j+1 is the vertical high-frequency component, representing the edge features in the vertical direction. HH j+1 is the diagonal high-frequency component, representing the detailed information in the diagonal direction. I(x, y) represents the pixel value of the feature map at position (i, j), φ(x) represents the scaling function of the low-pass filter used to extract the low-frequency components of the image, and ψ(x) represents the wavelet function of the high-pass filter used to extract the edge and texture information of the image.
[0059] The expressions of the two functions φ(x) and ψ(x) are as follows:
[0060]
[0061] Among them, h k is the coefficient of the low-pass filter, which determines the weighting method of the scaling function. g k is the coefficient of the high-pass filter, which determines the weighting method of the wavelet function. k is the index of the filter, representing the coefficient range of the filter;
[0062] In the embodiments of the present invention, Haar wavelets are used. Therefore, k = (0, 1),
[0063] In the third step, the present invention uses a three-level wavelet transform to perform three iterative decompositions on the feature map, and each iterative decomposition is performed on the low-frequency sub-band LL. Since each wavelet transform will reduce the height and width of the feature map by half, the finally obtained LL is smaller than the original input size. In order to maintain the same scale for subsequent calculations, bilinear interpolation is needed to restore the size. For each level of wavelet high-frequency components HL, LH, and HH, bilinear interpolation is also needed to enlarge them to the original size. After all levels of wavelet transforms, the finally obtained restored low-frequency component LL and the set of restored high-frequency components at each level are obtained.
[0064] In an embodiment of the present invention, the input of Holo-UFormer is a four-channel RGBD image of MIT-CGH-4K. After three-level DWT transformation, its high- and low-frequency feature maps are concatenated with the original input image in the channel dimension, and finally a 44-channel feature map is obtained. Subsequently, after convolution activation, it flows into the multi-domain Transformer module.
[0065] Step 2: Construct the Holo-UFormer network model.
[0066] Refer to Figure 2 for a further description of the result of constructing the Holo-UFormer network model of the present invention.
[0067] The structure of the Holo-UFormer network model is successively: three-level DWT, convolutional layer (padding set to 0, stride set to 1, kernel size set to 3); LeakyReLU activation. Flatten layer, an encoder-bottleneck-decoder structure composed of 9 multi-domain Transformer modules, convolutional layer (padding set to 1, stride set to 1, kernel size set to 3), Tanh activation.
[0068] Step 3: Obtain a dataset containing RGB images, depth images, amplitude images, and phase images, and each type of image contains two samples of high and low resolutions. Perform normalization on the RGB-D images, and the interval range is [-0.5, 0.5].
[0069] Step 4: Train the Holo-UFormer network model using a two-stage transfer training strategy with double resolutions.
[0070] In an embodiment of the present invention, the MIT-CGH-4K dataset is used. This dataset contains data pairs of two resolutions, 192 and 384. The process of training the Holo-UFormer network model using a two-stage transfer training strategy with double resolutions is as follows: First, pre-train the model at a low resolution (192) to learn the basic feature representation; then transfer the model to a high resolution (384) for fine-tuning to further optimize the model's ability to capture detailed features. This training strategy can not only accelerate model convergence but also effectively improve the generalization performance of the model on high-resolution data. Considering that the network structures used in the two-stage training are the same, there are two differences. One is the different resolutions of the input images, and the other is the physical device requirements. At a resolution of 192, the pixel pitch is 0.016 mm. At a resolution of 384, the pixel pitch is 0.008 mm.
[0071] Next, refer to Figure 2 for a further description of the process of training the Holo-UFormer network model using a two-stage transfer training strategy with double resolutions.
[0072] Low-resolution training stage: To make full use of the data and improve the model performance, the low-resolution samples in the training set are input into the Holo-UFormer network model for pre-training the network to learn the basic feature representation. In the network structure of Holo-UFormer, the input image RGBD and the feature map obtained through three-level wavelet transform are concatenated axially in the channel dimension and then flattened to obtain a 44-channel feature map with the same height and width as the input. After padding with 0 at the boundary, convolution with a kernel size of 3, a stride of 1, and the result after the LeakyReLU activation function, the height and width of the flattened feature map are flattened to one dimension while the number of channels remains unchanged. This format can be denoted as B, H*W, C. This result is fed into the multi-domain Transformer module, which internally consists of layer normalization, window multi-head attention, layer normalization, and multi-domain feed-forward network in sequence. There are two skip connections using channel-wise abstract concatenation to concatenate the initial input and the feature map processed by the attention mechanism to obtain a new feature map for subsequent layer normalization, and then concatenate it axially with the result of the multi-domain feed-forward network. This result is used as the input for the next iteration of the same process within the module. In the feed-forward network, after adjusting the data in the B, H*W, C format to the B, H, W, C format, in addition to performing depthwise separable convolution with a padding of 1, a kernel size of 3, a stride of 1, and an output channel number 4 times the input channel number in the spatial domain, there is also a frequency domain branch that includes performing three-level wavelet transform, convolution, Fourier transform, and convolution. Figure 1 The downsampling, upsampling, and magnification factors in [reference] are all 2, and all skip connections adopt the method of channel-wise concatenation. In the low-resolution training stage, after the model converges and is verified by the validation set, the embodiments of the present invention select the model with the best performance based on the evaluation metrics of peak signal-to-noise ratio and structural similarity index.
[0073] Then, the high-resolution samples in the training set are input into the pre-trained Holo-UFormer network model for secondary training, iteratively updating the network parameters until the total loss function of the network converges, obtaining a trained network model. The second training of the model is to fine-tune the pre-trained Holo-UFormer network model at high resolution to further optimize the model's ability to capture detailed features. This training strategy can not only accelerate the model convergence but also effectively improve the generalization performance of the model on high-resolution data.
[0074] Since the network structures used in the two-stage training are the same, due to the different resolutions of the input images, under the requirements of actual physical devices, it is necessary to ensure that the low-resolution * pixel pitch = high-resolution * pixel pitch.
[0075] During the two-stage training process, the same loss function is adopted. The following is the formula for the total loss function, which compares the L2 loss between the predicted hologram and the target hologram, the comparison after propagating different depths, the amplitude map comparison, and the phase map comparison. Generally speaking, it comprehensively considers all stages that the hologram can present.
[0076] The total loss function of the network is as follows:
[0077] L = L hologram + L propagation
[0078]
[0079] L propagation = σL amp + εL tv
[0080] L amp = ||AMP focal (H, D, d, d in ) - AMP focal (H gt , D, d, d in )||2
[0081]
[0082] L tv = ||TV(AMP focal (H, D, d, d in )) - TV(AMP focal (H gt , D, d, d in ))||1
[0083] Among them, L represents the total loss function used to train the Holo-UFormer network, L = L hologram + L propagation , L hologram represents the difference between the amplitude map and phase map in the dataset of interest and the amplitude map and phase map predicted by the network, and L propagation represents the difference between the amplitude maps after the hologram is propagated a certain distance, which is used to ensure the accurate display of the hologram at different depths. α, β, σ, and ε are all weight parameters, A and φ respectively represent the predicted results of the amplitude and phase forming the complex predicted hologram, e (·) represents the exponential function with the natural constant e as the base, i represents the imaginary unit symbol, π represents the circumference ratio, PhaseDifferenceCorrected represents the correction of the phase map predicted by the model to the true phase map in the dataset during the training process, which is calculated by and
[0084] The hologram H is calculated from the amplitude map and phase map in the dataset using the formula and H gt The amplitude map and phase map predicted by the Holo-UFormer network are also calculated using In the ASM angular spectrum method, m and n are the horizontal and vertical coordinates of the hologram in the discretized matrix form; h and w represent the height and width of the physical hologram respectively, ||·||1 represents the L1 norm operation, ||·||2 represents the L2 norm operation, d represents the focusing depth used for propagation, and d in represents the pixel value of the depth map provided by the dataset. The EXP exponential term and the parameter μ are used to adjust the attention weight between the in-focus and out-of-focus regions. D is the distance between the near and far clipping planes of the frustum. The predicted amplitude map with attention weight in the focus stack and the true amplitude map are propagated to the depth d through the ASM to predict the hologram and the true hologram.
[0085] Step 5: Use the trained network model to generate a high-fidelity hologram.
[0086] The effect of the present invention can be further demonstrated by the following simulation.
[0087] 1. Simulation experiment conditions.
[0088] The software platform for the simulation experiment of the present invention is: Ubuntu 22.04.5 LTS, CUDA Version: 12.4, python3.9.21, Pytorch 2.5.1
[0089] 2. Simulation content and result analysis.
[0090] The simulation experiment of the present invention is a simulation of using a network model to generate a hologram.
[0091] In the simulation experiment of the present invention, a classical U-Net network model and the Holo-UFormer network model proposed by the present invention are trained respectively, and the same loss function is used for comparative evaluation. The simulation experiment of the present invention is based on the 384×384 resolution test set of the MIT-CGH-4K dataset, and the reconstruction quality of the amplitude map and phase map is analyzed respectively.
[0092] Two indicators, the structural similarity index (SSIM) and the peak signal-to-noise ratio (PSNR), are used for performance evaluation to quantitatively compare the reconstruction quality of the two network models. The results are shown in the following table:
[0093] Network model SSIM PSNR(dB) U-Net 0.9213 30.89 Holo-UFormer 0.9731 35.34
[0094] The comparative results of the simulation experiments of the present invention show that on the MIT-CGH-4K test set, the present invention has achieved leading performance in terms of PSNR and SSIM image quality metrics, with a 4.45 dB improvement in PSNR and an approximately 0.05 improvement in SSIM compared to UNet. This proves that the Holo-UFormer proposed by the present invention is significantly superior to the traditional U-Net model in terms of the reconstruction accuracy of amplitude and phase maps, demonstrating stronger image restoration ability and network modeling performance.
Claims
1. A high-fidelity hologram generation method using multi-domain transformation and multi-resolution training, characterized in that, The steps of the hologram generation method are as follows: Step 1, construct a multi-domain Transformer module; Step 2, construct a Holo-UFormer network model; Step 3, obtain a dataset containing RGB images, depth images, amplitude images, and phase images, and each type of image contains two samples of high and low resolutions; Step 4, adopt a two-stage transfer training strategy with double resolutions to train the Holo-UFormer network model; Step 5, use the trained network model to generate high-fidelity holograms.
2. The hologram generation method according to claim 1, wherein The multi-domain Transformer module described in Step 1 is composed of a first normalization layer, a window multi-head attention layer, a second normalization layer, and a multi-domain feed-forward network connected in series in sequence.
3. The hologram generation method according to claim 2, wherein Each attention head in the window multi-head attention layer performs self-attention calculation using the following formula: Among them, Attention(Q, K, V) represents the calculation process of the attention mechanism used in the multi-domain Transformer module. When entering this module, the shape of the feature map is denoted as {batch size, height, width, number of channels}. After window partitioning and flattening, the shape becomes {total number of windows, window height * window width, number of channels}. The fully connected layer is used to calculate Q (Query), K (Key), and V (Value). k represents the serial number of the attention head, k = 1, 2,..., K, where K represents the total number of attention heads, and the dimension of k is d k = C / k, where C represents the number of channels, SoftMax represents the operation of normalizing the attention weights so that the sum of the values in each row is 1. The superscript T represents the transpose operation, and B represents the bias of the relative position encoding.
4. The hologram generation method according to claim 2, wherein In addition to performing depthwise separable convolution with a boundary padding of 1, a kernel size of 3, a stride of 1, and an output channel number 4 times the input channel number on the data in the spatial domain in the multi-domain feed-forward network, a three-level wavelet transform will be performed on the frequency domain branch, and then a Fourier transform will be performed after convolution.
5. The hologram generation method according to claim 4, wherein The three-level wavelet transform is a discrete wavelet transform, and its steps are as follows: The first step is to gradually decompose the feature map through a low-pass filter and a high-pass filter to extract feature information at different scales; The second step is to decompose the feature map through a low-pass filter to generate multiple sub-bands, including a low-frequency sub-band LL and high-frequency sub-bands {LH, HL, HH}; Among them, LL j+1 is the low-frequency subband, representing the global features of the extracted image. LH j+1 is the horizontal high-frequency component, representing the edge features in the horizontal direction. HL j+1 is the vertical high-frequency component, representing the edge features in the vertical direction. HH j+1 is the diagonal high-frequency component, representing the detailed information in the diagonal direction. I(x, y) represents the pixel value of the feature map at position (i, j). φ(x) represents the scaling function of the low-pass filter used to extract the low-frequency components of the image, and ψ(x) represents the wavelet function of the high-pass filter used to extract the edge and texture information of the image. The expressions of the two functions φ(x) and ψ(x) are as follows: where h k is the coefficient of the low-pass filter, which determines the weighting method of the scaling function, and g k is the coefficient of the high-pass filter, which determines the weighting method of the wavelet function. k is the index of the filter, representing the coefficient range of the filter; The third step is to continue to perform wavelet transform on the wavelet low-frequency component LL at each level, and a total of three wavelet transforms are used; for the wavelet low-frequency component LL at the first level and the wavelet high-frequency components HL, LH, and HH at each level, after being enlarged to the original size using bilinear interpolation, they are concatenated along the channel axis.
6. The hologram generation method according to claim 1, wherein The Holo-UFormer network model described in step 2 adopts an encoder-bottleneck-decoder architecture; the encoding layer is composed of four cascaded multi-domain Transformer modules, and each module downsamples the input features and extracts multi-scale spatio-frequency domain features; the bottleneck layer is composed of 1 multi-domain Transformer module for deep feature expression and information compression; the decoding layer is composed of four cascaded multi-domain Transformer modules, and each module gradually restores the spatial resolution through upsampling; among them, skip connections are provided between the corresponding levels of the encoding layer and the decoding layer for fusing low-level detail features and high-level semantic features; the dimension of the feature map accepted by the encoder is set to 16, downsampling is implemented using convolution, the convolution parameters are a stride of 2, a kernel size of 4, a padding size of 1, and the output channels are twice the number of input channels, upsampling uses transposed convolution, the parameters are a stride of 2 and a kernel size of 2, and the number of output channels is half the number of input channels; the number of input feature channels accepted by the multi-domain Transformer modules in the encoder are 16, 16*2, 16*4, 16*8 respectively; the corresponding levels of the decoder are the same as those of the encoder; the number of input feature channels accepted by the bottleneck layer is 16*8; the number of multi-head attention heads used in the multi-domain Transformer modules of the Holo-UFormer network model are 1, 2, 4, 8, 16, 16, 8, 4, 2 in sequence.
7. The hologram generation method according to claim 1, wherein The two-stage transfer training strategy described in step 4 means that the RGB map and the depth map in the dataset are concatenated and input into the Holo-UFormer network model, and the Holo-UFormer network model is pre-trained using the low-resolution samples in the dataset, and the amplitude map and the phase map are predicted by the network; the pre-trained Holo-UFormer network model is secondarily trained using the high-resolution samples in the dataset, and the network parameters are iteratively updated until the total loss function of the network converges, and a trained network model is obtained.
8. The hologram generation method according to claim 7, wherein, The total loss function of the network is as follows: L = L hologram + L propagation L propagation = σL amp + εL tv L amp = ||AMP focal (H, D, d, d in ) - AMP focal (H gt , D, d, d in ) ||2 L tv = ||TV(AMP focal (H,D,d,d in )) - TV(AMP focal (H gt ,D,d,d in ))||1 Among them, \(L\) represents the total loss function used for training the Holo-UFormer network, \(L\) hologram represents the differences between the amplitude map and phase map in the dataset and the amplitude map and phase map predicted by the network. \(\alpha\), \(\beta\), \(\sigma\), and \(\varepsilon\) are all weight parameters. \(A\) represents the amplitude map and phase map in the dataset. \(A\) gt , respectively represent the amplitude and phase of the predicted hologram. \(e\) (·) represents the exponential function with the natural constant \(e\) as the base. \(i\) represents the imaginary unit symbol. \(\pi\) represents the pi. PhaseDifferenceCorrected represents that during the training process, the phase map predicted by the model is corrected towards the true phase map in the dataset. It is obtained by It is calculated that the hologram H is derived from the amplitude map and phase map in the dataset using the formula It is calculated that H gt The amplitude map and phase map predicted by the Holo-UFormer network are also calculated using It is calculated that in the ASM angular spectrum method, m and n are the horizontal and vertical coordinates of the hologram in the form of a discretized matrix, h and w respectively represent the height and width of the physical hologram, ||·||1 represents the L1 norm operation, ||·||2 represents the L2 norm operation, d represents the focusing depth used for propagation, d in represents the depth image pixel matrix in the form of a two-dimensional matrix. The EXP exponential term and the parameter μ are used to adjust the attention weight of the network between the focal and defocus regions. D is the distance between the near and far shear planes of the viewing cone. The predicted amplitude map with attention weight and the real amplitude map in the focal stack are propagated to the depth d through ASM (angular spectrum method) to predict the hologram and the real hologram.