A three-dimensional computer hologram reconstruction method and device based on a gaze area optimization
By combining iterative optimization algorithms and end-to-end neural networks, the gaze region of holographic display is dynamically optimized, solving the problem of insufficient gaze region reconstruction quality in holographic display technology and achieving high-quality 3D display and low-latency real-time display effects.
Patent Information
- Application Number
- CN202511563568.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-10-30
AI Technical Summary
Existing computational holographic display technology, with limited spatial light modulator resources, struggles to provide high-quality reconstruction in the gaze area, while wavefront reconstruction errors in non-gaze areas are large, leading to visual interference and failing to match the characteristics of human vision.
By combining iterative optimization algorithms with end-to-end neural networks, display resources are dynamically allocated through gaze tracking to optimize holographic reconstruction of the gaze region. The iterative optimization algorithm optimizes the gaze region in the pre-loaded scene. The holographic generation process includes gaze region segmentation, weight matrix construction, light field propagation calculation, and weighted loss function optimization. The end-to-end neural network achieves high-quality reconstruction of the gaze region through complex amplitude field generation, gaze feature extraction, and holographic encoding.
With limited SLM resources, the hologram reconstruction quality of the gaze area is improved, the 3D display effect is maintained, and low-latency real-time display is achieved, thus enhancing the hologram generation capability.
Smart Images

Figure CN121033292B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computational holography (CGH) technology, specifically relating to a method and apparatus for reconstructing three-dimensional computational holograms based on gaze region optimization. Background Technology
[0002] In human visual perception, over 80% of external information is acquired through vision, and nearly half of the human brain's computational resources are used to process visual information. Current visual information presentation primarily relies on real-world scenes and display systems: the real world provides complete three-dimensional (3D) light field information, while traditional display systems can only present two-dimensional images, lacking depth cues. Holographic display technology, by reconstructing the wavefront information of objects, reproduces a true three-dimensional light field, and is therefore widely considered the ultimate solution for next-generation AR / VR / XR near-eye display devices.
[0003] Existing near-eye 3D display technologies are mainly divided into two categories:
[0004] The first type is stereoscopic display technology based on binocular parallax. This type of technology provides stereoscopic perception by offering images with parallax differences to the left and right eyes, utilizing the brain to fuse them and provide binocular depth cues. However, it has significant drawbacks: first, it cannot provide crucial monocular depth cues such as motion parallax and texture gradients; second, because the display focal plane is fixed, it is incompatible with the ciliary muscle accommodation mechanism of the human eye, leading to the well-known "vergence-accommodation conflict" (VAC), causing visual fatigue and dizziness.
[0005] The second category is computational holography (CGH) display technology. This technology reconstructs the target light field by calculating and controlling the amplitude and phase of light waves. Theoretically, it can provide all monocular and binocular depth cues, making it an ideal path to solve the VAC problem and achieve true 3D display. However, its development is limited by the physical performance of spatial light modulators (SLMs), such as the limited number of pixels, refresh rate, and low diffraction efficiency. With limited SLM pixel resources, it is difficult to generate high-quality holograms in real time, especially reconstructing high-resolution, wide-view 3D scenes, which remains a huge challenge.
[0006] To alleviate the computational burden of computational holography, foveated rendering technology has been introduced. This technology leverages the high resolution of the fovea and low resolution of the peripheral vision in the human visual system (HVS) to perform high-quality rendering only in the user's gaze area, while reducing rendering quality in non-gaze areas to save computational resources. However, existing research on foveated rendering almost entirely focuses on the goal of "accelerating graphics rendering," and its technical methods mostly involve reducing rendering resolution or model detail in non-gaze areas.
[0007] When directly transferring the concept of foveated rendering to the field of computational holography, a completely new and yet-to-be-solved technical problem arises:
[0008] Traditional foveated rendering sacrifices image resolution in non-foveated areas, while the core of holographic display is wavefront reconstruction. Simply allocating less computing resources or SLM pixels to non-foveated areas will lead to increased wavefront reconstruction errors in those areas, resulting not only in image artifacts but also in severe phase noise and speckle. These reconstruction defects do not match the characteristics of human peripheral vision and may even cause uncomfortable visual interference, thus completely violating the original intention of improving the experience by utilizing the characteristics of human vision.
[0009] Therefore, current technology lacks a method that can specifically target the characteristics of holographic wavefront reconstruction, dynamically improve the visual quality of the gaze region with limited SLM resources, and at the same time ensure that the reconstruction effect of the non-gaze region matches the characteristics of human visual perception. Summary of the Invention
[0010] To overcome the shortcomings of existing technologies, this invention provides a three-dimensional computational hologram reconstruction method and apparatus based on gaze region optimization. Combining an iterative optimization algorithm and an end-to-end neural network architecture, it dynamically allocates display resources through gaze tracking and focuses on optimizing the human eye's gaze region to achieve high dynamic range holographic near-eye display. The iterative optimization algorithm is suitable for pre-loading scenarios, while the end-to-end neural network architecture is suitable for real-time image calculation scenarios.
[0011] The technical solution adopted by this invention to solve its technical problem is:
[0012] A 3D computational hologram reconstruction method based on gaze region optimization includes iteratively optimized gaze region hologram generation and gaze region hologram generation using an end-to-end neural network.
[0013] The iterative optimization process for generating a gaze region hologram is as follows: First, gaze region segmentation and weight matrix construction are performed, then light field propagation calculation is performed, followed by multi-depth target image generation and weighted loss function construction, gradient optimization and parameter update are performed, and the optimization process is stopped when the preset number of optimizations is reached or the reconstructed image quality meets the set requirements; finally, hologram loading is performed.
[0014] The gaze region hologram generation process of the end-to-end neural network is as follows: First, a complex amplitude field generation module generates an intermediate plane complex amplitude light field from the initial RGBD image; then, an eye image is acquired using a near-eye camera to accurately extract the coordinates and related features of the user's gaze point; subsequently, the intermediate plane complex amplitude light field is propagated through the light field to obtain an initial hologram, which is then fused with gaze features to enhance the weight of the gaze region; finally, a hologram encoding module directly outputs a pure phase hologram that can modulate the phase of the incident light field. Simultaneously, to ensure neural network performance, differentiated training strategies are designed for different functional modules. The complex amplitude field generation module uses unsupervised training to autonomously learn the laws of multi-depth light fields; the gaze estimation module uses semi-supervised two-stage training to separate and strengthen gaze features; the feature fusion and hologram encoding modules use joint training to achieve weighted optimization of the pure phase hologram; finally, end-to-end joint training improves the overall end-to-end hologram optimization capability of the system.
[0015] Furthermore, the iteratively optimized gaze region hologram generation process includes the following sub-steps:
[0016] Step 1.1, Gaze region division: Determine the size and resolution of the holographic reconstructed image based on the physical parameters of the spatial light modulator, use the gaze estimation algorithm to obtain the user's gaze point position in real time, delineate a specific range centered on this point as the gaze region, and define the remaining part as the non-gaze region;
[0017] Step 1.2, Weight Matrix Construction: Generate a weight matrix of the same size as the hologram, assigning higher weight parameters to the gaze region and lower weight parameters to the non-gaze region, thus creating differentiated optimization priorities. In the initialization phase, a random phase distribution is used as the initial input to the hologram.
[0018] Step 1.3, Calculation of light field propagation: Based on the theory of angular spectrum diffraction, calculate the complex amplitude light field distribution of the hologram from the spatial light modulator plane to multiple target depth planes to simulate the propagation process of the light field in space.
[0019] The iteratively optimized gaze region hologram generation process also includes the following sub-steps:
[0020] Step 1.4, Multi-depth target image generation: Preprocess the target image containing depth information, and generate a set of target images at different depth positions through depth masking or out-of-focus rendering, so as to provide a reference standard for subsequent optimization;
[0021] Step 1.5: Construction of weighted loss function: Extract amplitude information from the reconstructed light field and convert it into a reconstructed image. Input the target image and the reconstructed image into the loss function, and introduce a weight matrix to focus on constraining the reconstruction error of the gaze region, forming a region-weighted total loss function.
[0022] The iteratively optimized gaze region hologram generation process also includes the following sub-steps:
[0023] Step 1.6, Gradient Optimization and Parameter Update: Using the phase distribution of the hologram as the optimization variable, the gradient of the loss function with respect to the hologram is calculated through the backpropagation algorithm. The phase parameters of the hologram are iteratively updated to gradually reduce the difference between the reconstructed image and the target image.
[0024] Step 1.7, Iteration Termination Condition: The optimization process stops when the preset number of optimizations is reached, or when the reconstructed image quality meets the set requirements;
[0025] Step 1.8, Hologram Loading: The optimized hologram is loaded into the spatial light modulator, and a three-dimensional light field containing wavefront information is reconstructed in the target space by modulating the incident light field, so as to realize the display of stereoscopic images with real depth clues.
[0026] Furthermore, the process of generating the gaze region hologram of the end-to-end neural network includes the following sub-steps:
[0027] Step 2.1: Multi-depth feature preprocessing: Perform depth mask processing or out-of-focus rendering on the RGBD target image to generate target images at different depth positions, and extract the complex amplitude light field distribution of the intermediate plane.
[0028] Step 2.2: Gait feature extraction. Real-time images of the user's eyes are captured using a near-eye camera, and eye features are analyzed to output accurate coordinates of the gaze point and related feature parameters.
[0029] Step 2.3: Light field propagation and feature fusion. The complex amplitude of the intermediate plane is propagated to the spatial light modulator plane through the angular spectrum diffraction algorithm to obtain the initial complex amplitude hologram; combined with the gaze characteristics, the gaze region is weighted and enhanced to generate fused features containing visual attention information.
[0030] Step 2.4: Generating pure phase holograms. The fused features are input into the neural network model, and the neural network model outputs pure phase holograms adapted to the spatial light modulator, thereby realizing phase modulation encoding of the incident light field.
[0031] The gaze region hologram generation process of the end-to-end neural network includes the following modules:
[0032] The gaze estimation module is used to acquire eye images in real time through a near-eye camera, extract the user's gaze characteristics using visual processing algorithms, and output coordinate information including the location of the gaze point.
[0033] The complex amplitude field generation module is used to generate the complex amplitude light field distribution of the intermediate target plane through a depth perception network, taking the RGBD target image as input, and capturing the multi-depth layer features of the scene;
[0034] The feature fusion module is used to spatiotemporally align complex amplitude light field information with gaze characteristics, enhance the feature representation of the gaze region through an attention mechanism, and generate enhanced fused features.
[0035] The hologram encoding module is used to map fused features into pure phase holograms that can be loaded by a spatial light modulator, realizing end-to-end mapping from scene input to hologram output.
[0036] Preferably, the neural network training includes the following sub-steps:
[0037] Step 3.1, Multi-depth target image preprocessing: During the neural network training process, the RGBD input image is first subjected to depth mask or defocus rendering to generate a set of target images at different depth positions, thereby providing multi-plane reference data for the training of subsequent modules and building a basic data support system;
[0038] Step 3.2, Unsupervised Training of the Complex Amplitude Field Generation Module: The RGBD dataset is then input into this module, causing it to output the complex amplitude distribution of the intermediate plane. Subsequently, the reconstructed light field distribution of each target plane is calculated using an angular spectral diffraction algorithm. After obtaining the reconstructed image, its amplitude information is extracted and compared with the target image to construct the loss function. Finally, the stochastic gradient descent algorithm is used to optimize the network parameters, realizing the unsupervised learning process and enabling the module to autonomously learn the features and patterns in the data.
[0039] Step 3.3, Semi-supervised Training of the Gaze Estimation Module: For the gaze estimation module, a two-stage training strategy is adopted to separate and enhance gaze features. In the feature separation stage, unlabeled eye images are input into the module to separate gaze features and appearance features. Then, the gaze and appearance features of different samples are recombined by the appearance restoration module. A loss function is constructed based on the difference between the restored image and the original image to optimize the feature separation capability, enabling the module to extract gaze-related features more accurately. In the gaze prediction stage, labeled eye images are input to extract gaze features and predict gaze coordinates. Then, a loss function is constructed based on the error between the predicted coordinates and the true label to further enhance the expression of gaze features and improve the accuracy of the module's gaze prediction.
[0040] The neural network training also includes the following sub-steps:
[0041] Step 3.4, Joint Training of Feature Fusion and Hologram Encoding Modules: In the joint training of the feature fusion and hologram encoding modules, the complex amplitude field is first propagated to the spatial light modulator plane to generate an initial hologram. Then, gaze features are fused, and the gaze region is weighted and enhanced through an attention mechanism. To better optimize this process, high-weight gaze regions and low-weight non-gaze regions are divided according to the gaze point. The weighted error between the reconstructed image and the target image is calculated, and a region weighted loss function is constructed. The parameters of the feature fusion module and the hologram encoding module are then jointly optimized, enabling the two modules to work collaboratively and improve the quality of hologram generation.
[0042] Step 3.5, End-to-end Joint Training: Finally, the labeled eye image and RGBD data are simultaneously input into the network. The training process of step 2.4.4 in the joint training of the feature fusion and hologram encoding modules is repeated according to the complete data flow. The parameters of all modules are optimized synchronously. Through this global training, the global optimum from input to output is achieved, enabling the entire neural network system to generate gaze region holograms efficiently and accurately.
[0043] A 3D computational hologram reconstruction device based on gaze region optimization includes an iteratively optimized gaze region hologram generation part and an end-to-end neural network gaze region hologram generation part.
[0044] The iteratively optimized gaze region hologram generation part includes the following modules:
[0045] The gaze region segmentation module is used to obtain the coordinates of the gaze point, segment the gaze region and the non-gaze region, and construct a weight matrix;
[0046] The multi-depth generation module is used to render 3D information into multi-depth target images using holographic optimization algorithms;
[0047] The light field propagation module is used to propagate the hologram loaded on the SLM plane to the target plane at various distances to obtain the reconstructed image of the hologram on the target plane;
[0048] The gradient optimization module is used to calculate the loss value between the hologram reconstructed image and the target image, and iteratively optimizes the hologram through the gradient optimization algorithm;
[0049] The gaze region hologram generation part of the end-to-end neural network includes the following modules:
[0050] The gaze estimation module is used to acquire eye images in real time through a near-eye camera, extract the user's gaze characteristics using visual processing algorithms, and output coordinate information including the location of the gaze point.
[0051] The complex amplitude field generation module is used to generate the complex amplitude light field distribution of the intermediate target plane through a depth perception network, taking the RGBD target image as input, and capturing the multi-depth layer features of the scene;
[0052] The feature fusion module is used to spatiotemporally align complex amplitude light field information with gaze characteristics, enhance the feature representation of the gaze region through an attention mechanism, and generate enhanced fused features.
[0053] The hologram encoding module is used to map fused features into pure phase holograms that can be loaded by a spatial light modulator, realizing end-to-end mapping from scene input to hologram output.
[0054] The beneficial effects of this invention are mainly reflected in:
[0055] 1. Higher quality hologram reconstruction images: By dynamically dividing the gaze region and non-gaze region through gaze estimation method, the pixel resources are allocated differently, which improves the hologram reconstruction quality in the visually sensitive area of the human eye, while maintaining the layered three-dimensional display effect;
[0056] 2. Faster real-time hologram reconstruction: An end-to-end neural network model is proposed to effectively integrate RGBD depth information and eye-tracking data to construct a hologram generation model that includes depth cues and visual attention. It has higher parallel computing capabilities and can achieve low-latency real-time display.
[0057] 3. Efficient training method: A phased training and end-to-end joint optimization method is proposed to enable the neural network model to maintain the ability of both the gaze estimation model and the hologram generation model, thus achieving better hologram generation capabilities. Attached Figure Description
[0058] Figure 1 This is a flowchart of the iteratively optimized gaze region hologram generation method of the present invention;
[0059] Figure 2 This is a flowchart of the end-to-end neural network gaze region hologram generation method of the present invention;
[0060] Figure 3 This is a flowchart of the training process of the end-to-end neural network gaze region hologram generation method of the present invention;
[0061] Figure 4 This is a schematic diagram of the optical path of holographic display in an embodiment of the present invention, wherein 1 is a laser source, 2 is a collimator, 3 is a polarizer, 4 is a beam splitter, 5 is a spatial light modulator, 6 is a 4f system, 7 is an aperture stop, 8 is an optical waveguide, 9 is an imaging camera, and 10 is a virtual image of the holographic reconstructed image.
[0062] Figure 5This is a comparison of simulation results of traditional holographic generation methods based on depth mask targets. The white squares represent the user's gaze area, and the numbers in the lower right corner of the box are the peak signal-to-noise ratio (PSNR) evaluation metrics of the images.
[0063] Figure 6 This is a comparison of simulation results based on the depth mask target of the present invention. The white square represents the user's gaze area, and the number in the lower right corner of the box is the PSNR evaluation index of the image.
[0064] Figure 7 This is a comparison of simulation results of traditional holographic generation methods based on defocused rendering of the target. The white square represents the user's gaze area, and the number in the lower right corner of the box is the PSNR evaluation index of the image.
[0065] Figure 8 This is a comparison chart of simulation results of the target rendering based on the present invention. The white square represents the user's gaze area, and the number in the lower right corner of the box is the PSNR evaluation index of the image. Detailed Implementation
[0066] The present invention will now be further described with reference to the accompanying drawings.
[0067] Reference Figures 1-8 A three-dimensional computational hologram reconstruction method based on gaze region optimization includes iterative optimization of gaze region hologram generation and end-to-end neural network-based gaze region hologram generation;
[0068] The iterative optimization process for generating a gaze region hologram is as follows: First, gaze region segmentation and weight matrix construction are performed, then light field propagation calculation is performed, followed by multi-depth target image generation and weighted loss function construction, gradient optimization and parameter update are performed, and the optimization process is stopped when the preset number of optimizations is reached or the reconstructed image quality meets the set requirements; finally, hologram loading is performed.
[0069] The iteratively optimized gaze region hologram generation process includes the following sub-steps:
[0070] Step 1.1, Fixation of the Gazing Region: A corneal reflection method based on a near-eye camera is used for gaze estimation. The camera resolution is 1280×720, and the frame rate is 60Hz. The camera captures the position of the user's pupil center and the corneal reflection point in real time, calculating the coordinates of the gaze point. The gaze region occupies 5%-15% of the hologram image and is dynamically adjusted according to the spatial light modulator (SLM) resolution. The SLM resolution is 1920×1080, and the radius of the gaze region is set to 100 pixels. The non-gazing region is the remaining portion.
[0071] Step 1.2, Weight Matrix Construction: Generate a two-dimensional weight matrix of the same size as the hologram. The gaze region is weighted using a Gaussian distribution:
[0072] ;
[0073] in, Here are the coordinates of the gaze point. , To define the radius of the fixation region, the center of the fixation region is weighted at 1, and the edges are smoothly transitioned; the initial hologram uses a random phase distribution, with a range of [missing information]. ;
[0074] Step 1.3, Calculation of light field propagation: Based on the angular spectrum diffraction theory, the propagation of the complex amplitude field is calculated using Fast Fourier Transform (FFT). Let the complex amplitude of the SLM plane be... (Pure phase SLM), where, The phase distribution on the SLM is given. The transfer function is: The propagation of the light field from the SLM plane to the target depth plane is calculated using FFT and inverse FFT. For spatial frequency, The target depth plane distance, For the wavelength, 532nm was selected.
[0075] Step 1.4, Multi-depth target image generation: The input RGBD image (resolution 1920×1080) is rendered out of focus and divided into 6 layers according to the depth value, representing the image observed when the human eye focuses on 6 depth planes. Each layer is generated using Gaussian blur to simulate a natural out-of-focus effect.
[0076] Step 1.5: Construction of the weighted loss function. Amplitude information is extracted from the reconstructed light field and converted into a reconstructed image. The target image and the reconstructed image are then fed into the loss function, and a weight matrix is introduced to specifically constrain the reconstruction error of the gaze region, forming a region-weighted total loss function:
[0077] ;
[0078] in, To reconstruct the amplitude, For the target amplitude, Total number of pixels This is a weight matrix, where the weight of the fixation region is 5 times that of the non-fixation region, thereby strengthening the fixation region constraint.
[0079] Step 1.6, Gradient Optimization and Parameter Update: The Adam optimizer is used, with the learning rate initialized to 0.01. The hologram phase is... As an optimization variable, the gradient is calculated through automatic differentiation, that is, the gradient of the loss function with respect to the hologram is calculated through the backpropagation algorithm. Iteratively update the phase parameters of the hologram, updating the phase in each iteration: ,in, For learning rate, This represents the number of iterations.
[0080] Step 1.7, Iteration Termination Condition: Set the maximum number of iterations to 1000, or stop optimization when the Structural Similarity Index (SSIM) of the gaze region is ≥0.8 and the SSIM of the non-gaze region is ≥0.6;
[0081] Step 1.8, Hologram Loading: The optimized hologram is loaded into a pure phase spatial light modulator for 3D reconstruction. Zero-order diffraction and conjugate images are eliminated by 4f system filtering, and the reconstructed image is received at the target depth.
[0082] The gaze region hologram generation process of the end-to-end neural network is as follows: First, a complex amplitude field generation module generates an intermediate plane complex amplitude light field from the initial RGBD image; then, an eye image is acquired using a near-eye camera to accurately extract the coordinates and related features of the user's gaze point; subsequently, the intermediate plane complex amplitude light field is propagated through the light field to obtain an initial hologram, which is then fused with gaze features to enhance the weight of the gaze region; finally, a hologram encoding module directly outputs a pure phase hologram that can modulate the phase of the incident light field. Simultaneously, to ensure neural network performance, differentiated training strategies are designed for different functional modules. The complex amplitude field generation module uses unsupervised training to autonomously learn the laws of multi-depth light fields; the gaze estimation module uses semi-supervised two-stage training to separate and strengthen gaze features; the feature fusion and hologram encoding modules use joint training to achieve weighted optimization of the pure phase hologram; finally, end-to-end joint training improves the overall end-to-end hologram optimization capability of the system.
[0083] The gaze region hologram generation process of the end-to-end neural network includes the following sub-steps:
[0084] Step 2.1: Multi-depth feature preprocessing. The RGBD target image undergoes depth masking or out-of-focus rendering, generating a complex amplitude light field in the intermediate plane. The 1920×1152 resolution RGBD input image is divided into 6 layers based on depth values (corresponding to the 6 depth planes of human eye focus). Depth masking involves dividing the image into 6 layers based on depth values and then using 6 binary masks to segment the target image into 6 focus regions, using only the focus regions as the target image. This method yields high-quality focus region images, but the out-of-focus regions lack a natural blur effect. Out-of-focus rendering uses Gaussian blur to render a focus blur effect on each of the 6 depth planes, using all of them as target planes. This allows the out-of-focus regions to have a natural blur effect, but the quality of the focus region images is relatively low. The rendered target image is then input into a neural network model to obtain the complex amplitude light field distribution in the intermediate plane.
[0085] Step 2.2: Gaze Feature Extraction. Real-time binocular images of the user are captured using a near-eye camera. Eye features are analyzed to output accurate gaze coordinates and related feature parameters. The input binocular images are resized to 320×320 and merged into a 2×320×320 three-dimensional matrix before being input into the neural network model. The neural network model extracts a 135×20×20 feature map from the binocular images through convolutional neural network downsampling. This feature map is then further compressed using a spatial attention module to obtain a 135×1×1 intermediate layer as the gaze feature. Finally, a fully connected layer is connected to output the predicted gaze coordinates.
[0086] Step 2.3: Fusion of Complex Amplitude Light Field and Gaze Attention Features. The complex amplitude light field of the intermediate plane is propagated to the SLM plane using an angular spectrum diffraction algorithm to obtain the SLM complex amplitude light field. A feature extraction neural network is used to obtain a holographic feature with a shape of 2048×15×9. Then, the gaze attention feature is adjusted to a shape of 1×15×9, and the gaze attention feature and the holographic feature are fused by feature stitching to obtain a fused feature with a shape of 2049×15×9.
[0087] Step 2.4: The hologram encoding module inputs the fused features and outputs a pure phase hologram adapted to the spatial light modulator through a neural network model. The fused features, with a shape of 2049×15×9, are upsampled through the neural network to obtain a 1×1920×1152 output matrix. The parameters in the matrix are multiplied by... The pure phase hologram distribution is obtained, and finally the output is constrained to a certain value using phase folding. This yields a hologram distribution suitable for SLM, ultimately enabling phase modulation coding of the incident light field;
[0088] The gaze region hologram generation process of the end-to-end neural network includes the following modules:
[0089] The gaze estimation module activates the near-eye infrared camera to acquire real-time images of the user's eyes at a resolution of 640×480. The acquired images are then resized to 320×320 and input into the gaze estimation module. This module consists of a convolutional neural network, a spatial attention module, and a fully connected layer. The convolutional neural network and spatial attention module form the backbone network to extract image features, while the fully connected layer outputs the gaze coordinates. Finally, it returns the image feature vector and gaze coordinate information.
[0090] The complex amplitude field generation module selects an RGBD target image containing an indoor scene from the publicly available NYU Depth V2 dataset. The RGB channel resolution of this image is 640×480, and the depth channel resolution corresponds to this. The RGBD image is resized to 1920×1152, and a Gaussian filter with a kernel size of 5×5 and a standard deviation of 1.0 is used to render the depth channel out of focus, simulating blur effects at different depths. This generates target images at six depth positions: 0.3m, 2m, 4m, 8m, 30m, and 100m from the camera. The RGBD image is then input into the complex amplitude field generation module, which is based on the U-Net architecture and contains 7 downsampling layers and 7 upsampling layers. Each layer has a 3×3 kernel size and a stride of 2. The module learns image features and outputs the complex amplitude light field distribution of the central target plane (4 meters from the camera). The complex amplitude light field data dimension is 1920×1152×2 (real and imaginary parts).
[0091] The feature fusion module propagates the complex amplitude light field distribution of the intermediate plane to the spatial light modulator plane using an angular spectrum diffraction algorithm. With a wavelength of 532nm and a focal length of 50mm for the holographic display system objective lens, the corresponding propagation distance across the intermediate plane is 50.31mm. Diffraction calculations yield the initial complex amplitude hologram. The gaze-attention features and the initial complex amplitude hologram are then input into the feature fusion module to generate a fused feature containing visual attention information. The fused feature data dimension is then compared with the initial complex amplitude hologram. Figure 1 To;
[0092] The hologram encoding module takes the fused features as input and, based on a convolutional neural network architecture, contains seven upsampling layers. The module learns to map the fused features into pure phase holograms, outputting pure phase hologram data within a certain range. The resolution is matched with the spatial light modulator. The generated pure phase hologram is loaded onto the Holoeye LUNA spatial light modulator to complete the phase modulation encoding of the incident light field, thereby realizing the generation and display of the hologram of the gaze region.
[0093] The neural network training includes the following sub-steps:
[0094] Step 3.1: Select 10,000 RGBD images from the SUN RGB-D dataset as training data. Perform depth masking on the depth channel of each image to generate target images at six depth positions: 0.3m, 2m, 4m, 8m, 30m, and 100m from the camera, providing multi-plane reference data for subsequent training.
[0095] Step 3.2: Unsupervised Training of the Complex Amplitude Field Generation Module: The preprocessed RGBD dataset is input into the complex amplitude field generation module, which outputs the complex amplitude distribution of the intermediate plane. The complex amplitude distribution is propagated to multiple target planes using an angular spectral diffraction algorithm, and the reconstructed light field distribution of each plane is calculated. The amplitude information of the reconstructed image is extracted and compared pixel-by-pixel with the RGB channels of the original target image. The mean squared error (MSE) is used to construct the loss function. ,in To reconstruct image pixel values, For the target image pixel values, The total number of pixels. The Adam optimizer was used with a learning rate of 0.001, trained for 100 epochs, and the network parameters were optimized to enable the module to accurately generate complex amplitude light field distributions.
[0096] Step 3.3, Semi-supervised Training of the Gaze Estimation Module: For the gaze estimation module, a two-stage training strategy is adopted to separate and enhance gaze features. In the feature separation stage, 5000 unlabeled eye images are selected from the GazeCapture dataset and input into the gaze estimation module. The module separates gaze features and appearance features, and the gaze and appearance features of different samples are reconstructed by the appearance restoration module. The structural similarity index (SSIM) between the restored image and the original image is calculated as the loss function. The stochastic gradient descent algorithm with a learning rate of 0.0001 was used for 50 epochs to optimize feature separation capabilities. In the gaze prediction stage, 3000 labeled eye images were selected from the MPIIGaze dataset. The input module extracted gaze features and predicted gaze coordinates. The loss function was constructed using the Euclidean distance between the predicted coordinates and the ground truth labels. To predict coordinates, For the actual coordinates, This represents the number of samples. Continue training for 50 more epochs to enhance the expression of gaze features and improve the accuracy of gaze prediction.
[0097] Step 3.4, Joint Training of Feature Fusion and Hologram Encoding Modules: The complex amplitude field output by the complex amplitude field generation module is propagated to the spatial light modulator plane to generate an initial hologram; gaze features are fused, and an attention mechanism is used to divide the gaze region (a circular region with a radius of 100 pixels centered on the gaze point) and the non-gaze region with low weights based on the gaze point; the weighted error between the reconstructed image and the target image in different regions is calculated, and a region-weighted loss function is constructed:
[0098] ;
[0099] in, , The Adam optimizer was used with a learning rate of 0.0001, and the two modules were jointly trained for 80 epochs to optimize their parameters.
[0100] Step 3.5, End-to-End Joint Training: Labeled eye images (from the Gaze360 dataset) and RGBD data (from the ScanNet dataset) are simultaneously input into the network. The network is then trained following the complete data flow repetition feature fusion and hologram encoding module joint training steps. Using the aforementioned region-weighted loss function, the Adam optimizer is employed with a learning rate of 0.0001 for 120 epochs. The parameters of all modules are optimized simultaneously to achieve global optimum from input to output, enabling the entire neural network system to efficiently and accurately generate gaze region holograms.
[0101] A 3D computational hologram reconstruction device based on gaze region optimization includes an iteratively optimized gaze region hologram generation part and an end-to-end neural network gaze region hologram generation part.
[0102] The iteratively optimized gaze region hologram generation part includes the following modules:
[0103] The gaze region segmentation module is used to obtain the coordinates of the gaze point, segment the gaze region and the non-gaze region, and construct a weight matrix;
[0104] The multi-depth generation module is used to render 3D information into multi-depth target images using holographic optimization algorithms;
[0105] The light field propagation module is used to propagate the hologram loaded on the SLM plane to the target plane at various distances to obtain the reconstructed image of the hologram on the target plane;
[0106] The gradient optimization module is used to calculate the loss value between the hologram reconstructed image and the target image, and iteratively optimizes the hologram through the gradient optimization algorithm;
[0107] The gaze region hologram generation part of the end-to-end neural network includes the following modules:
[0108] The gaze estimation module is used to acquire eye images in real time through a near-eye camera, extract the user's gaze characteristics using visual processing algorithms, and output coordinate information including the location of the gaze point.
[0109] The complex amplitude field generation module is used to generate the complex amplitude light field distribution of the intermediate target plane through a depth perception network, taking the RGBD target image as input, and capturing the multi-depth layer features of the scene;
[0110] The feature fusion module is used to spatiotemporally align complex amplitude light field information with gaze characteristics, enhance the feature representation of the gaze region through an attention mechanism, and generate enhanced fused features.
[0111] The hologram encoding module is used to map fused features into pure phase holograms that can be loaded by a spatial light modulator, realizing end-to-end mapping from scene input to hologram output.
[0112] In this embodiment, by optimizing the method and neural network model, and constructing the weight matrix, the limited SLM display resources are efficiently allocated, significantly improving the reconstructed image quality of the gaze region under the same conditions. Figure 5 and Figure 6 As shown, simulation results comparing the traditional holographic generation method and the method of this invention based on a depth mask target are presented. The white squares represent the user's gaze region, and the numbers in the lower right corner of the boxes are the PSNR evaluation metrics. It can be seen that within the gaze region, the PSNR of the traditional method is 25.35, while the PSNR of the method of this invention is 33.02. The method of this invention significantly improves the reconstructed image quality within the gaze region. Figure 7 and Figure 8 As shown, simulation results comparing the traditional holographic generation method and the method of this invention based on defocused rendering targets are presented. Defocused rendering targets can better reconstruct the natural blurring effect of the defocused region of the hologram; however, due to the limited pixel space of SLM, the image quality of the focused region is significantly reduced. Specifically, the PSNR of the traditional method within the gaze region is 23.27, while the PSNR of the method of this invention is 26.64. It can be seen that the method of this invention can improve the image quality of the gaze region to a relatively acceptable level.
[0113] The embodiments described in this specification are merely examples of implementations of the inventive concept and are for illustrative purposes only. The scope of protection of this invention should not be considered limited to the specific forms described in these embodiments; rather, it extends to those skilled in the art based on the inventive concept.
Claims
1. A method for reconstructing a three-dimensional computational hologram based on gaze region optimization, characterized in that, The three-dimensional computational hologram reconstruction method includes iteratively optimized gaze region hologram generation and end-to-end neural network gaze region hologram generation. The iterative optimization process for generating a gaze region hologram is as follows: First, gaze region segmentation and weight matrix construction are performed, then light field propagation calculation is performed, followed by multi-depth target image generation and weighted loss function construction, gradient optimization and parameter update are performed, and the optimization process is stopped when the preset number of optimizations is reached or the reconstructed image quality meets the set requirements; finally, hologram loading is performed. The process of generating the gaze region hologram of the end-to-end neural network is as follows: First, the complex amplitude field generation module generates an intermediate plane complex amplitude light field from the initial RGBD image; then, the near-eye camera acquires eye images and accurately extracts the coordinates of the user's gaze point and related features. The intermediate plane complex amplitude light field is then propagated through the light field to obtain an initial hologram, which is then fused with the gaze feature to enhance the weight of the gaze region. Finally, through the hologram encoding module, a pure phase hologram that can be phase modulated by the incident light field is directly output.
2. The three-dimensional computational hologram reconstruction method based on gaze region optimization as described in claim 1, characterized in that, The iteratively optimized gaze region hologram generation process includes the following sub-steps: Step 1.1, Gaze region division: Determine the size and resolution of the holographic reconstructed image based on the physical parameters of the spatial light modulator, use the gaze estimation algorithm to obtain the user's gaze point position in real time, delineate a specific range centered on this point as the gaze region, and define the remaining part as the non-gaze region; Step 1.2, Weight Matrix Construction: Generate a weight matrix of the same size as the hologram, assign high weight parameters to the gaze region and low weight parameters to the non-gaze region to form differentiated optimization priorities. In the initialization stage, a random phase distribution is used as the initial input of the hologram. Step 1.3, Calculation of light field propagation: Based on the theory of angular spectrum diffraction, calculate the complex amplitude light field distribution of the hologram from the spatial light modulator plane to multiple target depth planes to simulate the propagation process of the light field in space.
3. The three-dimensional computational hologram reconstruction method based on gaze region optimization as described in claim 2, characterized in that, The iteratively optimized gaze region hologram generation process also includes the following sub-steps: Step 1.4, Multi-depth target image generation: Preprocess the target image containing depth information, and generate a set of target images at different depth positions through depth mask or out-of-focus rendering, so as to provide a reference standard for subsequent optimization; Step 1.5: Construction of weighted loss function: Extract amplitude information from the reconstructed light field and convert it into a reconstructed image. Input the target image and the reconstructed image into the loss function, and introduce a weight matrix to focus on constraining the reconstruction error of the gaze region, forming a region-weighted total loss function.
4. The three-dimensional computational hologram reconstruction method based on gaze region optimization as described in claim 3, characterized in that, The iteratively optimized gaze region hologram generation process also includes the following sub-steps: Step 1.6, Gradient Optimization and Parameter Update: Using the phase distribution of the hologram as the optimization variable, the gradient of the loss function with respect to the hologram is calculated through the backpropagation algorithm. The phase parameters of the hologram are iteratively updated to gradually reduce the difference between the reconstructed image and the target image. Step 1.7, Iteration Termination Condition: The optimization process stops when the preset number of optimizations is reached, or when the reconstructed image quality meets the set requirements; Step 1.8, Hologram Loading: The optimized hologram is loaded into the spatial light modulator, and a three-dimensional light field containing wavefront information is reconstructed in the target space by modulating the incident light field, so as to realize the display of stereoscopic images with real depth clues.
5. A three-dimensional computational hologram reconstruction method based on gaze region optimization as described in any one of claims 1 to 4, characterized in that, The gaze region hologram generation process of the end-to-end neural network includes the following sub-steps: Step 2.1: Multi-depth feature preprocessing: Perform depth mask processing or out-of-focus rendering on the RGBD target image to generate target images at different depth positions, and simultaneously extract the complex amplitude light field distribution of the intermediate plane. Step 2.2: Gait feature extraction. Real-time images of the user's eyes are captured using a near-eye camera, and eye features are analyzed to output accurate coordinates of the gaze point and related feature parameters. Step 2.3: Light field propagation and feature fusion. The complex amplitude of the intermediate plane is propagated to the spatial light modulator plane through the angular spectrum diffraction algorithm to obtain the initial complex amplitude hologram; combined with the gaze characteristics, the gaze region is weighted and enhanced to generate fused features containing visual attention information. Step 2.4: Pure Phase Hologram Generation. The fused features are input into the hologram encoding module, and a pure phase hologram adapted to the spatial light modulator is output through the neural network model to realize the phase modulation encoding of the incident light field.
6. The three-dimensional computational hologram reconstruction method based on gaze region optimization as described in claim 5, characterized in that, The gaze region hologram generation process of the end-to-end neural network includes the following modules: The gaze estimation module is used to acquire eye images in real time through a near-eye camera, extract the user's gaze characteristics using visual processing algorithms, and output coordinate information including the location of the gaze point. The complex amplitude field generation module is used to generate the complex amplitude light field distribution of the intermediate target plane through a depth perception network, taking the RGBD target image as input, and capturing the multi-depth layer features of the scene; The feature fusion module is used to spatiotemporally align complex amplitude light field information with gaze characteristics, enhance the feature representation of the gaze region through an attention mechanism, and generate enhanced fused features. The hologram encoding module is used to map fused features into pure phase holograms that can be loaded by a spatial light modulator, realizing end-to-end mapping from scene input to hologram output.
7. The three-dimensional computational hologram reconstruction method based on gaze region optimization as described in claim 6, characterized in that, The neural network training includes the following sub-steps: Step 3.1, Multi-depth target image preprocessing: During the neural network training process, the RGBD input image is first subjected to depth mask or defocus rendering to generate a set of target images at different depth positions, thereby providing multi-plane reference data for the training of subsequent modules and building a basic data support system; Step 3.2 Unsupervised Training of Complex Amplitude Field Generation Module: Next, the RGBD dataset is input into this module, which outputs the complex amplitude distribution of the intermediate plane. Then, the reconstructed light field distribution of each target plane is calculated using the angular spectrum diffraction algorithm. After obtaining the reconstructed image, its amplitude information is extracted and compared with the target image to construct the loss function. Finally, the stochastic gradient descent algorithm is used to optimize the network parameters, realizing the unsupervised learning process, allowing the module to learn the features and patterns in the data autonomously. Step 3.3, Semi-supervised training of gaze estimation module: For gaze estimation module, a two-stage training strategy is adopted to separate and enhance gaze features; In the feature separation stage, unlabeled eye images are input into the module to separate gaze features and appearance features. Then, the gaze features and appearance features of different samples are recombined by the appearance restoration module. A loss function is constructed based on the difference between the restored image and the original image to optimize the feature separation capability, enabling the module to extract gaze-related features more accurately. In the gaze prediction stage, labeled eye images are input to extract gaze features and predict gaze coordinates. Then, a loss function is constructed based on the error between the predicted coordinates and the true label to further enhance the expression of gaze features and improve the accuracy of the module's gaze prediction.
8. The three-dimensional computational hologram reconstruction method based on gaze region optimization as described in claim 7, characterized in that, The neural network training also includes the following sub-steps: Step 3.4, Joint Training of Feature Fusion and Hologram Encoding Modules: In the joint training of the feature fusion and hologram encoding modules, the complex amplitude field is first propagated to the spatial light modulator plane to generate an initial hologram. Then, gaze features are fused, and the gaze region is weighted and enhanced through an attention mechanism. To better optimize this process, a high-weight gaze region and a low-weight non-gaze region are divided according to the gaze point. The weighted error between the reconstructed image and the target image is calculated, and a region weighted loss function is constructed. Then, the parameters of the feature fusion module and the hologram encoding module are jointly optimized so that the two modules can work together to improve the quality of hologram generation. Step 3.5, End-to-end Joint Training: Finally, the labeled eye image and RGBD data are simultaneously input into the network. The training process of step 2.4.4 in the joint training of the feature fusion and hologram encoding modules is repeated according to the complete data flow. The parameters of all modules are optimized synchronously. Through this global training, the global optimum from input to output is achieved, enabling the entire neural network system to generate gaze region holograms efficiently and accurately.
9. An apparatus for implementing the three-dimensional computational hologram reconstruction method based on gaze region optimization as described in claim 1, characterized in that, This includes an iteratively optimized gaze region hologram generation component and an end-to-end neural network gaze region hologram generation component. The iteratively optimized gaze region hologram generation part includes the following modules: The gaze region segmentation module is used to obtain the coordinates of the gaze point, segment the gaze region and the non-gaze region, and construct a weight matrix; The multi-depth generation module is used to render 3D information into multi-depth target images using holographic optimization algorithms; The light field propagation module is used to propagate the hologram loaded on the SLM plane to the target plane at various distances to obtain the reconstructed image of the hologram on the target plane; The gradient optimization module is used to calculate the loss value between the hologram reconstructed image and the target image, and iteratively optimizes the hologram through the gradient optimization algorithm; The gaze region hologram generation part of the end-to-end neural network includes the following modules: The gaze estimation module is used to acquire eye images in real time through a near-eye camera, extract the user's gaze characteristics using visual processing algorithms, and output coordinate information including the location of the gaze point. The complex amplitude field generation module is used to generate the complex amplitude light field distribution of the intermediate target plane through a depth perception network, taking the RGBD target image as input, and capturing the multi-depth layer features of the scene; The feature fusion module is used to spatiotemporally align complex amplitude light field information with gaze characteristics, enhance the feature representation of the gaze region through an attention mechanism, and generate enhanced fused features. The hologram encoding module is used to map fused features into pure phase holograms that can be loaded by a spatial light modulator, realizing end-to-end mapping from scene input to hologram output.
Citation Information
Patent Citations
Adjustable depth-of-field hologram reconstruction method based on gradient descent optimization algorithm
CN115857305A
Real fuzzy three-dimensional hologram reconstruction method based on joint training of multiple neural networks
CN117876591A