Hologram model training method and device based on label-free sample training set, equipment and medium
Through the iterative training of RGB-D images and parameter optimization of the label-free sample training set, the problem of taking into account both the quality and speed of the generation of hologram models is solved, and efficient and stable hologram generation and adaptability are achieved.
Patent Information
- Application Number
- CN202510465203.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, hologram model training methods are difficult to balance the generation quality and speed, and are highly dependent on data sets, resulting in increased training cost and time cost.
The label-free sample training set is adopted, and iterative training is performed by obtaining RGB-D images, multi-depth diffraction field matrix is calculated, the hologram is reconstructed using angular spectroscopy, and the parameters are optimized through model loss, so as to achieve efficient training of the hologram model.
Reliance on data sets is reduced, the generation quality and speed of holograms are improved, the generalization ability and coding efficiency of the model are enhanced, and the generated holograms are adapted to complex scenarios. The quality of the generated holograms is high and stable.
Smart Images

Figure CN120375121A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of three-dimensional imaging, and particularly to a hologram model training method, device, equipment and medium based on an unlabeled sample training set. Background Technique
[0002] Computer generated holography (CGH) combines computer technology with optical holography, and can achieve some functions that are difficult or even impossible to achieve by optical holography, such as controlling the size of the reproduced object, or reproducing some objects that do not exist in reality, etc. The device that replaces the photosensitive material in optical holography is mainly a spatial light modulator (SLM).
[0003] The current SLMs are mainly phase-type, which can only modulate the phase of the incident light and cannot modulate the amplitude. In related technologies, the generation methods for generating corresponding pure phase information from diffraction field information include the iterative Gerchberg-Saxton algorithm (GS algorithm) and the Wirtinger Holography (WH) algorithm, the one-step encoding Double-phase Holography algorithm (DPH algorithm) and the Stochastic Gradient Descent algorithm (SGD algorithm), etc. The above methods often cannot have both high reproduction quality and high generation speed. The DPH algorithm has the fastest generation speed, but the reproduction quality is often not high. And the WH algorithm requires a long iteration time to generate holograms with high reproduction quality.
[0004] Based on this, the present application provides a hologram model training method, device, equipment and medium based on an unlabeled sample training set. Summary of the Invention
[0005] The purpose of the present application is to provide a hologram model training method, device, equipment and medium based on an unlabeled sample training set, which can reduce the dependence on the data set and effectively improve the encoding quality and generation speed of three-dimensional holograms.
[0006] To achieve the above purpose, the present application provides the following solutions:
[0007] In a first aspect, the present application provides a hologram model training method based on an unlabeled sample training set, including:
[0008] Obtain an unlabeled sample training set; the unlabeled sample training set includes multiple RGB-D images; the RGB-D images are obtained by combining an intensity image and a depth image; the intensity image contains information of three color channels of red, green, and blue, and the depth image contains distance information of the object surface from the observation point;
[0009] Use the unlabeled sample training set to iteratively train the hologram model using the following steps until a preset number of iterations is reached;
[0010] Calculate the multi-depth diffraction field matrix of the target RGB-D image at a preset diffraction distance; the target RGB-D image is any one of the RGB-D images in the unlabeled sample training set;
[0011] Input the multi-depth diffraction field matrix into the hologram model to obtain a multi-depth hologram;
[0012] Use the angular spectrum method to reconstruct the multi-depth diffraction field of the multi-depth hologram layer by layer to obtain a focal stack image;
[0013] Calculate the model loss according to the target RGB-D image and the focal stack image;
[0014] Optimize the parameters of the hologram model according to the model loss.
[0015] In a second aspect, the present application provides a hologram model training device based on an unlabeled sample training set, including:
[0016] A data acquisition module for obtaining an unlabeled sample training set; the unlabeled sample training set includes multiple RGB-D images; the RGB-D images are obtained by combining an intensity image and a depth image; the intensity image contains information of three color channels of red, green, and blue, and the depth image contains distance information of the object surface from the observation point;
[0017] An iterative training module for using the unlabeled sample training set to iteratively train the hologram model using the following steps until a preset number of iterations is reached;
[0018] A diffraction field matrix calculation module for calculating the multi-depth diffraction field matrix of the target RGB-D image at a preset diffraction distance; the target RGB-D image is any one of the RGB-D images in the unlabeled sample training set;
[0019] A hologram acquisition module for inputting the multi-depth diffraction field matrix into the hologram model to obtain a multi-depth hologram;
[0020] A focal stack image acquisition module for using the angular spectrum method to reconstruct the multi-depth diffraction field of the multi-depth hologram layer by layer to obtain a focal stack image;
[0021] A loss calculation module, configured to calculate a model loss according to the target RGB-D image and the defocus stack image;
[0022] A parameter optimization module, configured to optimize the parameters of the hologram model according to the model loss.
[0023] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the hologram model training method based on an unlabeled sample training set described in any one of the above.
[0024] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the hologram model training method based on an unlabeled sample training set described in any one of the above are implemented.
[0025] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0026] The present application provides a hologram model training method, apparatus, device, and medium based on an unlabeled sample training set, which solves the problems of difficult large-scale acquisition of training data and high annotation costs by obtaining the unlabeled sample training set. This training set consists of multiple RGB-D images, which not only contain rich color information (red, green, and blue color channels), but also incorporate depth information, that is, the distance between the object surface and the observation point. These information together constitute the key inputs required for hologram model training, laying a solid foundation for realizing unlabeled training, enabling the full utilization of unannotated RGB-D image resources, and reducing the cost and difficulty of data preparation. By calculating the multi-depth diffraction field matrix of the target RGB-D image at a preset diffraction distance and inputting the multi-depth diffraction field matrix into the hologram model to obtain a multi-depth hologram, the problem of how to extract effective features from unlabeled data and generate high-quality holograms is solved. During the process of calculating the multi-depth diffraction field matrix, the depth information of the RGB-D image is fully utilized and transformed into a form that the hologram model can understand. Through the forward propagation of the model, a multi-depth hologram is generated, providing a basis for subsequent diffraction reconstruction and loss calculation. This not only improves the generation efficiency of the hologram, but also ensures that the generated hologram can accurately reflect the three-dimensional structure information of the original RGB-D image. By using the angular spectrum method to layer-by-layer diffractively reconstruct the multi-depth diffraction field of the multi-depth hologram to obtain a focal stack image, the problem of how to verify and evaluate the quality of hologram generation is solved. As the result of hologram diffractive reconstruction, the quality of the focal stack image directly reflects the performance of the hologram model. By comparing with the original target RGB-D image, the model loss can be calculated, and then the parameters of the hologram model can be optimized. By optimizing the parameters of the hologram model according to the model loss, the continuous improvement of the model performance is achieved. This step not only solves the problem of parameter adjustment in the model training process, but also ensures that the model can gradually adapt to the distribution characteristics of unlabeled data, thereby improving its generalization ability in real application scenarios.
[0027] In summary, the hologram model training method based on an unlabeled sample training set provided by the present application effectively solves the problem of data annotation dependence in traditional methods, improves the model training efficiency and the quality of hologram generation, and lays a solid foundation for the wide application of hologram technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0029] Figure 1Schematic flowchart of a hologram model training method based on an unlabeled sample training set provided by an embodiment of the present application;
[0030] Figure 2 Schematic diagram of random matrices with different numbers of pixels in height provided by an embodiment of the present application; wherein, Figure 2 (a) Schematic diagram of a random matrix with 4 pixels in height; Figure 2 (b) Schematic diagram of a random matrix with 64 pixels in height;
[0031] Figure 3 For Figure 2 Schematic diagrams of intensity images generated by two random matrices with different numbers of pixels in height provided in Figure 3 (a) Schematic diagram of an intensity image generated by a random matrix with 4 pixels in height through the NNU interpolation algorithm; Figure 3 (b) Schematic diagram of an intensity image generated by a random matrix with 64 pixels in height through the NNU interpolation algorithm; Figure 3 (c) Schematic diagram of an intensity image generated by a random matrix with 4 pixels in height through the BU interpolation algorithm; Figure 3 (d) Schematic diagram of an intensity image generated by a random matrix with 64 pixels in height through the BU interpolation algorithm; Figure 3 (e) Schematic diagram of an intensity image generated by a random matrix with 4 pixels in height through the BCU interpolation algorithm; Figure 3 (f) Schematic diagram of an intensity image generated by a random matrix with 64 pixels in height through the BCU interpolation algorithm;
[0032] Figure 4 Schematic diagrams of the spectral distributions of intensity images generated by random matrices with different numbers of pixels in height through different interpolation algorithms; wherein, Figure 4 (a) Schematic diagram of the spectral distribution of an intensity image generated through the NNU interpolation algorithm; Figure 4 (b) Schematic diagram of the spectral distribution of an intensity image generated through the BU interpolation algorithm; Figure 4 (c) Schematic diagram of the spectral distribution of an intensity image generated through the BCU interpolation algorithm;
[0033] Figure 5 Statistical chart of the depth distribution of depth images generated by different interpolation algorithms provided by an embodiment of the present application;
[0034] Figure 6 Schematic diagram of the network architecture of the hologram model provided by an embodiment of the present application;
[0035] Figure 7 Schematic diagram of the functional modules of a hologram model training device based on an unlabeled sample training set provided by an embodiment of the present application;
[0036] Figure 8 A structural schematic diagram of a computer device provided by an embodiment of the present application. Specific implementation manners
[0037] First, some technical terms involved in the embodiments of the present application are introduced.
[0038] Neural network training refers to the process of adjusting parameters such as weights and biases in a neural network so that the network can accurately extract features from input data and perform prediction or classification. This process highly depends on the quality and diversity of the dataset, and constructing a high-quality three-dimensional holographic dataset often requires huge time, manpower, and financial costs. The datasets directly applicable to the three-dimensional holographic field are very limited. A suitable dataset is extremely important for the training of the network model. In order to approach the limit performance of the CNN (Convolutional Neural Network) model and achieve better training effects, the training set used for model training should be as rich and comprehensive as possible. However, the data in the dataset is limited and cannot provide all possible training data, which results in the fact that the effect obtained by the trained network model on the test set is often not as good as that on the training set. The sample-free dataset has no data restrictions and can provide all possible training data, with a very high optimization ceiling.
[0039] The complex amplitude distribution of the diffraction field in a 3D scene is determined by its intensity and depth distributions. Different intensity and depth distributions will generate different complex amplitudes. In order to effectively ensure the encoding and generalization capabilities of the neural network model, the 3D dataset should cover diverse intensity distributions at all depth levels. Using a fully convolutional neural network to perform pure phase encoding on the complex amplitude and training with sample-free intensity images and uniformly distributed depth images to achieve flexible control of the ratio of high-frequency and low-frequency information in the training set has become a good idea for the problem of neural network encoding multi-depth holograms.
[0040] In view of this, the embodiments of the present application propose a hologram model training method based on an unlabeled sample training set. This method utilizes the concept of a sample-free dataset, simulates training data by generating diverse intensity images and uniformly distributed depth images, so as to approach the limit performance of the CNN model. At the same time, a fully convolutional neural network is used to perform pure phase encoding on the complex amplitude to achieve flexible control of the ratio of high-frequency and low-frequency information in the training set. This method not only reduces the cost of constructing a high-quality three-dimensional holographic dataset but also improves the encoding and generalization capabilities of the neural network model.
[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0042] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0043] In an exemplary embodiment, as Figure 1 shown, a hologram model training method based on an unlabeled sample training set is provided, including the following steps 101 to 107. Wherein:
[0044] Step 101, obtain an unlabeled sample training set; the unlabeled sample training set includes a plurality of RGB-D images; the RGB-D images are obtained by combining intensity images and depth images; the intensity images contain information of three color channels of red, green, and blue, and the depth images contain distance information of the object surface from the observation point.
[0045] Step 102, use the unlabeled sample training set to iteratively train the hologram model in the following steps until a preset number of iterations is reached.
[0046] Step 103, calculate the multi-depth diffraction field matrix of the target RGB-D image at a preset diffraction distance; the target RGB-D image is any RGB-D image in the unlabeled sample training set.
[0047] Step 104, input the multi-depth diffraction field matrix into the hologram model to obtain a multi-depth hologram.
[0048] Step 105, use the angular spectrum method to diffractively reconstruct the multi-depth diffraction field of the multi-depth hologram layer by layer to obtain a focal stack image.
[0049] Step 106, calculate the model loss according to the target RGB-D image and the focal stack image.
[0050] Step 107, optimize the parameters of the hologram model according to the model loss.
[0051] By implementing the above steps 101 to 107, the present application can make full use of unlabeled RGB-D image data, realize efficient and high-quality hologram generation through iterative training and optimizing model parameters, and also reduce the cost of data annotation.
[0052] In another exemplary embodiment of the present application, the method for generating the intensity image and the depth image in step 101 specifically includes:
[0053] Determine the number of rows h of the first original matrix and the preset width-to-height ratio of the intensity image.
[0054] After multiplying the number of rows h of the first original matrix by the preset width-to-height ratio of the intensity image, perform a rounding operation to obtain the number of columns w of the first original matrix.
[0055] Generate a first original matrix with the number of rows h and the number of columns w; wherein, each element value in the first original matrix is randomly distributed between 0 and 255. As Figure 2 shown, where δ represents the number of height pixels of the matrix, Figure 2 (a) is a schematic diagram of a random matrix with 4 height pixels; Figure 2 (b) is a schematic diagram of a random matrix with 64 height pixels.
[0056] Based on the first original matrix, use an interpolation algorithm to generate an intensity image with a width of w1 and a height of h1, where w1 and h1 satisfy the preset width-to-height ratio. A variety of interpolation algorithms can be used. Different interpolation algorithms have different spectral distributions, and different interpolation algorithms can also be mixed. The intensity images generated by different interpolation algorithms for different numbers of height pixels are as Figure 3 shown. Among them, Figure 3 (a) is a schematic diagram of the intensity image generated by a random matrix with 4 height pixels through the NNU (Nearest Neighbor Upsampling) interpolation algorithm; Figure 3 (b) is a schematic diagram of the intensity image generated by a random matrix with 64 height pixels through the NNU interpolation algorithm; Figure 3 (c) is a schematic diagram of the intensity image generated by a random matrix with 4 height pixels through the BU (Bilinear Upsampling) interpolation algorithm; Figure 3 (d) is a schematic diagram of the intensity image generated by a random matrix with 64 height pixels through the BU interpolation algorithm; Figure 3 (e) is a schematic diagram of the intensity image generated by a random matrix with 4 height pixels through the BCU (Bicubic Upsampling) interpolation algorithm; Figure 3 (f) is a schematic diagram of the intensity image generated by a random matrix with 64 height pixels through the BCU interpolation algorithm. Among them, the function of the interpolation algorithm is usually to increase the resolution or fineness of the image. Through interpolation processing, the intensity image contains more data points or finer information. Figure 3 (a) and Figure 3 (b) have a higher number of pixels (or resolution) than Figure 2 (a) andFigure 2 (b).
[0057] Among them, when implementing this implementation method, the spectral distribution of the intensity image obtained by analyzing the intensity images with different height pixel numbers δ and different interpolation algorithms through the spectral averaging method is as Figure 4 shown. Among them, Figure 4 (a) is a schematic diagram of the spectral distribution of the intensity image generated by the NNU interpolation algorithm with different height pixel numbers; Figure 4 (b) is a schematic diagram of the spectral distribution of the intensity image generated by the BU interpolation algorithm with different height pixel numbers; Figure 4 (c) is a schematic diagram of the spectral distribution of the intensity image generated by the BCU interpolation algorithm with different height pixel numbers. According to Figure 4 it can be found that the spectral distributions corresponding to different δ and different interpolation algorithms are quite different. The spectral distribution of the intensity image of NNU shows that as δ increases, the low-frequency part decreases significantly and the high-frequency part increases significantly; while the spectral differences of the intensity images generated by BU and BCU are not large, but upon careful observation, it can be found that there are still some differences in both the low-frequency and high-frequency components. The spectral distributions of the intensity images of BU and BCU show that as δ increases, the low-frequency components decrease, which is similar to the change of NNU, and the high-frequency components increase, but it is inconsistent with the change of nearest neighbor upsampling. Therefore, specific spectral distributions of the intensity image and uniform distributions of the depth image can be achieved through parameter control.
[0058] Determine the number of rows h2 and the number of columns w2 of the second original matrix; among them, h2 is less than or equal to h1, and w2 is less than or equal to w1.
[0059] Generate a second original matrix with h2 rows and w2 columns; among them, each element value in the second original matrix is randomly distributed between 0 and 255.
[0060] Based on the second original matrix, use an interpolation algorithm to generate a depth image with a width of w1 and a height of h1. A variety of interpolation algorithms can be used. Different interpolation algorithms have different spectral distributions, and different interpolation algorithms can also be mixed to obtain a depth image, and its depth distribution is as Figure 5 shown. And according to Figure 5 it can be known that compared with the NYU dataset collected by the depth camera and the 3D modeling dataset MIT-CGH-4K, the distribution of the unlabeled sample training set ZS-CGH in 10 depth layers is more uniform.
[0061] As an alternative implementation method, the number of rows h of the first original matrix is determined according to the frequency. If emphasizing the low frequency, it is recommended to set it to 1 to 16. If emphasizing the middle frequency, it is recommended to set it to 16 to 64. If emphasizing the high frequency, it is recommended to set it higher than 64.
[0062] In another exemplary embodiment of the present application, calculating the multi-depth diffraction field matrix of the target RGB-D image at a preset diffraction distance in step 103 specifically includes:
[0063] The target RGB-D image is evenly divided into N depth layers according to depth, and the complex amplitude distribution of each depth layer is calculated respectively using the angular spectrum method to obtain the multi-depth diffraction field matrix of the target RGB-D image; where N is an integer, and the depth intervals of each depth layer are equal. The iteration of the multi-depth diffraction process of the angular spectrum method (Angular Spectrum Method, ASM) can be extended to more general scenarios without sacrificing generality.
[0064] The calculation formula for the complex amplitude distribution is:
[0065]
[0066] where n represents the nth depth layer of the target RGB-D image, n = [0, N - 1], and N represents the total number of depth layers; U n (x, y) represents the complex amplitude distribution at the point (x, y) of the target RGB-D image in the nth depth layer; Λ{·} is the symbolic representation of the angular spectrum diffraction process, used to represent the process of the light wave in the target RGB-D image propagating from the current depth layer to the next depth layer; Δz′ represents the preset diffraction distance; I(x, y) represents the amplitude at the point (x, y) of the target RGB-D image; j is the imaginary unit, j 2 = -1; φ0 represents the initial wave state of the light wave; M n (x, y) represents the Boolean mask at the point (x, y) of the target RGB-D image in the nth depth layer, used to distinguish the light wave distributions of different depth layers; ASM{···}Δz′ represents the angular spectrum method calculation process with a preset diffraction distance, used to calculate the complex amplitude distribution after the light wave propagates in free space; F{} is the symbolic representation of the Fourier transform, used to convert the spatial domain signal into the frequency domain signal; F -1 {} is the symbolic representation of the inverse Fourier transform, used to convert the frequency domain signal back to the spatial domain signal; H Δz′ (f x , f y ) is the transfer function, representing the phase change during the propagation of the light wave in free space, where f x and f y respectively represent the spatial frequency components in the x and y directions; k is the wave number, k = 2π / λ, and λ is the light wave wavelength. where Δp is the pixel pitch of the spatial light modulator (SLM). Through layer-by-layer calculation, the complex amplitude U h of the hologram plane S h (x, y) can be obtained. With the hologram model constructed by CNN, the complex amplitude Uh (x, y) will be used to encode multi-depth holograms.
[0067] In another exemplary embodiment of the present application, step 104 specifically includes:
[0068] Convert the multi-depth diffraction field of the target RGB-D image to obtain a multi-channel input matrix.
[0069] Input the multi-channel input matrix into the hologram model to obtain a multi-depth hologram; the hologram model includes an input layer and an output layer; the input layer is used to extract features from the multi-channel input matrix to generate a phase matrix and input the phase matrix into the output layer; the output layer is used to recombine the phase matrix through a phase recombination algorithm to obtain a multi-depth hologram.
[0070] As an alternative implementation, converting the multi-depth diffraction field of the target RGB-D image to obtain a multi-channel input matrix specifically includes:
[0071] Use the following formula to calculate the complex amplitude distribution of the target RGB-D image from the (N - 1)-th depth layer to the holographic plane:
[0072] U h (x, y) = ASM{U N-1 (x, y)} -zh′-(N-1)Δz′ .
[0073] Where, U h (x, y) represents the complex amplitude distribution at the point (x, y) on the holographic plane; ASM{···}Δz′ represents the angular spectrum method calculation process for a preset diffraction distance, used to calculate the complex amplitude distribution after the light wave propagates in free space; U N-1 (x, y) represents the complex amplitude distribution at the point (x, y) on the (N - 1)-th depth layer of the target RGB-D image, N represents the total number of depth layers; -zh′-(N - 1) represents the reverse distance from the (N - 1)-th depth layer to the holographic plane.
[0074] Use Euler's formula to convert U h (x, y) into real part channel data and imaginary part channel data.
[0075] Based on the real part channel data and the imaginary part channel data, use a shuffle operation to generate a multi-channel input matrix containing 8 channels.
[0076] In another exemplary embodiment of the present application, the calculation formula of step 105 is specifically:
[0077]
[0078] where \(R(x, y)\) represents the pixel value of the focus-stacked image at point \((x, y)\) in the multi-depth hologram; \(R\) n (x, y) represents the focused light field at point \((x, y)\) in the \(n\)-th depth layer of the multi-depth hologram; \(n\) represents the \(n\)-th depth layer of the target RGB-D image, \(n = [0, N - 1]\), and \(N\) represents the total number of depth layers; \(U\) H (x, y) represents the complex amplitude at point \((x, y)\) in the multi-depth hologram; \(M\) n (x, y) represents the Boolean mask at point \((x, y)\) in the \(n\)-th depth layer of the target RGB-D image, which is used to distinguish the light wave distributions of different depth layers; \(ASM\{\cdots\}_z\) represents calculating the diffraction field of the multi-depth hologram at the reconstruction distance, \(z\) represents the reconstruction distance, \(z_h\) represents the reconstruction distance of the initial focal plane, and \(\Delta z\) represents the interval between adjacent focal planes; \(F\{\}\) is the symbol representation of the Fourier transform, which is used to convert the spatial-domain signal into the frequency-domain signal; \(F\) -1 \{\} is the symbol representation of the inverse Fourier transform, which is used to convert the frequency-domain signal back to the spatial-domain signal; \(H_z\) represents the transfer function, \(z\) h is the reconstruction distance of the initial focal plane, and \(\Delta z\) is the interval between adjacent focal planes; \(f\) x and \(f\) y respectively represent the spatial frequency components in the \(x\) and \(y\) directions.
[0079] In another exemplary embodiment of the present application, the calculation formula for calculating the model loss in step 106 according to the target RGB-D image and the focus-stacked image is specifically:
[0080]
[0081] where \(Loss\) represents the loss of the current iteration; \(L1\{\}\) represents the mean absolute error loss function; \(I(x, y)\) represents the amplitude of the target RGB-D image at point \((x, y)\); \(R(x, y)\) represents the pixel value of the focus-stacked image at point \((x, y)\) in the multi-depth hologram. represents a linear transformation, aiming to match the brightness of the target intensity image. By adjusting the laser power, this adjustment can be achieved in the optical experiment.
[0082] As an optional implementation manner, as Figure 6 shown, a fully convolutional neural network based on CNN is designed to encode the hologram model of the pure-phase hologram.
[0083] The input of the CNN is from the obtained complex amplitude \(U\) h(x, y) generation. For the convenience of holographic phase encoding, the complex amplitude is converted into two channels representing the real part and the imaginary part using Euler's formula. The "shuffle" operation is adopted to improve the model efficiency and generalization ability. This operation is applied to the data of the two channels to generate 8 channels. This conversion produces a matrix of size 1152*2048*8 as the input of the CNN.
[0084] Output of the CNN. The output consists of a phase matrix with a size of 1152*2048*4. By using the phase recombination technique, multi-depth holograms are obtained from the output of U-Net+.
[0085] Loss function of the CNN. The L1 (mean absolute error) loss function is used. Through the holographic reconstruction process using the method in step 105, the superimposed reconstructed focal stack image R(x, y) is generated. The residual (calculated as the difference between this reconstruction and the amplitude of the target intensity image) is used as the loss function, which helps to update the model parameters. The CNN consists of 3 convolutional modules, 2 2*2 max-pooling downsampling modules with convolutional kernels (size, padding, stride) set to (3, 1, 1), and two transposed convolutional upsampling modules with (2, 0, 4). The convolutional modules of the same size on the left and right are connected by skip connections to facilitate the transfer of feature maps. The residual of the input RGB-D image intensity is used as the loss function. Based on the CNN model, the model parameters are adjusted through gradient descent backpropagation to optimize the quality of the CNN model encoding the pure-phase hologram.
[0086] This application also provides an application scenario that applies the above hologram model training method based on an unlabeled sample training set. Specifically: The hologram model training method based on an unlabeled sample training set provided in this embodiment can be applied in a three-dimensional object digital reconstruction scenario. The three-dimensional object digital reconstruction scenario includes a data acquisition link, a model training and processing link, and a reconstruction output link; data enters the model training and processing link from the data acquisition link, and after iterative training and optimization processing of the hologram model trained with unlabeled samples, an accurate hologram model is obtained and enters the downstream reconstruction output link. The hologram model training method based on an unlabeled sample training set provided in this embodiment belongs to the hologram model iterative training and optimization link in the model training and processing link. Specifically, in the process of the digital reconstruction processing link for three-dimensional objects, based on an unlabeled RGB-D image dataset, the hologram model can be iteratively trained to continuously optimize the model parameters to improve the prediction and reconstruction ability of the multi-depth diffraction field, and finally achieve high-precision three-dimensional object digital reconstruction.
[0087] Compared with the traditional method of generating holograms using neural networks, this application has the following advantages:
[0088]
[0089] 1. The label-free sample training set based on this application is generated in real time by numerical calculation software, which does not occupy additional storage space. It is used to replace the real-scene training set in the model training process and can train a fully convolutional neural network in a simple and flexible "zero-sample" manner. By adjusting the number of rows of random data and the interpolation method, the ratio of high-frequency and low-frequency information in the intensity images of the training set can be controlled, and the depth images are evenly distributed across different depth layers to adapt to the actual application scenarios. The experimental results show that the numerical and full-color optical reconstruction effects of the model are verified through real scenes and test patterns. The model trained by ZS-CGH far outperforms the NYU dataset collected by depth cameras in various scenarios, and its performance in partial real scenes is comparable to that of the carefully designed 3D modeling dataset MIT-CGH-4K, and even exceeds the 3D modeling dataset in binary scenes.
[0090] 2. Realistic display quality and excellent 3D depth focusing effect. This application has achieved excellent performance in holographic display by using the CNN network. The experimental results show that the trained hologram model exhibits realistic display quality and excellent 3D depth focusing effect in multiple test sets (including real-scene test sets and binary pattern test sets). Compared with traditional methods, the encoding ability of CNN can better capture depth information, making the holographic display visually more realistic and accurate.
[0091] 3. Enhanced model generalization ability and stability. By geometrically modeling the intensity images of complex scenes and the evenly distributed depth images, this application enhances the generalization ability of the holographic encoding neural network, enabling it to still exhibit high stability and reliability when facing new or unseen data. Different from traditional methods, the newly designed dataset optimizes the intensity spectrum and depth distribution, enabling the hologram model to better handle complex scenes, thus improving the adaptability of the model in real scenes.
[0092] 4. High encoding speed and superior performance. Experiments show that the encoding time of this application at 4K resolution is only 10 ms, and it has achieved excellent results with an average PSNR of 35.55 dB and SSIM of 0.87 on the real-scene test set. Although there are differences between ZS-CGH and MIT-CGH-4K, the reconstruction effect differences are negligible, indicating that the holographic encoding method of this application has excellent performance and application potential. This provides a new technical path for the realization of real-time high-quality holographic display, significantly improving the encoding efficiency and display quality.
[0093] 5. The generated phase is smoother. Compared with traditional methods, for the holograms generated using the CNN-based hologram model, the phase change is gentler without problems such as phase mutations. Subsequently, problems such as the rings in the reconstructed image of the DPH algorithm and the speckle noise in the reconstructed image of the WH algorithm are avoided.
[0094] Based on the same inventive concept, an embodiment of the present application further provides a hologram model training device based on an unlabeled sample training set for implementing the above-mentioned hologram model training method based on an unlabeled sample training set. The implementation solutions for solving problems provided by this device are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more of the following embodiments of the hologram model training device based on an unlabeled sample training set can refer to the limitations on the hologram model training method based on an unlabeled sample training set in the above text and will not be elaborated here.
[0095] In an exemplary embodiment, as Figure 7 shown, a hologram model training device based on an unlabeled sample training set is provided, including:
[0096] A data acquisition module 201, configured to acquire an unlabeled sample training set; the unlabeled sample training set includes a plurality of RGB-D images; the RGB-D images are obtained by combining an intensity image and a depth image; the intensity image contains information of three color channels of red, green, and blue, and the depth image contains the distance information of the object surface from the observation point.
[0097] An iterative training module 202, configured to iteratively train the hologram model using the unlabeled sample training set by the following steps until a preset number of iterations is reached.
[0098] A diffraction field matrix calculation module 203, configured to calculate a multi-depth diffraction field matrix of a target RGB-D image at a preset diffraction distance; the target RGB-D image is any one of the RGB-D images in the unlabeled sample training set.
[0099] A hologram acquisition module 204, configured to input the multi-depth diffraction field matrix into the hologram model to obtain a multi-depth hologram.
[0100] A focal stack image acquisition module 205, configured to use the angular spectrum method to diffractively reconstruct the multi-depth diffraction field of the multi-depth hologram layer by layer to obtain a focal stack image.
[0101] A loss calculation module 206, configured to calculate a model loss according to the target RGB-D image and the focal stack image.
[0102] A parameter optimization module 207, configured to optimize the parameters of the hologram model according to the model loss.
[0103] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal, and its internal structure diagram may be as shown in Figure 8 . The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store unlabeled sample training data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a hologram model training method based on an unlabeled sample training set.
[0104] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0105] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0106] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0107] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0108] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0109] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0110] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0111] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A hologram model training method based on an unlabeled sample training set, characterized in that, The hologram model training method based on an unlabeled sample training set includes: Obtaining an unlabeled sample training set; the unlabeled sample training set includes a plurality of RGB-D images; the RGB-D images are obtained by combining an intensity image and a depth image; the intensity image contains information of three color channels of red, green, and blue, and the depth image contains distance information of the object surface from the observation point; Using the unlabeled sample training set to iteratively train the hologram model by the following steps until a preset number of iterations is reached; Calculating a multi-depth diffraction field matrix of a target RGB-D image at a preset diffraction distance; the target RGB-D image is any one of the RGB-D images in the unlabeled sample training set; Inputting the multi-depth diffraction field matrix into the hologram model to obtain a multi-depth hologram; Using the angular spectrum method to diffractively reconstruct the multi-depth diffraction field of the multi-depth hologram layer by layer to obtain a focal stack image; Calculating a model loss according to the target RGB-D image and the focal stack image; Optimizing the parameters of the hologram model according to the model loss.
2. The holographic model training method based on the tagless sample training set according to claim 1, wherein A method for generating an intensity image and a depth image specifically includes: Determining the number of rows h of a first original matrix and a preset width-to-height ratio of the intensity image; Multiplying the number of rows h of the first original matrix by the preset width-to-height ratio of the intensity image and then performing a rounding operation to obtain the number of columns w of the first original matrix; Generating a first original matrix with the number of rows h and the number of columns w; wherein, each element value in the first original matrix is randomly distributed between 0 and 255; Based on the first original matrix, using an interpolation algorithm to generate an intensity image with a width of w1 and a height of h1, where w1 and h1 satisfy the preset width-to-height ratio; Determining the number of rows h2 and the number of columns w2 of a second original matrix; wherein, h2 is less than or equal to h1, and w2 is less than or equal to w1; Generating a second original matrix with the number of rows h2 and the number of columns w2; wherein, each element value in the second original matrix is randomly distributed between 0 and 255; Based on the second original matrix, using an interpolation algorithm to generate a depth image with a width of w1 and a height of h1.
3. The hologram model training method based on an untagged sample training set according to claim 1, wherein Calculating a multi-depth diffraction field matrix of a target RGB-D image at a preset diffraction distance specifically includes: Uniformly dividing the target RGB-D image into N depth layers according to the depth, and using the angular spectrum method to calculate the complex amplitude distribution of each depth layer respectively to obtain the multi-depth diffraction field matrix of the target RGB-D image; wherein, N is an integer, and the depth intervals of each depth layer are equal; The calculation formula of the complex amplitude distribution is: ASM{U n-1 (x,y)}Δz′ = F -1 {F{U n-1 (x,y)}·H Δz′ (f x , f y )}; where n represents the n-th depth layer of the target RGB-D image, n = [0, N-1], and N represents the total number of depth layers; U n (x, y) represents the complex amplitude distribution at the point (x, y) of the target RGB-D image at the n-th depth layer; Λ{·} is the symbolic representation of the angular spectrum diffraction process, which is used to represent the process of the light wave in the target RGB-D image propagating from the current depth layer to the next depth layer; Δz′ represents the preset diffraction distance; I(x, y) represents the amplitude at the point (x, y) of the target RGB-D image; j is the imaginary unit, j 2 = -1; φ0 represents the initial wave state of the light wave; M n (x, y) represents the Boolean mask at the point (x, y) of the target RGB-D image at the n-th depth layer, which is used to distinguish the light wave distributions of different depth layers; ASM{···}Δz′ represents the calculation process of the angular spectrum method with the preset diffraction distance, which is used to calculate the complex amplitude distribution of the light wave after propagating in free space; F{} is the symbolic representation of the Fourier transform, which is used to convert the spatial domain signal into the frequency domain signal; F -1 {} is the symbolic representation of the inverse Fourier transform, which is used to convert the frequency domain signal back to the spatial domain signal; H Δz′ (f x , f y ) is the transfer function, which represents the phase change of the light wave during the propagation in free space, where f x and f y respectively represent the spatial frequency components in the x and y directions; k is the wave number, k = 2π / λ, and λ is the wavelength of the light wave.
4. The holographic model training method based on the tagless sample training set according to claim 1, characterized in that Inputting the multi-depth diffraction field matrix into the hologram model to obtain a multi-depth hologram specifically includes: Converting the format of the multi-depth diffraction field of the target RGB-D image to obtain a multi-channel input matrix; Inputting the multi-channel input matrix into the hologram model to obtain a multi-depth hologram; the hologram model includes an input layer and an output layer; the input layer is used to extract features from the multi-channel input matrix to generate a phase matrix and input the phase matrix into the output layer; the output layer is used to recombine the phase matrix through a phase recombination algorithm to obtain a multi-depth hologram.
5. The hologram model training method based on an unlabeled sample training set according to claim 1, wherein Convert the multi-depth diffraction field of the target RGB-D image to obtain a multi-channel input matrix, specifically including: Use the following formula to calculate the complex amplitude distribution of the target RGB-D image from the (N-1)th depth layer to the holographic plane: U h (x,y) = ASM{U N-1 (x,y)} -zh′-(N-1)Δz′ ; Among them, U h (x, y) represents the complex amplitude distribution of the holographic plane at the point (x, y); ASM{···}Δz′ represents the angular spectrum method calculation process of a preset diffraction distance, which is used to calculate the complex amplitude distribution of the light wave after propagating in free space; U N-1 (x, y) represents the complex amplitude distribution of the target RGB-D image at the point (x, y) of the (N - 1)-th depth layer, where N represents the total number of depth layers; -zh′ - (N - 1) represents the reverse distance from the (N - 1)-th depth layer to the holographic plane; Convert U h (x, y) into real part channel data and imaginary part channel data; Based on the real part channel data and the imaginary part channel data, use the shuffle operation to generate a multi-channel input matrix containing 8 channels.
6. The hologram model training method based on the tagless sample training set according to claim 1, wherein Use the angular spectrum method to diffractively reconstruct the multi-depth diffraction field of the multi-depth hologram layer by layer. The calculation formula for obtaining the focal stack image is specifically: Among them, R(x, y) represents the pixel value of the focus-stacked image at the point (x, y) in the multi-depth hologram; R n (x, y) represents the focused light field at the point (x, y) in the n-th depth layer of the multi-depth hologram; n represents the n-th depth layer of the target RGB-D image, n = [0, N - 1], and N represents the total number of depth layers; U H (x, y) represents the complex amplitude at the point (x, y) in the multi-depth hologram; M n (x, y) represents the Boolean mask at the point (x, y) in the n-th depth layer of the target RGB-D image, which is used to distinguish the light wave distributions of different depth layers; ASM{···}z represents calculating the diffraction field of the multi-depth hologram at the reconstruction distance, z represents the reconstruction distance, zh represents the reconstruction distance of the initial focal plane, and Δz represents the interval between adjacent focal planes; F{} is the symbolic representation of the Fourier transform, which is used to convert the spatial domain signal into the frequency domain signal; F -1 {} is the symbolic representation of the inverse Fourier transform, which is used to convert the frequency domain signal back to the spatial domain signal; Hz represents the transfer function, z h is the reconstruction distance of the initial focal plane, and Δz is the interval between adjacent focal planes; f x and f y respectively represent the spatial frequency components in the x and y directions.
7. The holographic model training method based on the untagged sample training set according to claim 1, characterized in that, The calculation formula for calculating the model loss according to the target RGB-D image and the focal stack image is specifically: Where Loss represents the loss of the current iteration; L1{} represents the mean absolute error loss function; I(x,y) represents the amplitude of the target RGB-D image at the point (x,y); R(x,y) represents the pixel value of the focal stack image at the point (x,y) of the multi-depth hologram.
8. A hologram model training device based on an unlabeled sample training set, characterized in that, The hologram model training device based on the unlabeled sample training set applies the hologram model training method based on the unlabeled sample training set according to any one of claims 1-7. The hologram model training device based on the unlabeled sample training set includes: A data acquisition module for acquiring an unlabeled sample training set; the unlabeled sample training set includes multiple RGB-D images; the RGB-D image is obtained by combining an intensity image and a depth image; the intensity image contains information on the red, green, and blue color channels, and the depth image contains information on the distance of the object surface from the observation point; An iterative training module for iteratively training the hologram model using the unlabeled sample training set by the following steps until a preset number of iterations is reached; A diffraction field matrix calculation module for calculating the multi-depth diffraction field matrix of the target RGB-D image at a preset diffraction distance; the target RGB-D image is any one of the RGB-D images in the unlabeled sample training set; A hologram acquisition module for inputting the multi-depth diffraction field matrix into the hologram model to obtain a multi-depth hologram; A focal stack image acquisition module for using the angular spectrum method to diffractively reconstruct the multi-depth diffraction field of the multi-depth hologram layer by layer to obtain a focal stack image; A loss calculation module for calculating the model loss according to the target RGB-D image and the focal stack image; A parameter optimization module for optimizing the parameters of the hologram model according to the model loss.
9. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the hologram model training method according to any one of claims 1-7 based on the unlabeled sample training set.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the hologram model training method according to any one of claims 1-7 based on the unlabeled sample training set.