Intelligent optimization method for compressing three-dimensional light field display

By constructing a deep learning-based optimal initial plane position prediction model, the problem of time-consuming and inaccurate initial plane position optimization in compressed 3D light field display is solved, achieving efficient and automated light field reconstruction and display optimization, thus improving display quality and efficiency.

CN116029919BActive Publication Date: 2026-02-17INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211584514.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2026-02-17
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

Existing compressed 3D light field display technology is time-consuming and inaccurate in optimizing the initial planar position, making it difficult to achieve real-time and efficient optimal position adjustment. This results in distortion of the reconstructed light field scene, failing to meet diverse user needs and promote widespread adoption.

Method used

A deep learning approach is used to construct an optimal initial plane position prediction model. Data is collected through a light field camera, and the model is trained offline and optimized online. A convolutional neural network is used to predict the optimal initial plane position, and the model is combined with a light field reconstruction model for efficient display optimization.

Benefits of technology

It has achieved automation and intelligence in compressed 3D light field display, improved the quality of light field reconstruction and display efficiency, and can accurately find the optimal initial plane position in different scenarios to enhance the display effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116029919B_ABST
    Figure CN116029919B_ABST
Patent Text Reader

Abstract

The application provides an intelligent optimization method for compressed three-dimensional light field display, which comprises data acquisition, offline training and online optimization; the data acquisition step refers to acquiring light field original images by using special equipment such as a light field camera, and performing analysis operation on the images, so as to obtain multi-viewpoint images and depth map information of different scenes; the offline training refers to finding the initial plane position corresponding to the best reconstruction quality, taking the depth map information and the corresponding data as a training set, and performing training; the online optimization refers to dynamically finding the optimal initial plane position; specifically, for light field source files of different scenes, after inputting a model, the optimal initial plane reconstruction position can be calculated in a short time, and the display on the light field display is shown, so that the best display effect is obtained. The application can improve the display quality of the compressed three-dimensional light field display, so that for a given arbitrary light field source file, the optimal data display scheme can be quickly obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of light field display methods, in particular to an intelligent optimization method for compressed three-dimensional light field display. BACKGROUND

[0002] Existing image acquisition and display lose multiple dimensions of visual information, which forces us to observe the three-dimensional world through a two-dimensional "window". When doctors perform abdominal surgery with a single-camera endoscope, they cannot determine the depth position of the tumor, so they need to observe from multiple angles multiple times to slowly cut; with the rapid development of electronic maps, traditional two-dimensional displays cannot provide occlusion relationships between buildings, which limits the observer's judgment of spatial accuracy; and the rise of the concept of the metaverse in recent years, etc. make the development of three-dimensional display technology particularly important. Among them, the compressed light field display using multi-layer spatial light modulators has the advantages of low cost, no loss of resolution, and small occupied area compared with holographic display and volumetric three-dimensional display, etc., making it have great application prospects.

[0003] However, current compressed three-dimensional light field displays mostly use fixed initial plane positions, ignoring the mapping relationship between the focal plane of light field acquisition and the initial plane position of reconstruction, which may cause distortion of the reconstructed light field scene. Currently, a few studies use optimization methods to find the optimal initial plane position, which takes a long time and cannot guarantee that the optimal initial plane position can be accurately found. Overall, the existing optimization methods are relatively complex, and a large amount of time is needed to analyze each picture, making it difficult to achieve real-time dynamic adjustment of the optimal initial plane position, making it difficult to meet the more efficient and accurate display optimization and improvement, and unable to meet the diversity of user needs and the promotion of compressed light field display. Therefore, the present application provides an intelligent optimization method for compressed three-dimensional light field display, which changes the optimization method to a deep learning method, which can greatly improve the display quality and optimization display processing time. SUMMARY

[0004] The present application provides an intelligent optimization method for compressed three-dimensional light field display, which improves the display quality of compressed three-dimensional light field display, so that for a given arbitrary light field source file, the optimal initial plane position can be obtained, and a more efficient and accurate high-quality compressed light field display image can be generated.

[0005] The purpose of the present application can be achieved by the following technical solutions:

[0006] An intelligent optimization method for compressed three-dimensional light field display, comprising the following steps:

[0007] S1, data acquisition: a light field camera acquires a light field original image, and performs an analysis operation on the image to obtain a different number of viewpoint images and depth map information through the above settings;

[0008] S2, offline training: input the multi-view image in S1 into the light field model to obtain the light field tensor. In the depth range of the depth map information in S1, the initial plane position is changed by an exhaustive method. The light field tensor and the initial plane position are input into the light field reconstruction model, and the number of display layers is set to perform light field reconstruction. The initial plane position is changed repeatedly to find the initial plane position corresponding to the best light field reconstruction quality, that is, the maximum PSANR value, as the optimal initial plane position. The depth map information and the optimal initial plane position are corresponded and packaged as a data set.

[0009] S3, online optimization: the training set in S2 is input, and the model outputs the optimal initial plane position value after learning, reconstructs the light field, and displays on the light field display.

[0010] Preferably, S1 includes the following steps:

[0011] S11, light field acquisition: use light field acquisition equipment such as a light field camera to collect scenes with different brightness, different scenes, and rich images with various contents, and add the currently widely used light field data set; the process of collecting the light field is the process of mapping the physical world coordinate system to the camera world coordinate system, as shown in formula 1, and the mathematical expression of the mapping function is:

[0012]

[0013] (x,y,1) T is a light field image matrix, K is a camera internal parameter determined by camera hardware, R represents a camera rotation angle, t is a camera translation parameter, (X,Y,Z,1) T is a space point determined by the scene structure.

[0014] S12, light field source file decoding, that is, solving accurate (x,y,1) T light field image matrix and corresponding field of view angle, generating a view point image: the file collected by the light field acquisition equipment such as a light field camera is stored in the format of LFR, LFP, and RAW, a general calibration is performed on different camera models with any lens number by using an existing model, a scale space analysis is performed on the micro image; the fitting method of the centroid network is used to determine the centroid spacing and projection mapping to ensure accurate decomposition of the sub-aperture image; the micro image is resampled to remove artifacts, the pixels are rearranged, the angle view position is adjusted, gamma is considered, and color correction is performed to decode the source file to obtain a view point image with high quality;

[0015] S13, light field depth map generation: the light field camera can simultaneously collect the spatial domain and angle domain information of the spatial light, and the depth of the spatial scene is recovered according to this characteristic. The light field depth map can be obtained by using a CNN model.

[0016] Preferably, S2 comprises the following steps:

[0017] S21, finding the optimal initial plane position: analyzing the light field depth map to obtain the maximum depth and the minimum depth, setting the minimum depth as the initial plane position. Modeling each layer of LCD as a spatially controllable polarization rotator, using the SART algorithm, setting the number of display layers, applying the best spatially varying polarization state rotation to each layer for tomographic solution, reconstructing the light field, and storing the PSNR value measuring the quality of the reconstructed light field; then taking a fixed step, constantly increasing the value of the initial plane position, reconstructing the light field until the initial plane position value is equal to the maximum depth, finding the initial plane position corresponding to the maximum PSNR value, defining it as the optimal initial plane position. This patent takes a three-layer compression display as an example;

[0018] The reconstructed light field model includes light field reconstruction and light field decomposition, wherein the light field reconstruction is as shown in formula 2, and the light field decomposition process is as shown in formulas 3, 4 and 5:

[0019]

[0020]

[0021]

[0022]

[0023]

[0024] wherein L is a light field image tensor composed of (x, y, 1) T , a field of view angle, W is a weight matrix, A, B and C are pixels of the three layers of liquid crystals before updating; A', B' and C' are pixels of the three layers of liquid crystals after updating; W (1) ,W (2) ,W (3) respectively represent the slices of W along the 1st, 2nd and 3rd dimensions, L (1) ,L (2) ,L (3) respectively represent the slices of L along the 1st, 2nd and 3rd dimensions, and d represents the Hadamard product.

[0025] Therefore, the light field image tensor L, the number of display layers and the initial plane position need to be input when the light field is reconstructed. The light field image tensor L can be obtained from the light field multi-view generated by S12 through the generated light field model.

[0026] ​Equations 6, 7, 8, 9 are the specific process of finding the optimal initial plane position of the exhaustive method combined with the reconstruction of the light field model. Among them, Mindepth is the minimum depth value, Maxdepth is the maximum depth value, Initial plane represents the initial plane position, and equation 6 represents the value range of the initial plane position; Equation 7 is to assign the minimum depth value to the initial plane position; In equation 8, Initial plane* represents the updated initial plane position, which is equal to the sum of the initial plane position of the last time and the step; Equation 9 represents the calculation of the step, and LOOP is the number of times of searching for the optimal initial plane position in a loop, which can be set according to the required accuracy.

[0027] Mindepth ≤ Initial plane ≤ Maxdepth Equation 6

[0028] Initial plane = Mindepth Equation 7

[0029] Initial plane* = Initial plane + step Equation 8

[0030] step = (Maxdepth - Mindepth) / LOOP Equation 9

[0031] S22, data set preparation: pack the depth map of the content to be displayed and the corresponding optimal initial plane position into a data set.

[0032] Preferably, S3 comprises the following steps:

[0033] S31, light field data preparation: from the data set in S22, select data for different brightness scenes, different scene scenes, and multiple attribute scenes, and randomly select 80% and 20% of the images as the training set and the verification set;

[0034] S32, optimal initial plane position prediction model training: based on a convolutional neural network, for the relationship between the depth distribution of the image and the correlation between the image features and the reconstruction quality, a multi-task learning model is constructed, and the training set obtained in S31 is used for model training until convergence and high accuracy is obtained on the validation set; the prediction model inputs the data set obtained in S22, and outputs the predicted optimal initial plane position. The prediction model structure mainly includes two parts. The first part is the feature extraction part, which is composed of convolutional layers and pooling layers. The convolutional layers are three layers, which are 5*5, 5*5 and 3*3 respectively. A 2*2 max pooling layer is periodically inserted between the consecutive convolutional layers, which can gradually reduce the spatial size of the data body, reduce the number of parameters in the network, reduce the consumption of computing resources, and effectively control overfitting. The second part is the prediction part composed of full connection, and the number of neurons in the three full connection layers is 2592, 1024 and 256 respectively.

[0035] The model selects an ELU activation function, and its expression is shown in formula 10:

[0036] ELU(x) = max(0, x) + min(0, a*(exp(x)-1)) Formula 10

[0037] Adding an activation function can make the neural network arbitrarily approximate any nonlinear function, thereby achieving better regression, where x is the output vector of the convolutional layer, that is, the input vector of the full connection layer, and ELU (x) is a regression function, and a plurality of ELU (x) functions are combined to obtain the final regression function model.

[0038] In order to improve the accuracy of the prediction model, the MSE function is used as the loss, and its expression is shown in formula 11:

[0039]

[0040]

[0041] where y m is the real optimal reconstruction initial plane position, y m is the predicted optimal reconstruction initial plane position, and the smaller the value of the test set loss, the closer the predicted optimal reconstruction initial plane position is to the real value. If the PSNR value of the corresponding reconstructed image also has a large improvement at this time, generally the learning effect of the model is better.

[0042] Because the number of learning samples is small, the model selects RAdam as the optimizer, which accelerates model convergence through adaptive learning rate variance. In addition, we use an exponentially decaying learning rate in each round. When the training round increases, the learning rate decays, accelerating the model to converge to a better model. ​

[0043] S33, optimal initial plane position model test: for a given input image, input its depth image into the optimal initial plane position prediction model trained in S32 to obtain the predicted optimal initial plane position, input the predicted initial plane position into the light field reconstruction model used in S21, and the PSNR value obtained after reconstructing the light field is the largest compared with the PSNR values reconstructed by other initial plane positions, which means that the model can successfully predict the optimal initial plane position.

[0044] S34, light field display: set the number of display layers, input the optimal initial plane position prediction model with any light field depth image to obtain the optimal initial plane position; according to the optimal initial plane position and the light field tensor, set the display layer number, reconstruct the light field, and obtain the image displayed by each layer of the display, and each layer of the display displays the corresponding generated reconstructed image.

[0045] Compared with the prior art, the intelligent optimization method for compressing three-dimensional light field display provided by the application constructs the corresponding optimal initial plane position prediction model for light field images of different scenes, so that the light field reconstruction quality of different scenes is better improved. The optimal initial plane position prediction model is used to automatically and efficiently generate the optimal initial plane position, so that the final better light field reconstruction and display results are obtained, and the whole process is more automated, intelligent and efficient. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 is a flowchart of the application.

[0047] Figure 2 is the display image after optimal initial plane position processing of the application.

[0048] Figure 3 is the loss graph of the validation set of the optimal initial plane position prediction model.

[0049] Figure 4 is the loss graph of the test set of the optimal initial plane position prediction model.

[0050] Figure 5 is the light field reconstruction graph of the light field data set collected under the fixed initial plane position and the tested optimal initial plane position.

[0051] Figure 6 is the light field reconstruction graph of the MITSynthetic data set collected under the fixed initial plane position and the tested optimal initial plane position.

[0052] Figure 7 is the light field reconstruction graph of the inria data set collected under the fixed initial plane position and the tested optimal initial plane position. DETAILED DESCRIPTION

[0053] The following is a specific embodiment of the present application and further describes the technical solutions of the present application in conjunction with the drawings, but the present application is not limited to these embodiments.

[0054] As shown in Figure 1 , Figure 2 , the embodiment provides an intelligent optimization method for compressing a three-dimensional light field display, including the following steps:

[0055] S1, data acquisition: a light field camera or other light field acquisition equipment acquires a light field original image, analyzes the image, and obtains different numbers of view graphs and depth graph information by setting; the specific steps are as follows:

[0056] S11, light field acquisition: use a light field camera to acquire rich images of day, night, close-up, long shot, simple scene, and complex scene, and add the currently widely used light field data set; mainly inria, MIT Synthetic, etc., 80% of the camera acquisition and mainstream light field data set are randomly selected as the training set of the model, and the remaining 20% are used as the verification set; wherein the light field acquisition principle is shown in formula 1, and the mathematical expression of the mapping function is:

[0057]

[0058] (x,y,1) T is a light field image matrix, K is a camera internal parameter determined by camera hardware, R represents a camera rotation angle, t is a camera translation parameter, (X,Y,Z,1) T is a space point determined by the scene structure.

[0059] S12, light field source file decoding, that is, solving accurate (x,y,1) TThe light field image matrix and the corresponding field of view angle generate a viewpoint graph: the file storage format collected by the light field camera is LFR, LFP and RAW, the existing model is used to calibrate different camera models with any lens number, the scale space analysis is carried out on the micro image; the fitting method of the centroid network is used to determine the centroid spacing and the projection mapping, the centroid network Levenberg-Marquardt optimization algorithm is used to globally reduce the least square error of the detected micro lens center, and the centroid spacing and the projection mapping are determined at the same time, so that the sub-aperture image is accurately decomposed; for accurate angle sampling, we provide micro image resampling, then remove the hexagonal artifact, rearrange the pixels, adjust the angle view position, consider gamma and color correction, so as to decode the source file to obtain a viewpoint graph with high quality, the viewpoint graph generation model of the application mainly learns the optimization algorithm in the PlenoptiCam model, and the unnecessary rendering module in the PlenoptiCam model is cut, mainly to speed up the decoding process.

[0060] S13, light field depth map generation: the light field camera can simultaneously collect the spatial domain and angle domain information of the spatial light, and the depth of the spatial scene is recovered according to this characteristic, and the light field depth map can be obtained by using the CNN model.

[0061] S2, offline training: in the depth map information range of S1, the initial plane position corresponding to the best reconstruction quality is found out by using the exhaustive method, and the depth map information and the corresponding optimal initial plane position are used as the training set; the specific steps are as follows:

[0062] S21, generate light field file, that is, compress the viewpoint graph information and then generate light field tensor: taking a 7x7 viewpoint graph as an example, first reduce the resolution of the viewpoint graph to 96x96, so as to reduce the generation and reconstruction time of the light field, then input the viewpoint graph into the light field model to obtain the field of view angle, pixel and other information of the light field, and generate the light field tensor and store it as a mat file;

[0063] S22, reconstruct the light field: according to the physical properties of the light field display, set the number of layers and the layer spacing in the light field model; analyze the light field depth map to obtain the maximum depth and the minimum depth, and set the minimum depth as the initial plane position; input the initial plane position and the file generated in S21 into the light field reconstruction model, output and store the PSNR value for measuring the quality of the reconstructed light field and the depth value of the initial plane position; then take a fixed step, constantly increase the value of the initial plane position, and reconstruct the light field until the depth value of the initial plane position is equal to the maximum depth, find the initial plane position corresponding to the maximum PSNR value, and define it as the optimal initial plane position; the patent takes a three-layer compression display as an example;

[0064] The reconstructed light field model includes light field reconstruction and light field decomposition, where light field reconstruction is as follows:

[0065] Equation 2 and the process of light field decomposition are shown in Equations 3, 4, and 5:

[0066]

[0067]

[0068]

[0069]

[0070] Where L is the light field image tensor, derived from (x, y, 1) T The image consists of a light field matrix and a field of view, where W is the weight matrix, A, B, and C are the pixels of the three liquid crystal layers before the update, and A'B'C' are the pixels of the three liquid crystal layers after the update; W (1) W (2) W (3) These represent slices taken along the 1st, 2nd, and 3rd dimensions of W, respectively, and L. (1) ,L (2) ,L (3) d represents slices taken along the 1st, 2nd, and 3rd dimensions of L, respectively, and d represents the Hadamard product.

[0071] Therefore, when reconstructing the light field, it is necessary to input the light field image tensor L, the number of display layers, and the initial plane position. The light field image tensor L can be obtained from the multi-view light field generated by S12 through the generated light field model.

[0072] Formulas 6, 7, 8, and 9 describe the specific process of using the exhaustive method to jointly reconstruct the light field model and find the optimal initial plane position. Here, Mindepth is the minimum depth value, Maxdepth is the maximum depth value, Initial plane represents the initial plane position, and Formula 6 indicates the range of values ​​for the initial plane position. Formula 7 assigns the minimum depth value to the initial plane position. In Formula 8, Initial plane* represents the updated initial plane position, which is equal to the sum of the previous initial plane position and the step size. Formula 9 calculates the step size, and LOOP represents the number of iterations to find the optimal initial plane position; this patent uses LOOP as an example of 100.

[0073] Mindepth ≤ Initial plane ≤ Maxdepth Formula 6

[0074] Initial plane = Mindepth Formula 7

[0075] Initial plane* = Initial plane + step Equation 8

[0076] step = (Maxdepth - Mindepth) / LOOP Equation 9

[0077] S23, Dataset preparation: Pack the light field depth map and the corresponding optimal initial plane position into a dataset.

[0078] S3, Online optimization: Take the training set in S2 as input, and the model will output the optimal reconstruction initial plane initial position depth after learning; the specific steps are as follows:

[0079] S31, Light field data preparation: From the S23 dataset, select data for different brightness scenes, different scene scenes, and content diverse scenes with multiple attributes, and randomly select 80% and 20% of the images as the training set and the validation set;

[0080] S32, Optimal initial plane position prediction model training: Based on the convolutional neural network, the relationship between the depth distribution of the image and the image features and the reconstruction quality is related, a multi-task learning model is constructed, and the training set obtained in S31 is used for model training until convergence and high accuracy is obtained on the validation set; in a specific implementation case, the prediction model structure is mainly divided into two parts, the first part is the feature extraction part, which is composed of convolutional layers and pooling layers, among which the convolutional layers are three layers, which are 5*5, 5*5, and 3*3, respectively. A 2*2 max pooling layer is periodically inserted between the consecutive convolutional layers, which can gradually reduce the spatial size of the data body, reduce the number of parameters in the network, reduce the consumption of computing resources, and effectively control overfitting. The second part is the prediction part composed of full connection, and the number of neurons in the three full connection layers is 2592, 1024, 256, and 1, respectively.

[0081] The model selects the ELU activation function, and its expression is shown in Equation 10:

[0082] ELU(x) = max(0, x) + min(0, a*(exp(x)-1)) Equation 10

[0083] Adding an activation function can make the neural network arbitrarily approximate any nonlinear function, thereby achieving better regression, where x is the output vector of the convolutional layer, that is, the input vector of the full connection layer, ELU (x) is a regression function, and a combination of multiple ELU (x) functions can obtain the final regression function model.

[0084] Because the number of learning samples is small, the model selects RAdam as the optimizer to accelerate model convergence through adaptive learning rate variance. In addition, we use an exponentially decaying learning rate in each round. When the training round increases, the learning rate decays, accelerating the model to converge to a better model.

[0085] In order to improve the accuracy of the prediction model, the MSE function is used as the loss, and its expression is shown in formula 11:

[0086]

[0087] Where y m is the real optimal reconstruction initial plane position, is the predicted optimal reconstruction initial plane position. The smaller the value of the test set Loss, the closer the predicted optimal reconstruction initial plane position is to the real value. The Loss of the embodiment is shown in Figure 3 、 4 , the horizontal coordinate is the training times, and the vertical coordinate is the loss value. It can be seen that the loss value is stable near 1.88, indicating that the prediction performance of the model is good and can accurately predict the optimal initial plane position.

[0088] S33, optimal initial plane position model test: for a given input image, input its depth image into the optimal initial plane position prediction model trained in S32 to obtain the predicted optimal initial plane position. Input the predicted initial plane position into the light field reconstruction model used in S21, and the PSNR value obtained after reconstructing the light field is the largest compared with the PSNR value reconstructed by other initial plane positions, which means that the model can successfully predict the optimal initial plane position. The specific implementation is shown in Figure 5 、 6 , 7, respectively, using the collected light field dataset, MIT Synthetic dataset, and inria dataset as examples, the PSNR value of light field reconstruction. The vertical coordinate is the PSNR value, and the horizontal coordinate is the picture number. Compared with other initial plane positions, the PSNR value using the predicted optimal initial plane position is increased by 4db, and the average is increased by 0.7db, while the PANR value is increased by 0.3db, which is considered to be optimized and effective, so it is sufficient to show that the method is very effective in optimization.

[0089] S34, light field display: set the number of display layers, input any light field depth image into the optimal initial plane position prediction model to obtain the optimal initial plane position; according to the optimal initial plane position and the light field tensor, set the display layer number, reconstruct the light field, and obtain the image displayed by each layer of the display. The display corresponding to the generated reconstruction image is shown in Figure 2 , the display image after optimal initial plane position processing.

[0090] The specific embodiments described herein are merely illustrative of the spirit of the application. Various modifications or changes in the specific embodiments described herein can occur to those skilled in the art to which the application pertains without departing from the spirit of the application, and it is understood that such modifications or changes are to be considered as within the scope of the application as defined by the appended claims.

Claims

1. An intelligent optimization method for compressing a three-dimensional light field display, comprising the following steps: S1, data acquisition, acquiring a light field original image using a light field acquisition device, and performing an analysis operation on the image to obtain multi-view images and depth map information of different scenes; S2, offline training: input the multi-view images in S1 into a light field model to obtain a light field tensor, within the depth range of the depth map information in S1, change the initial plane position using an exhaustive method, input the light field tensor and the initial plane position into a light field reconstruction model, set the number of display layers, perform light field reconstruction, and continuously change the initial plane position to find the initial plane position corresponding to the best light field reconstruction quality, i.e., the maximum PSNR value, as the optimal initial plane position; correspond the depth map information and the optimal initial plane position, and pack them into a data set; S3, online optimization: dynamically find the optimal initial plane position and then display, for different scene light field source files, input the optimal initial plane prediction model, wherein the optimal initial plane prediction model is based on a convolutional neural network, and is related to the relationship between the depth distribution of the image and the image features and the reconstruction quality, a multi-task learning model is constructed, the input of which is the data set, and the output is the predicted optimal initial plane position, the optimal initial plane prediction model structure mainly includes two parts, the first part is a feature extraction part, which is composed of convolutional layers and pooling layers, wherein the convolutional layers are three layers, which are 5*5, 5*5, and 3*3, respectively, a 2*2 max pooling layer is periodically inserted between the consecutive convolutional layers, which can gradually reduce the spatial size of the data body, reduce the number of parameters in the network, reduce the consumption of computing resources, effectively control overfitting, and the second part is a prediction part composed of fully connected layers, the number of neurons of the three fully connected layers is 2592, 1024, 256, and 1, respectively; the optimal initial plane prediction model can calculate the optimal initial plane position in a short time, reconstruct the light field, and display it on the light field display. 2.The intelligent optimization method of compressive three-dimensional light field display according to claim 1, wherein, S1 includes the following steps: S11, light field acquisition: use a light field acquisition device to acquire images with different brightness scenes, different scenes, and rich content, and add the currently widely used light field data set; wherein the process of acquiring the light field is the process of mapping the physical world coordinate system to the camera world coordinate system, as shown in formula 1, which is the mathematical expression of the mapping function: Formula 1 ; is the light field image matrix, K is the camera intrinsic parameter determined by the camera hardware, R represents the camera rotation angle, t is the camera translation parameter, is the spatial point determined by the scene structure; S12, light field source file decoding, i.e. solving accurate Light field image matrix and corresponding field of view angle, generating multi-viewpoint image: the file storage format collected by the light field acquisition device is LFR, LFP and RAW, the existing model is used to calibrate different camera models with arbitrary lens number, and the micro image is analyzed in scale space; the fitting method of the centroid network is used to determine the centroid spacing and projection mapping, so as to ensure the accurate decomposition of the sub-aperture image; the micro image is resampled to remove artifacts, the pixels are rearranged, the angle view position is adjusted, gamma and color correction are adopted, so as to decode the source file to obtain multi-viewpoint image with high quality; S13, light field depth map generation: the light field acquisition device can simultaneously acquire the spatial domain and angular domain information of the spatial light, according to this characteristic, the depth of the spatial scene is recovered, and the light field depth map can be obtained using a CNN model. 3.The method of claim 1, wherein, S2 includes the following steps: S21, finding the optimal initial plane position: analyzing the light field depth map to obtain the maximum depth and the minimum depth, setting the minimum depth as the initial plane position, modeling each layer of LCD as a spatially controllable polarization rotator, using the SART algorithm, setting the number of display layers, applying the best spatially varying polarization state rotation to each layer for tomographic solving, reconstructing the light field, and storing the PSNR value measuring the quality of the reconstructed light field; then taking a fixed step, constantly increasing the value of the initial plane position, reconstructing the light field, until the value of the initial plane position is equal to the value of the maximum depth, finding the value of the initial plane position corresponding to the maximum PSNR value, and defining it as the optimal initial plane position; The reconstructed light field model includes light field reconstruction and light field decomposition. The light field reconstruction is shown in formula 2, and the light field decomposition process is shown in formulas 3, 4, and 5: Formula 2: Equation 3; Formula 4; Equation 5; L is a light field image tensor, which is composed of a light field image matrix and a field of view angle, W is a weight matrix, A, B, C are pixels of the three layers of liquid crystals before updating; A', B', C' are pixels of the three layers of LCD after updating; , , W1, W2, W3 respectively represent slices of W along the 1st, 2nd, 3rd dimensions, , , L1, L2, L3 respectively represent slices of L along the 1st, 2nd, 3rd dimensions, represents a Hadamard product; When performing light field reconstruction, the light field image tensor L, the number of display layers, and the initial plane position need to be input. The light field image tensor L can be obtained from the generated light field model of the light field multi-view generated by S12, Formulas 6, 7, 8, and 9 are the specific process of finding the optimal initial plane position by the exhaustive method combined with the light field reconstruction model. Mindepth is the minimum depth value, Maxdepth is the maximum depth value, Initial plane represents the initial plane position, and formula 6 represents the value range of the initial plane position. Formula 7 assigns the minimum depth value to the initial plane position. In formula 8, Initialplane * represents the updated initial plane position, which is equal to the sum of the initial plane position of the last time and the step size. Formula 9 represents the calculation of the step size, and LOOP is the number of times of searching for the optimal initial plane position in a loop, which can be set according to the required accuracy. Equation 6; Equation 7; Formula 8: Equation 9; S22, data set preparation: packing the depth map of the content to be displayed and the corresponding optimal initial plane position into a data set.

4. The intelligent optimization method for compressive three-dimensional light field display according to claim 1, wherein, S3 includes the following steps: S31, light field data preparation: selecting data for scenes with different brightness, different scenes, different content, and different attributes from the data set in S22, and randomly selecting 80% and 20% of the images as the training set and the validation set; S32, optimal initial plane position prediction model training: using the training set obtained in S31 to train the model until it converges and achieves high accuracy on the validation set; The model selects An activation function, whose expression is shown in Equation 10: Formula 10; Adding activation function can make the neural network can arbitrarily approximate any nonlinear function, so as to realize better regression, wherein x is the output vector of convolution layer, that is, the input vector of full connection layer, For regression function, a plurality of The function combination can obtain the final regression function model, To improve the accuracy of the prediction model, the MSE function is used as the loss, and its expression is shown in formula 11: Formula 11; wherein is the true optimal reconstruction initial plane position, is the predicted optimal reconstruction initial plane position, the smaller the value of test set Loss, the closer the predicted optimal reconstruction initial plane position is to the true value; S33, optimal initial plane position model testing: for a given input image, input its depth image into the optimal initial plane position prediction model trained in S32 to obtain the predicted optimal initial plane position. Input the predicted initial plane position into the light field reconstruction model used in S21, and the PSNR value obtained after reconstructing the light field is the largest compared to other initial plane positions, which means that the model can successfully predict the optimal initial plane position. S34, light field display: input an arbitrary light field depth map to an optimal initial plane position prediction model to obtain an optimal initial plane position; according to the optimal initial plane position and a light field tensor, set a display layer number, reconstruct a light field, and obtain an image displayed by each layer of display, and each layer of display displays a corresponding generated reconstructed image.

Citation Information

Patent Citations

  • Methods for full parallax compressed light field 3D imaging systems

    CN106068645A

  • Algebraic iteration method and device for reestablishing optical field through focus stack under incomplete data

    CN106937044A