A light field 3D display slice source generation method based on deep learning

CN122597644APending Publication Date: 2026-08-18TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610705672.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

然而,现有方法在深度图的预测中受限于局部感受野,难以充分保留深度估计所需的局部细节信息,导致预测的深度信息不准确,影响EIA的生成质量

Benefits of technology

[0043] 1. This invention predicts high-precision depth maps through neural networks and achieves efficient generation of EIA through a geometric inverse mapping mechanism guided by depth information;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597644A_ABST
    Figure CN122597644A_ABST
Patent Text Reader

Abstract

The application discloses a light field 3D display slice source generation method based on deep learning, and the method comprises the following steps: constructing a multi-scale image feature learning module based on a double-branch semantic guide, and extracting multi-scale features with global structure information and local detail information; using a multi-scale step-by-step fusion decoding module to obtain a high-precision depth map suitable for light field 3D display slice source generation; constructing a depth-guided geometric inverse mapping EIA generation module, converting the solving of mapping coordinates into a fixed point problem, so as to improve the quality and efficiency of light field 3D display slice source generation; finally, using L1 loss to optimize the training process of the light field 3D display slice source generation network, so as to generate high-quality light field 3D display slice source; the application can effectively improve the 3D display effect of the light field 3D display system, shorten the generation time of the light field 3D slice source, and improve the real-time processing capability of the light field 3D display system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of light field 3D display source material, and more particularly to a method for generating light field 3D display source material based on deep learning. Background Technology

[0002] Light field 3D displays offer advantages such as full parallax and quasi-continuous viewing viewpoints, attracting significant attention from researchers and becoming a hot research topic in the field of 3D displays. Despite continuous progress in the hardware of light field 3D display technology, the acquisition of high-quality light field 3D display source materials (Elemental Image Array, EIA) remains a key bottleneck restricting its development.

[0003] Traditional EIA (Easy Aspect) acquisition typically relies on camera arrays or light field cameras to collect light field information from 3D scenes. However, these acquisition systems suffer from drawbacks such as large size, complex calibration, and limited dynamic range. To address these issues, researchers have proposed an EIA generation method based on a depth camera. This method generates EIA by establishing a mapping relationship between object surface pixels and the EIA plane, effectively simplifying the light field information acquisition process. However, while the introduction of a depth camera simplifies the process, it also increases the system's hardware cost, and its performance is susceptible to changes in ambient lighting and differences in object materials.

[0004] In recent years, deep learning, with its powerful nonlinear feature learning capabilities, has provided a new technical path for EIA generation. Ren et al. used convolutional neural networks to predict depth maps from single color images and then generated EIAs through pixel mapping algorithms. Jung et al. introduced deep neural networks with multi-view attention modules to predict depth maps and combined them with hierarchical virtual lens arrays to generate large field-of-view EIAs. Ma et al. used a depth generator to predict depth maps, obtained low-resolution EIAs through pixel mapping algorithms, and then used a super-resolution network to obtain high-resolution EIAs. However, existing methods are limited by the local receptive field in depth map prediction, making it difficult to fully retain the local detail information required for depth estimation, resulting in inaccurate predicted depth information and affecting the quality of generated EIAs.

[0005] In addition, existing methods suffer from insufficient accuracy in obtaining mapping coordinates and high computational cost during pixel mapping, which seriously affects the generation quality and efficiency of EIA. Summary of the Invention

[0006] This invention provides a method for generating light field 3D display source material based on deep learning. The method utilizes a multi-scale image feature learning module guided by dual-branch semantics to fully extract multi-scale features that combine global structural information and local detail information, thereby obtaining high-precision depth information and improving the generation quality of the image image inversion (EIA). Furthermore, a depth-guided geometric inverse mapping (EIA) generation module transforms the acquisition of mapping coordinates into fixed-point solving, effectively reducing computational load while improving the accuracy of mapping coordinates, thus enhancing the generation quality and efficiency of the EIA. See the description below for details.

[0007] A method for generating light field 3D display source material based on deep learning, the method comprising:

[0008] A light field 3D display source generation network is constructed, consisting of a multi-scale image feature learning module guided by dual-branch semantics, a multi-scale hierarchical fusion decoding module, and a depth-guided geometric inverse mapping (EIA) generation module.

[0009] The training process of the light field 3D display source generation network is optimized by using L1 loss, resulting in an optimized light field 3D display source generation network.

[0010] Input image information and generate light field 3D display source based on the optimized light field 3D display source generation network.

[0011] The multi-scale image feature learning module consists of a main branch, a detail branch, and a semantically guided shallow fusion unit.

[0012] The main branches employ a hierarchical feature encoder based on visual state space to extract shallow target features containing global structural information. and deep semantic features ;

[0013] The detail branch extracts local detail features from the image through a local detail encoder. ;

[0014] Semantic-guided shallow fusion units for local detail features and deep semantic features Perform channel alignment to integrate shallow target features. Local detail features after channel alignment Deep semantic features aligned with channels A common input gating mapping function is used to construct semantically gating guided weights. ;

[0015] Local detail features after channel alignment Deep semantic features aligned with channels By fusing the data, preliminary enhanced features are obtained, consisting of global structural information and local detailed information. ;

[0016] Using gated residual injection to analyze shallow target features Preliminary enhanced features of global structural information and local detailed information The features are then fused and refined using residual blocks to obtain enhanced features with both global structural information and local detail information. .

[0017] Among them, the enhanced features for:

[0018]

[0019] in, This represents the learnable scaling factor. This represents element-wise multiplication. This represents a residual block.

[0020] The depth-guided geometric inverse mapping (EIA) generation module is as follows:

[0021] This will be suitable for generating depth maps for EIA. Fit to surface equation By using geometric inverse mapping, the linear equation of a ray emitted from a pixel in EIA is obtained by passing it through the center of the corresponding microlens. ;

[0022] Solving the surface equation With linear equations The intersection points between the points determine the mapping coordinates of the pixel in the color image. The process iterates through each pixel in the EIA and calculates its pixel value to generate a high-quality EIA.

[0023] The aforementioned method will be suitable for generating depth maps for EIA. Fit to surface equation By using geometric inverse mapping, the linear equation of a ray emitted from a pixel in EIA is obtained by passing it through the center of the corresponding microlens. for:

[0024] The depth map is reconstructed using bicubic B-spline surfaces. Fit to a surface equation According to the geometric inverse mapping, the coordinates of a certain pixel on EIA are: The emitted light passes through the center of the corresponding microlens. The linear equation of the ray is obtained. , is represented as:

[0025]

[0026] in, This represents the distance between the microlens array and the EIA. These are the depth coordinates in three-dimensional space. When the pixel point on EIA... When an emitted ray intersects a depth surface, the intersection point satisfies both the linear equation of the ray and the surface equation of the depth map. The depth at the intersection point is... satisfy:

[0027]

[0028] The problem of finding the intersection point of a ray and a depth surface is transformed into solving for a fixed point, where the depth value is the intersection point. satisfy:

[0029]

[0030] in, This represents a fixed-point function.

[0031] Among them, the solution of the surface equation With linear equations The intersection point between them is:

[0032] We introduce Halley iteration based on the second derivative, using the center depth plane as the initial value. In the In the next iteration, based on the current depth value Inversely calculate the coordinates of the intersection point between the ray equation and the depth surface. for:

[0033]

[0034]

[0035] in, This represents the coordinates of the pixel to be determined in EIA. This represents the coordinates of the center of the corresponding microlens; the depth value at that point is obtained through a fixed-point function, and a residual function is constructed:

[0036]

[0037] in, This represents the depth value obtained by substituting the coordinates of the current candidate intersection point into the depth surface;

[0038] The Halley iterative update formula based on the second derivative is:

[0039]

[0040] in, For the first The depth value of the next iteration. and They represent the residual functions at... The first and second derivatives at point ;

[0041] When the change in depth value between two consecutive iterations satisfies Stop iteration when This is a preset threshold.

[0042] The beneficial effects of the technical solution provided by this invention are:

[0043] 1. This invention predicts high-precision depth maps through neural networks and achieves efficient generation of EIA through a geometric inverse mapping mechanism guided by depth information;

[0044] 2. This invention constructs a multi-scale image feature learning module based on bi-branch semantic guidance. By collaboratively modeling the global context information in multi-scale features, reliable depth estimation results are obtained, thereby improving the generation quality of EIA.

[0045] 3. This invention designs a depth-guided geometric inverse mapping (EIA) generation module, which transforms the solution of mapping coordinates into a fixed-point problem, significantly reducing computational complexity while improving mapping accuracy, thereby effectively improving the quality and efficiency of EIA generation.

[0046] 4. The high-quality EIA output of this invention can provide a light field data source for the light field 3D display system, thereby improving the display effect of the light field 3D display system; the high-efficiency EIA generation can shorten the light field 3D source generation time, reduce the computing burden on the device, and improve the real-time processing capability of the light field 3D display system. Attached Figure Description

[0047] Figure 1 A flowchart of a method for generating light field 3D display source material based on deep learning;

[0048] Figure 2 This is a comparison chart of the EIA reconstruction results of the method of this invention and the traditional method. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below.

[0050] The definitions of the modules and functions involved in the embodiments of this invention are given below, as detailed in the following description:

[0051] 1. Hierarchical feature encoder based on visual state space:

[0052] The hierarchical feature encoder consists of a patch embedding layer, four visual state space encoding stages, a downsampling layer, and a multi-scale feature output layer.

[0053] Given an input color image First, the initial image patch features are mapped to initial image patch features through a patch embedding layer, reducing the feature resolution to 1 / 4 of the input image. Then, the initial image patch features are sequentially input into four visual state space encoding stages. Each stage uses a VSS Block to perform global context modeling and hierarchical feature encoding on the initial image patch features, enhancing the expressive power of the image's global structural information. Between adjacent visual state space encoding stages, downsampling layers progressively reduce the spatial resolution and increase the channel dimension. Finally, a multi-scale feature output layer outputs shallow target features. and deep semantic features .

[0054] 2. Local detail feature encoder:

[0055] The local detail encoder takes a color image as input and first performs initial local feature mapping through a convolutional layer, followed by a nonlinear transformation using a GELU activation function. The activated features are then input into a residual block to obtain initial local detail features. These initial local detail features are further encoded through two levels of local detail coding layers. Each level first performs downsampling and channel mapping through a convolutional layer, then enhances the nonlinear expressive power using a GELU activation function, before inputting into the residual block for local feature refinement. Finally, channel attention (SE) is used to adaptively recalibrate the importance of different channels. Ultimately, the local detail encoder outputs local detail features at two different scales. .

[0056] 3. Semantic-guided shallow fusion unit:

[0057] Specifically, the semantically guided shallow fusion unit first processes local detail features through a convolutional layer. and deep semantic features Perform scale and channel alignment; then, align the shallow target features. The aligned local detail features and aligned deep semantic features are concatenated through channels, and semantic gating guidance weights are generated using a gating mapping function and an activation function (Sigmoid). Subsequently, the aligned local detail features and deep semantic features are fused through a fusion mapping function to obtain preliminary enhanced features. Finally, under the control of the semantic gating guidance weights, the preliminary enhanced features are injected into the shallow target features as residuals, and further refined through residual blocks to output enhanced features with both global structural information and local detail information. .

[0058] 4. Residual Block :

[0059] Among them, residual blocks It consists of depthwise convolutions, pointwise convolutions, normalization layers, nonlinear activation functions, and residual connections. Given input features... The specific formula is as follows:

[0060]

[0061] in, This represents the channel-reduction convolution operation. This indicates a channel-level up-dimension convolution operation. This represents the GELU activation function. Presentation layer normalization operation, This represents a 3×3 depthwise convolution.

[0062] 5. Gating mapping function :

[0063] Among them, the gated mapping function This generates mappings for learnable gated weights composed of convolutional layers and nonlinear activation functions. Given input features... The specific formula is as follows:

[0064]

[0065] in, This represents a 1×1 convolution operation. This represents a 3×3 convolution operation. This represents the GELU activation function.

[0066] 6. Fusion mapping function :

[0067] Among them, the fusion mapping function This is a learnable fusion mapping consisting of fused convolutions, residual blocks, and SE channel attention. Given input features... The specific formula is as follows:

[0068]

[0069] in, This indicates the attention of the SE channel. Represents the residual block. This represents a 1×1 convolution operation.

[0070] 7. Decoding Unit :

[0071] Among them, the decoding unit It consists of upsampling, skip connection concatenation, channel fusion convolution, two residual blocks, and SE channel attention. Given the current decoding features... and corresponding scale of skip connection features The specific formula is as follows:

[0072]

[0073] in, This indicates the attention of the SE channel. Represents the residual block. This represents a 1×1 convolution operation. Indicates an upsampling operation. This indicates a feature splicing operation.

[0074] 8. VSS Block: The input features of the VSS Block first pass through a normalization layer, and then the normalized features are input into the state space modeling branch and the gating branch. The state space modeling branch consists of a linear mapping layer, a deep convolutional layer, an activation function (SiLU), an SS2D module, and a normalization layer. The gating branch consists of a linear mapping layer and an activation function (SiLU) to generate adaptive modulation weights. Finally, the output of the state space modeling branch is multiplied element-wise with the weights generated by the gating branch, and the channel dimension is restored through the output mapping layer. Then, a residual connection is performed with the input features to obtain the output features of the VSS Block.

[0075] 9. SS2D module:

[0076] The SS2D module first unfolds the input features into a one-dimensional sequence along different spatial directions, including the horizontal direction, the reverse horizontal direction, the vertical direction, and the reverse vertical direction. Then, it performs selective state space scanning on each direction sequence. Next, it restores the sequence features obtained from each direction scan into a two-dimensional feature map and fuses the multi-directional features to obtain the output features of the SS2D module.

[0077] 10. GELU activation function :

[0078] Among them, the GELU activation function This is used to perform nonlinear mapping on input features, enhancing the network's ability to express complex feature relationships. Given input features... The specific formula is as follows:

[0079]

[0080] in, This represents the integral variable.

[0081] 11. SE Channel Attention :

[0082] Among them, SE channel attention The input features first pass through a global average pooling layer, compressing the spatial features of each channel into a channel description value. Then, the obtained channel description value is input into the first fully connected layer, and the feature is transformed by the activation function (GELU). Next, the second fully connected layer is used to generate channel attention weights in the range of (0,1) using the activation function (Sigmoid). Finally, the channel attention weights are multiplied with the original input features channel by channel to achieve channel recalibration of the input features.

[0083] 12. Sigmoid activation function :

[0084] Wherein, the Sigmoid activation function Used to map input features to the range (0,1). Given input features The specific formula is as follows:

[0085]

[0086] 13. SiLU activation function :

[0087] Wherein, the SiLU activation function This is used to perform nonlinear mapping on input features, enhancing the network's feature representation capability. Given input features... The specific formula is as follows:

[0088]

[0089] The following examples illustrate the specific implementation of the deep learning-based integrated imaging 3D video source generation method in this invention.

[0090] I. Constructing a multi-scale image feature learning module based on dual-branch semantic guidance

[0091] A multi-scale image feature learning module based on dual-branch semantic guidance is constructed, which consists of a main branch, a detail branch, and a semantically guided shallow fusion unit.

[0092] Specifically, given an input color image color image Simultaneously inputting the main branch and detail branches. The main branch employs a hierarchical feature encoder based on the Visual State Space (VSS Block) to extract shallow target features containing global structural information. and deep semantic features The detail branch extracts local detail features from the image using a local detail encoder. The calculation process can be expressed as:

[0093]

[0094]

[0095]

[0096] in, Represents initial local detail features. Indicates SE channel attention, Represents the residual block. This represents the GELU activation function. This represents a 3×3 convolution operation.

[0097] Subsequently, to obtain multi-scale image features that combine global structural information and local detail information, a semantically guided shallow fusion unit was designed. First, for local detail features... and deep semantic features Channel alignment can be calculated using the following formula:

[0098]

[0099]

[0100] in, and For channel-aligned convolution, For upsampling operation, This represents the local detail features after channel alignment. This represents the deep semantic features after channel alignment.

[0101] Then, the shallow target features Local detail features after channel alignment Deep semantic features aligned with channels A common input gating mapping function is used to construct semantically gating guided weights. Local detail features after channel alignment Deep semantic features aligned with channels By fusing the data, preliminary enhanced features are obtained, consisting of global structural information and local detailed information. Its calculation formula can be expressed as:

[0102]

[0103]

[0104] in, Represents the gated mapping function. This indicates a feature concatenation operation. This represents the Sigmoid activation function. This represents the fusion mapping function.

[0105] Using gated residual injection to analyze shallow target features Preliminary enhanced features of global structural information and local detailed information The features are then fused and refined using residual blocks to obtain enhanced features with both global structural information and local detail information. Its calculation formula can be expressed as:

[0106]

[0107] in, This represents the learnable scaling factor. This indicates an element-wise multiplication operation.

[0108] Finally, the multi-scale image feature learning module based on dual-branch semantic guidance outputs multi-scale image features. , , , ].

[0109] II. Constructing a multi-scale, step-by-step fusion decoding module

[0110] A multi-scale, stepwise fusion decoding module is constructed to generate high-precision depth maps. This module first concatenates the multi-scale image features output from the encoder with the corresponding level's output features from the decoder along the channel dimension. Then, it sequentially upsamples the concatenated features to gradually recover the high-resolution decoded features of the image. Subsequently, the high-resolution decoded features are upsampled, compressed, and depth-mapped to finally generate a high-precision depth map suitable for EIA generation.

[0111] Specifically, firstly, deep semantic features Perform channel mapping to obtain initial decoding features. Initial decoding features The input consists of multiple cascaded decoding units. Within each unit, skip connections are used to fuse multi-scale image features of the corresponding scale with the current decoding features. Convolution operations are then used to perform feature decoding and upsampling, thereby progressively restoring the image spatial resolution and ultimately outputting high-resolution decoded features. .

[0112] Based on this, the depth of the image is predicted. First, high-resolution decoded features are analyzed. Upsampling is performed, followed by channel compression and depth mapping operations to obtain a high-precision depth map suitable for EIA generation. .

[0113] III. Constructing a Deep-Guided Geometric Inverse Mapping EIA Generation Module

[0114] A depth-guided geometric inverse mapping (EIA) generation module is constructed. This module first generates a depth map suitable for generating EIA. Fit to surface equation Then, by performing geometric inverse mapping, the linear equation of the ray emitted from a pixel on EIA is obtained by passing it through the center of the corresponding microlens. Next, we solve the surface equations. With linear equations By identifying the intersection points between light rays and the depth surface, the problem of finding the intersection points is transformed into a fixed-point problem, thereby determining the mapped coordinates of the pixel in the color image. This improves the accuracy of the mapped coordinates while effectively reducing computational complexity. Finally, based on the above principle, each pixel in the EIA is traversed and its pixel value is calculated, thus generating a high-quality EIA.

[0115] First, for high-precision depth maps A depth inversion is performed, followed by proportional compression based on the depth range of the light field 3D display screen to obtain a depth map suitable for generating EIA. A spatial rectangular coordinate system is established with the plane containing the microlens array as the reference plane. ,in shaft and The axis is parallel to the EIA plane. The axis is perpendicular to EIA. A bicubic B-spline surface reconstruction method is used to reconstruct the depth map. Fit to a surface equation Based on the geometric inverse mapping, suppose the coordinates of a certain pixel on EIA are... The light emitted passes through the center of the corresponding microlens. The linear equation of the ray can then be obtained. , is represented as:

[0116]

[0117] in, This represents the distance between the microlens array and the EIA. These are the depth coordinates in three-dimensional space. When the pixel point on EIA... When an emitted ray intersects a depth surface, the intersection point satisfies both the linear equation of the ray and the surface equation of the depth map; therefore, the depth at the intersection point is... satisfy:

[0118]

[0119] Therefore, the problem of finding the intersection point of a ray and a depth surface is transformed into solving for a fixed point, where the depth value at the intersection point is... satisfy:

[0120]

[0121] in, This represents a fixed-point function.

[0122] To improve the accuracy and efficiency of the solution, a Halley iteration method based on the second derivative is introduced in the fixed-point solution process. First, the center depth plane is used as the initial value. In the In the next iteration, based on the current depth value Inversely calculate the coordinates of the intersection point between the ray equation and the depth surface. The calculation formula can be expressed as:

[0123]

[0124]

[0125] in, This represents the coordinates of the pixel to be determined in EIA. This represents the coordinates of the center of the corresponding microlens.

[0126] Subsequently, the depth value of the point is obtained through a fixed-point function, and a residual function is constructed:

[0127]

[0128] in, This represents the depth value obtained by substituting the coordinates of the current candidate intersection point into the depth surface.

[0129] The Halley iterative update formula based on the second derivative is:

[0130]

[0131] in, For the first The depth value of the next iteration. and They represent the residual functions at... The first and second derivatives at point .

[0132] When the change in depth value between two consecutive iterations satisfies Stop iteration when, where For the preset threshold (usually set to a value of...), This module effectively improves the calculation accuracy of mapped coordinates while reducing the number of iterations, thus obtaining precise mapped coordinates. Based on the precise mapped coordinates, from the color image... The pixel value at the corresponding position is obtained and assigned to the corresponding pixel in the EIA. By iterating through all pixels in the EIA using this principle, a high-quality EIA can be generated quickly.

[0133] IV. Training the Light Field 3D Display Source Generation Network

[0134] The light field 3D display image source generation network proposed in this embodiment includes a multi-scale image feature learning module based on dual-branch semantic guidance, a multi-scale progressive fusion decoding module, and a depth-guided geometric inverse mapping (EIA) generation module. Based on these three parts, the light field 3D display image source generation network is constructed. During the training process of the light field 3D display image source generation network, global structural information and local detail information of the input image are effectively extracted. High-precision depth information is obtained through progressive decoding and multi-scale fusion. Combined with the depth-guided geometric inverse mapping (EIA) generation module, high-quality EIA is obtained. Simultaneously, the training process is optimized using L1 loss to obtain an optimized light field 3D display image source generation network.

[0135] V. Application of the Optimized Light Field 3D Display Source Generation Network

[0136] The optimized light field 3D display source generation network based on Part 4 is applied to the light field 3D display source generation to obtain high-quality light field 3D display sources.

[0137] VI. Experimental Section

[0138] To verify the effectiveness of the proposed method in EIA generation quality, EIAs were generated using this method and the traditional forward pixel mapping method, respectively, and a 3D display visualization comparison experiment was conducted. The experiment selected the reconstructed 3D image and its representation in a color image. Comparing images from the same viewpoint, the PSNR of the EIA-reconstructed 3D image generated by the traditional forward pixel mapping method is 20.79 dB, and the SSIM is 0.7772, while the PSNR of the EIA-reconstructed 3D image generated by the method proposed in this embodiment is 21.33 dB, and the SSIM is 0.8217. The experimental results are as follows: Figure 2 As shown in the figure. Experimental results demonstrate that the method proposed in this embodiment of the invention can effectively improve the generation quality of EIA.

[0139] To verify the effectiveness of the proposed method in terms of EIA generation efficiency, a comparative experiment was conducted with the traditional forward pixel mapping method. The experiment was performed on a CPU platform with an 11th Gen Intel(R) Core(TM) i5-1135G7 @ 2.40 GHz processor. During the experiment, both the traditional forward pixel mapping method and the method proposed in this embodiment were run under the same hardware environment, the same input color image, the same depth map, and the same light field 3D display system parameters. The EIA resolution was 320×240, and the EIA consisted of 20×15 image pixels. The traditional method took 3034ms to generate the EIA, while the method proposed in this embodiment took only 24ms. The experiment only counted the runtime of the core EIA generation stage, excluding the time for color image reading, depth map reading, depth prediction, and result saving. The experimental results show that the EIA generation time of the method proposed in this embodiment is significantly lower than that of the traditional forward pixel mapping method; therefore, the method proposed in this embodiment can improve the EIA generation efficiency.

[0140] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the sequence numbers of the embodiments described above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0141] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating light field 3D display source material based on deep learning, characterized in that, The method includes: A light field 3D display source generation network is constructed, consisting of a multi-scale image feature learning module guided by dual-branch semantics, a multi-scale hierarchical fusion decoding module, and a depth-guided geometric inverse mapping (EIA) generation module. The training process of the light field 3D display source generation network is optimized by using L1 loss, resulting in an optimized light field 3D display source generation network. Input image information and generate light field 3D display source based on the optimized light field 3D display source generation network.

2. The method for generating light field 3D display source material based on deep learning according to claim 1, characterized in that, The multi-scale image feature learning module consists of a main branch, a detail branch, and a semantically guided shallow fusion unit. The main branches employ a hierarchical feature encoder based on visual state space to extract shallow target features containing global structural information. and deep semantic features ; The detail branch extracts local detail features from the image through a local detail encoder. ; Semantic-guided shallow fusion units for local detail features and deep semantic features Perform channel alignment to integrate shallow target features. Local detail features after channel alignment Deep semantic features aligned with channels A common input gating mapping function is used to construct semantically gating guided weights. ; Local detail features after channel alignment Deep semantic features aligned with channels By fusing the data, preliminary enhanced features are obtained, consisting of global structural information and local detailed information. ; Using gated residual injection to analyze shallow target features Preliminary enhanced features of global structural information and local detailed information The features are then fused and refined using residual blocks to obtain enhanced features with both global structural information and local detail information. .

3. The method for generating light field 3D display source material based on deep learning according to claim 2, characterized in that, The enhanced features for: ; in, This represents the learnable scaling factor. This represents element-wise multiplication. This represents a residual block.

4. The method for generating light field 3D display source material based on deep learning according to claim 1, characterized in that, The depth-guided geometric inverse mapping (EIA) generation module is as follows: This will be suitable for generating depth maps for EIA. Fit to surface equation By using geometric inverse mapping, the linear equation of a ray emitted from a pixel in EIA is obtained by passing it through the center of the corresponding microlens. ; Solving the surface equation With linear equations The intersection points between the points determine the mapping coordinates of the pixel in the color image. The process iterates through each pixel in the EIA and calculates its pixel value to generate a high-quality EIA.

5. The method for generating light field 3D display source material based on deep learning according to claim 4, characterized in that, The depth map to be used for generating EIA Fit to surface equation By using geometric inverse mapping, the linear equation of a ray emitted from a pixel in EIA is obtained by passing it through the center of the corresponding microlens. for: The depth map is reconstructed using bicubic B-spline surfaces. Fit to a surface equation According to the geometric inverse mapping, the coordinates of a certain pixel on EIA are: The emitted light passes through the center of the corresponding microlens. The linear equation of the ray is obtained. , is represented as: ; The distance between the microlens array and the EIA is... , For 3D spatial depth coordinates, when the pixel point on EIA... When an emitted ray intersects a depth surface, the intersection point satisfies both the linear equation of the ray and the surface equation of the depth map. The depth at the intersection point is... satisfy: ; The problem of finding the intersection point of a ray and a depth surface is transformed into solving for a fixed point, where the depth value is the intersection point. satisfy: ; in, This represents a fixed-point function.

6. The method for generating light field 3D display source material based on deep learning according to claim 4, characterized in that, The solution of the surface equation With linear equations The intersection point between them is: We introduce Halley iteration based on the second derivative, using the center depth plane as the initial value. In the In the next iteration, based on the current depth value Inversely calculate the coordinates of the intersection point between the ray equation and the depth surface. for: ; ; in, This represents the coordinates of the pixel to be determined in EIA. This represents the coordinates of the center of the corresponding microlens; the depth value at that point is obtained through a fixed-point function, and a residual function is constructed: ; in, This represents the depth value obtained by substituting the coordinates of the current candidate intersection point into the depth surface; The Halley iterative update formula based on the second derivative is: ; in, For the first The depth value of the next iteration. and They represent the residual functions at... The first and second derivatives at point ; When the change in depth value between two consecutive iterations satisfies Stop iteration when This is a preset threshold.