Real-time parallax optimization method for adaptive light field rendering, medium and equipment

By using a real-time parallax optimization method based on adaptive light field rendering, combined with data from microscopes and lidar, the parallax map is optimized and 3D reconstruction is performed. This solves the imaging problem of traditional binocular parallax algorithms in high magnification scenarios, achieving high-precision stereo reconstruction and fast rendering. It is suitable for naked-eye 3D microscopic imaging and industrial inspection.

CN121033239APending Publication Date: 2025-11-28NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511122055.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Traditional binocular parallax algorithms produce poor imaging results in high magnification scenarios, resulting in blurred edges and loss of details in stereo images, which affects the accuracy of cell morphology analysis and industrial inspection.

Method used

A real-time parallax optimization method using adaptive light field rendering is employed. Image sequences are acquired by a microscope equipped with a camera and depth maps are obtained by a LiDAR. By combining a spatiotemporal feature extraction network and a light field propagation model, the parallax map is optimized and 3D reconstruction is performed. High-fidelity 3D structures are generated using multi-mode imaging paths.

Benefits of technology

It achieves high-precision stereoscopic reconstruction in high magnification scenarios, eliminates edge blurring and detail loss, improves imaging quality and rendering speed, and is suitable for naked-eye 3D microscopy, biomedical, industrial inspection and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033239A_ABST
    Figure CN121033239A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time parallax optimization method for adaptive light field rendering, a medium and equipment. The method comprises the following steps: acquiring a dynamic microscopic image sequence and an auxiliary depth map; inputting the data into a double-flow architecture of a spatial-temporal feature extraction network for preprocessing, and outputting a standardized spatial-temporal feature tensor; inputting the standardized spatio-temporal feature tensor into a pre-trained spatio-temporal double-flow cascade network model for processing, and outputting an initial disparity map; a microscope optical parameter and an initial disparity map are input, AO-Net processing is utilized, a polynomial coefficient is generated, a light field propagation model is utilized, the polynomial coefficient is optimized based on the Fresnel diffraction principle, and corrected light field data are output; inputting an initial disparity map, processing the initial disparity map by adopting a grading strategy, and outputting an optimized disparity map; corrected light field data and an optimized disparity map are input, and a physically optimized 3D structure is output through 3D reconstruction and physical constraint; and inputting the physically optimized 3D structure and the optimized disparity map, and executing real-time rendering. According to the invention, the problem of poor 3D imaging effect is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer vision and optical imaging technology, and particularly relates to a real-time parallax optimization method for adaptive light field rendering, a medium and equipment. BACKGROUND

[0002] In the field of naked-eye 3D microscopic imaging, accurate parallax calculation and light field rendering are the core technologies to achieve high-fidelity stereoscopic display. The core goal is to reconstruct three-dimensional depth information conforming to the human eye's stereoscopic vision mechanism from two-dimensional microscopic images. This technology is widely used in biomedical diagnosis, such as cell morphology analysis, pathological section observation, and industrial detection, such as precision part defect identification and material microstructure analysis.

[0003] Currently, the traditional binocular parallax algorithm used in the field of naked-eye 3D microscopic imaging has a significant edge blur problem in high magnification scenarios. Limited by the accuracy of pixel-level parallax calculation, it is difficult to accurately reconstruct the boundaries of organelles, cell membranes and other fine structures, resulting in ghosting or loss of details in stereoscopic images. This may lead to misjudgment of cell morphology in medical diagnosis, and affect the identification accuracy of micron-level defects such as integrated circuit solder crack in industrial detection.

[0004] Therefore, there is an urgent need for a real-time parallax optimization method for adaptive light field rendering, a medium and equipment to solve the problem of poor imaging effect of traditional binocular parallax algorithm in high magnification scenarios. SUMMARY

[0005] The present application provides a real-time parallax optimization method for adaptive light field rendering, a medium and equipment to solve the problem of poor imaging effect of traditional binocular parallax algorithm in high magnification scenarios.

[0006] To achieve the above purpose, the present application adopts the following technical solutions:

[0007] A real-time parallax optimization method for adaptive light field rendering, comprising the following steps:

[0008] The dynamic microscopic image sequence is collected by a camera mounted on a microscope, and an auxiliary depth map is obtained by a laser radar; the collected dynamic microscopic image sequence and the auxiliary depth map are input into a double-flow architecture of a space-time feature extraction network for preprocessing, and a standardized space-time feature tensor is output; the standardized space-time feature tensor is input into a pre-trained space-time double-flow cascaded network model for processing, and an initial disparity map is output; the microscope optical parameters and the initial disparity map are input, and the AO-Net is used for processing and generating Zernike polynomial coefficients, and then a light field propagation model is used to optimize the Zernike polynomial coefficients based on the Fresnel diffraction principle, and corrected light field data is output; the initial disparity map is input and processed by a hierarchical strategy, and an optimized disparity map is output; the corrected light field data and the optimized disparity map are input, and a physically optimized 3D structure is output through 3D reconstruction and physical constraints; and real-time rendering is performed according to the input physically optimized 3D structure and the optimized disparity map.

[0009] To optimize the above technical solutions, the specific measures taken also include:

[0010] Further, the preprocessing of the collected dynamic microscopic image sequence and the auxiliary depth map by the double-flow architecture of the space-time feature extraction network and the output of the standardized space-time feature tensor comprise the following steps:

[0011] The dynamic microscopic image sequence and the auxiliary depth map are input, the double-flow architecture of the space-time feature extraction network performs pixel value mapping on the dynamic microscopic image sequence to generate normalized image data, the BM3D algorithm is used to process the normalized image data to generate denoised image data, the ST-ResNet-ResNet50 model is used to extract single-frame spatial features from the denoised image data, the Farneback optical flow algorithm is used to extract inter-frame motion vectors from the continuous frame denoised image data, and finally the single-frame spatial features and the inter-frame motion vectors are fused to output the standardized space-time feature tensor.

[0012] Further, the training of the space-time double-flow cascaded network model comprises the following steps:

[0013] The dilated convolution of the space-time double-flow cascaded network is used to process the standardized space-time feature tensor to generate multi-scale spatial features; at the same time, the bidirectional LSTM network is used to process the standardized space-time feature tensor to generate 16-frame time sequence features; then the CA-Transformer mechanism is used to fuse the multi-scale spatial features and the 16-frame time sequence features to generate dynamic enhancement features; finally, the photometric consistency loss and the structural similarity loss are used for joint optimization.

[0014] Further, the input microscope optical parameters and the initial disparity map are processed by the AO-Net to generate Zernike polynomial coefficients, and then an optical field propagation model is used to optimize the Zernike polynomial coefficients based on the Fresnel diffraction principle, and corrected optical field data are output, including the following steps:

[0015] The input microscope optical parameters and the initial disparity map are processed by the AO-Net to generate Zernike polynomial coefficients, and then an optical field propagation model is used to optimize the Zernike polynomial coefficients based on the Fresnel diffraction principle, and corrected optical field data are output.

[0016] Further, the input initial disparity map is processed using a hierarchical strategy to output an optimized disparity map, including the following steps:

[0017] The input initial disparity map is processed by the MobileNetV3 lightweight network to generate a global disparity map as a coarse resolution branch, and at the same time, the global disparity map edge is optimized by the progressive PUS upsampling to generate a sub-pixel edge map as a fine resolution branch, then the Fusion-Sharing mechanism is used to fuse the global disparity map and the sub-pixel edge map for feature sharing, finally, the CUDA kernel function is used for parallel calculation and the GPU memory pool is pre-allocated to realize real-time acceleration, and then the optimized disparity map is output.

[0018] Further, the input corrected optical field data and the optimized disparity map are output by 3D reconstruction and physical constraints to output a physically optimized 3D structure, including the following steps:

[0019] The input corrected optical field data and the optimized disparity map are processed by the curvature constraint to generate a non-penetrating membrane structure, and then an imaging mode selection branch is selected, including a transmission light mode and a fluorescence mode, wherein in the transmission light mode, the phase contrast enhancement processing is used to optimize the disparity map to generate a high-contrast disparity map, and in the fluorescence mode, the spectral unmixing technology is used to process the optimized disparity map to generate a split-channel disparity map, and finally the non-penetrating membrane structure and the high-contrast disparity map or the split-channel disparity map are fused to output the physically optimized 3D structure.

[0020] Further, the input physically optimized 3D structure and the optimized disparity map are used to perform real-time rendering, including the following steps:

[0021] In the naked eye 3D microscope path, the DIPRA algorithm is used to process the optimized disparity map to generate multi-view stereoscopic image data, and finally the stereoscopic image data is encoded into a 3D video stream in H.265 format in real time, and output.

[0022] Further, the real-time rendering according to the input physical optimization 3D structure and the optimized parallax map comprises the following steps:

[0023] In the endoscope path, the physical optimization 3D structure is converted into a real-time updated OBJ format dynamic mesh model through a moving cube algorithm and a Poisson surface reconstruction, and then is optimized for two scene branches:

[0024] In the industrial detection scene, material perception type particle filtering denoising is performed on the dynamic mesh model, a PBR industrial rendering model is used to generate a video stream with real-time detection superposition, and the output is a VP9 encoded video stream;

[0025] In the biomedical scene, the dynamic mesh model is intelligently cropped to generate a focal point area 3D structure through octree subdivision and implicit surface boundary detection, and then a Phong-BSSRDF hybrid shading model is used for real-time rendering, and a VP9 encoded video stream with a depth channel is output.

[0026] Further, a computer readable storage medium stores a computer program, and the computer program causes a computer to execute the real-time parallax optimization method for adaptive light field rendering.

[0027] Further, an electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the real-time parallax optimization method for adaptive light field rendering is realized.

[0028] The beneficial effects of the present application are:

[0029] The present application realizes high-precision stereo reconstruction of organelles, cell membranes and micron-level defects in a high magnification scene by fusing microscopic image sequences and laser radar auxiliary depth information, combining double-flow space-time feature extraction, effectively eliminating the edge blur, detail loss and ghosting problems of traditional binocular algorithms; the light field propagation model optimized based on the Fresnel diffraction principle is used to dynamically correct the Zernike polynomial coefficient, ensuring the spatial consistency and depth of field restoration of light field rendering; multi-scale hierarchical parallax optimization and CUDA parallel acceleration are used to output a high-fidelity optimized parallax map within milliseconds, and the curvature constraint and multi-mode imaging path are used to generate a physically consistent 3D structure, which not only meets the real-time needs of naked eye 3D microscopic display, but also takes into account the scalability of industrial detection and biomedical scenes, significantly improves the imaging quality, rendering speed and system adaptability, and provides stable and reliable technical support for microscopic stereo imaging and cross-field three-dimensional visualization.

[0030] The application is aimed at professional application scenarios such as naked eye 3D microscopes, endoscopes and the like, fuses micro image sequences and laser radar auxiliary depth information, combines a double-flow space-time feature extraction network and an optimized light field propagation model based on the Fresnel diffraction principle, designs a special light field modulation module and an output interface optimization scheme, breaks through the precision bottleneck and real-time limitation of traditional binocular disparity algorithms in high magnification and dynamic scenes, realizes high-fidelity stereoscopic reconstruction and real-time rendering of organelles, cell membranes and micron-level defects, significantly improves dynamic 3D imaging quality and system adaptability, and provides a revolutionary technical solution for the fields of life science research, precision industrial detection and the like. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 The whole flowchart of the real-time disparity optimization method of adaptive light field rendering is shown in the figure.

[0032] Figure 2 The flowchart of obtaining a standardized space-time feature tensor of the real-time disparity optimization method of adaptive light field rendering is shown in the figure.

[0033] Figure 3 The flowchart of outputting an initial disparity map of the real-time disparity optimization method of adaptive light field rendering is shown in the figure.

[0034] Figure 4 The flowchart of outputting corrected light field data of the real-time disparity optimization method of adaptive light field rendering is shown in the figure.

[0035] Figure 5 The flowchart of outputting an optimized disparity map of the real-time disparity optimization method of adaptive light field rendering is shown in the figure.

[0036] Figure 6 The flowchart of outputting a physically optimized 3D structure of the real-time disparity optimization method of adaptive light field rendering is shown in the figure.

[0037] Figure 7 The coordinate establishment diagram of a dynamic microscopic image of the real-time disparity optimization method of adaptive light field rendering is shown in the figure. DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0039] As shown in the drawingsFigure 1 As shown in the figure, a real-time parallax optimization method of adaptive light field rendering includes the following steps:

[0040] A dynamic microscopic image sequence is collected by a 4K scientific grade CMOS camera mounted on a microscope, and an auxiliary depth map is obtained by a laser radar for data collection; the collected dynamic microscopic image sequence and auxiliary depth map are input into a double-flow architecture of a space-time feature extraction network for preprocessing, and a standardized space-time feature tensor is output; the standardized space-time feature tensor is input into a pre-trained space-time double-flow cascaded network model for processing, and an initial parallax map is output; the microscope optical parameters and the initial parallax map are input, and the AO-Net is used to process and generate Zernike polynomial coefficients, and then a light field propagation model is used to optimize the Zernike polynomial coefficients based on the Fresnel diffraction principle, and corrected light field data is output; the initial parallax map is input and processed using a hierarchical strategy, and an optimized parallax map is output; the corrected light field data and the optimized parallax map are input, and a physical optimized 3D structure is output through 3D reconstruction and physical constraints; and real-time rendering is performed according to the input physical optimized 3D structure and the optimized parallax map.

[0041] The camera can have a resolution of 3840x2160 and a frame rate of 60fps.

[0042] As shown in the figure, Figure 2 In the further specific embodiments based on the above, the preprocessing of the collected dynamic microscopic image sequence and auxiliary depth map into the double-flow architecture of the space-time feature extraction network and the output of the standardized space-time feature tensor include the following steps:

[0043] The dynamic microscopic image sequence and the auxiliary depth map are input, the dynamic microscopic image sequence is subjected to pixel value mapping by the double-flow architecture of the space-time feature extraction network to generate normalized image data, the normalized image data is processed by the BM3D algorithm to generate denoised image data, the ST-ResNet-ResNet50 model is used to extract single-frame spatial features from the denoised image data, and the Farneback optical flow algorithm is used to extract inter-frame motion vectors from the continuous frame denoised image data, and finally the single-frame spatial features and the inter-frame motion vectors are fused to output the standardized space-time feature tensor. The value range of the normalized image data can be [-1, 1].

[0044] As shown in the figure, Figure 3 In the further specific embodiments based on the above, the training of the space-time double-flow cascaded network model includes the following steps:

[0045] The expansion rate of the spatio-temporal double-flow cascaded network is 1-5, and the dilated convolution is used to process the standardized spatio-temporal feature tensor to generate multi-scale spatial features; at the same time, a bidirectional LSTM network with a hidden layer dimension of 256 is used to process the standardized spatio-temporal feature tensor to generate 16-frame time sequence features; then the multi-scale spatial features and 16-frame time sequence features are fused through the CA-Transformer mechanism, and the dynamic enhancement features are generated by training on a synthetic dataset of 100,000+ related motion scenes; finally, the photometric consistency loss and the structural similarity loss are used for joint optimization.

[0046] The photometric consistency loss function is:

[0047]

[0048] As shown in the accompanying Figure 7 , in the formula, H represents the height of the dynamic microscopic image in the dynamic microscopic image sequence, that is, the number of pixels in the vertical direction, which is obtained by slicing the standardized spatio-temporal feature tensor; W represents the width of the dynamic microscopic image, that is, the number of pixels in the horizontal direction, which is obtained by slicing the standardized spatio-temporal feature tensor; I t represents the spatio-temporal feature tensor of the t-th frame of dynamic microscopic image, which is obtained by slicing the standardized spatio-temporal feature tensor; I t (a i,j ) represents the pixel value of the (i, j) position in the t-th frame of dynamic microscopic image data, which is contained in the standardized spatio-temporal feature tensor; I t-1 represents the spatio-temporal feature tensor of the t-1-th frame of dynamic microscopic image; d i,j represents the disparity value of the (i, j) pixel, which represents the position offset of the corresponding pixel in the two frames of dynamic microscopic image data in the unit of moving pixel value.

[0049] The structural similarity loss function is:

[0050] L SSIM =1-SSIM(I t ,I t+1 ))

[0051]

[0052] In the formula, μ x , μ y represent the mean of the pixel values of the incoming dynamic microscopic image, where x i is the i-th pixel value of the dynamic microscopic image; σ x , σ y represent the standard deviation of the pixel values of the incoming dynamic microscopic image, where x i is the i-th pixel value of the dynamic microscopic image; σ xyTo input the pixel value covariance of the dynamic microscopic image, c1 and c2 represent stability constants, where c1 = (K1 × L) 2 c2 = (K2 × L) 2 Where K1 = 0.01, K2 = 0.03, and L: the maximum pixel value of the input dynamic microscopic image.

[0053] As attached Figure 4 As shown, in a further specific embodiment based on the above, the input microscope optical parameters and initial parallax map are processed using AO-Net to generate Zernike polynomial coefficients, and then the Zernike polynomial coefficients are optimized based on the Fresnel diffraction principle using the light field propagation model to output corrected light field data, including the following steps:

[0054] Input the microscope optical parameters and initial parallax plot, where the microscope optical parameters include the microscope spherical aberration coefficient C. s Coma coefficient C c Then, the initial disparity map and optical parameters are processed using a 5-layer fully connected AO-Net to generate Zernike polynomial coefficients. The light field propagation model is then constructed by the DiffRender engine based on the phase filtering propagation formula in the Fourier domain. The Zernike polynomial coefficients are optimized based on the Fresnel diffraction principle, and the corrected light field data is output.

[0055] The above-mentioned Fourier domain phase filter propagation formula is: E out (a)=F -1 (F(E in (a)·e i·φ(a) )·H(u,v))

[0056] In the formula, F is the Fast Fourier Transform (FFT); F -1 For Inverse Fourier Transform (IFFT); E in (a) represents the complex amplitude of the input optical field obtained from the initial disparity map; φ(a) represents the phase delay of the spatial light modulator (SLM), and the phase modulation amount predicted by AO-Net. Among them, c k Zernike polynomial coefficients output by AO-Net, Z k (a) represents the Zernike basis functions; e i·φ(a) φ(a) is used in combination with i (the imaginary unit) to describe the phase delay of a light wave; H(u,v) is the optical transfer function (OTF), and (u,v) are the spatial frequency coordinates. θ m Let θ be the horizontal angle between the ray and the optical axis. n Let λ be the perpendicular angle between the ray and the optical axis, and λ be the wavelength of the light. H(u,v) = e-iπλz (u 2 +v 2 )·CTF(u,v;C s C c )·A(u,v), where z is the thickness, estimated from an auxiliary depth map. The contrast transfer function is calculated directly from the formula, and A(u,v) is the aperture function of the objective lens obtained from the microscope.

[0057] As attached Figure 5 As shown, in a further specific embodiment based on the above, the input initial disparity map is processed using a hierarchical strategy to output an optimized disparity map, including the following steps:

[0058] The initial disparity map is input and processed using the lightweight MobileNetV3 network to generate a global disparity map as a coarse-resolution branch. Simultaneously, progressive PUS upsampling and 4 layers of subpixel convolutions are used to optimize the edges of the global disparity map, generating a subpixel edge map as a fine-resolution branch. The global disparity map and the subpixel edge map are then fused using the Fusion-Sharing mechanism for feature sharing. Finally, real-time acceleration is achieved through parallel computation using CUDA kernel functions and pre-allocation of the GPU memory pool, and the optimized disparity map is output.

[0059] As attached Figure 6 As shown, in a further specific embodiment based on the above, the input corrected light field data and optimized disparity map, through 3D reconstruction and physical constraints, output a physically optimized 3D structure, including the following steps:

[0060] Input corrected light field data and optimized disparity map, process the corrected light field data using curvature constraint to generate a non-permeable film structure, then select a branch according to the imaging mode, including transmission light mode and fluorescence mode. In transmission light mode, phase contrast enhancement is used to optimize the disparity map and generate a high-contrast disparity map. In fluorescence mode, spectral unmixing technology is used to process and optimize the disparity map and generate a multi-channel disparity map. Finally, the non-permeable film structure is fused with the high-contrast disparity map or the multi-channel disparity map to output a physically optimized 3D structure.

[0061] In a further specific embodiment based on the above, the process of performing real-time rendering based on the input physically optimized 3D structure and optimized disparity map includes the following steps:

[0062] In the naked-eye 3D microscope path, the DIPRA algorithm is used to process and optimize the disparity map, generate multi-viewpoint stereo image data, and finally encode the stereo image data into H.265 format 3D video stream in real time and output it.

[0063] In further embodiments based on the above, the above-mentioned performing real-time rendering according to the input physical optimized 3D structure and the optimized disparity map further comprises the following steps:

[0064] In the endoscopic path, the physical optimized 3D structure is converted into a real-time updated OBJ format dynamic mesh model through a moving cube algorithm and a Poisson surface reconstruction, and then optimized for the two scene branches:

[0065] In the industrial detection scene, material-aware particle filtering denoising is performed on the dynamic mesh model, a PBR industrial rendering model is used to generate a video stream with real-time detection superimposition, and the output is a VP9 encoded video stream, supporting line-level 60fps real-time defect analysis.

[0066] In the biomedical scene, the dynamic mesh model is intelligently cropped to generate a focal area 3D structure through octree subdivision and implicit surface boundary detection, and then a Phong-BSSRDF hybrid shading model is used for real-time rendering, and a VP9 encoded video stream with a depth channel is output, supporting intraoperative real-time tissue layer visualization.

[0067] One embodiment of the present application is as follows:

[0068] The present application is described in detail taking the biomedical observation scene of the naked eye 3D microscope as an example.

[0069] First, an Olympus BX53 microscope is used to carry a 4K scientific grade CMOS camera with a resolution of 3840x2160 and a frame rate of 60fps, and the live cell samples in the culture dish are continuously observed; in the case of configuring auxiliary equipment such as laser radar, the auxiliary depth map with an accuracy of 0.1 μm and a resolution of 1024x1024 is synchronously acquired. The dynamic image sequence collected is first normalized to map the pixel value to the interval [-1, 1], and denoised to ≤5 gray levels using the BM3D algorithm. Subsequently, the normalized microscopic image and the depth map are input into the ST-ResNet-ResNet50 model to extract single-frame spatial features, and the Farneback optical flow algorithm is used to extract inter-frame motion vectors, and finally the single-frame spatial features and the inter-frame motion vectors are fused to generate a standardized spatio-temporal feature tensor.

[0070] Next, the standardized spatio-temporal feature tensor is input into the model, which uses a dilation rate of 1-5 to extract multi-scale spatial features, and a hidden layer dimension of 256 to extract 16-frame temporal features, and then fuses the multi-scale spatial features and the 16-frame temporal features through the CA-Transformer mechanism and trains on a synthetic dataset of 100,000+ live cell motion scenes, uses a combination of photometric consistency loss and structural similarity loss for optimization, and outputs an initial disparity map.

[0071] Then input the microscope optical parameters, including the spherical aberration coefficient C s , the coma coefficient C c and the initial parallax map, input the AO-Net of the 5-layer full connection structure, obtain the Zernike polynomial coefficient, and combine the light field propagation model constructed by the DiffRender engine to perform phase correction based on the Fresnel diffraction principle to obtain the corrected light field data.

[0072] Subsequently, the initial parallax map is input into the coarse resolution branch based on MobileNetV3 to generate a global parallax map, and is input into the fine resolution branch using 4 layers of sub-pixel convolution to generate a sub-pixel edge map by progressively optimizing edge details, and then the Fusion-Sharing mechanism is used to fuse the global parallax map and the sub-pixel edge map, and the CUDA kernel function and the video memory pool technology are used to realize real-time acceleration, and an optimized parallax map is output.

[0073] Then, according to the corrected light field data and the optimized parallax map, the corrected light field data is processed by curvature constraint (K) to generate a non-penetrating membrane structure. After selecting the imaging mode, if it is a transmission light mode, a phase contrast enhancement module is enabled to improve the contrast, and a high-contrast parallax map is generated; if it is a fluorescence mode, spectral unmixing is performed to realize independent optimization of multi-channel fluorescently labeled structures, and a split-channel parallax map is generated. Finally, the non-penetrating membrane structure and the high-contrast / split-channel parallax map are fused to output a physically optimized 3D structure.

[0074] Finally, in the naked eye 3D microscope path, the DIPRA algorithm is used to generate multi-view stereoscopic images, which are encoded into H.265 format 3D video stream output in real time, ensuring that real-time rendering and low-latency display are realized at 60Hz under 4K resolution.

[0075] Through the above embodiments, in the dynamic observation experiment of living cells, the three-dimensional tracking error of the algorithm for the mitochondrial migration trajectory is ≤0.5 μm, the depth reconstruction error of the cell membrane edge under a 100x objective is reduced from 200 nm to 80 nm, which meets the ISO10523 pathological image resolution requirement, the 4K real-time rendering delay is ≤16 ms, supporting 60fps high-speed dynamic imaging, effectively capturing the details of chromosome arrangement in the metaphase of cell division. The above only describes the preferred embodiments of the present application, and for those skilled in the art, the sensor configuration, network parameters, optimization algorithm, etc. can be improved without departing from the principles of the present application, and these improvements should be included in the protection scope of the present application.

[0076] The adaptive light field rendering real-time parallax optimization algorithm provided by the application effectively solves the problems of edge blur in traditional naked-eye 3D microscopic imaging, poor dynamic scene synchronization, insufficient optical aberration compensation and high hardware dependency, realizes a breakthrough in high-precision stereoscopic imaging capability, greatly reduces the edge blur degree through sub-pixel level parallax calculation and aberration correction, and can improve the equivalent resolution to 100 nanometer level, meets the high-precision analysis requirements of medicine and industry, improves the real-time performance and stability of dynamic scenes, realizes real-time rendering and low trajectory tracking error, adapts to living cell motion observation and industrial high-speed detection, enhances the physical reality and scene adaptability, ensures that the 3D structure topology correctness rate can be greater than or equal to 99%, and improves the display definition through automatic adaptation of transmission light / fluorescence mode, optimizes the hardware compatibility and cost, and reduces the deployment threshold by using a pure software solution with a cost of only 1 / 5-1 / 3 of the traditional solution, compatible with mainstream devices, expands the cross-field application, supports naked-eye 3D microscopes, industrial endoscopes and the like through standardized interfaces, realizes multi-scene coverage from biological observation to industrial detection, provides a revolutionary stereoscopic vision solution for life science and precision manufacturing, and promotes the transformation of three-dimensional imaging technology to intelligent algorithm driving.

[0077] In another embodiment, the application provides a computer readable storage medium storing a computer program, wherein the computer program enables a computer to execute the adaptive light field rendering real-time parallax optimization method as described above.

[0078] In another embodiment, the application provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the adaptive light field rendering real-time parallax optimization method as described above.

[0079] In the embodiments disclosed in the present application, the computer storage medium can be a tangible medium which can contain or store programs for use by or in connection with an instruction execution system, apparatus or device. The computer storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, apparatuses or devices, or any suitable combination of the above. More specific examples of computer storage media can include one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0080] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software manner depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0081] The above is only the preferred embodiment of the present application, and the protection scope of the present application is not limited to the above-mentioned embodiments. Any technical solution falling within the concept of the present application shall fall within the protection scope of the present application. It should be noted that, for ordinary skilled persons in the art, some improvements and refinements without departing from the principles of the present application shall be considered as the protection scope of the present application.

Claims

1. A real-time parallax optimization method for adaptive light field rendering, characterized in that, Includes the following steps: Dynamic microscopic image sequences are acquired using a camera mounted on a microscope, and auxiliary depth maps are obtained using a LiDAR scanner. The acquired dynamic microscopic image sequences and auxiliary depth maps are preprocessed using a two-stream architecture of a spatiotemporal feature extraction network, which outputs a standardized spatiotemporal feature tensor. This standardized spatiotemporal feature tensor is then processed by a pre-trained spatiotemporal two-stream cascaded network model, which outputs an initial disparity map. Microscope optical parameters and the initial disparity map are input, processed using AO-Net to generate Zernike polynomial coefficients, which are then optimized using a light field propagation model based on Fresnel diffraction principles, outputting corrected light field data. The initial disparity map is input and processed using a hierarchical strategy, outputting an optimized disparity map. The corrected light field data and optimized disparity map are input, and through 3D reconstruction and physical constraints, a physically optimized 3D structure is output. Real-time rendering is then performed based on the input physically optimized 3D structure and optimized disparity map.

2. The real-time parallax optimization method for adaptive light field rendering according to claim 1, characterized in that, The process of preprocessing the acquired dynamic microscopic image sequence and auxiliary depth map into a two-stream architecture of a spatiotemporal feature extraction network and outputting a standardized spatiotemporal feature tensor includes the following steps: The input consists of a dynamic microscopic image sequence and an auxiliary depth map. The two-stream architecture of the spatiotemporal feature extraction network performs pixel value mapping on the dynamic microscopic image sequence to generate normalized image data. The BM3D algorithm is then used to process the normalized image data to generate denoised image data. The ST-ResNet-ResNet50 model is then used to extract single-frame spatial features from the denoised image data. At the same time, the Farneback optical flow algorithm is used to extract inter-frame motion vectors from the continuous frame denoised image data. Finally, the single-frame spatial features and inter-frame motion vectors are fused to output a standardized spatiotemporal feature tensor.

3. The real-time parallax optimization method for adaptive light field rendering according to claim 1, characterized in that, The training of the spatiotemporal two-stream cascaded network model includes the following steps: A dilated convolution of a spatiotemporal dual-stream cascaded network is used to process the standardized spatiotemporal feature tensor to generate multi-scale spatial features. At the same time, a bidirectional LSTM network is used to process the standardized spatiotemporal feature tensor to generate 16 frames of temporal features. Then, the multi-scale spatial features and 16 frames of temporal features are fused through the CA-Transformer mechanism to generate dynamically enhanced features. Finally, a joint optimization was performed using photometric consistency loss and structural similarity loss.

4. The real-time parallax optimization method for adaptive light field rendering according to claim 1, characterized in that, The input microscope optical parameters and initial disparity map are processed using AO-Net to generate Zernike polynomial coefficients. Then, the Zernike polynomial coefficients are optimized using a light field propagation model based on the Fresnel diffraction principle, and the corrected light field data is output. The process includes the following steps: inputting microscope optical parameters and initial disparity map, processing the initial disparity map and optical parameters using AO-Net to generate Zernike polynomial coefficients, optimizing the Zernike polynomial coefficients using a light field propagation model constructed by the DiffRender engine based on the phase filtering propagation formula in the Fourier domain, and outputting the corrected light field data.

5. The real-time parallax optimization method for adaptive light field rendering according to claim 1, characterized in that, The initial disparity map is processed using a hierarchical strategy to output an optimized disparity map, including the following steps: The initial disparity map is input and processed using the lightweight MobileNetV3 network to generate a global disparity map as a coarse-resolution branch. Simultaneously, progressive PUS upsampling is used to optimize the edges of the global disparity map, generating a sub-pixel edge map as a fine-resolution branch. The global disparity map and the sub-pixel edge map are then fused using the Fusion-Sharing mechanism for feature sharing. Finally, real-time acceleration is achieved through parallel computation using CUDA kernel functions and pre-allocation of the GPU memory pool, and the optimized disparity map is output.

6. The real-time parallax optimization method for adaptive light field rendering according to claim 1, characterized in that, The input corrected light field data and optimized disparity map, through 3D reconstruction and physical constraints, output a physically optimized 3D structure, including the following steps: Input corrected light field data and optimized disparity map, process the corrected light field data using curvature constraint to generate a non-permeable film structure, then select a branch according to the imaging mode, including transmission light mode and fluorescence mode. In transmission light mode, phase contrast enhancement is used to optimize the disparity map and generate a high-contrast disparity map. In fluorescence mode, spectral unmixing technology is used to process and optimize the disparity map and generate a multi-channel disparity map. Finally, the non-permeable film structure is fused with the high-contrast disparity map or the multi-channel disparity map to output a physically optimized 3D structure.

7. The real-time parallax optimization method for adaptive light field rendering according to claim 1, characterized in that, The step of performing real-time rendering based on the input physically optimized 3D structure and optimized disparity map includes the following steps: In the naked-eye 3D microscope path, the DIPRA algorithm is used to process and optimize the disparity map, generate multi-viewpoint stereo image data, and finally encode the stereo image data into H.265 format 3D video stream in real time and output it.

8. The real-time parallax optimization method for adaptive light field rendering according to claim 1, characterized in that, The step of performing real-time rendering based on the input physically optimized 3D structure and optimized disparity map includes the following steps: In the endoscopic path, the physically optimized 3D structure is converted into a real-time updated OBJ format dynamic mesh model through the moving cube algorithm and Poisson surface reconstruction, and then optimized for two scene branches: In industrial inspection scenarios, material-aware particle filtering noise reduction is performed on the dynamic mesh model, and a PBR industrial rendering model is used to generate a video stream with real-time detection overlays. The output is a VP9 encoded video stream. In biomedical scenarios, a dynamic mesh model is intelligently clipped using octree partitioning and implicit surface boundary detection to generate a 3D structure of the focal area. Then, a Phong-BSSRDF hybrid shading model is used for real-time rendering to output a VP9 encoded video stream with a depth channel.

9. A computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to execute a real-time parallax optimization method for adaptive light field rendering as described in any one of claims 1-8.

10. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements a real-time parallax optimization method for adaptive light field rendering as described in any one of claims 1-8.