Methods and systems for optimizing stereoscopic vision in naked-eye 3D large screens

An ambient light field model is constructed by multi-sensor fusion and a lightweight CNN model. Glare suppression and perspective correction are performed by combining physical simulation and human visual sensitivity. An adaptive stereoscopic vision optimization for naked-eye 3D large screen is achieved by using a lightweight neural radiation field renderer. This solves the glare and perspective problems of naked-eye 3D large screen in complex environments and improves the picture quality and stereoscopic immersion.

CN121462742BActive Publication Date: 2026-05-26ANHUI SHENGZI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610010135.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-05-26
Estimated Expiration
2046-01-06

AI Technical Summary

Technical Problem

Existing naked-eye 3D large screens are susceptible to light interference in complex environments, resulting in glare and multi-viewpoint crosstalk noise, making it difficult to achieve high-quality stereoscopic visual effects. Furthermore, they cannot adjust the perspective in real time to match the viewer's perspective, affecting the immersive stereoscopic experience and adaptability to multiple viewers.

Method used

An ambient light field model is constructed by multi-sensor fusion and a lightweight CNN model. Glare suppression and perspective correction are performed by combining physical simulation and human visual sensitivity. Real-time re-rendering is achieved using a lightweight neural radiation field renderer. Image fusion and backlight adjustment are also performed to generate an adaptive stereoscopic vision optimization signal.

Benefits of technology

It significantly improves the picture quality and immersive experience of naked-eye 3D large screens in complex lighting environments, achieves dynamic adaptability to the environment and personalized optimization for multiple viewers, and ensures the protection of picture details and energy consumption balance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121462742B_ABST
    Figure CN121462742B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for optimizing stereoscopic vision on naked-eye 3D large screens, specifically relating to the field of naked-eye 3D vision optimization technology. The method includes: real-time perception of ambient light and viewer position through multi-sensor fusion to establish an ambient light field model; prediction of glare crosstalk noise based on physical simulation, and adaptive suppression and compensation combining human visual sensitivity and image content features; dynamic generation of a virtual camera based on the viewer's real-time binocular positions, and real-time re-rendering using a lightweight neural radiation field renderer to achieve motion parallax; finally, intelligent fusion and encoding of the glare compensation layer and perspective correction layer, and outputting the result to the screen. The system includes four main modules: environmental perception, glare compensation, perspective rendering, and fusion encoding. This invention effectively suppresses ambient light interference, improves the quality and immersion of stereoscopic images from different viewing angles, and is suitable for naked-eye 3D large-screen displays in outdoor and complex lighting environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of naked-eye 3D visual optimization technology, and more specifically, to a method and system for optimizing the stereoscopic vision of naked-eye 3D large screens. Background Technology

[0002] With its advantage of presenting stereoscopic visual effects without the need for auxiliary devices, naked-eye 3D large screens have been widely used in outdoor advertising, venue displays, and public information dissemination, becoming a core carrier for enhancing the visual communication experience. As application scenarios diversify, users are placing higher demands on the clarity, stereoscopic realism, environmental adaptability, and multi-viewer suitability of naked-eye 3D images. However, existing naked-eye 3D display technologies are mostly developed based on fixed scene preset parameters, lacking the ability to dynamically perceive and adaptively adjust to the actual deployment environment. This makes the stereoscopic visual effect susceptible to interference from external factors, making it difficult to meet the high-quality display needs in complex scenarios.

[0003] Existing glasses-free 3D large screens have significant shortcomings in adapting to ambient light. In practical applications, the screen is subject to various light interferences, such as direct sunlight and ambient light reflection. These rays, after being reflected by the screen surface, interact with the grating structure, easily generating glare and multi-viewpoint crosstalk noise, resulting in blurred image details, reduced contrast, and severely disrupting the continuity of stereoscopic vision. Traditional solutions often use fixed brightness adjustment or general noise reduction algorithms, without precise optimization based on the characteristics of ambient light sources, the optical properties of the screen surface, and the sensitivity of human vision. These solutions cannot completely suppress glare interference and may even cause image color distortion or loss of detail, making it difficult to balance the needs of glare suppression and image quality protection.

[0004] Existing technologies have significant shortcomings in terms of the realism and adaptability of stereoscopic vision. Traditional naked-eye 3D displays mostly use preset fixed viewing angles for image rendering. When viewers move, they cannot adjust their perspective in real time to match their viewing angle, resulting in problems such as missing motion parallax and perspective distortion, which weakens the immersive stereoscopic experience. At the same time, existing rendering technologies struggle to balance real-time performance and personalization. They either rely on complex calculations, leading to excessively high rendering latency, or sacrifice image accuracy by simplifying algorithms. Furthermore, in multi-viewer scenarios, there is a lack of priority allocation and image fusion mechanisms for different viewer positions, making it impossible to guarantee a high-quality viewing experience for multiple viewers simultaneously. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method and system for optimizing the stereoscopic vision of naked-eye 3D large screens.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] The method for optimizing the stereoscopic vision of naked-eye 3D large screens includes the following steps:

[0008] S1. Environmental Perception and Modeling: Through multi-sensor fusion and intelligent analysis, a digital understanding of the physical environment is constructed, capturing panoramic lighting information and audience position. A lightweight CNN model is used to decouple light source parameters, environmental diffuse reflection, and screen surface optical properties to form a real-time environmental light field model and audience geometric matrix.

[0009] S2. Glare Suppression and Compensation: Physical simulation is performed based on the ambient light field model and screen optical property database to generate a multi-viewpoint crosstalk noise map. The noise map is dynamically modulated by combining human visual sensitivity and image content features. Then, adaptive pre-compensation and perception enhancement are performed on the original image.

[0010] S3, Dynamic Perspective Correction: Define a dedicated virtual camera based on the real-time eye positions of each tracked viewer, and use a lightweight neural radiation field renderer to re-render the 3D scene in real time to generate left and right eye images that match the current viewing angle.

[0011] S4. Fusion Encoding Output: The multi-viewpoint image base layer after glare compensation is intelligently fused with the guide layer image generated by perspective correction. The display priority is assigned according to the viewer's position. The fused image generates sub-pixel driving signals through viewpoint encoding, which synchronously adjust the backlight and outputs it to the naked-eye 3D screen.

[0012] Specifically, the multiple sensors in S1 include:

[0013] Multispectral HDR camera, high-precision brightness meter and depth sensor;

[0014] The multispectral HDR camera captures a panoramic image of a 180-degree field of view in front of the screen, the high-precision luminance meter measures the light intensity in different normal directions, and the depth sensor tracks the three-dimensional coordinates and orientation of the viewer.

[0015] After training, the lightweight CNN model is fine-tuned for key parameters through an online incremental learning mechanism. The online incremental learning mechanism includes confidence assessment and sample selection, incremental fine-tuning execution steps, selecting high-confidence samples and storing them in a cache pool, freezing encoder weights during fine-tuning, and only updating the relevant weights of decoder branches.

[0016] Specifically, the lightweight CNN model:

[0017] The model employs an encoder-decoder architecture. The encoder uses MobileNetV3-Small as its backbone network, and the input is a four-channel stitched image provided by a multispectral HDR camera. The decoder includes three parallel branches, which output the light source parameters, the average RGB intensity value of the ambient diffuse light, and the optical property feature map of the screen surface, respectively. The training dataset of the model is constructed by simulating different light sources illuminating screen samples with known surface characteristics using a programmable LED array, and simultaneously acquiring HDR images and real light source parameters.

[0018] Specifically, in S2:

[0019] The physical simulation is based on a ray tracing and grating modulation model. It generates a multi-view crosstalk noise map by steps such as light source discretization, screen surface sampling, specular reflection path calculation, grating modulation query and noise value accumulation.

[0020] Dynamic modulation involves constructing a human visual sensitivity weight map and an image content preservation weight map, and combining the two weight maps with the original noise map to generate a modulated noise map.

[0021] The adaptive pre-compensation adjusts the color and brightness values ​​of the original image in reverse according to the modulated noise map. The perceptual enhancement includes local contrast stretching, edge sharpening, and intelligent color saturation enhancement of adjacent pixels in the disturbed area.

[0022] Specifically, the human visual sensitivity weighting map is obtained by multiplying the illumination adaptation factor, the spatial sensitivity function, and the color sensitivity factor;

[0023] The light adaptation factor is related to the current average ambient brightness and the dark adaptation threshold brightness.

[0024] The spatial sensitivity function assigns a higher weight to the central area of ​​the screen.

[0025] The chromaticity sensitivity factor is calculated based on the chromaticity information of the image;

[0026] The image content protection weight map is obtained by analyzing the texture complexity and edge importance of pixel blocks.

[0027] Specifically, in S3:

[0028] The optical center of the virtual camera is located at the actual position of the viewer's eyes, and the imaging plane coincides with the physical screen plane. The lightweight neural radiation field renderer is built based on the Instant Neural Graphics Primitives framework. In the offline preprocessing stage, a multi-resolution hash-encoded micro multilayer perceptron network is trained using a high-precision 3D model of the scene or dense viewpoint video sampling data. During online real-time rendering, rays are generated according to camera parameters, and the trained model is called to synthesize RGB colors through forward propagation. The synthesis time of a 1080p resolution image per eye is controlled within 8 milliseconds.

[0029] Specifically, in S4:

[0030] Intelligent fusion is achieved through a view area weight matrix. The main view area where the audience is located is set to 1, and the guide layer image is used entirely. Other areas are set to 0, and the base layer image is used. A transition zone is set at the boundary between the two types of areas. The weights within the transition zone change linearly. The final pixel value is calculated from the pixel values ​​of the guide layer and the base layer according to the weights.

[0031] The viewpoint encoding uses a diagonal interlacing algorithm that matches the screen's slit grating parameters. It extracts pixel lines in a specific tilt angle direction from each viewpoint image and rearranges them into an RGB sub-pixel matrix. The backlight adjustment sends adjustment commands based on the overall intensity of ambient light to achieve a balance between energy consumption and image quality.

[0032] The naked-eye 3D large screen stereoscopic vision optimization system includes the following modules:

[0033] The environmental perception and intelligent analysis module is used to simultaneously capture panoramic lighting information and audience spatial pose through multiple sensors, and use a lightweight CNN model to decouple relevant parameters to generate an environmental light field model and audience geometric matrix.

[0034] The glare modeling and adaptive compensation module is used to simulate and generate crosstalk noise maps based on the ambient light field model and the screen optical property database, and modulate the noise map by combining visual sensitivity and image content features to drive adaptive pre-compensation and perception enhancement.

[0035] The personalized perspective real-time rendering module is used to generate virtual camera parameters based on the real-time position of the viewer's eyes, and to achieve real-time re-rendering of 3D scenes through a lightweight neural radiation field renderer.

[0036] The multi-layer fusion and driving encoding module is used to achieve intelligent fusion of base layer and guide layer images, perform viewpoint encoding to generate driving signals and synchronously adjust the backlight.

[0037] The technical effects and advantages of this invention are as follows:

[0038] Significantly improves the environmental adaptability and image quality stability of naked-eye 3D large screens. Through the synergy of multi-sensor fusion and a lightweight CNN model, it achieves accurate perception and dynamic modeling of ambient light sources, diffuse reflection characteristics, and screen surface properties. Combined with an online incremental learning mechanism, it ensures that the model continuously adapts to the lighting characteristics of actual deployment scenarios. It innovatively adopts glare crosstalk prediction based on physical simulation, coupled with a dynamic noise modulation strategy that combines human visual sensitivity with image content characteristics. This effectively suppresses glare interference while maximally protecting image texture, edge, and color details, avoiding distortion caused by over-processing. It achieves a dynamic balance between glare suppression and image quality protection, significantly improving display performance in complex lighting environments.

[0039] Enhance the realism of stereoscopic vision and adaptability to multiple scenes, optimizing the overall viewing experience. Utilizing a lightweight neural radiation field renderer based on the Instant-NGP framework, a personalized virtual camera is dynamically generated based on the viewer's real-time binocular positions, and millisecond-level real-time re-rendering is completed, perfectly reproducing motion parallax and significantly improving the immersive stereoscopic experience. Intelligent fusion of the base and guide layers is achieved through a view zone weight matrix, allocating display priority according to viewer position. This balances personalized optimization for key viewers with the general viewing needs of other viewers. Simultaneously, viewpoint encoding matching screen optical parameters and adaptive backlight adjustment balance energy consumption while ensuring display quality, forming a complete closed loop from environmental perception to adaptive output. This adapts to diverse scenarios such as outdoor advertising and venue displays, comprehensively enhancing the practicality and competitiveness of naked-eye 3D large screens. Attached Figure Description

[0040] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0042] like Figure 1 As shown, the steps for optimizing the stereoscopic vision of a naked-eye 3D large screen are as follows:

[0043] Step 1: Environmental Perception and Modeling. Through multi-sensor fusion and intelligent analysis, a digital understanding of the physical environment is constructed. HDR cameras, luminance meters, and depth sensors are used to simultaneously capture panoramic lighting information and audience positions. A lightweight CNN model decouples precise light source parameters, ambient diffuse reflection, and screen surface optical properties from the images, forming a real-time ambient light field model and audience geometry matrix, providing the data foundation for all subsequent adaptive processing. The specific process is as follows:

[0044] After the system starts up, the environmental awareness modules deployed around the screen bezel work synchronously. A set of high dynamic range (HDR) multispectral cameras captures panoramic images of a 180-degree field of view in front of the screen, focusing on analyzing high-brightness areas. At the same time, a high-precision luminance meter array measures the illuminance in different normal directions.

[0045] All data is transmitted to the core processing server in real time. The server runs a lightweight convolutional neural network model based on an encoder-decoder architecture. The encoder uses a simplified MobileNetV3-Small as its backbone network, and the input is a stitched four-channel image (R, G, B, Near-IR) provided by a multispectral HDR camera, normalized to a resolution of 512x256 pixels. The decoding part is divided into three parallel branches:

[0046] Light source parameter branch: Through two fully connected layers, output a five-dimensional vector [u,v,ω,I] r ,I g ], representing the horizontal azimuth angle, pitch angle, angular radius (radians) of the main point light source, and the relative intensity ratio in the R and G channels, respectively;

[0047] Diffuse light branch: Outputs a three-dimensional vector representing the average RGB intensity value of ambient diffuse light (normalized to the range of 0-1).

[0048] Surface property branch: Outputs a 64x32 resolution dual-channel feature map, where the first channel represents the specular reflection coefficient of a local area of ​​the screen and the second channel represents the anisotropy index.

[0049] The training dataset for this model was constructed in a laboratory environment using a programmable LED array to simulate light sources of different directions, colors, and intensities, illuminating screen samples with various known surface properties (BRDF), while simultaneously acquiring HDR images and real light source parameters. The loss function is the sum of the smoothed L1 loss of each branch output and the true label. After completing the basic training, to enable the model to continuously adapt to the unique lighting characteristics of the deployment environment (such as specific building lighting, fixed lamp positions, etc.), an online incremental learning mechanism is used to fine-tune key parameters while ensuring that the basic capabilities of the model remain unchanged. The specific process is as follows:

[0050] Confidence Assessment and Sample Selection: Define the output confidence of the CNN model in the current frame. For each branch output, calculate its confidence score:

[0051] Branch confidence of light source parameters:

[0052] ;

[0053] in, This represents the confidence score of the predicted light source parameters, with values ​​between 0 and 1. A higher score indicates a more reliable prediction. It is the square of the L2 norm (Euclidean distance) of the vector. This is the currently predicted five-dimensional light source parameter vector; and These are the exponentially weighted moving average and variance of the parameters corresponding to the past F frames (e.g., F=300), respectively. To prevent division by zero, a small constant is included. This formula measures the degree of agreement between the current prediction and the recent historical distribution.

[0054] Confidence level of diffuse reflection branch:

[0055] ;

[0056] in, This represents the confidence score for the diffuse reflection light prediction results, with values ​​between 0 and 1. This represents the currently predicted diffuse RGB intensity. This is the predicted value from the previous frame. It is a small constant. This formula measures the temporal continuity of the prediction.

[0057] If and only if (For example =0.7, When the confidence level is 0.9, the current frame is determined to be a high-confidence sample. The four-channel HDR input image of the current frame is then processed. Decouple the output parameter triples from the corresponding model (Representing light source, diffuse reflection, and surface property parameters, respectively) as a new training sample pair The data is stored in a fixed-capacity (e.g., 5000) incremental learning cache pool. The cache pool is managed using a first-in, first-out (FIFO) strategy.

[0058] Incremental fine-tuning execution strategy: Periodically adjust the execution strategy in the background or when the number of samples in the cache pool reaches a threshold. When the value reaches 1000, initiate incremental fine-tuning. The fine-tuning process follows these principles to prevent catastrophic amnesia:

[0059] Parameter freezing: Keep all weights of the CNN encoder (MobileNetV3-Small backbone network) fixed, and only freeze the weights of the last two fully connected / convolutional layers of the three decoder branches. Update.

[0060] Loss function: Based on the original smoothed L1 loss, an elastic weight consolidation constraint term is added: ,in, The total loss function used in the incremental fine-tuning phase. This is the original smoothed L1 loss function, which measures the difference between the model's predicted values ​​and the true labels. These are the parameters that need to be fine-tuned. It is its initial value (anchor point) after the last offline training ended; It is a parameter The importance weights are approximated by calculating the diagonal Fisher information matrix of the gradient of this parameter on the offline training set: λ; λ is the strength coefficient of the Elastic Weight Consolidation (EWC) constraint term, which controls the degree of protection for historical knowledge.

[0061] Training process: Randomly sample small batches of data from the incremental learning cache pool, using a reduced learning rate (e.g., 1 / 10 of the initial learning rate), and combine this with the loss function described above. Perform gradient descent optimization a few times (e.g., 3-5 epochs).

[0062] Through this mechanism, the model can gradually learn and adapt to lighting patterns that persist in the actual operating environment but are not fully covered in the laboratory (such as sunlight entering a window at a specific angle in the afternoon) without destroying existing knowledge, thereby providing more stable and accurate environmental light field decoupling results in long-term operation.

[0063] A lightweight convolutional neural network model, once trained, is able to decouple the following from HDR panoramic images: ① the three-dimensional orientation, angular radius (size model), and spectral power distribution of the main ambient light sources (such as windows and lamps); ② the overall diffuse reflection intensity and color temperature of the ambient light; and ③ the anisotropic reflectance characteristics of the screen surface, which may be caused by non-ideal coatings or dust. Simultaneously, a separate depth vision sensor continuously tracks the three-dimensional coordinates and orientation of all viewers.

[0064] Step Two: Glare Suppression and Compensation. Based on the model from Step One, a multi-viewpoint crosstalk noise map generated by glare modulation after grating is predicted through physical simulation. Subsequently, a dynamic noise suppression optimization based on human visual sensitivity and image content features is innovatively introduced to intelligently modulate the noise map. Based on this, adaptive pre-compensation and perceptual enhancement are performed on the original image, thereby eliminating glare while maximizing the preservation of image quality. The specific process is as follows:

[0065] The core processing server performs physical simulation calculations using the ambient light field model generated in step one and a pre-stored database of screen optical properties (containing precise grating geometry parameters, lens curvature, and BRDF surface reflectivity models). First, the simulation calculates the glare highlight area that will form on the screen surface under the current ambient light illumination, along with its precise two-dimensional brightness distribution map.

[0066] Secondly, the simulation further calculates the degree of modulation and interference that these surface glare rays will cause to the optical paths of different viewpoint channels after passing through the grating structure, generating a multi-viewpoint crosstalk noise map. This physical simulation is based on a simplified ray tracing and grating modulation model, and the specific steps are as follows:

[0067] a) Light source discretization: The main light source obtained in step one is discretized into a set of virtual point light sources uniformly distributed within the corresponding solid angle, based on its angular radius ω. b) Screen surface sampling: The screen surface is uniformly divided into Mu×Nu micro-surface elements. c) Specular reflection path calculation: For each micro-surface element... and each virtual light source Calculate the direction vector of the reflected light according to the law of specular reflection. d) Grating modulation lookup: A pre-stored grating viewing angle-transmittance lookup table. Obtained through optical simulation or actual measurement, it describes the effect of viewing from the screen normal direction at different horizontal angles. and pitch angle Below, the percentage of light that can pass through the grating and enter a specific viewpoint channel. The direction of reflection... Switch to perspective And query its transmittance for all viewpoint channels k. e) Noise value accumulation: For each screen pixel position p (and micro-facets) Correspondingly), its predicted crosstalk noise value in viewpoint channel k. Calculate by summing the following formulas:

[0068] ;

[0069] in, This is the predicted crosstalk noise value calculated at screen pixel position p and viewpoint channel k. Light source The strength, micro-facet at the angle of incidence The specular reflection term below (from the surface properties obtained in step one). For micro-facets The transmittance (percentage) of the reflected light rays entering the viewpoint channel k after passing through the grating is obtained by looking up the pre-stored grating viewpoint-transmittance lookup table G. To start from micro-element To pixel The mapping function iterates through all i and j, generating a set of multi-view crosstalk noise maps with the same number of original viewpoints.

[0070] Next, adversarial glare reduction and content-enhanced rendering are introduced: before generating multi-view images for each frame, the dynamic rendering engine loads the noise map. After loading the noise map, dynamic noise suppression optimization is performed based on human visual sensitivity and image content features, as follows:

[0071] Human visual sensitivity weight mapping: Based on the current ambient light level (from the ambient light field model in step one) and the characteristics of human visual perception, a visual sensitivity weight map is constructed. The weight map is identical to the screen resolution, and each pixel position (x, y) has a weight value. Calculated by the following formula:

[0072] ;

[0073] in: Let be the weight of human visual sensitivity at pixel coordinates (x, y). As a light adaptation factor, Let be the spatial sensitivity function. This refers to the color sensitivity factor.

[0074] Light adaptation factor : ,in The average brightness of the current environment. The dark adaptation threshold luminance is 3 cd / m². This factor simulates the overall sensitivity change of the human eye under different ambient luminance conditions.

[0075] Spatial sensitivity function : ,in , Coordinates of the screen center Control the decay rate, =0.3, =0.7. This function simulates the higher sensitivity of the human eye to the center area of ​​the screen.

[0076] Color sensitivity factor Calculate based on the chromaticity information of the input image: ,in This represents the chromaticity component vector of a pixel in the YUV color space. This is the adjustment factor (typical value 0.5). This factor reflects the human eye's greater sensitivity to color distortion in high-saturation areas.

[0077] Image content feature analysis: Performs real-time feature analysis on the original content image to generate a content protection weight map. For each pixel block (e.g., 8×8), calculate the following features:

[0078] Texture complexity : ,in For the current pixel block, The number of pixels within the block. For gradient operators, Let L2 be the image gradient at pixel (i,j), used to measure the intensity of local change at that point (i.e., texture richness). High-texture regions are less sensitive to noise.

[0079] marginal importance Canny edge detection is used to identify significant edges, specifically edge pixels. =1, non-edge region It decays exponentially with distance from the edge.

[0080] Based on the above characteristics, content protection weight The calculation is as follows:

[0081] ;

[0082] in This represents the content protection weight at pixel coordinates (x, y). A higher value indicates that the content in that area is more important or more fragile, requiring more careful handling. represents the edge importance value of the element (x, y), which is 1 at salient edges and decays exponentially with distance from the edge. , , To adjust the parameters, To prevent small quantities from being excluded.

[0083] Adaptive noise modulation: Combine the two weight maps with the original noise map N (containing all viewpoint channels). Combined, a modulated noise map is generated. :

[0084] For each viewpoint channel And pixel position (x,y):

[0085] ;

[0086] in: This represents the final noise value at pixel (x,y) and viewpoint channel k after visual and content feature modulation. This is the original predicted crosstalk noise value. This is a normalization factor to ensure the stability of the range of the weight product. Content adaptive decay function: ,in For the luminance component, The average brightness of the image. The standard deviation of brightness, Located in [0, 0.3], this is the attenuation intensity coefficient. This function makes noise suppression more effective in medium brightness areas because the human eye is most sensitive to noise in these areas.

[0087] Dynamic adjustment of compensation parameters: Based on the modulated noise map, dynamically adjust the two original operations in step two:

[0088] Adaptive pre-compensation:

[0089] The compensation strength is no longer a fixed value, but is based on the modulation noise value. Dynamic determination of original image features:

[0090] ;

[0091] That As a nonlinear function, it ensures the use of more refined compensation strategies in highly sensitive regions.

[0092] Content-aware enhancement: The scope and intensity of the enhanced perception are based on Adjustment:

[0093] For high-texture areas ( (Higher values) reduce contrast stretching and avoid texture distortion;

[0094] For important peripheral areas ( (Higher value), enhancing selective protection for edge sharpening;

[0095] For visually sensitive areas ( (With higher values), color saturation increases should be approached with greater caution;

[0096] This optimization enables intelligent and refined noise suppression: more precise compensation is applied to areas sensitive to human vision and important content areas, while unnecessary processing is reduced in insensitive areas or areas with complex textures, thereby eliminating glare interference while maximizing the preservation of the original visual quality and naturalness of the content.

[0097] For pixel regions with severe crosstalk predicted in the noise map, the engine does not simply increase brightness (which would exacerbate glare), but instead performs two operations:

[0098] Pre-compensation: Based on the noise model, the color and brightness values ​​of the region in the original image are adjusted in reverse so that the final value reaching the human eye after glare interference is close to the expected value.

[0099] Perceptual enhancement: In a color space (such as CIELAB), local contrast stretching and edge sharpening are performed on adjacent pixels in the disturbed area, and color saturation is intelligently improved, using the characteristics of the human visual system to enhance the perception of details.

[0100] Step 3: Dynamic perspective correction. For each tracked viewer, a pair of dedicated virtual cameras are defined based on their real-time eye positions. Using a pre-trained lightweight neural radiation field renderer, the 3D scene is re-rendered in real-time using these camera parameters, generating left and right eye images that match the current viewing angle. This achieves "motion parallax," significantly enhancing the realism and immersion of stereoscopic vision. The specific process is as follows:

[0101] While addressing glare, perspective issues are handled in parallel. For each tracked viewer, two virtual cameras are redefined based on the precise position of their eyes in physical space. The optical centers of these two cameras are located at the actual positions of the viewer's eyes, and the image planes coincide with the physical screen plane. Multi-view dynamic perspective recorrection is introduced: instead of using pre-set, fixed rendering perspectives, the rendering engine receives a 3D scene model (or a dense viewpoint image plus a depth map) from the content source.

[0102] For each viewer, the engine uses its own pair of virtual camera parameters to re-render the 3D scene in real time. When a viewer moves to the left, the image generated for their left eye realistically shows more of the right side of the 3D objects on the screen, perfectly matching the perspective motion laws of observing objects in the real world. To achieve high-performance real-time rendering, a lightweight renderer based on neural radiation fields is used. This renderer has been pre-optimized and trained for the target scene and can synthesize high-quality, continuous parallax image pairs within milliseconds based on the new camera parameters.

[0103] The renderer is built on the Instant Neural Graphics Primitives framework. For each specific 3D scene content to be played, the following operations are performed during the offline preprocessing stage:

[0104] From high-precision 3D models of the scene or dense viewpoint videos (such as 32 viewpoints), tens of thousands of sets of RGB images and camera parameters corresponding to virtual cameras in different spatial positions and orientations are rendered or sampled.

[0105] These data were used to train a miniature multilayer perceptron network with multi-resolution hash encoding. The network inputs spatial location (x, y, z) and view orientation. Mapped to volume density σ and RGB color value c;

[0106] After training is complete, the weights and hash table of the network are stored in the GPU memory of the perspective re-rendering node.

[0107] During online real-time rendering, for any viewer's eye position (corresponding to a camera matrix) specified by the server, the renderer performs the following operations:

[0108] Based on the camera parameters, generate a ray for each pixel on the screen;

[0109] By calling the already solidified Instant-NGP model, hierarchical volume sampling is performed along the ray through a single efficient forward propagation, and the final RGB color of the ray is directly synthesized.

[0110] By leveraging the parallel computing capabilities and half-precision floating-point arithmetic of modern GPUs, the time to synthesize a 1080p resolution image for a single eye is controlled within 8 milliseconds, thereby meeting the frame rate requirements for real-time rendering (≥60fps).

[0111] Step 4: Fusion Encoding Output. The glare-compensated general image layer is intelligently fused with the perspective correction guidance layer generated for the viewer, allocating display priorities according to the viewer's position within the screen space. The fused multi-viewpoint image is then processed by a viewpoint encoding algorithm matched to the screen's optical parameters to generate the final sub-pixel driving signal, and the backlight is adjusted synchronously, ultimately presenting an adaptively optimized stereoscopic image on the naked-eye 3D large screen; the specific process is as follows:

[0112] After compensation and reprojection processing, two sets of data were obtained: one set is the multi-viewpoint image base layer after glare pre-compensation and perception enhancement; the other set is a high-quality image pair (guide layer) generated for the main audience through dynamic perspective correction.

[0113] In the final output stage, intelligent fusion is performed: in the screen display area, the guide layer image is used first for the main viewer's field of view to ensure optimal perspective accuracy;

[0114] The blending process is completed by a display driver compositor hardware or FPGA module; a viewport weight matrix W corresponding to the screen pixel grid is maintained;

[0115] For the view area marked as the primary viewer, the W value of the corresponding screen area is set to 1, and the guide layer image data is used entirely.

[0116] For other regions, the W value is set to 0, and the base layer image data is used;

[0117] At the boundary between the two regions, a transition band with a width of D pixels is set, where the W value linearly changes from 1 to 0. Final pixel value... From the formula Calculations show that and These are the pixel values ​​corresponding to the guiding layer and the base layer, respectively. The fusion weight value corresponding to the pixel is calculated based on the viewer's position and the transition zone. The value is derived from the distribution defined by the view weight matrix W and has undergone boundary smoothing.

[0118] After fusion, the viewpoint encoding uses a diagonal interleaving algorithm based on the screen's slit tilt grating parameters (e.g., 8 viewpoints, grating tilt angle arctan(1 / 3)). The fused multiple viewpoint images are extracted along the pixel lines of each viewpoint image along a specific tilt angle and rearranged into a two-dimensional RGB sub-pixel matrix. This matrix is ​​the final signal that can directly drive the screen.

[0119] For other non-primary areas or untracked viewer areas, a generic base layer image is displayed. All image data undergoes viewpoint encoding and sub-pixel mapping algorithms to generate the final drive signal.

[0120] Simultaneously, based on the overall intensity of ambient light, adjustment commands are sent to the screen backlight system to achieve a balance between global energy consumption and image quality. Finally, the synchronized drive signal is sent to the naked-eye 3D large screen for display, completing a full closed loop from environmental perception to adaptive content presentation.

[0121] The stereoscopic vision optimization system modules for naked-eye 3D large screens are as follows:

[0122] The environmental perception and intelligent analysis module is responsible for the real-time digital perception of the physical environment. Through integrated multispectral HDR cameras, high-precision luminance meters, and depth sensors, it simultaneously captures panoramic lighting information in front of the screen and the spatial poses of all viewers. At its core is a lightweight convolutional neural network used to decouple precise light source parameters, environmental diffuse reflection characteristics, and the optical properties of the screen surface from the raw data in real time, generating a dynamic environmental light field model and viewer geometry matrix, providing a precise perceptual foundation for all subsequent optimizations.

[0123] The glare modeling and adaptive compensation module is used to combat visual interference caused by ambient light. It receives an ambient light field model and combines it with a pre-stored screen optical database for physical simulation, accurately predicting the crosstalk noise map generated by glare after raster modulation on each viewpoint channel. Subsequently, this module innovatively introduces an intelligent optimization algorithm based on human visual sensitivity and image content characteristics to dynamically modulate the noise map and drive the rendering engine to adaptively pre-compensate and perceptually enhance the original content. This eliminates glare crosstalk while maximizing the preservation of image detail and naturalness.

[0124] The personalized perspective real-time rendering module provides moving viewers with stereoscopic images that conform to the laws of real-world vision. For each tracked viewer, it dynamically generates a pair of unique virtual camera parameters based on their real-time eye positions. Utilizing a lightweight neural radiation field renderer pre-trained on the Instant-NGP framework, this module can re-render the 3D scene in real-time based on the new camera parameters with millisecond latency, synthesizing personalized left and right eye views with correct motion parallax, significantly enhancing the immersion and realism of stereoscopic vision.

[0125] The multi-layer fusion and driving encoding module intelligently fuses the base layer image (after general glare compensation) with the perspective-corrected guiding layer image generated for a specific viewer, prioritizing and smoothly transitioning the images within the screen space based on the viewer's position. The fused multi-viewpoint image stream is then processed by a viewpoint encoding and sub-pixel mapping algorithm precisely matched to the physical parameters of the naked-eye 3D screen to generate a directly driveable display signal. This signal simultaneously controls the backlight system for overall brightness adjustment, ultimately outputting a high-quality, adaptively optimized naked-eye 3D image on the screen.

[0126] The above formulas are all dimensionless calculations. Dimensionless calculations can be performed using various methods such as standardization, which will not be elaborated here. The formulas are derived from software simulations based on a large amount of collected data, and the preset parameters in the formulas can be set by those skilled in the art according to the actual situation.

[0127] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, ATA hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state ATA hard disk.

[0128] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0129] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0130] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0131] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0132] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0133] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable ATA hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0134] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for optimizing stereoscopic vision on a naked-eye 3D large screen, characterized in that, Includes the following steps: S1. Environmental Perception and Modeling: Through multi-sensor fusion and intelligent analysis, a digital understanding of the physical environment is constructed, capturing panoramic lighting information and audience position. A lightweight CNN model is used to decouple light source parameters, environmental diffuse reflection, and screen surface optical properties to form a real-time environmental light field model and audience geometric matrix. S2. Glare Suppression and Compensation: Physical simulation is performed based on the ambient light field model and screen optical property database to generate a multi-viewpoint crosstalk noise map. The noise map is dynamically modulated by combining human visual sensitivity and image content features. Then, adaptive pre-compensation and perception enhancement are performed on the original image. S3, Dynamic Perspective Correction: Define a dedicated virtual camera based on the real-time eye positions of each tracked viewer, and use a lightweight neural radiation field renderer to re-render the 3D scene in real time to generate left and right eye images that match the current viewing angle. In S3: The optical center of the virtual camera is located at the actual position of the viewer's eyes, and the imaging plane coincides with the physical screen plane; the lightweight neural radiation field renderer is built based on the Instant Neural Graphics Primitives framework. In the offline preprocessing stage, a multi-resolution hash-encoded micro multilayer perceptron network is trained through high-precision 3D models of the scene or dense viewpoint video sampling data. During online real-time rendering, rays are generated according to camera parameters, and the trained model is called to synthesize RGB colors through forward propagation. The synthesis time of a 1080p resolution image per eye is controlled within 8 milliseconds. S4. Fusion Encoding Output: The multi-viewpoint image base layer after glare compensation is intelligently fused with the guide layer image generated by perspective correction. The display priority is assigned according to the viewer's position. The fused image generates sub-pixel driving signals through viewpoint encoding, which synchronously adjust the backlight and outputs it to the naked-eye 3D screen.

2. The method for optimizing stereoscopic vision in a naked-eye 3D large screen according to claim 1, characterized in that, The multiple sensors in S1 include: Multispectral HDR camera array, high-precision brightness meter and depth sensor; The multispectral HDR camera captures a panoramic image of a 180-degree field of view in front of the screen, the high-precision luminance meter measures the light intensity in different normal directions, and the depth sensor tracks the three-dimensional coordinates and orientation of the viewer. After training, the lightweight CNN model is fine-tuned for key parameters through an online incremental learning mechanism. The online incremental learning mechanism includes confidence assessment and sample selection, incremental fine-tuning execution steps, selecting high-confidence samples and storing them in a cache pool, freezing encoder weights during fine-tuning, and only updating the relevant weights of decoder branches.

3. The method for optimizing stereoscopic vision in a naked-eye 3D large screen according to claim 2, characterized in that, The lightweight CNN model: The model employs an encoder-decoder architecture. The encoder uses MobileNetV3-Small as its backbone network and takes a four-channel stitched image provided by a multispectral HDR camera as its input. The decoder consists of three parallel branches that output light source parameters, the average RGB intensity value of ambient diffuse light, and the optical property feature map of the screen surface, respectively. The training dataset of the model is constructed by simulating different light sources illuminating screen samples with known surface characteristics using a programmable LED array, and simultaneously acquiring HDR images and real light source parameters.

4. The method for optimizing stereoscopic vision in a naked-eye 3D large screen according to claim 1, characterized in that, In S2: The physical simulation is based on a ray tracing and grating modulation model. It generates a multi-view crosstalk noise map by steps such as light source discretization, screen surface sampling, specular reflection path calculation, grating modulation query and noise value accumulation. Dynamic modulation involves constructing a human visual sensitivity weight map and an image content preservation weight map, and combining the two weight maps with the original noise map to generate a modulated noise map. The adaptive pre-compensation adjusts the color and brightness values ​​of the original image in reverse according to the modulated noise map. The perceptual enhancement includes local contrast stretching, edge sharpening, and intelligent color saturation enhancement of adjacent pixels in the disturbed area.

5. The method for optimizing stereoscopic vision in a naked-eye 3D large screen according to claim 4, characterized in that, The human visual sensitivity weighting map is obtained by multiplying the illumination adaptation factor, the spatial sensitivity function, and the color sensitivity factor. The light adaptation factor is related to the current average ambient brightness and the dark adaptation threshold brightness. The spatial sensitivity function assigns a higher weight to the central area of ​​the screen. The chromaticity sensitivity factor is calculated based on the chromaticity information of the image; The image content protection weight map is obtained by analyzing the texture complexity and edge importance of pixel blocks.

6. The method for optimizing stereoscopic vision in a naked-eye 3D large screen according to claim 1, characterized in that, In S4: Intelligent fusion is achieved through a view area weight matrix. The main view area where the audience is located is set to 1, and the guide layer image is used entirely. Other areas are set to 0, and the base layer image is used. A transition zone is set at the boundary between the two types of areas. The weights within the transition zone change linearly. The final pixel value is calculated from the pixel values ​​of the guide layer and the base layer according to the weights. The viewpoint encoding uses a diagonal interlacing algorithm that matches the parameters of the screen's slit tilted grating; the backlight adjustment sends adjustment commands based on the overall intensity of the ambient light to achieve a balance between energy consumption and image quality.

7. A system applied to the stereoscopic vision optimization method for naked-eye 3D large screens according to any one of claims 1-6, characterized in that, Includes the following modules: The environmental perception and intelligent analysis module is used to simultaneously capture panoramic lighting information and audience spatial pose through multiple sensors, and use a lightweight CNN model to decouple relevant parameters to generate an environmental light field model and audience geometric matrix. The glare modeling and adaptive compensation module is used to simulate and generate crosstalk noise maps based on the ambient light field model and the screen optical property database, and modulate the noise map by combining visual sensitivity and image content features to drive adaptive pre-compensation and perception enhancement. The personalized perspective real-time rendering module is used to generate virtual camera parameters based on the real-time position of the viewer's eyes, and to achieve real-time re-rendering of 3D scenes through a lightweight neural radiation field renderer. The multi-layer fusion and driving encoding module is used to achieve intelligent fusion of base layer and guide layer images, perform viewpoint encoding to generate driving signals and synchronously adjust the backlight.

Citation Information

Patent Citations

  • Mask-variation human-eye tracking method of 3D (Three Dimensional) display for naked eyes

    CN103051909A

  • Commodity display interaction visualization method and device

    CN121033272A