Inner depth controllable cylindrical lens grating naked eye 3D display method and system

CN122525807APending Publication Date: 2026-08-07SHENZHEN VIEWON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN VIEWON CO LTD
Filing Date
2026-06-02
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]现有技术具有以下缺陷,通过固定负视差(crossed disparity)实现内景深,但深度范围固定,无法根据内容自适应调整

Benefits of technology

[0017] The present invention provides a method and system for naked-eye 3D display with controllable inner depth of field using a lenticular lens grating, comprising: receiving a single two-dimensional source image; analyzing the scene semantic features of the two-dimensional source image through a pre-trained neural network model; predicting the multi-layer depth distribution of the corresponding scene behind the display null plane based on the scene semantic features; generating inner depth control parameters; assigning differentiated negative parallax values ​​to different semantic regions of the two-dimensional source image according to the inner depth control parameters, so that each semantic region is imaged at different depth layers behind the display null plane; determining a target microstructure region from multiple microstructure regions of the lenticular lens grating based on the negative parallax values; wherein each microstructure region has a different optical pitch, corresponding to different inner depth control ranges; and performing pixel interleaving on a multi-view image sequence based on the negative parallax values ​​and the target microstructure region to generate naked-eye 3D video frames adapted for inner depth display. In this invention, a pre-trained neural network model analyzes the scene semantic features of the two-dimensional source image to predict the multi-layer depth distribution of the corresponding scene behind the zero plane of the display. Differential negative parallax values ​​are then assigned. Based on these negative parallax values, target microstructure regions are determined from multiple microstructure regions of the lenticular lens plate. Pixel interleaving is performed on the multi-view image sequence to generate glasses-free 3D video frames adapted for inner depth display. This invention achieves automatic calculation of multi-layer inner depth distribution through AI analysis of the scene semantic features of the two-dimensional source image, and realizes layered stereoscopic display of depth behind the zero plane of the display through pixel-level parallax control. This invention generates glasses-free 3D video frames that are fully adapted to the beam splitting characteristics of the lenticular lens plate, can be directly output to the display panel, and can present a multi-layered, controllable, and highly comfortable inner depth stereoscopic effect. Viewers intuitively feel that the image content is naturally embedded inside the screen, rather than jumping out of the screen, achieving an immersive, low-fatigue, and highly realistic inner depth glasses-free 3D display effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122525807A_ABST
    Figure CN122525807A_ABST
Patent Text Reader

Abstract

The application provides a lenticular grating naked-eye 3D display method and system with controllable internal depth of field, comprising: receiving a single two-dimensional source image, analyzing the scene semantic features of the two-dimensional source image through a pre-trained neural network model, predicting the multi-layer depth distribution of the corresponding scene behind the display zero plane based on the scene semantic features, and generating internal depth of field control parameters; assigning different semantic regions of the two-dimensional source image with differentiated negative parallax values according to the internal depth of field control parameters; determining a target microstructure region from a plurality of microstructure regions of a lenticular grating plate according to the negative parallax values; and performing pixel interleaving on a multi-view image sequence based on the negative parallax values and the target microstructure region, to generate a naked-eye 3D video frame suitable for internal depth of field display. In the application, the scene semantic features of the two-dimensional source image are analyzed through AI, the multi-layer internal depth of field distribution is automatically calculated, and depth layered stereoscopic display behind the display zero plane is realized through pixel-level parallax control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical technology, and in particular to a method and system for naked-eye 3D display using a cylindrical grating with controllable depth of field. Background Technology

[0002] Current glasses-free 3D display technology can be divided into two categories based on the position of the stereoscopic image: (1) Pop-out: The stereoscopic image is projected in front of the zero plane of the display and "jumps out" of the screen towards the viewer; this method requires a large convergence parallax and is prone to causing visual fatigue; it is usually suitable for emphasis, but it is uncomfortable to watch for a long time.

[0003] (2) Window Effect: The stereoscopic image is projected behind the zero plane of the display and "recessed" into the screen; this method has a smaller parallax and conforms to the natural viewing habits of the human eye; it is suitable for immersive display and has a high level of comfort for long-term viewing.

[0004] Current lenticular lens grating-based glasses-free 3D display technologies primarily focus on creating a strong sense of stereoscopic leaping from the outer depth of field, while research on the refined control of the inner depth of field is insufficient. In fact, inner depth of field is of significant value in the following scenarios: Museum artifact display: Artifacts should be "placed" inside the display case, rather than jumping out and disrupting the sense of space in the display case; Architecture / Interior Design Preview: The space should "extend" behind the screen, rather than jump out of the obscuring interface; In-vehicle information display: Navigation information should be "embedded" inside the dashboard to avoid distracting the driver's attention.

[0005] Existing technologies have the following drawbacks: they achieve interior depth by using fixed negative disparity (crossed disparity), but the depth range is fixed and cannot be adaptively adjusted according to content. Existing technologies only achieve a single layer of interior depth, failing to achieve multi-layered depth gradation and thus unable to present the depth layers of complex scenes. Furthermore, existing technologies do not incorporate AI content analysis into interior depth control, and therefore cannot automatically determine the optimal depth distribution based on scene semantics. Summary of the Invention

[0006] The main objective of this invention is to provide a naked-eye 3D display method and system with controllable inner depth of field using cylindrical gratings. The aim is to automatically calculate the distribution of multiple inner depth layers by analyzing the scene semantic features of two-dimensional source images using AI, and to achieve depth-layered stereoscopic display behind the zero plane of the display through pixel-level parallax control.

[0007] To achieve the above objectives, the present invention provides a method for naked-eye 3D display using a cylindrical grating with controllable depth of field, comprising the following steps: Receive a single two-dimensional source image, analyze the scene semantic features of the two-dimensional source image through a pre-trained neural network model, predict the multi-layer depth distribution of the corresponding scene behind the display zero plane based on the scene semantic features, and generate interior depth control parameters. Based on the inner depth control parameters, differentiated negative disparity values ​​are assigned to different semantic regions of the two-dimensional source image so that each semantic region is imaged at a different depth layer behind the display null plane. Based on the negative parallax value, the target microstructure region is determined from multiple microstructure regions of the cylindrical lens grating plate; wherein each microstructure region has a different optical pitch, corresponding to a different inner depth of field control range; Based on the negative parallax value and the target microstructure region, pixel interleaving is performed on the multi-view image sequence to generate naked-eye 3D video frames adapted for interior depth display.

[0008] Furthermore, the scene semantic features include: Scene type features: used to identify whether the two-dimensional source image belongs to an indoor scene, an outdoor scene, a product display scene, or an art display scene; Spatial hierarchy features: used to analyze the spatial layout relationships between foreground, midground, and background objects in a scene; Perspective depth features: used to extract linear perspective cues and atmospheric perspective cues in a scene and estimate the depth of the scene.

[0009] Furthermore, based on scene semantic features, the multi-layer depth distribution of the corresponding scene behind the display zero plane is predicted, generating interior depth control parameters, including: Based on the semantic features of the scene, determine the number of interior depth layers N, where N≥2; Assign a depth range [d_min, d_max] to each layer, where d represents the virtual depth distance behind the display zero plane; Based on the expected spatial location of each semantic region in the scene of the two-dimensional source image, it is mapped to the corresponding depth layer.

[0010] Furthermore, the negative parallax value is positively correlated with the virtual depth distance; the upper limit of the negative parallax value is determined based on the range of interior depth for comfortable viewing by the human eye; the negative parallax difference between adjacent depth layers is determined based on the interlayer depth interval and the minimum resolvable depth difference of the human eye.

[0011] Furthermore, the lenticular lens grating plate is attached to the display zero plane, including: The first microstructure region, with an optical pitch of P1, corresponds to the near-field depth control range, which is 0~50mm behind the display null plane. The second microstructure region, with an optical pitch of P2, corresponds to the field depth control range, which is 50~150mm behind the display zero plane. The third microstructure region, with an optical pitch of P3, corresponds to the far-field depth-of-field control range, which is 150~300mm behind the display zero plane; Among them, P1 < P2 < P3, and the smaller the pitch, the finer the depth resolution.

[0012] Furthermore, it also includes: Based on the distribution ratio of each depth layer in the inner depth control parameters, the target microstructure region is determined before pixel interleaving; Align the switching action of the target microstructure region with the display timing of the naked-eye 3D video frame.

[0013] Furthermore, after assigning differentiated negative disparity values ​​to different semantic regions of the two-dimensional source image, the method further includes: Detect whether the negative disparity value exceeds a preset safety threshold; When the negative disparity value of one of the semantic regions exceeds the safety threshold, the corresponding semantic region is automatically compressed to the nearest safe depth layer, and a visual cue is generated.

[0014] Furthermore, the source scenarios for the single two-dimensional source image include: digital display images of museum artifacts, interior renderings of real estate projects, screenshots of vehicle information systems, product display images of commercial retail terminals, three-dimensional scene projection images of digital twin systems, interactive pet scenarios, and animation content displays.

[0015] Furthermore, the pixel interleaving includes: An independent pixel interlacing mask is generated for each depth layer, and the period and phase of the interlacing mask are determined based on the negative parallax value of the corresponding layer and the optical pitch of the target microstructure region. Overlay and blend all the interlaced masks of the depth layers to generate a multi-layered joint interlaced mask of inner depth of field. The pixel arrangement of the multi-view image sequence is performed using the joint interleaving mask.

[0016] This invention also provides a naked-eye 3D display system with controllable depth of field using lenticular lenses, comprising: The analysis unit is used to receive a single two-dimensional source image, analyze the scene semantic features of the two-dimensional source image through a pre-trained neural network model, predict the multi-layer depth distribution of the corresponding scene behind the display zero plane based on the scene semantic features, and generate interior depth control parameters. The allocation unit is used to allocate differentiated negative disparity values ​​to different semantic regions of the two-dimensional source image according to the inner depth control parameters, so that each semantic region is imaged at different depth layers behind the display null plane. The determining unit is used to determine the target microstructure region from multiple microstructure regions of the cylindrical lens grating plate based on the negative parallax value; wherein each microstructure region has a different optical pitch and corresponds to a different inner depth of field control range. The generation unit is used to perform pixel interleaving on the multi-view image sequence based on the negative parallax value and the target microstructure region to generate naked-eye 3D video frames adapted for interior depth display.

[0017] The present invention provides a method and system for naked-eye 3D display with controllable inner depth of field using a lenticular lens grating, comprising: receiving a single two-dimensional source image; analyzing the scene semantic features of the two-dimensional source image through a pre-trained neural network model; predicting the multi-layer depth distribution of the corresponding scene behind the display null plane based on the scene semantic features; generating inner depth control parameters; assigning differentiated negative parallax values ​​to different semantic regions of the two-dimensional source image according to the inner depth control parameters, so that each semantic region is imaged at different depth layers behind the display null plane; determining a target microstructure region from multiple microstructure regions of the lenticular lens grating based on the negative parallax values; wherein each microstructure region has a different optical pitch, corresponding to different inner depth control ranges; and performing pixel interleaving on a multi-view image sequence based on the negative parallax values ​​and the target microstructure region to generate naked-eye 3D video frames adapted for inner depth display. In this invention, a pre-trained neural network model analyzes the scene semantic features of the two-dimensional source image to predict the multi-layer depth distribution of the corresponding scene behind the zero plane of the display. Differential negative parallax values ​​are then assigned. Based on these negative parallax values, target microstructure regions are determined from multiple microstructure regions of the lenticular lens plate. Pixel interleaving is performed on the multi-view image sequence to generate glasses-free 3D video frames adapted for inner depth display. This invention achieves automatic calculation of multi-layer inner depth distribution through AI analysis of the scene semantic features of the two-dimensional source image, and realizes layered stereoscopic display of depth behind the zero plane of the display through pixel-level parallax control. This invention generates glasses-free 3D video frames that are fully adapted to the beam splitting characteristics of the lenticular lens plate, can be directly output to the display panel, and can present a multi-layered, controllable, and highly comfortable inner depth stereoscopic effect. Viewers intuitively feel that the image content is naturally embedded inside the screen, rather than jumping out of the screen, achieving an immersive, low-fatigue, and highly realistic inner depth glasses-free 3D display effect. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the steps of a naked-eye 3D display method with controllable inner depth of field using a cylindrical grating in one embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a naked-eye 3D display device with controllable inner depth of field using a cylindrical lens grating, according to an embodiment of the present invention. Figure 3 This is a structural block diagram of a naked-eye 3D display system with controllable inner depth of field using a cylindrical lens grating, according to one embodiment of the present invention.

[0019] The implementation, functional features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0021] It is particularly important to note that all technical steps, algorithm applications, and parameter settings in the technical solution of this application have clear technical objectives and application value. They do not utilize complex steps and algorithmic formulas to achieve simple functions. To provide detailed explanations of each step and avoid ambiguity, some conventional algorithms are used for illustration. However, this does not mean that the algorithms and technical features listed herein are the only way to implement the technical solution of this application, nor is it intended to limit the scope of protection of this application. This application is not a combination or stacking of the listed algorithms and technical features; its essence is to exemplify the implementation methods of this application to fully explain it. It does not pursue formal complexity by adding meaningless technical steps, nor does it involve the accumulation of technologies divorced from practical needs; it conforms to the conventional logic of technical improvement and design.

[0022] Reference Figure 1 One embodiment of the present invention provides a naked-eye 3D display method with controllable inner depth of field using a cylindrical grating, comprising the following steps: Step S1: Receive a single two-dimensional source image, analyze the scene semantic features of the two-dimensional source image through a pre-trained neural network model, predict the multi-layer depth distribution of the corresponding scene behind the display zero plane based on the scene semantic features, and generate interior depth control parameters. Step S2: Based on the inner depth control parameters, assign differentiated negative disparity values ​​to different semantic regions of the two-dimensional source image so that each semantic region is imaged at different depth layers behind the display null plane. Step S3: Based on the negative parallax value, determine the target microstructure region from multiple microstructure regions of the cylindrical lens grating plate; wherein each microstructure region has a different optical pitch, corresponding to a different inner depth of field control range; Step S4: Based on the negative parallax value and the target microstructure region, perform pixel interleaving on the multi-view image sequence to generate naked-eye 3D video frames adapted for interior depth display.

[0023] In this embodiment, the provided method for naked-eye 3D display with controllable inner depth using a cylindrical grating is based on a single two-dimensional source image. Through scene semantic intelligent analysis, multi-layer depth planning, differentiated negative parallax allocation, optical structure matching, and pixel interleaving processing, it achieves multi-level, high-precision, and high-comfort naked-eye 3D imaging of inner depth behind the zero plane. The complete execution flow is as follows: First, step S1 is executed, receiving a single 2D source image to be converted into 3D depth of field. This 2D source image can be a planar material suitable for depth of field presentation, such as a digital display image of museum artifacts, a real estate interior rendering, an in-vehicle information system interface, a commercial product display image, a digital twin scene image, an interactive pet scene image, or an animation content display scene image. This 2D source image is input into a deep neural network model trained with a large number of labeled samples. The model performs comprehensive and refined scene semantic feature analysis on the image. On the one hand, it identifies and determines the scene type corresponding to the image, distinguishing between indoor scenes, outdoor scenes, product display scenes, art display scenes, or in-vehicle interactive scenes. On the other hand, it extracts the relative positional relationships, occlusion relationships, and distribution structure of foreground objects, midground objects, and background objects in the image through a spatial hierarchy detection algorithm, forming a corresponding spatial hierarchy mask. Simultaneously, based on lines... Visual cues such as perspective, atmospheric perspective, texture gradient, and relative object size are used to estimate the overall depth and spatial scale of a scene. After fully acquiring scene type features, spatial hierarchy features, and perspective depth features, the semantic features of the scene are used as the direct and sole basis for calculation. The multi-layer depth distribution of the scene content in the virtual imaging space extending inward along the display normal from the light-emitting surface of the display panel to the zero plane is predicted. The total number of layers required for the inner depth, the virtual depth distance range corresponding to each layer, and the depth layer to which each semantic region in the image should belong are determined. Finally, based on the above multi-layer depth distribution results, inner depth control parameters are generated, including the number of inner depth layers, the depth range of each layer, the mapping relationship between each semantic region and the depth layer, and the overall depth safety constraint threshold. This provides complete, unified, and precise control instructions for subsequent parallax allocation and raster region matching.

[0024] After generating the interior depth control parameters, step S2 is executed. Using the interior depth control parameters output in step S1 as the core processing basis, differentiated negative disparity values ​​are allocated according to preset rules for different semantic regions obtained through semantic segmentation and spatial partitioning in the two-dimensional source image. The negative disparity value, as a key optical parameter for achieving concave imaging behind the display null plane, follows strict quantization rules: the magnitude of the negative disparity value is positively correlated with the virtual depth distance behind the display null plane; that is, the greater the virtual depth and the closer the imaging position is to the inside of the screen, the larger the allocated negative disparity value. Simultaneously, the maximum value of the negative disparity value is limited to within a certain range. Within the safe threshold range for comfortable viewing of the inner depth, problems such as visual fatigue, dizziness, or failure of stereoscopic perception caused by excessive parallax are avoided. The negative parallax difference between adjacent depth layers is precisely set according to the interlayer depth interval and the minimum resolvable depth difference of the human eye to ensure smooth depth transition, no jumps, and no optical crosstalk. Through the above-mentioned refined negative parallax allocation, each semantic region in the two-dimensional source image can be accurately and stably imaged on preset different depth levels behind the zero plane of the display after being beam-splitting and imaging by the cylindrical lens grating. This constructs a multi-layered inner depth stereoscopic structure with distinct layers, accurate spatial relationships, and conforms to the natural viewing habits of the human eye.

[0025] Subsequently, step S3 is executed. The cylindrical lens grating plate used in this invention is not a single optical structure, but integrates at least two microstructure regions with different optical pitches. Different microstructure regions correspond to different inner depth-of-field control ranges. Among them, the small pitch region corresponds to the near-field inner depth-of-field control range, the medium pitch region corresponds to the mid-field inner depth-of-field control range, and the large pitch region corresponds to the far-field inner depth-of-field control range. The smaller the optical pitch, the higher the corresponding depth control accuracy and imaging resolution. This step is based on the negative disparity values ​​of each semantic region determined in step S2, the virtual depth corresponding to each negative disparity value, and the overall depth layer distribution ratio, from the cylindrical lens... The grating plate matches, filters, and determines target microstructure regions that are highly compatible with the current display requirements across multiple microstructure regions. The selection logic is based on the proportion of depth distribution as the core criterion. When the proportion of near-field depth region is high, small-pitch microstructure regions are selected; when the proportion of mid-field depth region is high, medium-pitch microstructure regions are selected; and when the proportion of far-field depth region is high, large-pitch microstructure regions are selected. Through this optical parameter matching process, the optical characteristics of the lenticular grating are made to perfectly match the requirements of inner depth control, ensuring that the image after subsequent pixel interleaving can form a clear, distortion-free, and depth-accurate inner depth stereoscopic image under the target microstructure region.

[0026] Finally, step S4 is executed, using the negative parallax value determined in step S2 and the target microstructure region determined in step S3 as dual core control conditions to perform pixel interleaving processing on the multi-view image sequence necessary for naked-eye 3D display of the lenticular lens grating. This multi-view image sequence is a common input material in the field of naked-eye 3D display and can be obtained through viewpoint rendering based on two-dimensional source images and depth information, multi-view shooting, or multi-view synthesis. During the pixel interleaving process, the pixel offset, arrangement weight, and layered display rules of each viewpoint image are determined according to the negative parallax value. At the same time, the period, phase, and mask of pixel interleaving are determined according to the optical pitch of the target microstructure region. The model shape and superposition method first generate independent pixel interlacing masks for each depth region, and then superimpose and merge the masks into a unified multi-layer inner depth joint interlacing mask. Using this joint interlacing mask, the multi-view image sequence is rearranged pixel by pixel. After the above precise interlacing processing, the final naked-eye 3D video frame is generated that is fully adapted to the beam splitting characteristics of the lenticular lens grating, can be directly output on the display panel, and can present a multi-level, controllable, and highly comfortable inner depth stereoscopic effect. This allows the viewer to intuitively feel that the picture content is naturally embedded in the screen, rather than jumping out of the screen, thus achieving an immersive, low-fatigue, and highly realistic inner depth naked-eye 3D display effect.

[0027] Reference Figure 2 In this embodiment, the layered structure of the naked-eye 3D display device with controllable depth of field consists of three core stacked components from the inside out (i.e., from the light signal generation to the emission direction): a display panel, a lenticular lens plate, and an outer protective / optical modulation layer. The structure and function of each part are as follows: The display panel, as the innermost component, is the carrier for generating naked-eye 3D images, used to display multi-view image sequences after pixel interlacing processing. The light output from the panel carries pixel information from each viewpoint, providing the basic light signal for subsequent lenticular grating beam splitting imaging.

[0028] The lenticular lens grating, fitted to the display's zero plane, is the core optical component for achieving inner depth-of-field beam splitting control. This grating integrates multiple microstructure regions with different optical pitches. These different microstructure regions correspond to different inner depth-of-field control ranges. Based on preset negative parallax values ​​and depth layer distribution, it can directionally deflect multi-view light emitted from the display panel, projecting light from different angles to the viewer's eyes to create a stereoscopic perception. Simultaneously, it enables inner depth imaging at different depth levels behind the display's zero plane.

[0029] The outer protective / optical modulation layer, as the outermost component, covers the surface of the lenticular lens plate and has the dual functions of structural protection and optical optimization: on the one hand, it provides physical protection for the lenticular lens plate to prevent it from being damaged by external forces and environmental factors; on the other hand, it optimizes the light intensity distribution of the emitted light and reduces optical crosstalk through the optical modulation structure, thereby improving the uniformity of the image and the overall display quality.

[0030] In this embodiment, the pre-trained neural network model consists of five core modules: an input layer, a scene classifier, a spatial hierarchy detector, a perspective depth estimator, and a negative disparity calculator. These modules work together sequentially to achieve end-to-end generation of multi-layer interior depth control parameters from a two-dimensional image. The workflow is as follows: The input layer takes a single two-dimensional source image to be processed as input. The image is in the format of a three-channel color image (H×W×3, where H is the image height and W is the image width), providing raw data input for all subsequent modules.

[0031] The scene classifier, with ResNet-18 network as its backbone, receives two-dimensional source images and extracts global semantic features, outputting the scene type to which the image belongs (including indoor, outdoor, product display, art display, etc.), providing a scene-level control strategy basis for subsequent deep distribution planning.

[0032] The spatial hierarchy detector, also based on DenseNet, outputs two key pieces of information: a pixel-level depth map, reflecting the relative depth relationships of different regions in the image; and a spatial hierarchy mask for the foreground, midground, and background, combined with the FPN feature pyramid structure, which clarifies the hierarchical affiliation of different semantic regions and provides a spatial structural basis for multi-layer depth segmentation.

[0033] A perspective depth estimator receives scene type, spatial hierarchy mask, and depth. Figure 3 The class information is comprehensively processed using DenseNet as the backbone network, and finally outputs complete interior depth control parameters, including: the number of interior depth layers N; the depth range corresponding to each depth layer; and the mapping relationship between each semantic region and the depth layer, providing accurate depth control instructions for subsequent negative disparity calculation.

[0034] The negative parallax calculator takes the depth range output by the perspective depth estimator and the cylindrical grating parameters as input to calculate the differential negative parallax values ​​corresponding to each depth layer, providing core control parameters for subsequent pixel interleaving and optical imaging.

[0035] In addition, the training data for the neural network model is a 2D image-3D display effect pairing dataset labeled with multi-layer depth distribution, which is used to perform end-to-end supervised training on the entire neural network architecture to ensure that the depth control parameters output by the model are highly matched with the expected display effect.

[0036] In one embodiment, the scene semantic features include: Scene type features: used to identify whether the two-dimensional source image belongs to an indoor scene, an outdoor scene, a product display scene, or an art display scene; Spatial hierarchy features: used to analyze the spatial layout relationships between foreground, midground, and background objects in a scene; Perspective depth features: used to extract linear perspective cues and atmospheric perspective cues in a scene and estimate the depth of the scene.

[0037] In this embodiment, the scene type feature is the classification and identification result of the overall application scenario and display purpose of the two-dimensional source image, serving as the top-level decision-making basis for the interior depth control strategy. This feature uses a pre-trained neural network to discriminate the overall semantic information of the image, accurately identifying the scene category to which the two-dimensional source image belongs. Specifically, it includes typical types such as indoor scenes, outdoor scenes, product display scenes, and art display scenes. Different scene types correspond to different interior depth design principles, depth layers, depth range, and safety constraints: for example, indoor scenes emphasize spatial extension and layer separation, art display scenes emphasize subject prominence and detail reproduction, vehicle scenes need to meet lightweight, low-interference, and high-safety interior depth constraints, and product display scenes emphasize the clear presentation of the subject's outline and spatial layers. Through the extraction and identification of scene type features, a basic control strategy matching scene attributes can be provided for subsequent multi-layer depth distribution planning, ensuring a high degree of adaptation between the interior depth effect and the image application scenario.

[0038] Spatial hierarchy features are a quantitative description of the relative positions, occlusion relationships, and hierarchical structure of objects within a scene, serving as the core basis for achieving multi-layered depth of field. This feature, through spatial structure analysis and semantic segmentation algorithms using neural networks, hierarchically divides visual elements in a 2D source image. It is used to analyze and determine the spatial layout, stacking, size proportions, and visual hierarchy of foreground, midground, and background objects in the scene. Foreground objects are typically the main subject or focus of attention and should be assigned a shallower depth of field; midground objects, as spatial transitions, correspond to a medium depth of field; and background objects contribute to the overall spatial atmosphere and correspond to a deeper depth of field. By extracting spatial hierarchy features, the priority and affiliation of each object behind the zero plane can be clearly defined, providing a clear structural basis for multi-layered depth allocation and ensuring that the final depth of field image conforms to the natural perceptual logic of human vision of real space.

[0039] Perspective depth features are objective visual cues extracted from two-dimensional images to reflect the depth of three-dimensional space, and are a key basis for accurate depth numerical calculation. This feature uses a neural network to analyze the geometric structure and optical attenuation laws of the image to extract linear and atmospheric perspective cues within the scene, and quantitatively estimates the overall depth range of the scene based on these cues. Linear perspective cues include geometric features such as the convergence trend of parallel lines, vanishing point positions, object size gradients, and texture density changes, directly reflecting spatial extension direction and distance changes. Atmospheric perspective cues include optical features such as brightness gradations, contrast attenuation, color shifts, and sharpness changes, used to determine the depth distribution of distant areas. By comprehensively analyzing linear and atmospheric perspective features, the maximum depth of the scene, the relative depth values ​​of each region, and the spatial scale relationship can be accurately calculated. This provides precise data support for the division of distance intervals for each depth layer behind the zero plane and the calculation of negative parallax values, ensuring that the sense of depth in the interior scene is realistic, natural, and quantifiable.

[0040] In this invention, the scene semantic features extracted by the neural network model include three components: scene type features, spatial hierarchy features, and perspective depth features. Scene type features are used to classify the overall scene of the 2D source image, identifying whether it belongs to an indoor scene, outdoor scene, product display scene, or art display scene, thereby determining the interior depth control strategy and safety constraint rules that match the application scene. Spatial hierarchy features are used to perform structural analysis of visual objects within the scene, analyzing the spatial layout, occlusion, and hierarchical relationships of foreground, midground, and background objects, providing a spatial structural basis for multi-layered interior depth planning. Perspective depth features are used to extract linear perspective cues and atmospheric perspective cues from the image, quantitatively estimating the overall depth range of the scene through geometric convergence and optical attenuation laws, providing a precise numerical basis for depth interval division and negative parallax calculation. These three features work together to form a complete scene semantic description system, providing a comprehensive, accurate, and reliable analytical basis for predicting the multi-layered depth distribution of the scene behind the zero plane of display.

[0041] In one embodiment, the depth distribution of the corresponding scene behind the display zero plane is predicted based on scene semantic features, and interior depth control parameters are generated, including: Based on the semantic features of the scene, determine the number of interior depth layers N, where N≥2; Assign a depth range [d_min, d_max] to each layer, where d represents the virtual depth distance behind the display zero plane; Based on the expected spatial location of each semantic region in the scene of the two-dimensional source image, it is mapped to the corresponding depth layer.

[0042] In this embodiment, the scene type features, spatial hierarchy features, and perspective depth features extracted in step S1 are used as a comprehensive judgment criterion. Combined with the scene's spatial complexity, the number of object layers, and depth extension characteristics, the total number of inner depth layers N required behind the zero plane is automatically determined. The number of layers N is not less than 2 layers to ensure a clear and distinguishable three-dimensional layered effect. The scene type determines the basic strategy for inner depth layering; the spatial hierarchy features directly reflect the inherent number of foreground, midground, and background layers in the scene; and the perspective depth features are used to determine whether the scene has sufficient depth to support a multi-layered distribution. For vehicle-mounted display scenes with simple structures and high safety requirements, a basic layering of 2 layers can be set. For indoor display and cultural relic display scenes with complex spatial structures and rich layers, a layering of 3 or more layers can be set, thus ensuring a high degree of matching between the number of layers and the scene's spatial structure.

[0043] After determining the number of inner depth layers N, based on the overall depth scale reflected by the scene's semantic features, the safe range of inner depth for comfortable human viewing, and the depth adjustment range supported by the lenticular lens grating optics, a corresponding virtual depth interval [d_min, d_max] is assigned to each inner depth layer. Here, d is the virtual depth distance extending inwards along the normal direction from the display zero plane (i.e., the light-emitting surface of the display panel) as the reference, d_min is the minimum virtual depth value of the current layer, and d_max is the maximum virtual depth value of the current layer. Following a spatial progression, each depth layer corresponds to near-field inner depth, mid-field inner depth, and far-field inner depth in ascending order. Each layer interval is continuous and non-overlapping, collectively forming a complete and continuous inner depth space, ensuring that the imaging positions of different layers form an orderly, reasonable, and visually pleasing depth arrangement behind the screen.

[0044] After allocating the regions for each depth layer, based on the inherent spatial relationships, visual hierarchy, and occlusion relationships of each semantic region in the 2D source image within the real scene, and combined with the foreground, midground, and background attributes determined by spatial hierarchy features, each semantic region is mapped one-to-one to the pre-divided depth layers. Specifically, foreground objects, which are the main visual subjects in the scene, are mapped to shallow depth layers closer to the display zero plane; midground objects, which serve as spatial transitions, are mapped to intermediate depth layers; and background objects or distant spatial elements, which serve as environmental backdrops, are mapped to deep depth layers farther from the display zero plane. This mapping rule establishes a precise, stable, and logically consistent correspondence between planar semantic regions in the 2D image and the 3D depth layers behind the display zero plane, providing a clear hierarchical basis for subsequent negative parallax allocation and pixel interleaving.

[0045] After determining the number of interior depth layers, allocating the depth range of each layer, and mapping the depth layers of each semantic region, the information such as the number of layers, depth ranges, and region-depth layer correspondences is integrated to generate standardized interior depth control parameters. These parameters serve as a unified control basis for subsequent multi-layer parallax allocation, selection of lenticular lens microstructure regions, and pixel interleaving processing. They can completely and accurately guide the entire process of interior depth stereoscopic imaging, ensuring that the final display effect is highly consistent with scene semantics, spatial structure, and human visual habits.

[0046] In one embodiment, the negative parallax value is positively correlated with the virtual depth distance; the upper limit of the negative parallax value is determined based on the range of interior depth for comfortable viewing by the human eye; the negative parallax difference between adjacent depth layers is determined based on the interlayer depth interval and the minimum resolvable depth difference of the human eye.

[0047] In this embodiment, the negative parallax value is positively correlated with the virtual depth distance. This means that in the interior depth imaging system behind the display null plane, the larger the virtual depth distance, the larger the configured negative parallax value; conversely, the smaller the virtual depth distance, the smaller the configured negative parallax value. The virtual depth distance represents the distance the image imaging position extends into the screen relative to the display null plane, while the negative parallax value is the core parameter controlling the lenticular grating beam splitting imaging. This rule ensures that the deeper an object is "recessed" inside the screen, the larger its corresponding negative parallax value, the stronger the parallax signal received by the viewer's eyes, and the clearer the perceived interior depth. This ensures that depth perception remains consistent with the virtual spatial distance, conforming to the natural visual laws of human observation of real three-dimensional space.

[0048] The upper limit of negative parallax must be strictly limited according to the safe range for comfortable viewing of the inner depth of field by the human eye. This is to avoid visual fatigue, dizziness, ghosting, or loss of stereoscopic perception due to excessive negative parallax. Since the inner depth image is located behind the zero plane of the display, the human visual system has a fixed physiological threshold for tolerance to negative parallax. Exceeding this threshold will impair binocular fusion, causing viewing discomfort. Therefore, this invention, by pre-setting a maximum safe depth range for comfortable viewing of the inner depth of field by the human eye, constrains the negative parallax value within the corresponding upper limit of this range. This ensures that, under prolonged viewing conditions, the inner depth stereoscopic display remains comfortable, stable, and without discomfort, meeting the usage requirements of long-term viewing scenarios such as museum displays, in-vehicle displays, and indoor previews.

[0049] The negative parallax difference between adjacent depth layers needs to be precisely calculated based on two indicators: the interlayer depth interval and the minimum resolvable depth difference of the human eye. This ensures smooth transitions between depth layers, clear distinctions, and the absence of abrupt changes or crosstalk. The interlayer depth interval is the virtual distance difference between two consecutive depth layers behind the zero plane of the display, determining the spatial span between the two layers. The minimum resolvable depth difference is the minimum depth change threshold that the human visual system can perceive; below this value, layers cannot be distinguished, while above it, layer breaks are likely to occur. By linking the negative parallax difference with these two indicators, the depth changes between adjacent depth layers can meet the requirement of clear distinction by the human eye while maintaining a natural and continuous transition effect. This results in a delicate, uniform, and realistic sense of spatial depth across multiple layers of interior scene.

[0050] In this embodiment of the invention, the allocation and calculation of negative parallax values ​​follow three strict rules: First, the negative parallax value is positively correlated with the virtual depth distance; the greater the virtual depth, the greater the negative parallax value, to ensure that depth perception is consistent with the virtual spatial distance. Second, the upper limit of the negative parallax value is determined based on the safe range of the human eye's comfortable viewing of the interior depth, constraining the negative parallax within the physiological tolerance threshold to avoid visual fatigue from prolonged viewing. Third, the negative parallax difference between adjacent depth layers is determined jointly based on the interlayer depth interval and the minimum resolvable depth difference of the human eye, making the depth layers both clearly distinguishable and smoothly transitioned, thereby achieving accurate, comfortable, and stable multi-layer interior depth stereoscopic display.

[0051] In one embodiment, the lenticular lens plate is attached to the display zero plane, comprising: The first microstructure region, with an optical pitch of P1, corresponds to the near-field depth control range, which is 0~50mm behind the display null plane. The second microstructure region, with an optical pitch of P2, corresponds to the field depth control range, which is 50~150mm behind the display zero plane. The third microstructure region, with an optical pitch of P3, corresponds to the far-field depth-of-field control range, which is 150~300mm behind the display zero plane; Among them, P1 < P2 < P3, and the smaller the pitch, the finer the depth resolution.

[0052] In this embodiment, the lenticular lens grating plate is closely attached to the light-emitting surface of the display panel and is the core optical component for achieving inner depth-of-field beam splitting imaging. To adapt to the display requirements of different levels and depths of inner depth, this invention divides the lenticular lens grating plate into three microstructure regions with different optical parameters and clearly defined functions. Each region corresponds to the inner depth control of different distance ranges behind the zero plane of the display, and each region adopts a different optical pitch design to achieve layered, high-precision inner depth-of-field stereo imaging.

[0053] The specific structure and functions are as follows: The optical pitch of the first microstructure region is P1, which specifically corresponds to the near-field depth control range, i.e., the virtual depth range of 0mm to 50mm behind the display null plane. This region is mainly used to present foreground subjects that are close to the screen surface, have prominent details, and require fine resolution. Due to its short depth distance and small spatial span, the optical system is required to have higher depth resolution accuracy. Therefore, a small pitch structure is used to achieve delicate and stable near-field depth representation.

[0054] The optical pitch of the second microstructure region is P2, specifically corresponding to the mid-field depth control range, i.e., the virtual depth range where the imaging position is located 50mm to 150mm behind the zero plane of the display. This region is mainly used to present mid-field objects in the scene that are in the middle layer and play a role in spatial transition. With a moderate depth distance and a large spatial span, a medium optical pitch can balance depth control accuracy and depth coverage, ensuring clear mid-field layers and natural transitions.

[0055] The optical pitch of the third microstructure region is P3, specifically corresponding to the far-field depth control range, i.e., the virtual depth range where the imaging position is located 150mm to 300mm behind the display null plane. This region is used to present background or distant content in the scene that is far from the screen surface and is used to construct the overall spatial atmosphere. It has the greatest depth distance and the largest spatial span, and has lower requirements for fine resolution. Therefore, a larger pitch structure is used to achieve complete coverage of a large range of depth.

[0056] The optical pitches of the three microstructure regions satisfy the relationship P1 < P2 < P3, meaning the pitch is smallest in the near field, followed by the mid field, and largest in the far field. The smaller the optical pitch, the finer the corresponding depth resolution: smaller pitches allow for denser and more precise control of light deflection, enabling the resolution of smaller depth variations, making them suitable for high-precision near-field depth representation; larger pitches cover a wider range of light deflection, covering a greater depth space, but with relatively lower depth resolution, making them suitable for large-scale far-field depth representation.

[0057] Through the division of labor and cooperation among the three microstructure regions mentioned above, the cylindrical grating plate can simultaneously support precise control of the depth of field in the near, middle, and far ranges, enabling content at different depth levels to present a clear, comfortable, and layered three-dimensional effect of the depth of field under the corresponding optimal optical structure.

[0058] In one embodiment, it further includes: Based on the distribution ratio of each depth layer in the inner depth control parameters, the target microstructure region is determined before pixel interleaving; Align the switching action of the target microstructure region with the display timing of the naked-eye 3D video frame.

[0059] In this embodiment, the multi-region switching control of the lenticular lens grating also includes the following refined processing steps to ensure stable inner depth display effect, accurate optical matching, and no misalignment, flicker, or crosstalk in the image output.

[0060] First, based on the depth-of-field distribution characteristics of the current scene, the optimal target microstructure region is precisely selected from multiple microstructure regions of the lenticular lens grating to ensure a high degree of matching between the depth-of-field imaging effect and the optical structure. Specifically, the distribution information of each depth layer recorded in the depth-of-field control parameters is read first. The area ratio, pixel ratio, and visual weight of the near-field depth layer, mid-field depth layer, and far-field depth layer in the current two-dimensional source image are statistically analyzed, serving as the core basis for region selection. If the near-field depth layer dominates, the first microstructure region with higher near-field depth-of-field and depth resolution is selected first. If the mid-field depth layer has the highest proportion, the second microstructure region adapted to the mid-field depth-of-field is selected as the main working region. If the far-field depth layer is dominant, the third microstructure region corresponding to the far-field depth-of-field is selected. The determination of the target microstructure region needs to be completed before pixel interleaving processing is performed on the multi-view image sequence. This ensures that the period, phase, mask parameters, etc. of the subsequent pixel interleaving can be completely matched with the optical pitch of the target microstructure region, avoiding problems such as incompatible imaging parameters, depth distortion, or optical crosstalk caused by the lag in region selection.

[0061] Furthermore, after identifying the target microstructure region, to ensure that the switching of the raster region is synchronized with the image display without misalignment, flicker, or gaps, strict timing synchronization control must be implemented for the switching action. Specifically, a synchronization signal that is completely consistent with the refresh rate and frame output rhythm of the display panel is generated by the timing synchronization module. This locks the electronically controlled switching action of the target microstructure region within the frame blanking period or frame switching interval of the naked-eye 3D video frame, ensuring that the completion time of the raster region switching is strictly aligned with the start time of the display of the new 3D image. Through this timing alignment process, it can be ensured that the viewer cannot visually perceive the switching action of the raster region, avoiding phenomena such as image tearing, ghosting, depth jumps, uneven brightness, or instantaneous flicker. This ensures that the display process of multiple inner depths remains continuous, stable, and smooth, meeting the needs of long-term, highly immersive viewing.

[0062] In one embodiment, after assigning differentiated negative disparity values ​​to different semantic regions of the two-dimensional source image, the method further includes: Detect whether the negative disparity value exceeds a preset safety threshold; When the negative disparity value of one of the semantic regions exceeds the safety threshold, the corresponding semantic region is automatically compressed to the nearest safe depth layer, and a visual cue is generated.

[0063] In this embodiment, after assigning differentiated negative disparity values ​​to different semantic regions of the 2D source image, an interior depth safety constraint and adaptive correction step are further set to ensure that the stereoscopic imaging is always within a range that is comfortable, safe, and stably fused to the human eye, avoiding visual fatigue, dizziness, ghosting, or depth distortion caused by excessive negative disparity. This safety protection process specifically includes the following two closely linked processing steps: First, a full-area, full-coverage safety check is performed on the inner depth imaging. The core objective is to ensure that the negative parallax values ​​corresponding to all semantic regions are within a physiologically safe range that is comfortable for the human eye. Using the inner depth safety constraint module as the execution unit, the negative parallax values ​​obtained after allocation for each semantic region are retrieved and compared item by item with a pre-set negative parallax safety threshold. This safety threshold is determined comprehensively based on the physiological characteristics of human vision, the limitations of the cylindrical grating optical structure, and the requirements for long-term viewing comfort; it is a critical value that ensures that the inner depth stereoscopic display does not cause visual discomfort. By detecting the negative parallax values ​​of each semantic region one by one, abnormal semantic regions exceeding the safety range can be accurately identified, providing an accurate basis for subsequent automatic correction processing.

[0064] Furthermore, when the interior depth safety constraint module detects that the negative disparity value of any semantic region exceeds the preset safety threshold, it immediately activates the automatic safety correction mechanism. Without disrupting the overall scene spatial structure and hierarchical relationships, it performs depth compression processing on the over-limit region. Specifically, based on the original depth position of the over-limit semantic region, it searches and matches the legal safe depth layer with the closest spatial attributes and shortest depth distance among multiple safe depth layers behind the display zero plane. The imaging depth of the semantic region is then forcibly constrained to this safe depth layer, while simultaneously updating its corresponding negative disparity value to within the safety threshold range. During the automatic depth compression correction, corresponding visual prompts are generated simultaneously. These prompts identify the location, degree of over-limit, and correction result of the over-limit region, facilitating system debugging, parameter calibration, and display status monitoring, ensuring the entire interior depth display process is safe, controllable, and traceable. Through the above safety detection and adaptive correction processing, the interior depth imaging of all semantic regions meets visual safety requirements, significantly improving the comfort and stability of long-term viewing.

[0065] In one embodiment, the source scenarios of the single two-dimensional source image include: digital display images of museum artifacts, interior renderings of real estate projects, screenshots of vehicle information systems, product display images of commercial retail terminals, three-dimensional scene projection images of digital twin systems, interactive pet scenes, and animation content displays.

[0066] In one specific embodiment, it is applied to the display of interior depth of museum artifacts.

[0067] Application scenario: Naked-eye 3D displays are installed inside museum showcases to display stereoscopic images of cultural relics.

[0068] Input image: Photograph of a bronze artifact (single 2D image).

[0069] The specific processing flow is as follows: S1: AI interior depth analysis; Scene type identification: Art exhibition scene → Interior depth N=3; Spatial hierarchy analysis: main body of the cultural relic (foreground), inscription area (middle ground), base / background (distant view); Perspective depth estimation: The scene depth is approximately 200mm.

[0070] S2: Multi-layer depth distribution; First layer (near field): Surface details of the main body of the cultural relic, depth range 0~30mm, negative parallax d1=0.8mm; Second layer (midfield): Inscription recessed area, depth range 30~80mm, negative parallax d2=2.1mm; Third layer (far field): base relief, depth range 80~150mm, negative parallax d3=3.5mm.

[0071] S3: Microstructure region selection; The three-layer depth distribution ratio is as follows: near field 40%, mid field 35%, far field 25% → Select the second microstructure region (P2, mid field depth of field) as the main activation region; Near-field and far-field are finely adjusted and adapted through pixel-level interleaving parameters.

[0072] S4: Pixel interlacing and display; The main surface of the cultural relic: intricately interwoven, presenting an inner depth that can be felt within reach; Inscription area: medium interweaving, presenting an inner depth recessed into the surface of the vessel; The base relief is intricately woven, creating a sense of depth and spaciousness.

[0073] Technical effect: Viewers feel that the cultural relics are "placed" inside the display case, rather than jumping out of the screen and breaking the boundaries of the display case, which improves the immersion by 60% and the comfort rating for long-term viewing is 4.7 / 5.

[0074] In one specific embodiment, it is applied to the display of interior depth in real estate interior renderings.

[0075] Application scenario: Naked-eye 3D display screen in sales offices, generating a three-dimensional spatial preview from a two-dimensional rendering.

[0076] Input image: Living room rendering (single 2D image).

[0077] Processing flow: S1: AI interior depth analysis; Scene type recognition: Indoor scene → Interior depth N=4; Spatial hierarchy: foreground furniture (sofa), midground furniture (coffee table), background space (windows / balcony), ceiling.

[0078] S2: Multi-layer depth distribution; First layer: Sofa surface, depth 0~40mm, d1=1.0mm; Second layer: coffee table area, depth 40~90mm, d2=2.3mm; Third layer: Window area, depth 90~180mm, d3=4.2mm; Fourth layer: Ceiling area, depth 180~250mm, d4=5.8mm.

[0079] S3: Microstructure region selection; Four layers are evenly distributed → Select the first microstructure region (P1, near-field depth of field, high resolution); Smooth transitions between layers are achieved through inter-frame temporal interpolation.

[0080] S4: Display effect; The sofa is located behind the screen surface and can sense the movement of the seat cushion; The coffee table is placed in front of the sofa, creating a clear sense of spatial hierarchy; The windows "extend" into the distance, creating a strong sense of perspective; The ceiling "surges" overhead, creating a realistic sense of spatial scale.

[0081] Technical benefits: Customers can intuitively perceive the true spatial dimensions of the room, improving decision-making efficiency by 45%.

[0082] In one specific embodiment, it is applied to depth display within an in-vehicle information system, specifically a naked-eye 3D display screen on a car dashboard, to display navigation information.

[0083] Input image: Screenshot of the navigation interface.

[0084] Processing flow: S1: AI interior depth analysis; Scene type recognition: In-vehicle information system → Interior depth N=2 (safety constraint); Spatial hierarchy: navigation arrow (foreground layer), map background (background layer).

[0085] S2: Depth distribution; Navigation arrow layer: depth 0~20mm, d1=0.5mm (slight interior depth, does not interfere with driving); Map background layer: depth 20~50mm, d2=1.3mm (embedded inside the dashboard).

[0086] S3: Safety constraints; Interior depth safety constraint module detection: All negative parallax values ​​are below the safety threshold (2.0mm). Display is allowed after verification.

[0087] S4: Display effect; The navigation arrows are slightly "floating" on the map surface, protruding but not jumping out; The map is "embedded" inside the dashboard, without obstructing the view of the road ahead; The overall information display is clear, and the time it takes for the driver to switch eyes is reduced by 30%.

[0088] In one embodiment, the pixel interleaving includes: An independent pixel interlacing mask is generated for each depth layer, and the period and phase of the interlacing mask are determined based on the negative parallax value of the corresponding layer and the optical pitch of the target microstructure region. Overlay and blend all the interlaced masks of the depth layers to generate a multi-layered joint interlaced mask of inner depth of field. The pixel arrangement of the multi-view image sequence is performed using the joint interleaving mask.

[0089] In this embodiment, firstly, for each independent depth layer that has been divided behind the zero plane, a unique pixel interlacing mask is generated. This pixel interlacing mask serves as template parameters to control the pixel arrangement, output order, and beam splitting direction of multi-view images. Its core function is to ensure that the image content of the corresponding depth layer can be accurately imaged at a preset virtual depth position under the beam splitting effect of the lenticular lens grating. During the generation process, the period and phase parameters of the interlacing mask corresponding to each depth layer are jointly determined by the negative parallax value allocated to that depth layer and the optical pitch of the target microstructure region of the lenticular lens grating: the negative parallax value determines the pixel offset and stereoscopic imaging intensity of that depth layer, and the optical pitch of the target microstructure region determines the beam splitting period and light deflection pattern, ensuring that each interlacing mask can match the inner depth display requirements of the corresponding depth layer, guaranteeing clear, crosstalk-free, and accurately positioned single-layer depth imaging.

[0090] After generating the independent pixel interlacing masks for each depth layer, all the independent interlacing masks corresponding to all depth layers are superimposed and blended according to the depth hierarchy to form a unified multi-layer inner depth joint interlacing mask that can simultaneously carry out multi-layer inner depth display control functions. The superimposition and blending process follows the spatial front-back relationship, semantic region distribution, and display priority rules of each depth layer, preserving the independent control capability of each mask layer over the corresponding depth region, while eliminating possible pixel conflicts, arrangement interference, or optical crosstalk between layers. This allows the joint interlacing mask to simultaneously achieve parallel, coordinated, and stable control over multiple layers of inner depth, ensuring that image content from different depth layers can be independently imaged in the same frame without interfering with each other, thus presenting a continuous, natural, and layered overall inner depth stereoscopic effect.

[0091] The generated multi-layered depth-of-field joint interlacing mask is used as the control basis for pixel arrangement, and the multi-view image sequence required for naked-eye 3D display is rearranged row by row, column by column, and pixel by pixel. During the arrangement process, the joint interlacing mask distributes the pixel information of different viewpoints in the multi-view image sequence to the corresponding pixel positions on the display panel according to preset period, phase, and layering rules. This ensures that after passing through the target microstructure area of ​​the lenticular lens grating, light from each viewpoint can be projected onto the viewer's eyes in a preset direction, thereby forming a stereoscopic imaging space matching each depth layer behind the zero plane of the display. Through this pixel arrangement processing based on the joint interlacing mask, ordinary multi-view images can be converted into a dedicated image format that adapts to the optical characteristics of the lenticular lens grating and supports the synchronous display of multiple layers of depth-of-field. The final output is a naked-eye 3D video frame that can achieve a stable, clear, and highly comfortable depth-of-field effect.

[0092] In one embodiment, it further includes: The system acquires parameters such as current viewing distance, ambient light intensity, and human eye fusion range in real time, and dynamically and adaptively corrects the negative parallax value of each semantic region based on these parameters. Specifically, when the viewing distance decreases or the ambient brightness decreases, the negative parallax value is reduced by a preset coefficient; when the viewing distance increases or the ambient brightness increases, the negative parallax value is increased by a preset coefficient, and the corrected negative parallax value never exceeds a safe threshold. Based on the semantic features of the scene, a depth protection priority is assigned to each semantic region. Depth locking protection is performed on high-priority main regions to prevent them from being compressed to an unexpected depth layer during the security constraint process.

[0093] In this embodiment, after completing the negative parallax value allocation and safety threshold detection for each semantic region, the present invention further sets up an interior depth dynamic optimization and subject depth protection mechanism to adapt to different viewing conditions, improve long-term viewing comfort, and ensure that the stereoscopic display effect of the core subject of the scene is not destroyed by safety compression. The specific implementation process is as follows: Through external or integrated sensing units, real-time parameters of the current viewing environment and viewing status are collected, including the viewing distance obtained by the distance sensor, the ambient light intensity obtained by the light sensor module, and the human eye fusion range parameters preset according to the scene type and the visual characteristics of the audience. After acquiring the above three key parameters, the inner depth dynamic compensation unit is the core, and the negative parallax value of each semantic region is dynamically and finely adaptively corrected based on real-time environmental and physiological parameters. This ensures that the negative parallax configuration always maintains an optimal match with the current viewing conditions, avoiding problems such as decreased stereoscopic perception, visual fatigue, or fusion difficulties caused by changes in viewing distance and ambient light interference. This significantly improves the versatility and comfort of inner depth display in different usage scenarios.

[0094] During the dynamic adaptive correction process, a strict viewing condition-negative parallax linkage adjustment rule is followed: When the viewing distance decreases, the human eye's sensitivity to stereoscopic parallax increases, and excessive negative parallax can easily cause dizziness and discomfort. Therefore, the negative parallax value is reduced according to a preset compensation coefficient. When the ambient brightness decreases, the image contrast and outline sharpness decrease, and excessive stereoscopic parallax will exacerbate the fusion burden. Therefore, the negative parallax value is also reduced according to a preset coefficient. Conversely, when the viewing distance increases, the human eye's need for depth perception increases, and the negative parallax value can be appropriately increased according to a preset coefficient to ensure stereoscopic effect. When the ambient brightness increases, visual comfort and anti-interference ability increase, and the negative parallax value can be appropriately increased according to a preset coefficient to enhance the sense of layering. Throughout the entire dynamic adjustment process, the corrected negative parallax value is always constrained within a preset safety threshold, ensuring that any adjustment result does not exceed the comfortable viewing range of the human eye, fundamentally avoiding problems such as visual fatigue, dizziness, and ghosting.

[0095] Based on the extracted scene semantic features and spatial hierarchy features, a depth protection priority division is performed on each semantic region in the 2D source image: the core display objects, key interactive elements, and main information carriers (such as cultural relics, navigation arrows, core indoor furniture, and product subjects) are set as high-priority regions; background, decorations, and secondary environmental elements are set as low-priority regions. For high-priority subject semantic regions, the system implements a depth-locking protection mechanism, that is, in the subsequent processes of interior depth safety constraints, over-limit compression, and dynamic adjustment, the preset depth value and negative parallax value are kept unchanged and not compressed to other depth layers unintentionally, and only the low-priority background region is subject to adaptive depth adjustment. Through this subject depth-locking strategy, under the premise of satisfying visual safety, the interior depth and spatial expression of the core subject are preserved to the maximum extent, so that the overall picture hierarchy is reasonable, the key points are highlighted, and the display effect is stable and controllable.

[0096] In one embodiment, the pixel interleaving further includes: Based on the semantic complexity and detail richness of each depth layer, the precision and sampling density of the corresponding pixel interleaving mask are dynamically adjusted. High-precision dense sampling masks are used for high-detail semantic regions, and low-precision sparse sampling masks are used for low-detail background regions. Transition interlaced submasks are inserted between adjacent depth layers to make the pixel arrangement period and phase of adjacent depth layers change smoothly and gradually, eliminating interlayer depth jumps and optical crosstalk, and reducing moiré interference.

[0097] In this embodiment, to further improve the clarity, sense of layering, and visual comfort of the inner depth naked-eye 3D display, and to solve industry pain points such as detail loss, interlayer crosstalk, and moiré interference in the traditional fixed-mode pixel interleaving, this invention adds two refined optimization mechanisms in the pixel interleaving process. The specific implementation process is as follows: First, based on the scene semantic features extracted in step S1, the semantic complexity and detail richness of the semantic regions carried by each depth layer behind the zero plane are quantitatively evaluated. Semantic complexity is evaluated using the number of object contours, texture density, and edge detail richness within the region as core indicators, while detail richness is quantified by pixel grayscale gradient changes, color transition smoothness, and feature point distribution density. For different evaluation results, the system dynamically adjusts the precision and sampling density of the pixel interleaving mask corresponding to each depth layer, achieving fine-grained pixel control on demand.

[0098] Specifically, for depth layers containing high-detail semantic regions such as artifact textures, product details, and human faces, a high-precision dense-sampling interlaced mask is used. This mask maximizes the preservation of edge details, texture features, and color transitions of semantic regions by increasing pixel sampling frequency, reducing sampling intervals, and optimizing phase calibration accuracy, avoiding detail blurring or contour distortion caused by insufficient sampling. For depth layers containing only low-detail background regions such as solid-color backgrounds, distant views, and simple decorative elements, a low-precision sparse-sampling interlaced mask is used. While ensuring the basic spatial hierarchy, the sampling frequency is appropriately reduced and the sampling interval is expanded to reduce unnecessary computation and improve processing efficiency, while avoiding redundant pixel interference caused by oversampling.

[0099] This dynamically adaptable mask generation strategy ensures both the display accuracy of the core semantic region and the overall processing efficiency, achieving an optimal balance between accuracy and efficiency, and significantly improving the detail reproduction and visual quality of the inner depth image.

[0100] To address issues such as depth jumps, optical crosstalk, and moiré patterns caused by abrupt changes in arrangement parameters between adjacent depth layers in traditional multi-layer pixel interleaving, this invention, after generating independent interleaving masks for each depth layer, further incorporates an interlayer transition optimization mechanism. It identifies the boundary positions of all adjacent depth layers and inserts a transitional interleaving sub-mask between the interleaving masks of adjacent layers. The pixel arrangement period and phase parameters of this sub-mask are not fixed values ​​but are linearly gradient-designed based on the core parameters of the upper and lower masks. The period / phase values ​​gradually transition from the upper mask to the lower mask, forming a continuous and smooth parameter transition band.

[0101] During pixel arrangement, the transitional interlaced sub-mask seamlessly connects with the master masks of the upper and lower layers, enabling a gradual change in the projection direction and beam splitting angle of pixels in adjacent depth layers, rather than an abrupt switch. This design fundamentally eliminates the visual discontinuity caused by abrupt changes in inter-layer depth, suppresses optical crosstalk between light from different depth layers, and breaks the fixed interference conditions caused by moiré patterns through parameter gradients, significantly reducing the interference of moiré patterns on the display effect. Ultimately, it achieves a seamless and smooth transition between depth layers, enhancing the overall coherence and comfort of the image.

[0102] In the above embodiments, this application incorporates some existing algorithms and technical features for explanation and description to make the specification more detailed, clear, and complete, thus complying with the provisions of the Patent Law. However, this is not achieved by using a series of complex steps and algorithmic formulas, nor by complicating the technical solution, nor by combining or stacking conventional or simple features. The existing algorithms and technical features listed are for the purpose of disclosing the specific implementation methods of each step of this application (not to limit this application) and to avoid situations where this application cannot be implemented.

[0103] Reference Figure 3 In another embodiment of the present invention, a naked-eye 3D display system with controllable depth of field using lenticular lenses is also provided, comprising: The analysis unit is used to receive a single two-dimensional source image, analyze the scene semantic features of the two-dimensional source image through a pre-trained neural network model, predict the multi-layer depth distribution of the corresponding scene behind the display zero plane based on the scene semantic features, and generate interior depth control parameters. The allocation unit is used to allocate differentiated negative disparity values ​​to different semantic regions of the two-dimensional source image according to the inner depth control parameters, so that each semantic region is imaged at different depth layers behind the display null plane. The determining unit is used to determine the target microstructure region from multiple microstructure regions of the cylindrical lens grating plate based on the negative parallax value; wherein each microstructure region has a different optical pitch and corresponds to a different inner depth of field control range. The generation unit is used to perform pixel interleaving on the multi-view image sequence based on the negative parallax value and the target microstructure region to generate naked-eye 3D video frames adapted for interior depth display.

[0104] In this embodiment, the specific implementation of each unit in the above system embodiment is described in the above method embodiment, and will not be repeated here.

[0105] In summary, the naked-eye 3D display method and system with controllable inner depth of lenticular lens grating provided in this embodiment of the invention includes: receiving a single two-dimensional source image; analyzing the scene semantic features of the two-dimensional source image through a pre-trained neural network model; predicting the multi-layer depth distribution of the corresponding scene behind the display null plane based on the scene semantic features; generating inner depth control parameters; assigning differentiated negative parallax values ​​to different semantic regions of the two-dimensional source image according to the inner depth control parameters, so that each semantic region is imaged at different depth layers behind the display null plane; determining a target microstructure region from multiple microstructure regions of the lenticular lens grating plate according to the negative parallax values; wherein each microstructure region has a different optical pitch, corresponding to different inner depth control ranges; and performing pixel interleaving on a multi-view image sequence based on the negative parallax values ​​and the target microstructure region to generate naked-eye 3D video frames adapted for inner depth display. In this invention, a pre-trained neural network model analyzes the scene semantic features of the two-dimensional source image to predict the multi-layer depth distribution of the corresponding scene behind the display zero plane. Differential negative parallax values ​​are then assigned. Based on these negative parallax values, target microstructure regions are determined from multiple microstructure regions of the lenticular lens plate. Pixel interleaving is performed on the multi-view image sequence to generate naked-eye 3D video frames adapted for inner depth display. This achieves the automatic calculation of multi-layer inner depth distribution through AI analysis of the scene semantic features of the two-dimensional source image, and enables layered stereoscopic display of depth behind the display zero plane through pixel-level parallax control.

[0106] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the present invention and embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0107] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0108] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for naked-eye 3D display with controllable depth of field using a cylindrical grating, characterized in that, Includes the following steps: Receive a single two-dimensional source image, analyze the scene semantic features of the two-dimensional source image through a pre-trained neural network model, predict the multi-layer depth distribution of the corresponding scene behind the display zero plane based on the scene semantic features, and generate interior depth control parameters. Based on the inner depth control parameters, differentiated negative disparity values ​​are assigned to different semantic regions of the two-dimensional source image so that each semantic region is imaged at a different depth layer behind the display null plane. Based on the negative parallax value, the target microstructure region is determined from multiple microstructure regions of the cylindrical lens grating plate; wherein each microstructure region has a different optical pitch, corresponding to a different inner depth of field control range; Based on the negative parallax value and the target microstructure region, pixel interleaving is performed on the multi-view image sequence to generate naked-eye 3D video frames adapted for interior depth display.

2. The naked-eye 3D display method with controllable inner depth of field using a cylindrical grating as described in claim 1, characterized in that, The scene semantic features include: Scene type features: used to identify whether the two-dimensional source image belongs to an indoor scene, an outdoor scene, a product display scene, or an art display scene; Spatial hierarchy features: used to analyze the spatial layout relationships between foreground, midground, and background objects in a scene; Perspective depth features: used to extract linear perspective cues and atmospheric perspective cues in a scene and estimate the depth of the scene.

3. The naked-eye 3D display method with controllable inner depth of field using a cylindrical grating as described in claim 1, characterized in that, Based on scene semantic features, predict the multi-layer depth distribution of the corresponding scene behind the display zero plane, and generate interior depth control parameters, including: Based on the semantic features of the scene, determine the number of interior depth layers N, where N≥2; Assign a depth range [d_min, d_max] to each layer, where d represents the virtual depth distance behind the display zero plane; Based on the expected spatial location of each semantic region in the scene of the two-dimensional source image, it is mapped to the corresponding depth layer.

4. The naked-eye 3D display method with controllable inner depth of field using a cylindrical grating according to claim 3, characterized in that, The negative parallax value is positively correlated with the virtual depth distance; the upper limit of the negative parallax value is determined based on the range of interior depth for comfortable viewing by the human eye; the negative parallax difference between adjacent depth layers is determined based on the interlayer depth interval and the minimum resolvable depth difference of the human eye.

5. The naked-eye 3D display method with controllable inner depth of field using a cylindrical grating as described in claim 1, characterized in that, The lenticular lens grating plate is attached to the zero plane of the display, including: The first microstructure region, with an optical pitch of P1, corresponds to the near-field depth control range, which is 0~50mm behind the display null plane. The second microstructure region, with an optical pitch of P2, corresponds to the field depth control range, which is 50~150mm behind the display zero plane. The third microstructure region, with an optical pitch of P3, corresponds to the far-field depth-of-field control range, which is 150~300mm behind the display zero plane; Among them, P1 < P2 < P3, and the smaller the pitch, the finer the depth resolution.

6. The naked-eye 3D display method with controllable inner depth of field using a cylindrical grating according to claim 1, characterized in that, Also includes: Based on the distribution ratio of each depth layer in the inner depth control parameters, the target microstructure region is determined before pixel interleaving; Align the switching action of the target microstructure region with the display timing of the naked-eye 3D video frame.

7. The naked-eye 3D display method with controllable inner depth of field using a cylindrical grating according to claim 1, characterized in that, After assigning differentiated negative disparity values ​​to different semantic regions of the two-dimensional source image, the method further includes: Detect whether the negative disparity value exceeds a preset safety threshold; When the negative disparity value of one of the semantic regions exceeds the safety threshold, the corresponding semantic region is automatically compressed to the nearest safe depth layer, and a visual cue is generated.

8. The naked-eye 3D display method with controllable inner depth of field using a cylindrical grating according to claim 1, characterized in that, The single two-dimensional source image sources include: digital display images of museum artifacts, interior renderings of real estate projects, screenshots of vehicle information systems, product display images of commercial retail terminals, three-dimensional scene projection images of digital twin systems, interactive pet scenes, and animation content displays.

9. The naked-eye 3D display method with controllable inner depth of field using a cylindrical grating according to claim 1, characterized in that, The pixel interleaving includes: An independent pixel interlacing mask is generated for each depth layer, and the period and phase of the interlacing mask are determined based on the negative parallax value of the corresponding layer and the optical pitch of the target microstructure region. Overlay and blend all the interlaced masks of the depth layers to generate a multi-layered joint interlaced mask of inner depth of field. The pixel arrangement of the multi-view image sequence is performed using the joint interleaving mask.

10. A naked-eye 3D display system with controllable depth of field using lenticular lenses, characterized in that, include: The analysis unit is used to receive a single two-dimensional source image, analyze the scene semantic features of the two-dimensional source image through a pre-trained neural network model, predict the multi-layer depth distribution of the corresponding scene behind the display zero plane based on the scene semantic features, and generate interior depth control parameters. The allocation unit is used to allocate differentiated negative disparity values ​​to different semantic regions of the two-dimensional source image according to the inner depth control parameters, so that each semantic region is imaged at different depth layers behind the display null plane. The determining unit is used to determine the target microstructure region from multiple microstructure regions of the cylindrical lens grating plate based on the negative parallax value; wherein each microstructure region has a different optical pitch and corresponds to a different inner depth of field control range. The generation unit is used to perform pixel interleaving on the multi-view image sequence based on the negative parallax value and the target microstructure region to generate naked-eye 3D video frames adapted for interior depth display.