Image display method for naked eye 3D large screen

By tracking audience positions in real time and managing dynamic viewing areas, combined with optical compensation and depth-of-field optimization, the display problem of naked-eye 3D large screens in multi-audience environments has been solved, achieving a high-quality and comfortable stereoscopic visual experience.

CN121771376BActive Publication Date: 2026-06-16XIXIAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIXIAN TECH CO LTD
Filing Date
2025-12-30
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing naked-eye 3D large screens struggle to achieve high-quality display in multi-viewer and dynamic environments, exhibiting problems such as viewpoint loss, image crosstalk, and visual fatigue. Furthermore, they lack real-time response and optical compensation capabilities based on viewer position and content importance.

Method used

By deploying a network of depth cameras to track the viewer's position in real time, dynamically dividing the view area and calculating rendering priorities, and combining optical pre-compensation and adaptive depth-of-field optimization, multi-view rendering and pixel-weighted fusion are achieved, and parallax is dynamically adjusted to improve the display effect.

Benefits of technology

It achieves a stable stereoscopic visual experience in multi-viewer environments, improves the efficiency of key information transmission, reduces image crosstalk and visual fatigue, and enhances display quality and comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121771376B_ABST
    Figure CN121771376B_ABST
Patent Text Reader

Abstract

The application discloses a naked eye 3D large screen image display method, and particularly relates to the technical field of intelligent display, which comprises the following steps: firstly, the positions of the eyes of multiple audiences in front of the screen are collected and tracked in real time through a depth camera network, and a unique ID is allocated to each audience to realize continuous tracking; then, the audiences are divided into multiple viewing zones based on a dynamic clustering algorithm, and the rendering priority of each viewing zone is calculated by comprehensively considering the number of audiences in each viewing zone and the importance of the screen content concerned by the audiences; then, left and right eye views with correct parallax are rendered in parallel for each activated viewing zone, and digital pre-compensation is performed by using a pre-stored optical model to suppress crosstalk and uneven brightness; finally, the comfortable parallax threshold is obtained by querying a human factors engineering database according to the distance and angle of the viewing zone of the audiences, and extreme depth of field exceeding the threshold in the scene is dynamically compressed. The application realizes multi-audience adaptive high-quality 3D display, and effectively improves the viewing experience, resource utilization efficiency and visual comfort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent display technology, and more specifically, to a method for displaying images on a naked-eye 3D large screen. Background Technology

[0002] As a cutting-edge technology that presents stereoscopic visual effects without the need for auxiliary glasses, naked-eye 3D display technology has shown broad application prospects in outdoor advertising, exhibitions, and immersive experiences in recent years. Existing naked-eye 3D large screens mostly employ visual impairment or lenticular lens technology, projecting images from different perspectives onto the viewer's left and right eyes through optical structures to create stereoscopic vision. However, these systems typically render images based on preset viewing zones, lacking dynamic response capabilities to the real-time position and number of viewers. When viewers move or multiple people view the screen simultaneously, problems such as loss of perspective, image crosstalk, and visual fatigue easily occur, severely impacting the viewing experience and display effect.

[0003] Current image rendering technologies often overlook the differences in the importance of content itself, employing an equal rendering strategy for all viewports, resulting in poor display of critical information areas. Furthermore, due to the inherent non-ideal characteristics of optical systems, such as brightness attenuation and viewpoint-dependent crosstalk, the display quality of images varies significantly across different areas of the screen. Although some research has attempted optimization through optical compensation or viewport segmentation, most methods remain at the static or semi-static processing stage, struggling to adapt to dynamically changing viewing environments and failing to achieve truly multi-viewer, multi-viewport, high-quality glasses-free 3D display.

[0004] Furthermore, traditional glasses-free 3D systems lack dynamic consideration of human visual comfort in depth control. Excessive or insufficient parallax can easily trigger accommodation-convergence conflict, leading to visual fatigue and even dizziness. Although some systems attempt to mitigate this problem by limiting the maximum parallax, they often use a fixed threshold, failing to adaptively adjust based on the viewer's actual viewing distance and angle. Therefore, there is an urgent need for a glasses-free 3D large-screen image display method capable of real-time perception of viewer distribution, dynamic division of viewing areas, intelligent allocation of rendering resources, and possessing optical compensation and depth-adaptive capabilities, to improve the overall viewing experience and visual comfort in multi-viewer environments. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method for displaying images on a naked-eye 3D large screen.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] The image display method for naked-eye 3D large screen includes the following steps:

[0008] S1. Viewpoint Acquisition and Tracking: Through a deployed network of depth cameras, the three-dimensional coordinates of each viewer's eyes in front of the screen are detected and located in real time, and a unique ID is assigned and maintained for each viewer. Continuous tracking is performed across frames using motion models and appearance features.

[0009] S2. View Zone Division and Priority Calculation: Dynamically cluster all collected viewpoints to form several view zones with similar viewing angles; calculate the final rendering priority of each view zone by combining the number of viewers in each view zone and the importance of the screen content they are interested in.

[0010] S3, Multi-view rendering and optical compensation: For each active view area, a pair of left and right eye view layers with correct parallax are rendered in parallel; before mapping the view layer pixels to the physical screen, digital pre-compensation is performed according to the pre-stored optical model to cancel the inherent crosstalk and brightness unevenness of the screen.

[0011] S4. Depth of field adaptive optimization: Real-time analysis of the depth information of the 3D scene; the average interpupillary distance of the viewer within the view area is obtained through the depth camera network, and the maximum comfortable parallax threshold is obtained by querying the 3D human factors engineering database in combination with the distance and horizontal viewing angle of the view area; extreme depth of field in the scene that exceeds the threshold is dynamically compressed by nonlinear mapping, so as to control the parallax within the visually comfortable range.

[0012] Specifically, the cross-frame continuous tracking in S1 using motion models and appearance features includes:

[0013] Maintain a motion state vector for each viewer ID and predict the expected position in the next frame based on the motion model;

[0014] Calculate the association cost between the new detected target and the existing ID prediction location. The association cost is determined by both the location distance and the similarity of the apparent features.

[0015] The matching algorithm is used to associate the target, and a recovery mechanism is established for temporarily lost IDs.

[0016] Specifically, in S2:

[0017] Dynamic clustering employs a dynamic density clustering algorithm, with the neighborhood radius set based on the interpupillary distance of an adult and adaptively adjusted according to the audience distribution density.

[0018] Specifically, in S2, when calculating the final rendering priority:

[0019] A content importance heatmap is introduced, which reflects the relative importance of content in different areas of the screen. The priority of the viewing area is determined by both the number of viewers and the importance of the content in the area of ​​focus.

[0020] Specifically, the digital pre-compensation in S3 includes:

[0021] Based on a pre-stored optical transfer function library, a pre-compensated filter is designed for each view area;

[0022] The original view layer obtained from the rendering is subjected to frequency domain filtering.

[0023] Simultaneously, brightness uniformity compensation is performed on the processed image.

[0024] Specifically, S3 also includes:

[0025] Pixel-weighted fusion is performed on screen areas where multiple viewports overlap. The fusion weight is determined by the priority of the viewport and the distance from the pixel to the center of the viewport.

[0026] Specifically, in S4:

[0027] Depth-of-field adaptive optimization: Real-time analysis of depth information in 3D scenes; obtaining the average interpupillary distance of viewers within the view area through a depth camera network, and combining the distance and horizontal viewing angle of the view area to query the 3D human factors engineering database to obtain the maximum comfortable parallax threshold;

[0028] A non-linear mapping method is used to dynamically compress extreme depth of field in the scene that exceeds the threshold, thereby controlling parallax within a visually comfortable range.

[0029] Specifically, in S4:

[0030] Dynamic compression uses a non-linear mapping method to map parallax that exceeds the comfort range to the comfort range.

[0031] The technical effects and advantages of this invention are as follows:

[0032] This invention achieves a fundamental shift from a screen-centric to a viewer-centric approach, enabling real-time tracking of the eye positions of multiple viewers and dynamically dividing the viewing areas to ensure a stable, seamless stereoscopic visual experience for each viewer. Simultaneously, by introducing a content importance heatmap and a priority calculation mechanism, limited rendering resources are intelligently allocated to viewing areas with a large number of viewers or those focusing on important content, thereby significantly improving the efficiency of conveying key information and overall display performance in multi-viewer scenarios.

[0033] Through sophisticated digital pre-compensation technology and adaptive depth-of-field optimization, the inherent technical bottlenecks of naked-eye 3D displays are effectively overcome. Optical pre-compensation effectively suppresses crosstalk and brightness unevenness caused by the optical characteristics of lenses or gratings, improving image clarity and uniformity. Comfort parallax control based on a human factors engineering database can dynamically compress the scene depth within a visually comfortable range, fundamentally reducing visual fatigue and dizziness that may be caused by prolonged viewing. Attached Figure Description

[0034] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] like Figure 1 As shown, the image display method module for naked-eye 3D large screen is as follows:

[0037] Step 1: Viewpoint Acquisition and Tracking: Using a deployed depth camera network, the 3D coordinates of each viewer's eyes are detected and located in real time. Simultaneously, a unique ID is assigned and maintained for each viewer, and continuous cross-frame tracking is performed using motion models and appearance features. Even in the event of brief occlusion, the association can be restored, ensuring the continuity of viewer identity. Specific execution process:

[0038] Multiple high-resolution, wide-field-of-view depth-sensing cameras (such as structured light or TOF cameras) are deployed above and / or below the naked-eye 3D large screen to form a stereoscopic vision monitoring network covering the entire potential viewing area in front of the screen.

[0039] These cameras synchronously acquire RGB images and depth information of the scene at a fixed frequency (e.g., 30Hz). The acquired data is sent to a central processing unit, which first separates all moving human targets using background modeling and motion detection algorithms. Then, for each human target, the pixel coordinates of its eyes are located using a facial landmark detection model (such as a deep learning-based facial landmark detection model).

[0040] By combining depth information with camera calibration parameters, the pixel coordinates of both eyes are converted into three-dimensional spatial coordinates (X, Y, Z) in the world coordinate system through coordinate transformation. Finally, a real-time data stream containing a unique ID, the three-dimensional coordinates of the left eye, and the three-dimensional coordinates of the right eye is generated for each viewer.

[0041] While generating the real-time data stream, the following continuous audience identity tracking and loss recovery process is executed:

[0042] Cross-frame tracking based on motion models: Maintaining a motion state vector for each active viewer ID. , in The world coordinates of the current frame. The instantaneous velocity of the current frame.

[0043] Prediction algorithms such as Kalman filtering are used to predict the state of the previous frame. Predict the expected state of the current frame .

[0044] After receiving all viewer detection results for a new frame, calculate the predicted position of each newly detected target relative to all tracked IDs. The cost of association between them. It is determined by both the Mahalanobis distance of location and the cosine distance of appearance features (such as the HSV color histogram based on RGB images):

[0045] ;

[0046] in, It is Mahalanobis distance, which can effectively take into account the uncertainty of motion trajectory; It is the cosine similarity of appearance features; This is a coefficient that balances the weights of the two (e.g., 0.7); the optimal matching is achieved through the Hungarian algorithm, associating the newly detected target with an existing ID and updating its motion state. For successful detections that fail to match, initialize them with a new tracking ID.

[0047] Loss recovery based on trajectory and appearance features: When the tracking of a viewer ID is lost due to temporary occlusion, the ID is not immediately deleted, but marked as temporarily lost, and the position is continued to be predicted using the motion model. At the same time, a recovery threshold region is established around the predicted position.

[0048] In subsequent frames, any newly detected target that appears within the "recovery threshold" region but is not successfully associated will have its appearance feature similarity calculated with the temporarily lost ID. If the similarity exceeds a preset threshold... If the value is 0.8 (e.g.), and its direction of movement is consistent with the historical trajectory before the ID was lost, then the ID is determined to be recovered and re-associated with this new detection target, thereby maintaining the continuity of the viewer's identity.

[0049] Set a loss timer for each ID. If an ID is lost in consecutive... If the ID cannot be recovered within 15 frames (e.g., 0.5 seconds), it is determined that the viewer has left the viewing area, and the ID and all its status data are eventually deleted.

[0050] Step 2: Viewpoint Segmentation and Priority Calculation: All collected viewpoints are dynamically clustered to form several "viewer group viewpoints" with similar viewing angles. Then, based on the number of viewers in each viewpoint and the importance of the screen content they are focusing on, the final rendering priority of each viewpoint is calculated. Specific execution process:

[0051] The central processing unit receives the real-time viewpoint data stream from step one. First, all viewers' left and right eye viewpoints are treated as independent two-dimensional points (primarily considering their projected coordinates on a plane parallel to the screen). Then, a dynamic density clustering algorithm (such as DBSCAN) is used to perform cluster analysis on these viewpoints. Two key parameters in the DBSCAN algorithm are set as follows:

[0052] Neighborhood radius (eps): Set to 0.15 meters. This is based on more than twice the typical adult interpupillary distance (approximately 6.5 cm) and takes into account a certain position tracking error. It ensures that the binocular viewpoints of the same viewer and the viewpoints of adjacent standing viewers are correctly grouped into the same cluster, while avoiding incorrect clustering of viewers who are too far away.

[0053] Minimum number of samples (min) samples ): Set to 2. A viewport must contain at least one viewer (because it has two viewpoints, left and right) to be confirmed. This setting effectively filters out isolated noise (such as false detections).

[0054] Every 10 frames of data processed, the EPS value is adaptively fine-tuned based on the spatial distribution density of the current viewpoint, with the adjustment not exceeding 20% ​​of the initial value, in order to cope with different scenarios of audience concentration and sparseness.

[0055] Each cluster represents the view area of ​​a group of viewers with similar viewing angles.

[0056] Then, a rendering priority is assigned to each viewport based on the number of viewers within it, with viewports containing more viewers having higher priority. When assigning viewport priorities, content importance is introduced as a dynamic weighting factor, as follows:

[0057] Content importance heatmap generation: The content management module or graphics rendering engine preloads or generates in real time a content importance heatmap corresponding to the screen pixel space. This image is a single-channel matrix with the same resolution as the screen, showing the value of each pixel. This defines the relative importance of content within a corresponding screen area. For example:

[0058] Product logo and core slogan area in the advertisement Set the value to 1.0; for the main character's face or key interactive areas. Set the value to 0.8 ~ 1.0; for background and secondary information areas. Set the value to 0.1 ~ 0.3;

[0059] This heatmap is based on predefined content metadata or dynamically generated through real-time image recognition algorithms (such as saliency detection);

[0060] Viewport content attention calculation: For each viewport defined in step two... Calculate its content attention score The content attention score reflects the importance of the content in the screen area that the audience in that viewpoint pays attention to;

[0061] First, according to the view area By calculating the projection coordinates of the center point on the screen and combining this with the average viewing angle of that visual area, the collective attention focus area of ​​the audience in that visual area can be estimated. (Typically, it is a circular area with the intersection of the line of sight and the screen as its center, and the radius is inversely proportional to the viewing distance.)

[0062] Then, the attention focus area is calculated. All pixels within the content importance heatmap The average of the above values ​​is used as the content attention score for that view area:

[0063] ;

[0064] in, Score based on content attention. For the region The area refers specifically to the total number of pixels contained within that area;

[0065] In practical discrete computation, this value is obtained by applying the region... All pixels within The arithmetic mean of the values ​​is obtained;

[0066] Overall priority calculation: View area Final rendering priority Due to its audience size Content attention score The calculation formula is as follows, jointly determined:

[0067] ;

[0068] in: For the view area The final rendering priority, It is the view area The number of viewers inside; It is the view area Content attention score; It is the content weight coefficient ( `content importance` is an adjustable system parameter used to control the strength of the impact of content importance on the final priority. At that time, it degenerates into prioritizing the sheer number of viewers; when At that time, there is a tendency to allocate more resources to view areas that focus on important content.

[0069] Based on this calculation All viewports are sorted and resources are allocated, and then used for pixel-weighted fusion in step three.

[0070] At the same time, the movement trend of each view area in the next moment is predicted (e.g., short-term trajectory prediction based on Kalman filtering) to ensure the continuity and stability of view area division and avoid screen flicker.

[0071] Step 3, Multi-view Rendering and Optical Compensation: For each active viewport, a pair of left and right eye view layers with correct parallax are rendered in parallel. Before mapping the view layer pixels to the physical screen, digital pre-compensation is performed based on a pre-stored optical model to counteract inherent crosstalk and brightness unevenness on the screen, thereby improving the sharpness and uniformity of the final image. Specific execution process:

[0072] The graphics rendering engine receives the viewport partitioning results and priority information from step two. For each activated viewport, a separate rendering thread is started.

[0073] Each rendering thread calculates the corresponding binocular parallax based on the average spatial coordinates of the center point of the viewport, and renders two images with subtle differences in perspective—a left-eye view and a right-eye view—in real time. The rendering process does not generate a complete screen image, but rather generates a view layer corresponding to that viewport.

[0074] Subsequently, a pixel mapping module will allocate the pixels of the left and right eye views rendered in each view zone to the corresponding physical pixel subsets on the screen, based on the optical characteristics of the naked-eye 3D screen (such as the directionality of lenses or gratings).

[0075] After pixel mapping is completed and before viewport fusion is performed, the view layer signal that is about to be sent to the physical pixels of the screen undergoes digital pre-compensation based on an optical model; the process is as follows:

[0076] Optical transfer function modeling and storage: A library of optical transfer functions related to the viewing angle, obtained through precise calibration, is pre-stored. The function library describes the optical transfer functions at a specific viewing angle. Below, the actual light field distribution formed in the viewer's eyes after an ideal pixel signal passes through the optical system.

[0077] For each view area Its optical effect simplifies to a point spread function (PSF), which is the spatial representation of the OTF, denoted as . PSF quantifies how light energy from an ideal pixel leaks into its neighboring region (i.e., crosstalk).

[0078] Pre-compensation filter design: To cancel optical crosstalk, for each viewport... View layer image Design a pre-compensated filter.

[0079] The pre-compensated filter in the frequency domain modulates the optical transfer function. This is achieved through inverse processing. To avoid noise amplification caused by direct inversion, the Wiener filter method is used, and its frequency domain expression is:

[0080] ;

[0081] in, For the pre-compensation filter in the frequency domain The value of the point, yes The complex conjugate, It is a regularization parameter (usually taken as...) (1% to 10% of the average energy spectrum) is used to stabilize the solution process and control noise.

[0082] Real-time pre-compensation processing: for each view area The original view layer obtained from rendering Before sending in physical pixels, the image is processed as follows to obtain the compensated image. :

[0083] ;

[0084] in, For pre-compensation filters, and These represent the Fourier transform and the inverse Fourier transform, respectively. In the spatial domain, this operation is equivalent to convolution with a pre-compensated kernel, which pre-blurs and inversely crosstalks the image, ensuring that the image perceived by the viewer after passing through the actual optical system is closer to the ideal image. .

[0085] Brightness uniformity compensation: Simultaneously, a pre-calibrated screen brightness attenuation map is loaded. ( This figure records the maximum brightness attenuation coefficient in different areas of the screen due to optical structure. During the pre-compensation phase, simultaneously... Perform brightness gain compensation:

[0086] ;

[0087] in, For at pixel The final image pixel value after complete optical pre-compensation (including crosstalk compensation and brightness compensation); For at pixel The pixel value of the image after crosstalk pre-compensation but before brightness compensation;

[0088] After optical pre-compensation processing This information is then used as input to the subsequent weighted fusion module, thereby ensuring that the final displayed image achieves optimal performance in terms of crosstalk suppression and brightness uniformity.

[0089] For screen areas where multiple viewports overlap, pixels from different view layers are weighted and fused according to viewport priority, with higher-priority viewports receiving greater weight to ensure optimal viewing for the primary audience. The specific calculation method for weighted fusion is as follows:

[0090] For any physical pixel P on the screen that is covered by multiple viewports, its final displayed pixel value It is determined by the following formula:

[0091] ;

[0092] in: It is the pixel value rendered by the i-th viewport for pixel point P; It is the fusion weight of the view region, determined by the view region priority coefficient. The Euclidean distance from pixel P to the center of the viewing area on the screen projection. The decision was made jointly, and the calculation formula is as follows: ; To prioritize the main audience groups based on the normalized number of viewers within the viewing area. Set it to 1.0, and reduce the rest proportionally; The distance attenuation coefficient is set to an empirical value of 0.5 to ensure a smooth transition at the viewport boundary and avoid abrupt image cropping.

[0093] Step 4: Depth-of-Field Adaptive Optimization: Real-time analysis of the 3D scene's depth information, and based on the distance and angle of the viewer's current field of view, querying the human factors engineering database to obtain the maximum comfortable parallax threshold. Dynamically compressing extreme depths of field exceeding this threshold in the scene, controlling parallax within a comfortable range, effectively reducing visual fatigue; Specific execution process:

[0094] While rendering is being performed in step three, the depth analysis module will perform real-time analysis of the depth map of the current 3D scene to calculate the average depth of field and the maximum depth of field of the scene.

[0095] A pre-built database of comfortable parallax ranges is established, linking the distance between the viewer and the screen, the viewing angle, and the maximum tolerable parallax. This database is a pre-generated two-dimensional lookup table constructed based on human factors engineering experimental data. The experiment involved participants viewing 3D images with different parallaxes at varying viewing distances (D, from 1 meter to 20 meters, in 1-meter intervals) and different horizontal viewing angles (θ, from -60 degrees to +60 degrees, in 5-degree intervals), recording the critical parallax values ​​at which they reported visual fatigue, thereby fitting the maximum comfortable parallax threshold (…). Functions relating distance and viewpoint: ;in The function is fitted based on human factors engineering experimental data, where D is the viewing distance. A horizontal perspective;

[0096] In real-time operation, based on the average distance of the view area obtained in step one... and average perspective The current view area is retrieved in real time from this two-dimensional lookup table using bilinear interpolation. For example, when the system calculates a certain view area... If the parallax of an object in the scene rendered in step three reaches 70 pixels, the depth optimization module will compress the depth information of the object and non-linearly map its parallax to a comfortable range of 0-50 pixels.

[0097] If the scene parallax calculated by the rendering engine exceeds this threshold, it dynamically performs non-linear compression on the parallax of objects that are too far or too close in the scene, controlling extreme depth of field within a comfortable range, rather than simply cropping it. This adaptive adjustment ensures that viewers will not experience accommodation-convergence conflict due to excessive parallax when viewing from any position, thus effectively reducing visual fatigue.

[0098] The average interpupillary distance (IPD) is obtained by calculating the IPD of each viewer in real time based on the three-dimensional coordinates of their eyes located in step one (formula: To simplify the calculation by ignoring the Z-axis difference, in the formula... This is the interpupillary distance (i.e., the distance between the eyes) of a single viewer currently being calculated. , These are the X-axis and Y-axis coordinates of the left eye in a two-dimensional plane (ignoring the Z-axis depth direction) (these coordinates come from the X and Y dimensions of the viewer's two eyes in the three-dimensional coordinates located in step one). , These are the X-axis and Y-axis coordinates of the right eye on the same two-dimensional plane, respectively. The average interpupillary distance of all viewers within the same visual region is taken as the average interpupillary distance of that visual region. .

[0099] 3D Database Query: The parameters of the 3D lookup table are viewing distance (1-20 meters), horizontal angle of view (-60° to +60°), and average interpupillary distance (58-72mm, in 4mm increments). The table was constructed using calibration data from a visual fatigue experiment with 100 adult subjects. Linear interpolation was used to obtain precise thresholds during the query. .

[0100] The above formulas are all dimensionless calculations. Dimensionless calculations can be performed using various methods such as standardization, which will not be elaborated here. The formulas are derived from software simulations based on a large amount of collected data, and the preset parameters in the formulas can be set by those skilled in the art according to the actual situation.

[0101] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, ATA hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state ATA hard disk.

[0102] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0103] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0104] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0106] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0107] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable ATA hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0108] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for displaying images on a naked-eye 3D large screen, characterized in that, Includes the following steps: S1. Viewpoint Acquisition and Tracking: Through a deployed network of depth cameras, the three-dimensional coordinates of each viewer's eyes in front of the screen are detected and located in real time, and a unique ID is assigned and maintained for each viewer. Continuous tracking is performed across frames using motion models and appearance features. S2. View Zone Division and Priority Calculation: Dynamically cluster all collected viewpoints to form several view zones with similar viewing angles; calculate the final rendering priority of each view zone by combining the number of viewers in each view zone and the importance of the screen content they are interested in. S3, Multi-view rendering and optical compensation: For each active view area, a pair of left and right eye view layers with correct parallax are rendered in parallel; before mapping the view layer pixels to the physical screen, digital pre-compensation is performed according to the pre-stored optical model to cancel the inherent crosstalk and brightness unevenness of the screen. The digital pre-compensation in S3 includes: Based on a pre-stored optical transfer function library, a pre-compensated filter is designed for each view area; The original view layer obtained from the rendering is subjected to frequency domain filtering. Simultaneously, brightness uniformity compensation is performed on the processed image; S3 further includes: Pixel-weighted fusion is performed on screen areas where multiple view areas overlap. The fusion weight is determined by the priority of the view area and the distance of the pixel from the center of the view area. Higher priority view areas have higher pixel weights to ensure the viewing experience for the main audience. S4. Depth of field adaptive optimization: Analyzes the depth information of the 3D scene in real time, and queries the human factors engineering database to obtain the maximum comfortable parallax threshold based on the distance and angle of the current viewer's field of view; dynamically compresses extreme depth of field in the scene that exceeds this threshold, and controls the parallax within a comfortable range.

2. The image display method for a naked-eye 3D large screen according to claim 1, characterized in that, The cross-frame continuous tracking in S1 using motion models and appearance features includes: Maintain a motion state vector for each viewer ID and predict the expected position in the next frame based on the motion model; Calculate the association cost between the new detected target and the existing ID prediction location. The association cost is determined by both the location distance and the similarity of the apparent features. The matching algorithm is used to associate the target, and a recovery mechanism is established for temporarily lost IDs.

3. The image display method for a naked-eye 3D large screen according to claim 1, characterized in that, In S2: Dynamic clustering employs a dynamic density clustering algorithm, with the neighborhood radius set based on the interpupillary distance of an adult and adaptively adjusted according to the audience distribution density.

4. The image display method for a naked-eye 3D large screen according to claim 1, characterized in that, When calculating the final rendering priority in S2: A content importance heatmap is introduced, which reflects the relative importance of content in different areas of the screen. The priority of the viewing area is determined by both the number of viewers and the importance of the content in the area of ​​focus.

5. The image display method for a naked-eye 3D large screen according to claim 1, characterized in that, In S4: Depth-of-field adaptive optimization: Real-time analysis of depth information in 3D scenes; obtaining the average interpupillary distance of viewers within the view area through a depth camera network, and combining the distance and horizontal viewing angle of the view area to query the 3D human factors engineering database to obtain the maximum comfortable parallax threshold; A non-linear mapping method is used to dynamically compress extreme depth of field in the scene that exceeds the threshold, thereby controlling parallax within a visually comfortable range.

6. The image display method for a naked-eye 3D large screen according to claim 1, characterized in that, In S4: Dynamic compression uses a non-linear mapping method to map parallax that exceeds the comfort range to the comfort range.

Citation Information

Patent Citations

  • Video display system and video display method

    JP2017188715A

  • Naked-eye 3D display method and system for 2d game

    US20240414309A1