Airborne naked-eye three-dimensional display device, method and medium based on two-dimensional video

CN122679263APending Publication Date: 2026-09-01LIUZHOU DINGYU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610790106.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0002]目前,裸眼三维显示技术主要分为两大类:一类是基于视差屏障或柱状透镜的裸眼3D屏幕,其本质是在屏幕平面上呈现左右眼视差图像,观众仍需在特定位置观看,且必须使用专门拍摄或渲染的3D片源,技术实现成本极高;另一类是空中成像技术,通过逆反射或等效负折射率透镜将影像悬浮于空中,但同样仅能显示平面或预制的三维内容,无法将普通2D视频进行实时影像悬浮显示

Benefits of technology

[0015] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122679263A_ABST
    Figure CN122679263A_ABST
Patent Text Reader

Abstract

This application relates to an aerial levitation naked-eye 3D display device, method, and medium based on 2D video. The method involves decoding 2D video to obtain video frame images, scaling and color space conversion of these images, performing depth analysis to obtain a corresponding depth map, converting the video frame images into multi-layer depth representations based on a set number of depth layers and the depth map, rendering them as left-eye and right-eye view images, and setting a parallax baseline. The left-eye and right-eye view images are then output side-by-side to an aerial imaging panel based on the parallax baseline and driven for display. The aerial imaging panel displays the left-eye and right-eye view images as an aerial levitation naked-eye 3D image using incoherent light based on configured levitation distance and brightness information. This technical solution can generate naked-eye levitation 3D images from common 2D video signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video display technology, and in particular to an aerial levitation naked-eye 3D display device based on 2D video, an aerial levitation naked-eye 3D display method based on 2D video, and a computer-readable storage medium. Background Technology

[0002] Currently, naked-eye 3D display technology is mainly divided into two categories: one is naked-eye 3D screens based on parallax barriers or lenticular lenses, which essentially present left and right eye parallax images on the screen plane. Viewers still need to watch from a specific position and must use specially shot or rendered 3D content, making the technology extremely expensive to implement; the other is aerial imaging technology, which suspends images in the air through retroreflection or equivalent negative refractive index lenses, but it can only display planar or pre-made 3D content and cannot display ordinary 2D videos in real time.

[0003] It is evident that existing naked-eye 3D display technologies are mainly based on 3D image display methods, which are costly to implement and difficult to use for floating displays of 2D videos. Summary of the Invention

[0004] To address one of the aforementioned technical deficiencies, this application provides an aerial levitation naked-eye 3D display device based on 2D video, an aerial levitation naked-eye 3D display method based on 2D video, and a computer-readable storage medium.

[0005] A naked-eye 3D display device based on 2D video, comprising: The video input and preprocessing module is used to decode the input two-dimensional video using a decoder to obtain video frame images, and to perform image scaling and color space conversion on the video frame images; A real-time depth estimation engine is used to perform image depth analysis on the video frame images to obtain the corresponding depth maps; The multi-plane image layering synthesis module is used to convert the video frame image into a multi-layer depth representation according to the set number of depth layers and the depth map, and render it into a left-eye view image and a right-eye view image, and set the disparity baseline of the left-eye view image and the right-eye view image. An aerial imaging drive module is used to output the left-eye view image and the right-eye view image side by side to the aerial imaging panel according to the parallax baseline and drive it to be displayed. The aerial imaging panel displays the left-eye and right-eye perspective images as non-coherent three-dimensional images suspended in the air using incoherent light, based on the configured suspension distance and brightness information.

[0006] In some embodiments, the aerial levitation naked-eye 3D display device based on 2D video further includes: A black-and-white video colorization module is used to input black-and-white video images and convert them into color two-dimensional images for input into the video input and preprocessing module.

[0007] In some embodiments, the black-and-white video colorization module of the aerial levitation naked-eye 3D display device based on 2D video is further used to downsample the previous frame color image and fuse the previous frame color image during the decoding stage of the black-and-white video image to obtain a colored 2D video image.

[0008] In some embodiments, the video input and preprocessing module is further configured to execute an adaptive deinterlacing algorithm to convert the colored two-dimensional video image into a progressive scan signal, and dynamically adjust the output refresh rate by detecting the field frequency of the input signal; The real-time depth estimation engine is also used in a block-matching-based dense optical flow algorithm to perform pixel-level alignment of motion between adjacent frames.

[0009] In some embodiments, the aerial levitation naked-eye 3D display device based on 2D video further includes: The tracking and adaptive optimization module is used to capture the user's face image in real time using the camera and adaptively and dynamically adjust the parallax baseline based on the position information of the face image.

[0010] In some embodiments, the aerial levitation naked-eye 3D display device based on 2D video further includes: The dynamic thermal management module is used to reduce the input resolution of the real-time depth estimation engine or turn off the black-and-white video colorization module when the temperature exceeds a set temperature threshold.

[0011] In some embodiments, the aerial levitation naked-eye 3D display device based on 2D video includes an image depth estimation engine that is used to access video frame images, perform depth estimation on the video frame images using a pre-trained lightweight convolutional neural network, and output a depth image of the same resolution; wherein the lightweight convolutional neural network includes a bilinear interpolation layer and a convolutional layer.

[0012] A method for displaying 3D levitation in mid-air without glasses based on 2D video includes the following steps: The input two-dimensional video is decoded using a decoder to obtain video frame images, and the video frame images are then scaled and converted in color space. Image depth analysis is performed on the video frame images to obtain the corresponding depth maps; The video frame image is converted into a multi-layer depth representation according to the set depth layer and the depth map, and rendered as a left-eye view image and a right-eye view image. The parallax baseline of the left-eye view image and the right-eye view image is set. The left-eye and right-eye view images are output side-by-side to the aerial imaging panel based on the parallax baseline and driven to be displayed; wherein, the aerial imaging panel displays the left-eye and right-eye view images as aerial suspended naked-eye 3D images according to the configured suspension distance and brightness information.

[0013] In some embodiments, before decoding the input two-dimensional video using a decoder to obtain video frame images, the method further includes: When the input is detected to be a black and white video image, the black and white video image is converted into a color two-dimensional image.

[0014] A computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded by a processor and executed to perform the steps of the described method for aerial levitation naked-eye 3D display based on 2D video. The technical solution of the above embodiment utilizes a decoder to decode the input two-dimensional video to obtain video frame images, and performs image scaling and color space conversion on the video frame images; performs image depth analysis on the video frame images to obtain corresponding depth maps; converts the video frame images into multi-layer depth representations according to the set number of depth layers and the depth maps, and renders them as left-eye view images and right-eye view images, and sets the disparity baseline of the left-eye view images and right-eye view images; outputs the left-eye view images and right-eye view images side by side to the aerial imaging panel according to the disparity baseline and drives it for display; the aerial imaging panel displays the left-eye view images and right-eye view images as aerial suspended naked-eye 3D images in an incoherent light manner according to the configured suspension distance and brightness information; this technical solution, by receiving two-dimensional video signals of common formats in real time, performs depth estimation and multi-view synthesis on the video frame images, drives the incoherent light aerial imaging panel, generates color, naked-eye viewable suspended 3D images, and supports extended functions such as automatic colorization of black and white films and single-view adaptive tracking, and has the characteristics of low cost, low latency, and scalability.

[0015] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description

[0016] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a schematic diagram of the structure of an aerial levitation naked-eye 3D display device based on 2D video, according to one embodiment. Figure 2This is a schematic diagram of the structure of a naked-eye 3D display device based on two-dimensional video that is suspended in mid-air according to another embodiment; Figure 3 This is a flowchart of an embodiment of a method for displaying naked-eye 3D levitation in mid-air based on 2D video; Figure 4 This is a hardware principle block diagram of an aerial levitation naked-eye 3D display device based on two-dimensional video, according to one embodiment. Figure 5 This is a software hierarchy diagram of an embodiment of the naked-eye 3D display scheme based on 2D video for aerial levitation according to this application; Figure 6 This is a data stream and delay breakdown diagram of an example of an aerial levitational naked-eye 3D display device based on 2D video; Figure 7 This is a timing diagram of the interface between the aerial imaging drive module and the aerial imaging panel in one embodiment; Figure 8 This is an example assembly structure diagram of a naked-eye 3D display device based on 2D video that is suspended in mid-air. Detailed Implementation

[0017] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0018] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the word “comprising” as used in this application’s specification means the presence of the stated feature, integer, step, or operation, but does not preclude the presence or addition of one or more other features, integers, steps, or operations.

[0019] The technical solution of this application provides an aerial levitation naked-eye 3D display device based on 2D video, which can receive 2D video signals of common formats in real time. By performing depth estimation and multi-view synthesis on video frame images, it drives an incoherent light aerial imaging panel to generate color, naked-eye viewable levitation 3D images. It also supports extended functions such as automatic colorization of black and white films and single-viewer adaptive tracking, and has the characteristics of low cost, low latency and scalability.

[0020] refer to Figure 1 As shown, Figure 1 This is a schematic diagram of the structure of an aerial levitation naked-eye 3D display device based on 2D video, according to one embodiment, including: (1) Video input and preprocessing module, which is used to decode the input two-dimensional video using a decoder to obtain video frame images, and to perform image scaling and color space conversion on the video frame images.

[0021] Specifically, the two-dimensional video is input to the decoder and decoded frame by frame to obtain video frame images. Then, the video frame images are scaled to a set size and the corresponding color space conversion is performed. Color space conversion can change the color representation of the video frame images, such as converting other video formats to RGB format.

[0022] For example, the video input and preprocessing module can use the Longxun LT8619C as the HDMI receiver chip, which supports 4K@30fps input. Video decoding can use FFmpeg (running on an ARM core) and supports H.264 / H.265 hardware decoding. During preprocessing, the input video frame image can be scaled to 1080p, and ARM NEON instructions can be used for color space conversion.

[0023] (2) Real-time depth estimation engine, used to perform image depth analysis on the video frame image to obtain the corresponding depth map.

[0024] Specifically, image depth analysis is performed on the video frame images input by the video input and output by the preprocessing module to obtain the corresponding depth map of the video frame image. Image depth estimation technology can infer the distance information of each pixel point in the scene to the camera imaging plane from a single or multiple two-dimensional images. The depth map is a two-dimensional matrix image, where each pixel value represents the distance of the corresponding scene point to the camera, which can be used in subsequent 3D reconstruction and stereo matching scenes in the synthesis process.

[0025] Preferably, the corresponding depth map can be obtained from the video frame image based on the depth estimation technique of deep learning; for example, the convolutional neural network in deep learning technology can be used for depth estimation, and methods such as supervised learning and self-supervised learning can be used, which have considerable accuracy.

[0026] (3) Multi-plane image layering synthesis module, used to convert the video frame image into a multi-layer depth representation according to the set depth layer number and the depth map, and render it as a left-eye view image and a right-eye view image, and set the parallax baseline of the left-eye view image and the right-eye view image.

[0027] Specifically, according to the set number of depth layers and their corresponding depth maps, the video frame images are converted into multi-layer depth representation images. Since aerial imaging is a three-dimensional perspective image, the parallax baseline of the left-eye and right-eye perspective images is set according to the imaging position and the distance between the user's viewing position, and the video frame images are rendered as left-eye and right-eye perspective images.

[0028] For example, the depth layer number N can be set to 16 to render two perspectives (left eye and right eye); regarding the parallax baseline, it can be set to a user-adjustable range of 0 to 32 pixels (at 1080p), with a default of 16 pixels; using ARM NEON acceleration, the frame takes about 5ms.

[0029] (4) An aerial imaging drive module is used to output the left eye view image and the right eye view image to the aerial imaging panel side by side according to the parallax baseline and drive it to be displayed.

[0030] Specifically, after generating left-eye and right-eye view images, they are output to the aerial imaging panel in a side-by-side format, driving the aerial imaging panel to display the images, thereby producing a 3D display effect that can be viewed with the naked eye.

[0031] For example, the aerial imaging driver module outputs side-by-side left and right view images (the resolution can be 1920×1080, i.e., 960×1080 for the left eye and 960×1080 for the right eye), and sends them to an external aerial imaging panel for playback. During playback, the hovering distance and brightness of the aerial imaging panel can be configured in advance to match a suitable viewing position.

[0032] (5) The aerial imaging panel displays the left-eye view image and the right-eye view image in an incoherent light manner as an aerial suspended naked-eye three-dimensional image based on the configured suspension distance and brightness information.

[0033] In this embodiment, the aerial imaging panel can project left-eye and right-eye view images into the air using light emitted from its internal light source. The two light signals are incoherently imaged at a set position, thus displaying a naked-eye 3D image suspended in the air.

[0034] For example, the aerial imaging panel can suspend images directly in the air without any screen or medium using a negative refraction flat plate lens, and can support direct finger control in the air.

[0035] Preferably, the aerial levitation naked-eye 3D display device based on 2D video of this application may also include a human-computer interaction module, which can be used to display a user control and interaction interface, and further realize the device through a remote control / mobile app, such as adjusting 3D depth intensity, color style, etc.

[0036] The technical solutions described above can be used to achieve real-time naked-eye 3D conversion in common 2D video formats without relying on dedicated 3D content sources, thus reducing dependence on 3D content sources and expanding the range of playable content. Furthermore, the use of an incoherent light aerial imaging scheme avoids the risks of coherent light speckle and retinal hotspots compared to laser holographic schemes. The output signal is a side-by-side ordinary video stream, without generating high-intensity ultrasound, eliminating laser speckle, and offering adjustable parallax depth to prevent dizziness. The entire device can be implemented by adding some hardware to a home playback device, making it suitable for widespread home adoption.

[0037] To further clarify the technical solution of this application, more embodiments are described below, with reference to [reference needed]. Figure 2 As shown, Figure 2 This is a schematic diagram of another embodiment of a naked-eye 3D display device based on 2D video that is suspended in mid-air.

[0038] In some embodiments, in order to improve the playback effect of old black and white videos, the aerial levitation naked-eye 3D display device based on 2D video of this application may also include a black and white video colorization module connected to the video input and preprocessing module, for inputting black and white video images and converting the black and white video images into color 2D images for input to the video input and preprocessing module.

[0039] Furthermore, the black-and-white video colorization module can also be used to downsample the previous frame's color image and fuse the previous frame's color image during the decoding stage of the black-and-white video image to obtain a color two-dimensional video image.

[0040] For example, a black-and-white video to color conversion module may include a black-and-white detection module and a color conversion module; wherein, the black-and-white detection module is used to detect whether the currently input two-dimensional video is a black-and-white video, and if so, the color conversion module is activated to perform color conversion processing on the black-and-white video to obtain a color two-dimensional video image.

[0041] Specifically, in order to achieve the above-mentioned process of fusing and generating color two-dimensional video images, the colorization module can generate the colorization by constructing a lightweight fully convolutional semantic segmentation network (U-Net). For example, the grayscale image (540p) decoded from the black and white video and the color image of the previous frame (downsampled to 540p) can be input into the fully convolutional semantic segmentation network to output a color image (540p).

[0042] The lightweight U-Net can be designed to include a 4-layer encoder, a 4-layer decoder, skip connections, and approximately 1.2M parameters. After downsampling the color image of the previous frame, it is concatenated and fused with the feature image of the current frame in the decoder stage. The loss function of the lightweight U-Net includes color reconstruction loss and temporal consistency loss. The optical flow can be calculated using the pre-trained FlowNet2 (simplified version). The temporal loss is defined as a penalty for the difference between the color image of the previous frame after optical flow deformation and the color image of the current frame. During inference, the first frame has no historical information, and the colorization result is directly generated from the grayscale image.

[0043] Preferably, the lightweight U-Net can be deployed before the video input and preprocessing modules, with a model size of approximately 10MB and an NPU inference time of approximately 20ms. It is only enabled when black and white video content is detected and the user turns it on.

[0044] In some embodiments, in order to achieve better processing effects on black and white videos, the aerial levitation naked-eye 3D display device based on 2D video of this application; the video input and preprocessing module can also be used to execute an adaptive deinterlacing algorithm to convert the color 2D video image into a progressive scan signal, and dynamically adjust the output refresh rate by detecting the field frequency of the input signal; the real-time depth estimation engine can also be used to perform pixel-level alignment of large-scale motion between adjacent frames based on a block matching dense optical flow algorithm to eliminate image jitter.

[0045] Specifically, since old black-and-white videos typically use interlaced scanning and their frame rates do not match those of modern 2D videos, the video input and preprocessing module first executes an adaptive deinterlacing algorithm (such as the YADIF algorithm) to convert the input source into a progressive scan signal. This involves first performing deinterlacing and frame rate unification processing, and then performing optical flow calculation. Simultaneously, by detecting the field frequency of the input signal, the output refresh rate (24Hz / 30Hz / 60Hz) is dynamically adjusted, and optical flow motion compensation is introduced into the real-time depth estimation engine to eliminate image jitter caused by frame rate conversion.

[0046] For example, motion compensation employs a block-matching-based dense optical flow algorithm to perform pixel-level alignment of large-scale motion between adjacent frames, avoiding depth map flicker and thus eliminating motion artifacts.

[0047] In some embodiments, the aerial levitation naked-eye 3D display device based on two-dimensional video of this application may further include: a tracking and adaptive optimization module, used to capture the user's face image in real time using a camera, and adaptively and dynamically adjust the parallax baseline according to the position information of the face image.

[0048] Specifically, the tracking and adaptive optimization module can capture real-time images of the user through an external camera, recognize facial images through real-time images, detect the position of the user's head, and dynamically adjust the parallax baseline according to needs.

[0049] The technical solution of the above embodiments can support multi-view projection encoding by fixing dual viewing angles and utilizing the difference in human eye position, so that family members can watch at the same time and achieve a multi-viewer sharing effect.

[0050] In some embodiments, in order to better manage the temperature of the device and ensure its stability, the aerial levitation naked-eye 3D display device based on 2D video of this application may further include: a dynamic thermal management module, used to reduce the input resolution of the real-time depth estimation engine or turn off the black-and-white video colorization module when the temperature exceeds a set temperature threshold.

[0051] As described in the above embodiments, temperature control management can ensure video playback stability, thereby improving playback quality and enhancing the user's viewing experience.

[0052] In some embodiments, a relative depth map can be inferred from a single two-dimensional video frame image using a monocular depth estimation technique based on deep learning.

[0053] Specifically, the image depth estimation engine can access video frame images, use a pre-trained lightweight convolutional neural network to perform depth estimation on the video frame images, and output a depth image of the same resolution.

[0054] For example, a lightweight convolutional neural network can use a simplified MobileNetV3-small+ decoder, taking a 540p (960×540) input and outputting an 8-bit depth map of the same resolution; the specific parameters are shown in the table below:

[0055] For the decoder of a lightweight convolutional neural network, each Upsample+Conv layer enlarges the feature map size by a factor of 2 (using bilinear interpolation), followed by a 3×3 convolution, with skip connections passing features from the corresponding encoder layer.

[0056] Preferably, the lightweight convolutional neural network can use INT8 quantization. INT8 quantization is a method that can compress floating-point tensors into INT8 quantization. The model size is about 3MB, and the running platform can be RK3588 NPU (6 TOPS computing power). The actual performance is 540p@30fps, and the single frame time is about 8ms.

[0057] For example, the lightweight convolutional neural network in this embodiment can include self-labeled data from a public dataset for training. The public dataset can use NYU Depth V2 (15,000 frames), KITTI (5,000 frames), or Scannet (25,000 frames). The self-labeled data can use the Blender Cycles rendering engine to generate 5,000 frames of diverse 3D scenes (including indoor, outdoor, animation, and live-action styles), while outputting color images and accurate depth maps. 500 frames are manually sampled, and the automatically generated depth maps are compared with the manually corrected depth maps. The average absolute relative error is <1%. The total training set is approximately 55,000 frames, and the validation set is 2,000 frames.

[0058] The following describes an embodiment of a method for displaying naked-eye 3D levitation in mid-air based on 2D video.

[0059] refer to Figure 3 As shown, Figure 3 This is a flowchart of an embodiment of a naked-eye 3D display method based on 2D video for aerial levitation, including the following steps: Step S101: The input two-dimensional video is decoded using a decoder to obtain video frame images, and the video frame images are then scaled and converted in color space.

[0060] Specifically, a two-dimensional video can be input into a decoder to decode it frame by frame to obtain video frame images. Then, the video frame images are scaled to a set size and subjected to corresponding color space conversion. Color space conversion can change the color representation of the video frame images.

[0061] Step S102: Perform image depth analysis on the video frame image to obtain the corresponding depth map.

[0062] Specifically, depth analysis can be performed on two-dimensional video frame images to obtain the corresponding depth map. The depth map is a two-dimensional matrix image, where each pixel value represents the distance from the corresponding scene point to the camera. It can be used in subsequent 3D reconstruction and stereo matching scenes in the synthesis process.

[0063] Furthermore, a relative depth map can be inferred from a single 2D video frame image using monocular depth estimation techniques based on deep learning.

[0064] Specifically, a pre-trained lightweight convolutional neural network can be used to estimate the depth of video frame images and output a depth image of the same resolution.

[0065] Step S103: Convert the video frame image into a multi-layer depth representation according to the set depth layer number and the depth map, and render it as a left-eye view image and a right-eye view image, and set the parallax baseline of the left-eye view image and the right-eye view image.

[0066] For example, the depth layer number N can be set to 16 to render two viewpoints (left eye and right eye); regarding the parallax baseline, the user can adjust it from 0 to 32 pixels (at 1080p), with a default of 16 pixels. Step S104: Output the left-eye view image and the right-eye view image side by side to the aerial imaging panel according to the parallax baseline and drive it to be displayed; wherein, the aerial imaging panel displays the left-eye view image and the right-eye view image as an aerial suspended naked-eye 3D image according to the configured suspension distance and brightness information.

[0067] For example, the aerial imaging driver module outputs side-by-side view images (with a resolution of 1920×1080, i.e., 960×1080 for the left eye and 960×1080 for the right eye), and can be sent to an external aerial imaging panel for display via an HDMI interface.

[0068] In some embodiments, to improve the playback effect of old black and white videos, the aerial levitation naked-eye 3D display method based on 2D video of this application may further include: S100: When the input is detected to be a black and white video image, the black and white video image is converted into a color two-dimensional image.

[0069] Furthermore, colorization can employ a time-consistent U-Net, which downsamples the previous frame's color image and fuses it during the decoding stage of the black-and-white video image to obtain a color two-dimensional video image.

[0070] In some embodiments, in order to achieve better processing effects on black and white videos, the aerial levitation naked-eye 3D display method based on 2D video of this application may further include the following steps: In step S101, during the process of decoding the input two-dimensional video using a decoder to obtain video frame images, an adaptive deinterlacing algorithm can also be used to convert the color two-dimensional video image into a progressive scan signal, and the output refresh rate can be dynamically adjusted by detecting the field frequency of the input signal.

[0071] In step S103, when converting the video frame image into a multi-layer depth representation according to the set number of depth layers and the depth map, pixel-level alignment of large-scale motion between adjacent frames can also be performed based on the block matching dense optical flow algorithm to eliminate image jitter.

[0072] In some embodiments, to better manage the temperature of the device and ensure its stability, the aerial levitation naked-eye 3D display method based on 2D video of this application may further include the following steps: The device temperature is monitored in real time. When the temperature exceeds the set temperature threshold, in step S102, the input resolution of the video frame image is reduced or the process of converting the black and white video image into a color two-dimensional image in step S100 is turned off.

[0073] In some embodiments, the aerial levitation naked-eye 3D display method based on 2D video of this application may further include the following steps in step S104, during which the left-eye view image and the right-eye view image are output side-by-side to the aerial imaging panel according to the parallax baseline and the panel is driven to display the image: The system uses a camera to capture real-time images of the user's face and dynamically adjusts the parallax baseline based on the positional information of the face image.

[0074] For example, a user's real-time image can be captured by an external USB camera, and the face image can be recognized in real time through a face tracking algorithm. Then, the user's face position information can be detected. The parallax baseline can be dynamically adjusted according to needs. For example, the parallax baseline can be adjusted for a specific user based on the detected faces of different users, which can be adapted to multi-user applications in home scenarios.

[0075] Based on the solutions of the foregoing embodiments, the hardware principle block diagram of the aerial levitation naked-eye 3D display solution based on 2D video in this application can be referred to. Figure 4 As shown, Figure 4 This is a hardware principle block diagram of an aerial levitation naked-eye 3D display device based on 2D video, as shown in the figure. Its hardware structure mainly includes the following main parts: Main control SoC: can use RK3588 (including 4-core Cortex-A76, 4-core Cortex-A55, NPU 6TOPS, Mali-G610 GPU); runs on Linux and is responsible for video decoding, depth estimation, MPI synthesis and other functions.

[0076] Memory unit: It can use 4GB LPDDR4, connected to RK3588 via a 64-bit bus.

[0077] Storage unit: Can use 32GB eMMC 5.1, connected to RK3588 SDMMC interface.

[0078] HDMI input port: The LT8619C chip (HDMI 2.0 to MIPI CSI) can be used. The HDMI input port is connected to the LT8619C chip, and its MIPI CSI output is connected to the CSI interface of the RK3588, supporting 4K@30fps input.

[0079] HDMI output port: It can use the IT66121 chip (HDMI 2.0 transmitter) to connect to the HDMITX interface of the RK3588, and can output 1080p@60Hz side by side to the air imaging panel.

[0080] USB Hub Interface: Utilizing the GL852G chip, it provides two USB 2.0 ports for connecting to the RK3588's USB Host. One port is for a USB camera (optional), and the other is reserved.

[0081] I2C bus: Connected to the LT8619C and IT66121 via I2C2 of RK3588, which can be used for configuration.

[0082] Power management module: 12V DC input, converted to 5V, 3.3V, 1.8V, and 0.9V via DC-DC converter (MP9942). Note: Total power consumption <10W.

[0083] Clock module: It can output a 24MHz crystal oscillator for use by RK3588, and output a 27MHz clock signal for use by HDMI input port and HDMI output port.

[0084] refer to Figure 5 As shown, Figure 5 This is a software hierarchy diagram of an embodiment of the naked-eye 3D display scheme based on 2D video for aerial levitation, as described in this application. The main hierarchical structure is as follows: From top to bottom, it is divided into: user interface layer, application logic layer, algorithm layer, and hardware abstraction layer.

[0085] (a) User Interface Layer: It can include a remote control / App interface module to receive user input (3D intensity, color on / off switch, etc.).

[0086] (ii) Application Logic Layer: ① Main control thread: Manages the entire pipeline.

[0087] ② Video capture thread: Reads frames from HDMI input and sends them to the preprocessing queue.

[0088] ③ Output thread: Outputs the rendered frames via HDMI.

[0089] (III) Algorithm Layer: ① Real-time depth estimation engine (NPU): Input RGB frames and output depth maps (see the table in the manual for the specific network structure, which is shown here).

[0090] ② Black and white detection module: Histogram analysis to mark black and white content.

[0091] ③ Colorization module (optional): Grayscale to RGB conversion, using U-Net+ timing loss.

[0092] ④ Multi-plane image layering synthesis module (MPI): 16-layer decomposition, rendering left-eye view image and right-eye view image.

[0093] (iv) Hardware Abstraction Layer: This can include HDMI drivers, NPU drivers, I2C drivers, USB camera drivers, etc.

[0094] refer to Figure 6 As shown, Figure 6 This is a data flow and latency breakdown diagram of an example of a 2D video-based, airborne, glasses-free 3D display device; its complete pipeline from input to output and the latency budget for each stage are shown in the table below:

[0095] As shown in the figure, the total latency of the entire processing flow is ≤45ms (colorization off) and ≤65ms (colorization on). Colorization is off by default, which meets the <60ms requirement.

[0096] refer to Figure 7 As shown, Figure 7 This is a timing diagram of the interface between an aerial imaging driver module and an aerial imaging panel according to one embodiment. The aerial imaging driver module outputs video data via an HDMI output port at 1920×1080@60Hz, in a side-by-side format. The left half of the screen (columns 0-959) displays the image from the left eye's perspective, and the right half (columns 960-1919) displays the image from the right eye's perspective. The pixel clock is 148.5MHz. Control signals are transmitted via an I2C bus to control the parameters of the aerial imaging panel. Address 0x3C, and the registers may include: 0x01: hover distance (0-100, corresponding to 20cm-2m); 0x02: brightness (0-255); 0x03: contrast ratio (0-255). After the aerial imaging panel is powered on, the aerial imaging driver module first configures the parameters of the aerial imaging panel via the I2C bus, and then outputs the video stream via the HDMI output port.

[0097] For the user interface, a remote control can be used, which can be designed as a regular button. Alternatively, it can be controlled via a mobile app. The control interface can include multiple interface elements, such as: a slider for "3D Depth Intensity" (0%~100%); buttons for "Movie Mode", "Action Mode", and "Old Film Restoration Mode"; a checkbox for "Auto-coloring in Black and White"; an indicator light for "Audience Tracking Enabled"; and functions for "Real-time Adjustment, Instant Effect" and "Depth Intensity Default 30% Gentle 3D".

[0098] refer to Figure 8 As shown, Figure 8 This is an example assembly diagram of a 2D video-based, levitating, glasses-free 3D display device. The display device is designed within a main unit box. The rear interface of the main unit box may include the following: DC 12V input (5.5×2.1mm), HDMI input (Type A), i.e., HDMI IN, HDMI output (Type A), i.e., HDMI OU, USB-A (camera), and I2C debugging interface (4-pin). The wiring steps are as follows: Connect the video source (such as a TV box) to the HDMI IN of the main unit box using an HDMI cable; connect the HDMI OUT of the main unit box to the HDMI IN of the levitating display panel using another HDMI cable; connect a USB camera if head tracking is required; connect a 12V power adapter.

[0099] For thermal management, the RK3588 main control chip of the SoC is surface-mounted with graphene thermal pads and an aluminum alloy heat sink. The system software layer has a dynamic frequency adjustment function: when the chip temperature exceeds 80℃, the input resolution of the real-time depth estimation engine is automatically downgraded from 540p to 360p, or the black-and-white video colorization module is temporarily shut down to prioritize smooth video playback (no less than 24fps), and automatically restored after the temperature drops. At the same time, the NPU core voltage can be dynamically adjusted to reduce power consumption in low-load scenarios, thereby reducing the overall power consumption of the system.

[0100] The following describes an embodiment of a computer-readable storage medium.

[0101] This application provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded by a processor and executes the steps of the aerial levitation naked-eye 3D display method based on two-dimensional video according to any embodiment of this application.

[0102] In an exemplary embodiment, the computer-readable storage medium may be a non-transitory computer-readable storage medium that includes instructions, such as a memory that includes instructions. For example, a non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0103] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A naked-eye 3D display device based on 2D video, characterized in that, include: The video input and preprocessing module is used to decode the input two-dimensional video using a decoder to obtain video frame images, and to perform image scaling and color space conversion on the video frame images; A real-time depth estimation engine is used to perform image depth analysis on the video frame images to obtain the corresponding depth maps; The multi-plane image layering synthesis module is used to convert the video frame image into a multi-layer depth representation according to the set number of depth layers and the depth map, and render it into a left-eye view image and a right-eye view image, and set the disparity baseline of the left-eye view image and the right-eye view image. An aerial imaging drive module is used to output the left-eye view image and the right-eye view image side by side to the aerial imaging panel according to the parallax baseline and drive it to be displayed. The aerial imaging panel displays the left-eye and right-eye perspective images as non-coherent three-dimensional images suspended in the air using incoherent light, based on the configured suspension distance and brightness information.

2. The aerial levitation naked-eye 3D display device based on 2D video according to claim 1, characterized in that, Also includes: A black-and-white video colorization module is used to input black-and-white video images and convert them into color two-dimensional images for input into the video input and preprocessing module.

3. The aerial levitation naked-eye 3D display device based on 2D video according to claim 2, characterized in that, The black-and-white video colorization module is further used to downsample the previous frame color image and fuse the previous frame color image during the decoding stage of the black-and-white video image to obtain a color two-dimensional video image.

4. The aerial levitation naked-eye 3D display device based on 2D video according to claim 3, characterized in that, The video input and preprocessing module is also used to execute an adaptive deinterlacing algorithm to convert the color two-dimensional video image into a progressive scan signal, and dynamically adjust the output refresh rate by detecting the field frequency of the input signal; The real-time depth estimation engine is also used in a block-matching-based dense optical flow algorithm to perform pixel-level alignment of motion between adjacent frames.

5. The aerial levitation naked-eye 3D display device based on 2D video according to claim 1, characterized in that, Also includes: The tracking and adaptive optimization module is used to capture the user's face image in real time using the camera and adaptively and dynamically adjust the parallax baseline based on the position information of the face image.

6. The aerial levitation naked-eye 3D display device based on 2D video according to claim 5, characterized in that, Also includes: The dynamic thermal management module is used to reduce the input resolution of the real-time depth estimation engine or turn off the black-and-white video colorization module when the temperature exceeds a set temperature threshold.

7. The aerial levitation naked-eye 3D display device based on 2D video according to claim 1, characterized in that, The image depth estimation engine is used to input video frame images, perform depth estimation on the video frame images using a pre-trained lightweight convolutional neural network, and output a depth image of the same resolution; wherein, the lightweight convolutional neural network includes a bilinear interpolation layer and a convolutional layer.

8. A method for naked-eye 3D display based on 2D video with aerial suspension, characterized in that, Includes the following steps: The input two-dimensional video is decoded using a decoder to obtain video frame images, and the video frame images are then scaled and converted in color space. Image depth analysis is performed on the video frame images to obtain the corresponding depth maps; The video frame image is converted into a multi-layer depth representation according to the set depth layer and the depth map, and rendered as a left-eye view image and a right-eye view image. The parallax baseline of the left-eye view image and the right-eye view image is set. The left-eye and right-eye view images are output side-by-side to the aerial imaging panel based on the parallax baseline and driven to be displayed; wherein, the aerial imaging panel displays the left-eye and right-eye view images as aerial suspended naked-eye 3D images according to the configured suspension distance and brightness information.

9. The method for aerial levitation naked-eye 3D display based on 2D video according to claim 8, characterized in that, Before using a decoder to decode the input two-dimensional video to obtain video frame images, the process also includes: When the input is detected to be a black and white video image, the black and white video image is converted into a color two-dimensional image.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or instruction set is loaded by a processor and executed according to the steps of the aerial levitation naked-eye 3D display method based on two-dimensional video as described in claim 8 or 9.