Naked eye 3D end-to-end real-time communication method and system
By simplifying hardware and software synchronization technology and combining raster display and depth estimation, low-cost, high-quality naked-eye 3D end-to-end communication is achieved, solving the problems of high hardware cost, complex system and poor viewing experience in existing technologies, and providing a flexible 3D visual experience.
Patent Information
- Application Number
- CN202510881811.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-03
AI Technical Summary
Existing naked-eye 3D end-to-end communication technology has problems such as high hardware cost, complex system, poor viewing experience, and high technical threshold. In particular, in complex scenarios, crosstalk is severe and the comfortable viewing area is narrow.
It uses two cameras, a switch and a grating display, combined with hardware and software synchronization technology, and dynamically adjusts the grating projection angle through image stitching, depth estimation and light distribution algorithms to achieve naked-eye 3D effects.
Significantly reduce hardware costs and technical barriers, improve viewing comfort and flexibility, provide a natural and convenient 3D visual experience, and adapt to complex scenes and user movements.
Smart Images

Figure CN120751107A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D communication technology, and in particular to a naked-eye 3D end-to-end real-time communication method and system. Background Art
[0002] With the rapid advancement of technology, people's demand for a superior visual experience has reached a new level, and traditional 2D communication methods no longer meet these needs. Against this backdrop, 3D communication technology has emerged and become a research hotspot, demonstrating valuable application prospects in numerous fields. In the medical field, 3D communication technology enables precise and efficient remote 3D consultations, helping doctors obtain a comprehensive and three-dimensional view of patient lesions and formulate more precise treatment plans. In education, 3D communication technology can create immersive teaching environments, allowing students to feel as if they were immersed in historical sites, the microcosm, or the distant universe, greatly enhancing the fun of learning. In the industrial sector, 3D communication technology allows maintenance personnel to clearly visualize the complex internal structures of equipment and accurately locate fault points, significantly improving maintenance efficiency and reducing downtime losses.
[0003] Currently, virtual reality (VR) technology is a relatively mature approach to achieving 3D effects. However, VR technology also has significant limitations. Users must wear additional equipment such as VR glasses and a VR helmet. Prolonged wear often causes discomfort, such as dizziness and a sense of restraint, which severely limits user usage duration and frequency. In contrast, glasses-free 3D technology, by eliminating the need for additional equipment, completely eliminates these discomforts, providing a more natural and convenient 3D visual experience. Therefore, it holds irreplaceable potential in the 3D communications field.
[0004] However, existing naked-eye 3D end-to-end communication technology has significant drawbacks: (1) High hardware cost and complex system: They generally rely on multiple (usually ≥3) high-precision cameras, dedicated depth sensors (such as structured light, ToF), complex synchronization equipment, and high-end computing units. This not only leads to extremely high system construction costs and complex deployment and maintenance, but also greatly limits its popularization and application.
[0005] (2) Limited viewing experience: Naked-eye 3D displays using fixed-parameter gratings have serious crosstalk in complex scenes, and the comfortable viewing area is extremely narrow. Users need to maintain a specific position and posture to obtain a good 3D effect, resulting in a poor user experience.
[0006] (3) High technical threshold: The system integrates multiple complex technologies, which makes technical implementation difficult and has poor flexibility. Summary of the Invention
[0007] In order to solve the above problems, the present invention proposes a naked-eye 3D end-to-end real-time communication method and system, which effectively solves the core problems of the above-mentioned existing technologies such as high hardware cost, complex system, poor viewing experience, and high technical threshold.
[0008] According to some embodiments, the present invention adopts the following technical solutions: A naked-eye 3D end-to-end real-time communication method, comprising: The transmitter collects two video streams and a single audio stream of the target scene, and synchronizes them through a combination of hardware and software. The synchronized two video streams are then spliced frame by frame to generate three queues of image frames, audio frames, and timestamps, which are then sent to the receiver item by item according to the queues. The receiving end synchronizes the received image frames and audio frames based on the timestamps, and uses the grating display and speakers to perform naked-eye 3D presentation of the target scene for the synchronized image frames and audio frames; Among them, the naked-eye 3D presentation of the image frame is based on the depth information in the image frame, calling the light distribution algorithm, dynamically adjusting the light projection angle and distribution range of the grating in the grating display screen, and accurately projecting the corresponding pixel points in the redrawn image frame into the user's left and right eyes.
[0009] According to some embodiments, the present invention adopts the following technical solutions: A naked-eye 3D end-to-end real-time communication system, comprising: The sending module is configured to: collect two video streams and a single audio stream of the target scene at the sending end, synchronize the two video streams and the single audio stream through a combination of hardware and software, perform frame-by-frame image splicing on the two synchronized video streams, and obtain three queues of image frames, audio frames, and timestamps, and send them to the receiving end one by one according to the queues; The receiving module is configured to: synchronize the received image frames and audio frames based on the timestamp at the receiving end, and use the grating display screen and the speaker to perform naked-eye 3D presentation of the target scene on the synchronized image frames and audio frames; Among them, the naked-eye 3D presentation of the image frame is based on the depth information in the image frame, calling the light distribution algorithm, dynamically adjusting the light projection angle and distribution range of the grating in the grating display screen, and accurately projecting the corresponding pixel points in the redrawn image frame into the user's left and right eyes.
[0010] According to some embodiments, the present invention adopts the following technical solutions: A computer program product includes a computer program, which implements the naked-eye 3D end-to-end real-time communication method when executed by a processor.
[0011] According to some embodiments, the present invention adopts the following technical solutions: A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the naked-eye 3D end-to-end real-time communication method is implemented.
[0012] According to some embodiments, the present invention adopts the following technical solutions: An electronic device includes: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the naked-eye 3D end-to-end real-time communication method.
[0013] Compared with the prior art, the present invention has the following beneficial effects: The present invention innovatively implements a simple and efficient naked-eye 3D end-to-end communication method. Compared with existing systems, it does not rely on multiple cameras, dedicated depth sensors and complex synchronization equipment, greatly reducing hardware costs and system complexity; effectively solves the serious crosstalk and narrow comfortable viewing area caused by fixed gratings, and improves viewing comfort; does not require complex multimodal fusion, synchronization and control technologies, reducing technical requirements and improving flexibility.
[0014] In the present invention, only four peripheral devices, namely two cameras, a small switch, and a grating screen display, and a small amount of algorithms are needed to build a single-ended system that can realize naked-eye 3D end-to-end communication. On the one hand, it greatly reduces the hardware cost and construction difficulty, reduces the maintenance cost and technical threshold, and improves the reliability and stability of the system. On the other hand, it avoids the discomfort of users wearing the device, brings a more natural and convenient 3D visual interaction experience, and increases users' willingness to use it. This invention effectively solves the shortcomings of existing technologies, promotes the widespread application of naked-eye 3D technology in multiple fields, promotes its market penetration and industry development, promotes the progress of related industries, and has significant economic and social benefits.
[0015] In addition, the present invention can also flexibly replace equipment according to the implementation effect. For example, if only 720P or 1080P naked-eye 3D communication effect is required, a low-end camera, a low-end CPU and a low-resolution raster display can be selected to further reduce hardware costs; if 4K or 8K naked-eye 3D communication effect is required, a high-end camera, a high-end graphics card and a high-resolution raster display need to be replaced. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0017] Figure 1 This is a framework diagram of existing naked-eye 3D technology. Figure 2 This is a flow chart of the method of Example 1.
[0018] Figure 3 Schematic diagram of the camera and switch of Example 1.
[0019] Figure 4 This is the original structure diagram of the stereo matching model in Example 1.
[0020] Figure 5 This is a diagram of the improved structure of the stereo matching model in Example 1.
[0021] Figure 6 This is a technical idea diagram of Example 1. DETAILED DESCRIPTION
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0023] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0024] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "comprising" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0025] Example 1 An embodiment of the present invention provides a naked-eye 3D end-to-end real-time communication method, comprising: The transmitter collects two video streams and a single audio stream of the target scene, and synchronizes them through a combination of hardware and software. The synchronized two video streams are then spliced frame by frame to generate three queues of image frames, audio frames, and timestamps, which are then sent to the receiver item by item according to the queues. The receiving end synchronizes the received image frames and audio frames based on the timestamps, and uses the grating display and speakers to perform naked-eye 3D presentation of the target scene for the synchronized image frames and audio frames; Among them, the naked-eye 3D presentation of the image frame is based on the depth information in the image frame, calling the light distribution algorithm, dynamically adjusting the light projection angle and distribution range of the grating in the grating display screen, and accurately projecting the corresponding pixel points in the redrawn image frame into the user's left and right eyes.
[0026] As an embodiment, the present invention provides a glasses-free 3D end-to-end real-time communication method. This method uses minimal general-purpose hardware, combined with innovative data processing processes and optimization algorithms, to achieve a high-quality, low-crosstalk, and wide-comfort-zone glasses-free 3D end-to-end communication experience while significantly reducing cost and complexity. This effectively addresses the core issues of the prior art, such as high hardware cost, complex system, poor viewing experience, and high technical barriers. The specific implementation process is as follows: Generally speaking, to achieve naked eye 3D effect, it is necessary to use a variety of complex technologies and equipment, such as Figure 1 As shown, to achieve 3D model reconstruction, multiple cameras are required to collect color images and depth images from different orientations (depth images can be obtained using algorithms such as binocular reconstruction or by processing infrared images using algorithms). The resulting images from multiple orientations and 3D model reconstruction algorithms are then used to further refine the 3D model. This places a huge strain on the computer's computing power, storage performance, network transmission capabilities, and technical performance. Furthermore, existing naked-eye 3D systems suffer from severe crosstalk and a narrow viewing comfort zone in complex scenes. The root cause is that the grating parameters are fixed and cannot dynamically adapt to changes in scene depth.
[0027] In view of the problems existing in the current naked-eye 3D communication methods and systems and the above content, this embodiment proposes a method for simply realizing naked-eye 3D end-to-end communication. By synchronously acquiring, processing, compressing, transmitting, and decompressing 2D video streams and audio streams, naked-eye 3D end-to-end communication can be realized. The main process is as follows: Figure 2 As shown, specifically: 1. Both parties in communication ensure that the network is unobstructed and that cameras, microphones, monitors and other equipment are functioning properly.
[0028] Second, the communicating party, the sender, captures the video streams captured by the two cameras and the audio stream from the microphone, decodes them synchronously, performs image processing operations such as distortion correction on the image frames, and then splices them into a single image frame in the left-right direction. This is then synchronized with the audio and finally compressed for transmission. Specifically: As attached Figure 3As shown, two cameras with identical performance parameters and synchronization interfaces are placed closely side by side, ensuring that the optical axes of their lenses are approximately parallel. This arrangement closely simulates the human eye's binocular parallax, laying the foundation for achieving high-quality glasses-free 3D effects. Sync cables are connected to the synchronization interfaces of the two cameras in sequence. The Sync In interface of camera 1 and the Sync Out interface of camera 2 remain unconnected. Camera 1 is designated as the synchronization source for the entire camera system, leading the synchronization of the two cameras and avoiding image tearing or ghosting caused by asynchrony.
[0029] Use network cables to connect each camera to a switch, and then to the host computer, establishing a stable, high-speed local area network communication environment. This facilitates the rapid and efficient transmission of captured audio and video data to the host computer for processing. Use RTSP (Real-Time Streaming Protocol) to synchronously capture the video streams from two cameras and the audio stream from a single microphone. The captured audio and video streams are decoded and stored with the timestamp and decoded data. During the capture process, ensure that the camera and microphone sampling rate, resolution, and other parameters are set appropriately to obtain high-quality raw audio and video data. The decoded audio and video streams provide the foundation for subsequent processing, and their quality directly impacts the final naked-eye 3D effect.
[0030] Correct and process each decoded image frame. Due to the distortion characteristics of the camera lens (including radial distortion, tangential distortion, etc.), the image needs to be accurately corrected to eliminate the adverse effects caused by the optical characteristics of the camera lens and improve image quality. The checkerboard calibration method can be used to obtain the camera's internal and external parameters to correct the image distortion.
[0031] Synchronize the corrected image and decoded audio. It's known that the two decoded image frames and audio frames are stored in three queues. Each storage unit also stores the timestamp of the image frame / audio frame (the system time at the time of acquisition). Use the timestamp to synchronize the image and audio frames in a separate thread. Since camera 1 uses a synchronization line for hard synchronization, the delay mainly comes from transmission delay. Therefore, the synchronization algorithm can be performed according to the following steps: (1) Obtain the data of the head position of the two image storage queues, calculate the difference of the timestamps of the two image frames, and find two image frames whose difference is less than a threshold. If the difference of the timestamps of the two image frames is greater than a certain threshold, then the frame with the smaller timestamp value is the older frame and should be discarded. The data of the head position of the corresponding image storage queue in the image storage queue should be re-identified immediately; (2) Use the average of the timestamps of the two synchronized image frames and the timestamp of the audio frame to calculate the difference, and find the image frame and audio frame whose difference is less than the threshold. If the timestamp of the image frame is greater than the timestamp of the audio frame, then discard the current time frame, reread the data at the head position of the audio frame storage queue and make another judgment. Otherwise, discard the image frame and re-execute (1); (3) Splice the two image frames horizontally to eliminate the possibility that the receiver may receive two frames with the same or similar timestamps within a large time difference due to subsequent concurrent uploads and network delays; (4) Re-assign the timestamp according to the current system time and store it in the queue in the format of {image frame, audio frame, timestamp}; (5) Wait 1ms and then execute (1).
[0032] The processed image and audio are synchronously encoded and compressed, and finally sent to the receiver to reduce the amount of transmitted data and reduce network bandwidth usage. Third, the other party in communication, that is, the receiver, captures the audio and video streams, and after synchronous decoding, uses the depth estimation module to obtain the depth map, rearranges the image pixels according to the grating screen, and uses the human eye tracking algorithm to obtain information such as the pupil position and pupil distance of both eyes. Combined with the depth information, the light distribution algorithm is used to project the image light into the left and right eyes respectively, thereby producing a more realistic naked-eye 3D effect. Specifically: 1. After receiving the audio and video streams, the receiver stores and decodes the compressed audio and video data frames, then stores the decoded image and audio frames. The decoded image and audio frames are synchronized using timestamps. The timestamps here are sent by the sender as part of the content. Due to network transmission delays, the synchronization algorithm follows the following steps: The data at the head position of the image frame and audio frame queues are taken out respectively, and the timestamp difference is calculated; if the difference is less than the threshold, the two are synchronized and stored as {image frame, audio frame}; otherwise, if the timestamp of the image frame is greater than the timestamp of the audio frame, then the data is stored as {image frame, audio frame}. If the timestamp of the image frame is less than the timestamp of the audio frame, the current image frame is discarded and a new image frame is read and compared with the current audio frame again; the running time Tms of (1) to (2) is calculated. If T < 16, the algorithm needs to sleep for (16-T) ms and then re-execute (1) to meet 60FPS. Otherwise, execute (1) directly. 2. Real-time depth map generation and parameter optimization
[0033] The decoded image frame can be divided into left and right images horizontally, and the epipolar lines are horizontal. Therefore, a lightweight neural network-based stereo matching model is directly used to estimate the disparity map. This is then combined with known information such as camera parameters to obtain the depth map, thereby obtaining the depth information of the image frame.
[0034] This stereo matching model is based on a neural network and has been optimized through model optimization techniques such as pruning and parameter adjustment, as well as fine-tuning training. Compared with commonly used depth estimation algorithms, it is not only more accurate, but can also achieve 1080P 60FPS and ensure more image details. The architectures of the original model and the optimized model are shown below. Figure 4 、 Figure 5 As shown, the model optimization content is as follows: The faster MobileNetV2 replaces Feature Extraction Network 1 in the original model. Testing has shown that processing 1080P image pairs on a 4070TiSuper graphics card increases network processing speed by 7 times in this phase. Feature Extraction Network 2 in the original model has been removed. This feature extraction network primarily provides shallow texture features of the left image for the refinement module. Using shallow features from Feature Extraction Network 1 or MobileNetV2 has similar functionality and reduces model computational overhead. A full-resolution disparity map optimization module has been added. This module optimizes the full-resolution disparity map using the original image and 2D convolutions. Compared to the 3D convolutions in the refinement module, 2D convolutions require less computation and result in faster model computation. Furthermore, this module can be used in conjunction with the refinement module to alternately optimize the disparity map globally and locally. For large-resolution image pairs, the backbone network processes the small image pairs, while the full-resolution disparity map optimization module processes the full-resolution image and disparity map. By splitting the large-resolution image into smaller blocks, the parallel computing power of the graphics card is fully utilized, improving model computational speed. Using the optimized model, the disparity map is estimated as follows: The left and right images are input into the optimized model. If the resolution of the left and right images is lower than 1080P, they are directly input into the optimized model. Otherwise, they are cropped into image blocks and sent into the optimized model to obtain multi-scale cost volumes and feature pairs. After obtaining the multi-scale cost volumes and feature pairs, the refinement module is used to cyclically optimize the disparity map or disparity map blocks. In particular, for large-resolution image pairs, after the disparity map blocks are reassembled into full-resolution disparity maps, the full-resolution disparity map is optimized using the full-resolution disparity map optimization module, and then the full-resolution disparity map is split. Small-resolution disparity maps do not need to be spliced or split. The full-resolution disparity map optimization module and the refinement module are optimized alternately until the specified number of cycles are executed and the final full-resolution disparity map is output. 3. The image frame to be displayed is redrawn according to the grating width. With the help of the built-in camera of the grating display, the human eye tracking algorithm is used to monitor the position of the user's pupil and the direction of sight in real time. The light distribution algorithm is called in combination with the depth information to dynamically adjust the projection angle and distribution range of the grating light. The corresponding pixel points in the redrawn image are accurately projected into the user's left and right eyes, even when the user is moving, thus producing a more realistic naked-eye 3D effect and bringing users an immersive visual experience.
[0035] The main ideas of this embodiment are as follows: Figure 6 As shown, the captured video and audio streams are first decoded, and the images are corrected, spliced, and synchronized in turn to ensure the quality of subsequent effects; then the images and audio are synchronously encoded and compressed for transmission respectively. After decoding, the receiving end applies a synchronization algorithm to ensure the synchronization of audio and video. Combined with the grating display screen and depth estimation, human eye tracking, and light distribution algorithm, pixels are accurately projected to the user's eyes to form a naked-eye 3D effect, ultimately achieving naked-eye 3D end-to-end communication.
[0036] Instead of using the simple method for achieving naked-eye 3D end-to-end communication in this embodiment, other methods are used. Then: (1) More complex hardware: such as increasing the number of cameras (≥3) or using dedicated depth cameras (structured light, ToF), which leads to a sharp increase in cost, an increase in system size, and possible reduction in reliability; (2) More complex synchronization: Multiple signals with different modes need to be synchronized, increasing technical complexity and failure points; (3) Using static pre-calculation adaptation cannot adapt to dynamic scenes and user movements; (4) Computationally intensive reconstruction: Depth acquisition relies on traditional multi-view reconstruction algorithms or large, unoptimized models, which are difficult to run in real time on hardware.
[0037] Example 2 In one embodiment of the present invention, a naked-eye 3D end-to-end real-time communication system is provided, comprising: The sending module is configured to: collect two video streams and a single audio stream of the target scene at the sending end, synchronize the two video streams and the single audio stream through a combination of hardware and software, perform frame-by-frame image splicing on the two synchronized video streams, and obtain three queues of image frames, audio frames, and timestamps, and send them to the receiving end one by one according to the queues; The receiving module is configured to: synchronize the received image frames and audio frames based on the timestamp at the receiving end, and use the grating display screen and the speaker to perform naked-eye 3D presentation of the target scene on the synchronized image frames and audio frames; Among them, the naked-eye 3D presentation of the image frame is based on the depth information in the image frame, calling the light distribution algorithm, dynamically adjusting the light projection angle and distribution range of the grating in the grating display screen, and accurately projecting the corresponding pixel points in the redrawn image frame into the user's left and right eyes.
[0038] Example 3 In one embodiment of the present invention, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements the naked-eye 3D end-to-end real-time communication method.
[0039] Example 4 In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the naked-eye 3D end-to-end real-time communication method is implemented.
[0040] Example 5 In one embodiment of the present invention, an electronic device is provided, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the naked-eye 3D end-to-end real-time communication method.
[0041] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0042] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0043] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A naked-eye 3D end-to-end real-time communication method, characterized in that: include: The transmitter collects two video streams and a single audio stream of the target scene, and synchronizes them through a combination of hardware and software. The synchronized two video streams are then spliced frame by frame to generate three queues of image frames, audio frames, and timestamps, which are then sent to the receiver item by item according to the queues. The receiving end synchronizes the received image frames and audio frames based on the timestamps, and uses the grating display and speakers to perform naked-eye 3D presentation of the target scene for the synchronized image frames and audio frames; Among them, the naked-eye 3D presentation of the image frame is based on the depth information in the image frame, calling the light distribution algorithm, dynamically adjusting the light projection angle and distribution range of the grating in the grating display screen, and accurately projecting the corresponding pixel points in the redrawn image frame into the user's left and right eyes.
2. The naked-eye 3D end-to-end real-time communication method according to claim 1, characterized in that: The combination of hardware and software includes: Use two cameras with the same performance parameters and synchronization interfaces to capture two video streams. Use synchronization cables to connect the synchronization interfaces of the two cameras in sequence. Designate one camera as the synchronization source to lead the synchronization operation of the two cameras. Based on the timestamps of two image frames in the two video streams, the two image frames whose timestamp difference is less than a threshold are used as synchronized image frames. The difference is calculated using the average value of the timestamps of the two synchronized image frames and the timestamp of the audio frame. The image frame and audio frame whose difference is less than the threshold are used as the two synchronized image frames and one audio frame.
3. The naked-eye 3D end-to-end real-time communication method according to claim 1, wherein: The image splicing is to splice the two synchronized image frames in the horizontal direction.
4. The naked-eye 3D end-to-end real-time communication method according to claim 1, wherein: The depth information in the image frame is represented by a depth map. The depth map is constructed by estimating the disparity map using a lightweight neural network-based stereo matching model, and then combining it with camera parameter information to obtain the depth map.
5. The naked-eye 3D end-to-end real-time communication method according to claim 4, characterized in that: The lightweight neural network-based stereo matching model is used to estimate the disparity map, specifically: Using MobileNetV2, we extract multi-scale cost volumes and feature pairs from the left and right images of the image frame. After obtaining the multi-scale cost volume and feature pairs, the full-resolution disparity map optimization module and the refinement module are used to alternately optimize the disparity blocks and disparity maps until the specified number of cycles are executed and the final full-resolution disparity map is output.
6. The naked-eye 3D end-to-end real-time communication method according to claim 1, wherein: The light distribution algorithm also includes using an eye tracking algorithm to monitor the position of the user's pupil and the direction of sight in real time, and dynamically adjusting the light projection angle and distribution range of the grating in combination with depth information.
7. A naked-eye 3D end-to-end real-time communication system, characterized in that: include: The sending module is configured to: collect two video streams and a single audio stream of the target scene at the sending end, synchronize the two video streams and the single audio stream through a combination of hardware and software, perform frame-by-frame image splicing on the two synchronized video streams, and obtain three queues of image frames, audio frames, and timestamps, and send them to the receiving end one by one according to the queues; The receiving module is configured to: synchronize the received image frames and audio frames based on the timestamp at the receiving end, and use the grating display screen and the speaker to perform naked-eye 3D presentation of the target scene on the synchronized image frames and audio frames; Among them, the naked-eye 3D presentation of the image frame is based on the depth information in the image frame, calling the light distribution algorithm, dynamically adjusting the light projection angle and distribution range of the grating in the grating display screen, and accurately projecting the corresponding pixel points in the redrawn image frame into the user's left and right eyes.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for naked-eye 3D end-to-end real-time communication according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the naked-eye 3D end-to-end real-time communication method according to any one of claims 1 to 6 is implemented.
10. An electronic device, characterized in that: include: A processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement a naked-eye 3D end-to-end real-time communication method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Smart television with 3D (three-dimensional) video playing function and processing method thereof
CN102946544A
Two-viewpoint stereo image synthesizing method and system
CN105430368A
Naked-eye 3D live broadcast method and device, 3D screen client and streaming media cloud server
CN108900928A
Naked eye 3D image processing method, device and equipment
CN109522866A
Naked eye 3D medical video image display method and device, storage medium and equipment
CN109951694A