Remote holographic telepresence system based on light field imaging and low-latency synchronization method thereof
Patent Information
- Application Number
- CN202611084126.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]光场纹理流、深度流、音频流、交互指令流编码运算量差异显著,同一触发周期采集的四路数据编码完成时刻不同步;现有技术采用固定等待最慢流打包机制,引入固定排队时延;或直接混序打包造成接收端时序错乱
本发明通过北斗及IEEE1588v2双冗余时钟搭配相机固有时延硬件补偿实现20ns内超高精度采集同步,从根源消除全息重影;依托前景背景差异化分层压缩使传输带宽降低60%以上,结合异构流动态弹性预对齐打包控制排队时延;打破传统传输时延对称假设构建自适应缓冲模型,搭配分层拥塞与差异化FEC提升跨城、5G弱网场景传输鲁棒性;接收端将解码、渲染、SLM刷新、全息屏VSync硬件锁相绑定,把四流同步偏差控制在20ms以内,杜绝全息画面撕裂、残影与视觉眩晕;配合全链路分段时延监测自适应调控,最终将端到端总时延稳定控制在180ms以内,能够满足跨校区全息授课、医学解剖等专业实操类远程教学的低延迟、高稳定使用需求。
Smart Images

Figure CN122845781A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of holographic 3D display online teaching technology, and is particularly applicable to cross-campus immersive holographic teaching, medical anatomy training, and remote 3D interactive demonstration scenarios. Specifically, it is a remote holographic on-site teaching system based on light field imaging and its low-latency synchronization method. Background Technology
[0002] Current mainstream distance learning relies on two-dimensional video live streaming systems, which can only transmit planar images and lack parallax information, resulting in a lack of immersive experience for both teachers and students. Furthermore, it cannot provide three-dimensional, intuitive demonstrations of objects such as three-dimensional teaching aids, human anatomy, and mechanical assembly parts, leading to extremely poor remote teaching effectiveness for practical courses. To address this issue, existing technologies are gradually incorporating light field acquisition and holographic reconstruction schemes. These schemes use multi-viewpoint camera arrays to acquire spatial light information and reconstruct naked-eye three-dimensional holographic images at the receiving end, enabling remote, on-site teaching.
[0003] Traditional solutions use a single FPGA pulse to synchronously trigger multiple light field cameras, assuming that all photosensitive sensors have completely identical exposure times. This fails to account for inherent shutter opening delays, ADC readout delays, and individual differences in transmission cable delays. The actual exposure timing deviation across multiple cameras can reach tens to hundreds of microseconds. This timing deviation translates into sub-pixel offset at the light field sub-viewpoint, resulting in ghosting at the edges of the image, parallax abrupt changes, and image jitter after holographic reconstruction at the receiver. Software post-processing alignment compensation is insufficient in accuracy, computationally expensive, and cannot eliminate the original timing deviation. Furthermore, conventional GPS / PPS single-channel timing modes are prone to clock drift after forwarding through a local area network switch, leading to continuous degradation of synchronization consistency over long-term operation.
[0004] The computational complexity of encoding light field texture streams, depth streams, audio streams, and interactive command streams varies significantly, and the encoding completion times of the four data streams acquired in the same trigger cycle are not synchronized. Existing technologies use a fixed-wait slowest stream packaging mechanism, introducing a fixed queuing delay; or they directly mix-and-match packaging, causing timing disorder at the receiving end. At the same time, a general global uniform compression strategy compresses the teacher, teaching aids foreground, blackboard, and wall background at the same rate, which either puts enormous pressure on bandwidth and causes a surge in synchronization delay, or results in the loss of foreground holographic details, making it unsuitable for the content characteristics of teaching scenarios.
[0005] Existing WebRTC and SRT low-latency transmission synchronization models assume equal uplink and downlink unidirectional transmission delays. However, cross-city fiber optic and 5G URLLC heterogeneous links generally exhibit asymmetric uplink and downlink delays, requiring excessive redundancy depth in the receiver's synchronization buffer, thus preventing further reduction in end-to-end latency. Conventional congestion control strategies are not designed for the layered characteristics of the optical field, which can easily lead to foreground keyframe loss, synchronization queue overflow, and holographic image stuttering and tearing during network congestion.
[0006] The existing synchronization strategy only verifies the logical alignment of the timestamp in the header of the data packet before sending it directly into the rendering pipeline. It does not perform hardware phase-locked binding of the light field decoding completion timing, holographic rendering operation phase, spatial light modulator (SLM) refresh clock, and naked-eye holographic screen VSync vertical synchronization signal. As a result, the rendering time fluctuates randomly, the SLM refresh phase shifts, and the timing of multiple spliced holographic screens is inconsistent. Dynamic holographic images are prone to screen tearing, motion blur, and dizziness. The synchronization error of multi-stream audio and video exceeds the human eye's perception threshold.
[0007] The existing remote interaction mode matches the previously rendered light field frames after the student's operation command is uploaded, which is a post-correction mode. There is an inherent time difference between the interactive action and the holographic image. There is no command look-ahead scheduling mechanism. The synchronization error of teaching interactions such as 3D model rotation, sectioning, disassembly, and annotation is large, which cannot meet the needs of high-precision on-site interactive training.
[0008] In summary, existing light field remote holographic teaching systems suffer from systemic synchronization defects across the entire chain, including acquisition synchronization, encoding and encapsulation, transmission compensation, rendering output, and two-way interaction. End-to-end latency is generally higher than 350ms, multi-stream synchronization deviation exceeds the standard, holographic imaging quality is poor, and the interaction feels disjointed, making it difficult to scale up for professional immersive remote teaching scenarios. Summary of the Invention
[0009] In view of this, the purpose of this invention is to overcome the shortcomings of the prior art and provide a remote holographic on-site teaching system based on light field imaging and its low-latency synchronization method.
[0010] To achieve the above objectives, one of the present invention provides the following technical solution: A remote holographic on-site teaching system based on light field imaging includes: The system includes a global nanosecond-level dual-redundant timing management unit, a high-precision hardware synchronous acquisition layer, a content-aware light field preprocessing layered encoding layer, a latency-asymmetric adaptive transmission synchronization layer, a rendering phase-VSync linkage multi-stream alignment and restoration layer, a bidirectional interactive forward-looking closed-loop synchronization control layer, and a global latency monitoring adaptive correction module. The global nanosecond-level dual-redundant timing management unit is deployed on two edge servers, equipped with a Beidou 1PPS timing module and an IEEE 1588v2 transparent clock module, which divides the dual clock domains and periodically corrects clock drift, and outputs a unified timing reference. The high-precision hardware synchronous acquisition layer is configured with a multi-channel light field camera array, and integrates other types of acquisition devices, FPGA synchronous controller and time delay storage register. It relies on the inherent time delay compensation of the device to output trigger pulses, synchronously acquires various raw data and embeds timing information. The content-aware light field preprocessing layer is deployed on the edge GPU server on the main speaker side. It performs foreground and background differential encoding on the image, monitors the encoding time of multiple data streams, and encapsulates the data after dynamic elastic caching alignment. Interactive commands are transmitted in-band piggybacking. The delay-asymmetric adaptive transmission synchronization layer adopts 5GuRLLC and fiber optic dual links to calculate uplink and downlink one-way delays in real time, and combines network jitter classification configuration buffer, congestion control and differentiated forward error correction strategies. The rendering phase-VSync linkage multi-stream alignment restoration layer is deployed on the holographic terminal in the classroom to perform timing calibration on multiple data streams and hardware phase lock the decoding, rendering, spatial light modulator, and holographic screen refresh signals to achieve synchronous output of holographic images. The bidirectional interactive forward-looking closed-loop synchronous control layer receives remote interactive commands, predicts command timing and schedules them in advance, and forms a timing self-correction closed loop by combining interpolation compensation and error statistics. The global latency monitoring and adaptive correction module performs latency monitoring across the entire link layer, dynamically adjusts system operating parameters based on preset thresholds, and stores timing logs.
[0011] More preferably, the formula for calculating the total inherent time delay of the single-path light field camera is: ,in For the first Road camera shutter opening delay For ADC analog-to-digital conversion readout delay, This refers to the data transmission delay.
[0012] Even better, the global nanosecond-level dual-redundant timing management unit automatically completes clock frequency offset and phase drift correction every 200ms.
[0013] More preferably, the delay-asymmetric adaptive transmission synchronization layer adjusts according to network jitter. Three-level dynamic buffering strategy: low jitter Employs a minimal buffer mode; Dynamically adjusting cache depth to achieve timing compensation; Background redundant frames are discarded based on content priority.
[0014] Even better, the rendering phase-VSync linked multi-stream alignment restoration layer only starts the holographic reconstruction operation when it matches the rising edge of the holographic screen VSync vertical synchronization at the moment of decoding completion, and the audio stream is set to a ±15ms human ear perception error tolerance range.
[0015] Even better, the two-way interactive forward-looking closed-loop synchronization control layer uses local optical field deformation interpolation to correct the delayed commands, and calculates the two-way interactive synchronization error every 500ms and adjusts the system synchronization parameters in a closed loop.
[0016] The second aspect of this invention provides the following technical solution: A low-latency synchronization method, applied to the aforementioned remote holographic on-site teaching system based on light field imaging, includes the following steps: Step S1: Construct a global dual-redundant nanosecond clock reference, divide the clock domain into two levels and continuously correct clock drift; deploy the Beidou 1PPS timing module on the main speaker and the listener to generate global GTS absolute time, deploy the IEEE1588v2 transparent clock on the local area network to eliminate switch forwarding offset, divide the nanosecond-level acquisition clock domain and the microsecond-level network system clock domain, and correct the clock frequency offset and phase offset every 200ms; Step S2: Offline calibration of inherent camera delay at the acquisition end + online dynamic compensation synchronous acquisition; offline calibration of the total inherent delay of each light field camera and storage; FPGA generates advance compensation trigger pulses for each camera based on the global reference time, synchronously triggering all light field cameras, depth sensors, audio arrays, and gesture capture modules to acquire data synchronously; each frame of raw data is embedded with a GTS acquisition timestamp to control the global acquisition synchronization deviation within 20ns; Step S3: Sender performs content-aware layered light field compression + heterogeneous stream dynamic pre-alignment and packaging; performs differentiated light field compression based on depth threshold segmentation of foreground and background regions; monitors the encoding time of four heterogeneous data streams (light field texture stream, depth data stream, audio stream, and interactive command stream) in real time, sets an elastic buffer based on the estimated completion time of the slowest stream, aligns the four data streams, and then encapsulates and packages them uniformly, with interactive command streams being piggybacked into the packaging to control the packaging queuing delay; Step S4: Asymmetric uplink and downlink delay calculation at the transport layer + adaptive buffer synchronous transmission; periodic probing to solve independent unidirectional uplink and downlink delays, constructing a dynamic synchronous buffer model based on network jitter and delay difference, and hierarchical matching buffer strategy; optical field hierarchical congestion control combined with differentiated FEC configuration, seamless switching of dual-link timing to transmit encoded data packets; Step S5: Refined timing calibration of multi-stream receiver + hardware phase-locked synchronization restoration of rendering and display; based on the GTS alignment of the data packets, the timing of the four heterogeneous data streams is buffered, interpolated, or expired frames are discarded to compensate for timing deviations; hardware phase-locking is performed on the decoding timing, rendering phase, SLM refresh, and holographic screen VSync signal, and only the synchronization phase is used to start the holographic reconstruction rendering output, and unified timing correction is performed on multiple spliced large screens; Step S6: Two-way interactive forward-looking timing closed-loop synchronization; student interaction commands are sent back in reverse, and the main speaker predicts the timing position of the command and inserts it into the rendering queue in advance for synchronization; local light field interpolation is used to compensate for slight command lag; the two-way synchronization error is periodically counted and the system synchronization parameters are adjusted in the closed loop. Step S7: End-to-end segmented latency monitoring and adaptive self-correction; segmented statistics of latency at each level, triggering multi-level latency threshold control strategies, dynamically optimizing encoding, buffering, and transmission parameters, and logging timing data for long-term clock self-correction.
[0017] Even better, in step S3, the maximum threshold for setting the elastic cache is 8ms.
[0018] Even better, in step S5, the maximum allowed synchronization deviation threshold for the four data streams is 20ms.
[0019] Even better, in step S2, the synchronization deviation of the global acquisition is controlled within 20ns.
[0020] The beneficial effects of this invention are as follows: This invention achieves ultra-high precision acquisition and synchronization within 20ns by combining BeiDou and IEEE1588v2 dual redundant clocks with hardware compensation for the camera's inherent latency, eliminating holographic ghosting at its source. It reduces transmission bandwidth by over 60% through foreground-background differentiated layered compression, and controls queuing latency by combining heterogeneous stream dynamic elastic pre-alignment packaging. It breaks the traditional assumption of symmetrical transmission latency by constructing an adaptive buffer model, and improves transmission robustness in cross-city and 5G weak network scenarios with layered congestion and differentiated FEC. At the receiving end, decoding, rendering, SLM refresh, and holographic screen VSync hardware phase-locked loops are bound together, controlling the four-stream synchronization deviation within 20ms, eliminating holographic image tearing, ghosting, and visual dizziness. Combined with end-to-end segmented latency monitoring and adaptive adjustment, the end-to-end total latency is ultimately stabilized within 180ms, meeting the low-latency and high-stability requirements for cross-campus holographic teaching and professional practical remote teaching such as medical anatomy. Attached Figure Description
[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 Block diagram of the remote holographic real-time teaching system of the present invention; Figure 2 A flowchart of the low-latency synchronization method of the present invention.
[0022] In the diagram: 1-Global nanosecond-level dual-redundant timing management unit; 2-High-precision hardware synchronous acquisition layer; 3-Content-aware light field preprocessing layered encoding layer; 4-Delay asymmetric adaptive transmission synchronization layer; 5-Rendering phase-VSync linkage multi-stream alignment and restoration layer; 6-Two-way interactive forward-looking closed-loop synchronization control layer. Detailed Implementation
[0023] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0024] like Figure 1 As shown, this invention provides a remote holographic on-site teaching system based on light field imaging, which is divided into a five-layer architecture, including: a global nanosecond-level dual-redundant timing management unit 1; a high-precision hardware synchronous acquisition layer 2; a content-aware light field preprocessing layered encoding layer 3; a time-delay asymmetric adaptive transmission synchronization layer 4; a rendering phase-VSync linked multi-stream alignment and restoration layer 5; and a bidirectional interactive forward-looking closed-loop synchronization control layer 6. The global nanosecond-level dual-redundant timing management unit is deployed on the edge server of the main lecture hall and the edge server of the classroom, respectively. It includes the Beidou 1PPS nanosecond timing module, the optimized IEEE1588v2 transparent clock submodule, and the two-level clock domain mapping and calibration submodule. The Beidou 1PPS outputs the global reference time GTS, constructing a hardware acquisition nanosecond clock domain and a network transmission microsecond system clock domain. The IEEE1588v2 transparent clock penetrates the switch port to eliminate forwarding offset, automatically corrects the clock frequency offset and phase drift every 200ms, and outputs a unified timing reference to all modules of the system.
[0025] The high-precision hardware synchronous acquisition layer includes a ring light field camera array, a depth sensing module, a ring microphone array, a gesture interaction capture module, an FPGA synchronous trigger controller, and a camera inherent delay storage register; the FPGA synchronous trigger controller receives the reference time from the global timing management unit. Pre-calibrate the total inherent time delay of each light field camera offline. Stored in a register; FPGA for the first The camera outputs a lead compensation trigger pulse. This ensures that the actual exposure times of all cameras are strictly aligned; it synchronously acquires raw light field views, depth maps, audio samples, and gesture interaction data, with each frame's header embedding a GTS acquisition timestamp, device number, and single-channel compensation delay parameters; and it controls the global sensor synchronization timing deviation to within 20ns.
[0026] The content-aware light field preprocessing layered coding layer is deployed on the edge GPU server on the main speaker side. It includes a foreground and background segmentation submodule, a light field layered compression submodule, a four-way heterogeneous stream coding time monitoring submodule, a dynamic elastic pre-alignment packing submodule, and an in-band piggybacking packaging unit. Depth threshold segmentation distinguishes between the foreground (teacher's body, 3D teaching aids, and hand interaction areas) and the background (static walls, blackboard, and desks / chairs). The foreground light field fully preserves parallax information using AV1 high-bitrate fine encoding, while the background view is downsampled and frame-skipping compression reduces redundancy. Real-time monitoring of the encoding time of four streams—light field texture stream, depth stream, audio stream, and interactive command stream—is used. The estimated completion time of the slowest data stream in the current cycle is used as the encapsulation benchmark. Data streams that have been encoded in advance are dynamically aligned and encapsulated after short-term elastic buffering. Data packets carry the acquisition GTS, encoding completion timestamp, estimated one-way transmission latency, and stream type identifier. Interactive command packets use in-band piggyback encapsulation to reduce protocol overhead, and the maximum threshold for the elastic buffer is set to 8ms to suppress packet queuing latency.
[0027] The delay-asymmetric adaptive transmission synchronization layer includes a 5G URLLC / fiber dual-link transmission unit, a real-time uplink and downlink unidirectional delay measurement submodule, an adaptive dynamic synchronization buffer control submodule, an optical field layered differentiated congestion control submodule, and a forward error correction differentiated configuration unit; it periodically sends probe packets to solve for uplink unidirectional delay. Downlink one-way delay Breaking the assumption of symmetric round-trip time delay; based on real-time network jitter A dynamic buffering calculation model is constructed, which includes basic jitter compensation and latency asymmetry correction terms, and a hierarchical buffering strategy is configured. In the low jitter mode, the buffer is aligned with minimal buffering; in the medium jitter mode, the buffer depth is dynamically adjusted for timing compensation; and in the high jitter mode, background redundant frames are discarded first to avoid queue overflow. Congestion control sets high transmission priority for the foreground stream and actively reduces the bit rate of the background stream to alleviate congestion. Audio and interactive commands are configured with high-redundancy FEC, while the background light field is configured with low-redundancy FEC. During seamless switching between dual links, the data packet sequence number and GTS timing are continued, and the switching timing error is less than 15ms.
[0028] The rendering phase-VSync linkage multi-stream alignment restoration layer is deployed on the holographic terminal in the classroom. It includes a multi-stream fine timing calibration submodule, a rendering pipeline phase-locked unit, an SLM spatial light modulator synchronization control unit, a holographic screen VSync hardware synchronization unit, a Gaussian sputtering lightweight holographic reconstruction engine, and a multi-screen splicing timing correction submodule. The timing deviation of the four streams is checked based on the data packet GTS. The leading frame is flexibly buffered and waited, the slightly lagging frame is interpolated in the time domain to complete, and the expired frame exceeding the threshold is smoothly discarded. The maximum synchronization deviation of the four streams is constrained to 20ms. The timing of holographic decoding completion, rendering start phase, SLM refresh clock and holographic screen VSync rising edge are hardware phase-locked, and single-frame holographic rendering is started only in the VSync synchronization phase. The reconstruction engine's fragmented parallel operation is bound to the global clock beat to suppress the fluctuation of rendering time. Multiple holographic splicing screens are allocated a unified global refresh synchronization signal to eliminate splicing timing misalignment and brightness unevenness.
[0029] The two-way interactive forward-looking closed-loop synchronization control layer includes a reverse interaction command feedback unit, an interaction timing prediction and insertion submodule, a local light field deformation interpolation compensation unit, and a two-way time delay closed-loop statistical correction unit. Student-side touch, gesture, and 3D model operation commands are fed back with local timestamps. The lecturer-side timing control unit predicts the command matching GTS timing position and inserts the interaction modification command in advance into the corresponding light field rendering queue to achieve synchronous effect between the command and the corresponding light field frame. When the command is slightly delayed, local light field deformation interpolation is used to correct the teaching aid posture. The total synchronization error of the two-way interaction is periodically calculated, and the buffer depth and encoding bitrate parameters at both ends are adjusted to form a timing negative feedback self-correction closed loop, which is suitable for teaching interaction scenarios such as 3D sectioning, assembly and disassembly, and blackboard annotation.
[0030] The global latency monitoring and adaptive correction module performs segmented latency tracking and statistics on the acquisition layer, encoding layer, transmission layer, and rendering layer, and sets a three-level latency threshold strategy. When the total latency exceeds the limit, it automatically performs view clipping, background bitrate reduction, compression of receiver buffer depth, and switching to low-latency transmission parameters. The synchronization deviation log is persistently stored for long-term clock drift correction and parameter self-learning optimization.
[0031] like Figure 2 As shown, this invention provides a low-latency synchronization method for remote holographic on-site teaching based on light field imaging, comprising the following steps: Step S1: Construct a global dual-redundant nanosecond clock reference, divide the clock domain into two levels and continuously correct clock drift; deploy the Beidou 1PPS timing module on the main speaker and the listener to generate global GTS absolute time, deploy the IEEE1588v2 transparent clock on the local area network to eliminate switch forwarding offset, divide the nanosecond-level acquisition clock domain and the microsecond-level network system clock domain, and correct the clock frequency offset and phase offset every 200ms.
[0032] Step S2: Offline calibration of the inherent delay of the acquisition camera + online dynamic compensation synchronous acquisition; offline calibration of the total inherent delay of each light field camera and storage; FPGA generates advance compensation trigger pulses for each camera based on the global reference time, synchronously triggering all light field cameras, depth sensors, audio arrays, and gesture capture modules to acquire data synchronously; each frame of raw data is embedded with a GTS acquisition timestamp to control the global acquisition synchronization deviation within 20ns.
[0033] Step S3: Sender performs content-aware layered light field compression + dynamic pre-alignment and packaging of heterogeneous streams; performs differentiated light field compression based on depth threshold segmentation of foreground and background regions; monitors the encoding time of the four heterogeneous data streams in real time, sets an elastic buffer based on the estimated completion time of the slowest stream, aligns the four data streams, and then encapsulates and packages them uniformly, with the interactive instruction packet being piggybacked into the encapsulation, controlling the packaging queuing delay.
[0034] Step S4: Asymmetric uplink and downlink delay measurement at the transport layer + adaptive buffer synchronous transmission; periodic probing to solve for independent unidirectional uplink and downlink delays, constructing a dynamic synchronous buffer model based on network jitter and delay difference, and hierarchical matching buffer strategy; optical field hierarchical congestion control combined with differentiated FEC configuration, and seamless switching of dual-link timing to transmit encoded data packets.
[0035] Step S5: Refined timing calibration of multiple streams at the receiving end + hardware phase-locked synchronization restoration of rendering and display; based on the GTS alignment of the data packets, the timing deviation is buffered, interpolated, or expired frames are discarded; the decoding timing, rendering phase, SLM refresh, and holographic screen VSync signal are hardware phase-locked, and only the synchronized phase is used to start the holographic reconstruction rendering output, and the timing of multiple spliced large screens is uniformly corrected to eliminate holographic tearing and ghosting problems.
[0036] Step S6: Two-way interactive forward-looking timing closed-loop synchronization; student interaction commands are sent back in reverse, and the main speaker predicts the timing position of the command and inserts it into the rendering queue in advance to take effect synchronously; local light field interpolation is used to compensate for slight delays in commands; the two-way synchronization error is periodically counted, and the system synchronization parameters are adjusted in the closed loop to continuously optimize the interaction synchronization accuracy.
[0037] Step S7: End-to-end segmented latency monitoring and adaptive self-correction; segmented statistics of latency at each level, triggering multi-level latency threshold control strategies, dynamically optimizing encoding, buffering, and transmission parameters, and logging timing data for long-term clock self-correction.
[0038] The following is a specific example: 1. Hardware configuration of the main lecture classroom acquisition terminal Ring-shaped light field camera array: 10-channel microlens structure 4K@60fps light field camera, horizontal total field of view 135°, pitch angle ±30°, focusing on the teacher's teaching area; equipped with 3-channel TOF depth camera; 8-channel ring microphone array; 1 set of infrared gesture capture module; 1 FPGA synchronization control board; 1 set of Beidou 1PPS timing module; 1 IEEE1588v2 clock slave module; 1 edge GPU server (RTX A8000); 5G industrial gateway + gigabit fiber optic dual-link routing equipment.
[0039] 2. Hardware configuration of the classroom playback terminal Same specifications Beidou timing module, IEEE1588v2 clock slave module; receiver edge GPU server; spatial light modulator (SLM) module, naked-eye holographic splicing screen (3-screen); VSync synchronization control board; interactive touch terminal, gesture acquisition unit; dual-link receiving gateway.
[0040] 3. Networking Mode: Campuses within the same city use gigabit fiber optic leased lines, while cross-city campuses deploy 5G uRLLC low-latency slicing networks, with basic one-way latency ranging from 30 to 100ms.
[0041] 4. Specific Implementation Steps Step 1 Global Clock Deployment and Calibration The main and listening edge servers are connected to the Beidou 1PPS module to output nanosecond-level GTS absolute time; the local area network switch is enabled with IEEE1588v2 transparent clock mode, and the port embeds forwarding timestamps to eliminate store-and-forward offset; the clock correction period is set to 200ms, and the clock frequency offset and phase deviation are measured in real time and automatically compensated to suppress clock drift during long-term operation; the acquisition side is divided into a 1ns resolution hardware clock domain and the host side is divided into a 1μs system clock domain, and a periodic mapping conversion relationship is established.
[0042] Step 2: Camera inherent time delay calibration and synchronous acquisition (1) Offline calibration: A checkerboard dot matrix calibration board is set up, and each light field camera is triggered individually in sequence to measure the shutter opening delay of a single camera. ADC readout delay Data cable transmission delay Calculate the total inherent delay (2) Online synchronous triggering: The FPGA receives the global reference exposure time. , for the Road camera output trigger time All cameras are exposed synchronously under the control of compensation pulses; depth camera, microphone, and gesture sensor are synchronously acquired using the same time delay compensation logic; the header of each data frame is written with the current GTS acquisition timestamp, device number, and compensation time delay parameters, and the timing deviation of the global acquisition is stable at ≤18ns.
[0043] Step 3: Layered compression and dynamic packaging at the sending end (1) Foreground and background segmentation: Based on depth threshold segmentation, the area with a depth of 0~1.5m is identified as the foreground (teacher, hands, desktop teaching aids), and the area beyond 1.5m is identified as the static background; (2) Differentiated encoding: AV1 hardware encoding of the complete foreground light field view, high bit rate to preserve parallax details; 1 frame is extracted every 4 frames of the background view and downsampled and compressed from adjacent viewpoints, reducing the overall data volume by 62% after compression; the depth stream uses dedicated inter-frame compression, the audio uses AAC low-latency encoding, and the interactive instructions are encapsulated in binary packets. (3) Dynamic pre-alignment packing: real-time acquisition of single-frame encoding time of four data streams Take the maximum value As a packaging benchmark, pre-encoded data is stored in an elastic cache queue, with cache duration matching. The buffer limit is set to 8ms to avoid latency accumulation; the acquisition GTS, encoding completion time, estimated one-way latency, and stream type identifier are written into the header of the packaged data packet; interactive commands are embedded into the data packet using in-band piggybacking and are not separately encapsulated in the packet header; the packaged data packet is then sent to the transport layer.
[0044] Step 4: Delay-Asymmetric Adaptive Transmission Synchronization (1) One-way delay calculation: Insert probe packets every 100ms to calculate the uplink delay from lecturer to listener. Listening to the lecture → Downlink delay of the main speaker Real-time recording of link jitter ; (2) Dynamic buffer calculation: Set buffer baseline The target buffer depth at the receiver is obtained by superimposing the uplink and downlink delay asymmetry correction term. (3) Hierarchical buffering strategy: ① Low vibration Minimal buffer mode, microsecond-level timing alignment only, queuing latency ≤12ms; ② Medium jitter : Dynamically adjust the buffer depth, buffer the leading frame to wait for the lagging frame of the same source, and use temporal linear interpolation to compensate for small timing deviations. ③ High jitter Priority-based frame dropping strategy: discard redundant background frames, retain key foreground light field frames, and prevent queue overflow. (4) Layered congestion and FEC: Foreground stream bandwidth has the highest priority, and only the background stream bit rate is reduced when congestion occurs; audio and interactive command FEC redundancy ratio is 25%, and background light field FEC redundancy ratio is 8%; 5G / fiber dual link real-time monitoring latency, and the sequence number and GTS timing are continued during handover, with a handover timing error ≤14ms.
[0045] Step 5: Receiver multi-stream alignment + rendering and display phase-locked synchronization (1) Multi-stream timing calibration: Extract the common GTS of data packets and calculate the arrival timing difference of the four streams; if the timing difference is less than 20ms: buffer the leading frame and wait, and interpolate the slightly lagging frame to make up the difference; if the timing difference is greater than 20ms, determine the expired frame and discard it smoothly; set the audio to ±15ms human ear perception fault tolerance range to optimize the auditory synchronization experience; (2) Hardware phase-locked synchronization: The rendering engine decoding completion signal, rendering start phase signal, SLM refresh clock, and holographic screen VSync vertical synchronization signal are connected to the synchronization control board hardware linkage; a frame of holographic rendering is started only when the decoding data is ready and aligned with the rising edge of VSync; the start time of the next frame rendering is predicted in real time, and the decoding thread is pre-scheduled to load the corresponding timing data packet to eliminate the delay between decoding and rendering. (3) Holographic reconstruction output: The parallel reconstruction of Gaussian sputtering is adopted, and the operation cycle is bound to the global clock to suppress the fluctuation of rendering time. The three-panel holographic splicing screen is allocated a unified global synchronous refresh signal, and the geometric and timing correction is performed screen by screen to eliminate the timing offset of splicing seams and uneven brightness, and output continuous and tear-free naked-eye holographic images.
[0046] Step 6: Two-way interactive forward-looking closed-loop synchronization (1) Student-side gesture, touch, and 3D model rotation / section commands are encapsulated with local timestamps and sent back to the lecturer; (2) The timing control unit of the main end predicts and matches the timing position of GTS based on the instruction transmission delay, and inserts the interactive modification instruction in advance into the corresponding timing light field rendering queue. The instruction takes effect synchronously with the corresponding light field frame. (3) In scenarios where instructions are slightly delayed, local light field deformation interpolation is used to correct the posture of the three-dimensional teaching aids and weaken the visual effect of operation delay. (4) The total synchronization error of bidirectional interaction is counted every 500ms. If the error exceeds the standard, the buffer depth, encoding rate and FEC ratio at both ends are automatically adjusted to form a time-series negative feedback self-correction closed loop, which is suitable for teaching interaction scenarios such as dissection, parts assembly and remote annotation.
[0047] Step 7: End-to-end latency adaptive monitoring and correction Latency monitoring points are set at four levels: acquisition, encoding, transmission, and rendering, to statistically analyze segmented latency and total end-to-end latency in real time. Three thresholds are set: Level 1 warning for total latency > 180ms, Level 2 adjustment for total latency > 220ms, and Level 3 emergency strategy for total latency > 280ms. Upon triggering, background view clipping, background bitrate reduction, receiver buffer depth shrinking, and switching to low-latency transmission parameters are automatically executed. Timing deviation data is stored locally in logs for long-term clock drift correction and system parameter self-learning optimization.
[0048] 5. Actual Measured Results Data The actual test results of the cross-city 500km network in this embodiment are as follows: Average camera acquisition synchronization deviation: 17.2 ns; Maximum synchronization deviation for the four streams: light field / depth / audio / interaction: 17.3ms; Average end-to-end latency across the entire link: 156ms, maximum 175ms, consistently below the design limit of 180ms; The overall transmission bandwidth of the optical field decreased by 62.7% compared to the original uncompressed scheme; With a weak network jitter of 50ms, there is no screen tearing or continuous stuttering, interactive operations are synchronized and smooth, and holographic imaging is free of ghosting and afterimages, meeting the requirements for immersive remote teaching.
[0049] Compared with the prior art, the advantages of this embodiment are: 1. Synchronization deviation of multi-channel sensor acquisition <20ns; synchronization deviation of light field, depth, audio, and interaction four streams <20ms; end-to-end controllable total latency ≤180ms, significantly better than the existing conventional latency level of over 350ms; 2. Naked-eye holography is free of ghosting, tearing, and dynamic afterimages; it offers continuous visibility with a large 120° parallax, and provides long-term viewing without dizziness. 3. It can be stably deployed across cities, 5G uRLLC, and ordinary campus broadband, and has strong robustness against jitter and packet loss in weak networks. 4. Content-aware layered compression reduces the overall transmission bandwidth by more than 60%, and hardware encoding and decoding parallel operation meets the real-time processing requirement of 60fps; 5. It enables cross-campus holographic teaching by renowned teachers, remote medical anatomy teaching, mechanical structure assembly training, and immersive 3D experimental demonstrations, balancing high-quality educational resources and solving the pain points of insufficient intuitiveness and lagging interaction in practical remote teaching.
[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A remote holographic on-site teaching system based on light field imaging, characterized in that, It includes a global nanosecond-level dual-redundant timing management unit, a high-precision hardware synchronous acquisition layer, a content-aware light field preprocessing layered encoding layer, a latency-asymmetric adaptive transmission synchronization layer, a rendering phase-VSync linkage multi-stream alignment and restoration layer, a bidirectional interactive forward-looking closed-loop synchronization control layer, and a global latency monitoring adaptive correction module. The global nanosecond-level dual-redundant timing management unit is deployed on two edge servers, equipped with a Beidou 1PPS timing module and an IEEE 1588v2 transparent clock module, which divides the dual clock domains and periodically corrects clock drift, and outputs a unified timing reference. The high-precision hardware synchronous acquisition layer is configured with a multi-channel light field camera array, and integrates other types of acquisition devices, FPGA synchronous controller and time delay storage register. It relies on the inherent time delay compensation of the device to output trigger pulses, synchronously acquires various raw data and embeds timing information. The content-aware light field preprocessing layer is deployed on the edge GPU server on the main speaker side. It performs foreground and background differential encoding on the image, monitors the encoding time of multiple data streams, and encapsulates the data after dynamic elastic caching alignment. Interactive commands are transmitted in-band piggybacking. The delay-asymmetric adaptive transmission synchronization layer adopts 5GuRLLC and fiber optic dual links to calculate uplink and downlink one-way delays in real time, and combines network jitter classification configuration buffer, congestion control and differentiated forward error correction strategies. The rendering phase-VSync linkage multi-stream alignment restoration layer is deployed on the holographic terminal in the classroom to perform timing calibration on multiple data streams and hardware phase lock the decoding, rendering, spatial light modulator, and holographic screen refresh signals to achieve synchronous output of holographic images. The bidirectional interactive forward-looking closed-loop synchronous control layer receives remote interactive commands, predicts command timing and schedules them in advance, and forms a timing self-correction closed loop by combining interpolation compensation and error statistics. The global latency monitoring and adaptive correction module performs latency monitoring across the entire link layer, dynamically adjusts system operating parameters based on preset thresholds, and stores timing logs.
2. The remote holographic on-site teaching system based on light field imaging according to claim 1, characterized in that, The formula for calculating the total inherent time delay of the single-path light field camera is as follows: ,in For the first Road camera shutter opening delay, For ADC analog-to-digital conversion readout delay, This refers to the data transmission delay.
3. The remote holographic on-site teaching system based on light field imaging according to claim 1, characterized in that, The global nanosecond-level dual-redundant timing management unit automatically completes clock frequency offset and phase drift correction every 200ms.
4. The remote holographic on-site teaching system based on light field imaging according to claim 1, characterized in that, The time-delay asymmetric adaptive transmission synchronization layer adjusts according to network jitter. Three-level dynamic buffering strategy: low jitter Employs a minimal buffer mode; Dynamically adjust cache depth to achieve timing compensation; Background redundant frames are discarded based on content priority.
5. The remote holographic on-site teaching system based on light field imaging according to claim 1, characterized in that, The rendering phase-VSync linked multi-stream alignment restoration layer only starts the holographic reconstruction operation when the decoding is completed and the rising edge of the holographic screen VSync vertical synchronization is matched. The audio stream is set with a ±15ms human ear perception error tolerance range.
6. The remote holographic on-site teaching system based on light field imaging according to claim 1, characterized in that, The bidirectional interactive forward-looking closed-loop synchronization control layer uses local optical field deformation interpolation to correct lagging commands, and calculates the bidirectional interactive synchronization error every 500ms and adjusts the system synchronization parameters in a closed loop.
7. A low-latency synchronization method, characterized in that, The method applied to the remote holographic on-site teaching system based on light field imaging as described in any one of claims 1-6 includes the following steps: Step S1: Construct a global dual-redundant nanosecond clock reference, divide the clock domain into two levels and continuously correct clock drift; deploy the Beidou 1PPS timing module on the main speaker and the listener to generate global GTS absolute time, deploy the IEEE1588v2 transparent clock on the local area network to eliminate switch forwarding offset, divide the nanosecond-level acquisition clock domain and the microsecond-level network system clock domain, and correct the clock frequency offset and phase offset every 200ms; Step S2: Offline calibration of inherent delay of the acquisition camera + online dynamic compensation synchronous acquisition; offline calibration of the total inherent delay of each light field camera and storage; FPGA generates advance compensation trigger pulses for each camera based on the global reference time, and synchronously triggers all light field cameras, depth sensors, sound pickup arrays and gesture capture modules to acquire data synchronously. Each frame of raw data is embedded with a GTS acquisition timestamp, keeping the global acquisition synchronization deviation within 20ns; Step S3: Sender performs content-aware layered light field compression + heterogeneous stream dynamic pre-alignment and packaging; performs differentiated light field compression based on depth threshold segmentation of foreground and background regions; monitors the encoding time of four heterogeneous data streams (light field texture stream, depth data stream, audio stream, and interactive command stream) in real time, sets an elastic buffer based on the estimated completion time of the slowest stream, aligns the four data streams, and then encapsulates and packages them uniformly, with interactive command streams being piggybacked into the packaging to control the packaging queuing delay; Step S4: Asymmetric uplink and downlink delay calculation at the transport layer + adaptive buffer synchronous transmission; periodic probing to solve independent unidirectional uplink and downlink delays, constructing a dynamic synchronous buffer model based on network jitter and delay difference, and hierarchical matching buffer strategy; optical field hierarchical congestion control combined with differentiated FEC configuration, seamless switching of dual-link timing to transmit encoded data packets; Step S5: Refined timing calibration of multi-stream at the receiving end + hardware phase-locked loop synchronization restoration for rendering and display; Based on the timing of the four heterogeneous data streams aligned by the data packet GTS, timing deviations are buffered, interpolated, or expired frames are discarded; the decoding timing, rendering phase, SLM refresh, and holographic screen VSync signal are hardware phase-locked, and only the synchronous phase is used to start the holographic reconstruction rendering output, and the timing of multiple spliced large screens is uniformly corrected. Step S6: Two-way interactive forward-looking timing closed-loop synchronization; student interaction commands are sent back in reverse, and the main speaker predicts the timing position of the command and inserts it into the rendering queue in advance for synchronization; local light field interpolation is used to compensate for slight command lag; the two-way synchronization error is periodically counted and the system synchronization parameters are adjusted in the closed loop. Step S7: End-to-end segmented latency monitoring and adaptive self-correction; The system segments and calculates latency at each level, triggers multi-level latency threshold adjustment strategies, dynamically optimizes encoding, buffering, and transmission parameters, and logs time-series data for long-term clock self-calibration.
8. The low-latency synchronization method for remote holographic on-site teaching based on light field imaging according to claim 7, characterized in that, In step S3, the maximum threshold for setting the elastic cache is 8ms.
9. The low-latency synchronization method according to claim 7, characterized in that, In step S5, the maximum allowed synchronization deviation threshold for the four data streams is 20ms.
10. The low-latency synchronization method according to claim 7, characterized in that, In step S2, the synchronization deviation of the global acquisition is controlled within 20ns.