Low-delay real-time control method and device for teaching screen projection equipment
Through frame-by-frame analysis of teacher behavior and adaptive transmission of network state, the delay and synchronization problems of teaching screen projection equipment in complex network environments are solved, and the low-latency and high-responsive teaching screen projection effect is achieved, improving the efficiency and continuity of teaching interaction.
Patent Information
- Application Number
- CN202510936890.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing teaching screen projection equipment has problems such as screen delay, frame rate drop, content out of synchronization and image screen when teachers operate quickly. It is difficult to achieve true intelligent scheduling and prediction, and cannot adapt to changes in complex network states, resulting in poor continuity of teaching content.
By collecting real-time behavior monitoring images of teachers to analyze the action behaviors on a frame-by-frame basis, high-level control instructions are generated, content complexity calculation and preload buffering decisions are performed, multi-dimensional network state is identified and adaptive area transmission selection is performed, local priority pixel stream filling and lagging transmission pixel stream rendering is performed, and intelligent real-time screen projection control model is finally built.
It realizes non-invasive monitoring and real-time recognition of teachers' natural teaching actions, improves the semantic understanding ability of system response, optimizes control delay and information redundancy, ensures timely display of key teaching areas, reduces the real-time transmission requirements of non-core areas, improves the system's energy saving and transmission agility, and ensures rapid response and complete display of teaching images.
Smart Images

Figure CN120499424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of screen projection control, and in particular to a low-latency real-time control method and device for teaching screen projection equipment. Background Art
[0002] With the deep integration of information technology and education, digital teaching has gradually become the mainstream mode of classroom teaching. In particular, the use of teaching projection devices has increased significantly in scenarios such as multimedia teaching, smart classrooms, and remote learning. Teaching projection devices transmit and project the teacher's operating interface, teaching courseware, video resources, or interactive whiteboard content to the student's display terminal in real time, achieving synchronous visualization of teaching content, improving the efficiency of teaching interaction and classroom accessibility. Especially in the context of the epidemic, the surge in demand for online teaching and blended learning has further increased the requirements for the stability, real-time performance, and intelligence of teaching projection systems.
[0003] However, in practical applications, educational screen projection equipment still faces a number of technical bottlenecks. Among them, "low latency" and "real-time control" are key factors affecting the teaching experience. Traditional screen projection methods often rely on static frame synchronization mechanisms based on fixed-interval transmission or use common screen mirroring protocols (such as Miracast and AirPlay) for data mirroring. While these methods are generally feasible in ordinary home entertainment or office settings, they exhibit significant shortcomings in teaching scenarios. For example, when teachers engage in frequent activities such as rapidly writing on the blackboard, turning pages, explaining key points, or switching between operations, problems such as image delay, frame rate drops, content desynchronization, and even image distortion can occur, seriously impacting teaching continuity and maintaining student attention. Furthermore, existing educational screen projection systems typically rely on fixed content transmission structures and lack the ability to perceive and predict teacher actions. They often only passively push content based on changes in image content, ignoring the dynamic nature of the teacher's behavior during teaching, making it difficult to achieve truly intelligent scheduling and pre-emptive buffering. Furthermore, traditional systems have difficulty adapting to complex changes in network status during image transmission. In environments with bandwidth fluctuations, sudden delays, or severe packet loss, they cannot guarantee the continuous display of teaching content, which can easily lead to experience problems such as content freezes and synchronization failures. Therefore, a more intelligent real-time control method for screen projection is needed. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes a low-latency real-time control method and device for teaching screen projection equipment to solve at least one of the above technical problems.
[0005] To achieve the above object, the present invention provides a low-latency real-time control method for a teaching screen projection device, comprising the following steps: Step S1: The teacher's real-time behavior monitoring image is collected by the teaching terminal, and frame-by-frame action analysis is performed, and operation instructions are intelligently generated to obtain high-level control instructions; Step S2: Identify the teaching image content of the next frame based on the high-level control instruction, perform content complexity calculation and preload buffering decision, and build a preload image content buffering strategy; Step S3: identifying multi-dimensional network status indicators, and performing adaptive regional transmission selection on the teaching image content to obtain a local priority pixel stream and a delayed transmission pixel stream; Step S4: performing local filling processing on the local priority pixel stream based on the preloaded image content buffering strategy, and performing fast temporary screen rendering to obtain a local pixel rendering image; Step S5: performing secondary transmission rendering on the delayed transmission pixel stream, and performing global rendering reconstruction in combination with the local pixel rendering image to construct a global teaching rendering image; Step S6: Render the global teaching rendering image to complete self-correction, calculate the real-time frequency of teaching image switching, and perform pre-loaded frequency resonance optimization to build an intelligent real-time projection control model.
[0006] In this specification, a low-latency real-time control device for a teaching screen projection device is provided, which is used to execute the low-latency real-time control method for the teaching screen projection device as described above, including: The teaching behavior analysis module is used to collect real-time teacher behavior monitoring images based on the teaching terminal, perform frame-by-frame action analysis, and intelligently generate operation instructions to obtain high-level control instructions; The preloading buffer module is used to identify the teaching image content of the next frame based on high-level control instructions, perform content complexity calculation and preloading buffer decision-making, and build a preloading image content buffering strategy; An adaptive transmission module, configured to identify multi-dimensional network status indicators and perform adaptive regional transmission selection on the teaching image content to obtain a local priority pixel stream and a delayed transmission pixel stream; The local rendering module is used to perform local filling processing on the local priority pixel stream based on the preloaded image content buffering strategy and perform fast temporary screen rendering to obtain a local pixel rendering image; The global rendering reconstruction module is used to perform secondary transmission rendering on the delayed transmission pixel stream, and to perform global rendering reconstruction in combination with the local pixel rendering image to construct a global teaching rendering image; The same-frequency resonance optimization module is used to perform self-correction on the rendering of the global teaching rendering image, calculate the real-time frequency of teaching image switching, perform pre-loaded same-frequency resonance optimization, and build an intelligent real-time projection control model.
[0007] The beneficial effects of the present invention include: enabling non-invasive monitoring and real-time recognition of teachers' natural teaching movements without the need for additional equipment; frame-by-frame motion analysis technology accurately captures interactive teaching behaviors such as tapping, dragging, circling, and annotation; utilizing an intelligent command generation mechanism to abstract underlying physical actions into high-level teaching control intent (such as "page turning" and "region emphasis"), effectively improving the semantic understanding of system responses; avoiding the control delays and information redundancy caused by traditional event-level frame-by-frame data transmission, thereby optimizing the system control chain; and layering teaching images based on their criticality to prioritize the timely display of key teaching areas. Dynamically identifying network status (bandwidth, packet loss, latency, etc.) enables "content-selective" data compression, breaking through network bottlenecks. This reduces the need for immediate transmission of non-core areas, making the system more energy-efficient and agile. Local filling allows the teaching screen to quickly respond to teacher actions, avoiding the phenomenon of "content updated but the screen not moving." Even if the entire image is not yet rendered, the teacher's intended areas are prioritized for display, improving student comprehension efficiency. Local rendering frees up time for secondary pixel stream processing, making the overall image smoother and avoiding abrupt transitions. The system gradually completes the remaining areas, ultimately achieving a complete image display, ensuring that the panoramic teaching information is conveyed completely and accurately. Through interpolation and image reconstruction techniques, the content filling process is smooth and natural, reducing the user's perceived "rendering frame jumps." The rendering of non-critical content is deferred to idle periods to improve the performance of the main processing thread. Through timestamp identification and visual recall, the system can self-confirm, self-check, and self-adjust its own rendering status, establishing a closed-loop screen projection feedback mechanism. After analyzing the teacher's operating rhythm, the system rhythm dynamically adapts to it, achieving a consistent rhythm among "human, machine, and image." Over long-term operation, the system can gradually form a memory of the teacher's habitual rhythm, achieve personalized adaptation, and optimize the response speed and visual rhythm matching over the long term. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 This is a schematic flow chart of the steps of a low-latency real-time control method for a teaching screen projection device according to the present invention; Figure 2 Detailed implementation flow chart of step S1; Figure 3 Detailed implementation flow chart of step S2; Figure 4 Schematic diagram of the detailed implementation steps of step S3. DETAILED DESCRIPTION
[0009] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0010] This application example provides a low-latency real-time control method and device for teaching screen projection equipment. The execution subjects of the low-latency real-time control method and device for teaching screen projection equipment include but are not limited to: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. equipped with the system, which can be regarded as general computing nodes of this application. The data processing platform includes but is not limited to: at least one of an audio and image management system, an information management system, and a cloud data management system.
[0011] See also Figures 1 to 4 The present invention provides a low-latency real-time control method for a teaching screen projection device, comprising the following steps: Step S1: The teacher's real-time behavior monitoring image is collected by the teaching terminal, and frame-by-frame action analysis is performed, and operation instructions are intelligently generated to obtain high-level control instructions; Step S2: Identify the teaching image content of the next frame based on the high-level control instruction, perform content complexity calculation and preload buffering decision, and build a preload image content buffering strategy; Step S3: identifying multi-dimensional network status indicators, and performing adaptive regional transmission selection on the teaching image content to obtain a local priority pixel stream and a delayed transmission pixel stream; Step S4: performing local filling processing on the local priority pixel stream based on the preloaded image content buffering strategy, and performing fast temporary screen rendering to obtain a local pixel rendering image; Step S5: performing secondary transmission rendering on the delayed transmission pixel stream, and performing global rendering reconstruction in combination with the local pixel rendering image to construct a global teaching rendering image; Step S6: Render the global teaching rendering image to complete self-correction, calculate the real-time frequency of teaching image switching, and perform pre-loaded frequency resonance optimization to build an intelligent real-time projection control model.
[0012] In the embodiment of the present invention, see Figure 1 , is a flowchart of the steps of a low-latency real-time control method for a teaching screen projection device of the present invention. In this example, the steps of the low-latency real-time control method for a teaching screen projection device include: Step S1: The teacher's real-time behavior monitoring image is collected by the teaching terminal, and frame-by-frame action analysis is performed, and operation instructions are intelligently generated to obtain high-level control instructions; In this embodiment, the teaching terminal uses a front-facing HD camera or a depth camera to capture the teacher's frontal movements in real time. The camera configuration generally uses a 60fps frame rate and 1080p resolution to ensure no image blur or recognition delay during rapid movements. The captured data is a chronological stream of image frames, forming a continuous time series image input of the teacher's movements. The system then performs standardized preprocessing on the raw image frames, primarily including portrait region extraction, background filtering, brightness normalization, and motion compensation. This step uses a multi-scale YOLOv5 model to locate and identify the teacher's skeleton. The recognition results are mapped into a two-dimensional coordinate sequence of skeletal key points (such as shoulders, elbows, wrists, and knees), which serves as the basic input for subsequent action recognition. During the action analysis phase, the system uses a temporal graph neural network (ST-GCN) or a long short-term memory network (LSTM) to analyze these skeleton point sequences frame by frame. Recognition targets include over ten typical teaching actions, such as waving left / right (page turning), pointing to the screen (command selection), overlapping hands (pause command), and fast forward (play / resume). The system achieves over 93% recognition accuracy through supervised learning on a training set (consisting of approximately 12,000 action samples), supporting sub-second action switching responses. After an action is recognized, the system enters the "command mapping" phase. This phase matches the analysis results with the system's pre-set action library to generate basic commands (such as "NextSlide," "ZoomIn," and "PauseVideo"). High-level control commands are then derived by overlaying context, rhythm, and environmental conditions to create decisive commands. For example, if the system recognizes three consecutive right-waving gestures, it identifies this as a rapid jump command rather than a standard page turn and generates a high-level command, "JumpToSectionX." These high-level control commands are then sent in real time to the projection device module via a lightweight messaging channel (such as ZeroMQ or WebSocket), enabling intelligent response and display control of the teaching content. Experimental verification shows that the average processing time of the entire processing chain from image acquisition to command generation is less than 90ms, which effectively supports the needs of near-real-time teaching interaction and significantly improves the teaching projection equipment's ability to understand and respond to teachers' natural behaviors.
[0013] Step S2: Identify the teaching image content of the next frame based on the high-level control instruction, perform content complexity calculation and preload buffering decision, and build a preload image content buffering strategy; In this embodiment, after a high-level control instruction is generated, the system quickly matches the instruction with the content index table in the teaching resource library. For example, for the "NextSlide" instruction, the system retrieves the next PowerPoint presentation or teaching video frame and performs image analysis on the resource content. Teaching content can include diverse structural information such as text, charts, images, and embedded videos. For this purpose, the system constructs a content structure tree model, annotating the type, size, position, and hierarchical relationships of the content elements as the basis for complexity calculation. Next, content complexity calculation is implemented using a multi-factor fusion algorithm. Specifically, it includes three core metrics: image pixel density, layer depth, and the number of dynamic elements (such as embedded animations or video frames). In experiments, the system introduced a weighted complexity scoring function, in which pixel density accounts for 40%, layer depth accounts for 30%, and dynamic elements account for 30%. For example, a complex PowerPoint page with an image density of 220 DPI, a layer depth of 5 layers, and two GIF animations has an overall complexity score of 0.86 (out of a maximum score of 1), classifying it as "high complexity content." After determining the content complexity score, the system enters the preload buffering decision phase. The system dynamically adjusts the preload buffer window size and buffering priority based on the complexity score. In the experimental configuration, pages with a score above 0.75 trigger pre-loading within 50ms of the control command generation, allocating at least 512MB of local video memory for image and video frame pre-rendering. Content with a score below 0.45 uses a lazy loading mechanism, buffering only within the frame before presentation to conserve resources. The buffering strategy is managed by the Content Buffer Scheduler, which adjusts the time window based on historical user activity (such as the teacher's average page-turning frequency and pause intervals) to improve buffer hit rates. The system also supports a dynamic sliding window mechanism, enabling parallel prediction and buffering of multiple pages of instructional content. Experiments show that using this buffering strategy, the system maintains an average response latency of less than 100ms, even with high content complexity, and a cache hit rate exceeding 93%, effectively improving the smoothness and stability of teaching projection. The preloaded image content buffering strategy is written into the GPU pre-cache or local rendering queue, and waits for subsequent instructions to trigger image output, laying the foundation for the next step of teaching projection image scheduling and real-time rendering, ensuring that the projection system has low-latency, high-concurrency teaching image processing capabilities under continuous operation by teachers.
[0014] Step S3: identifying multi-dimensional network status indicators, and performing adaptive regional transmission selection on the teaching image content to obtain a local priority pixel stream and a delayed transmission pixel stream; In this embodiment, the current network environment is monitored and features extracted in real time. Network status metrics collected include, but are not limited to, bandwidth (Mbps), round-trip time (RTT) (ms), packet loss rate (%), and jitter (ms). For example, in experiments, the system automatically invoked the network status detection module every 100ms, combining it with QUIC protocol or TCP statistics packets to achieve millisecond-level sampling. The results of one acquisition were: bandwidth 8.5Mbps, RTT 42ms, packet loss rate 0.8%, and jitter 2ms. These metrics are fed into the "Transmission Scheduling Optimization Module" as input variables. The system identifies content areas and assigns priority scores to instructional images. Using a deep attention mechanism and visual saliency maps, the system identifies key instructional elements in instructional images, such as text, graphic boundaries, areas indicated by teacher gestures, and areas where the cursor is pointing. These elements are then divided into several "high-attention areas" and "background areas." These areas are labeled with transmission priority tags before image encoding. Combining current network status indicators with image content region labels, the system dynamically generates a "local priority pixel stream" and a "delayed pixel stream." The priority pixel stream uses differential compression and forward error correction (FEC) to rapidly transmit data in key areas of the first frame, ensuring that data can be rendered at the highest frame rate (≥25 fps) even under poor network conditions. The delayed pixel stream, consisting of supplementary content such as background fills and edge images in non-interactive areas, is transmitted slowly and at a lower priority after the main frame is displayed, using a secondary compression strategy (such as delayed GOP encoding). In experiments, when network status assessment results indicate an RTT exceeding 70ms and a packet loss rate exceeding 2%, the system identifies the "text content + teacher cursor area" as the priority stream and utilizes AV1's "block-by-block parallel encoding" and fast uplink scheduling transmission mechanism to keep its average arrival latency under 50ms. Background layers and decorative elements are treated as delayed pixel streams, complemented by a delay of 300-500ms, making them virtually imperceptible to the user. This adaptive regional transmission strategy enables rapid response and fault-tolerant transmission of critical content in complex teaching images and fluctuating network environments. The final output image is rendered in real time on the terminal side using a local stitching mechanism, significantly reducing perceived latency during teaching image transmission, improving the continuity and reliability of teaching interactions, and providing foundational data support for subsequent image synchronization and behavioral control strategies.
[0015] Step S4: performing local filling processing on the local priority pixel stream based on the preloaded image content buffering strategy, and performing fast temporary screen rendering to obtain a local pixel rendering image; In this embodiment, based on the generated preloaded image content buffering strategy, image content buffer blocks are constructed. Each buffer block is divided into regions and prioritized to correspond to high-importance areas of the teaching screen (such as the teacher's explanation section, handwriting area, and real-time mouse position area). When a localized priority pixel stream arrives at the teaching terminal via the network transmission channel, the system first determines the region to which it belongs and fills the corresponding buffer block. Simultaneously, for missing segments, the system uses a compensation mechanism based on time-series frame prediction to call upon data from the corresponding region in previously rendered image frames to create a temporary visual fill-in. This then enters the fast temporary screen rendering phase. The system utilizes a GPU parallel acceleration framework (such as the Vulkan or Metal API) to rapidly load the received local pixel data into the rendering pipeline. To further reduce latency, the system utilizes a tile-based rendering mechanism, first rendering high-priority areas block by block, then gradually filling in the edge areas. This process eliminates the need to wait for all pixel data to arrive, significantly reducing first-frame response time. Experiments show that, using this method, the average first-frame local rendering time can be controlled within 38ms, given a bandwidth limited to 6Mbps and a frame size of 1920x1080 for standard HD teaching content. To improve smoothness of transitions and continuity of teaching information, the system also incorporates a neural network-based image edge restoration model (such as a small and lightweight model based on Fast Edge Completion) to fill in detailed transitions between local priority pixel data, ensuring that even if some areas are temporarily delayed in transmission, there is no visual breakage or tearing. The output local pixel rendering image is a "temporary image" composed of the current local priority pixel data, generated through compensation and restoration. This image is displayed on the teaching terminal with a seamless transition and is dynamically overwritten and restored with subsequent full image data. This progressive rendering method has demonstrated strong network jitter resistance in experiments, with each local restoration covering an average of over 92% of the key teaching area. This ensures millisecond-level synchronization between the teacher's operation and the projected display, providing stable image support for terminal behavior linkage control.
[0016] Step S5: performing secondary transmission rendering on the delayed transmission pixel stream, and performing global rendering reconstruction in combination with the local pixel rendering image to construct a global teaching rendering image; In this embodiment, the network end continues to receive delayed transmission pixel streams that have not been fully transmitted. Such pixel streams usually belong to low-priority areas in the image, such as background layers, edge areas, interface borders, etc., and their delayed transmission during the initial screen projection process will not interfere with key teaching information. After receiving these delayed pixels, the system maps them to the corresponding positions of the full-frame image through the content address indexing mechanism (Content Hash Mapping) and marks them as "areas to be filled." Next, it enters the secondary transmission rendering stage. The system adopts a deferred rendering strategy to cache the delayed pixels in an independent texture buffer, load them through the GPU concurrent rendering pipeline, and perform pixel-level alignment and superposition with the rendered local pixel images. To avoid screen flickering or ghosting effects, the system introduces an image fusion transition module during rendering, and adopts a hybrid strategy based on local weighted fusion to fuse old pixels and newly received pixels. In order to achieve consistent reconstruction of the picture, the system also needs to process the possible local image splicing edges. After filling in the lagging pixels, a full-frame seam blending operation is performed on the entire image to correct transitional artifacts that may occur during the local fast rendering phase. The system then uses a lightweight image inpainting model (such as an edge smoothing model based on a U-Net architecture) to fill in pixels at the boundaries of each image region, achieving a natural transition across the entire frame. Ultimately, this process enables seamless transition from a fast temporary image to a complete image. In an experimental setup, using 1920×1080 resolution, 30 frames per second video content as an example, the amount of lagging pixels transferred was kept within 35% of the total pixels in the original image. The average global rendering reconstruction time was less than 72ms, and the visual integrity score exceeded 92%. The system refreshes the image reconstruction status twice per second, performing an MD5 checksum and rendering version stamp on the reconstructed image to ensure the consistency and integrity of each frame.
[0017] Step S6: Render the global teaching rendering image to complete self-correction, calculate the real-time frequency of teaching image switching, and perform pre-loaded frequency resonance optimization to build an intelligent real-time projection control model.
[0018] In this embodiment, after generating a global teaching rendered image, the system performs a "rendering completion self-correction" process on the image. This process involves embedding an invisible digital watermark timestamp (e.g., perturbation implantation based on the DCT domain) within the rendered frame. This timestamp is consistent with the image generation record within the projection system. After the projection terminal completes image rendering, it analyzes the timestamp via a feedback loop or an embedded feedback mechanism and compares it with the currently preset content scheduling timeline. If the rendering output time exceeds the expected range (e.g., inter-frame delay exceeds 50ms), the frame is determined to have potential lag or scheduling offset, and the system triggers a "reschedule" or "preload" strategy to improve image output stability. Secondly, the system performs real-time frequency analysis of the switching behavior of continuous teaching images. This analysis, based on the high-level control command sequence derived from the teacher's behavior analysis, extracts the temporal characteristics of image changes between each frame or teaching segment. Using a sliding window method (with a window size of 35 seconds), the system calculates the frequency of image content changes. This allows the system to identify dynamic teaching rhythms in high-frequency bands (>2 times / second) and low-frequency bands (<0.5 times / second). For example, during actual teaching, the page-turning rhythm of a PowerPoint presentation typically remains stable at 0.81.2 times per second, while blackboard writing scenarios vary more slowly. These frequency analysis results are directly used in the subsequent development of a frequency-synchronization optimization strategy. Finally, based on the real-time image switching frequency and the previously established preloaded image content buffering strategy, the system introduces a "frequency-synchronization" optimization mechanism to intelligently adjust preloaded images. "Frequency-synchronization" means that the system synchronizes the image preloading rhythm on the projection screen with the teacher's actual operation frequency, preemptively loading the required image areas based on the teacher's intended operation. For example, when the teacher rapidly turns pages, the system can prefetch the next page of the PowerPoint presentation and cache it in the GPU cache pool, achieving seamless page switching. During periods of extended intervals between operations, the preloading frequency can be appropriately reduced to conserve system resources. This optimization strategy combines a dynamic frequency adjustment algorithm (such as dynamic threshold adjustment of the FFT spectrum) with a content prioritization model to construct a feedback-driven, rhythm-adaptive intelligent real-time control model. In a 2024 experiment using projection hardware based on an Intel Core i7 and RTX 3060, this model reduced image switching latency by approximately 38%, increased average rendering stability to 98.2%, and reduced system CPU utilization by 14.6%. While maintaining high image quality, the system significantly improved the smoothness and responsiveness of image switching.
[0019] In this embodiment, refer to Figure 2 , is a flowchart of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Collect teachers' real-time behavior monitoring images based on the camera of the teaching terminal; Performing global image brightness enhancement on the teacher's real-time behavior monitoring image to obtain a global brightness enhanced behavior image; Perform deep convolution optimization on the global brightness enhanced behavior image to construct a resolution convolution optimized image; Based on the resolution convolution optimized image, frame-by-frame action behavior analysis is performed, and operation instructions are intelligently generated to obtain high-level control instructions.
[0020] In this embodiment, in teaching scenarios, a high-definition camera built into or connected to a teaching terminal (such as a smart podium, interactive whiteboard, or all-in-one teaching machine) is used to continuously and stably capture video images of the teacher's behavior. It is recommended that the camera used have a resolution of 1080p or higher and a frame rate of 30fps or higher to ensure image clarity and sampling timeliness. The system acquires the raw video stream through an access port (such as USB 3.0 or an embedded CSI channel) and decodes it into individual frames for subsequent analysis. During the data acquisition process, to ensure stability, a multi-threaded video acquisition module is used in conjunction with a frame buffering mechanism to avoid frame drops and sampling asynchrony. Furthermore, to enhance the effectiveness of monitoring teacher behavior, image foreground extraction algorithms, such as background subtraction, are introduced. This is combined with the MOG2 model in OpenCV to accurately segment the teacher's activity area, thereby filtering out irrelevant background information and improving the payload efficiency of subsequent processing stages. The image acquisition module may also incorporate color correction and automatic white balancing mechanisms to adapt to the varying lighting conditions present in classrooms and improve image quality. This step ultimately produces a sequence of real-time teacher behavior monitoring images, providing continuous and valid raw input for subsequent processing. Uneven classroom lighting, backlighting, or screen reflections can cause insufficient brightness and low contrast in the captured images, hindering subsequent action recognition and behavior analysis. To address this, this step performs global brightness enhancement on the raw behavior images, optimizing the overall brightness distribution and making the target areas (such as the teacher's gestures and body movements) more distinct and prominent. The primary method employed is based on adaptive histogram equalization (AHE) and its improved version, contrast-limited adaptive histogram equalization (CLAHE). This method divides the image into multiple small regions (tiles), performs contrast enhancement on each region, and then performs boundary fusion. In this experiment, the tile grid size was set to 8×8 and the clip limit was set to 4.0 to suppress noise amplification while enhancing detail in both dark and bright areas of the image. Furthermore, to preserve color information, the image is converted from RGB to YCbCr before processing, and only the luminance channel, Y, is enhanced to avoid affecting the color channels. This enhanced image brightness makes the teacher's movements more distinct, especially in low-light conditions such as hand movements or standing areas. The enhanced image significantly improves the clarity of movement edges, laying a solid foundation for subsequent movement recognition and feature extraction.This step is accelerated by GPUs, ensuring that image enhancement processing latency is kept below 10ms, maintaining the low-latency requirements of the entire control chain. A deep convolutional neural network is used to extract features and optimize the resolution of the global brightness-enhanced image. This step employs lightweight super-resolution network models, such as ESPCN (Efficient Sub-Pixel Convolutional Network) or EDSR (Enhanced Deep Residual Networks), along with an upsampling strategy to achieve image resolution enhancement.
[0021] In the experimental setup, the input image (typically 720p) was upscaled to approximately 1080p resolution using a deep network. The network employed a three-layer convolutional architecture with 64, 32, and 16 channels, respectively. Each layer had a 3×3 kernel size and used ReLU as the activation function. The model was trained using Mean Squared Error (MSE) as the loss function, Adam as the optimizer, and an initial learning rate of 1e-4. During inference, TensorRT was used for model quantization and acceleration, keeping the average frame processing time under 12ms. This high-resolution deep convolutional optimization not only enhances image detail but also improves the ability to extract spatial semantic information. For example, when processing detailed gestures of the instructor (such as finger placement and palm position), high-resolution images significantly reduce blurring and feature loss, improving the accuracy of the subsequent action recognition model. This step also uses denoising mechanisms (such as bilateral filtering and a residual network architecture) to reduce artifacts caused by image enhancement, ensuring that the output image maintains both high resolution and stability and authenticity. Based on the high-resolution images output from the previous stage, the system performs frame-by-frame action recognition and motion parsing. The core technology employed combines time series models (such as LSTM or 3D-CNN) with pose estimation algorithms (such as OpenPose or MediaPipe) to perform multi-dimensional modeling and semantic understanding of the teacher's actions. The system can recognize teaching behaviors such as gesture commands (pointing on the screen, turning a page with a gesture), body movements (moving left, moving right), and semantic states (starting a demonstration, pausing a lecture). The pose estimation stage uses the MediaPipe Hands and Pose models to extract keypoints, resulting in 21 hand keypoints and 33 human skeleton keypoints per frame, with a data dimension of (33+21)×3 (including x, y coordinates and confidence scores). For action recognition, a two-stream architecture is employed: the spatial branch uses a 2D CNN to extract static features from each frame, while the temporal branch uses an LSTM to capture the dynamic characteristics of actions over time. Combined with a classification network, each frame is labeled (e.g., "gesture up" = turning the page up) and a time window is used to enhance robustness. The action generation module is based on a dual mechanism of rule engines and machine learning. Predefined command templates are used to match behavioral labels, while reinforcement learning strategies are used to optimize the response efficiency of command generation. Ultimately, the identified behaviors are converted into high-level system control commands (such as "CMD_PAGE_NEXT" and "CMD_POINTER_MOVE(X,Y)") and transmitted to the teaching screen projection control module via a low-latency communication interface (such as UDP or MQTT). Throughout this process, the end-to-end latency of the control link is kept below 100ms, enabling real-time synchronization between teacher actions and screen projection control, providing technical support for low-latency interactive teaching.
[0022] In this embodiment, the specific steps of performing frame-by-frame action behavior analysis based on the resolution convolution optimized image and intelligently generating operation instructions to obtain high-level control instructions are as follows: Decompose the resolution convolution optimized image into multiple frames to obtain a time-series frame image sequence; Perform frame-by-frame action behavior visual recognition on the time-series frame image sequence and extract the action behavior features of each frame; Perform teaching control semantic analysis on the action behavior characteristics of each frame to obtain the semantic characteristics of teacher control behavior; Abstract behavioral intention mining is performed on the semantic features of teacher control behavior to obtain behavioral control intention signals; Based on the behavioral control intention signal, operation instructions are intelligently generated to obtain high-level control instructions.
[0023] In this embodiment, a sequence of image frames with time-series properties is extracted from a continuously input stream of resolution-convolution-optimized images to support subsequent action recognition and control command generation based on temporal information. Multi-frame decomposition primarily relies on a frame buffer mechanism and time window sampling techniques. The system sets a fixed sliding time window (e.g., 500ms) within which consecutive frames are extracted (e.g., 15 frames sampled at a 30fps frame rate) to construct a complete time-series frame image sequence. In the specific implementation, the input image stream is asynchronously sampled using a double-buffered queue structure (producer-consumer model) to avoid sampling delays that affect real-time performance. In experiments, the maximum buffer capacity was set to 60 frames, and the time window step size was 200ms (overlapping sampling) to ensure temporal coherence and inter-frame transition characteristics for action recognition. Each frame retains a timestamp and unique identifier to ensure logical and temporal order consistency within the image sequence. The original continuous image stream is structured into a time-series input format suitable for processing by the action recognition model. By maintaining temporal consistency across image frames, the system effectively captures the dynamic evolution of the teacher's continuous movements, providing high-quality, low-latency data support for subsequent modeling and behavioral intent recognition. Each frame in the temporal image sequence undergoes independent action recognition and feature extraction, forming a complete sequence of semantic information about the action. The method employed primarily relies on deep convolutional neural networks (such as ResNet-50 or MobileNetV2) in computer vision, combined with pose estimation frameworks (such as MediaPipe Pose or OpenPose), to extract behavioral features such as the teacher's body posture, limb movements, and hand gestures in each frame. In specific operation, the system performs keypoint detection on each frame, identifying 33 skeletal keypoints and 21 hand gesture keypoints, and obtaining spatial (x, y) positions and action confidence scores. In the experimental setup, each frame's action features are represented as a 54×3 matrix (3D coordinates + confidence scores). Data normalization and noise reduction (such as Gaussian smoothing) are used to mitigate the impact of image noise on feature extraction. In addition, to further improve the accuracy of frame-level action recognition, a lightweight image feature extraction network, such as ShuffleNet or EfficientNet-B0, will be introduced to perform feature encoding on the action area in the image and output a high-dimensional vector as an action embedding representation. Ultimately, each frame of the image will form a set of structured behavioral feature data, including skeletal key point vectors, action area feature embeddings, and environmental context (such as whether it is close to the screen). The low-level action behavior features extracted frame by frame are converted into high-level control semantic features related to the teaching context, realizing an abstract transition from perception to semantics. This process is mainly completed by combining a semantic classification model with a rule engine, that is, semantically mapping actions with predefined teaching behavior intentions, such as "raising hands to turn pages", "pointing to the screen", "gesture zooming" and other common teaching control behaviors.First, a semantic tag library (Teaching Semantic Dictionary) containing common teacher actions is constructed, comprising 30 categories of basic semantic action labels. A multi-label classification model (such as a Bi-LSTM-based attention network) is trained to perform multi-dimensional feature fusion on the feature vector of each frame, outputting a corresponding set of semantic action labels. Using a cross-entropy loss function during model training, the labeling accuracy reaches 87.6% (validated on a real-world on-campus teaching dataset). A teaching context-aware mechanism is also introduced. This mechanism considers the teacher's current position (e.g., proximity to the screen), body orientation, and hand position to determine whether an action is a valid teaching operation. For example, a "pointing at the screen" is only considered effective when the teacher is facing the screen and their finger is at the edge of the display area, thus preventing false triggering. This stage provides a critical semantic bridge for the control system. By mapping action features into recognizable teaching control action labels, subsequent intent modeling and command generation are both operational and accurate, ensuring the system's strong contextual adaptability in understanding the teacher's operational intent. By fusing and reasoning on the semantic features of continuous control actions in a time series, the teacher's overall behavioral intention within a specific time period is discovered. For example, recognizing “waving your hand to the right three times in a row” can be abstracted as “wanting to turn pages continuously”, or “pointing to the lower right corner of the screen and staying there for 1 second” can be abstracted as “clicking on a certain area”.
[0024] The implementation primarily utilizes temporal modeling techniques, combined with a Transformer architecture or Temporal Convolutional Network (TCN) to model multi-frame semantic features. Each set of semantic feature sequences serves as input to the model (e.g., a set of 10-frame semantic label vectors). A temporal attention mechanism is used to mine inter-action dependencies, thereby generating behavioral control intent signals. The system pre-defines 10 types of control intent signals, including "page turn intention," "pause playback," "region click," and "annotation start." In the experimental setup, the model uses a 256-dimensional input embedding, a three-layer Transformer encoder, and a 500ms time window. By sliding modeling multiple label sequences, the system achieves contextual understanding of the teacher's actions over continuous time periods, effectively improving intent recognition accuracy. The model achieved a test accuracy of 89.2%, with inference time under 30ms. This step, which translates semantic labels into abstract control intent, is a key step in enabling teachers to freely and naturally operate teaching equipment. It significantly improves the system's responsiveness to complex teaching behaviors, enhances user experience, and reduces the rate of misoperation. The system intelligently converts behavioral control intent signals mined by the system into high-level control commands that can be directly parsed and executed by the teaching projection device. This conversion process takes into account practical operating conditions such as device interface protocols, command standardization, and latency minimization. The system uses an intelligent mapping module (Action-to-Command Engine) to dynamically generate structured control commands (e.g., in JSON format) based on predefined mappings between control intents and device operation commands. This approach utilizes a hybrid rule-based and model-based command generation mechanism. Basic mappings are quickly matched using a hash dictionary (e.g., "page turn intent" → "CMD_NEXT_PAGE"). For complex actions, customized commands are generated through logical reasoning based on intent confidence and contextual state. For example, when recognizing the "pause lecture" intent, the system also determines whether a video or animation is currently playing; if not, the command is not issued. To reduce command latency, the system uses lightweight communication protocols such as MQTT or WebSocket-based data channels to push commands. The command structure includes information such as timestamp, control type, and command parameters, enabling the teaching projection device to quickly parse and respond. The command issuance process is kept within 50ms, ensuring the total delay from motion recognition to device response is less than 150ms. This step completes the entire closed-loop control chain, enabling the entire process from the teacher's natural behavior to intelligent screen projection control. This ensures that the system possesses core capabilities such as low latency, high precision, and strong interactivity, fully meeting the real-time control requirements of modern intelligent teaching environments.
[0025] In this embodiment, refer to Figure 3 , is a flowchart of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Identify the content of the next frame of teaching images based on high-level control instructions; Calculating the number of elements and the degree of detail of the teaching image content to obtain the visual complexity of the image content; Predicting an image switching time point based on the teaching image content to obtain the image switching time point; Calculate the available preloading time based on the image switching time point to obtain the preloading time window length; A preloading buffering decision is made based on the visual complexity of the image content based on the length of the preloading time window, and a preloading image content buffering strategy is constructed.
[0026] In this embodiment, high-level control instructions generated based on teacher behavior accurately predict the teaching image content to be displayed or interacted with. This step inherits the aforementioned intelligent control instruction generation module. Instructions typically include operations such as "turn page," "open link," and "load animation demonstration." The system quickly locates the corresponding teaching image resource in the background resource management module by parsing the instruction type and its accompanying parameters (such as the target page number and region location). The system maintains a teaching content index table that corresponds to the structured information of the teaching courseware (such as PPT page structure, PDF page, video frame markup, etc.). After each high-level control instruction is recognized, it triggers the content query module. The corresponding resource path, file identifier, and image type are quickly retrieved and their physical location on the device is located in the cache directory. For example, when the control instruction is "CMD_PAGE_NEXT," the system automatically retrieves the image content of the "next frame" based on the current page number, such as "slide_12.png." To further accelerate response times, the system uses an LRU (least recently used) caching strategy to keep previously accessed content in a hot cache in memory. Simultaneously, a multi-threaded loading mechanism is employed for predicted instruction content, enabling asynchronous extraction of images and interactive content. This step ensures that the teaching projection system can acquire target image resources in a very short time, providing a foundation for subsequent loading, display, and preprocessing. After identifying the teaching image content to be displayed, the system assesses its visual processing complexity to provide a basis for loading optimization, video memory management, and preprocessing scheduling. Visual complexity assessment primarily focuses on two dimensions: the number of image elements and detail density. This assessment directly influences subsequent buffering strategies and adjustments to the switching prediction model.
[0027] First, the image is rapidly segmented (e.g., using the SLIC superpixel algorithm) to roughly determine the number of blocks within the image. Edge detection (e.g., Canny or Laplacian operators) is then used to calculate the image's structural complexity. In experiments, images segmented into more than 500 superpixel blocks and with a greater than 15% edge percentage are considered "high complexity." The system also calculates the image's Shannon Entropy (Shannon Entropy), with higher values indicating greater information density. An empirical threshold for image entropy greater than 6.5 indicates a high level of detail. For multimedia content (e.g., animation frames and embedded videos), a keyframe recognition mechanism is used. Inter-frame differences (based on SSIM or MSE) are calculated to calculate scene change frequency, which is used to estimate the rate of visual change and serve as a dynamic factor of visual complexity. Visual complexity is quantitatively recorded as a three-dimensional tuple (number of elements, edge density, and image entropy) for subsequent cache resource scheduling strategies. This process takes an average of approximately 25ms, ensuring it does not impact overall response latency and provides highly accurate input for prediction algorithms. During the teaching process, image switching timing is a key reference for achieving low-latency loading and resource prefetching. This step predicts the next likely image switching time based on historical behavior and the current instruction context. The system uses a time series prediction model (such as LSTM or Temporal Attention) and combines the teacher's current operation frequency, content complexity, and course pacing to estimate the probability distribution of future image switching. Input features include: the duration of the previous image, the teacher's last control instruction type, the current content complexity level, and interaction frequency. During experimental training, the prediction model was trained using real-world teaching data (collected from 30 classes totaling 1500 minutes of video). The average prediction error was less than 500ms, and the prediction confidence threshold was set to 0.85. When the prediction model outputs a switching probability exceeding the threshold for a certain time period (e.g., t+2.3 seconds), that moment is marked as the "estimated image switching time point." This step also integrates a fast heuristic prediction rule: for example, if the average interval between three consecutive "page-turning gestures" is 1.5 seconds, the interval between subsequent switching points is automatically set to 1.5 seconds. This prediction mechanism ensures that the system can complete resource loading before the switch is about to occur, thereby significantly reducing the occurrence of screen projection delays or freezes. The image switching time point is updated in real time as a core parameter, and an early warning mechanism is provided to downstream modules, providing a time benchmark for image buffering and preloading operations. After the image switching prediction point is determined, the system needs to calculate the idle period from the current moment to the next image switching point, that is, the "preloading time window length". This window is defined as the safe time interval between the current moment and the predicted switching point that can be used for resource loading and processing. This window length is a key basis for formulating buffering strategies and resource scheduling priorities.
[0028] The system first calculates the time available for preloading by obtaining the current system clock (system timestamp t_now) and the predicted switching time (t_pred). The difference, Δt = t_pred - t_now, represents the available preloading window. In experiments, Δt typically fluctuated between 300ms and 2500ms, with an average of 1450ms. To ensure system stability, a loading safety margin (Loading Safety Margin) was set at 200ms. If Δt < 300ms, preloading was not performed to avoid conflicts or redundant loading resources. Taking into account task parallelism, the system estimates the preprocessing time (T_preprocess) required to load the current image content, such as 10ms for image decoding, 25ms for complex image enhancement, and 20ms for video memory allocation. The sum of these times is then compared with Δt to ensure that resource loading tasks do not exceed the preloading window. This calculation ensures time-sensitive system resource scheduling, enabling precise control of loading task priority and concurrency, improving overall image switching responsiveness and maintaining stable operation of the teaching projection device. While minimizing system latency, adaptive buffering scheduling schemes are developed for content of varying complexity. Buffering strategies require a comprehensive consideration of two key variables: the visual complexity of the image content and the length of the available preloading window. The system employs a hierarchical loading strategy, classifying image content into three levels of complexity (low, medium, and high) and categorizing preloading window lengths into short (<800ms), medium (800-1600ms), and long (>1600ms). Within these nine combinations, the system configures different loading priorities, processing flows, and video memory management strategies. For example: high complexity + short window: skip preprocessing and directly cache the original image; medium complexity + medium window: perform compression decoding and brightening; high complexity + long window: perform full image enhancement and edge refinement. To enforce these strategies, the buffering module manages preloading task queues (such as the PriorityQueue) through a multi-queue asynchronous mechanism, ensuring that images that are about to be displayed are loaded first. To reduce video memory pressure, the system implements an image hotness elimination strategy (content access frequency × last access time) to dynamically clear the cache of low-priority images. This pre-loaded buffering strategy reduced the average image switching response time from 480ms to 180ms, and the lag rate decreased by approximately 63%. The system can still stably support smooth image switching in high-load teaching environments, fully ensuring the real-time and consistency of low-latency teaching projection.
[0029] In this embodiment, reference Figure 4 The above is a flowchart of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include: Real-time detection of the current network's transmission bandwidth, delay, and packet loss rate to obtain multi-dimensional network status indicators; Perform real-time delay status evaluation based on multi-dimensional network status indicators to obtain a real-time network delay status evaluation value; Divide the teaching image content of the next frame into key areas, calculate the regional priority, and mark the priority transmission level of each area; Adaptive area transmission selection is performed based on the real-time network delay status evaluation value and priority transmission level to obtain the local priority image area and the delayed transmission image area; Local pixel particle compression is performed on the local priority image area and the delayed transmission image area to obtain a local priority pixel stream and a delayed transmission pixel stream.
[0030] In this embodiment, comprehensive monitoring of the current network link status allows for dynamic adaptation of the transmission strategy for teaching image content, ensuring low-latency, high-quality transmission under varying network conditions. The system monitors three key network performance indicators in real time: transmission bandwidth, communication latency, and packet loss rate. These three indicators constitute a multi-dimensional set of indicators for the current network status and serve as a core reference for subsequent policy scheduling. The system integrates a lightweight network monitoring module, employing a combination of active probing and passive statistics. The active probing component periodically sends test packets (UDP ping) to measure round-trip time (RTT) to assess latency. It also assesses available bandwidth based on the test packet transmission rate and response time (e.g., sending 64KB of test data per second and calculating the ACK response time). Passive statistics monitor the ACK and NACK feedback signals in the real-time data stream, recording the frequency of packet loss events and dynamically calculating the packet loss rate. In the experimental setup, the system set a detection frequency of 5 times per second, a packet size of 128B, a default baseline latency tolerance threshold of 80ms, and a packet loss rate warning threshold of 5%. The detection results are normalized and stored in a network status cache pool, where they are updated in a sliding time window, forming a dynamic metric library for latency assessment and image transmission policy adjustment. This step ensures the system's millisecond-level awareness of network changes, providing the foundation for subsequent transmission adaptation and image content scheduling. After completing the network bandwidth, latency, and packet loss rate detection, the system conducts a comprehensive assessment of the current network quality to quantify the network latency status, providing data for image content scheduling and transmission priority determination. This assessment process combines a weighted scoring model with a sliding time window mechanism to calculate the current network latency assessment value (Real-Time Network Latency Score, R-NLS) in real time. In the model settings, latency (RTT) is weighted as 0.5, packet loss rate is weighted as 0.3, and bandwidth is weighted as 0.2. The comprehensive calculation formula is: R-NLS = 0.5 × f(RTT) + 0.3 × f(PacketLoss) + 0.2 × f(1 / Bandwidth), where f represents a normalization function (such as Z-score or Min-Max). To enhance robustness, the system uses a sliding window mechanism to record the evaluation results for each second within the last five seconds and employs a time-weighted average method to generate the current effective evaluation value. Experimental data demonstrates that this evaluation method can smooth sudden network fluctuations and maintains stability in mobile teaching scenarios (such as tablets or wireless projection). Evaluation values are categorized into four levels: excellent (<0.25), medium (0.25), poor (0.75), and extremely poor (>0.75). The system dynamically switches image region transmission strategies and compression levels based on these values. This delay status evaluation value serves as the scheduling signal for subsequent region priority transmission decisions, significantly improving the system's adaptability to network uncertainties and ensuring a continuously smooth projection process.The next teaching image is divided into spatial regions based on content importance and visual attention, and each region is assigned a different transmission priority. This ensures that core teaching content is presented in a timely manner even when network conditions are poor. Image region division primarily utilizes a combination of content awareness and semantic analysis to identify and categorize key elements in the teaching image.
[0031] First, the system performs perception-driven saliency detection on the image, such as using UNet or U2Net to extract salient regions, identifying areas where students are most likely to focus. Combined with structured courseware data or optical character recognition (OCR) results, the system identifies key teaching content, such as text blocks, titles, formulas, and diagrams. In the experimental setup, the image was divided into 4×4 grid cells (16 cells total). For each cell, its saliency score, text density, and interactive annotation overlap were calculated. Each metric was normalized and weighted, resulting in a priority score for that region. The system sets three transmission priority levels: high (P1), medium (P2), and low (P3). For example, areas containing instructional instructions or key content are marked as P1, while background patterns or auxiliary illustrations are marked as P3. This grading results in a transmission priority mask map, which is used to guide subsequent image data segmentation and resource scheduling. This hierarchical transmission control of image regions ensures that students receive critical teaching information in a timely manner, even when latency assessments are high, ensuring low-latency core teaching information security. Based on the delay state evaluation value (R-NLS) and image region priority labels, the system performs adaptive region selection and batch transmission for each frame. This divides image content into "local priority image regions" and "delayed transmission image regions," ensuring that critical areas are presented first even in poor network conditions. The system also designs a region scheduling strategy matrix that determines transmission strategies based on a cross-match between network status levels (excellent, medium, poor, and extremely poor) and image region priorities (P1, P2, and P3). For example, in a "poor" state, only the P1 region is transmitted, delaying the P2 and P3 regions; in an "extremely poor" state, ultra-low-resolution P1 content compression is enabled. The system constructs a transmission task list based on the image block structure, attaching a priority field to each image region task and entering it into an asynchronous transmission queue (using a priority queue data structure). The system dynamically adjusts the number of concurrent transmission threads based on real-time bandwidth estimation (for example, when bandwidth is less than 2Mbps, the number is limited to one thread, sending only the core region stream).
[0032] Ultimately, each image frame is logically divided into two categories: an immediate priority image region (encoded immediately and transmitted first) and a deferred image region (encoded asynchronously and cached). This ensures that critical teaching images are presented first during network congestion or unexpected delays, optimizing the interactive experience between teachers and students and enhancing the system's robustness and adaptability. Microblock-level compression technology is employed to differentially compress image regions, constructing two independent transmission streams: the immediate priority stream and the deferred stream. This optimizes the structure of image content at the network transport layer. At the encoding level, image compression employs a perceptual quality-oriented encoding strategy. Priority regions (such as P1) are compressed with high fidelity (e.g., H.265 encoding with residual enhancement) to ensure detail clarity and display latency within 100ms. Deferred regions are compressed at a low bitrate, even employing color downsampling (YUV420) and block interpolation. Each image area is divided into micro-blocks (such as 8×8 or 16×16 pixels), and different compression ratios are assigned according to their priority levels. In the experimental setting, the compression ratio of the P1 area is 1:4, and the compression ratio of the P3 area can reach up to 1:20. The system integrates an asynchronous encoding thread pool, and the local priority pixel stream directly enters the "real-time transmission buffer", while the lagging pixel stream is written to the "delayed cache queue" and waits for transmission after the network conditions improve. This differentiated compression and diversion transmission strategy effectively reduces the pressure on the main link data flow and reduces the overall image loading time by 20~40%, ensuring that key content is displayed first and secondary content is supplemented later, thereby significantly improving the screen projection response speed and network adaptability without sacrificing teaching integrity. It is one of the important supporting modules for the low-latency control system of teaching screen projection equipment.
[0033] In this embodiment, step S4 includes the following steps: Pre-constructing the teaching scene according to the teaching image content of the next frame to obtain a pre-constructed loaded teaching scene; Extracting buffered interaction paths based on preloaded image content buffering strategy; The projection terminal receives the local priority pixel stream based on the buffered interaction path; Performing local filling processing on the pre-built loaded teaching scene according to the local priority pixel flow to obtain a local image filling scene; Perform fast temporary screen rendering on the local image filling scene to obtain a local pixel rendering image.
[0034] In this embodiment, after the image content is identified and extracted, in order to reduce the loading delay before the image is rendered, the system needs to build a teaching scene framework for pre-rendering before the image arrives, that is, pre-construct and load the teaching scene. The core goal of this step is to build an image skeleton (Scene Skeleton) based on the known content in advance, and retain the area placeholders, structural logic and interactive element framework, and then perform local filling after receiving the actual pixel content. The system parses the image structure data, such as page hierarchical information, text block layout, graphic container, dynamic area position, etc., and builds a structured scene tree (Scene DOM Tree) through content tags (ContentMetadata). The scene tree includes node types (text, graphics, annotation areas, etc.), spatial positions (coordinates, dimensions), hierarchical relationships and interactive properties (whether it can respond to gestures, etc.).
[0035] For a standard educational image (1920×1080), containing an average of 32 layer nodes, the system completes scene pre-construction in approximately 40ms. Scene elements use vector graphics as placeholders (such as SVG paths or Canvas outlines), and grayscale rectangles mark areas where pixels have yet to be loaded. This pre-constructed scene serves as a mounting container for subsequent partial pixel filling, helping to reduce initial black screen latency and providing structural support for rapid responsiveness.
[0036] Through a pre-built loading mechanism, the teaching projection terminal can pre-display the page structure before pixel data is fully transmitted, improving the perceived speed of screen loading and reserving space for asynchronous filling of local pixels, achieving a more natural image transition experience. Based on the previously established image region priority mask and buffering strategy, an interactive path for image data transmission and rendering is established. The so-called "buffer interactive path" refers to the logical order and priority channel settings for pixel data transmission from the content source (such as the teaching server or local cache) to the projection terminal's various buffers (image structure buffer, transmission buffer, decoding buffer, and rendering buffer). The system analyzes the distribution map of the current image's local priority regions (such as the P1 and P2 marked areas), combines the pre-load time window with the network assessment level, and determines the processing nodes and channel priorities that the priority regions must pass through. For example: image structure → compression processing → transmission scheduling → decoding queue → temporary rendering buffer → rendering engine. In the experimental setup, the system set a maximum of eight concurrent processing threads for the buffer nodes, of which three high-priority channels were exclusively used for local priority pixel data. A streaming protocol (such as UDP-based QUIC) was used to reduce TCP handshake latency. Path extraction also includes rate-limiting parameters, such as a maximum processing rate of 50ms per frame for image region segmentation to avoid thread blocking. The buffered interaction path forms a dynamic scheduling model that guides the system's handling of data packets and pixel blocks of varying priorities, maximizing the utilization of limited network and computing resources and prioritizing the processing and display of locally important content. After obtaining the buffered interaction path, the projection terminal activates the corresponding pixel reception mechanism, prioritizing pixel streams marked "local priority." The core of this mechanism is to identify and prioritize pixel blocks in important regions, ensuring that the teacher's key content is presented on screen as quickly as possible, reducing perceived latency and improving real-time interaction. The system employs an event-driven data flow scheduling model. Based on the channel priority set in the interaction path, it uses a non-blocking receive buffer (such as Zero-Copy or DMA memory mapping) to monitor data packet streams on the corresponding transmission port. Each pixel packet carries region identification information (such as region ID, coordinate range, and priority level), and the system caches them according to priority into different decoding queues.
[0037] The projection terminal can receive approximately 180 image area data blocks (128×128 pixels) per second. Given a 10Mbps network bandwidth and a latency of less than 100ms, the local priority area image can be initially displayed within 150ms. The system employs a differentiated buffer queue mechanism, whereby high-priority data blocks are set with a 5ms timer and immediately transferred into a temporary rendering buffer once they are decodable, avoiding the need to wait for subsequent blocks and delay overall rendering progress. This mechanism significantly improves local display efficiency, enabling the projection terminal to implement a "display the important first, add details later" image transmission strategy even when network fluctuations or resource constraints persist. This is a key guarantee for low-latency teaching images. Once the projection terminal successfully receives the local priority pixel stream, the system immediately performs a local fill operation based on the pre-loaded teaching scene. The essence of the local fill process is to map the priority pixel data to the corresponding area of a predefined image structure according to its spatial location, enabling rapid visualization of key visual areas within the structured scene. The system relies primarily on the Region Mapping Table, which records the spatial index relationship between each logical partition of the image and the scene node (including coordinate range, scaling factor, rotation angle, etc.). The system can directly index the target rendering area based on the region ID contained in the pixel block. For the received pixel stream data, the system first decodes it (such as JPEG frames or H.265 compressed blocks), and then maps the decoded results to the corresponding grid area in the image Canvas buffer.
[0038] Partial fill rates can reach up to 90 FPS, primarily limited by the decoding and GPU synchronous write rates of image blocks. To avoid visual discontinuity in certain areas due to data latency, the system supports edge softening and smooth transitions between blocks to enhance fill continuity. Placeholders are retained during the fill process for subsequent completion of low-priority pixel streams to ensure final image consistency. After completing partial pixel fill, the system immediately performs a temporary screen rendering of the image content to generate a displayable "partial pixel rendered image." While not yet complete, this image already captures the core teaching content, enabling a "preemptive response" effect in user perception and a key component in low-latency image control. The fast rendering process utilizes lightweight GPU instruction sets (such as OpenGL ES or Metal) for image synthesis, calling Canvas or FrameBuffer components to draw the filled area to the target screen. The system sets a frame priority flag in the render queue, prioritizing the push of partially rendered frames to the primary display buffer, implementing a "partially displayable" strategy. A double-buffering mechanism also prevents render tearing: the currently rendered frame is switched to the foreground display after the backend buffer is complete, minimizing visual presentation time. The local rendering process for a single frame takes approximately 12ms, and the GPU can handle over 100 regional updates per second at peak speed. To further enhance consistency, the system integrates a partial frame repair mechanism. When key areas are not fully filled, they are temporarily replaced with low-resolution blurred blocks, which are automatically redrawn when new data blocks arrive. This partial filling strategy allows students to see key teaching content before the entire image is fully transmitted, avoiding cognitive interruptions caused by blank interfaces or "jumpy" presentations, and greatly improving the user's visual experience.
[0039] In this embodiment, step S5 includes the following steps: Identify the transmission resource occupancy information of the local priority pixel stream; Based on multi-dimensional network status indicators, idle transmission resources are calculated based on transmission resource occupancy information to obtain real-time idle transmission resources; Perform secondary transmission rendering on the delayed transmission pixel stream according to real-time idle transmission resources, and perform fuzzy interpolation processing to obtain a rendered image of the delayed area; The global rendering is reconstructed based on the lagging area rendering image and the local pixel rendering image to construct a global teaching rendering image.
[0040] In this embodiment, the resource usage of the current local priority pixel stream at the network transport layer is dynamically assessed, including but not limited to bandwidth utilization, channel concurrency, buffer utilization, and transmission queue load. By monitoring these metrics, the system can determine the actual transmission channel consumption of the current image backbone data stream, providing room for decision-making when scheduling subsequent delayed data. The system assigns a unique flow identifier (such as flow ID, flow type, and regional level) to each image stream in the data channel. A real-time counter (e.g., based on a ring buffer or memory-mapped I / O) is also implemented in the underlying network I / O module to periodically sample and categorize traffic data. Key monitoring metrics include the current priority pixel stream data rate (bps), average packet size, TCP window size change, UDP retransmission count, and buffer frame loss rate. The system refreshes resource utilization status every 100ms, calculating the data throughput trend over the last second using a sliding window approach. In a 10Mbps network environment, the local priority pixel stream averages approximately 5.8Mbps of the main channel bandwidth, while maintaining a concurrent buffer of approximately 20 regional frames. This utilization information is ultimately aggregated into a set of structured metrics for use by the subsequent idle resource assessment module. Furthermore, combined with the real-time monitoring results of the network status, the remaining network resources available for data transmission in the lagging area, namely the "real-time idle transmission resources", are dynamically calculated. This calculation process integrates network bandwidth evaluation, a dynamic adjustment model for packet loss rate, and a buffer share balancing mechanism to achieve intelligent resource perception and dynamic allocation. The calculation method is based on two types of inputs: one is the "priority flow resource occupancy table" output from the previous step; the other is the current system's network status indicator set (such as RTT, available bandwidth, packet loss rate, etc.). The system has designed a set of weighted resource margin models, defined as: Available transmission bandwidth = current bandwidth × (1 - packet loss rate) × network status coefficient Real-time idle resources = available transmission bandwidth - average bandwidth consumption of local priority flows The network status coefficient automatically adjusts based on the aforementioned latency assessment level (for example, it decreases to 0.7 when the latency status is "poor") to improve system robustness. In experiments, the system maintained an average of approximately 2.5-3.0 Mbps of remaining transmission resources when network latency was kept below 100ms and packet loss was below 2%. The system periodically updates the idle resource assessment value every 50ms and uses it as a scheduling signal to dynamically determine whether to activate secondary image streams, set compression levels, and prioritize supplementary transmissions. This mechanism enables dynamic awareness and on-demand utilization of transmission resources, providing real-time conditions for smooth loading of delayed areas. Once the system confirms the availability of real-time idle resources, it activates the secondary transmission and rendering process for the pixel streams in the delayed areas. Because the delayed areas are not core to the teaching visual experience, their transmission and presentation strategies emphasize "quick completion" rather than "precise reproduction." Therefore, moderate compression transmission is required while ensuring system stability, combined with fuzzy interpolation techniques for visual infill to produce acceptable rendered images in the delayed areas. The specific process involves first applying high-compression processing (e.g., H.265 low-bitrate segmentation, WebP image compression, or AI-based semantic coding) to the pixel content in the lagging area, with a compression ratio between 1:10 and 1:20. Secondly, the transmission concurrency is set based on available resources (e.g., a maximum of 50 lagging blocks per second) and sent to the terminal decoding queue. Upon arrival at the terminal, the system initiates real-time interpolation rendering without waiting for all blocks to be complete. The fuzzy interpolation method uses a boundary-preserving bidirectional interpolation algorithm (e.g., Bilateral Interpolation) and fills missing areas with the mean value of surrounding pixels. In the experimental setup, for image blocks marked as P3 in a 1080p image, the average filling time was less than 20ms, and the perceived error ΔSSIM was less than 0.05, meeting the visual quality tolerance required for teaching. The main purpose of the lagging rendered image is to fill in gaps, mitigate visual jumps, and avoid disruptive learning experiences caused by unloaded areas on the screen. It, along with the locally prioritized pixel-rendered image, constitutes the two main sources of the final image reconstruction. On the terminal side, the local pixel rendering image (composed of the P1 / P2 areas) is merged with the lagging area rendering image (the P3 area composed of fuzzy interpolation) to form a complete frame of global teaching rendering image. This global image must ensure that the teaching information is complete, the visual hierarchy is natural, and the picture connection is smooth. The reconstruction process adopts a multi-channel image synthesis strategy. The system uses the pre-constructed scene as the layer basis, first loads the local rendering image content, and then maps the lagging pixel layer to the corresponding area based on the area index. To avoid the sense of boundary caused by block-level transitions, the system introduces Alpha blending and boundary softening mechanisms: a 12-pixel transition band is set at the junction of each area, and Gaussian blur (σ=1.5) is used to generate the blending area to achieve a natural transition.In addition, the synthesis engine supports a dynamic frame rate compensation strategy: when the lagging area has not been fully loaded, the content of the previous frame can be used to partially fill in the frame through time interpolation or the residual method of the previous frame to avoid the flickering of black blocks or empty frames. In the test environment, the average time required to synthesize a complete global image frame is 27ms, and the error rate of the reconstructed image (ΔPSNR) is controlled within 2.5dB, making it difficult for the naked eye to distinguish the difference. The final generated global teaching rendering image is the complete picture displayed to students by the projection terminal. It not only ensures the timeliness of the core teaching content, but also ensures the overall integrity and aesthetics of the picture, building a complete low-latency, high-perceptual quality teaching projection experience closed loop.
[0041] In this embodiment, step S6 includes the following steps: A digital watermark timestamp is implanted into the global teaching rendering image, and a self-correction is performed based on the timestamp on the screen fragment captured by the projection terminal to generate a self-correction strategy for the rendering image; Based on the teaching terminal, the teacher's continuous operation is recognized and the real-time frequency of the teaching image switching is extracted; Performing a teacher's real-time behavior rhythm analysis on the real-time frequency to generate a teacher's real-time behavior rhythm feature; According to the real-time behavior rhythm characteristics of teachers, the pre-loaded image content buffering strategy is optimized for pre-loaded frequency resonance, and an intelligent real-time screen projection control model is constructed.
[0042] In this embodiment, to enhance the self-monitoring capabilities of teaching projection image rendering and detect abnormalities in terminal rendering quality without manual intervention, the system incorporates a "digital watermark timestamp embedding + camera self-correction" mechanism. This method embeds a timestamp, recognizable by the terminal, into the final rendered image, achieving closed-loop verification of the entire process from image output to screen display. When synthesizing the "global teaching rendering image," the system embeds the current frame generation time (e.g., UTC timestamp) into an invisible layer of the image. Using frequency-domain watermark encoding techniques (e.g., DCT transform + spread spectrum coding), the timestamp is converted to a 16-bit binary code group and then embedded in the low-frequency region of the image's frequency domain. This embedding is performed in areas with low local texture intensity to minimize visual perception. The camera on the projection terminal periodically captures fragments of the current screen display (e.g., a 256×256 pixel block in the center every 2 seconds). Frequency-domain analysis is used to decode the watermark timestamp and compare it with the current system time. If the deviation exceeds a threshold (e.g., 500ms), a "rendering delay warning" is triggered, and the image content features and device status are recorded for subsequent tracing. In the experimental setup, watermark embedding took less than 3ms, had a negligible impact on image quality (ΔPSNR < 0.15dB), and the average misidentification rate of the embedded position was less than 1.2%. This mechanism implements a low-cost self-verification strategy for image display, providing the terminal with basic image self-diagnosis and performance fallback capabilities. It is an essential component for ensuring low-latency and highly reliable teaching rendering. To more accurately grasp the teacher's operational rhythm and image transformation patterns in actual teaching, the system needs to identify the teacher's continuous operational behavior in real time and extract the frequency parameter of image content switching. This frequency parameter not only reflects the intensity and rhythm of the teacher's operation but also directly affects the update strategy and resource scheduling efficiency of the projected content. The teaching terminal continuously monitors the teacher's action flow (such as gestures to turn pages, clicks to switch, and pointing slides) and records these action events with a timestamp. Each action that triggers an image content switch is marked as a "valid switch event" by the system. The number of events is counted within a sliding time window (e.g., the past 30 seconds) to calculate the image switching frequency per unit time (times / second or frames / second). The manipulation recognition method utilizes a temporal deep convolutional network (e.g., 3D-ConvNet) and a temporal action recognition model (e.g., TCN-Temporal Convolutional Network), combining dual-channel data fusion of image features (frame-to-frame change) and interactive events (touch, cursor movement). The system sets a trigger threshold; a valid switch is considered when the movement amplitude exceeds 15 pixels and the local image difference rate exceeds 20%. Experimental results show that in standard multimedia teaching scenarios, the teacher's image switching frequency generally ranges from 0.25 to 0.45 times per second, while high-intensity manipulation can exceed 0.7 times per second.The extracted frequency serves as an important input parameter for subsequent rhythm analysis and pre-loading synchronization, achieving precise alignment between the teaching operation rhythm and the system response rhythm. After extracting the image switching frequency, the system further performs dynamic behavioral modeling on the time series data to extract the teacher's operation rhythm pattern during specific teaching stages. Rhythm analysis can predict the time window for the next switching, optimize image scheduling strategies, and avoid delay mutations such as "operation before loading." The rhythm analysis method is based on multi-scale time series modeling. The system uses a combination of weighted Fourier transform (WFT) and wavelet decomposition (DWT) to extract time-frequency features of the image switching frequency. The goal is to identify the "rhythm fundamental frequency" (i.e., the average operation period) and "rhythm fluctuation rate" (i.e., the degree of rhythm variation) in teaching activities. In actual deployment, the system sets two analysis windows: a short window of 5 seconds to capture instantaneous changes and a long window of 30 seconds to capture periodic trends. Combined with real-time teacher operation data, the system can extract five typical rhythm features, such as "constant page turning pattern" and "pause-burst operation pattern." Each rhythm is encoded as a high-dimensional vector (e.g., [period mean, fluctuation amplitude, peak density, number of local extreme points]) for use by the scheduling module. This approach aligns image loading behavior with the teacher's operation rhythm, building an intelligent projection control model with real-time perception, prediction, and adaptability to achieve an optimal low-latency projection experience. The specific approach involves two levels of policy optimization: first, image preloading window adjustment. The system dynamically adjusts the number of buffered images and the order of loading based on the current rhythm state. For example, in a fast-paced state, the forward preloading depth is increased (e.g., from 2 to 5 pages) and pixel priority in high-frequency interactive areas is increased. In a paused state, the loading window is compressed to conserve resources for detail filling. Second, the loading policy structure is optimized. A predictive scheduling mechanism is introduced. The system uses rhythm characteristics to predict the next operation timeframe and pre-allocate network and cache resources to ensure that images are ready before the expected time, avoiding loading interruptions caused by sudden operations. In experimental scenarios, "co-frequency resonance optimization" improved image switching response speed by 18% to 35% compared to static buffering strategies, achieving particularly significant results in high-frequency operation instruction. Ultimately, the system integrates the teacher's behavioral rhythm and image loading process into a unified control model, forming a closed-loop intelligent projection control strategy. This model, integrated into the transmission coordination module between the teaching terminal and the projection device, features cross-frame prediction, dynamic scheduling, and strategy adaptation, making it one of the core mechanisms for achieving a high-quality, low-latency, stable, and smooth teaching projection experience.
[0043] In this embodiment, a low-latency real-time control device for a teaching screen projection device is provided, which is used to execute the low-latency real-time control method for the teaching screen projection device as described above, including: The teaching behavior analysis module is used to collect real-time teacher behavior monitoring images based on the teaching terminal, perform frame-by-frame action analysis, and intelligently generate operation instructions to obtain high-level control instructions; The preloading buffer module is used to identify the teaching image content of the next frame based on high-level control instructions, perform content complexity calculation and preloading buffer decision-making, and build a preloading image content buffering strategy; An adaptive transmission module, configured to identify multi-dimensional network status indicators and perform adaptive regional transmission selection on the teaching image content to obtain a local priority pixel stream and a delayed transmission pixel stream; The local rendering module is used to perform local filling processing on the local priority pixel stream based on the preloaded image content buffering strategy and perform fast temporary screen rendering to obtain a local pixel rendering image; The global rendering reconstruction module is used to perform secondary transmission rendering on the delayed transmission pixel stream, and to perform global rendering reconstruction in combination with the local pixel rendering image to construct a global teaching rendering image; The same-frequency resonance optimization module is used to perform self-correction on the rendering of the global teaching rendering image, calculate the real-time frequency of teaching image switching, perform pre-loaded same-frequency resonance optimization, and build an intelligent real-time projection control model.
[0044] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.
[0045] The foregoing description is intended only to provide specific embodiments of the present invention, which are intended to enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest manner consistent with the principles and novel features disclosed herein.
Claims
1. A low-latency real-time control method for a teaching screen projection device, characterized in that: The teaching screen projection device includes a teaching terminal and a screen projection terminal and includes the following steps: Step S1: The teacher's real-time behavior monitoring image is collected by the teaching terminal, and frame-by-frame action analysis is performed, and operation instructions are intelligently generated to obtain high-level control instructions; Step S2: Identify the teaching image content of the next frame based on the high-level control instruction, perform content complexity calculation and preload buffering decision, and build a preload image content buffering strategy; Step S3: identifying multi-dimensional network status indicators, and performing adaptive regional transmission selection on the teaching image content to obtain a local priority pixel stream and a delayed transmission pixel stream; Step S4: performing local filling processing on the local priority pixel stream based on the preloaded image content buffering strategy, and performing fast temporary screen rendering to obtain a local pixel rendering image; Step S5: performing secondary transmission rendering on the delayed transmission pixel stream, and performing global rendering reconstruction in combination with the local pixel rendering image to construct a global teaching rendering image; Step S6: Render the global teaching rendering image to complete self-correction, calculate the real-time frequency of teaching image switching, and perform pre-loaded frequency resonance optimization to build an intelligent real-time projection control model.
2. The low-latency real-time control method for a teaching screen projection device according to claim 1 is characterized in that: The specific steps of step S1 are: Collect teachers' real-time behavior monitoring images based on the camera of the teaching terminal; Performing global image brightness enhancement on the teacher's real-time behavior monitoring image to obtain a global brightness enhanced behavior image; Perform deep convolution optimization on the global brightness enhanced behavior image to construct a resolution convolution optimized image; Based on the resolution convolution optimized image, frame-by-frame action behavior analysis is performed, and operation instructions are intelligently generated to obtain high-level control instructions.
3. The low-latency real-time control method for a teaching screen projection device according to claim 2 is characterized in that: The specific steps of performing frame-by-frame action behavior analysis based on the resolution convolution optimized image and intelligently generating operation instructions to obtain high-level control instructions are as follows: Decompose the resolution convolution optimized image into multiple frames to obtain a time-series frame image sequence; Perform frame-by-frame action behavior visual recognition on the time-series frame image sequence and extract the action behavior features of each frame; Perform teaching control semantic analysis on the action behavior characteristics of each frame to obtain the semantic characteristics of teacher control behavior; Abstract behavioral intention mining is performed on the semantic features of teacher control behavior to obtain behavioral control intention signals; Based on the behavioral control intention signal, operation instructions are intelligently generated to obtain high-level control instructions.
4. The low-latency real-time control method for a teaching screen projection device according to claim 1 is characterized in that: The specific steps of step S2 are: Identify the content of the next frame of teaching images based on high-level control instructions; Calculating the number of elements and the degree of detail of the teaching image content to obtain the visual complexity of the image content; Predicting an image switching time point based on the teaching image content to obtain the image switching time point; Calculate the available preloading time based on the image switching time point to obtain the preloading time window length; A preloading buffering decision is made based on the visual complexity of the image content based on the length of the preloading time window, and a preloading image content buffering strategy is constructed.
5. The low-latency real-time control method for a teaching screen projection device according to claim 1 is characterized in that: The specific steps of step S3 are: Real-time detection of the current network's transmission bandwidth, delay, and packet loss rate to obtain multi-dimensional network status indicators; Perform real-time delay status evaluation based on multi-dimensional network status indicators to obtain a real-time network delay status evaluation value; Divide the teaching image content of the next frame into key areas, calculate the regional priority, and mark the priority transmission level of each area; Adaptive area transmission selection is performed based on the real-time network delay status evaluation value and priority transmission level to obtain the local priority image area and the delayed transmission image area; Local pixel particle compression is performed on the local priority image area and the delayed transmission image area to obtain a local priority pixel stream and a delayed transmission pixel stream.
6. The low-latency real-time control method for a teaching screen projection device according to claim 1, characterized in that: The specific steps of step S4 are: Pre-constructing the teaching scene according to the teaching image content of the next frame to obtain a pre-constructed loaded teaching scene; Extracting buffered interaction paths based on preloaded image content buffering strategy; The projection terminal receives the local priority pixel stream based on the buffered interaction path; Performing local filling processing on the pre-built loaded teaching scene according to the local priority pixel flow to obtain a local image filling scene; Perform fast temporary screen rendering on the local image filling scene to obtain a local pixel rendering image.
7. The low-latency real-time control method for a teaching screen projection device according to claim 1, characterized in that: The specific steps of step S5 are: Identify the transmission resource occupancy information of the local priority pixel stream; Based on multi-dimensional network status indicators, idle transmission resources are calculated based on transmission resource occupancy information to obtain real-time idle transmission resources; Perform secondary transmission rendering on the delayed transmission pixel stream according to real-time idle transmission resources, and perform fuzzy interpolation processing to obtain a rendered image of the delayed area; The global rendering is reconstructed based on the lagging area rendering image and the local pixel rendering image to construct a global teaching rendering image.
8. The low-latency real-time control method for a teaching screen projection device according to claim 1, characterized in that: The specific steps of step S6 are: A digital watermark timestamp is implanted into the global teaching rendering image, and a self-correction is performed based on the timestamp on the screen fragment captured by the projection terminal to generate a self-correction strategy for the rendering image; Based on the teaching terminal, the teacher's continuous operation is recognized and the real-time frequency of the teaching image switching is extracted; Performing a teacher's real-time behavior rhythm analysis on the real-time frequency to generate a teacher's real-time behavior rhythm feature; According to the real-time behavior rhythm characteristics of teachers, the pre-loaded image content buffering strategy is optimized for pre-loaded frequency resonance, and an intelligent real-time screen projection control model is constructed.
9. A low-latency real-time control device for a teaching screen projection device, characterized in that: The low-latency real-time control method for executing the teaching screen projection device according to claim 1 comprises: The teaching behavior analysis module is used to collect real-time teacher behavior monitoring images based on the teaching terminal, perform frame-by-frame action analysis, and intelligently generate operation instructions to obtain high-level control instructions; The preloading buffer module is used to identify the teaching image content of the next frame based on high-level control instructions, perform content complexity calculation and preloading buffer decision-making, and build a preloading image content buffering strategy; An adaptive transmission module, configured to identify multi-dimensional network status indicators and perform adaptive regional transmission selection on the teaching image content to obtain a local priority pixel stream and a delayed transmission pixel stream; The local rendering module is used to perform local filling processing on the local priority pixel stream based on the preloaded image content buffering strategy and perform fast temporary screen rendering to obtain a local pixel rendering image; The global rendering reconstruction module is used to perform secondary transmission rendering on the delayed transmission pixel stream, and to perform global rendering reconstruction in combination with the local pixel rendering image to construct a global teaching rendering image; The same-frequency resonance optimization module is used to perform self-correction on the rendering of the global teaching rendering image, calculate the real-time frequency of teaching image switching, perform pre-loaded same-frequency resonance optimization, and build an intelligent real-time projection control model.