A video synchronous acquisition method and device based on optical distant images
By using optical telephoto imaging and synchronous acquisition technology, the problem of visual fatigue caused by screen size limitations in learning devices has been solved. This has enabled a stable and clear telephoto viewing experience and behavior monitoring on small-sized display devices, generating learning reports to improve learning efficiency.
Patent Information
- Application Number
- CN202511518183.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing learning display devices suffer from visual fatigue and shortened attention span when viewed at close range for extended periods due to limited screen size. Furthermore, their limited information collection capabilities make it difficult to simultaneously monitor learners' behavior and provide feedback.
A video synchronization acquisition method based on optical distant images is adopted. By calculating the geometric mapping matrix and point spread function, geometric pre-distortion and frequency domain compensation are performed on the content to be synchronized. Combined with attitude parameters to predict the rotation range, multi-view rendering is performed to generate sub-frame sequences to achieve a stable distant image viewing experience under small desktop conditions.
It significantly alleviates the visual burden on learners from prolonged close-up viewing, ensures the stability and clarity of the virtual image, expands the exit pupil range, allows learners to obtain a continuous viewing experience with natural head movements, and provides learning reports to improve learning efficiency.
Smart Images

Figure CN121000964B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video acquisition technology, and in particular to a method and apparatus for synchronous video acquisition based on optical distant images. Background Technology
[0002] Existing learning display devices are mostly electronic whiteboards, tablets, or projection terminals, requiring learners to stare at the screen for extended periods at close range. Due to the limited screen size, learning content often appears cramped during presentation, resulting in low viewing comfort. Prolonged use can easily lead to eye strain, shortened attention span, and decreased learning efficiency.
[0003] Furthermore, existing devices have limited information collection capabilities during the learning process, mostly limited to one-way output of displayed content, making it difficult to simultaneously monitor the learner's behavioral state and thus unable to provide timely feedback during the learning process. Even with cameras or external acquisition devices, their functions are mostly limited to image capture, lacking the ability to analyze the learning content synchronously.
[0004] To address the above issues, this application presents a video synchronous acquisition method and device based on optical distant images. Summary of the Invention
[0005] The technical problem this application aims to solve is to address the shortcomings of existing technologies by providing a video synchronization acquisition method and apparatus based on optical distant images. This method acquires calibration images of the content to be synchronized and user posture parameters. By using the calibration images and a preset calibration sequence, a geometric mapping matrix and point spread function are calculated. Geometric pre-distortion and frequency domain compensation are then performed on the content to be synchronized to obtain a compensated frame. Combined with the rotation range predicted by the posture parameters, time-multiplexed multi-view rendering is used to generate a sub-frame sequence, which is then output from the display end. This achieves a stable distant image viewing experience under small desktop conditions.
[0006] To achieve the above objectives, this application provides the following technical solution:
[0007] A video synchronization acquisition method based on optical distant images is applied to a video synchronization acquisition device. The video synchronization acquisition device includes an acquisition end and a display end. The acquisition end includes a first acquisition device and a second acquisition device. The first acquisition device is used to acquire a calibration image of the content to be synchronized, and the second acquisition device is used to acquire the user's posture parameters. The method includes:
[0008] Based on the acquired calibration image, and combined with the preset calibration sequence, calculate the geometric mapping matrix used to map the calibration image to the display end and the point spread function characterizing spatially variable imaging degradation;
[0009] Geometric pre-distortion is performed on the content to be synchronized according to the geometric mapping matrix, and frequency domain compensation is performed on the geometrically pre-distorted content to be synchronized according to the point spread function to obtain a compensated frame.
[0010] The rotation range is predicted based on the collected attitude parameters, and the compensation frame is rendered using time-reused multi-view rendering based on the rotation range to generate a sub-frame sequence and output it through the display terminal.
[0011] The video synchronization acquisition device further includes a cloud platform, which is communicatively connected to the acquisition terminal. The method further includes:
[0012] During the operation of the display terminal, the first video stream corresponding to the display terminal is acquired, and the second video stream corresponding to the user is acquired through the second acquisition device. The first video stream includes the virtual image video displayed on the display terminal.
[0013] Upload the first video stream and the second video stream to the cloud;
[0014] Analyze the first video stream and the second video stream to extract user behavior data and display content data;
[0015] The user behavior data and display content data are input into a preset large model, and a learning report is generated through the large model. The learning report includes at least one of the following: user focus, learning progress, and behavior pattern.
[0016] The geometric mapping matrix is calculated as follows:
[0017] The background image corresponding to the content to be synchronized is obtained based on the calibration image, the anisotropic power spectrum of the background image is calculated, and the edge lines and rectangle vertices of the device are obtained through morphological detection and Hough transform, wherein the device is the terminal corresponding to the content to be synchronized.
[0018] A sparse geometric constraint set is constructed based on the edge lines and rectangle vertices, and two corresponding vanishing lines are obtained. The initial value geometric mapping matrix corresponding to the sparse geometric constraint set is calculated by direct linear transformation.
[0019] Using the principal direction of the anisotropic power spectrum as a regular a priori constraint, the initial value geometric mapping matrix is globally nonlinearly refined to obtain the geometric mapping matrix.
[0020] The calculation of the initial value geometric mapping matrix corresponding to the sparse geometric constraint set through direct linear transformation includes:
[0021] Based on the field of view of the first acquisition device, the edge line length, straightness, corner confidence and local contrast of the sparse geometric constraint set are weighted and corrected to obtain the corrected constraint set.
[0022] Based on the modified constraint set, construct consistency constraints for point-to-point correspondence, line-to-line correspondence, and parallelism of two vanishing lines;
[0023] Calculate the system of linear equations based on the aforementioned consistency constraints;
[0024] The linear equations are solved using the least squares method, and the initial mapping result is obtained by combining the results with singular value decomposition.
[0025] Based on orthogonality consistency, the initial mapping result is scaled and sign-consistent to obtain the initial value geometric mapping matrix.
[0026] The preset calibration sequence is based on the calibration pattern of the narrowband RGB three channels of the calibration image. The calibration pattern includes a sinusoidal stripe grid in four directions, where the directions include 0 degrees, 45 degrees, 90 degrees and 135 degrees. The calibration pattern includes at least three spatial frequencies corresponding to the three channels.
[0027] The point spread function is calculated as follows:
[0028] The response frames of each channel in the calibration image are projected onto the virtual image plane to obtain the response space;
[0029] In the response space, the calibration patterns of the three narrowband RGB channels are output sequentially, and phase demodulation is performed on each calibration pattern to obtain multiple modulation responses;
[0030] The response space is divided into multiple spatial partitions according to the sinusoidal fringe grid. The modulation response is weighted and fitted by combining the spatial frequency corresponding to the spatial partition to obtain the contrast transfer curve corresponding to the spatial partition.
[0031] The contrast transfer curve is nonlinearly fitted to obtain the principal axis length, orientation angle and energy normalization value of the point diffusion kernel corresponding to the spatial partition;
[0032] The point diffusion function is obtained by weighted summation of the principal axis length, orientation angle, and energy normalization value of the point diffusion kernel corresponding to each spatial partition through global smoothing and spatial cubic interpolation.
[0033] The obtained compensation frame includes:
[0034] The content to be synchronized is remapped in reverse coordinates according to the geometric mapping matrix, and the coordinates are filtered by Lanzos interpolation to obtain the remapping result.
[0035] The remapping result is subjected to mirror boundary expansion and block overlap processing to obtain multiple mapping partitions with different spatial frequencies;
[0036] For each mapping partition, deconvolution is performed in the frequency domain based on the point spread function to obtain the partition compensation result;
[0037] The partition compensation results are weighted and fused to output the compensation frame, wherein the weights are set according to the channel co-address of each mapping partition, and the channel co-address is calculated based on the edge position offset of the three color channels of the mapping partition in the remapping result.
[0038] Multi-view rendering of the compensated frame based on the rotation range, including time multiplexing:
[0039] During at least one refresh cycle of the display, a viewing angle domain is generated based on the rotation range;
[0040] The core viewpoint cluster and boundary viewpoint cluster of the viewpoint domain are calculated by clustering algorithm. The core viewpoint cluster and boundary viewpoint cluster are optimized by combining preset viewpoint coverage threshold and brightness equalization constraint to obtain the allocation duty cycle and exposure share of each viewpoint.
[0041] Based on the allocated duty cycle and exposure share, the compensation frame is parallax remapped, and the display position of the compensation frame is adjusted in combination with the time sequence of the rotation range. A subframe sequence arranged alternately according to the core viewpoint cluster and the boundary viewpoint cluster is generated and output.
[0042] A video synchronization acquisition device based on optical distant images, the device comprising:
[0043] The acquisition end is used to acquire calibration images of the content to be synchronized and to acquire the user's posture parameters;
[0044] The display end is used to calculate the geometric mapping matrix and point spread function corresponding to the calibration image, perform geometric pre-distortion, frequency domain compensation and multi-view rendering on the content to be synchronized, and display the virtual image video;
[0045] The cloud-based system communicates with the acquisition and display terminals to receive and analyze the first video stream corresponding to the display terminal and the second video stream corresponding to the acquisition terminal, extract user behavior data and display content data, and generate a learning report through a preset large model.
[0046] The acquisition terminal includes:
[0047] The first acquisition device is used to acquire the calibration image of the content to be synchronized;
[0048] The second acquisition device is used to collect the user's posture parameters and behavioral feature data, and output posture information and video streams for rotation range prediction and user attention analysis.
[0049] Compared with the prior art, the beneficial effects of this application are:
[0050] This application introduces an optical far-image imaging and synchronous acquisition scheme within a small desktop footprint, enabling small-sized display content to be presented on a far-image virtual plane, thus significantly alleviating the visual burden on learners during prolonged close-range viewing. Through joint calculation of the geometric mapping matrix and point spread function, precise compensation for imaging distortion and spatial degradation is achieved, ensuring the stability and clarity of the virtual image. Combining posture parameters to predict the rotation range and performing multi-viewpoint rendering expands the exit pupil range, allowing learners to maintain a continuous viewing experience even with natural head movements. Attached Figure Description
[0051] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0052] Figure 1 This is an exemplary application scenario diagram of an embodiment of this application;
[0053] Figure 2 This is a schematic diagram of the structure of a video synchronous acquisition device based on optical distant images according to an embodiment of this application;
[0054] Figure 3 This is a flowchart illustrating a video synchronization acquisition method based on optical distant images, according to an embodiment of this application.
[0055] Reference numerals: 100, device body; 101, display screen; 102, second acquisition device; 103, first acquisition device. Detailed Implementation
[0056] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0057] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0058] In real classroom and home learning scenarios, learners are consistently faced with viewing conditions of "close-up content and short viewing distance":
[0059] While tablets, electronic whiteboards, and short-distance projection are easy to deploy, the prolonged close-range viewing conditions make it almost impossible to avoid issues such as visual fatigue, fluctuating attention, and posture shifts in users.
[0060] At the same time, the teaching process is not just about displaying content; it also involves the generation and passive recording of learning behaviors, such as facial expressions, slight head turns, gaze paths, and small actions like holding a pen and turning pages. These all constitute key clues for teaching feedback.
[0061] In the example technologies, either the projection distance is increased, sacrificing usable desktop space, or the content acquisition and behavior acquisition are placed on different devices, which often leads to a disconnect between geometric relationships and the timeline. Especially in environments with strong reflections, the display and capture interfere with each other, making it difficult to achieve both clear images and reliable synchronization at the same time.
[0062] This application focuses on a trade-off that is closer to everyday working conditions:
[0063] Under the constraints of small desktop size, short optical path and limited brightness budget, the near-screen light field is reconstructed into a far-image perception through the optical far-image principle, so as to reduce the learner's adjustment burden.
[0064] It is easy to understand that the distant image in this application does not rely on large optical devices or long projection strokes, but rather, within a limited shape envelope, with the goal of virtual image formation and angular tolerance management, combined with a calibration sequence that can be stably identified by the imaging link and weakly intrusive visual cues, the originally scattered display, calibration and acquisition processes are converged into the same virtual image plane coordinate system.
[0065] Unlike hardware optimization solutions that often involve replacing lenses or adding optical screens, the directional degradation is more pronounced in short optical paths, the eyebox is smaller, and chromatic aberration and geometric distortion are rapidly amplified at the edges. Any slight head movement is enough to make small print that is just barely legible cross the readability threshold. And as long as the estimation of geometric mapping and dot spread drifts even slightly, the correspondence between content and behavior will be distorted.
[0066] Those skilled in the art will understand that existing electronic whiteboards or conventional screen projections take visibility as the end goal, ignoring the experiential threshold of seeing correctly and for a longer period of time; while AR / VR devices can create a sense of distance, they introduce wearing burdens and social isolation, and are not conducive to the interaction of open classrooms; the soft splicing solution of screen recording plus external camera will drift in geometric alignment and time reference when there are changes in lighting, screen reflections and slight movement of the device, and the back-end analysis has to rely on fragile feature matching, and the stability is difficult to meet the needs of daily teaching.
[0067] It is important to emphasize that in actual remote teaching, there is no rigid drive of the common data surge. The problem is not that the video cannot be transmitted, but how to continuously and reliably use the resolution budget on key content without increasing the complexity of deployment and maintenance costs.
[0068] Based on the foregoing, the proposed solution does not rely on large-scale hardware modifications, but rather takes the physical consequences of a small desktop size as its starting point:
[0069] Stable geometric and imaging degradation priors are constructed using observable near-field textures and channel dispersion under short optical paths; the sharpness and directionality loss of the virtual image plane are quantified in extremely short time slots using narrow-band, multi-directional calibration sequences; within the predicted head rotation range, angular tolerance is extended in a limited brightness budget through time-reused multi-view rendering, so that key text and formulas remain readable under natural head movements.
[0070] refer to Figure 1 , Figure 1 This is an exemplary application scenario diagram provided for an embodiment of this application.
[0071] Figure 1 The application scenario is shown to include the acquisition end, the display end, and the cloud, which interact with each other through wired or wireless transmission networks; each unit can be set up separately or integrated into different functional modules in the same device.
[0072] In the first aspect, the acquisition end obtains data for imaging calibration and behavior analysis. The acquisition end can also perform lightweight coding and privacy desensitization, send synchronization anchors, feature summaries and optional short backtracking fragments for cloud alignment to the cloud, and send calibration images and pose parameters to the display end for real-time rendering control.
[0073] Secondly, the display end is used to realize optical image display and synchronous rendering control in a small desktop layout. Based on the calibration image and its response frame provided by the acquisition end, the geometric mapping matrix and spatially variable point spread function for virtual image plane mapping are calculated. Based on this, geometric pre-distortion and frequency domain compensation are performed on the content to be synchronized to obtain the compensation frame. Combined with the rotation range given by the acquisition end, the compensation frame is time-multiplexed and rendered from multiple viewpoints. Invisible synchronization anchor points are embedded in the output subframe for subsequent video alignment. The display end outputs the virtual image video as the first video stream and can selectively transmit frame-level metadata back.
[0074] Thirdly, the cloud is used to complete the alignment, analysis and result generation of multi-source videos, and, under compliant conditions, calls large models to evaluate the learning process and generate interpretable results, outputting learning reports including focus, progress and behavior patterns; cloud results can be sent back to the display end for rendering strategy loop, or sent to the user terminal for viewing and management.
[0075] It is understood that the names and divisions of the aforementioned functional modules are merely examples, and in actual implementation, they can be trimmed, merged, or replaced according to application requirements.
[0076] refer to Figure 2 , Figure 2 This is a schematic diagram of a video synchronous acquisition device based on optical distant images, provided as an embodiment of this application.
[0077] Figure 2 The video synchronization acquisition device shown includes a device body 100, a display screen 101, a second acquisition device 102, and a first acquisition device 103. The lower dashed line in the figure indicates the installation position of the display screen 101 inside the device body 100, and the upper dashed line in the figure indicates the user's line of sight.
[0078] Furthermore, the device body 100 is designed to achieve integrated packaging for optical image display and synchronous acquisition within a small desktop footprint. Its interior can accommodate the display screen 101 and its associated optical components. The display screen 101, as the main output component, presents content after geometric pre-distortion and frequency domain compensation. The second acquisition device 102 is located below the device body 100 and is used to acquire the user's posture parameters and behavioral characteristics for rotation range prediction and learning process analysis. The first acquisition device 103 is located in the mounting support area of the device body 100 and is used to acquire calibration images and calibration sequence response frames output by the display screen 101, providing data support for the calculation of the geometric mapping matrix and point spread function.
[0079] With the cooperation of the aforementioned components, the device in this application embodiment can achieve simultaneous content display and user behavior collection within a limited volume, and effectively alleviate visual fatigue of learners by combining the optical far-image imaging principle.
[0080] It is understood that the first and second acquisition devices of this application may be image sensors, depth cameras, infrared cameras, or composite sensing modules with attitude tracking functions, or acquisition units that combine optical and inertial sensing. This application does not impose further limitations on these devices in this and subsequent embodiments.
[0081] In an embodiment not shown in the figure, the device further includes a processor, which is used to preprocess the calibration image output by the first acquisition device and calculate the geometric mapping matrix and point spread function based on the calibration sequence; the processor is also used to predict the rotation range according to the user posture parameters acquired by the second acquisition device, and control the display screen to output the subframe sequence after geometric pre-distortion, frequency domain compensation and time multiplexing rendering, thereby realizing real-time linkage between display and acquisition locally.
[0082] In an embodiment not shown in the figure, the device is also connected to the cloud, which embeds a large language model. The large language model is used to perform semantic analysis and behavior recognition on the user video stream uploaded by the acquisition terminal and the video stream of the display terminal, and generate a learning report including learning focus, learning progress and behavior patterns. At the same time, the large language model can also perform long-term modeling of the user's learning habits and send personalized rendering optimization parameters or learning feedback suggestions back to the display terminal to improve the learning experience and teaching assistance effect.
[0083] Next, with reference to the accompanying drawings, a video synchronization acquisition method based on optical distant images provided in this application will be further described. Figure 3 The method shown is applied to a video synchronization acquisition device, which includes an acquisition end and a display end. The acquisition end includes a first acquisition device and a second acquisition device. The first acquisition device is used to acquire a calibration image of the content to be synchronized, and the second acquisition device is used to acquire the user's posture parameters. The method includes:
[0084] S1: Based on the acquired calibration image, calculate the geometric mapping matrix used to map the calibration image to the display end and the point spread function characterizing spatially variable imaging degradation, combined with the preset calibration sequence.
[0085] In this embodiment, the preset calibration sequence is composed of a narrowband RGB three-channel sinusoidal stripe pattern. Stripes of different directions and frequencies are used to extract the geometric distortion and spatial degradation features of the imaging system.
[0086] The first acquisition device acquires these calibration images within a limited sampling time and extracts the geometric constraint relationship between the display end and the acquisition end through power spectrum directionality analysis and corner detection, thereby constructing a geometric mapping matrix. Simultaneously, it estimates the point spread function using frequency response variations and obtains spatially variable degradation features by combining partitioned fitting. In this way, the extraction of mapping relationships and degradation functions can be completed quickly and stably under the constraints of a small desktop volume and short optical path.
[0087] S2: Perform geometric pre-distortion on the content to be synchronized according to the geometric mapping matrix, and perform frequency domain compensation on the geometrically pre-distorted content to be synchronized according to the point spread function to obtain a compensated frame;
[0088] In this embodiment, the compensation process first performs inverse mapping on the original teaching content to eliminate the inherent geometric distortion of the display end, and then combines the point spread function to perform frequency domain deconvolution to restore the blurred spatial frequency components.
[0089] In practical applications, a combination of Lanzos interpolation and regularized deconvolution can be used to ensure the sharpness of the compensation while avoiding excessive noise amplification or ringing artifacts. The final output compensated frame maintains the clarity of edges and details, ensuring that the virtual image content output to the display remains highly readable even in complex desktop environments such as low brightness and strong reflections. This concentrates limited display resources on the clear presentation of key content, thereby reducing visual fatigue for learners.
[0090] S3: Predict the rotation range based on the collected attitude parameters, perform time-multiplexed multi-view rendering on the compensation frame based on the rotation range, generate a sub-frame sequence and output it through the display terminal;
[0091] In this embodiment, the second acquisition device captures the user's head posture and rotation trend to predict the user's possible viewing angle range within a refresh cycle, classifying it into core viewpoints and boundary viewpoints. During output, the display terminal performs parallax remapping on the compensation frames based on these viewpoints, obtaining multiple subframes, which are then output sequentially using time multiplexing. To ensure a good viewing experience, core viewpoints are allocated a higher duty cycle and brightness, while boundary viewpoints are guaranteed sufficient angular coverage. This ensures that even if the learner deviates from the center position during normal head movements, key content remains clearly perceived, achieving stable presentation of optically distant virtual images. This solves the problem of limited exit pupil in small-volume display devices, allowing learners to enjoy a comfortable viewing experience without maintaining a rigid posture.
[0092] Before delving into the specific technical details of the steps, the embodiments of this application need to be emphasized again.
[0093] This application uses the virtual image plane as the sole reference coordinate, and all display, acquisition and alignment processes are normalized to the virtual image plane; during the startup phase, only a narrow-band, multi-directional stripe sequence is output in a very short time slot to complete the quantization of geometric mapping and direction-related degradation without interrupting the teaching presentation; during the subsequent operation phase, it no longer relies on the complete video stream, but continuously generates synchronization anchors with timestamps and behavioral feature summaries to maintain consistent content-behavior alignment, and only triggers second-level backtracking segments for review when necessary.
[0094] Those skilled in the art will understand that posture data can be collected according to the actual situation. Specific types can include posture sequences, blink rate, ambient illumination, edge intensity, etc. It is sufficient to support the calculation of mapping update, corner domain allocation and behavior alignment to a minimum. This application does not impose any further limitations.
[0095] In this embodiment, narrowband RGB stripes are output during the interval between courseware page switching or still frame pause. The first acquisition device performs phase demodulation and power spectrum directionality analysis on the response frame to obtain the mapping relationship of the virtual image plane and the degradation parameters of each partition. When a slight pose drift is detected between the display end and the acquisition end, only a single-direction, single-frequency minimal subsequence is replayed to complete the rapid refinement, avoiding disruption to the viewing continuity.
[0096] It is easy to understand that the aforementioned content transforms the directional degradation and dispersion differences caused by small volume and short optical path into a weak perturbation that can be continuously tracked. Once a drift occurs, it can be corrected in milliseconds. Without additional optical structures or site modifications, it can maintain the long-term stability of the virtual image plane under desktop constraints.
[0097] Furthermore, the core of content presentation in this application is not limited to sharpening, but rather uses corner-domain arrangement as the driving force to arrange subframes:
[0098] The attitude parameters generated by the second acquisition device are integrated into a range of achievable rotation. The rendering end constructs a set of viewpoints based on this and solves the brightness duty cycle, exposure share, and angular windowing together into a timing table. The high-probability viewpoints obtain denser subframes and more stringent angular windowing, while the boundary viewpoints maintain necessary coverage and flexible brightness. Under a limited brightness budget, the effective angular domain of the exit pupil is expanded, so that the key text remains in the clear zone during natural head movements. Stable readability can be maintained without forcing a fixed posture, while avoiding flickering and crosstalk caused by high-frequency subframe switching.
[0099] Furthermore, in some optional implementations, the display end embeds synchronization anchors in the invisible layer of the subframe, and the first acquisition device uses this to recover the content time base and geometric position. What is uploaded to the cloud is not the original frame, but the anchor sequence and behavioral feature summary. Only when there is a conflict between salient content and focus is a very short evidence fragment sent back and erased locally upon expiration. The most critical alignment and interpretation are completed with the minimum amount of data, while providing auditable evidence for backend evaluation, reducing bandwidth and compliance risks, and ensuring that the learning report generated by the large model can be traced back to a clear content fragment and time point.
[0100] In this application, the module division is not limited to physical form:
[0101] The first acquisition device can be an imaging module for virtual images, and the second acquisition device can be an RGB / IR or depth module for users. The processing unit can reside in the device or work in collaboration with the cloud.
[0102] Those skilled in the art will understand that different deployment methods can be replaced as long as they meet the minimum requirements of unified virtual image plane coordinates, executable angular domain timing, and decodeable anchor points.
[0103] The common thread in the aforementioned approaches is that they transform the physical constraints of small volume and short optical path into computable priors and executable timing tables, thereby achieving stable far-image rendering in a desktop environment.
[0104] Next, we will further elaborate on the technical content of the geometric mapping matrix in this application.
[0105] It is understood that the virtual image plane described in this application refers to the equivalent display plane formed after refraction or reflection by the optical image structure at the display end. It does not depend on the physical position of the display screen, but is the plane where the virtual image is perceived by the learner's eyes under certain viewing distance conditions.
[0106] Those skilled in the art will understand that the virtual image plane can be determined by the focal point and projection path of the optical system, and its position has a stable geometric mapping relationship with the actual display screen.
[0107] The geometric mapping matrix is used to describe the correspondence between the calibration image acquired by the acquisition end and the virtual image plane of the display end. Through this matrix, the pixels in the coordinate system of the acquisition end can be mapped to the unified coordinate system of the virtual image plane.
[0108] In this embodiment, the process of obtaining the geometric mapping matrix not only includes conventional corner detection and vanishing line calculation, but also introduces the anisotropic power spectrum of the background texture as a regularization prior. This ensures that a stable mapping relationship can still be obtained even under conditions of small desktop volume and short optical path, where the acquisition angle is limited or there is local occlusion. This application can effectively avoid the matrix drift problem caused by insufficient edge detection or sparse point set, thereby ensuring the geometric consistency of the virtual image content in the user's eyes.
[0109] In one example, the geometric mapping matrix is calculated as follows:
[0110] S1.1: Obtain the background image corresponding to the content to be synchronized based on the calibration image, calculate the anisotropic power spectrum of the background image, and obtain the edge lines and rectangle vertices of the device through morphological detection and Hough transform, wherein the device is the terminal corresponding to the content to be synchronized.
[0111] Specifically, in order to obtain stable geometric constraints under the conditions of small desktop volume and short optical path, it is necessary to peel off the content layer from the calibration image and construct background features that can reflect the relationship between the virtual image plane and the terminal bounding box. The background image is used to provide two types of information:
[0112] One is the geometric elements of the terminal frame and the display opening, which facilitates the subsequent establishment of point-line sparse constraints;
[0113] Secondly, the texture statistics, which exhibit directional bias due to the short optical path and small aperture, are used to determine the main direction of the power spectrum and play a regularization role in the refinement stage.
[0114] In some alternative implementations, the background image can be created using a spatial mask, for example, by selecting the frame set with the least content change, suppressing the illumination base through brightness stabilization and local contrast equalization, and then removing transiently superimposed content elements using temporal median calculation, so that the remaining textures and edges can better represent the real geometric relationships and optical degradation trends.
[0115] In this embodiment, the anisotropic power spectrum of the background image is processed by a block window function and then statistically analyzed in the frequency domain, specifically including the following three steps:
[0116] In the first step, the background image is first mapped from color to brightness, and guided filtering is used to suppress large-scale shadows. Then the image is divided into several overlapping small blocks, and each block is multiplied by a cosine window to reduce the impact of edge leakage.
[0117] In the second step, the frequency domain amplitude distribution of each block is calculated, and the cumulative amplitude distribution is distributed in several angular sectors and ring frequency bands to form a direction histogram. In order to avoid the direction bias caused by text residue, before the angular sector statistics, the slender high contrast region is removed by the content mask based on the connected component. At the same time, the specular highlights are marked and removed from the statistics using the saturation and high gradient ratio detection method.
[0118] In the third step, the orientation histograms of each block are spatially robustly averaged to obtain the global principal orientation and the confidence interval. Based on the global principal orientation and the confidence interval, the anisotropic power spectrum is calculated.
[0119] Furthermore, edges and corners are extracted using a combination of multi-threshold Canny and morphological closing operations. After probabilistic Hough transform, two sets of approximately parallel long straight lines are obtained. Within the straight line sets, iterative segment growth and least squares fitting are used to improve straightness. Intersection points are obtained using a sub-pixel corner point localization algorithm. Finally, candidate sets of the four corners of the terminal rectangle and set of edge straight lines are obtained.
[0120] It is understood that the device described in this application refers to a terminal used to carry content to be displayed, including but not limited to mobile phones, tablets, laptops, and other display devices capable of outputting image or video signals. It should be noted that the device's role in this application is limited to a content source; there are no restrictions on the applications running within it or the files stored there, as long as they can generate teaching materials, learning documents, courseware demonstrations, or other learning-related content, they can be transmitted to the display terminal via the display link. Under this definition, the display terminal receives the screen signal output by the device, not the device's own operating interface or operating environment. Therefore, the method of this application is applicable to various types of content terminals and does not require modification for specific models.
[0121] S1.2: Construct a sparse geometric constraint set based on the edge lines and rectangle vertices, and obtain the corresponding two vanishing lines. Calculate the initial value geometric mapping matrix corresponding to the sparse geometric constraint set through direct linear transformation.
[0122] Specifically, the construction of sparse geometric constraints aims to establish a preliminary correspondence between the image at the acquisition end and the virtual image plane using the fewest and most stable geometric elements. Based on the two sets of parallel lines obtained in S1.1, the corresponding vanishing lines are estimated respectively; the four corners and two long sides of the rectangle are used as the main point-line constraints, supplemented by the visible inner frame or edge lines within the display area to form additional line constraints.
[0123] In one example, the specific steps of S1.2 are as follows:
[0124] S1.2.1: Based on the field of view of the first acquisition device, the edge line length, straightness, corner confidence and local contrast of the sparse geometric constraint set are weighted and corrected to obtain the corrected constraint set;
[0125] Specifically, the small size of the desktop and the short optical path limit the working distance between the first acquisition device and the terminal display plane. The edge field area shows obvious perspective shortening and directional degradation in imaging. The visible length and contrast of the actual edge of the same line in the online and offline directions are not symmetrical. The positioning reliability of corner points differs greatly between positions close to and far from the optical axis.
[0126] In this embodiment, to avoid the appearance of longer and brighter dominant initial values under such limited field of view conditions, a weighting strategy coupled with the field of view angle needs to be introduced:
[0127] For each edge line, first estimate its angle with the optical axis and its polar position on the imaging plane, and reduce the length weight of the edges that are more affected by perspective shortening in advance.
[0128] The confidence score for corner points is determined by joint scoring, which automatically reduces the weight of corner points that are close to the edge field and whose contrast is affected by the illuminance roll-off.
[0129] For local contrast, the contrast enhancement after adaptive histogram equalization and guided filtering is used as a reference. When the enhancement is too large, it reflects that the original contrast is insufficient. The geometric constraints of this area are also weighted and suppressed.
[0130] It is easy to understand that the aforementioned method can solve the asymmetric imaging problem caused by the difference between the horizontal and vertical fields of view, and avoid the angular offset caused by the small volume of the desktop being mistakenly treated as the real geometry in the solution.
[0131] Furthermore, the field of view is stored in two dimensions: one is the horizontal / vertical field of view range given by the manufacturing parameters, and the other is the effective field of view obtained by rapid backtesting through the calibration sequence during operation, which is used to reflect the actual coverage area caused by the current placement and pitch.
[0132] Furthermore, the weighting adjustment retrieves coefficients from the cross-tabulation of these two types of parameters, applying a quaternary weighting to edge line length, straightness, corner confidence, and local contrast. The cross-tabulation can be set empirically.
[0133] In some optional implementations, straightness is measured by the residual area of a piecewise linear fit, taking into account local bends caused by tabletop reflections. Samples at bends are removed from the straightness calculation window to avoid local reflection errors from inflating the fitting residuals. After weight correction, the contribution of the constraint set in the edge field and paraxial region tends to be balanced, reducing the impact of observation bias caused by short optical paths on the initial value matrix.
[0134] S1.2.2: Based on the modified constraint set, construct consistency constraints for point-to-point correspondence, line-to-line correspondence, and parallelism of two vanishing lines;
[0135] Specifically, in limited field of view, issues such as edge occlusion, missing corners, or reflection stripes often arise, and single-type constraints easily degenerate into ill-conditioned equations. To ensure stable and solvable linear constraints in small desktop layouts, point-to-point and line-to-line constraints need to be combined, along with a consistency constraint of two parallel vanishing lines. This constraint ensures that the two sets of opposite edges in the terminal display area still correspond to the same parallel direction after mapping. Point-to-point constraints are used to anchor the topological relationship between the four corners of the rectangle and the standard rectangle of the virtual image plane; line-to-line constraints are used to maintain the geometric consistency of the entire straight line of the edge; and the parallel consistency of vanishing lines is used to constrain the parallel image from shearing drift even when corners are missing. Since the visible length difference between the two sets of opposite edges is significant in small volumes, parallel consistency can compensate for the length imbalance, preventing large-angle twisting even when one corner is missing.
[0136] S1.2.3: Calculate the system of linear equations based on the aforementioned consistency constraints;
[0137] In this embodiment, the linear equation system is constructed in blocks for parallel verification: point-to-point submatrices, line-to-line submatrices, and parallel consistency submatrices are generated and then the condition number is evaluated. If the condition number of a certain submatric exceeds a set threshold, the corresponding constraint is resampled and replaced.
[0138] S1.2.4: Solve the linear equation system using the least squares method and obtain the initial mapping result by combining it with singular value decomposition;
[0139] Specifically, linear equation systems under limited field of view often exhibit approximate rank deficiency or excessively large condition numbers, making conventional least squares susceptible to noise amplification and leading to unstable initial mappings. Orthogonal decomposition of the linear system using singular value decomposition (SVD) can numerically separate directions that contribute less to the solution, suppressing the corresponding small singular value components and thus reducing the amplification effect of noise in the solution. Least squares solutions are then performed based on SVD, prioritizing the retention of components corresponding to dominant singular vectors, thereby reducing the sensitivity of the initial mapping results to observation noise and outliers.
[0140] It is easy to understand that the least squares method and singular value decomposition are existing technologies. This application will not elaborate on how to solve the linear equation system by the least squares method and obtain the initial mapping result by combining it with singular value decomposition.
[0141] It should be noted that, in order to avoid information loss caused by excessive truncation, the singular value threshold in this application is not fixed, but is related to the field of view angle and the distribution of constraint weights: when the field of view is narrow and the constraints are concentrated in a single direction, the threshold is raised appropriately to prevent the single-direction constraint from overwhelmingly dominating; when the field of view coverage is relatively balanced, the threshold can be lowered appropriately to retain more geometric details.
[0142] S1.2.5: Based on orthogonality consistency, the initial mapping result is scaled and sign-consistent to obtain the initial value geometric mapping matrix;
[0143] Specifically, the initial mapping obtained by linear solution has scale and sign uncertainties, and is prone to slight mirror or rotation errors under limited field of view. Therefore, normalization and consistency correction are required in the post-processing stage. Scale normalization uses the known side length ratio of the terminal display area as a reference, aligning the mapped standard rectangle with the observed edge in length ratio, ensuring that the scaling relationship conforms to the physical shape of the terminal. Sign consistency is determined by checking the local sign of the mapped Jacobian and the monotonicity of the pixel arrangement to determine if mirror flips exist. Once a flip is found, the sign of the corresponding column or row of the mapping matrix is corrected to ensure that the mapping maintains a right-hand direction.
[0144] S1.3: Using the principal direction of the anisotropic power spectrum as a regular a priori constraint, the initial value geometric mapping matrix is globally nonlinearly refined to obtain the geometric mapping matrix;
[0145] Specifically, the initial geometric mapping matrix can only guarantee that the mapping satisfies the constraints within a local region, but may cause offsets globally due to insufficient edge points or detection errors. Therefore, further correction is needed through global nonlinear refinement. In this embodiment, the core idea of refinement is to introduce the principal direction of the power spectrum as a regular prior into the optimization process to ensure that the overall directionality on the virtual image plane is consistent with the mapping relationship.
[0146] In this embodiment, optimization is performed iteratively, with the initial value being the initial geometric mapping matrix. In each iteration, the result of the mapping matrix applied to the calibration image is first calculated, and gradient direction histograms are extracted in blocks on the virtual image plane. These histograms are then compared with the pre-obtained principal directions of the power spectrum, and the angle difference is calculated as a regularization constraint term. The regularization constraint term and the traditional geometric reprojection error term together constitute the objective function, and the two are combined in a weighted sum form. To ensure the stability of the optimization, the weights in the objective function are determined by the energy concentration of each block; regions with higher energy concentration contribute more to the objective function, thereby avoiding noise dominating the optimization process.
[0147] Furthermore, this embodiment employs the Levenberg-Marquardt algorithm for optimization, which combines the fast convergence of Gauss-Newton's method with the stability of gradient descent, making it suitable for handling nonlinear problems with local minima.
[0148] In each iteration, the Jacobian approximation of the objective function with respect to the matrix parameters is first calculated, and then the update step size is adaptively adjusted in conjunction with the damping factor. If an update decreases the residual, the damping factor is reduced to accelerate convergence; if the residual increases, the damping factor is increased to reduce the step size and avoid oscillations. Through this dynamic adjustment, the optimization can quickly approach the optimal solution in the early stages and maintain stable convergence in the later stages.
[0149] To ensure the reliability of the optimization results, this embodiment sets two types of termination conditions. One type is an upper limit on the number of iterations, which forcibly stops when the iteration reaches a preset number of iterations without convergence; the other type is a convergence threshold, which is considered to have converged when the objective function decreases by less than a set threshold and the parameter update amount is less than a predetermined value in two adjacent iterations.
[0150] It is understood that the preset number of iterations and convergence threshold in this application can be determined by those skilled in the art through a large number of experiments, and will not be elaborated here.
[0151] Next, we will further elaborate on the technical content of the point spread function in this application.
[0152] In this embodiment, the preset calibration sequence is constructed based on the calibration pattern of the narrowband RGB three channels of the calibration image. The calibration pattern includes a sinusoidal stripe grid in four directions, wherein the directions include 0 degrees, 45 degrees, 90 degrees and 135 degrees. The calibration pattern includes at least three spatial frequencies corresponding to the three channels.
[0153] In one example, the generation of the calibration pattern follows a composite constraint of "narrow-band color separation, fixed orientation, graded frequency, phase shift coding, and short-time playback".
[0154] Specifically, to quantify orientation-dependent image degradation under the conditions of small desktop volume and short optical path, the pattern needs to simultaneously possess orientational discernibility, frequency coverage, and the ability to suppress ambient light on the virtual image plane, while minimizing disruption to the continuous presentation of teaching content. The calibration pattern was thus designed as a narrowband RGB three-channel sinusoidal stripe grid, with each channel playing independently. Each channel contains four fixed directions: 0°, 45°, 90°, and 135°, with at least three spatial frequencies configured in each direction.
[0155] In some optional implementations, the absolute values of the three spatial frequencies are not fixed constants, but rather selected from the reconstructable frequency bandwidth based on the pixel pitch of the display screen, the sampling resolution of the first acquisition device, and the effective field-of-view matching relationship between the two. This ensures that the lowest level is far from low-frequency illumination fluctuations and backlight unevenness, the middle level covers common text strokes and thin line layers, and the highest level touches the energy roll-off zone of the point spread function, so as to stably sample the direction-frequency two-dimensional space.
[0156] Those skilled in the art will understand that specific data can be collected according to the actual situation, and the specific types can be pixel pitch, focal length estimation, shooting distance, effective field of view, etc. It is sufficient to deduce the upper limit of the displayable frequency and the safe bandwidth to the minimum extent. This application does not impose any further limitations.
[0157] In this embodiment, narrowband color separation is achieved through the RGB native sub-pixels or equivalent narrowband display paths of the display end. Only one color channel is activated within the same time window to avoid spectral crosstalk between channels affecting the estimation of their respective image degradation. To ensure that each channel is within the approximately linear photoelectric response range, the color gamut / grayscale measurement results of the display end are first called to map the desired stripe contrast to the panel input code value. Linearization is performed using an inverse EOTF lookup table, followed by bit depth dithering and spatial error diffusion to reduce additional harmonics caused by quantization steps. To further reduce the beat frequency phenomenon with the pixel grid, the stripe period is avoided from being aligned with the sub-pixel period or an integer multiple thereof. Since dispersion is more easily observed under small volume and short optical path, the difference in peak wavelength of the three channels helps to separate the degradation characteristics of different bands. Therefore, the frequency levels of the three channels are kept consistent while the brightness duty cycle is allowed to be finely adjusted in exchange for modulation signal-to-noise alignment across channels.
[0158] Furthermore, the calibration pattern employs multiphase shift encoding and DC removal in the time domain to suppress the influence of ambient light and backlight fluctuations. Stripes of each direction and frequency are played in short bursts in a three- or five-phase sequence with uniformly distributed phase intervals. During playback, a zero-mean micro-amplitude brightness balance is superimposed on the stripe texture to prevent a single frame pattern from introducing a global brightness bias. The first acquisition device maintains uniform exposure and gain settings within the time interval between adjacent phase exposures and automatically increases the number of phase frames or extends the single-phase dwell time when the modulation depth falls below the threshold, ensuring that the modulation error after phase demodulation meets the subsequent fitting requirements.
[0159] In one example, the point spread function is calculated as follows:
[0160] S1.4: Project the response frames of each channel in the calibration image onto the virtual image plane to obtain the response space;
[0161] S1.5: Output the calibration patterns of the three narrowband RGB channels sequentially in the response space, and perform phase demodulation on each calibration pattern to obtain multiple modulation responses;
[0162] S1.6: Divide the response space into multiple spatial partitions according to the sinusoidal fringe grid, and perform weighted fitting on the modulation response in combination with the spatial frequency corresponding to the spatial partition to obtain the contrast transfer curve corresponding to the spatial partition.
[0163] S1.7: Perform nonlinear fitting on the contrast transfer curve to obtain the principal axis length, orientation angle and energy normalization value of the point diffusion kernel corresponding to the spatial partition;
[0164] S1.8: The point spread function is obtained by weighted summation of the principal axis length, orientation angle and energy normalization value of the point spread kernel corresponding to each spatial partition through global smoothing and spatial cubic interpolation;
[0165] In this embodiment, the point spread function is calculated using the virtual image plane as the sole reference: First, the three-channel response frames are unified to the virtual image plane using the aforementioned geometric mapping matrix, eliminating the inconsistency between the sampling geometry of the acquisition end and the display end, so that the subsequent frequency and direction estimations are performed in the same coordinate system; then, the narrowband red, green, and blue sinusoidal fringes are phase demodulated to obtain the modulation response at each spatial frequency and each fixed direction. The bias caused by ambient illuminance fluctuations and backlight unevenness is suppressed by DC removal and phase averaging; to characterize spatially variable degradation, the virtual image plane is divided into multiple partitions according to the fringe grid, and the modulation samples are weighted according to the signal-to-noise ratio, fringe visibility, and frequency reliability within each partition, and the contrast transfer curve of that partition is fitted to obtain the contrast transfer curve of that partition, thereby characterizing different positions and directions. The upward detail attenuation law is studied. After obtaining the curve, a parametric model of the anisotropic kernel is used to perform nonlinear fitting, directly obtaining parameters such as the principal axis length, principal axis direction, and energy normalization of the point spread kernel for each partition. This ensures that the shape and orientation of the kernel can correspond to the directional blurring and dispersion differences commonly seen under short optical paths and small aperture conditions. Furthermore, the estimates of the three channels are aligned through inter-channel consistency constraints to separate color-related degradation components. Finally, global smoothing and spatial cubic interpolation are applied to the parameters of each partition to ensure that the point spread transitions continuously at the partition boundaries and maintains physical interpretability. A continuous point spread function covering the entire virtual image plane is reconstructed according to the position for subsequent frequency domain compensation and multi-view rendering. This ensures that the compensation process conforms to local imaging characteristics while remaining stable and usable globally.
[0166] Next, we will further elaborate on the technical content of the compensation frame in this application.
[0167] It is understandable that the geometric mapping matrix and point spread function calculated in the above content are the prior parameter set for runtime rendering. That is, the pre-modeling is completed once or with low-frequency incremental updates before entering the content presentation stage. Its output serves as the direct driving force for subsequent geometric remapping and frequency domain compensation. It is not solved repeatedly frame by frame with ordinary content frames, so as to ensure that the processing chain still has low latency and stable clarity under the conditions of small desktop volume and short optical path.
[0168] In one example, obtaining the compensated frame includes:
[0169] S2.1: Remap the content to be synchronized in reverse coordinates according to the geometric mapping matrix, and filter the coordinates through Lanzos interpolation to obtain the remapping result;
[0170] Specifically, the inverse coordinate remapping takes a geometric mapping matrix as input to restore the pixel coordinates of the content source frame to the virtual image plane coordinate system, ensuring that subsequent degradation compensation corresponds one-to-one with this coordinate system. To avoid aliasing during scaling, rotation, and shearing, electro-optical conversion linearization and chroma decoupling are performed on the content source frame before remapping, with interpolation prioritized in the luminance domain. The coordinate inversion adopts a pixel-wise reverse lookup strategy to ensure that each output position can find a corresponding sub-pixel position in the source frame. For sampling points with sub-pixel displacement, a convolution approximation is performed using a kernel function with finite support to obtain a smooth transition and suppress high-frequency aliasing.
[0171] In this embodiment, the interpolation kernel adopts a Lanzos third-order or fifth-order configuration, and the kernel width is related to the magnification: when the magnification is close to one times, a narrower kernel is used to improve edge sharpness; when the magnification increases, the kernel is appropriately widened to improve anti-aliasing capability. To reduce pseudo-high frequencies before interpolation, a direction-adaptive pre-filter is first performed on the source frame. The pre-filter intensity is determined by the anisotropy of the structure tensor. The weakening degree is lower in the strong texture direction and higher in the flat area. For transparent overlay subtitles and thin line elements, a content saliency mask is added so that these high-value edges use a higher-order kernel and restrict cross-edge sampling during interpolation to avoid edge smearing.
[0172] S2.2: Perform mirror boundary expansion and block overlap processing on the remapping result to obtain multiple mapping partitions with different spatial frequencies;
[0173] Specifically, mirror boundary expansion provides sufficient convolutional support for frequency domain processing and eliminates edge-wrapping artifacts. After expansion, the area is divided into blocks with a fixed stride, and overlapping bands are set between the blocks for seamless subsequent stitching. To ensure that each partition has a distinguishable dominant frequency characteristic in the frequency domain, the partitions are not equally divided, but adaptively defined based on local spectrum and structural strength: densely textured areas form smaller partitions for refined compensation, while flat areas form larger partitions to reduce computational overhead.
[0174] In this embodiment, the spatial frequency is measured by two paths: one is the local energy ratio of Fast Fourier Transform (FFT), and the other is the multi-scale gradient energy and angular consistency. The two paths are combined to obtain the pixel's dominant frequency estimate, which is then used to form candidate partitions through watershed or superpixel aggregation. The candidate partitions are aligned with preset minimum and maximum block sizes and snapped to a fixed fast transform size. The overlap width is set according to the effective radius of the dot diffusion kernel; the larger the radius, the wider the overlap, typically ranging from 24 to 48 pixels. Each partition stores its own dominant frequency label and orientation label as the basis for subsequent selection of gain curves and angular windowing.
[0175] S2.3: For each mapping partition, perform deconvolution in the frequency domain based on the point spread function to obtain the partition compensation result;
[0176] Specifically, deconvolution is performed independently within each partition, and the direction-dependent recovery is achieved using the point spread function corresponding to that partition. To reduce leakage from the boundaries, the partition image is multiplied by a cosine or Hanning window before entering the frequency domain. The frequency domain gain curve is obtained by the point spread function transformation and is limited according to the signal-to-noise ratio, dominant frequency, and angular label of the partition. The gain gradually converges to a safe upper limit at the high frequency end. The direction-dependent relationship is grounded through angular windowing, with the principal axis of the windowing consistent with the principal axis of the point spread function. The sidelobe directional gain is suppressed to reduce crosstalk.
[0177] S2.4: Perform weighted fusion on the partition compensation results and output the compensation frame, wherein the weights are set according to the channel co-address of each mapping partition, and the channel co-address is calculated based on the edge position offset of the three color channels of the mapping partition in the remapping result;
[0178] In this embodiment, the calculation of channel co-address is based on edge position offset: sub-pixel edge detection is performed on the three color channels in the remapping result, and the edge positioning adopts a two-level method of phase consistency or parabolic fitting to obtain the edge position along the main direction in each partition; the offset between the three color channels is formed into a co-address vector after removing extrema and smoothing a small range. The smaller the co-address vector, the better the alignment of the three channels, and the higher the weight of the partition; when the co-address vector of a certain channel deviates too much from that of other channels in the partition, the weight of that channel in the fusion is reduced to avoid the enhancement of color edges or misalignments.
[0179] Next, we will further elaborate on the technical aspects of multi-view rendering in this application.
[0180] It is easy to understand that how to predict the rotation range based on the collected attitude parameters is a matter of existing technology. For example, it can be achieved through head attitude estimation, eye tracking, infrared structured light array or inertial measurement unit fusion, etc. This application will not elaborate on it here.
[0181] It is important to emphasize that the rotation range in this application refers to the reachable visual field range of the user during the learning process. This range is predicted by the acquisition device by combining head posture, eye movement trends, and short-term historical trajectories. Specifically, it is used to limit the multi-view rendering angular domain distribution of the compensation frame, that is, within the limited refresh cycle of the display, which angles need to generate corresponding subframes and which angles can use sparse coverage. In other words, the rotation range determines the upper and lower boundaries of the angular domain division in the virtual image plane, thereby affecting the width of the exit pupil expansion and the sharpness bandwidth.
[0182] Understandably, given the limited size of a desktop, the optical path of the display is restricted, resulting in a smaller natural exit pupil within a single frame. If traditional single-viewpoint compensation is still used, even slight head movements by the user will cause them to leave the sharp focus area. To address this issue, this embodiment does not distribute subframes evenly at equal angles. Instead, it performs dense rendering in the high-probability central region and sparse rendering in the low-probability edge region based on the range of rotation. Specifically, when the compensation frame enters the rendering stage, it is converted into multiple parallax versions. Different versions are augmented with pose-related angular offsets on the virtual image plane. Each offset version is time-multiplexed and output sequentially, thereby physically expanding the exit pupil.
[0183] In one example, multi-view rendering of the compensated frame based on the rotation range includes:
[0184] S3.1: During at least one refresh cycle of the display terminal, a viewing angle domain is generated based on the rotation range;
[0185] In this embodiment, the rotation range is analyzed using the head posture angle and eye movement trend obtained from the acquisition end, forming an angle interval in two dimensions: horizontal and vertical. To facilitate rendering engine processing, this interval is discretized into multiple angle sampling points. The sampling interval is determined by the display refresh rate and subframe timing capabilities; for example, at a 120Hz refresh rate, one cycle can be divided into 8 to 12 discrete viewpoints. An orientation resolution scheduling mechanism is introduced during discretization: dense viewpoints are generated at smaller intervals in the central region, and sparse viewpoints are generated at larger intervals in the boundary region, thus forming a set of viewing angle domains covering the rotation range. Each angle sampling point stores its corresponding offset vector, sampling weight, and time tag for subsequent subframe generation and sorting.
[0186] S3.2: Calculate the core viewpoint cluster and boundary viewpoint cluster of the viewpoint domain using a clustering algorithm, and optimize the core viewpoint cluster and boundary viewpoint cluster by combining a preset viewpoint coverage threshold and brightness equalization constraint to obtain the allocation duty cycle and exposure share of each viewpoint.
[0187] In this embodiment, K-means or density peak clustering methods are used to divide the sampling points within the angular domain into core viewpoint clusters and boundary viewpoint clusters. Core viewpoint clusters correspond to high-density sampling points at the center of the angular domain, while boundary viewpoint clusters correspond to sparse sampling points at the edges of the angular domain. Subsequently, the angular coverage within each cluster is checked according to a preset coverage threshold. If the coverage is insufficient, virtual sampling points are inserted within the cluster to ensure continuity; if the coverage is too dense, redundancy is reduced through cluster shrinkage. Brightness equalization constraints normalize the exposure share within each cluster to ensure consistent global brightness after multiple subframes are superimposed. In the final output, a duty cycle and exposure share are assigned to each viewpoint, with core viewpoints typically occupying a higher refresh share, while boundary viewpoints maintain a lower but non-zero exposure share.
[0188] S3.3: Perform parallax remapping on the compensation frame according to the allocated duty cycle and exposure share, adjust the display position of the compensation frame in combination with the time sequence of the rotation range, generate and output a subframe sequence arranged alternately according to the core viewpoint cluster and the boundary viewpoint cluster;
[0189] In this embodiment, the rendering engine copies the compensation frames and performs parallax remapping based on the allocated duty cycle for each viewpoint. During remapping, geometric transformations are performed based on the offset vector within the rotation range, causing different subframes to exhibit angular differences. To ensure visual continuity of boundary subframes, bilinear or Lanzos interpolation is used in both the horizontal and vertical dimensions during remapping, along with a boundary smoothing filter to reduce jumps during viewpoint switching. According to chronological order, core viewpoint clusters and boundary viewpoint clusters are alternately inserted into the refresh cycle, with core viewpoints occupying the first half of the cycle and boundary viewpoints filling the end, thus generating a complete subframe sequence. When the subframe sequence is output to the display, the display duration is adjusted according to a predetermined duty cycle, and a corresponding brightness share is set in the exposure controller, ensuring that the final multi-viewpoint synthesis forms a continuous and stable virtual image of the distance in the user's eyes.
[0190] In one example, the video synchronization acquisition device further includes a cloud platform, which is communicatively connected to the acquisition terminal, and the method further includes:
[0191] During the operation of the display terminal, the first video stream corresponding to the display terminal is acquired, and the second video stream corresponding to the user is acquired through the second acquisition device. The first video stream includes the virtual image video displayed on the display terminal.
[0192] Upload the first video stream and the second video stream to the cloud;
[0193] Analyze the first video stream and the second video stream to extract user behavior data and display content data;
[0194] The user behavior data and display content data are input into a preset large model, and a learning report is generated through the large model. The learning report includes at least one of the following: user focus, learning progress, and behavior pattern.
[0195] In one example, this application embodiment provides a video synchronization acquisition device based on optical distant images, the device comprising:
[0196] The acquisition end is used to acquire calibration images of the content to be synchronized and to acquire the user's posture parameters;
[0197] The display end is used to calculate the geometric mapping matrix and point spread function corresponding to the calibration image, perform geometric pre-distortion, frequency domain compensation and multi-view rendering on the content to be synchronized, and display the virtual image video;
[0198] The cloud-based system communicates with the acquisition and display terminals to receive and analyze the first video stream corresponding to the display terminal and the second video stream corresponding to the acquisition terminal, extract user behavior data and display content data, and generate a learning report through a preset large model.
[0199] The acquisition terminal includes:
[0200] The first acquisition device is used to acquire the calibration image of the content to be synchronized;
[0201] The second acquisition device is used to collect the user's posture parameters and behavioral feature data, and output posture information and video streams for rotation range prediction and user attention analysis.
[0202] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method for video synchronous acquisition based on optical far-field imaging, applied to a video synchronous acquisition device, and characterized in that, The video synchronous acquisition device includes an acquisition end and a display end, the acquisition end includes a first acquisition device and a second acquisition device, the first acquisition device is used to acquire a calibration image of content to be synchronized, and the second acquisition device is used to acquire a gesture parameter of a user, and the method includes the following steps: According to the acquired calibration image, a geometric mapping matrix for mapping the calibration image to the display end and a point spread function representing spatial variable imaging degradation are calculated in combination with a preset calibration sequence; Geometric pre-distortion is performed on the content to be synchronized according to the geometric mapping matrix, and frequency domain compensation is performed on the content to be synchronized after geometric pre-distortion according to the point spread function, to obtain a compensation frame; According to the acquired gesture parameter, a rotation range is predicted, multi-view rendering of the compensation frame is performed according to the rotation range, a sub-frame sequence is generated, and the sub-frame sequence is output through the display end; The calculation method of the geometric mapping matrix is as follows: A background image corresponding to the content to be synchronized is acquired according to the calibration image, an anisotropic power spectrum of the background image is calculated, and edge straight lines and rectangular vertices of a device are acquired through morphological detection and Hough transformation, wherein the device is a terminal corresponding to the content to be synchronized; A sparse geometric constraint set is constructed according to the edge straight lines and the rectangular vertices, and two vanishing lines are obtained, and an initial value geometric mapping matrix corresponding to the sparse geometric constraint set is calculated through direct linear transformation; The main direction of the anisotropic power spectrum is taken as a regular prior constraint, and the initial value geometric mapping matrix is globally and nonlinearly refined to obtain the geometric mapping matrix. 2.The video synchronous acquisition method based on optical far-field image according to claim 1, wherein, The video synchronous acquisition device further includes a cloud end, the cloud end is in communication connection with the acquisition end, and the method further includes the following steps: During the working process of the display end, a first video stream corresponding to the display end is acquired, and a second video stream corresponding to a user is acquired through the second acquisition device, wherein the first video stream includes a virtual image video displayed by the display end; The first video stream and the second video stream are uploaded to the cloud end; The first video stream and the second video stream are analyzed, and user behavior data and display content data are extracted; The user behavior data and the display content data are input into a preset large model, and a learning report is generated through the large model, wherein the learning report at least includes one of user concentration, learning progress and behavior mode. 3.The video synchronous acquisition method based on optical far-field image according to claim 1, wherein, The initial value geometric mapping matrix corresponding to the sparse geometric constraint set is calculated through direct linear transformation, including: The edge straight line length, straightness, corner confidence and local contrast of the sparse geometric constraint set are weighted and corrected according to the field angle of the first acquisition device, to obtain a corrected constraint set; According to the corrected constraint set, a consistency constraint condition of point-to-point correspondence, line-to-line correspondence and parallelism of two vanishing lines is constructed; A linear equation set is calculated according to the consistency constraint condition; The linear equation set is solved through the least square method, and an initial mapping result is obtained through singular value decomposition; The initial mapping result is corrected in scale and sign consistency according to orthogonal consistency, to obtain the initial value geometric mapping matrix.
4. The video synchronous acquisition method based on optical far-field image according to claim 1, characterized in that, The preset calibration sequence is composed according to a calibration pattern of narrow-band RGB three channels of the calibration image, the calibration pattern includes a sinusoidal fringe grid of four directions, wherein the directions include 0 degrees, 45 degrees, 90 degrees and 135 degrees, and the calibration pattern includes at least three spatial frequencies corresponding to three channels.
5. The video synchronous acquisition method based on optical far-field image according to claim 4, characterized in that, The point spread function is calculated in the following manner: Project the response frame of each channel in the calibration image to a virtual image plane to obtain a response space; Output the calibration pattern of the narrow-band RGB three channels in the response space in sequence, and perform phase demodulation on each calibration pattern to obtain a plurality of modulation degree responses; Divide the response space into a plurality of spatial partitions according to the sinusoidal fringe grid, and perform weighted fitting on the modulation degree responses in combination with the spatial frequencies corresponding to the spatial partitions to obtain a contrast transfer curve corresponding to the spatial partitions; Perform nonlinear fitting on the contrast transfer curve to obtain a point spread kernel main axis length, a direction angle and an energy normalization value corresponding to the spatial partitions; Perform weighted summation on the point spread kernel main axis length, the direction angle and the energy normalization value corresponding to each spatial partition through global smoothing and spatial cubic interpolation to obtain the point spread function.
6. The video synchronous acquisition method based on optical far-field image according to claim 1, characterized in that, The compensated frame is obtained by: Performing inverse coordinate remapping on the content to be synchronized according to the geometric mapping matrix, and performing coordinate filtering through Lanczos interpolation to obtain a remapping result; Performing mirror boundary extension and block overlap processing on the remapping result to obtain a plurality of mapping partitions with different spatial frequencies; For each mapping partition, performing deconvolution processing in the frequency domain according to the point spread function to obtain a partition compensation result; Performing weighted fusion on the partition compensation result to output the compensated frame, wherein the weights are set according to the channel co-locations of each mapping partition, and the channel co-locations are calculated according to the edge position offset of the three color channels of the mapping partition in the remapping result.
7. The video synchronous acquisition method based on optical far-field image according to claim 1, characterized in that, The time-multiplexed multi-view rendering of the compensated frame according to the rotation range includes: In at least one refresh cycle of the display end, generating a view angle angle domain according to the rotation range; Calculating the core view cluster and the boundary view cluster of the view angle angle domain through a clustering algorithm, and optimizing the core view cluster and the boundary view cluster in combination with a preset view coverage threshold and a brightness balance constraint to obtain an allocation duty ratio and an exposure share of each view; Performing parallax remapping on the compensated frame according to the allocation duty ratio and the exposure share, adjusting the display position of the compensated frame in combination with the time sequence of the rotation range, generating a subframe sequence alternately arranged according to the core view cluster and the boundary view cluster, and outputting the subframe sequence.
8. An optical far-field video synchronous acquisition device for implementing an optical far-field video synchronous acquisition method according to any one of claims 1-7, characterized in that, The device includes: The acquisition end is configured to acquire a calibration image of the content to be synchronized and a posture parameter of a user; The display end is configured to calculate a geometric mapping matrix and a point spread function corresponding to the calibration image, perform geometric pre-distortion, frequency domain compensation and multi-view rendering on the content to be synchronized, and display a virtual image video; The cloud end is in communication connection with the acquisition end and the display end, configured to receive and analyze a first video stream corresponding to the display end and a second video stream corresponding to the acquisition end, extract user behavior data and display content data, and generate a learning report through a preset large model.
9. The video synchronous acquisition device based on optical far-field image according to claim 8, characterized in that, The acquisition end includes: The first acquisition device is used for acquiring a calibration image of the content to be synchronized. The second acquisition device is used for acquiring gesture parameters and behavior characteristic data of the user and outputting gesture information and a video stream for rotation range prediction and user concentration analysis.
Citation Information
Patent Citations
Bayes' rule based multi-frame blind convolution super-resolution reconstruction method and device
CN107680040A
Eye protection learning method, system and device for far image display and storage medium
CN119025061A