Image processing method and display method applied to video see-through head-mounted display device, device and medium

By acquiring the user's gaze point and depth in a video perspective head-mounted display device, and using a mapping relation library to deblur real images and correct virtual image fusion, the problem of inconsistency between virtual and real image fusion is solved, thus improving image quality and user experience.

CN121391664BActive Publication Date: 2026-04-28YONGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YONGJIANG LAB
Filing Date
2025-12-24
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing video see-through head-mounted display devices, different types of aberrations are introduced when light passes through the camera optical system and the near-eye display optical system, resulting in a decrease in image quality. In particular, when the user's gaze point and depth change, the virtual image and the real image are not fused consistently, resulting in problems such as blurring, distortion, and artifacts.

Method used

The eye-tracking module and depth perception module acquire the user's gaze point and gaze depth. The point spread function corresponding to the gaze point and gaze depth is determined using a pre-calibrated mapping relationship library. The real image is deblurred, and the deblurred image is fused and pre-corrected with the virtual image to reduce the aberration of the display optical path and ensure the consistency of the clarity between the virtual object and the real scene.

Benefits of technology

It improves the clarity and consistency of images in video perspective head-mounted displays, reduces blur, distortion, and artifacts in fused images, and enhances the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121391664B_ABST
    Figure CN121391664B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method and display method applied to a video perspective head-mounted display device, a device and a medium. The device comprises an eye movement tracking module, a depth perception module and an external scene camera. The processing method comprises the following steps: obtaining a gaze point of a user through the eye movement tracking module, obtaining a gaze depth through the depth perception module, and shooting a real image through the external scene camera; determining a first target point spread function corresponding to the gaze point and the gaze depth according to a first mapping relationship library, wherein the first mapping relationship library comprises a first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths; deblurring the real image based on the first target point spread function to obtain a deblurred image; and fusing and pre-correcting the deblurred image and a virtual image to obtain a pre-corrected fused image. The scheme disclosed by the application solves the problem of multiple optical path aberration coupling through deblurring of the real image and display light path pre-correction, and improves the quality of the fused image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video perspective mixed reality technology, and in particular to an image processing method, display method, device and medium for use in video perspective head-mounted display devices. Background Technology

[0002] When a user observes the real world through a VST (Video See Through) head-mounted display, light needs to pass through both the camera optical system and the near-eye display optical system. These two optical systems introduce different types of aberrations and there are multiple optical path aberration coupling problems, which reduce image quality.

[0003] In conclusion, improving the image quality in VST head-mounted display devices is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide an image processing method, display method, device and medium for use in video see-through head-mounted display devices, for improving image quality in VST head-mounted display devices.

[0005] To achieve the above objectives, this application provides the following technical solution:

[0006] An image processing method for a video perspective head-mounted display device, the video perspective head-mounted display device including an eye-tracking module, a depth sensing module, and an external scene camera, the method comprising: acquiring a user's gaze point through the eye-tracking module, acquiring the user's corresponding gaze depth through the depth sensing module, and capturing a corresponding real-world image through the external scene camera; determining a first target point diffusion function corresponding to the gaze point and the gaze depth according to a pre-calibrated first mapping relationship library, wherein the first mapping relationship library includes the first point diffusion function of the external scene camera under different simulated gaze points and simulated gaze depths; deblurring the real-world image based on the first target point diffusion function to obtain a deblurred image; and fusing and pre-correcting the deblurred image and the virtual image to obtain a pre-corrected fused image.

[0007] Optionally, determining the first target point spread function corresponding to the fixation point and the fixation depth according to a pre-calibrated first mapping relation library includes: interpolating the first point spread function in the first mapping relation library according to the fixation point and the fixation depth to obtain the first target point spread function; or, predicting the first target point spread function using a pre-trained prediction model according to the fixation point and the fixation depth; wherein the prediction model is trained based on the simulated fixation point and simulated fixation depth and the corresponding first point spread function in the first mapping relation library.

[0008] Optionally, before capturing a real-world image using the external scene camera, the method further includes: determining whether the frequency of changes in the user's gaze depth exceeds a threshold;

[0009] The step of deblurring the real image based on the first target point diffusion function includes: deblurring the real image based on the first target point diffusion function in response to determining that the frequency of user gaze depth changes exceeds a threshold; and / or, pre-correcting the deblurred image and the virtual image includes: pre-correcting the deblurred image and the virtual image in response to determining that the frequency of user gaze depth changes exceeds a threshold.

[0010] Optionally, the external scene camera is a zoom camera, the video perspective head-mounted display device includes an optical engine, the optical engine is used to display the pre-corrected fused image, and the method further includes: adjusting the focal length of the external scene camera based on the gaze depth in response to determining that the frequency of changes in the user's gaze depth does not exceed a threshold; and / or adjusting the focal length of the optical engine based on the gaze depth in response to determining that the frequency of changes in the user's gaze depth does not exceed a threshold.

[0011] Optionally, the method further includes: placing the subject at different simulated gaze points and simulated gaze depths; capturing multiple calibration images by pointing the external scene camera at the subject respectively; and determining the first point spread function of the external scene camera at different simulated gaze points and simulated gaze depths based on the multiple calibration images.

[0012] Optionally, based on the multiple calibration images, determining the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths includes: based on any calibration image, determining the second point spread function under the corresponding simulated gaze point and simulated gaze depth, wherein the data volume of the second point spread function is greater than the data volume of the first point spread function; and performing data volume reduction processing on each of the second point spread functions to obtain the corresponding first point spread function.

[0013] The step of deblurring the real-world image based on the first target point diffusion function to obtain a deblurred image includes: restoring the first target point diffusion function to obtain a corresponding second target point diffusion function; and deblurring the real-world image based on the second target point diffusion function to obtain the deblurred image.

[0014] Optionally, the video see-through head-mounted display device includes an optical engine, which is used to display the pre-corrected fused image. The process of fusing and pre-correcting the deblurred image and the virtual image to obtain the pre-corrected fused image includes: fusing the deblurred image and the virtual image to obtain a fused image; pre-correcting the fused image based on the optical characteristics of the optical engine to obtain the pre-corrected fused image; or, pre-correcting the deblurred image and the virtual image separately based on the optical characteristics of the optical engine to obtain a pre-corrected deblurred image and a pre-corrected virtual image; and fusing the pre-corrected deblurred image and the pre-corrected virtual image to obtain the pre-corrected fused image.

[0015] Optionally, it further includes: obtaining the eye movement angle corresponding to the user and the fixation point; determining a third target point diffusion function corresponding to the fixation point and the eye movement angle according to a pre-calibrated second mapping relationship library, wherein the second mapping relationship library includes the third point diffusion function of the optical engine under different simulated fixation points and simulated eye movement angles, and the third point diffusion function is used to characterize the optical characteristics of the optical engine.

[0016] An image display method for a video perspective head-mounted display device, the video perspective head-mounted display device including an optical engine, the optical engine including a display screen, the method comprising: inputting a pre-corrected fused image obtained using any of the above-described image processing methods for a video perspective head-mounted display device onto the display screen for display.

[0017] A video see-through head-mounted display device includes: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the image processing method applied to the video see-through head-mounted display device as described above and / or the steps of the image display method applied to the video see-through head-mounted display device as described above.

[0018] A readable storage medium storing a computer program that, when executed by a processor, implements the steps of the image processing method applied to a video perspective head-mounted display device as described above and / or the steps of the image display method applied to a video perspective head-mounted display device as described above.

[0019] This application provides an image processing method, display method, device, and medium for a video perspective head-mounted display device. The video perspective head-mounted display device includes an eye-tracking module, a depth sensing module, and an external scene camera. The image processing method for the video perspective head-mounted display device includes: acquiring the user's gaze point through the eye-tracking module, acquiring the user's corresponding gaze depth through the depth sensing module, and capturing a corresponding real-world image through the external scene camera; determining a first target point spread function corresponding to the gaze point and gaze depth according to a pre-calibrated first mapping relation library, wherein the first mapping relation library includes the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths; deblurring the real-world image based on the first target point spread function to obtain a deblurred image; and fusing and pre-correcting the deblurred image and the virtual image to obtain a pre-corrected fused image.

[0020] The technical solution disclosed in this application determines a first target point diffusion function corresponding to the user's gaze point and gaze depth based on a pre-calibrated first mapping relationship library of external scene cameras under different simulated gaze points and simulated gaze depths. Based on the first target point diffusion function, the real-world image captured by the external scene camera is deblurred to solve the blurring problem caused by the external scene camera, improve the quality of the real-world image, and ensure consistency in clarity between the real-world image and the subsequent virtual image. The blurring occurs because existing external scene cameras are typically fixed-focus cameras, unable to adjust their focal length according to changes in gaze depth, resulting in blurred images. After deblurring the real-world image, the deblurred real-world image and the virtual image are fused and pre-corrected. Pre-correction reduces aberrations in the display optical path, lessening the blur, distortion, and artifacts in the final fused image, ensuring that the pre-corrected fused image forms a clear fused image at the user's eye after optical-mechanical display. In other words, the application solves the problem of multiple optical path aberration coupling by deblurring the real image and pre-correcting the display optical path, and ensures the consistency of the clarity between the virtual object and the real scene, thereby improving the quality of the fused image and enhancing the user's viewing experience.

[0021] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0022] Figure 1 A flowchart illustrating an image processing method applied to a video perspective head-mounted display device, provided as an embodiment of this application;

[0023] Figure 2 A flowchart illustrating a first mapping relation library calibration method provided in this application embodiment;

[0024] Figure 3 A flowchart illustrating another first mapping relation library calibration method provided in this application embodiment;

[0025] Figure 4 A flowchart is provided for another image processing method applied to a video perspective head-mounted display device, as an embodiment of this application;

[0026] Figure 5 A flowchart illustrating another image processing method applied to a video perspective head-mounted display device provided in this application embodiment;

[0027] Figure 6 A flowchart illustrating another image processing method applied to a video perspective head-mounted display device provided in this application embodiment;

[0028] Figure 7 A flowchart illustrating another image processing method applied to a video perspective head-mounted display device provided in this application embodiment;

[0029] Figure 8 An overall data roadmap for an image processing method applied to a video perspective head-mounted display device, provided in an embodiment of this application;

[0030] Figure 9 A flowchart of another image processing method applied to a video perspective head-mounted display device provided in this application embodiment;

[0031] Figure 10 A flowchart of another image processing method applied to a video perspective head-mounted display device provided in this application embodiment;

[0032] Figure 11 The diagrams provided for comparative examples 1, 2 and embodiments of this application are schematic diagrams. Detailed Implementation

[0033] With the rise of the metaverse concept, head-mounted displays are gradually becoming more widespread in the consumer market. However, the visual quality of existing VST (Virtual-Sight) head-mounted displays still faces significant bottlenecks: when users observe the real world through a VST system, light must pass through both the camera optical system and the near-eye display optical system. These two optical systems introduce different types of blurring aberrations (such as spherical aberration and field curvature), leading to blurring, distortion, and artifacts in the final displayed image. More complexly, aberration characteristics dynamically change with the following factors:

[0034] 1. Spatial position variation: Aberrations in optical systems typically exhibit field curvature characteristics, with aberrations at the edges of the field of view being significantly greater than those in the central region;

[0035] 2. Depth variation: The VST camera produces images of objects at different distances with varying degrees of blur.

[0036] 3. Eye movement changes: The movement of the user's eyeballs will change the path of light propagation in the display light path.

[0037] As can be seen from the above, the image processing pipeline of existing video see-through head-mounted display devices suffers from a serious aberration superposition effect: the aberrations of the VST camera in the real scene imaging path are coupled with the aberrations of the optomechanical system in the display optical path.

[0038] Furthermore, if the aberrations of the VST camera in the real-world imaging path are not corrected, the resulting image quality will be poor when fused with a clear virtual image due to the inconsistent degree of blur between the two.

[0039] Therefore, this application provides an image processing method, display method, device, and medium for video perspective head-mounted display devices. Based on the different optical path aberrations of the virtual and real scenes in the video perspective head-mounted display device, a virtual-real aberration correction algorithm is designed. This algorithm fully considers the changes in blur introduced by the external scene camera with depth and spatial coordinates, and simultaneously performs good optical aberration correction on both the virtual and real scenes, aligning their display effects and reducing the blur, distortion, and artifacts in the final fused image. This ensures that the pre-corrected fused image can form a clear fused image at the user's eye after being displayed via an optomechanical system. In other words, this application solves the problem of multiple optical path aberration coupling through real image deblurring and display optical path pre-correction, ensuring consistency in sharpness between virtual objects and real scenes, improving the quality of the fused image, and thus enhancing the user's viewing experience.

[0040] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0041] See Figure 1 This is a flowchart illustrating an image processing method applied to a video perspective head-mounted display device, as provided in this application embodiment. The video perspective head-mounted display device may include an eye-tracking module, a depth sensing module, and an external scene camera. The method may include:

[0042] S11: Obtain the user's gaze point through the eye-tracking module, obtain the user's corresponding gaze depth through the depth perception module, and capture the corresponding real-world image through the external scene camera.

[0043] It should be noted that the video-perspective head-mounted display device in some embodiments of this application may include an eye-tracking module, a depth sensing module, and an external scene camera. The execution entity of the image processing method applied to the video-perspective head-mounted display device provided in some embodiments of this application may be a processor or similar component within the video-perspective head-mounted display device. Therefore, the execution entity of the image processing method applied to the video-perspective head-mounted display device provided in some embodiments of this application may also be the video-perspective head-mounted display device itself.

[0044] The eye-tracking module acquires the user's gaze point. For example, the center point of the user's gaze area can be used as the gaze point, and the eye-tracking module can obtain the coordinates of this center point. Specifically, the eye-tracking module acquires real-time images of the user's eyes captured by an eye-tracking camera (included within the module). Then, through iris localization and pupil center detection, the user's gaze direction is calculated in real time. The user's gaze direction is mapped to the gaze point coordinates (x', y') in the current display coordinate system. In this process, methods such as Kalman filtering can be used to smooth the eye movement trajectory, avoid jitter, and improve the accuracy of the user's gaze point acquisition. The eye-tracking module can precisely align aberration correction with the user's visual focus, ensuring maximum perceptual quality.

[0045] The depth perception module acquires the gaze depth corresponding to the user's gaze point. Specifically, depth information can be collected using an external scene camera or an external depth sensor and input into the depth perception module. Then, the depth perception module can calculate the gaze depth z' corresponding to the current gaze point using binocular stereo matching or Time-of-Flight (ToF) technology. Furthermore, the obtained depth can be filtered and temporally smoothed to improve robustness. By introducing depth dimension information through the depth perception module, aberration correction can adaptively adjust for objects at different distances, improving the accuracy of deblurring real-world images and thus enhancing the quality of the resulting deblurred image.

[0046] An external scene camera can capture corresponding real-world images. The video perspective head-mounted display device can acquire the corresponding real-world images captured by the external scene camera, so as to adaptively restore the corresponding real-world images captured by the external scene camera based on the user's gaze point and gaze depth, thereby obtaining a clear real-world image.

[0047] S12: Determine the first target point spread function corresponding to the gaze point and gaze depth according to the pre-calibrated first mapping relationship library, wherein the first mapping relationship library may include the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths.

[0048] In some embodiments of this application, a first mapping relation library corresponding to the external scene camera can be pre-calibrated. This first mapping relation library includes the first point spread function (PSF(x,y,z)) of the external scene camera under different simulated gaze points and simulated gaze depths. Exemplarily, the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths can be obtained experimentally or through simulation to obtain the first mapping relation library. Utilizing the first mapping relation library provides benchmark data for deblurring real-world images, enabling the video perspective head-mounted display device to accurately estimate aberrations at any gaze point and gaze depth, and to deblur the real-world images captured by the external scene camera based on the corresponding first point spread function. By pre-calibrating the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths, and establishing the first mapping relation library, combined with eye tracking and depth perception to generate the corresponding first target point spread function in real time, joint modeling of the three key factors of spatial position, depth, and eye movement can be achieved. This accurately reflects the dynamic blurring characteristics of actual optical imaging, providing high-precision priors for deblurring and compensation.

[0049] Based on the above, the first target point spread function PSF(x',y',z') corresponding to the user's gaze point and gaze depth can be determined according to the pre-calibrated first mapping relation library to realize the three-dimensional aberration modeling of "space-depth-eye movement", providing an accurate blur kernel prior for real image deblurring, thereby removing the aberration of the external scene camera by deblurring the real image based on the first target point spread function.

[0050] The dynamic point spread function modeling achieved by the above method can adapt to the user's eye movement and depth in real time, thereby improving the deblurring effect of real images and improving the quality of the obtained deblurred images.

[0051] S13: Based on the first target point diffusion function, the real image is deblurred to obtain a deblurred image.

[0052] After determining the first target point diffusion function corresponding to the user's gaze point and gaze depth according to the pre-calibrated first mapping relationship library, the corresponding real-world image captured by the external scene camera can be deblurred based on the first target point diffusion function to obtain a deblurred image, thereby eliminating the blur introduced by the external scene camera and restoring the details of the real world.

[0053] Specifically, based on the first target point diffusion function, an AI (Artificial Intelligence) model can be used to deblur real-world images to obtain a clearer image of the real-world scene (i.e., a deblurred image). The first target point diffusion function serves as a conditional input, guiding the AI ​​model to adaptively restore the current gaze region. For example, convolutional neural networks or improved Transformer structures (such as NAFNet (Nonlinear Activation-Free Network) / an improved version of Restormer) can be used to deblur real-world images; alternatively, UNet-like structures (lightweight UNet, ResUNet) can be used for image-to-image aberration compensation; or, GAN (Generative Adversarial Network) can be used to deblur real-world images to enhance the clarity and perceptual quality of the deblurred image; or, a frequency-domain convolutional network can be used to compensate for low-frequency blur and high-frequency aberrations separately in the frequency space; or, PINN (Physics-informed Neural Network) can be used to introduce an optical system model as regularization. Guided by the diffusion function of the first target point, the AI ​​model is used to perform spatial adaptive restoration of real-world images captured by external scene cameras, which significantly improves the clarity of real-world images in edge scenes and at different depths of field, thus solving the problem of blurry real-world scenes.

[0054] For example, a non-uniform deconvolution can be performed on a real-world image based on a deep learning-based deblurring network (such as a UNet variant that fuses convolution and attention), with the first target point spread function serving as a conditional input to guide the AI ​​model to adaptively restore the current gaze region.

[0055] Another example is that a modified UNet network can be used to deblur real-world images based on a first target point diffusion function: the encoder part uses multi-scale convolution to extract blur features; the decoder part combines a channel attention mechanism to achieve fine restoration; and the first target point diffusion function is used as a conditional input.

[0056] Of course, real-world images can also be deblurred using non-AI methods based on the first target point diffusion function. For example, optimized algorithms for deconvolution (such as Wiener filtering and non-blind deconvolution) can be used to deblur real-world images, enabling spatial adaptive restoration of real-world images captured by external scene cameras in a non-AI manner. This significantly improves the clarity of real-world images in edge scenes and at different depths of field, thus solving the problem of blurry real-world scenes.

[0057] S14: Fuse and pre-correct the deblurred image and the virtual image to obtain the pre-corrected fused image.

[0058] Furthermore, the video perspective head-mounted display device can also render and generate virtual images (i.e., virtual scene rendering results), and perform registration, fusion, and pre-correction on the deblurred image and the virtual image to obtain a pre-corrected fused image (i.e., a virtual-real image). Specifically, for pre-correction, the input image can be pre-compensated based on the fuzzy aberration characteristics (such as coma and spherical aberration) of the display optical system (i.e., the optomechanical system in the video perspective head-mounted display device). This allows for pre-compensation of fuzzy aberrations introduced by the optomechanical system before image display, resulting in a clearer and more natural image reaching the human eye. By fusing and pre-correcting the deblurred image and the virtual image, simultaneous optimization of the sharpness and geometric consistency between the virtual and real images is ensured, and fuzzy aberrations introduced by the optomechanical system can be pre-compensated before image display, resulting in a clearer and more natural image reaching the human eye.

[0059] For the aforementioned registration, fusion, and pre-correction, the deblurred image and the virtual image can be registered and fused first to obtain a fused image. Then, the fused image is pre-corrected to obtain a pre-corrected fused image. During registration, the virtual image can be matched with its coordinate system to ensure consistency with the deblurred image. Then, the virtual and deblurred images can be fused to generate a composite image of the virtual and real scenes. Registration and fusion ensure the visual consistency between the virtual and real images, providing a foundational image for the final pre-correction.

[0060] Alternatively, the deblurred image and the virtual image can be pre-corrected separately to obtain pre-corrected deblurred images and pre-corrected virtual images. Then, the pre-corrected deblurred image and the pre-corrected virtual image can be registered and fused to obtain a pre-corrected fused image. The process of registering and fusing the pre-corrected deblurred image and the pre-corrected virtual image is the same as the process of registering and fusing the deblurred image and the virtual image described above, and will not be repeated here.

[0061] By registering and fusing the deblurred image with the virtual image in a unified coordinate system, the consistency of the virtual object and the real scene in terms of geometric position and sharpness can be ensured, avoiding the common problems of "virtual sharpness, real blurriness" or "virtual and real misalignment." Furthermore, pre-correction before the rendered result is sent to the optical system can compensate for blur-like aberrations in the display optical path (such as spherical aberration, coma, etc.), ensuring that the virtual image entering the human eye remains sharp and aligned with the real image, thereby improving the quality of the virtual and real images entering the human eye.

[0062] As can be seen from the above, some embodiments of this application form a closed-loop process through "calibration-perception (i.e., eye movement-depth)-dynamic point spread function-real image deblurring-joint rendering and pre-correction", which can perceive the user's gaze point and gaze depth in real time and dynamically adjust the compensation, significantly improving the user's viewing experience under large field of view and dynamic gaze conditions, and effectively solving the problem of inconsistent virtual and real image aberrations in video perspective head-mounted display devices.

[0063] The technical solution disclosed in this application determines a first target point diffusion function corresponding to the user's gaze point and gaze depth based on a pre-calibrated first mapping relationship library of external scene cameras under different simulated gaze points and simulated gaze depths. The first target point diffusion function is used to deblur the real-world image captured by the external scene camera, thereby solving the blurring problem caused by the external scene camera and improving the quality of the real-world image, ensuring consistency in clarity between the real-world image and the subsequent virtual image. Blurring occurs because existing external scene cameras are typically fixed-focus cameras, unable to adjust their focal length according to changes in gaze depth, resulting in blurred images. After deblurring the real-world image, the deblurred real-world image and the virtual image are fused and pre-corrected. Pre-correction reduces aberrations in the display optical path, lessening the blur, distortion, and artifacts in the final fused image, ensuring that the pre-corrected fused image forms a clear fused image at the user's eye after optical-mechanical display. In other words, the application solves the problem of multiple optical path aberration coupling by deblurring the real image and pre-correcting the display optical path, and ensures the consistency of the clarity between the virtual object and the real scene, thereby improving the quality of the fused image and enhancing the user's viewing experience.

[0064] This application provides an image processing method for a video perspective head-mounted display device, which determines a first target point diffusion function corresponding to the gaze point and gaze depth based on a pre-calibrated first mapping relationship library, and may include:

[0065] In the first mapping relation library, the first point diffusion function is interpolated based on the gaze point and gaze depth to obtain the first target point diffusion function;

[0066] Alternatively, based on the fixation point and fixation depth, a pre-trained prediction model can be used to predict the spread function of the first target point; the prediction model is trained based on the simulated fixation point and simulated fixation depth and the corresponding first point spread function in the first mapping relation library.

[0067] In some embodiments of this application, the specific implementation of determining the first target point diffusion function corresponding to the fixation point and fixation depth based on a pre-calibrated first mapping relationship library can be any one of the following two methods:

[0068] (1) The first mapping relation library stores the relationship between discrete simulated gaze points and simulated gaze depths and the corresponding first point spread functions. Based on this, it can be determined whether there are target simulated gaze points (target simulated gaze points are the same as gaze points) and target simulated gaze depths (target simulated gaze depths are the same as gaze depths) corresponding to gaze points and gaze depths in the first mapping relation library. If there are target simulated gaze points and target simulated gaze depths corresponding to gaze points and gaze depths in the first mapping relation library, the first point spread function corresponding to the target simulated gaze points and target simulated gaze depths is determined as the first target point spread function; if there are no target simulated gaze points and target simulated gaze depths corresponding to gaze points and gaze depths in the first mapping relation library, a selected simulated gaze point and a selected simulated gaze depth whose distance from the gaze point does not exceed the first preset distance and whose distance from the gaze depth does not exceed the second preset distance are selected from the first mapping relation library. An interpolation method (such as trilinear interpolation or spline interpolation) is used to interpolate the first point spread function corresponding to the selected simulated gaze point and the selected simulated gaze depth to obtain the first target point spread function corresponding to the gaze point and gaze depth, thereby reducing the amount of computation.

[0069] Alternatively, based on the relationships between discrete simulated gaze points and simulated gaze depths and their corresponding first point spread functions stored in the first mapping relation library, the relationships between continuous simulated gaze points and simulated gaze depths and their corresponding first point spread functions can be obtained through interpolation, or by fitting. It should be noted that the aforementioned processes can be performed in advance, for example, after the first mapping relation library has been pre-calibrated, so that the first target point spread function corresponding to the user's gaze point and gaze depth can be determined promptly, quickly, and efficiently during image processing. Based on this, the first target point spread function corresponding to the user's gaze point and gaze depth can be determined according to the first mapping relation library.

[0070] (2) After the first mapping relation library is pre-calibrated, the model can be trained based on the simulated gaze points, simulated gaze depths and corresponding first point spread functions in the first mapping relation library to obtain the prediction model. After obtaining the user's gaze points and gaze depths, the user's gaze points and gaze depths can be input into the pre-trained prediction model to predict the corresponding first target point spread function, thereby reducing the amount of computation.

[0071] For example, an MLP (Multi-Layer Perceptron) model can be trained based on simulated gaze points, simulated gaze depths, and corresponding first point spread functions in a first mapping relation library to obtain an MLP prediction model. Of course, other models can also be trained to obtain a prediction model for predicting the first target point spread function; this application embodiment does not limit this approach.

[0072] The above method enables reliable and accurate determination of the first target point diffusion function corresponding to the user's gaze point and gaze depth based on a pre-calibrated first mapping relation library, providing a high-precision prior for deblurring real images, thereby improving the deblurring effect of real images and enhancing the quality of the obtained deblurred and fused images.

[0073] This application provides an image processing method for a video perspective head-mounted display device. Before capturing a real-world image using an external scene camera, the method may further include:

[0074] Determine whether the frequency of changes in the user's gaze depth exceeds a threshold;

[0075] Deblurring of real-world images based on the first target point diffusion function can include:

[0076] In response to the determination that the frequency of changes in the user's gaze depth exceeds a threshold, the real image is deblurred based on the first target point diffusion function;

[0077] And / or, pre-correction of the deblurred image and the virtual image may include:

[0078] In response to the determination that the frequency of changes in the user's gaze depth exceeds a threshold, pre-correction is performed on the deblurred image and the virtual image.

[0079] In some embodiments of this application, before capturing a real-world image using an external scene camera, the frequency of changes in the user's gaze depth can be obtained to determine whether the frequency of changes in the user's gaze depth exceeds a threshold.

[0080] If the frequency of user gaze depth changes exceeds a threshold, it indicates that the user's gaze depth changes frequently. In this case, in response to the determination that the frequency of user gaze depth changes exceeds the threshold, the real image can be deblurred based on the first target point diffusion function, and / or, in response to the determination that the frequency of user gaze depth changes exceeds the threshold, the deblurred image and the virtual image can be pre-corrected. This enables fast, efficient, and accurate deblurring of the real image and / or pre-correction of the deblurred image and the virtual image using software compensation when the frequency of user gaze depth changes is relatively fast.

[0081] As an optional method to determine the frequency of changes in a user's gaze depth, the user's gaze depth change frequency can be obtained by performing differential / filtering on the gaze depth sequence within a preset time window and then counting the number of changes or the magnitude of changes.

[0082] This application provides an image processing method for a video perspective head-mounted display device. The external scene camera is a zoom camera, and the video perspective head-mounted display device may include an optical engine for displaying a pre-corrected fused image. The method may further include:

[0083] In response to the determination that the frequency of changes in the user's gaze depth does not exceed a threshold, the focal length of the external scene camera is adjusted based on the gaze depth;

[0084] And / or, in response to determining that the frequency of changes in the user's gaze depth does not exceed a threshold, adjust the focal length of the optical engine based on the gaze depth.

[0085] In some embodiments of this application, the external scene camera included in the video see-through head-mounted display device can be a variable-focus camera. Exemplarily, the external scene camera may include a liquid crystal variable mirror, a liquid crystal adaptive lens, a deformable mirror, a movable lens, etc. Additionally, the video see-through head-mounted display device may also include an optical engine for displaying a pre-corrected fused image. Exemplarily, the optical engine may include a near-eye display optics system and a display screen, and the optical engine may also be a variable-focus optical engine. Exemplarily, a deformable mirror, a liquid crystal adaptive projector, a movable lens, etc., may be incorporated into the optical engine.

[0086] The system determines whether the frequency of changes in the user's gaze depth exceeds a threshold. If the frequency does not exceed the threshold, it indicates that the user's gaze depth changes relatively slowly. In this case, in response to the user's gaze depth change frequency not exceeding the threshold, the focal length of the external scene camera can be adjusted based on the gaze depth, and / or the focal length of the optical engine can be adjusted based on the gaze depth. This improves image sharpness by adjusting the focal length of the external scene camera and / or the optical engine. Furthermore, aberration compensation by adjusting the focal length of the external scene camera and / or the optical engine reduces the computational overhead of the software algorithm.

[0087] The above can achieve hardware and software synergy: the division of labor between hardware variable focus modules such as liquid crystal variable lenses or deformable lenses and software compensation. When the frequency of changes in the user's gaze depth is low, the hardware variable focus module can be used to adjust the sharpness, and when the frequency of changes in the user's gaze depth is high, the software method can be used for compensation.

[0088] See Figure 2This is a flowchart of a first mapping relation library calibration method provided in an embodiment of this application. The image processing method applied to a video perspective head-mounted display device provided in an embodiment of this application may further include:

[0089] The subject is placed at different simulated gaze points and simulated gaze depths;

[0090] Multiple calibration images are captured by pointing an external scene camera at the subject.

[0091] Based on multiple calibration images, the first point spread function of the external scene camera is determined under different simulated gaze points and simulated gaze depths.

[0092] In some embodiments of this application, the first mapping relation library can be obtained by the following method:

[0093] (1) Place the subject to be photographed at different simulated gaze points and simulated gaze depths.

[0094] (2) Multiple calibration images are collected by taking pictures of the subject at different simulated gaze points and simulated gaze depths using an external scene camera.

[0095] (3) Based on each acquired calibration image, determine the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths to obtain the first mapping relationship library.

[0096] That is, the first point spread function of the external scene camera can be recorded experimentally under different simulated gaze points and simulated gaze depths. These first point spread functions can be organized into a three-dimensional matrix form PSF(x,y,z), forming a first mapping relationship library indexed by image plane coordinates (x,y) and depth z. This provides benchmark data for dynamic aberration modeling and interpolation, enabling video perspective head-mounted display devices to accurately estimate and deblur aberrations at any gaze point and gaze depth.

[0097] Of course, the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths can also be obtained through optical simulation methods.

[0098] For example, this can be achieved using point light source targets or Zemax optical simulation methods. Under different simulated gaze points and simulated gaze depths, the first point spread function (PSF) of the external scene image captured by the camera is recorded experimentally or through simulation. These PSFs are then organized into a three-dimensional matrix form PSF(x,y,z), forming a first mapping relationship library indexed by image plane coordinates (x,y) and depth z. See details... Figure 3The flowchart illustrates another first mapping relation library calibration method provided in this application embodiment: A calibration sampling grid is set with Δx=0.1m, Δy=0.1m, and Δz=0.25m; the target is controlled to move sequentially on the grid; an external scene camera sequentially captures target images at various positions; clustering and compression are performed on the captured second point diffusion function, and the first point diffusion function obtained by compressing the second point diffusion function is stored. The storage amount can be reduced by downsampling, principal component decomposition, or clustering compression.

[0099] Another example is that point light source arrays can be arranged at multiple locations with different simulated gaze depths (the specific values ​​of the simulated gaze depth can be set according to needs or experience, such as 0.5m, 1m, 2m, 3m, etc.) to acquire images of the external scene camera within the entire field of view, and obtain the first point spread function under different (x,y,z) (i.e., different simulated gaze points and simulated gaze depths).

[0100] This application provides an image processing method for a video perspective head-mounted display device, which, based on multiple calibration images, determines the first point spread function of an external scene camera under different simulated gaze points and simulated gaze depths, and may include:

[0101] Based on any calibration image, determine the second point spread function at the corresponding simulated gaze point and simulated gaze depth. The amount of data for the second point spread function is greater than the amount of data for the first point spread function.

[0102] For any second-point diffusion function, the data volume is reduced to obtain the corresponding first-point diffusion function;

[0103] Based on the first target point diffusion function, the real image is deblurred to obtain a deblurred image, which may include:

[0104] The diffusion function of the first target point is restored to obtain the corresponding diffusion function of the second target point;

[0105] Based on the second target point diffusion function, the real image is deblurred to obtain a deblurred image.

[0106] In some embodiments of this application, the method for determining the first point spread function of an external scene camera under different simulated gaze points and simulated gaze depths based on multiple calibration images can be as follows: For each calibration image, a second point spread function (i.e., the original point spread function under the corresponding simulated gaze point and simulated gaze depth) is determined for that calibration image, wherein the data volume of the second point spread function is greater than that of the first point spread function; for each second point spread function, data volume reduction processing is performed to obtain the first point spread function corresponding to each second point spread function, thereby reducing the data volume of the first mapping relation library and reducing the storage requirements of the video perspective head-mounted display device. The aforementioned data volume reduction processing can be downsampling, principal component decomposition, or clustering, etc., and this application embodiment does not limit this to these methods.

[0107] Based on the above, the process of deblurring the real image based on the first target point diffusion function to obtain the deblurred image can be as follows: First, the first target point diffusion function corresponding to the user's gaze point and gaze depth is restored (i.e., the reverse processing of reducing data volume) to obtain the second target point diffusion function corresponding to the first target point diffusion function; then, the real image is deblurred based on the second target point diffusion function to obtain the deblurred image, thereby improving the reliability and accuracy of real image deblurring.

[0108] The above methods can reduce the data storage requirements of video see-through head-mounted displays to lower costs, while also using precise point spread functions to deblur real-world images, thereby improving the clarity and quality of the deblurred images.

[0109] See Figure 4 and Figure 5 ,in, Figure 4 A flowchart is provided for another image processing method applied to a video perspective head-mounted display device, as an embodiment of this application. Figure 5 This is a flowchart illustrating another image processing method applied to a video perspective head-mounted display device, as provided in this application embodiment. The video perspective head-mounted display device may include an optical engine, which is used to display a pre-corrected fused image. The method involves fusing and pre-correcting a deblurred image and a virtual image to obtain the pre-corrected fused image, and may include:

[0110] The deblurred image and the virtual image are merged to obtain the merged image;

[0111] Based on the optical characteristics of the optomechanics, the fused image is pre-corrected to obtain the pre-corrected fused image;

[0112] or,

[0113] Based on the optical characteristics of the optomechanic, the deblurred image and the virtual image are pre-corrected respectively to obtain the pre-corrected deblurred image and the pre-corrected virtual image;

[0114] The pre-corrected deblurred image and the pre-corrected virtual image are fused to obtain the pre-corrected fused image.

[0115] In some embodiments of this application, the video see-through head-mounted display device may further include an optical engine that can be used to display a pre-corrected fused image. Exemplarily, the optical engine may include a near-eye display optics system and a display screen.

[0116] In obtaining a pre-corrected fused image by fusing and pre-correcting the deblurred image and the virtual image, the deblurred image and the virtual image can be fused first, and then pre-corrected uniformly based on the optical characteristics of the optomechanical system. Alternatively, a branch-type compensation method can be used: the deblurred image and the virtual image can be pre-corrected separately based on the optical characteristics of the optomechanical system, and then fused. These two paths are parallel implementation methods, and the device can choose one based on its computing power. For example, when computing power is limited, the deblurred image and the virtual image can be fused first, and then pre-corrected uniformly based on the optical characteristics of the optomechanical system; when computing power is sufficient, a branch-type compensation method can be used: the deblurred image and the virtual image can be pre-corrected separately based on the optical characteristics of the optomechanical system, and then fused.

[0117] Specifically, for the first method described above, the deblurred image and the virtual image can be fused (specifically, registration can be performed first, followed by fusion) to obtain a fused image. Then, based on the optical characteristics of the optomechanical system, the fused image is pre-corrected to obtain a pre-corrected fused image.

[0118] For the second method described above, the deblurred image and the virtual image can be pre-corrected based on the optical characteristics of the optomechanical system to obtain the pre-corrected deblurred image and the pre-corrected virtual image respectively. Then, the pre-corrected deblurred image and the pre-corrected virtual image are fused (specifically, registration can be performed first, followed by fusion) to obtain the pre-corrected fused image.

[0119] The above method enables flexible pre-correction of deblurred images and virtual images, and can achieve global pre-correction. It can also ensure that the virtual image and the real image are optimized in terms of sharpness and geometric consistency at the same time, so as to offset the residual aberrations of the real optical path and ensure that the virtual image that finally enters the human eye is clear and aligned with the real image.

[0120] See Figure 6 and Figure 7 ,in, Figure 6A flowchart illustrating another image processing method applied to a video perspective head-mounted display device provided in this application embodiment. Figure 7 A flowchart illustrating another image processing method applied to a video perspective head-mounted display device provided in this application embodiment. The image processing method applied to a video perspective head-mounted display device provided in this application embodiment may further include:

[0121] Obtain the eye movement angle corresponding to the user's gaze point;

[0122] Based on a pre-calibrated second mapping relation library, the third target point diffusion function corresponding to the fixation point and eye movement angle is determined. The second mapping relation library may include the third point diffusion function of the optical engine under different simulated fixation points and simulated eye movement angles. The third point diffusion function is used to characterize the optical features of the optical engine.

[0123] In some embodiments of this application, the eye movement angle corresponding to the user's gaze point can also be obtained. For example, the eye movement angle corresponding to the user's gaze point can be obtained simultaneously with the eye tracking module. The eye movement angle can be expressed as ( ',θ'), where, θ' is the angle of deflection of the human eye relative to the optical axis; specifically, ' represents the azimuth angle, and θ' represents the elevation angle / polar angle.

[0124] In addition, in some embodiments of this application, a second mapping relation library corresponding to the optical engine can be pre-calibrated. This second mapping relation library may include the third-point spread function of the optical engine under different simulated gaze points and simulated eye movement angles. For example, the third-point spread function of the optical engine under different gaze points and eye movement angles can be obtained experimentally or through simulation to obtain the second mapping relation library. The second mapping relation library provides benchmark data for pre-correction of deblurred images and virtual images, enabling video perspective head-mounted display devices to accurately estimate and pre-correct near-eye display optical engine aberrations under arbitrary gaze points and eye movement angles. It should be noted that the specific implementation of pre-calibrating the second mapping relation library corresponding to the optical engine is similar to that of pre-calibrating the first mapping relation library corresponding to the external scene camera. For example, it can be implemented using a point light source target or Zemax optical simulation method. Under different simulated gaze points and simulated eye movement angles, the third-point spread function of the optical engine is recorded experimentally or through simulation, and these third-point spread functions are organized into a four-dimensional matrix form PSF(x,y, ,θ), forming an image plane coordinate system (x,y) and an eye movement angle (θ). ,θ) is the second mapping relation library indexed by.

[0125] Based on the above, the third target point spread function PSF(x',y',) corresponding to the user's gaze point and eye movement angle can be determined according to the pre-calibrated second mapping relation library. The aforementioned steps can specifically be as follows: In the second mapping relation library, interpolate the third point spread function based on the fixation point and eye movement angle to obtain the third target point spread function; or, based on the fixation point and eye movement angle, use a pre-trained third target point spread function prediction model to predict the third target point spread function, where the prediction model is trained based on simulated fixation points and simulated eye movement angles in the second mapping relation library. It should be noted that the specific implementation of the aforementioned process is similar to the implementation of determining the first target point spread function corresponding to the fixation point and fixation depth based on a pre-calibrated first mapping relation library, and will not be elaborated further here.

[0126] The point spread function (PSF) describes the intensity distribution of light from an ideal point light source on the image plane after passing through an imaging system. Essentially, it reflects the blurring characteristics and resolution performance of the imaging system when imaging a point object. In other words, the PSF reveals how the imaging system will blur a previously sharp point into a blurry spot, and how closely it can distinguish two points. Therefore, the third target PSF can be used to characterize the optical features of the optomechanical system. Specifically, pre-correcting the fused image based on the optical characteristics of the optomechanical system can be understood as pre-correcting the fused image based on the third target PSF. Pre-correcting the deblurred image and the virtual image separately based on the optical characteristics of the optomechanical system can be understood as pre-correcting the deblurred image and the virtual image separately based on the third target PSF, thereby improving the pre-correction effect and enhancing the quality of the fused image.

[0127] To reduce the data storage size of the video perspective head-mounted display device, multiple fourth-point spread functions (i.e., the original point spread functions under the corresponding simulated gaze points and simulated eye movement angles) can be determined based on different simulated gaze points and simulated eye movement angles. The data size of the fourth-point spread function is larger than that of the third-point spread function. Data reduction processing is performed on each fourth-point spread function to obtain the corresponding third-point spread function. Based on this, pre-correction based on the third target point spread function can be specifically performed as follows: the third target point spread function is restored to obtain the corresponding fourth target point spread function, and pre-correction is performed based on the fourth target point spread function. It should be noted that the method for reducing the data size of the fourth-point spread function is the same as the method for reducing the data size of the second-point spread function described above, and will not be repeated here.

[0128] Specifically, when performing pre-correction based on the third target point spread function, an AI optical pre-correction network can be used (with the third target point spread function serving as a conditional input to guide the AI ​​optical pre-correction network) to perform blur compensation corresponding to the blur characteristics of the optomechanical system. This allows for pre-compensation of blur-like aberrations introduced by the optomechanical system before image display, resulting in a clearer and more natural image reaching the human eye. For example, the AI ​​optical pre-correction network could be a ResNet-based blur compensation network.

[0129] Specifically, the AI ​​optical pre-calibration network can be a convolutional neural network or an improved Transformer structure, or a UNet-like structure (lightweight UNet, ResUNet), or a GAN (Generative Adversarial Network), or a frequency-domain convolutional network, PINN, etc. Alternatively, pre-calibration can be performed using non-AI methods based on a third target point spread function; for example, optimized deconvolution algorithms (such as Wiener filtering, non-blind deconvolution, etc.) can be used for pre-calibration.

[0130] Of course, existing AI aberration blind compensation algorithms can also be used directly for pre-correction (without needing to calibrate the optical-mechanical point spread function, i.e., without needing to rely on a second mapping relation library).

[0131] The following is combined with Figures 8 to 10 The above embodiments will be further described, wherein, Figure 8 This application provides an overall data roadmap for an image processing method applied to a video perspective head-mounted display device. Figure 9 The flowchart of another image processing method applied to a video perspective head-mounted display device provided in this application embodiment is as follows: Figure 10 This application provides another image processing method for a video perspective head-mounted display device. The camera calibration module is responsible for calibrating the first point spread function of the camera in the external scene under different simulated gaze points and simulated gaze depths, and constructing a first mapping relationship library. The eye tracking module and the depth perception module respectively acquire the user's gaze point and eye movement angle (exemplarily, (x', y', ... The processing module takes a gaze point (z', θ') and a gaze depth (z') as input. The processing module determines the first target point diffusion function corresponding to the gaze point and gaze depth based on a first mapping relation library. This first target point diffusion function is then input to the deblurring module. The deblurring module deblurs the real-world image corresponding to the gaze point and gaze depth captured by the external scene camera based on the first target point diffusion function, resulting in a deblurred image. This deblurring module adaptively restores the real-world image, obtaining a clearer image of the real-world scene. Simultaneously, the rendering module renders a virtual image and registers and fuses it with the deblurred image to obtain a fused image. This fused image is input to the optical path pre-correction module. Specifically, the rendering module performs lighting and geometric matching on the virtual image to ensure it aligns with the coordinates of the real-world scene. Then, the two images are fused to generate a composite image of the virtual and real-world scenes. The rendering module ensures visual consistency between the virtual and real images, providing a foundational image for the final pre-correction. The optical path pre-calibration module utilizes a deep learning-based optical pre-calibration network for joint pre-calibration (specifically, pre-compensating the input image based on the fuzzy aberration characteristics of the display optical system) to generate a pre-calibrated fused image that is consistent between virtual and real. This allows the optical path pre-calibration module to compensate for fuzzy aberrations introduced by the optical mechanism before image display, resulting in a clearer and more natural image reaching the human eye. Within the optical path pre-calibration module, there are two options: one is to calibrate the third-point spread function of the optical mechanism under different simulated gaze points and simulated eye movement angles using the optical mechanism calibration module, constructing a second mapping relation library. The optical mechanism PSF processing module then obtains the third target point spread function corresponding to the user's gaze point and eye movement angle based on the second mapping relation library, and uses the third target point spread function to pre-compensate for display optical path aberrations. The other option is to directly use the existing AI aberration blind compensation algorithm without calibrating the optical mechanism point spread function. The final fused image is projected to the human eye through the display optical system (i.e., the final fused image is input to the display module through the optical mechanism and perceived by the human eye), allowing the user to perceive a clear fused virtual and real image. The input to the display module is the final fused image, which is projected onto the human retina through a near-eye optical system, allowing the user to perceive a clear MR image that is consistent with the virtual image, thus completing a closed loop from data processing to final human perception. This method effectively solves the problem of inconsistent virtual and real image aberrations in video perspective head-mounted display devices through a closed-loop process of "calibration-perception-dynamic point spread function modeling-deblurring-joint rendering-optical path pre-correction".

[0132] This application also provides an image display method for a video perspective head-mounted display device. The video perspective head-mounted display device may include an optical engine, which may include a display screen. The method may include:

[0133] The pre-corrected fused image obtained using any of the above image processing methods applied to video perspective head-mounted display devices is input onto the display screen for display.

[0134] In some embodiments of this application, a pre-corrected fused image can be obtained using any of the above-mentioned image processing methods applied to video perspective head-mounted display devices, and the pre-corrected fused image can be input to the display screen included in the optical engine for display, so that the displayed pre-corrected fused image can be projected onto the human retina through the near-eye optical system in the optical engine, so that the user can perceive a clear MR image that is consistent with the virtual reality.

[0135] The following description, in conjunction with comparative examples, illustrates some embodiments of this application:

[0136] Embodiment of this application: AI-based dynamic aberration pre-correction video perspective head-mounted display device

[0137] 1. Experimental equipment

[0138] One VST-type mixed reality head-mounted display device with a resolution of 1920×1080 and a field of view of approximately 90°; the VST camera has a resolution of 1920×1080.

[0139] Binocular eye-tracking camera with a sampling rate of 120Hz;

[0140] ToF depth camera with a depth resolution of 640×480;

[0141] High-spot light source array as calibration target;

[0142] The RTX 4090 GPU serves as the computing platform.

[0143] 2. System Calibration

[0144] Point light source arrays were arranged at four different depths (0.5m, 1m, 2m, 3m);

[0145] Camera images are acquired throughout the entire field of view to obtain PSF (Point Spread Function) at different (x,y,z) values.

[0146] Principal component analysis was used to compress the first mapping relation library and store it as a lookup table.

[0147] 3. Operation Process

[0148] After the user wears the device, the VST camera captures real-time images, the eye-tracking module outputs the corresponding gaze point (x', y') in real time, and the depth perception module outputs the corresponding gaze depth z'.

[0149] The processing module retrieves and interpolates the first target point diffusion function PSF(x',y',z') from the first mapping relation library and passes it to the deblurring module;

[0150] The deblurring module uses an improved UNet network: the encoder part uses multi-scale convolution to extract blurred features; the decoder part combines channel attention mechanism to achieve fine restoration; the first target point spread function PSF(x',y',z') is used as conditional input.

[0151] The output clear, deblurred image is blended with the virtual image in the rendering module;

[0152] The optical path pre-correction module uses a ResNet-based blur compensation network to pre-compensate the fused image;

[0153] The final image is projected onto the human eye, and the user experiences a clear, aligned, and seamlessly integrated virtual-real image.

[0154] Comparative Example 1:

[0155] 1. Experimental equipment

[0156] One VST type mixed reality head-mounted display device with a resolution of 1920×1080 and a field of view of approximately 90°. The VST head-mounted display device is also equipped with an external camera with a resolution of 3840×3840.

[0157] The RTX 4090 GPU serves as the computing platform.

[0158] 2. Operation Process

[0159] After the user wears the device, the high-resolution VST camera captures clear real-world images in real time. These clear real-world images are not processed in any way, preserving the original input.

[0160] The virtual road image and the unprocessed real image are first fused, and then the fused image is pre-corrected; the fuzzy aberration distribution (including spherical aberration, coma, etc.) of the display optical path is measured using optical modeling methods; a mapping function is constructed to pre-compensate the fuzzy aberration of the MR synthesized image before display.

[0161] Comparative Example 2: Pre-correction of MR synthetic images

[0162] 1. The experimental equipment and system calibration are the same as those in the embodiments of this application.

[0163] 2. Operation process

[0164] After the user wears the device, the VST camera captures real-time images without processing them, preserving the original input.

[0165] The virtual road image and the unprocessed real image are first fused, and then the fused image is pre-corrected; the fuzzy aberration distribution (including spherical aberration, coma, etc.) of the display optical path is measured using optical modeling methods; a mapping function is constructed to pre-compensate the fuzzy aberration of the MR synthesized image before display.

[0166] 3. Experimental Results

[0167] like Figure 11 The diagram shown is a schematic representation of Comparative Example 1, Comparative Example 2, and the embodiments provided in this application. Wherein, Figure 11 In Comparative Example 1, (a) is the ideal fused image obtained by fusing a high-definition real image (captured by a high-definition camera with higher resolution) with a virtual image and then pre-correcting it. Figure 11 (b) in Comparative Example 2 is a blurred fused image obtained by fusing a blurred real image and a virtual image and then pre-correcting it. Because the virtual image is clearer and has better geometric correction, but the depth-related spatial variation PSF of the real image is not processed, an inconsistency of "virtual image is clear and real image is blurry" appears when the image is fused and displayed, which is quite different from the ideal fused image in Comparative Example 1. Figure 11 (c) in this application is a clear fused image obtained by fusing the deblurred real image and the virtual image and then pre-correcting them. Because dynamic blur aberration pre-correction is performed on both the virtual image and the real image at the same time, combined with PSF(x,y,z), eye movement and depth input, a clear display that is consistent with the virtual and real images is achieved, which is close to the ideal fused image.

[0168] This application ensures that virtual and real images are optimized in terms of sharpness and geometric consistency, and can adapt to eye movements and depth in real time through dynamic PSF modeling, with the effect being superior to static compensation.

[0169] This application embodiment also provides a video perspective head-mounted display device, which may include:

[0170] Memory, used to store computer programs;

[0171] The processor, when executing a computer program stored in memory, can implement the steps of any of the above-described image processing methods applied to a video perspective head-mounted display device and / or the steps of the above-described image display methods applied to a video perspective head-mounted display device.

[0172] This application also provides a readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps of any of the above-described image processing methods applied to a video perspective head-mounted display device and / or the steps of the above-described image display methods applied to a video perspective head-mounted display device.

[0173] The description of the image display method, video see-through head-mounted display device and readable storage medium provided in this application embodiment can be found in the detailed description of the corresponding part of the image processing method for video see-through head-mounted display device provided in this application embodiment, and will not be repeated here.

[0174] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0175] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0176] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0177] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0178] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "joining," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise expressly limited. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0179] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. An image processing method applied to a video perspective head-mounted display device, characterized in that, The video-guided head-mounted display device includes an eye-tracking module, a depth sensing module, and an external scene camera; the method includes: The eye-tracking module obtains the user's gaze point, the depth perception module obtains the user's corresponding gaze depth, and the external scene camera captures the corresponding real-world image. According to a pre-calibrated first mapping relationship library, a first target point spread function corresponding to the gaze point and the gaze depth is determined, wherein the first mapping relationship library includes the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths; Based on the first target point diffusion function, the real image is deblurred to obtain a deblurred image; The deblurred image and the virtual image are fused and pre-corrected to obtain a pre-corrected fused image; Before capturing a real-world image using the external scene camera, the method further includes: Determine whether the frequency of changes in the user's gaze depth exceeds a threshold; The process of deblurring the real-world image based on the first target point diffusion function includes: In response to determining that the frequency of changes in the user's gaze depth exceeds a threshold, the real-world image is deblurred based on the first target point diffusion function; And / or, pre-correcting the deblurred image and the virtual image, including: In response to the determination that the frequency of changes in the user's gaze depth exceeds a threshold, the deblurred image and the virtual image are pre-corrected.

2. The image processing method for a video perspective head-mounted display device according to claim 1, characterized in that, Based on a pre-defined first mapping database, the first target point diffusion function corresponding to the fixation point and the fixation depth is determined, including: In the first mapping relationship library, the first point diffusion function is interpolated according to the gaze point and the gaze depth to obtain the first target point diffusion function; Alternatively, based on the fixation point and the fixation depth, a pre-trained prediction model is used to predict the spread function of the first target point; the prediction model is trained based on the simulated fixation point and simulated fixation depth and the corresponding first point spread function in the first mapping relation library.

3. The image processing method for a video perspective head-mounted display device according to claim 1, characterized in that, The external scene camera is a zoom camera, the video perspective head-mounted display device includes an optical engine, the optical engine is used to display the pre-corrected fused image, and the method further includes: In response to the determination that the frequency of changes in the user's gaze depth does not exceed a threshold, the focal length of the external scene camera is adjusted based on the gaze depth; And / or, in response to determining that the frequency of changes in the user's gaze depth does not exceed a threshold, the focal length of the optical engine is adjusted based on the gaze depth.

4. The image processing method for a video perspective head-mounted display device according to claim 1, characterized in that, The method further includes: The subject is placed at different simulated gaze points and simulated gaze depths; Multiple calibration images are captured by pointing the external scene camera at the subject respectively; Based on the multiple calibration images, the first point spread function of the external scene camera under different simulated gaze points and simulated gaze depths is determined.

5. The image processing method for a video perspective head-mounted display device according to claim 4, characterized in that, Based on the multiple calibration images, the first point spread function of the external scene camera is determined under different simulated gaze points and simulated gaze depths, including: Based on any calibration image, determine the second point spread function at the corresponding simulated gaze point and simulated gaze depth, wherein the data volume of the second point spread function is greater than the data volume of the first point spread function; For each of the second point diffusion functions, the data volume is reduced to obtain the corresponding first point diffusion function; The step of deblurring the real-world image based on the first target point diffusion function to obtain a deblurred image includes: The diffusion function of the first target point is restored to obtain the corresponding diffusion function of the second target point. Based on the second target point diffusion function, the real image is deblurred to obtain the deblurred image.

6. The image processing method applied to a video perspective head-mounted display device according to any one of claims 1 to 5, characterized in that, The video-guided head-mounted display device includes an optical engine, which is used to display the pre-corrected fused image. The process of fusing and pre-correcting the deblurred image and the virtual image to obtain the pre-corrected fused image includes: The deblurred image and the virtual image are fused to obtain a fused image; Based on the optical characteristics of the optical engine, the fused image is pre-corrected to obtain a pre-corrected fused image; or, Based on the optical characteristics of the optical engine, the deblurred image and the virtual image are pre-corrected respectively to obtain the pre-corrected deblurred image and the pre-corrected virtual image; The pre-corrected deblurred image and the pre-corrected virtual image are fused to obtain a pre-corrected fused image.

7. The image processing method for a video perspective head-mounted display device according to claim 6, characterized in that, Also includes: Obtain the eye movement angle corresponding to the user's gaze point; According to a pre-calibrated second mapping relationship library, a third target point diffusion function corresponding to the fixation point and the eye movement angle is determined. The second mapping relationship library includes the third point diffusion function of the optical engine under different simulated fixation points and simulated eye movement angles. The third point diffusion function is used to characterize the optical properties of the optical engine.

8. An image display method applied to a video see-through head-mounted display device, characterized in that, The video-guided head-mounted display device includes an optical engine, the optical engine includes a display screen, and the method includes: The pre-corrected fused image obtained by the image processing method for video perspective head-mounted display devices according to any one of claims 1 to 7 is input to the display screen for display.

9. A video see-through head-mounted display device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the image processing method for a video perspective head-mounted display device as described in any one of claims 1 to 7, or the steps of the image display method for a video perspective head-mounted display device as described in claim 8.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the image processing method for a video perspective head-mounted display device as described in any one of claims 1 to 7, or the steps of the image display method for a video perspective head-mounted display device as described in claim 8.

Citation Information

Patent Citations

  • Image filtering method based on eye movement tracking, near-eye display method and system, head-mounted display, medium and product

    CN119599905A

  • Extended depth-of-field correction using reconstructed depth map

    US20240169494A1