Methods and system for characterizing and calibrating a wearable helmet-mounted digital visual augmentation device with an ar / VR display

The method and system use stereo imaging and machine learning to calibrate helmet-mounted AR/VR displays, addressing weight distribution and spatial awareness issues by generating a digital twin, thus optimizing headborne weight and enhancing user experience.

WO2025171485A1PCT designated stage Publication Date: 2025-08-21NAT RES COUNCIL OF CANADA +1

Patent Information

Application Number
PCT/CA2025/050189
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-13
Filing Date
2025-02-13
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing helmet-mounted digital visual augmentation devices suffer from imbalanced headborne weight distribution, leading to neck strain and reduced spatial awareness due to offset viewpoints, which are not effectively addressed by current calibration methods that fail to account for complex display shapes and distortions.

Method used

A method and system using stereo imaging, structured-light patterns, and machine learning to calibrate helmet-mounted AR/VR displays, incorporating a neural network to correct deviations and generate a digital twin, ensuring accurate alignment and preserving 3D spatial awareness.

Benefits of technology

The solution optimizes headborne weight distribution, enhances spatial awareness, and reduces cybersickness by accurately calibrating the display system, supporting a wide range of AR/VR devices with complex shapes and distortions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025050189_21082025_PF_FP_ABST
    Figure CA2025050189_21082025_PF_FP_ABST
Patent Text Reader

Abstract

A method for characterizing and calibrating a helmet-mounted digital visual augmentation device for a user, wherein the method comprises: with cameras mounted on the helmet, acquiring stereo images of a calibration target; calibrating a position of the user's eyes; with a projection system, projecting structured-light patterns associated with the target on an AR / VR display; acquiring a 3D model of the AR / VR display and obtaining a shape of the AR / VR display; mapping between AR / VR display pixel coordinate and the position in 3D on the calibration target in front of the user's eyes where the pixel is perceived by the user; encoding deviations of the projection system based on the mapping, wherein the 3D model of the AR / VR display is in a same reference frame as the cameras mounted on the helmet; using at least one machine learning model to correct deviations between an image of the target as seen from a point of view of the user's eyes and an image of the target captured by the cameras; and generating a digital twin of the digital visual augmentation device.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEM FOR CHARACTERIZING AND CALIBRATING A WEARABLE HELMET-MOUNTED DIGITAL VISUAL AUGMENTATION DEVICE WITH AN AR / VR DISPLAYFIELD

[0001] The present disclosure relates to methods and systems for calibrating wearable helmet-mounted digital visual augmentation devices with AR / VR displays. BACKGROUND

[0002] Visual augmentation devices, such as night vision goggles (NVG), display visors, and augmented reality systems are designed to enhance performance through improvements in situational awareness, day / night vision and target acquisition. However, there are concerns about the potentially adverse effects associated with these devices on a helmet due to the increased headbome weight of the system which forces users to modify their neck posture such that greater deviations from a neutral upright position was adopted, which can lead to fatigue and injury over prolonged use.

[0003] Current TRL-9 NVGs are comprised of Image Intensification (12) tubes mounted in front of the helmet, as best practice dictates that goggle apertures should be coaligned with the optical axes of the eyes (inline configuration). The result is unbalanced headbome weight. Helmet-mounted monoculars (ex: AN / PVS-14), bioculars (ex: AN / PVS-7) and binoculars (ex: AN / PVS-31) produce a forward shift of the centre of mass (CM). Extensive studies have been conducted on neck muscle strain relating to helmet-mounted NVGs for helicopter pilots [1] and dismounted soldiers [2].

[0004] An emerging category of future night vision goggles is expected to be fully digital (DNVG), with 12 tube technology being replaced with imaging sensors, electronics and microdisplays. Despite advancements (currently TRL-6 / 7), any future digital goggle based on an inline configuration is likely to suffer from similar weight and balance issues.

[0005] According to present knowledge, there are two alternatives to improve the CM of the headbome weight. One such proposal uses ballasts for inline NVGs / DNVGs, which entails adding weight on the back of the helmet to balance the load, such as a counterweight or battery pack. Another proposal entails rearranging the goggle components in a distributed apertures configuration [3], The Steiner OpticsAN / PVS-21 Low Profile Night Vision Goggle is an example of a distributed aperture binocular NVG. It is claimed that: "by using a see-through beam-combiner display, protrusion is reduced by up to 10cm. The improved lever arm enhances capabilities in confined spaces and permits aggressive use in rough terrain. "[4], Evidently, the AN / PVS-21 has not been widely adopted by the operational community and end-user appreciation is mixed, with complaints including bumping into door frames and issues with reaching for door knobs. This suggests some level of alteration of 3D spatial awareness that negatively impacts mobility. Additionally, further improving the centre of mass of a system such as the AN / PVS-21 is complex and impractical, due to the necessity of optically guiding the light from the scene to the 12 tubes, and then to the eyes.

[0006] Digital visual augmentation devices with distributed apertures can be designed to improve the CM of the headbome weight. However, as with the AN / PVS- 21, moving the cameras away from the optical axes of the eyes introduces offset viewpoints with six degrees of freedom (horizontal, vertical and longitudinal translations, and roll, pitch and yaw rotations). Vision systems with offset viewpoints are known to alter depth perception and spatial awareness as the offset increases. These alterations degrade soldier mobility and targeting capabilities.

[0007] For future digital night vision systems, system architectures are inherently modular (sensing, processing, projection) and lend themselves more naturally to a distributed configuration. Digital vision systems only require that the projection module (display) be centered with the optical axes of the eye. Examples of distributed digital systems can be found in aviation for helicopter pilots (e.g. Thales TopOwl).

[0008] One survey of calibration methods for AR helmets classifies calibration methods into three categories: manual, semi-automatic and automatic [5], All of the manual calibration methods cannot be applied in the context of night vision systems as the burden of calibration is too high. The automatic and semi-automatic method separates the calibration model into, two components, the display model and the eye model. Once the display model is known, the eye model is composed of 3 parameters per eye for a total of 6 parameters. When the interpupilary distance of a subject is known, this reduced to 5 parameters. In some situations, the 5 parameters can be fixed without inducing too much parallax error, and there exist methods to recover the 6parameters using the eye-tracking feature of some helmets to calibrate the position of the eyes [6, 7, 8, 17], Semi-automatic methods typically use a camera to replace the human eye and assume a planar display without distortion [9, 10, 11, 12, 14], Other methods take into account some amount distortion [10, 14], while others uses a light field to model to estimate a 4D to 4D mapping [15, 16],

[0009] One of the drawbacks of the existing methods is that these methods do not generalize to multiple types of helmets. Furthermore, there is no 3D reconstruction of the display and thus no-complex shape, and the different methods assume a flat display which does not apply to a large field of view displays; and it is also challenging to model complex lenses distortion induced by the display in a device independent way. SUMMARY

[0010] In one of its aspects, a method for characterizing and calibrating a helmetmounted digital visual augmentation device for a user, wherein the method comprises: with cameras mounted on the helmet, acquiring stereo images of a calibration target; calibrating a position of the user’s eyes; with a projection system, projecting structured-light patterns associated with the target on an AR / VR display; acquiring a 3D model of the AR / VR display and obtaining a shape of the AR / VR display; mapping between AR / VR display pixel coordinate and the position in 3D on the calibration target in front of the user’s eyes where the pixel is perceived by the user; encoding deviations of the projection system based on the mapping, wherein the 3D model of the AR / VR display is in a same reference frame as the cameras mounted on the helmet; using at least one machine learning model to correct deviations between an image of the target as seen from a point of view of the user’s eyes and an image of the target captured by the cameras; and generating a digital twin of the digital visual augmentation device

[0011] In another of its aspects, an image rendering system for a visual augmentation device associated with a head-mounted display (HMD) on a helmet, the system comprising: a set of cameras mounted on the helmet configured to acquire stereo images of a calibration target; a projection system for projecting structured-light patterns associated with the target on an AR / VR display; acquiring a 3D model of the AR / VR display and obtaining a shape of the AR / VR display; a mapping between AR / VR display pixel coordinate and a position in 3D on the calibration target in front of the user’s eyes where the pixel is perceived by the user; encoding deviations of the projection system based on the mapping, wherein the 3D model of the AR / VR display is in a same reference frame as the set of cameras mounted on the helmet; a neural network module comprising at least one machine learning model trained to correct deviations between an image of the target as seen from a point of view of the user’s eyes and an image of the target captured by the cameras.

[0012] In another of its aspects, a method for calibrating a visual augmentation device with a virtual reality (VR) or augmented reality (AR) system comprising operations performed by an apparatus comprising a processor coupled to at least one memory device storing instructions executable by the processor to at least perform the operations of:(i) with a first set of image capture devices, capturing a first image of a target comprising a series of known patterns;(ii) displaying the captured first image on a display of the wearable device, wherein the displayed first image comprises a plurality of first pixels of the display;(iii) with a second set of image capture devices within the wearable device, capturing the displayed image of the target to form a second image with a plurality of second pixels;(iv) using the series of known patterns, perform pixel by pixel mapping of the second image from second set of image capture devices with the displayed first image on the display to generate a pixel-by-pixel map;(iv) with the pixel-by-pixel map, obtain a 3-D mapping of the location of the plurality of second pixels and orientation of the plurality of second pixels in relation to the plurality of first pixels; and(v) based on the 3-D mapping, correcting deviations between the second image and the first image, thereby calibrating the visual augmentation device.

[0013] In another of its aspects, a system for calibrating a helmet-mounted visual augmentation device with an AR / VR display, the system comprising: an apparatus comprising one or more processors coupled to at least one memory device storing instructions executable by the processor; a neural network module comprising at least one machine learning model comprising a set of instructions executable by the processor to at least: receive a pair of stereo images of a target captured by a first imaging apparatus and a second imaging apparatus; extract a plurality of first feature points from the image of the first imaging apparatus and extract a plurality of second feature points from the image of the second imaging apparatus; match the plurality of first feature points from the image of the first imaging apparatus to the plurality of second feature points from the image of the second imaging apparatus; generate a disparity map based on a plurality of matching results; and calculate parameters of the first imaging apparatus and the second imaging apparatus based on the disparity map.

[0014] In another of its aspects, a computer system for calibrating a helmetmounted visual augmentation device with an AR / VR display, the computer system comprising: one or more computer processors; at least a first set of image capture devices configured to capture a first image of a target;a display for displaying the captured first image, wherein the displayed first image comprises a plurality of first pixels of the display; at least a second set of image capture devices configured to capture the displayed image to form a second image with a plurality of second pixels; a computer-readable storage medium comprising a first set of program instructions stored thereon, the first set of program instructions executable by the one or more processors to perform pixel by pixel mapping of the second image with the displayed first image on the display to generate a pixel-by-pixel map; the computer-readable storage medium comprising a second set of program instructions stored thereon, the second set of program instructions executable by the one or more processors to obtain a 3D mapping of the location of the plurality of second pixels and orientation of the plurality of second pixels in relation to the plurality of first pixels; and based on the 3D mapping, the computer-readable storage medium comprising a third set of program instructions stored thereon, the third set of program instructions executable by the one or more processors to align the second image with the first image such that there is substantial alignment between the second image and the first image, thereby calibrating the visual augmentation device.

[0015] Advantageously, the methods and systems described herein work with a large set of AR / VR displays and support curve display and distortion, as they comprise a correction methodology that accounts for offset viewpoints and preserves 3D spatial awareness while avoiding undesired effects such as cybersickness; allow digital twining of the helmet; allow onsite maintenance of the helmet; are less dependent to the user's visual system; further constrain the eye model; have increased flexibility in the number of cameras and modality used; and optimize the CM of the head bome weight.BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 shows an example system for characterizing and calibrating an AR / VR device for digital night vision using structured light with various implementations described herein;

[0017] Figure 2 shows a mannequin wearing a helmet with two exterior cameras and two interior cameras;

[0018] Figure 3a shows a calibration bench with the mannequin with the helmet of Figure 2;

[0019] Figure 3b shows the mannequin with the helmet and a simulated AR / VR display with structured-light patterns;

[0020] Figure 3 c shows a target;

[0021] Figure 3d shows a laser tracker used as a reference instrument;

[0022] Figure 3e shows an example screenshot of a calibration software application;

[0023] Figure 3f shows multi-spectral calibration targets used for the calibration;

[0024] Figure 3g shows a target being measured by the laser tracker using a spherically mounted retroreflector (SMR);

[0025] Figure 4 shows example steps for a method for characterizing and calibrating a digital visual augmentation device with AR / VR display using structured light with various implementations described herein;

[0026] Figure 5 shows an example workflow for a calibration stage;

[0027] Figures 6a-f show examples of structured-light patterns generated by the simulated AR / VR display;

[0028] Figure 7a shows an example workflow for the compensation stage comprising an inference phase;

[0029] Figure 7b shows an example workflow for the compensation stage comprising the training phase;

[0030] Figure 8 shows a result of the calibration corresponding to the simulated display shown in Figure 3b;

[0031] Figures 9 and 10 show further example screenshots of the calibration software application;

[0032] Figures 11 and 12 show graphical user interfaces displaying the calibrated cameras of the helmet; and

[0033] Figure 13 illustrates a block diagram of an example machine or apparatus upon which any one or more of the techniques (e.g., methodologies) discussed herein may be performed.DETAILED DESCRIPTION

[0034] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the following description to refer to the same or similar elements. While embodiments of the disclosure may be described, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the elements illustrated in the drawings, and the methods described herein may be modified by substituting, reordering, or adding stages to the disclosed methods. Accordingly, the following detailed description does not limit the disclosure. Instead, the proper scope of the disclosure is defined by the appended claims.

[0035] Moreover, it should be appreciated that the particular implementations shown and described herein are illustrative of the invention and are not intended to otherwise limit the scope of the invention in any way. Indeed, for the sake of brevity, certain sub-components of the individual operating components, and other functional aspects of the systems may not be described in detail herein. Furthermore, the connecting lines shown in the various figures contained herein are intended to represent exemplary functional relationships and / or physical couplings between the various elements. It should be noted that many alternative or additional functional relationships or physical connections may be present in a practical system.

[0036] In Figure 1, there is shown an example system for characterizing and calibrating a digital visual augmentation device 12 with an AR / VR display using various implementations described herein, designated by reference numeral 10. The system 10 comprises a digital visual augmentation device 12 associated with a helmet 13 with two head-mounted cameras 14a, 14b on the exterior of the helmet 13, and two cameras 16a, 16b within the helmet 13, and an AR / VR display 18 within the helmet 13; one or more targets 20 in the field of view of the two head-mounted cameras 14a, 14b; and an apparatus 22 upon which any one or more of the techniques (e.g., methodologies) discussed herein may be performed. Figure 2 shows a mannequin 24 with a helmet 13 having two exterior cameras 14a, 14b, two interiors cameras 16a, 16b (display 18 not shown). In another example, more than two exterior cameras 14a, 14b or more than two interior cameras 16a, 16b may be employed. The AR / VR display18 allows for an optical see-through display i.e. augmented reality (AR) and a video see-through display i.e. virtual reality (VR).

[0037] The apparatus 14 may be communicatively coupled to a server machine 26 via a communication network 28, and one or more databases 30 may be communicatively coupled to the server machine 26.

[0038] Figure 3a shows a calibration bench 26 with the mannequin 24 with the helmet 13 having two exterior cameras 14a, 14b, two interiors cameras 16a, 16b, while Figure 3b shows the mannequin 24 with the helmet 13 and a simulated AR / VR display 18 with structured-light patterns. Figure 3c shows one or more targets 20. The output of the cameras 14a, 14, 16a, 16b and the display 18 are coupled to the apparatus 22 which comprises calibration software application calibrating digital visual augmentation device 12. An example screenshot of the calibration software application is shown in Figure 3e.

[0039] Figure 4 shows a flow diagram 100 depicting a computer-implemented method for characterizing and calibrating a digital visual augmentation device 12 using structured light with various implementations described herein. The method comprises at least a calibration stage and a compensation stage. It should be understood that while the operational flow diagram indicates a particular order of execution of the operations, in other implementations, the operations might be executed in a different order. Further, in some implementations, additional operations or blocks may be added to the method. Likewise, some operations or blocks may be omitted.

[0040] In one example, the AR / VR display 18 is used as a structured-light proj ector and a 3D scan or model of the display 18 is acquired. Such an arrangement allows for the acquisition of the shape of the display 18 and a mapping between the AR pixel coordinate or feature point and the position in 3D in front of the eyes where the pixel is perceived by the user. In addition, such a mapping encodes all the deviations of the projection system. Moreover, the 3D model of the display 18 is in the same reference frame as the cameras 14a, 14b mounted on the helmet 13. Accordingly, the two interior cameras 16a, 16b do not replace the user’s eyes, and are merely used to perform the 3D re-construction.

[0041] In step 102, at least three types of data are prepared for the calibration stage,. The first type of input data is composed of a set of at least 5 quadruplets and preferably of the form (id,X,Y,Z) where (X,Y,Z) is a 3D coordinate that corresponds to physical target id (Targets are shown in Figure 3f and 3g). In Figure 3d, a laser tracker 34 is used as a reference instrument, and Figure 3f shows multi-spectral calibration targets 20 used for the calibration, in which the multi-spectral calibration targets 20 are arranged in a general configuration, viewable by, or in the field of vision of, a camera 14a, 14b mounted on the helmet 13, such that all the 3D coordinates contained in the set quadruplets are not on a single plane in 3D space. Figure 3g shows one target 20 being measured by the laser tracker 34. Specifically, the laser tracker34 measures the position of the centre of a spherically mounted retroreflector (SMR)35 which is shown within the dashed circle. In one example, using structured light, a plane of light, created by passing a laser beam through a cylindrical lens, is swept across the multi-spectral calibration targets 20 using a galvanometer-driven mirror. At each position of the plane, a light stripe is created, which is sensed by the two- dimensional cameras 16a, 16b. The intersection of the known plane and the line of sight from the cameras 16a, 16b determines the three-dimensional coordinates. Accordingly, a second type of data is composed of images of the target 20 that are captured by cameras 16a, 16b within the helmet 13. The images in the set are captured while the cameras 16a, 16b are static (not moving) with respect to the targets 20.

[0042] The third type of data is composed of a set of quintuplets, in which the first three elements of a quintuplet are composed of a projected image and two images from the two cameras 16a, 16b mounted within the helmet 13 and facing the display 18. Those cameras 16a, 16b and associated images are referred to as the structured- light ones and are pointed at one of the two displays 18 contained in the augmented or virtual reality helmet, as shown in Figures 2 and 3a. Moreover, the quintuplet also contains two matrices that represent the position and orientation of the reference camera with respect to the structured-light cameras. Figure 5 shows an example workflow for the calibration stage in more detail.

[0043] Figures 6a-f show examples of structured-light patterns generated by the simulated AR / VR display 18 comprising horizontal and vertical fringes. Figures 6a, 6c, 6e show images from the left mannequin camera 16b and Figures 6b, 6d, 6f showimages from the right mannequin camera 16a, such that a stereo image pair consists of left and right images, which are captured by the cameras 16a, 16b from both views at the same time.

[0044] In step 104, the data is input into calibrations models comprising families of algorithms associated with the calibration stage. One family of algorithms comprises numerical optimization routines that determine the best model parameters that fit a set of observations. In some cases, the models are linear and can be solved using linear algebra. In other cases, the models are non-linear and require the use of iterative algorithms such as Levenberg-Marquardt method and alternatives to compute intrinsic and extrinsic camera parameters. Another family of algorithms comprises a structured-light encoding and decoding routine. Yet another family of algorithms comprises image processing algorithms that extracts the position of an imaged physical target in a camera’s image. Those skilled in the art can with ease select an algorithm suitable to the type of target used.

[0045] In step 106, an output of the calibration stage is generated and comprises two sets of data that are used at the compensation stage. The first set of data is composed of matrices and vectors of parameters that represent the intrinsic and extrinsic parameters of all the cameras, and the intrinsic and extrinsic parameters related to the structured-light and reference camera can be discarded as they are not used at the compensation stage. The second set of data is a quintuplet of parameters composed of the 3D position of each pixel of the display 18 of an augmented or virtual reality helmet and its associated pixel coordinates. The formatting of this set of parameters is as follows: (i,j,X,Y,Z) where (I,j) are the pixel coordinates and (X,Y,Z) are the 3D coordinates.

[0046] Generally, the algorithms used at the calibration station are typically executed on a workstation, such as the machine 22, however, the algorithms may also be executed on a remote machine 26, or by a remote machine 26 in a cloud computing environment.

[0047] Next, the output of the calibration stage is input into the compensation stage.In step 108, an inference process is initiated as part of the compensation calibration stage. The images of the inference camera pair and the associated intrinsic andextrinsic parameters are used by a low-power stereo matching subsystems to generate a sparse disparity map, as a way to relate the same coordinates in a 2D space and a 3D space. The map points for a given location may obtained using various methods, for example, some approaches may generate a large number of (dense) points, lower resolution depth points and other approaches may generate a much smaller number of (sparse) points. The low-power stereo matching subsystems may comprise an Intel® Xeon® Processor E5 V4, from Intel Corporation, U.S.A. The parameters of each display and the associated 3D position of the eye can be set manually or using an eye tracker. The inference phase may be performed on low-power edge computing platform that could be used for a wearable application. Figure 7a shows an example workflow for the compensation stage comprising an inference phase in more detail.

[0048] In step 110, a training process is initiated as part of the compensation stage. As part of the training phase, a stereo matching process which entails finding pixels corresponding to the same 3D point in a scene is undertaken. A low-power stereo matching subsystem, such as the Intel V4, is used to generate a power efficient disparity map. A high-quality disparity map of the same scene may be obtained by adding an active light source to the system.

[0049] In certain embodiments, a processor 201 (as shown in Figure 13) of machine 22, 26 operates one or more machine learning models. Machine learning models may be operated using any combination of hardware and / or software (e.g., program instructions) located in processor 201 and / or on machine 22, 26. In some embodiments, one or more neural network modules are used to operate the machine learning models on machine 22, 26. In one example, different machine learning classifiers or algorithms are used for building machine learning models or predictive models, such as, supervised learning algorithms, unsupervised learning algorithms and reinforcement learning algorithms. Examples of supervised learning algorithm systems include support vector machine, decision tree, linear regression, logistic regression, naive Bayes, k-nearest neighbor, random forest, AdaBoost, XGBoost, and neural network methods. Examples of unsupervised learning algorithm systems include K-means, mean shift, affinity propagation, hierarchical clustering, DBSCAN (density-based spatial clustering of applications with noise), Gaussian mixture modeling, Markov random fields, ISODATA (iterative self-organizing data), andfuzzy C-means systems. Examples of reinforcement learning algorithm systems include Maja and Teaching-Box systems. Generally, training the predictive models involves optimizing the parameters of a predictive system to minimize the loss function. In addition to the training step, the predictive models also undergoes validation using test datasets.

[0050] The neural network module may include neural network circuitry installed or configured with operating parameters that have been learned by the neural network module or a similar neural network module (e.g., a neural network module operating on a different processor or device). For example, a neural network module may be trained using training images (e.g., reference images) and / or other training data to generate operating parameters for the neural network circuitry. The operating parameters generated from the training may then be provided to neural network module installed on apparatus 22, 26.

[0051] In one example, the machine learning models are trained using a pair of high-quality and sparse disparity maps. In one example, the neural network module is a depth convolutional neural network. The neural network module may be based on an architecture that is usually associated with the problem of image super resolution. The fact that the neural network generates a disparity map allows a more compact network than if images were generated. Furthermore, generating the disparity map allows the system to significantly reduce the problem of hallucination associated with generative models. The output of the neural network requires further processing.

[0052] In step 112, an accuracy measure for the candidate classifiers trained by the training process is determined, and the model is evaluated through feedback provided to the feature process, or to the training process, or to both. Further, the feedback preferably continues as an iterative process until a predetermined stopping criterion is met, that is, until the CNN best model is found (step 114). A subset of features or attributes for making the best predictions are selected. In one example, each input feature is multiplied by some value, or weight, such that during the training step, the weights are updated until the best model is found.

[0053] In step 116, a post-processing process is initiated. The post-processing step comprises sub-steps for replacing all disparities in the output map that have a validvalue in the input one. This is done to ensure that the values at the input of the network are preserved.

[0054] Second, if the neural network is applied to a pair of cameras with a different baseline (distance between cameras) than the training one, the disparity values are scaled by the factor defined by the training baseline and the one used at inference.

[0055] Third, with the knowledge of the parameters of the displays and the 3D position of the eyes, it is possible to reproject the images of the inference cameras into the display using the disparity map via a number of steps as outlined below.

[0056] The first step is the compensation of the parallax. The disparity map obtained from the energy efficient stereo matching subsystem is reprojected into the left and right inference images. For the pixels in the left and right inference image there is an associated disparity and intensity (i.e., color). Using this information, it is possible to reproject the images into a virtual camera having the same orientation but a different position from the inference ones. Furthermore, each virtual camera is located at the position of one of the eyes. The second step is the compensation of the variation of orientation between the displays (virtual cameras) with respect to the inference camera. This is achieved by applying a projective transform (homography). The third step applies the final correction using the display parameters and positions of the user eyes to indicate correction for the point where the disparity should be zero. This corresponds to a different rotation applied to each image, such a rotation is computed based on the distance of the point of interest that is observed by the user. This last step permits preservation of the unity magnification and requires the application of a projective transform to the image (i.e., homography). Generally, the processing is designed to be executed on a GPU platform such as a NVIDIA Jetson low-power system, from Nvidia Corporation, U.S.A. Figure 7b shows an example workflow for the compensation stage comprising the training phase in more detail.

[0057] In step 118, calibration is completed. The result of the calibration is shown in Figure 8 corresponding to the simulated display 18 shown in Figure 3b. As can be seen in Figure 8, at the top left portion there is depicted three cameras 14a’, 14b’ and 14c’ from the helmet 13 (one camera 14c’ is occluded by the other two cameras 14a’, 14b’). At the bottom left portion of Figure 8, the 3D model of each pixel (position andnormal) of the display 18’ for each eye. Note that the helmet’s cameras and display’s pixels are in the same global reference frame. On the right portion of Figure 8, there is shown mannequin’s two cameras 16a’, 16b’ which are used to calibrate the visual augmentation device 12.

[0058] Figures 9 and 10 show graphical user interfaces displaying screenshots of a calibration software application.

[0059] Figures 11 and 12 show graphical user interfaces displaying the calibrated cameras of the helmet 13.

[0060] In more detail, Figure 13 illustrates a block diagram of an example machine or apparatus 200, such as machine 22 or 26, upon which any one or more of the techniques (e.g., methodologies) discussed herein may be performed. Apparatus 200 (e.g., computer system) may include a hardware processor 201 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 202 and a static memory 206, connected via an interlink 203 (e.g., link or bus), as some or all of these components may constitute hardware for systems or related implementations discussed above.

[0061] Generally, the hardware processor 201 may, for example, include at least one of a Central Processing Unit (CPU), a Reduced Instruction Set Computing (RISC) Processor, a Complex Instruction Set Computing (CISC) Processor, a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), a Tensor Processing Unit (TPU), a Neural Processing Unit (NPU), aVision Processing Unit (VPU), a Machine Learning Accelerator, an Artificial Intelligence Accelerator, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), a Radio- Frequency Integrated Circuit (RFIC), a Neuromorphic Processor, a Quantum Processor, or any combination thereof. A processor circuit may further be a multicore processor having two or more independent processors (sometimes referred to as "cores") that may execute instructions contemporaneously. Multi-core processors contain multiple computational cores on a single integrated circuit die, each of which can independently execute program instructions in parallel. Parallel processing on multi-core processors may be implemented via architectures like superscalar, VLIW, vector processing, or SIMD that allow each core to run separate instruction streams concurrently. A processor circuit may be emulated in software, running on a physicalprocessor, as a virtual processor or virtual circuit. The virtual processor may behave like an independent processor but is implemented in software rather than hardware.

[0062] Specific examples of main memory 202 include Random Access Memory (RAM), and semiconductor memory devices, which may include storage locations in semiconductors such as registers. Specific examples of static memory 206 include non-volatile memory, such as semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; RAM; or optical media such as CD-ROM and DVD-ROM disks.

[0063] The apparatus 200 may further include a display device 210, an input device 212 (e.g., a keyboard), and a user interface (UI) navigation device 214 (e.g., a mouse). In an example, the display device 210, input device 212, and UI navigation device 214 may be a touch-screen display. The apparatus 200 may include a mass storage device 216 (e.g., drive unit), a signal generation device 218 (e.g., a speaker), a network interface device 220. The apparatus 200 may include an output controller 228, such as a serial (e.g., universal serial bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection to communicate or control one or more peripheral devices (e.g., a printer, card reader, cameras, etc.). The apparatus 200 may include sensors 229 coupled thereto.

[0064] The mass storage device 216 may comprise a machine-readable medium 222 on which is stored one or more sets of data structures or instructions 204 (e.g., software) embodying or utilized by any one or more of the techniques or functions described herein. The instructions 204 may also reside, completely or at least partially, within the main memory 202, within static memory 206, or within the hardware processor 201 during execution thereof by the apparatus 200. In an example, one or any combination of the hardware processor 201, the main memory 202, the static memory 206, or the mass storage device 216 comprises a machine readable medium.

[0065] Specific examples of machine-readable media include, one or more of nonvolatile memory, such as semiconductor memory devices (e.g., EPROM or EEPROM) and flash memory devices; magnetic disks, such as internal hard disks andremovable disks; magneto-optical disks; RAM; or optical media such as CD-ROM and DVD-ROM disks. While the machine-readable medium is illustrated as a single medium, the term "machine readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) configured to store the one or more instructions 204.

[0066] The term “machine readable medium” includes, for example, any medium that is capable of storing, encoding, or carrying instructions for execution by the apparatus 200 and that cause the apparatus 200 to perform any one or more of the techniques of the present disclosure or causes another apparatus or system to perform any one or more of the techniques, or that is capable of storing, encoding or carrying data structures used by or associated with such instructions. Non-limiting machine- readable medium examples include solid-state memories, optical media, or magnetic media. Specific examples of machine-readable media include: non-volatile memory, such as semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; Random Access Memory (RAM); or optical media such as CD-ROM and DVD-ROM disks. In some examples, machine readable media includes non-transitory machine-readable media. In some examples, machine readable media includes machine readable media that is not a transitory propagating signal.

[0067] The instructions 204 may be transmitted or received, for example, over a communications network 205 using a transmission medium via the network interface device 220 utilizing any one of a number of transfer protocols (e.g., frame relay, internet protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), etc.). Example communication networks include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), mobile telephone networks (e.g., cellular networks), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards known as WiFi®), IEEE 802.15.4 family of standards, a Long Term Evolution (LTE) 4G or 5G family of standards, a Universal Mobile Telecommunications System (UMTS) familyof standards, peer-to-peer (P2P) networks, satellite communication networks, among others.

[0068] In an example, the network interface device 220 includes one or more physical jacks (e.g., Ethernet, coaxial, or other interconnection) or one or more antennas to access the communications network 205. In an example, the network interface device 220 includes one or more antennas to wirelessly communicate using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) techniques. In some examples, the network interface device 220 wirelessly communicates using Multiple User MIMO techniques. The term “transmission medium” shall be taken to include any intangible medium that is capable of storing, encoding or carrying instructions for execution by the apparatus 200, and includes digital or analog communications signals or other intangible medium to facilitate communication of such software.

[0069] Examples, as described herein, can include, or can operate on, logic or a number of components, modules, or mechanisms (all referred to hereinafter as “modules”). Modules are tangible entities (e.g., hardware) capable of performing specified operations and is configured or arranged in a certain manner. In an example, circuits are arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner as a module. In an example, the whole or part of one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware processors are configured by firmware or software (e.g., instructions, an application portion, or an application) as a module that operates to perform specified operations. In an example, the software can reside on a non- transitory computer readable storage medium or other machine-readable medium. In an example, the software, when executed by the underlying hardware of the module, causes the hardware to perform the specified operations.

[0070] Accordingly, the term “module” is understood to encompass a tangible entity, be that an entity that is physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., transitorily) configured (e.g., programmed) to operate in a specified manner or to perform part or all of any operation described herein. Considering examples in which modules are temporarily configured, each of the modules need not be instantiated at any one moment in time. For example, wherethe modules comprise a general-purpose hardware processor configured using software, the general-purpose hardware processor is configured as respective different modules at different times. Software can accordingly configure a hardware processor, for example, to constitute a particular module at one instance of time and to constitute a different module at a different instance of time.

[0071] Implementations of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non- transitory computer-storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer-storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0072] A computer program, which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site ordistributed across multiple sites and interconnected by a communication network. While portions of the programs illustrated in the various figures are shown as individual modules that implement the various features and functionality through various objects, methods, or other processes, the programs may instead include a number of sub-modules, third-party services, components, libraries, and such, as appropriate. Conversely, the features and functionality of various components can be combined into single components, as appropriate.

[0073] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., a CPU, a GPU, an FPGA, or an ASIC.

[0074] The term “graphical user interface,” or “GUI,” may be used in the singular or the plural to describe one or more graphical user interfaces and each of the displays of a particular graphical user interface. Therefore, a GUI may represent any graphical user interface, including but not limited to, a web browser, a touch screen, or a command line interface (CLI) that processes information and efficiently presents the information results to the user. In general, a GUI may include a plurality of user interface (UI) elements, some or all associated with a web browser, such as interactive fields, pull-down lists, and buttons operable by the user. These and other UI elements may be related to or represent the functions of the web browser.

[0075] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of wireline and / or wireless digital data communication, e.g., a communications network 205.

[0076] Various Notes

[0077] Each of the non-limiting aspects in this document can stand on its own or can be combined in various permutations or combinations with one or more of the other aspects or other subject matter described in this document.

[0078] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, specific embodiments in which the invention can be practiced. These embodiments are also referred to generally as “examples.” Such examples can include elements in addition to those shown or described. However, the present inventors also contemplate examples in which only those elements shown or described are provided. Moreover, the present inventors also contemplate examples using any combination or permutation of those elements shown or described (or one or more aspects thereof), either with respect to a particular example (or one or more aspects thereof), or with respect to other examples (or one or more aspects thereof) shown or described herein.

[0079] In this document, the terms “a” or “an” are used, as is common in patent documents, to include one or more than one, independent of any other instances or usages of “at least one” or “one or more.” In this document, the term “or” is used to refer to a nonexclusive or, such that “A or B” includes “A but not B,” “B but not A,” and “A and B,” unless otherwise indicated. In this document, the terms “including” and “in which” are used as the plain-English equivalents of the respective terms “comprising” and “wherein.” Also, in the following claims, the terms “including” and “comprising” are open-ended, that is, a system, device, article, composition, formulation, or process that includes elements in addition to those listed after such a term in a claim are still deemed to fall within the scope of that claim. Moreover, in the following claims, the terms “first,” “second,” and “third,” etc., are used merely as labels, and are not intended to impose numerical requirements on their objects.

[0080] Method examples described herein can be machine or computer- implemented at least in part. Some examples can include a computer-readable medium or machine-readable medium encoded with instructions operable to configure an electronic device to perform methods as described in the above examples. An implementation of such methods can include code, such as microcode, assembly language code, a higher-level language code, or the like. Such code can include computer readable instructions for performing various methods. The codemay form portions of computer program products. Such instructions can be read and executed by one or more processors to enable performance of operations comprising a method, for example. The instructions are in any suitable form, such as but not limited to source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like.

[0081] Further, in an example, the code can be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, such as during execution or at other times. Examples of these tangible computer-readable media can include, but are not limited to, hard disks, removable magnetic disks, removable optical disks (e.g., compact disks and digital video disks), magnetic cassettes, memory cards or sticks, random access memories (RAMs), read only memories (ROMs), and the like.

[0082] The above description is intended to be illustrative, and not restrictive. For example, the above-described examples (or one or more aspects thereof) may be used in combination with each other. Other embodiments can be used, such as by one of ordinary skill in the art upon reviewing the above description. The Abstract is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Also, in the above Detailed Description, various features may be grouped together to streamline the disclosure. This should not be interpreted as intending that an unclaimed disclosed feature is essential to any claim. Rather, inventive subject matter may lie in less than all features of a particular disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description as examples or embodiments, with each claim standing on its own as a separate embodiment, and it is contemplated that such embodiments can be combined with each other in various combinations or permutations. The scope of the invention should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

[0083] REFERENCES1. Thuresson, M., Jan L., and Karin H.-R., "Neck muscle activity in helicopter pilots: effect of position and helmet-mounted equipment." Aviation, space, and environmental medicine 74, no. 5, pp.527-532, 2003.2. Manoogian, S. J., Kennedy, E.A., Duma, S.M., "A literature review of musculoskeletal injuries to the human neck and the effects of head-supported mass worn by soldier", USAARL Contract Report No. CR-2006-01, 2005.3. Bull, G.C., "Helmet display options: a route map." In Helmet-Mounted Displays II, vol. 1290, pp. 81-92. International Society for Optics and Photonics, 1990.4. Steiner Optics, https: / / www.steiner-optics.com / imaging-systems / anpvs-21-low- profde-nvg, product page accessed March 19, 2021.5. Grubert et al. A Survey of Calibration Methods for Optical See-Through Head- Mounted Displays 2017 (https: / / arxiv.org / pdf / 1709.04299.pdf).6. Cutolo, F., Cattari, N., Fontana, U., & Ferrari, V. (2020). Optical See-Through Head-Mounted Displays With Short Focal Distance: Conditions for Mitigating Parallax-Related Registration Error. Frontiers in Robotics and Al, 7, 196. https: / / pdfs.semanticscholar.org / e7dc / 269c368aa2f72afb6812ed3f04f9017ed0ab.pdf7. Plopski, A., Itoh, Y., Nitschke, C., Kiyokawa, K., Klinker, G., & Takemura, H. (2015). Comeal -imaging calibration for optical see-through head-mounted displays. IEEE transactions on visualization and computer graphics, 21(f), 481-4908. Jones, J. A., Edewaard, D., Tyrrell, R. A., & Hodges, L. F. (2016, March). A schematic eye for virtual environments. In 2016 IEEE Symposium on 3D User Interfaces (3DUI) (pp. 221-230). IEEE9. Ballestin, Giorgio & Chessa, Manuela & Solari, Fabio. “Assessment of Optical See-Through Head Mounted Display Calibration for Interactive Augmented Reality”. Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV) Workshops, 201910. X. Hu, F. R. Y. Baena and F. Cutolo, "Alignment-Free Offline Calibration of Commercial Optical See-Through Head-Mounted Displays with Simplified Procedures," in IEEE Access, vol. 8, pp. 223661223674, 2020, doi:10.1109 / ACCESS.2020.3044184. https: / / ieeexplore.ieee.org / abstract / document / 929137711. Gilson, S. J., Fitzgibbon, A. W., & Glennerster, A. (2008). Spatial calibration of an optical see-through head-mounted display. Journal of neuroscience methods, 173(\), 140-14612. Makibuchi, Naoya & Kato, Haruhisa & Sugano, Masaru. (2016). Field of view extension for augmented reality on binocular optical see-through head-mounted displays. 1-2. 10.1145 / 3005274.3005289.13. Owen, C. B., Zhou, J., Tang, A., & Xiao, F. (2004, November). Display -relative calibration for optical see-through head-mounted displays. In Third IEEE and ACM international symposium on mixed and augmented reality (pp. 70-78). IEEE.14. Lee, S., & Hua, H. (2014). A robust camera-based method for optical distortion calibration of head-mounted displays. Journal of Display Technology, 77(10), 845- 853.15. Itoh, Y., & Klinker, G. (2015). Light-field correction for spatial calibration of optical see-through head-mounted displays. IEEE Transactions on Visualization and Computer Graphics, 21(A), 471-48016. Itoh, Y., & Klinker, G. (2015, September). Simultaneous direct and augmented view distortion calibration of optical see-through head-mounted displays. In 2015 IEEE International Symposium on Mixed and Augmented Reality (pp. 43-48). IEEE17. V. Ferrari, N. Cattari, U. Fontana and F. Cutolo, "Parallax Free Registration for Augmented Reality OpticalSee-through Displays in the Peripersonal Space," in IEEE Transactions on Visualization and Computer Graphics, doi:10.1109 / TVCG.2020.3021534. https : / / ieeexplore. ieee. ore / abstract / document / 9186170

Claims

CLAIMS:

1. A computer system for calibrating a helmet-mounted visual augmentation device with an AR / VR display, the computer system comprising: one or more computer processors; at least a first set of image capture devices configured to capture a first image of a target; a display for displaying the captured first image, wherein the displayed first image comprises a plurality of first pixels of the display; at least a second set of image capture devices configured to capture the displayed image to form a second image with a plurality of second pixels; a computer-readable storage medium comprising a first set of program instructions stored thereon, the first set of program instructions executable by the one or more processors to perform pixel by pixel mapping of the second image with the displayed first image on the display to generate a pixel-by-pixel map; the computer-readable storage medium comprising a second set of program instructions stored thereon, the second set of program instructions executable by the one or more processors to obtain a 3D mapping of the location of the plurality of second pixels and orientation of the plurality of second pixels in relation to the plurality of first pixels; and based on the 3D mapping, the computer-readable storage medium comprising a third set of program instructions stored thereon, the third set of program instructions executable by the one or more processors to align the second image with the first image such that there is substantial alignment between the second image and the first image, thereby calibrating the visual augmentation device.

2. The system of claim 1, wherein first set of image capture devices are mounted on the wearable device, and the second set of image capture devices are mounted within the helmet with a view of the display.

3. The system of claim 3, wherein the display comprises structured-light patterns.

4. The system of any one of claims 1 to 3, wherein the target comprises a plurality of targets.

5. The system of claim 4, further comprising an eye-tracking system configured to determine a location of eyes of a user in relation to the display, thereby triangulating and calibrating the system based on the location of the eyes.

6. The system of claim 6, wherein the system is calibrated for each individual user, thereby allow for personalization to minimize perceptual alterations and motion sickness.

7. The system of claim 6, further comprising a neural network module comprising at least one machine learning model comprising a set of instructions executable by the processor to at least: receive a first pair of stereo images of the target captured by first set of image capture devices and a second pair of stereo images of the target captured by second set of image capture devices; extract a plurality of first feature points from the first pair of stereo images and extract a plurality of second feature points from the first pair of stereo images; match the plurality of first feature points to the plurality of second feature points; generate a disparity map based on a plurality of matching results; and calculate parameters of the first set of image capture devices and the second set of image capture devices based on the disparity map.

8. The system of claim 7, further comprising a low-power stereo matching subsystems to generate the disparity map with the same coordinates in a 2D space and a 3D space.

9. The system of claim 7, further comprising generating a digital twin of the helmet-mounted visual augmentation device with AR / VR display.

10. A method for characterizing and calibrating a helmet-mounted digital visual augmentation device for a user, wherein the method comprises: with cameras mounted on the helmet, acquiring stereo images of a calibration target; calibrating a position of the user’s eyes; with a projection system, projecting structured-light patterns associated with the target on an AR / VR display; acquiring a 3D model of the AR / VR display and obtaining a shape of the AR / VR display; mapping between AR / VR display pixel coordinate and the position in 3D on the calibration target in front of the user’s eyes where the pixel is perceived by the user; encoding deviations of the projection system based on the mapping, wherein the 3D model of the AR / VR display is in a same reference frame as the cameras mounted on the helmet; using at least one machine learning model to correct deviations between an image of the target as seen from a point of view of the user’s eyes and an image of the target captured by the cameras; and generating a digital twin of the digital visual augmentation device.

11. The method of claim 10, wherein the 3D model of the AR / VR display and 3D position, and orientation of each pixel are used to correct parallax and thereby minimize perceptual alterations and motion sickness.

12. The method of claim 11, wherein the AR / VR display comprises structured- light patterns.

13. The method of claim 10, further comprising steps for determining a location of eyes of a user in relation to the AR / VR display for calibrating the digital visual augmentation device based on the location of the eyes.

14. The method of claim 10, further comprising steps for calibrating the digital visual augmentation device for an individual user, thereby allowing for personalization to minimize perceptual alterations and motion sickness.

15. The method of claim 10, further comprising a neural network module comprising the at least one machine learning model comprising a set of instructions executable by the processor to at least: receive a first pair of stereo images of the target captured by a first set of cameras and a second pair of stereo images of the target captured bya second set of cameras; extract a plurality of first feature points from the first pair of stereo images and extract a plurality of second feature points from the first pair of stereo images; match the plurality of first feature points to the plurality of second feature points; generate a disparity map based on a plurality of matching results; and calculate parameters of the first set of cameras and the second set of cameras based on the disparity map.

16. The method of claim 15, further comprising steps for generating the disparity map with the same coordinates in a 2D space and a 3D space.

17. The method of claim 10, further comprising steps for generating a digital twin of the helmet-mounted visual augmentation device with the AR / VR display.

18. An image rendering system for a visual augmentation device associated with a head-mounted display (HMD) on a helmet, the system comprising: a set of cameras mounted on the helmet configured to acquire stereo images of a calibration target; a projection system for projecting structured-light patterns associated with the target on an AR / VR display;acquiring a 3D model of the AR / VR display and obtaining a shape of the AR / VR display; a mapping between AR / VR display pixel coordinate and a position in 3D on the calibration target in front of the user’s eyes where the pixel is perceived by the user; encoding distortions of the projection system based on the mapping, wherein the 3D model of the AR / VR display is in a same reference frame as the set of cameras mounted on the helmet; a neural network module comprising at least one machine learning model trained to correct deviations between an image of the target as seen from a point of view of the user’s eyes and an image of the target captured by the cameras.

19. A method for calibrating a visual augmentation device with a virtual reality (VR) and / or augmented reality (AR) system comprising operations performed by an apparatus comprising a processor coupled to at least one memory device storing instructions executable by the processor to at least perform the operations of:(i) with a first set of image capture devices, capturing a first image of a target comprising a series of known patterns;(ii) displaying the captured first image on a display of the wearable device, wherein the displayed first image comprises a plurality of first pixels of the display;(iii) with a second set of image capture devices within the wearable device, capturing the displayed image of the target to form a second image with a plurality of second pixels;(iv) using the series of known patterns, perform pixel by pixel mapping of the second image from second set of image capture devices with the displayed first image on the display to generate a pixel-by-pixel map;(iv) with the pixel-by-pixel map, obtain a 3-D mapping of the location of the plurality of second pixels and orientation of the plurality of second pixels in relation to the plurality of first pixels; and(v) based on the 3-D mapping, correcting deviations between the second image and the first image, thereby calibrating the visual augmentation device.

20. A system for calibrating a helmet-mounted visual augmentation device with an AR / VR display, the system comprising: an apparatus comprising one or more processors coupled to at least one memory device storing instructions executable by the processor; a neural network module comprising at least one machine learning model comprising a set of instructions executable by the processor to at least: receive a pair of stereo images of a target captured by a first imaging apparatus and a second imaging apparatus; extract a plurality of first feature points from the image of the first imaging apparatus and extract a plurality of second feature points from the image of the second imaging apparatus; match the plurality of first feature points from the image of the first imaging apparatus to the plurality of second feature points from the image of the second imaging apparatus; generate a disparity map based on a plurality of matching results; and calculate parameters of the first imaging apparatus and the second imaging apparatus based on the disparity map.

Citation Information

Patent Citations

  • Method, device, equipment and system for calibrating augmented reality equipment and storage medium

    CN116309854A

  • Calibration for immersive content systems

    US20160253795A1

  • Methods and systems for creating virtual and augmented reality

    US20190094981A1

  • Aligning multiple coordinate systems for information model rendering

    WO2022167505A1

Cited By

  • Image display method of head-mounted display, construction system and computing control unit

    CN120980201A