Learning-based ar-assisted dental treatment automatic calibration and navigation method

By employing a learning-based AR-assisted method, utilizing a depth RGB camera and a feature point detection network for label-free calibration and navigation, the high cost and limited line of sight of optical markers in traditional AR treatment are resolved, achieving efficient and accurate 3D image navigation and simplifying the dental surgery process.

CN116919586BActive Publication Date: 2026-07-31SHENZHEN INST OF ADVANCED TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH
Filing Date
2023-07-03
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In the current technology of digital dental diagnosis and treatment in oral and maxillofacial surgery, traditional AR-based treatment relies on optical markers and trackers, resulting in a simple and costly treatment process. Furthermore, the use of reference markers and bulky tracking devices can easily lead to visual limitations and inaccuracies, and the lack of three-dimensional image navigation results in high complexity of surgical navigation.

Method used

A learning-based AR-assisted method is adopted, which uses a virtual reality display device to acquire a 3D mandibular model. Multiple feature points are identified through a depth RGB camera and a feature point detection network to achieve label-free calibration and navigation. Combined with a multi-feature iterative nearest-point algorithm, the virtual environment and the real environment are accurately aligned.

Benefits of technology

It enables markerless calibration and navigation, improves calibration efficiency and accuracy, avoids the cumbersome operation and equipment cost of optical marking, provides reliable three-dimensional image navigation, and reduces surgical complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116919586B_ABST
    Figure CN116919586B_ABST
Patent Text Reader

Abstract

This invention discloses a learning-based AR-assisted automatic calibration and navigation method for dental treatment. The method includes: acquiring a 3D mandibular model of the target using a virtual reality display device; inputting the 3D mandibular model into a trained feature point detection network to identify multiple corresponding feature points; aligning the multiple feature points with feature points of a corresponding real 3D mandibular model to calibrate and navigate the virtual model under the virtual reality display device onto the real model, wherein the feature points of the real 3D mandibular model are detected using a depth camera; and combining the virtual and real environments based on a multi-feature iterative nearest-point algorithm to complete the display of the virtual projection in the real environment. This invention improves the efficiency and accuracy of automatic calibration and navigation in dental treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical engineering technology, and more specifically, to a learning-based AR-assisted automatic calibration and navigation method for dental treatment. Background Technology

[0002] Oral and maxillofacial surgery is a discipline primarily focused on surgical treatment, with a focus on the prevention and treatment of diseases of the oral organs, facial soft tissues, maxillofacial bones, temporomandibular joint, and certain neck conditions. Computer-assisted therapy (CAT) is a commonly used treatment method that has transformed treatment methods in many different medical fields, including digital dentistry, improving efficiency and accuracy, reducing patient impact, and assisting with preoperative and intraoperative procedures. Augmented reality (AR) technology is gaining increasing popularity in the field of CAT. AR technology guides surgery by overlaying virtual anatomical structures onto the real patient.

[0003] To improve the accuracy and reliability of display calibration procedures, existing solutions are based on optical tracking systems. These systems use multiple markers as dynamic reference frames, firmly fixed to the target anatomical structure to track the direction and posture of the target as the dynamic reference frames move. During tracking, it is typically necessary to calibrate the relative posture between the dynamic reference frame and the offline anatomical structure. Since errors in marker-to-target registration can be distributed throughout the process and can introduce unnecessary errors during marker insertion, registration based on error points or contours can be utilized. Therefore, achieving safety and avoiding invasiveness when using markers in computer-assisted therapy is challenging.

[0004] In existing technologies, Kellner et al. proposed a geometric calibration method with a two-stage concept (Kellner F, Bolte B, Bruder G, et al. Geometric calibration of head-mounted displays and its effects on distance estimation[J].IEEE transactions on visualization and computer graphics,2012,18(4):589-596.), which tracks six-DOF head-mounted markers and three-DOF manual markers, improving the user's interaction. However, this method requires manual marking, reducing the efficiency and accuracy of calibration. For another example, Jun et al. proposed a calibration method (Jun H, Kim GA calibration method for optical see-through head-mounted displays with a depth camera[C] / / 2016IEEE Virtual Reality(VR).IEEE,2016:103-111.), which utilizes a low-cost time-of-flight depth camera to perform two stages: full calibration and simplified calibration to calculate key calibration parameters. However, this method requires the user to point to a virtual circle with their fingertip, which is not only cumbersome but also has a high probability of drawing the wrong virtual circle, thus reducing the accuracy of calibration. .

[0005] Analysis reveals that in the field of digital dental treatment in oral and maxillofacial surgery, traditional AR-based treatments rely on optical markers and trackers. This makes the treatment process simplistic and costly. Furthermore, creating tooth models using baseline markers, retaining reference markers, or employing large optical tracking devices is prone to errors. These baseline markers and bulky tracking devices present difficulties in use, such as limiting the dentist's line of sight or causing inaccuracies due to marker displacement. This increases the complexity of subsequent techniques and necessitates significant modifications to surgical planning. Moreover, traditional computer-aided treatment relies on two-dimensional (2D) imaging rather than perceiving three-dimensional (3D) images for guidance and navigation. This results in a lack of perception of oral cavity depth information by the dentist and causes hand-eye coordination problems, thus surgical navigation in dental treatment remains challenging. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a learning-based AR-assisted automatic calibration and navigation method for dental treatment. This method includes the following steps:

[0007] A 3D mandibular model of the target was obtained using a virtual reality display device;

[0008] The 3D mandibular model is input into a trained feature point detection network to identify multiple corresponding feature points;

[0009] Align the multiple feature points with the feature points of the corresponding real 3D mandibular model to calibrate and navigate the virtual model under the virtual reality display device onto the real model, wherein the feature points of the real 3D mandibular model are detected using a depth camera.

[0010] Based on the multi-feature iterative nearest point algorithm, the virtual environment and the real environment are combined to complete the display of virtual projection in the real environment.

[0011] Compared with the prior art, the advantages of the present invention are that it realizes markerless calibration and navigation based on augmented reality (AR) in digital oral treatment, which can be used for visualization augmented reality based on head-mounted displays. By improving the quality of the depth map between virtual and real distances, the calibration process of virtual models in head-mounted displays is fully automated, thereby improving the efficiency and accuracy of calibration.

[0012] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.

[0014] Figure 1 This is a flowchart of a learning-based AR-assisted dental treatment automatic calibration and navigation method according to an embodiment of the present invention;

[0015] Figure 2 This is a schematic diagram of the structure of a feature point detection network using depth RGB data according to an embodiment of the present invention;

[0016] Figure 3 This is a schematic diagram of three-dimensional model alignment according to an embodiment of the present invention;

[0017] In the attached diagram, Landmarks detection; Conv; Max pool; dropout; upsampling; Output Probability Map; Pixel-wise labeling. Detailed Implementation

[0018] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0019] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.

[0020] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0021] In all the examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0022] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0023] This invention provides a learning-based AR-assisted automatic calibration and navigation method for dental treatment. It is an AR-assisted label-free calibration and navigation scheme that uses an additional depth RGB stereo camera to automatically calibrate a virtual 3D mandibular model under a head-mounted display. In general, a pre-trained convolutional neural network model is first used to detect anatomical feature points using depth RGB images. Then, based on these feature points, the coordinate system of the virtual model is automatically aligned with the world coordinate system to calibrate the virtual model under the head-mounted display. Finally, label-free navigation is performed to overlay the virtual and real environments, thereby enabling accurate surgical navigation for dental treatment. The scheme mainly includes three core components: virtual image or environment modeling, virtual image and real-world space registration, and display technology that combines the virtual and real environments. The final display technology is achieved through a head-mounted display. To enhance the depth navigation and tracking of the 3D model, a depth RGB camera is integrated on top of the head-mounted display. This depth RGB camera utilizes active stereo imaging technology with unstructured light to achieve accurate reconstruction within a limited field of view.

[0024] Specifically, see Figure 1 As shown, the provided learning-based AR-assisted dental treatment automatic calibration and navigation method includes the following steps:

[0025] Step S110: Construct a training dataset, which includes RGB images of the mandible, depth information, and multiple feature points of the mandible.

[0026] In one embodiment, during the preoperative surgical planning phase, patient sample data is acquired using CBCT (cone-beam computed tomography). Based on regions of interest (RoIs), the data is segmented and reconstructed according to the diagnostic and treatment plan. Then, a depth RGB stereo camera is used to record and generate a dataset to train a deep neural network for anatomical landmark detection.

[0027] For example, a dataset reflects the correspondence between RGB images and depth information of the mandible and multiple feature points of the mandible. Feature points are basic features used to identify the mandible, such as the head, neck, base, and beak. The number of feature points can be set according to accuracy and efficiency requirements.

[0028] Furthermore, to enhance the diversity of the dataset, the virtual geometry can employ a variety of different textures and materials. To ensure that the training dataset includes diverse camera poses, moderate random camera position and rotation offsets are introduced before capturing each image. Additionally, various lighting configurations with different intensities, color temperatures, and locations are set to maximize the diversity of the dataset.

[0029] Step S120: Construct a neural network model as a feature point detection network and train it using the training dataset.

[0030] In one embodiment, the mandibular anatomical model feature point detection network is constructed based on an FCN (fully convolutional network), see [link to relevant documentation]. Figure 2 As shown, this feature point detection network consists of an encoder and a decoder. The encoder mainly comprises convolutional layers and max-pooling layers, which gradually reduce the feature map size and capture higher-level semantic information through convolutional operations. The decoder mainly contains upsampling layers and convolutional layers, which gradually recover image detail information through upsampling or deconvolution. Skip connections are designed to connect global and local information to produce more accurate and refined detection results. The purpose of detecting anatomical model feature points is to identify the basic features of the mandibular 3D model, such as the head, neck, base, and beak.

[0031] Furthermore, for feature point detection, the Fully Convolutional Network (FCN) was optimized to predict 3D tooth models by generating dense probability maps after separating the foreground and background. The FCN output was used to label feature points of the mandible and create labels for real-time pixel-by-pixel tracking. The FCN encoder used convolution and pooling operations to compute feature maps with decreasing spatial resolution and increasing depth information. The decoder used transposed convolution and element-wise fusion to generate class score maps with the same spatial dimensions as the input image.

[0032] To perform initial calibration of the HMD (Head-Mounted Display), a depth RGB camera mounted on the HMD captures and identifies features of the mandibular phantom. The captured frames are processed to conform to the input modality required by the FCN (Fully Convolutional Network), which identifies a 3-channel probability map of the probability value for each category (e.g., background = 0, mandible = 1, and feature point = 2) at each pixel location. By combining the probabilities with depth information from the input frames, the mandibular feature points are densely labeled. These labeled pixels are then used to create a data correspondence for model-based position tracking. This tracking method matches the object model with the labeled data, enabling the determination of the object's 6D pose.

[0033] Subsequently, the network was trained using depth RGB images generated from the 3D model, and tested using real 3D models. Furthermore, this invention employs data augmentation techniques to improve performance.

[0034] Step S130: Analyze anatomical feature points using a trained feature point detection network and align them with the feature points of the real 3D model to achieve label-free calibration between the virtual model and the real model.

[0035] Because the feature point detection network is pre-trained for mandibular feature detection, when feature points of a real model are captured using a head-mounted display, the feature point detection network can align the virtual model onto the real model to complete label-free calibration and navigation.

[0036] Specifically, the detected feature points can describe the characteristics of the mandible. Using these detected anatomical feature points, a method for automatically calibrating the virtual model and navigation was further designed. After detecting feature points on the real 3D object using a depth stereo camera, multiple feature points of the virtual model are aligned with feature points of its real 3D model. This method preserves the feature point movement distance between the virtual model and the real object and provides accurate calibration. Figure 3 As shown, in the virtual environment, the six detected feature points are automatically aligned with the real object, among which... Figure 3 (a) Corresponds to a real 3D model. Figure 3 (b) Corresponding virtual 3D model.

[0037] It should be noted that three non-collinear points in space are sufficient to determine the pose of a 3D virtual object, but it is preferable to use six points as they are easier to align and provide better depth cues.

[0038] In one embodiment, the corresponding transformation matrix is ​​calculated, and the RANSAC (Random Sample Consensus) algorithm is used to remove outliers with reprojection errors, refining the most accurate transformation. RANSAC can estimate the parameters of a mathematical model iteratively from a set of observation datasets containing "outsiders," and remove "outsiders" or "noise" by setting relevant thresholds.

[0039] Due to the depth difference between the target and the virtual display, parallax that could lead to misalignment was not considered. For markerless registration, in order to achieve display of pixel p... (s) and the observed 3D real point V (M) Spatial consistency between them is calibrated using the following equation:

[0040]

[0041] Where, p (s) To display the position of the pixels, K is the camera's projection matrix. For the camera's 3D pose, M is the internal transformation matrix of the HMD (Head-Mounted Display). (D) To track attitude, V (M) These are 3D solid points.

[0042] Step S140: Based on feature points, the virtual environment and the real environment are combined using a multi-feature iterative nearest point algorithm to achieve markerless navigation.

[0043] A crucial step in developing head-mounted display-based applications is obtaining the transformation matrix between a real-world stereo camera and a 3D virtual model. In one embodiment, feature point detection techniques from a depth RGB camera are used to obtain the overall transformation and reprojection matrix.

[0044] For example, a multi-feature iterative nearest-neighbor (ICP) algorithm is provided for point-to-point navigation based on detected feature points. This method uses a three-dimensional mandible model. The system automatically estimates the navigation error between the feature points of the real object and the feature points of the 3D virtual model detected from the depth RGB camera. The coordinates of the feature points of the 3D virtual model have been obtained from the depth camera coordinate system, which displays and aligns the 3D model to the coordinate system of the real object to complete markerless navigation. The following transformation matrix is ​​used:

[0045]

[0046] in, It is the reprojected feature point matrix. It is the internal transformation matrix of the head-mounted display, f i(D) This is the original feature point matrix, where n is the number of feature points.

[0047] Specifically, the RANSAC algorithm is used to remove outlier samples based on the reprojection error to achieve a more accurate transformation. The reprojection error is calculated as follows:

[0048]

[0049] To further verify the effectiveness of the invention, experiments were conducted. In the experiments, a depth RGB camera was connected to a commercial head-mounted AR display to detect and calibrate mandibular feature points in real time. The system developed in this invention was verified to achieve a reprojection error of 1.09 ± 0.23 mm pixels from the virtual model to the real object. Furthermore, the calibration achieved a display error of 5.33 ± 1.89 arcmin. In addition, markerless navigation was performed on a dental treatment experiment (mandibular bone) based on real anatomical structures to verify the use of the system in digital dentistry. According to user feedback, the surgical navigation based on the commercial head-mounted display features reliable and stable tracking, low display latency, and rapid alignment with real anatomical structures, achieving overall translational and rotational surgical navigation errors of 3.85 ± 0.62 mm and 2.65 ± 1.47°, respectively. The invention successfully achieves learning-based AR-assisted automatic calibration and navigation for dental treatment.

[0050] In summary, compared with the prior art, the present invention has the following advantages:

[0051] 1) A novel calibration and navigation method for a highly reconfigurable head-mounted display virtual model is proposed. It uses a depth RGB sensing stereo camera to automatically detect feature points in the region of interest (RoI) of the mandible, thereby avoiding the use of optical markers and sensors and enabling the tracking of overlapping virtual models and real objects.

[0052] 2) This invention designs a large-scale depth RGB dataset with multiple occlusions and random regions with different poses, which conforms to real-world scenarios to overcome data problems.

[0053] 3) Addressing the current label-based calibration and navigation methods, this invention proposes a novel learning-based label-free head-mounted display AR system for computer-aided digital dentistry. Furthermore, by improving the quality of the depth map between virtual and real distances, the virtual model calibration process in the head-mounted display is fully automated, improving calibration efficiency and accuracy.

[0054] 4) This invention eliminates user input and avoids external sensors, making the measurement process more accurate and avoiding complex equipment.

[0055] 5) Based on the identified feature points, this invention proposes a navigation method that has a shorter reprojection time, higher accuracy, and enables rapid and accurate projection and quantitative analysis.

[0056] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0057] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0058] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0059] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, Python, etc., and conventional procedural programming languages ​​such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0060] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0061] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0062] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0063] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.

[0064] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.

Claims

1. A learning-based AR-assisted automatic calibration and navigation method for dental treatment, comprising the following steps: A 3D mandibular model of the target was obtained using a virtual reality display device; The 3D mandibular model is input into a trained feature point detection network to identify multiple corresponding feature points; Align the multiple feature points with the feature points of the corresponding real 3D mandibular model to calibrate and navigate the virtual model under the virtual reality display device onto the real model, wherein the feature points of the real 3D mandibular model are detected using a depth camera. Based on the multi-feature iterative nearest point algorithm, the virtual environment and the real environment are combined to complete the display of virtual projection in the real environment; The feature point detection network is a fully convolutional network model, which includes an encoder and a decoder, with skip connections between them. The encoder uses convolution and pooling operations to compute feature maps with decreasing spatial resolution and increasing depth information. The decoder uses transposed convolution and element-wise fusion to generate class score maps with the same spatial dimensions as the input image. The output of the fully convolutional network model is used to label the feature points of the 3D mandibular model. The fully convolutional network model identifies a multi-channel probability map of multiple category probability values ​​at each pixel location for the input RGB image. By combining the multi-channel probability map with depth information from the input frame, the mandibular feature points are labeled. The multiple categories include background, mandible, and feature points. The training dataset for the feature point detection network is constructed according to the following steps: Sample data of the patient's mandible were obtained using cone-beam computed tomography (CBCT) of the oral and maxillofacial region. Based on the region of interest, the sample data is segmented and reconstructed according to the diagnosis and treatment plan; A dataset is generated using a depth RGB stereo camera, the dataset containing RGB images of the mandible, depth information, and multiple feature points of the mandible; The dataset is augmented to construct a training dataset, wherein the augmentation includes: for virtual geometry, using a variety of different textures and materials; and for RGB images, including different camera poses and setting various lighting configurations with different intensities, color temperatures, and positions.

2. The method according to claim 1, characterized in that, The registration between the display pixel points and the 3D real points is achieved according to the following formula: and 3D real points between the display pixel points and the 3D real points is achieved according to the following formula: in, To display the position of the pixel, K The projection matrix of the depth camera. For the stereo pose of the depth camera, It is the internal transformation matrix of a virtual reality display device. To track attitude, These are 3D solid points.

3. The method according to claim 1, characterized in that, The following formula is used to display the virtual projection in the real environment: in, It is the reprojected feature point matrix. It is the internal transformation matrix of a virtual reality display device. It is the original feature point matrix, where n is the number of feature points.

4. The method according to claim 3, characterized in that, During the process of displaying the virtual projection in the real environment, points with reprojection errors greater than a set threshold are discarded as outlier samples. The reprojection error is calculated according to the following formula: 。 5. The method according to claim 1, characterized in that, The virtual reality display device is a head-mounted AR display with a depth RGB camera integrated on top.

6. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

7. A computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.