Track point pose information sequence generation method and augmented reality display equipment

By collecting and optimizing image frames and inertial measurement data of augmented reality display devices, the optimized trajectory position information sequence is generated, which solves the trajectory drift problem and improves the accuracy and calculation efficiency of the trajectory position sequence.

CN120471985AActive Publication Date: 2025-08-12HANGZHOU LINGBAN TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510527058.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-12
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the prior art, when generating the trajectory position pose sequence of the augmented reality display device, there is a serious problem of trajectory drift, which is mainly due to the large cumulative error caused by sensor noise, visual feature matching error and IMU drift, which affects the accuracy of the trajectory position pose sequence.

Method used

By collecting image frame sequences and inertial measurement data sequences of preset scenes containing three-dimensional feature points, an initial trajectory point pose information sequence is generated, and the sliding window loop optimization process is used to reduce cumulative errors, and optimize trajectory point pose information sequences are generated to improve accuracy.

Benefits of technology

It effectively reduces trajectory drift, improves the accuracy of trajectory point pose sequences, reduces the waste of computer computing resources, and enhances the adaptability and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471985A_ABST
    Figure CN120471985A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a trajectory point pose information sequence generation method and augmented reality display equipment. One specific embodiment of the method comprises the following steps: for a preset scene containing three-dimensional feature points, acquiring an image frame sequence and an inertial measurement data sequence corresponding to the preset scene; based on the image frame sequence and the inertial measurement data sequence, generating an initial trajectory point pose information sequence; performing sliding window loopback optimization processing on the image frame sequence to generate at least one piece of optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence; and based on the at least one piece of optimized trajectory point pose information, updating the initial trajectory point pose information sequence to obtain a trajectory point pose information sequence. According to the embodiment, the trajectory drift degree of the equipment trajectory displayed based on the trajectory point pose sequence is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a method for generating a trajectory point pose information sequence and an augmented reality display device. Background Art

[0002] Track point pose information sequence generation is a technology that generates a track point pose sequence for an augmented reality display device (e.g., an AR device) during its movement. Currently, this is typically done by using a VIO algorithm to estimate the pose sequence of the AR display device during its movement using environmental and motion data.

[0003] However, the inventors have discovered that when the above method is used to generate a trajectory point pose sequence of an augmented reality display device during movement, the following technical problems often arise:

[0004] VIO relies on data association between adjacent frames for pose estimation, but sensor noise, visual feature matching errors and IMU (inertial measurement unit) drift will accumulate over time, causing the trajectory deviation to gradually increase, which in turn leads to poor accuracy of the trajectory point pose sequence generated by the augmented reality display device during movement, resulting in a high degree of trajectory drift in the device trajectory displayed based on the trajectory point pose sequence.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention

[0006] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] Some embodiments of the present disclosure propose a method for generating a trajectory point pose information sequence and an augmented reality display device to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating a trajectory point pose information sequence, the method comprising: for a preset scene containing three-dimensional feature points, collecting an image frame sequence and an inertial measurement data sequence corresponding to the preset scene; generating an initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence; performing sliding window loop optimization processing on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence; and updating the initial trajectory point pose information sequence based on the at least one optimized trajectory point pose information to obtain a trajectory point pose information sequence.

[0009] In a second aspect, some embodiments of the present disclosure provide an augmented reality display device, which is an augmented reality head-mounted display device or an augmented reality handheld display device, and the above-mentioned augmented reality display device includes: one or more processors; a camera device, for acquiring image frames corresponding to the above-mentioned preset scenes according to a preset acquisition frequency, and in response to determining that the acquired image frames meet preset labeling conditions, marking the acquired image frames as loop candidate image frames, thereby obtaining an image frame sequence containing at least one loop candidate image frame; an inertial measurement component, for acquiring an inertial measurement data sequence; a storage device, for storing one or more programs; a display module, for displaying the trajectory corresponding to the above-mentioned trajectory point posture information sequence during imaging; when the above-mentioned one or more programs are executed by the above-mentioned one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.

[0010] The above-described various embodiments of the present disclosure have the following beneficial effects: Through the trajectory point pose information sequence generation method of some embodiments of the present disclosure, the accuracy of the generated trajectory point pose sequence is improved, and the degree of trajectory drift of the device trajectory displayed based on the trajectory point pose sequence is reduced. Specifically, the poor accuracy of the generated trajectory point pose sequence and the high degree of trajectory drift of the device trajectory displayed based on the trajectory point pose sequence are caused by: VIO relies on data correlation between adjacent frames for pose estimation, but sensor noise, visual feature matching errors, and IMU (inertial measurement unit) drift accumulate over time, causing the trajectory deviation to gradually increase. This in turn leads to poor accuracy of the trajectory point pose sequence generated by the augmented reality display device during movement, resulting in a high degree of trajectory drift of the device trajectory displayed based on the trajectory point pose sequence. Based on this, the trajectory point pose information sequence generation method of some embodiments of the present disclosure first collects an image frame sequence and an inertial measurement data sequence corresponding to a preset scene containing three-dimensional feature points. This can obtain the image frame sequence and inertial measurement data sequence collected by the augmented reality display device during movement. Next, based on the image frame sequence and the inertial measurement data sequence, an initial trajectory point pose information sequence is generated. This generates an initial trajectory point pose information sequence that initially represents the trajectory of the augmented reality display device during movement. Subsequently, a sliding window loop closure optimization process is performed on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence. Thus, a sliding window loop closure optimization process can be used to generate at least one optimized trajectory point pose information for updating the initial trajectory point pose information sequence. Finally, based on the at least one optimized trajectory point pose information, the initial trajectory point pose information sequence is updated to obtain a trajectory point pose information sequence. This allows the initial trajectory point pose information sequence to be optimized using the at least one optimized trajectory point pose information, reducing the cumulative error in the initial trajectory point pose information sequence, reducing the difference between the trajectory represented by the generated trajectory point pose information sequence and the actual trajectory of the augmented reality display device during movement, and reducing the degree of trajectory drift of the device trajectory displayed based on the trajectory point pose sequence. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0012] Figure 1 is a flowchart of some embodiments of the method for generating trajectory point pose information sequences using some embodiments of the present disclosure;

[0013] Figure 2 1 is a schematic structural diagram of other embodiments of the method for generating trajectory point pose information sequences according to the present disclosure;

[0014] Figure 3 is an internal test diagram of the trajectory point pose information sequence generation method disclosed herein;

[0015] Figure 4 Schematic diagram of the hardware structure of the augmented reality display device according to the present disclosure. DETAILED DESCRIPTION

[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0017] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0019] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0020] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0021] Figure 1 The process 100 of some embodiments of the method for generating a trajectory point pose information sequence according to the present disclosure is shown. The method for generating a trajectory point pose information sequence includes the following steps:

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Step 101 : For a preset scene containing three-dimensional feature points, an image frame sequence and an inertial measurement data sequence corresponding to the preset scene are collected.

[0024] In some embodiments, the execution subject (e.g., an AR device) of the trajectory point pose information sequence generation method can collect an image frame sequence and an inertial measurement data sequence corresponding to a preset scene containing three-dimensional feature points. The execution subject may be an augmented reality display device (e.g., an AR device). The augmented reality display device may be an augmented reality head-mounted display device (e.g., a user-mounted AR device) or an augmented reality handheld display device (e.g., a handheld AR device). The image frame sequence may be an image sequence collected by the augmented reality display device during movement in the preset scene. Each inertial measurement data in the inertial measurement data sequence corresponds to a corresponding image frame in the image frame sequence. The execution subject may collect its own inertial measurement data while collecting image frames. The inertial measurement data includes acceleration and angular velocity.

[0025] In some optional implementations of some embodiments, the execution entity may collect an image frame sequence and an inertial measurement data sequence corresponding to the preset scene through the following steps:

[0026] In the first step, the camera device captures image frames corresponding to the above-mentioned preset scene at a preset capture frequency, and in response to determining that the captured image frames meet the preset labeling conditions, the captured image frames are marked as loop candidate image frames to obtain an image frame sequence containing at least one loop candidate image frame. The preset labeling conditions may be, but are not limited to, one of the following: the number of times the image frame is repeatedly observed exceeds a preset number (for example, the number of times the image frame successfully matches the historical frame is greater than 3), the similarity between the environmental features (such as planes, straight lines, key points) similar to the historical frames is greater than a preset value (for example, the similarity between the wall features in the image frame and the wall features of the preset historical frames is greater than a preset value), and there is corresponding user interaction information (for example, the interaction information of the user clicking the image frame as an anchor point). The camera device may be a camera on an augmented reality display device.

[0027] The second step is to collect an inertial measurement data sequence using the inertial measurement assembly. Each image frame in the image frame sequence corresponds to a piece of inertial measurement data in the inertial measurement data sequence. The image frames and the inertial measurement data are collected at the same time. The inertial measurement assembly includes an accelerometer and a gyroscope.

[0028] Step 102: Generate an initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence.

[0029] In some embodiments, the execution entity may generate an initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence. In practice, the execution entity may generate the initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence using visual inertial odometry (VIO) technology. The initial trajectory point pose information sequence may represent the motion trajectory and pose of the augmented reality display device in three-dimensional space. Each initial trajectory point pose information in the initial trajectory point pose information sequence corresponds to an image frame in the image frame sequence. The initial trajectory point pose information may represent the position and pose of the augmented reality display device in three-dimensional space. For example, the initial trajectory point pose information may be "position (m): (0, 0, 0), pose (quaternion): (0.0, 0.0, 0.0, 1.0)." Alternatively, in practice, the execution entity may also generate the initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence using tight coupling and loose coupling methods.

[0030] Step 103: Perform sliding window loop closure optimization processing on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence.

[0031] In some embodiments, the execution entity may perform a sliding window loop closure optimization process on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence. In practice, the execution entity may use a preset image as a loop closure image. Then, the execution entity may input at least one image frame in the image frame sequence into the sliding window. Then, for each image frame in the sliding window, the execution entity may calculate the relative pose between the loop closure image and the image frame as a loop closure constraint through feature matching and geometric verification, and at the same time, by minimizing the loop closure constraint, visual reprojection error, and IMU pre-integration error, solve the optimized pose information corresponding to the image frame as the optimized trajectory point pose information through a nonlinear optimization algorithm (such as Ceres Solver or g2o). The optimized trajectory point pose information may represent the position and posture of the augmented reality display device that captures the image frame in three-dimensional space. The preset image may be a pre-selected scene image containing the three-dimensional feature points.

[0032] In some optional implementations of some embodiments, the execution entity may perform sliding window loop optimization on the image frame sequence through the following steps to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence:

[0033] In the first step, the image frame sequence is divided according to a sliding window method to obtain an initial image frame queue sequence. The first initial image frame queue in the initial image frame queue sequence includes at least one loop candidate image frame. In practice, the execution subject can divide the initial image frame queue sequence according to a sliding window with a sliding window size of a first preset value and a sliding step of a second value, and input each initial image frame divided into a group into the queue to obtain an initial image frame queue sequence. As an example, the image frame sequence can be "{image frame 1, image frame 2, image frame 3, image frame 4, image frame 5, image frame 6, image frame 7, image frame 8}", the sliding window size can be 3, and the sliding step can be 1. Then the initial image frame queue sequence can be "{{image frame 1, image frame 2, image frame 3}, {image frame 2, image frame 3, image frame 4}, {image frame 3, image frame 4, image frame 5}, {image frame 4, image frame 5, image frame 6}, {image frame 5, image frame 6, image frame 7}, {image frame 6, image frame 7, image frame 8}}".

[0034] In the second step, for each initial image frame queue in the above initial image frame queue sequence, perform the following steps:

[0035] The first sub-step is to determine the last initial image frame in the initial image frame queue as the reference image frame.

[0036] In a second sub-step, at least one loop closure candidate image frame located before the reference image frame in the image frame sequence is determined as a loop closure candidate image frame sequence.

[0037] In a third sub-step, in response to determining that there is a loop candidate image frame in the loop candidate image frame sequence in the above-mentioned initial image frame queue, based on the above-mentioned initial image frame queue, optimized trajectory point pose information corresponding to the last initial image frame in the above-mentioned initial image frame queue is generated.

[0038] Therefore, through the sliding window and the loop closure candidate image frame sequence, the pose drift of the reference frame can be corrected through the loop closure constraint within the local window, and at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence can be generated.

[0039] In the process of adopting technical solutions to solve the problems mentioned in the background technology, the following problems often arise:

[0040] In the process of generating the optimized trajectory point pose information corresponding to the last initial image frame in the initial image frame queue based on the loop candidate image frames and the initial image frame queue in the initial image frame queue, when the loop candidate image frames in the initial image frame queue do not have a common view with the last initial image frame, the system will still attempt to perform feature matching (such as SIFT, ORB) and geometric verification (such as RANSAC). These operations require CPU and GPU resources, resulting in a waste of computer computing resources. At the same time, when there are multiple loop candidate image frames in the initial image frame queue, a loop candidate image is randomly selected to generate the optimized trajectory point pose information. The randomly selected loop image frame may only have a small common view area with the last initial image frame, or even no overlap at all. In the process of generating the optimized trajectory point pose information, the optimization algorithm requires more iterations to barely converge, resulting in a waste of computer computing resources.

[0041] Faced with the above technical problems, the inventors decided to adopt the following solutions:

[0042] In some optional implementations of some embodiments, the execution entity may generate optimized trajectory point pose information corresponding to the last initial image frame in the initial image frame queue based on the initial image frame queue through the following steps:

[0043] In the first step, in response to determining that the number of loop closure candidate image frames included in the initial image frame queue is one and that the included loop closure candidate image frame has a common view with the last initial image frame, the loop closure candidate image frame included in the initial image frame queue is determined as a reference loop closure image frame. In practice, the execution entity may perform feature point extraction processing on the loop closure candidate image frame and the last initial image frame using a feature point extraction algorithm (e.g., a SIFT algorithm) to obtain a first feature point extraction information set corresponding to the loop closure candidate image frame and a second feature point extraction information set corresponding to the last initial image frame. Thereafter, the execution entity may input the first feature point extraction information set and the second feature point extraction information set into a pre-trained feature point matching model (e.g., a SuperGlue model, a LoFTR model) to obtain feature point extraction information matching pairs. In response to determining that the number of feature point extraction information matching pairs in each of the feature point extraction information matching pairs is greater than a preset value, the execution entity may determine that the loop closure candidate image frame has a common view with the last initial image frame. Each piece of first feature point extraction information in the first feature point extraction information set may represent a feature point in a loop closure candidate image frame. The first feature point extraction information may be represented by a feature descriptor. Each piece of second feature point extraction information in the second feature point extraction information set may represent a feature point in the last initial image frame. The second feature point extraction information may be represented by a feature descriptor.

[0044] In the second step, in response to determining that the number of loop closure candidate image frames included in the above-mentioned initial image frame queue is greater than one, each loop closure candidate image frame included in the above-mentioned initial image frame queue is determined as each loop closure candidate image frame to be screened.

[0045] In the third step, the last initial image frame is determined as the target image frame.

[0046] In the fourth step, one of the above-mentioned loop candidate image frames to be screened that meets the preset maximum common view area screening condition is determined as the maximum common view image as the reference loop candidate image frame. The above-mentioned preset maximum common view area screening condition can be the largest common view area with the target image frame. The size of the common view area with the target image frame can be represented by the number of feature point pairs that match the loop candidate image frame to be screened and the target image frame. Optionally, the size of the common view area with the target image frame can also be represented by the number of pixels in the overlapping area between the loop candidate image frame to be screened and the target image frame. The above-mentioned execution entity can find the overlapping area between the loop candidate image frame to be screened and the target image frame through image registration (such as homography transformation or basic matrix).

[0047] In a fifth step, the inertial measurement data corresponding to the initial image frames included in the initial image frame sequence in the inertial measurement data sequence are arranged to obtain a target inertial measurement data sequence. In practice, the execution entity may first determine the inertial measurement data corresponding to the initial image frames included in the initial image frame sequence in the inertial measurement data sequence as the target inertial measurement data. The execution entity may then arrange the target inertial measurement data in the order in which they appear in the inertial measurement data sequence to obtain the target inertial measurement data sequence.

[0048] The sixth step is to generate optimization objective function information based on the target inertial measurement data sequence, the initial image frame queue, the reference loop image frame and the target image frame. The optimization objective function information can be a mathematical expression that combines the IMU constraint information, the visual constraint information (reprojection error information) and the loop error term. The expression describes the comprehensive influence of all constraints on the trajectory point pose (position and attitude). In practice, the execution entity can generate IMU constraint information based on the target inertial measurement data sequence and the preset inertial measurement component noise parameter. Then, based on the initial image frame queue, the three-dimensional coordinates of the three-dimensional feature points and the camera parameter information of the camera device, the reprojection error information is generated as the visual constraint information. Then, based on the reference loop image frame and the target image frame, the loop error term is generated. Finally, based on the IMU constraint information, the visual constraint information and the loop error term, the optimization objective function information is generated.

[0049] The seventh step is to iteratively optimize the objective function corresponding to the above-mentioned optimization objective function information to generate optimized trajectory point pose information corresponding to the above-mentioned target image frame. In practice, the above-mentioned execution entity can use a nonlinear optimization algorithm such as the Gauss-Newton method or the Levenberg-Marquardt method to iteratively optimize the objective function to obtain the optimized trajectory point pose information corresponding to the target image frame. The above-mentioned optimized trajectory point pose information can represent the position and posture of the augmented reality display device that captured the target image frame in three-dimensional space.

[0050] The above technical solution and its related contents, as an inventive point of an embodiment of the present disclosure, solve the technical problem of "waste of computer computing resources". The factors that lead to the waste of computer computing resources are often as follows: in the process of generating the optimized trajectory point pose information corresponding to the last initial image frame in the above initial image frame queue based on the loop candidate image frames and the initial image frame queue in the initial image frame queue, when the loop candidate image frames in the initial image frame queue do not have a common view with the last initial image frame, the system will still try to perform feature matching (such as SIFT, ORB) and geometric verification (such as RANSAC). These operations require CPU and GPU resources, resulting in a waste of computer computing resources. At the same time, when there are multiple loop candidate image frames in the initial image frame queue, a loop candidate image is randomly selected to generate the optimized trajectory point pose information. The randomly selected loop image frame may only have a small common view area with the last initial image frame, or even no overlap at all. In the process of generating the optimized trajectory point pose information, the optimization algorithm requires more iterations to barely converge, resulting in a waste of computer computing resources. If the above factors are solved, the effect of reducing the waste of computer computing resources can be achieved. To achieve this effect, first, in response to determining that the number of loop closure candidate image frames included in the initial image frame queue is one and that the included loop closure candidate image frames have co-viewing relationships with the last initial image frame, the loop closure candidate image frame included in the initial image frame queue is determined as a reference loop closure image frame. Thus, if the initial image frame queue includes a loop closure candidate image frame that has co-viewing relationships with the initial image frame, the loop closure candidate image frame can be determined as a reference loop closure image frame for generating optimization objective function information and optimized trajectory point pose information. Next, in response to determining that the number of loop closure candidate image frames included in the initial image frame queue is greater than one, each loop closure candidate image frame included in the initial image frame queue is determined as a loop closure candidate image frame to be screened. Thus, each loop closure candidate image frame to be screened included in the initial image frame queue can be determined. Thereafter, the last initial image frame is determined as a target image frame. Then, among the loop closure candidate image frames to be screened, a loop closure candidate image frame to be screened that meets a preset maximum co-viewing region screening condition is determined as a maximum co-viewing image, serving as a reference loop closure candidate image frame. Thus, a candidate loop closure image with the largest common view area with the target image frame can be selected from each candidate loop closure image frame to be screened as a reference loop closure image frame for generating the optimization objective function information and the optimized trajectory point pose information. Subsequently, the inertial measurement data corresponding to each initial image frame included in the initial image frame queue in the above-mentioned inertial measurement data sequence are arranged to obtain a target inertial measurement data sequence. Thus, a target inertial measurement data sequence for generating the optimization objective function information can be obtained.Next, optimization objective function information is generated based on the target inertial measurement data sequence, the initial image frame queue, the reference loop closure image frame, and the target image frame. The objective function corresponding to the optimization objective function information is then iteratively optimized to generate optimized trajectory point pose information corresponding to the target image frame. This allows the optimization of the objective function information to be iteratively optimized to generate optimized trajectory point pose information. Furthermore, when generating optimized trajectory point pose information corresponding to the target image frame based on the initial image frame queue, when the initial image frame queue contains only one loop closure candidate image frame that is co-visually visible with the last initial image frame, that frame is determined as the reference loop closure image frame for generating the optimization objective function information and optimized trajectory point pose information. Alternatively, when the initial image frame queue contains multiple loop closure candidate image frames, the last initial image frame is first determined as the target image frame. Then, from each of the candidate loop closure image frames to be screened, the one with the largest co-visually visible area with the target image frame is selected as the reference loop closure image frame, thereby ensuring that the selected reference loop closure image frame has a larger co-visually visible area with the target image frame, rather than being randomly selected. Furthermore, when generating the optimized trajectory point pose information based on the selected reference loop closure image frame, unnecessary generation operations are avoided when there is no common view. At the same time, among multiple loop closure candidate image frames, the image frame with the largest common view area is selected as the reference loop closure image frame, ensuring that the objective function can converge faster when generating the optimized trajectory point pose information, reducing the waste of computer computing resources.

[0051] In some optional implementations of some embodiments, after determining at least one loop closure candidate image frame preceding the reference image frame in the image frame sequence as a loop closure candidate image frame sequence, the execution entity may further perform the following steps:

[0052] In the first step, in response to determining that the initial image frame queue does not contain the loop closure candidate image frame in the loop closure candidate image frame sequence, the following steps are performed:

[0053] In a first sub-step, a loop closure candidate image frame in the loop closure candidate image frame sequence that meets a preset screening condition is replaced with the first initial image frame in the initial image frame queue to update the initial image frame queue. The preset screening condition may be that the common viewing area with the last initial image frame in the initial image frame queue is the largest.

[0054] The second sub-step is to generate, based on the updated initial image frame queue, optimized trajectory point pose information corresponding to the last initial image frame in the updated initial image frame queue.

[0055] Therefore, in the case where there is no loop closure candidate image frame in the initial image frame queue, the initial image frame queue can be updated by replacing the loop closure candidate image frame, and the optimized trajectory point pose information is generated based on the updated queue, thereby enhancing the adaptability and robustness of the system in the absence of direct loop closure matching.

[0056] In some optional implementations of some embodiments, the execution entity may generate optimized trajectory point pose information corresponding to the last initial image frame in the updated initial image frame queue based on the updated initial image frame queue through the following steps:

[0057] In the first step, the last initial image frame in the updated initial image frame queue is determined as the target image frame.

[0058] In the second step, the loop closure candidate image frame in the updated initial image frame queue is determined as the reference loop closure candidate image frame.

[0059] In the third step, a common view detection process is performed on the target image frame and the reference loop closure candidate image frame to obtain common view detection information. In practice, the execution entity may perform feature point extraction on the target image frame and the reference loop closure candidate image frame respectively through a feature point extraction algorithm (e.g., SIFT algorithm) to obtain a target image frame feature point extraction information set corresponding to the target image frame and a reference loop closure candidate image frame feature point extraction information set corresponding to the reference loop closure candidate image frame. Afterwards, the execution entity may input the target image frame feature point extraction information set and the reference loop closure candidate image frame feature point extraction information set into a pre-trained feature point matching model (e.g., SuperGlue model, LoFTR model) to obtain each feature point extraction information matching pair. In response to determining that the number of feature point extraction information matching pairs in each feature point extraction information matching pair is greater than a preset value, the execution entity may determine the information characterizing the existence of common view between the target image frame and the reference loop closure candidate image frame as the common view detection information. In response to determining that the number of feature point extraction information matching pairs in each feature point extraction information matching pair is less than or equal to a preset value, the execution entity may determine information indicating that there is common view between the target image frame and the reference loop closure candidate image frame as common view detection information. The common view detection information may be a Boolean value, for example, True indicating that there is common view between the target image frame and the reference loop closure candidate image frame, and False indicating that there is no common view between the target image frame and the reference loop closure candidate image frame.

[0060] The fourth step is to generate optimized trajectory point pose information corresponding to the target image frame based on the common view detection information and the updated initial image frame queue.

[0061] Therefore, the target image frame and the reference loop closure candidate image frame in the updated initial image frame queue can be used for common view detection, and the trajectory point pose information corresponding to the target image frame can be optimized based on the detection results to correct the accumulated error in the trajectory and improve the consistency between the trajectory corresponding to the optimized trajectory point pose information and the true trajectory.

[0062] In some optional implementations of some embodiments, the execution entity may generate optimized trajectory point pose information corresponding to the target image frame based on the common view detection information and the updated initial image frame queue through the following steps:

[0063] In the first step, in response to determining that the common view detection information indicates that the target image frame and the reference loop closure candidate image frame have common view, the inertial measurement data corresponding to the initial image frames included in the initial image frame queue in the inertial measurement data sequence are arranged to obtain a target inertial measurement data sequence, and the reference loop closure candidate image frame is determined as a reference loop closure image frame.

[0064] The second step is to generate optimization objective function information based on the target inertial measurement data sequence, the initial image frame queue, the reference loop image frame and the target image frame.

[0065] The third step is to iteratively optimize the objective function corresponding to the optimization objective function information to generate optimized trajectory point pose information corresponding to the target image frame. In practice, the execution entity can iteratively optimize the objective function corresponding to the optimization objective function information using a nonlinear optimization algorithm such as the Gauss-Newton method or the Levenberg-Marquardt method to obtain the optimized trajectory point pose information.

[0066] Therefore, an objective function can be constructed based on inertial measurement data, the initial image frame queue, the reference loop image frame and the target image frame, and then the objective function can be iteratively optimized to generate more accurate optimized trajectory point pose information corresponding to the target image frame.

[0067] In some optional implementations of some embodiments, the execution entity may generate optimization objective function information based on the target inertial measurement data sequence, the initial image frame queue, the reference loop-back image frame, and the target image frame through the following steps:

[0068] The first step is to generate IMU constraint information based on the target inertial measurement data sequence and preset inertial measurement unit noise parameters. The preset inertial measurement unit noise parameters may be a noise covariance matrix, representing the random error introduced by the inertial measurement unit when measuring angular velocity and acceleration. In practice, the execution entity may generate IMU constraint information based on the target inertial measurement data sequence and the preset inertial measurement unit noise parameters using an inertial navigation algorithm. The IMU constraint information may represent an IMU constraint.

[0069] In the second step, based on the above-mentioned initial image frame queue, the three-dimensional coordinates of the above-mentioned three-dimensional feature points and the camera parameter information of the camera device, reprojection error information is generated as visual constraint information. In practice, for each initial image frame in the above-mentioned initial image frame queue, the above-mentioned execution subject can identify the pixel position of the three-dimensional feature point in the initial image frame through the SIFT algorithm. Then the above-mentioned execution subject can use the camera pose, camera external parameters and camera internal parameters to convert the three-dimensional coordinates of the three-dimensional feature point into projection coordinates. Afterwards, the above-mentioned execution subject can determine the Euclidean distance between the above-mentioned pixel position and the above-mentioned projection coordinates as the reprojection error corresponding to the above-mentioned initial image frame. Then, the above-mentioned execution subject can construct the above-mentioned various reprojection errors in the form of residual terms (or factor graph edges) as visual constraint information through the Bundle Adjustment (BA) in visual SLAM. Among them, the above-mentioned visual constraint information can represent visual constraints.

[0070] The third step is to generate a loop closure error term based on the reference loop closure image frame and the target image frame. In practice, the execution entity can detect feature point pairs between the target image frame and the reference loop closure image frame through feature matching or a deep learning model. Then, for each feature point pair, the execution entity can determine its projection error in the two frames. The execution entity can then call a general nonlinear optimization library (such as Ceres Solver, g2o, GTSAM, etc.) to generate a loop closure error term based on each feature point pair. The loop closure error term can represent a loop closure constraint.

[0071] The fourth step is to generate optimization objective function information based on the IMU constraint information, the visual constraint information, and the loop closure error term. The optimization objective function information may be a weighted sum of the IMU constraint information, the visual constraint information, and the loop closure error term. In practice, the execution entity may determine the optimization objective function information as the weighted sum of the IMU constraint information, the visual constraint information, and the loop closure error term.

[0072] Step 104 : Based on at least one optimized trajectory point pose information, the initial trajectory point pose information sequence is updated to obtain a trajectory point pose information sequence.

[0073] In some embodiments, the execution entity may update the initial trajectory point pose information sequence based on the at least one optimized trajectory point pose information to obtain a trajectory point pose information sequence.

[0074] In some optional implementations of some embodiments, the execution entity may update the initial trajectory point pose information sequence based on the at least one optimized trajectory point pose information to obtain a trajectory point pose information sequence:

[0075] In the first step, for each optimized trajectory point pose information in the at least one optimized trajectory point pose information, the following steps are performed:

[0076] In the second step, the image frame corresponding to the above-mentioned optimized trajectory point posture information in the above-mentioned image frame sequence is determined as the reference image frame.

[0077] In the second step, the optimized trajectory point pose information is replaced with the initial trajectory point pose information corresponding to the reference image frame in the initial trajectory point pose information sequence, so as to update the initial trajectory point pose information sequence.

[0078] The fourth step is to determine the updated initial trajectory point pose information sequence as the trajectory point pose information sequence.

[0079] In some optional implementations of some embodiments, after the initial trajectory point pose information sequence is updated based on the at least one optimized trajectory point pose information to obtain the trajectory point pose information sequence, the execution entity may further perform the following steps:

[0080] Displaying the trajectory corresponding to the trajectory point position information sequence. In practice, the execution subject can perform imaging through a display module to display the trajectory corresponding to the trajectory point position information sequence.

[0081] In this way, the trajectory corresponding to the trajectory point posture information sequence can be displayed.

[0082] The above-described various embodiments of the present disclosure have the following beneficial effects: Through the methods for generating trajectory point pose information sequences according to some embodiments of the present disclosure, the accuracy of the generated trajectory point pose sequences is improved, and the degree of trajectory drift of the device trajectory displayed based on the trajectory point pose sequences is reduced. Specifically, the poor accuracy of the generated trajectory point pose sequences and the high degree of trajectory drift of the device trajectory displayed based on the trajectory point pose sequences are caused by the fact that VIO relies on data correlation between adjacent frames for pose estimation. However, sensor noise, visual feature matching errors, and IMU (inertial measurement unit) drift accumulate over time, causing trajectory deviations to gradually increase. This, in turn, leads to poor accuracy of the trajectory point pose sequences generated by the augmented reality display device during movement, and high degree of trajectory drift of the device trajectory displayed based on the trajectory point pose sequences. Based on this, the methods for generating trajectory point pose information sequences according to some embodiments of the present disclosure first collect, for a preset scene containing three-dimensional feature points, a sequence of image frames and a sequence of inertial measurement data corresponding to the preset scene. This results in a sequence of image frames and a sequence of inertial measurement data collected by the augmented reality display device during movement. Next, based on the image frame sequence and the inertial measurement data sequence, an initial trajectory point pose information sequence is generated. This generates an initial trajectory point pose information sequence that initially represents the trajectory of the augmented reality display device during movement. Subsequently, a sliding window loop closure optimization process is performed on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence. Thus, a sliding window loop closure optimization process can be used to generate at least one optimized trajectory point pose information for updating the initial trajectory point pose information sequence. Finally, based on the at least one optimized trajectory point pose information, the initial trajectory point pose information sequence is updated to obtain a trajectory point pose information sequence. This allows the initial trajectory point pose information sequence to be optimized using the at least one optimized trajectory point pose information, reducing the cumulative error in the initial trajectory point pose information sequence, reducing the difference between the trajectory represented by the generated trajectory point pose information sequence and the actual trajectory of the augmented reality display device during movement, and reducing the degree of trajectory drift of the device trajectory displayed based on the trajectory point pose sequence.

[0083] Further references Figure 2 , which shows a process 200 of another embodiment of a method for generating a trajectory point pose information sequence. The process 200 of the method for generating a trajectory point pose information sequence includes the following steps:

[0084] Step 201: For a preset scene containing three-dimensional feature points, an image frame corresponding to the preset scene is captured by a camera device at a preset capture frequency, and in response to determining that the captured image frame meets a preset labeling condition, the captured image frame is marked as a loop candidate image frame, thereby obtaining an image frame sequence including at least one loop candidate image frame.

[0085] In some embodiments, the above-mentioned execution entity can capture image frames corresponding to a preset scene containing three-dimensional feature points through a camera device at a preset capture frequency, and in response to determining that the captured image frames meet preset labeling conditions, mark the captured image frames as loop candidate image frames, thereby obtaining an image frame sequence containing at least one loop candidate image frame.

[0086] Step 202: Collect an inertial measurement data sequence through an inertial measurement assembly.

[0087] In some embodiments, the execution entity may collect an inertial measurement data sequence through the inertial measurement component, wherein each image frame in the image frame sequence corresponds to an inertial measurement data in the inertial measurement data sequence.

[0088] Step 203: Generate an initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence.

[0089] In some embodiments, the execution entity may generate an initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence.

[0090] Step 204 : Perform sliding window loop closure optimization processing on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence.

[0091] In some embodiments, the execution entity may perform sliding window loop optimization processing on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence.

[0092] Step 205 : Based on at least one optimized trajectory point pose information, the initial trajectory point pose information sequence is updated to obtain a trajectory point pose information sequence.

[0093] In some embodiments, the execution entity may update the initial trajectory point pose information sequence based on the at least one optimized trajectory point pose information to obtain a trajectory point pose information sequence.

[0094] In from Figure 2 It can be seen that Figure 1 Compared with the description of some corresponding embodiments, Figure 2The process 200 of the method for generating a trajectory point pose information sequence in some corresponding embodiments embodies capturing image frames corresponding to the above-mentioned preset scene at a preset capture frequency by a camera device, and in response to determining that the captured image frames meet a preset labeling condition, marking the captured image frames as loop candidate image frames, thereby obtaining an image frame sequence including at least one loop candidate image frame. The inertial measurement data sequence is captured by the above-mentioned inertial measurement component, wherein each image frame in the above-mentioned image frame sequence corresponds to an inertial measurement data in the above-mentioned inertial measurement data sequence. Also, because when capturing image frames, the captured image frames are marked as loop candidate image frames in real time when the image frames meet the preset labeling condition, since the image frame capture and the labeling of the loop candidate image frames are both performed automatically, the system can quickly process a large amount of image data, quickly generate an image frame sequence including loop candidate image frames, and promptly provide input for the generation of the trajectory point pose information sequence, reduce manual intervention in image labeling and human participation in labeling, and improve the response speed of the system when generating the trajectory point pose information sequence.

[0095] In some embodiments, the specific implementation of steps 203-205 and the resulting technical effects can be referred to Figure 1 The corresponding steps 102-104 in the embodiments are not described in detail here.

[0096] like Figure 3 As shown, Figure 3 In (a), the actual trajectory can be the motion trajectory of the augmented reality display device in three-dimensional space, represented by the orange trajectory line. In (b), the orange trajectory line can be the actual trajectory, and the blue trajectory line can be the trajectory corresponding to the initial trajectory point pose information sequence. In (c), the orange trajectory line can be the actual trajectory, and the green trajectory line can be the trajectory corresponding to the corrected trajectory point pose information sequence.

[0097] Reference below Figure 4 , which shows a hardware structure diagram of an augmented reality display device 400 with display function.

[0098] like Figure 4As shown, the augmented reality display device 400 includes a processing device (CPU) 401, a memory (ROM) 402, an input unit 403, and an output unit 404, wherein the processing device 401, the memory 402, the input unit 403, and the output unit 404 are interconnected via a bus 405. Here, the methods according to some embodiments of the present disclosure can be implemented as a computer program and stored in the memory 402. The processing device 401 in the augmented reality display device 400 specifically implements the trajectory point pose information sequence generation function defined in the methods of some embodiments of the present disclosure by calling the computer program stored in the memory 402. In some implementations, the input unit 403 may include a camera, a microphone, a gyroscope, an accelerometer, a magnetometer, or other devices, and the output unit 404 may be a display module or other device that can be used to display the trajectory corresponding to the trajectory point pose information sequence. The display module may include an optical machine and optical elements. The optical elements may include prisms, free-form surfaces, bird baths, optical waveguides, and other optical elements. Therefore, when the processing device 401 calls the above-mentioned computer program to execute the trajectory point posture information sequence generation function, it can control the input unit 403 to collect the image frame sequence and inertial measurement data sequence corresponding to the above-mentioned preset scene, and control the output unit 404 to display the trajectory corresponding to the above-mentioned trajectory point posture information sequence.

[0099] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0100] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0101] The computer-readable medium may be included in the augmented reality display device, or may exist independently and not be incorporated into the augmented reality display device. The computer-readable medium carries one or more programs. When executed by the augmented reality display device, the augmented reality display device: for a preset scene containing three-dimensional feature points, collects an image frame sequence and an inertial measurement data sequence corresponding to the preset scene; generates an initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence; performs a sliding window loop optimization process on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence; and updates the initial trajectory point pose information sequence based on the at least one optimized trajectory point pose information to obtain a trajectory point pose information sequence.

[0102] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0104] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0105] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for generating a trajectory point pose information sequence, comprising: For a preset scene containing three-dimensional feature points, collecting an image frame sequence and an inertial measurement data sequence corresponding to the preset scene; Generating an initial trajectory point pose information sequence based on the image frame sequence and the inertial measurement data sequence; Performing a sliding window loop closure optimization process on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence; Based on the at least one optimized trajectory point pose information, the initial trajectory point pose information sequence is updated to obtain a trajectory point pose information sequence.

2. The method according to claim 1, wherein After updating the initial trajectory point pose information sequence based on the at least one optimized trajectory point pose information to obtain the trajectory point pose information sequence, the method further includes: Display the trajectory corresponding to the trajectory point pose information sequence.

3. The method according to claim 1, wherein The collecting of the image frame sequence and the inertial measurement data sequence corresponding to the preset scene includes: capturing, by a camera device, image frames corresponding to the preset scene at a preset capturing frequency, and in response to determining that the captured image frames satisfy a preset labeling condition, marking the captured image frames as loop closure candidate image frames, thereby obtaining an image frame sequence including at least one loop closure candidate image frame; An inertial measurement data sequence is collected by the inertial measurement assembly, wherein each image frame in the image frame sequence corresponds to an inertial measurement data in the inertial measurement data sequence.

4. The method according to claim 1, wherein Each optimized trajectory point pose information in the at least one optimized trajectory point pose information corresponds to an image frame in the image frame sequence, each initial trajectory point pose information in the initial trajectory point pose information sequence corresponds to an image frame in the image frame sequence, and the updating processing of the initial trajectory point pose information sequence based on the at least one optimized trajectory point pose information to obtain the trajectory point pose information sequence includes: For each optimized trajectory point pose information in the at least one optimized trajectory point pose information, perform the following steps: Determining an image frame in the image frame sequence corresponding to the optimized trajectory point pose information as a reference image frame; Replacing the optimized trajectory point pose information with the initial trajectory point pose information corresponding to the reference image frame in the initial trajectory point pose information sequence to update the initial trajectory point pose information sequence; The updated initial trajectory point pose information sequence is determined as the trajectory point pose information sequence.

5. The method according to claim 1, wherein The performing a sliding window loop closure optimization process on the image frame sequence to generate at least one optimized trajectory point pose information corresponding to at least one image frame in the image frame sequence includes: Dividing the image frame sequence in a sliding window manner to obtain an initial image frame queue sequence, wherein a first initial image frame queue in the initial image frame queue sequence includes at least one loop candidate image frame; For each initial image frame queue in the sequence of initial image frame queues, performing the following steps: Determining the last initial image frame in the initial image frame queue as a reference image frame; determining at least one loop closure candidate image frame located before the reference image frame in the image frame sequence as a loop closure candidate image frame sequence; In response to determining that a loop closure candidate image frame in a loop closure candidate image frame sequence exists in the initial image frame queue, based on the initial image frame queue, optimized trajectory point pose information corresponding to the last initial image frame in the initial image frame queue is generated.

6. The method according to claim 5, wherein: After determining at least one loop closure candidate image frame located before the reference image frame in the image frame sequence as a loop closure candidate image frame sequence, the method further includes: In response to determining that the initial image frame queue does not contain a loop closure candidate image frame in the loop closure candidate image frame sequence, performing the following steps: Replacing the loop closure candidate image frame that meets a preset screening condition in the loop closure candidate image frame sequence with the first initial image frame in the initial image frame queue to update the initial image frame queue; Based on the updated initial image frame queue, optimized trajectory point pose information corresponding to the last initial image frame in the updated initial image frame queue is generated.

7. The method according to claim 6, wherein: The step of generating, based on the updated initial image frame queue, optimized trajectory point pose information corresponding to the last initial image frame in the updated initial image frame queue comprises: Determine the last initial image frame in the updated initial image frame queue as the target image frame; Determine the loop closure candidate image frame in the updated initial image frame queue as a reference loop closure candidate image frame; Performing common view detection processing on the target image frame and the reference loop closure candidate image frame to obtain common view detection information; Based on the common view detection information and the updated initial image frame queue, optimized trajectory point pose information corresponding to the target image frame is generated.

8. The method according to claim 7, wherein: The generating, based on the common view detection information and the updated initial image frame queue, optimized trajectory point pose information corresponding to the target image frame includes: In response to determining that the common view detection information indicates that the target image frame and the reference loop closure candidate image frame have common view, arranging the inertial measurement data corresponding to the initial image frames included in the initial image frame queue in the inertial measurement data sequence to obtain a target inertial measurement data sequence, and determining the reference loop closure candidate image frame as a reference loop closure image frame; generating optimization objective function information based on the target inertial measurement data sequence, the initial image frame queue, the reference loop-back image frame, and the target image frame; The objective function corresponding to the optimization objective function information is iteratively optimized to generate optimized trajectory point pose information corresponding to the target image frame.

9. The method according to claim 8, wherein The generating of optimization objective function information based on the target inertial measurement data sequence, the initial image frame queue, the reference loop-back image frame, and the target image frame includes: generating IMU constraint information based on the target inertial measurement data sequence and preset inertial measurement unit noise parameters; generating reprojection error information as visual constraint information based on the initial image frame queue, the three-dimensional coordinates of the three-dimensional feature points, and camera parameter information of the camera device; generating a loop closure error term based on the reference loop closure image frame and the target image frame; Based on the IMU constraint information, the visual constraint information and the loop error term, optimization objective function information is generated.

10. An augmented reality display device, which is an augmented reality head-mounted display device or an augmented reality handheld display device, and the augmented reality display device comprises: one or more processors; a camera device configured to capture image frames corresponding to the preset scene at a preset capture frequency, and in response to determining that the captured image frames satisfy a preset labeling condition, mark the captured image frames as loop closure candidate image frames, thereby obtaining an image frame sequence including at least one loop closure candidate image frame; An inertial measurement unit for collecting inertial measurement data sequences; a storage device for storing one or more programs; A display module is used to display the trajectory corresponding to the trajectory point posture information sequence in imaging; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Unmanned aerial vehicle scene dense reconstruction method based on VI-SLAM and depth estimation network

    CN112435325A

  • SLAM omnidirectional loopback correction method based on multi-camera panoramic vision

    CN113506342A

  • Method for estimating motion trail of engineering machinery, processor and engineering machinery

    CN117782077A

  • Quadruped robot track generation method, device and equipment and storage medium

    CN118089728A

  • Data processing method and device based on monocular vision inertial odometer, electronic equipment and storage medium

    CN118243134A