Pose estimation method based on multiple cameras and head-mounted display equipment
By using the matching points of 2D points and 3D point clouds of multiple cameras under multi-camera conditions, the problem of insufficient repositioning accuracy in traditional methods is solved, and higher repositioning accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202311626587.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-06
AI Technical Summary
The traditional PnP algorithm can only use the 2D feature points of a single camera for pose estimation under multi-camera conditions, resulting in a decrease in repositioning accuracy when the number of 2D points is insufficient, and even leads to repositioning failure or inaccuracy.
By acquiring the current frame sequence acquired by multiple cameras, find historical images similar to the environmental images in the current frame sequence, and obtain the 3D point cloud corresponding to the historical image. Then, for each environmental image, the 2D point set is matched with the 3D point cloud to obtain multiple matching point pairs. Combined with pre-calibrated inter-camera parameters, the PnP algorithm is used to estimate the pose of the current frame.
By increasing the number of matching point pairs, relocation failure or inaccuracy caused by insufficient 2D point extraction is avoided, and the accuracy and robustness of relocation are improved.
Smart Images

Figure CN120107348A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of virtual reality (VR) / augmented reality (AR) technology, and provides a multi-camera-based pose estimation method and a head-mounted display device. Background Art
[0002] In the field of VR, relocalization technology or loopback technology is crucial. Relocalization technology can estimate the current pose of the head mounted display (HMD) after it is restarted or the screen is turned off and on again, while loopback technology can eliminate the accumulated positioning error of the HMD over a long period of time. In general, in addition to loopback detection, relocalization technology and loopback technology also need to match the historical 3D point cloud and the current 2D feature points, and use the PnP (Perpective-n-Points) algorithm to calculate the relative pose of the current frame.
[0003] However, the traditional PnP algorithm only uses the 2D feature points collected by one camera for pose estimation, that is, when there are multiple cameras installed on the HMD, only the 2D feature points of one camera can be selected for pose estimation. In this way, when the number of matches between the 2D feature points of a single camera and the 3D point cloud is small, the pose cannot be calculated or the calculated pose is inaccurate, thus affecting the relocation accuracy of the HMD. Summary of the invention
[0004] The embodiments of the present application provide a multi-camera-based pose estimation method and a head-mounted display device, which are used to improve the accuracy of repositioning of the head-mounted display device.
[0005] On the one hand, an embodiment of the present application provides a multi-camera-based pose estimation method, comprising:
[0006] Acquire a current frame sequence captured by a plurality of cameras, wherein the current frame sequence includes an environment image captured synchronously by the plurality of cameras, and the plurality of cameras have different directions and positions;
[0007] Searching for a historical image similar to the environment image in the current frame sequence from the generated historical map, and obtaining a historical 3D point cloud corresponding to the historical image;
[0008] For each environment image in the current frame sequence, matching the 2D point set in the environment image with the historical 3D point cloud to obtain a matching point pair;
[0009] According to multiple matching point pairs corresponding to the same 3D point and pre-calibrated extrinsic parameters between multiple cameras, a current pose of the current frame relative to the historical image is estimated.
[0010] On the other hand, an embodiment of the present application provides a head-mounted display device, including a processor, a memory, a display, and a plurality of cameras, wherein the plurality of cameras have different orientations and positions, and the plurality of cameras, the display, the memory, and the processor are connected via a bus;
[0011] The multiple cameras are used to collect surrounding environment images;
[0012] The display is used to display a screen containing a virtual object;
[0013] The memory stores a computer program, and the processor performs the following operations according to the computer program:
[0014] Acquire a current frame sequence captured by the multiple cameras, wherein the current frame sequence includes an environment image captured synchronously by the multiple cameras;
[0015] Searching for a historical image similar to the environment image in the current frame sequence from the generated historical map, and obtaining a historical 3D point cloud corresponding to the historical image;
[0016] For each environment image in the current frame sequence, matching the 2D point set in the environment image with the historical 3D point cloud to obtain a matching point pair;
[0017] According to multiple matching point pairs corresponding to the same 3D point and pre-calibrated extrinsic parameters between multiple cameras, a current pose of the current frame relative to the historical image is estimated.
[0018] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer device to execute the steps of a multi-camera-based pose estimation method provided in an embodiment of the present application.
[0019] The beneficial effects of the multi-camera-based pose estimation method and head-mounted display device provided in the embodiments of the present application are as follows:
[0020] When multiple cameras are installed, through loop closure detection, historical images similar to environmental images in the current frame sequence collected by multiple cameras are obtained in the historical map, and for each environmental image in the current frame sequence, the 2D point set in the environmental image is matched with the 3D point cloud corresponding to the historical image to obtain matching point pairs. In this way, the same 3D point will form multiple matching point pairs with 2D points in multiple cameras, and then combined with the pre-calibrated extrinsic parameters between multiple cameras, the pose estimation of 2D point pairs in multiple cameras is realized. Compared with using only 2D points in one camera that match 3D points, the number of matching point pairs is larger, which can prevent the problem of relocation failure or inaccuracy caused by insufficient 2D point extraction, and improve the accuracy and robustness of relocation.
[0021] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0023] Figure 1A An example of a head mounted display device provided in an embodiment of the present application;
[0024] Figure 1B Another example of a head mounted display device provided in an embodiment of the present application;
[0025] Figure 2 The overall framework diagram of the multi-camera-based pose estimation method provided in the embodiment of the present application;
[0026] Figure 3 A flowchart of a multi-camera-based pose estimation method provided in an embodiment of the present application;
[0027] Figure 4 A schematic diagram of loop closure matching based on multiple cameras provided in an embodiment of the present application;
[0028] Figure 5 A flowchart of PnP pose estimation using matching point pairs of multiple cameras provided in an embodiment of the present application;
[0029] Figure 6 A flow chart of a method for calculating the current posture provided in an embodiment of the present application;
[0030] Figure 7 A loop optimization flow chart provided for an embodiment of the present application;
[0031] Figure 8 A loop optimization effect diagram provided in an embodiment of the present application;
[0032] Fig. 9 A structural diagram of a head-mounted display device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solution of the present application, rather than all of the embodiments. Based on the embodiments recorded in the application documents, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the technical solution of the present application.
[0034] At present, head-mounted display devices on the market (such as AR devices, VR devices, etc.) usually use multiple time points for simultaneous localization and mapping (SLAM). In the open source SALM algorithm, the ORB-SLAM3 algorithm uses historical 3D points and 3D points of the current frame for matching, and then uses the Iterative Closest Point (ICP) algorithm to estimate the pose of the current frame. This algorithm has nothing to do with the number of cameras, but the 3D points of the current frame used are generated by triangulating the 2D points of the current frame, so the number of 3D points of the current frame must be less than the total number of 2D points of multiple cameras. Monocular Visual-Inertial State Estimator (VINS-MONO) and Direct Sparse Odometry with LoopClosure (LDSO) both use PnP pose estimation based on a single camera.
[0035] The PnP algorithm is a method for solving the motion from 3D points to 2D points. Its purpose is to solve the position of the camera coordinate system relative to the world coordinate system. The formula is expressed as in, is the rotation matrix from the world coordinate system to the camera coordinate system. It is the rotation matrix from the world coordinate system to the camera coordinate system. It can transform the same coordinate vector in the world coordinate system into the camera coordinate system. is the translation vector from the world coordinate system to the camera coordinate system, that is, the vector from the origin of the camera coordinate system to the origin of the world coordinate system. Among them, the pose solution methods include but are not limited to direct linear transformation (DLT), P3P, and EPnP. However, the traditional PnP algorithm only uses a camera's 3D point to 2D point matching pair for pose calculation. In this way, when the number of 2D points extracted is insufficient, the success rate of pose calculation will be greatly reduced, which will eventually lead to a decrease in trajectory accuracy during loop optimization. Especially in AR applications, the virtual objects displayed after long-term use of the HMD will be offset from the initial position.
[0036] In view of this, an embodiment of the present application provides a pose estimation method based on multiple cameras. For the case where multiple cameras are installed, the method obtains historical images in the historical map that are similar to the environmental images in the current frame sequence collected by multiple cameras through loop detection, and for each environmental image in the current frame sequence, the 2D point set in the environmental image is matched with the 3D point cloud corresponding to the historical image to obtain matching point pairs. In this way, the same 3D point will form multiple matching point pairs with the 2D points in multiple cameras, and then combined with the pre-calibrated external parameters between multiple cameras, the matching results of the 2D points and 3D points in multiple cameras are used to estimate the relative pose of the current frame relative to the historical image using the PnP algorithm. Compared with using the 2D points of a single camera for PnP solution, the number of matching point pairs is increased, which can prevent the repositioning failure or inaccuracy caused by insufficient 2D point extraction, and improve the accuracy and robustness of repositioning.
[0037] like Figure 1A and Figure 1B , which is a schematic diagram of an HMD provided in an embodiment of the present application, wherein: Figure 1A For AR devices, Figure 1B As VR devices, both HMDs are equipped with multiple cameras, which are in different directions and positions, as indicated by circles in the figure. Through these cameras, the PnP algorithm is used to estimate the pose of the HMD, thereby achieving the relocation of the HMD.
[0038] The overall framework of pose estimation based on multiple cameras provided in the embodiment of the present application is as follows: Figure 2As shown in the figure, after the HMD is restarted or the screen is turned off and then turned on again, multiple cameras on the HMD synchronously collect surrounding environment images in real time, and perform loop detection in combination with the pre-built history map to find the history image most similar to the current frame. If the loop detection is successful, the 3D point cloud corresponding to the history image is loop matched with the 2D point set extracted from the environment image collected by each camera. If the loop matching is successful, the current pose of the current frame relative to the history image is estimated based on the matching point pairs corresponding to the multiple cameras, so as to perform loop optimization based on the current pose, and update the pose in the history map based on the optimized result, so as to realize accurate and robust relocalization using the 2D points of multiple cameras.
[0039] based on Figure 2 The framework shown in FIG. 1 is a flow chart of a multi-camera pose estimation method provided in an embodiment of the present application. Figure 3 As shown, it mainly includes the following steps:
[0040] S301: Acquire a current frame sequence captured by multiple cameras, wherein the current frame sequence includes environment images captured synchronously by multiple cameras.
[0041] In practical applications, multiple cameras on the HMD synchronously capture surrounding environment images in real time. Therefore, the environment images synchronously captured by multiple cameras at the current moment constitute a current frame sequence, and the number of environment images in each frame sequence is equal to the number of cameras.
[0042] Since the orientations and positions of multiple cameras on the HMD are different, the 3D points observed by each camera are different, but there are overlapping areas. Therefore, the content of the environment image collected by each camera in each frame sequence is different.
[0043] S302: Searching for a historical image similar to the environment image in the current frame sequence from the generated historical map, and obtaining a historical 3D point cloud corresponding to the historical image.
[0044] Loop detection is the similarity calculation of image features, which detects whether the current frame forms a closed loop with the historical images in the constructed historical map. Among them, HI-SLAM supports multi-camera loop detection. You can select any camera for loop detection, or select multiple cameras for loop detection together.
[0045] Taking loop closure detection with a single camera as an example, in the specific implementation, a reference image is selected from the environmental image of the current frame sequence, and the reference image is matched with multiple historical images in the constructed historical map respectively. The historical image with the largest number of matching points is taken as the historical image most similar to the reference image, and the historical 3D point cloud corresponding to the historical pose of the historical image in the historical map is obtained.
[0046] It should be noted that the reference image may be an environment image containing the richest environment features, or may be any environment image.
[0047] Taking multi-camera loop detection as an example, in the specific implementation, the environmental images in the current frame sequence are stitched into one image, so that multiple cameras are treated as one camera for processing. In this way, the stitched image is matched with multiple historical images in the constructed historical map respectively, and the historical image with the largest number of matching points is taken as the historical image most similar to the stitched image, and the historical 3D point cloud corresponding to the historical pose of the historical image in the historical map is obtained.
[0048] S303: For each environment image in the current frame sequence, the 2D point set in the environment image is matched with the historical 3D point cloud to obtain a matching point pair.
[0049] In the case of multiple cameras, the traditional PnP algorithm only uses the environment image captured by one camera for loop matching. Specifically, the historical 3D point cloud is matched with the 2D point set in the environment image captured by each camera, and then the result of the camera with the largest number of successful matches is selected for PnP calculation to estimate the position and posture of the HMD. However, since the HMD is in a constant state of motion, the field of view of each camera will change. When the collected environment image has fewer features (such as when the collected environment image contains large areas of white walls, smooth desktops, etc.), the number of matching point pairs obtained by a single camera may be insufficient, resulting in relocalization failure or inaccuracy.
[0050] In one example, feature points are extracted from each environment image in the current frame sequence, and the 2D point sets extracted from each environment image are brute-force matched with the historical 3D point clouds to obtain matching point pairs, thereby increasing the number of matching point pairs and improving the accuracy and robustness of relocalization.
[0051] Since the fields of view of multiple cameras are stored in overlapping areas, the same 3D point in the historical 3D point cloud may form multiple matching point pairs with 2D points in the environment images captured by multiple cameras.
[0052] Take two cameras as an example, Figure 4 As shown, the loop matching process provided in an embodiment of the present application matches the 2D point sets extracted from the first environment image captured by the first camera and the second environment image captured by the second camera with the historical 3D point cloud, respectively, wherein the two 3D points in the historical 3D point cloud are both observed by the first camera and the second camera, that is, the two 3D points in the historical 3D point cloud have matching 2D points in the first environment image and the second environment image.
[0053] S304: Estimate the current pose of the current frame relative to the historical image based on multiple matching point pairs corresponding to the same 3D point and pre-calibrated extrinsic parameters between multiple cameras.
[0054] In the embodiment of the present application, the PnP algorithm is used to directly use the matching point pairs consisting of the historical 3D point cloud and the 2D point sets collected by multiple cameras to calculate the current pose of the current frame relative to the historical image. For the specific process, see Figure 5 , mainly includes the following steps:
[0055] S3041: For multiple matching point pairs corresponding to the same 3D point, normalize the image coordinates of the 2D points in the matching point pairs to obtain a scale coefficient of the image coordinates.
[0056] Assume that a 3D point in the historical 3D point cloud has a world coordinate P in the world coordinate system W =[x W ,y W , z W ]The image coordinates of the 2D point matched by the 3D point in the environment images captured by multiple cameras are Wherein, i=0, 1, 2, ..., is the number of matching point pairs corresponding to the 3D point. The image coordinates of each 2D point are normalized to obtain the scale coefficient s between the world coordinate system and the camera coordinate system.
[0057] S3042: Select a reference camera from multiple cameras, and establish a linear equation for solving the current posture based on the rotation matrix and translation vector between other cameras in the multiple cameras and the reference camera that are pre-calibrated, and the obtained scale coefficients.
[0058] In the case of multiple cameras, the intrinsic parameters of each camera and the extrinsic parameters between the cameras can be obtained through calibration, where the extrinsic parameters include the rotation matrix and the translation vector. Through the extrinsic parameters between the cameras, the image coordinates of multiple cameras can be unified into the same coordinate system, and the scale coefficients corresponding to each camera can be used to establish the linear equation for solving the current posture.
[0059] Taking two cameras as an example, assuming that the first camera is the reference camera, the relative pose (i.e., external parameters) of the second camera to the first camera is obtained by calibration: P W The normalized coordinates of the matched 2D points in the first environment image captured by the first camera are [u 0 , v 0 , 1], the corresponding scale factor is s 0 , the normalized coordinates of the matched 2D points in the first environment image captured by the first camera are [u 1 , v 1 , 1], the corresponding scale factor is s 1, then:
[0060]
[0061]
[0062] in, is a 3*4 matrix, which is the position of the reference camera in the world coordinate system. The following relationship is established through formula 1 and formula 2:
[0063]
[0064]
[0065] Among them, Formula 3 and Formula 4 can be transformed into linear equations:
[0066] Ax=b Formula 5
[0067] Where x = [r 00 r 01 r 02 r 10 r 11 r 12 r 30 r 31 r 32 t x t y t z ].
[0068] S3043: According to the linear equation, the world coordinates of the 3D point and the image coordinates of the 2D point in the multiple matching point pairs corresponding to the same 3D point are used to solve the rotation matrix and translation vector of the reference camera in the world coordinate system to obtain the current posture.
[0069] The world coordinates of the 3D point and the image coordinates of the 2D point in the multiple matching point pairs corresponding to the same 3D point are taken as known quantities. After substituting them into Formula 5, we can solve The 12 parameters in get a 3*4 matrix. The first three columns contain the rotation information, and the fourth column is the translation vector.
[0070] In one example, the current posture can be solved by using the bundle adjustment method (BA). Figure 6 , mainly includes the following steps:
[0071] S3043_1: Select some candidate matching point pairs from multiple matching point pairs corresponding to the same 3D point, and substitute the world coordinates of the 3D points and the image coordinates of the 2D points in the candidate matching point pairs into the linear equation to obtain the initial pose.
[0072] Since an initial value is required when estimating the pose using the BA algorithm, since Formula 5 contains 12 unknown parameters, 6 matching point pairs are required to calculate the initial pose.
[0073] S3043_2: For each other matching point pair in the multiple matching point pairs, according to the initial posture, the world coordinates of the 3D point in the other matching point pair are projected into the environment image corresponding to the matching 2D point to obtain the projection coordinates, and the coordinate error between the projection coordinates and the image coordinates of the 2D point is calculated.
[0074] Theoretically, if the initial pose is accurate enough, the projection point of the 3D point in the image should completely coincide with the matching 2D point. However, due to the different orientations and positions of the cameras, the rotation of the 2D points will be inconsistent, so there will be errors in the initial pose.
[0075] S3043_3: Determine whether the coordinate error is greater than a preset error threshold, if so, execute S3043_4, if not, execute S3043_5.
[0076] When the error between the projection coordinates of the 3D point and the image coordinates of the 2D point in a matching point pair is large, it indicates that there may be a matching error, so it can be deleted to improve the accuracy of the pose calculation.
[0077] S3043_4: Delete the other matching point pairs.
[0078] S3043_5: Keep the other matching point pairs.
[0079] S3043_6: Take half of the sum of the squares of the coordinate errors of each remaining matching point pair as the reprojection error.
[0080] The squares of the coordinate errors of the remaining matching point pairs are summed and multiplied by 1 / 2, and the current pose solution is transformed into a nonlinear least squares problem, in which the current pose and 3D points are used as optimization objects.
[0081] S3043_7: Obtain the current pose by minimizing the reprojection error.
[0082] By minimizing the reprojection error, the accurate current pose can be obtained.
[0083] Among them, the first three columns of the pose matrix constitute the rotation matrix, and the fourth column is the translation vector. Therefore, the first three columns of the pose matrix must satisfy unit orthogonality. Therefore, for the pose matrix The first three columns of are subjected to singular value decomposition (SVD), and the formula is expressed as follows:
[0084] R′=U∑V TFormula 6
[0085] Among them, matrices U and V are orthogonal matrices of rotation transformation, and ∑ is a diagonal matrix of scaling transformation.
[0086] Let the rotation matrix R = UV T , obtain the orthogonal rotation matrix, and then obtain the current pose of the current frame relative to the historical image based on the 3*3 rotation matrix and the 3*1 translation vector.
[0087] In the embodiments of the present application, each camera on the HMD is fully utilized, and the 2D points in the environment image captured by each camera are matched with the historical 3D point cloud to obtain the matching point pairs corresponding to each camera. The matching point pairs corresponding to multiple cameras are used to perform PnP pose estimation using the pre-calibrated extrinsic parameters between the cameras. Compared with the 2D points captured by a single camera, the number of matching point pairs is enriched when the environmental features are insufficiently captured (such as white walls, smooth desktops, etc.), thereby improving the accuracy and success rate of pose estimation, and further improving the accuracy and robustness of repositioning after the HMD is restarted or the screen is turned off and then on.
[0088] In one example, after the current pose of the current frame relative to the historical image, repositioning can be performed through loop closure optimization. The specific loop closure optimization (i.e., repositioning) process is as follows: Figure 7 As shown, it mainly includes the following steps:
[0089] S305: Obtain the relative posture between the current frame and the adjacent frame, and obtain the relative posture between the adjacent frame and the historical image by combining the current posture of the current frame relative to the historical image.
[0090] After the HMD is restarted or the screen is turned off and then turned on again, the current pose of each frame can be calculated in real time through S301-S401. Therefore, the relative pose between frames can be obtained through the pose of the current frame and the pose of its adjacent frames. Combined with the current pose of the current frame relative to the historical image, the relative pose between the adjacent frame and the historical image can be obtained.
[0091] S306: Optimizing the posture trajectory in the historical map according to the relative posture between the adjacent frame and the historical image, and the current posture of the current frame relative to the historical image.
[0092] Wherein, N is an integer greater than or equal to 1.
[0093] like Figure 8 As shown in FIG. 1 , a comparison chart of the historical pose trajectories in the historical map before and after loop optimization is shown, where L1 is the historical pose before optimization, represented by a dotted line, and L2 is the historical pose trajectory after optimization, represented by a solid line.
[0094] In the embodiments of the present application, through loop optimization, the positioning error accumulated over a long period of time in SLAM positioning can be eliminated and the accuracy of repositioning can be improved.
[0095] Based on the same technical concept, the embodiment of the present application provides a head-mounted display device, which can be Figure 1A AR devices in the Figure 1B The VR device in the head-mounted display device can implement the steps of the above-mentioned multi-camera based pose estimation method and can achieve the same technical effect, which will not be repeated here.
[0096] See also Fig. 9 The head mounted display device includes a processor 901, a memory 902, a display 903 and a plurality of cameras 904, wherein the plurality of cameras 904 have different orientations and positions, and the plurality of cameras 904, the display 903, the memory 902 and the processor 901 are connected via a bus 905;
[0097] The multiple cameras 904 are used to collect surrounding environment images;
[0098] The display 903 is used to display a screen containing a virtual object;
[0099] The memory 902 stores a computer program, and the processor 901 performs the following operations according to the computer program:
[0100] Acquire a current frame sequence captured by the multiple cameras, wherein the current frame sequence includes an environment image captured synchronously by the multiple cameras;
[0101] Searching for a historical image similar to the environment image in the current frame sequence from the generated historical map, and obtaining a historical 3D point cloud corresponding to the historical image;
[0102] For each environment image in the current frame sequence, matching the 2D point set in the environment image with the historical 3D point cloud to obtain a matching point pair;
[0103] According to multiple matching point pairs corresponding to the same 3D point and pre-calibrated extrinsic parameters between multiple cameras, a current pose of the current frame relative to the historical image is estimated.
[0104] Optionally, the processor 901 estimates the current pose of the current frame relative to the historical image according to multiple matching point pairs corresponding to the same 3D point and pre-calibrated extrinsic parameters between multiple cameras, and the specific operation is:
[0105] For multiple matching point pairs corresponding to the same 3D point, normalize the image coordinates of the 2D points in the matching point pairs to obtain a scale coefficient of the image coordinates;
[0106] Selecting a reference camera from the multiple cameras, and establishing a linear equation for solving the current pose according to pre-calibrated rotation matrices and translation vectors between other cameras in the multiple cameras and the reference camera, and the obtained scale coefficients;
[0107] According to the linear equation, the world coordinates of the 3D point and the image coordinates of the 2D point in multiple matching point pairs corresponding to the same 3D point are used to solve the rotation matrix and translation vector of the reference camera in the world coordinate system to obtain the current posture.
[0108] Optionally, the processor 901 solves the rotation matrix and translation vector of the reference camera in the world coordinate system according to the linear equation and uses the world coordinates of the 3D point and the image coordinates of the 2D point in multiple matching point pairs corresponding to the same 3D point to obtain the current posture. The specific operation is:
[0109] Selecting some candidate matching point pairs from multiple matching point pairs corresponding to the same 3D point, and substituting the world coordinates of the 3D points and the image coordinates of the 2D points in the candidate matching point pairs into the linear equation to obtain an initial pose;
[0110] For each other matching point pair in the multiple matching point pairs, projecting the world coordinates of the 3D point in the other matching point pair into the environment image corresponding to the matching 2D point according to the initial pose to obtain the projection coordinates, and calculating the coordinate error between the projection coordinates and the image coordinates of the 2D point;
[0111] If the coordinate error is greater than the preset error threshold, deleting the other matching point pairs;
[0112] Take half of the sum of the squares of the coordinate errors of the remaining matching point pairs as the reprojection error;
[0113] The current pose is obtained by minimizing the reprojection error.
[0114] Optionally, after obtaining the current relative pose of the current frame relative to the historical image, the processor 901 further executes:
[0115] Acquire the relative posture between the current frame and the adjacent frame, and obtain the relative posture between the adjacent frame and the historical image by combining the current posture of the current frame relative to the historical image;
[0116] The posture trajectory in the historical map is optimized according to the relative posture between the adjacent frame and the historical image, and the current posture of the current frame relative to the historical image.
[0117] Optionally, each posture is a 3*4 matrix, the first three columns of the matrix are rotation matrices, the fourth column of the matrix is a translation vector, and the rotation matrix is obtained by singular value decomposition.
[0118] It should be noted that Fig. 9 This is just an example, and provides the necessary hardware for the head mounted display device to execute the steps of a multi-camera based pose estimation method provided in an embodiment of the present application. If not shown, the head mounted device may also include a two-hand handle, an IMU, a speaker, a pickup, a power supply, etc.
[0119] In the head-mounted display device of the present application, the memory may be a volatile memory, such as a random-access memory (RAM); the memory may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, and the memory may be a combination of the above memories, but is not limited thereto. The processor may include one or more central processing units (CPUs) or a general-purpose processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof.
[0120] An embodiment of the present application also provides a computer-readable storage medium for storing some instructions, which, when executed, can complete a multi-camera-based pose estimation method in the aforementioned embodiment.
[0121] An embodiment of the present application also provides a computer program product for storing a computer program, which is used to execute a multi-camera-based pose estimation method in the aforementioned embodiment.
[0122] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0123] The present application is described with reference to the flowchart and / or block diagram of the method, device (system), and computer program product according to the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram, and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart or multiple flows and / or one box or multiple boxes in the block diagram.
[0124] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0125] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0126] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A pose estimation method based on multiple cameras, It is characterized in that include: Acquire a current frame sequence captured by a plurality of cameras, wherein the current frame sequence includes an environment image captured synchronously by the plurality of cameras, and the plurality of cameras have different directions and positions; Searching for a historical image similar to the environment image in the current frame sequence from the generated historical map, and obtaining a historical 3D point cloud corresponding to the historical image; For each environment image in the current frame sequence, matching the 2D point set in the environment image with the historical 3D point cloud to obtain a matching point pair; According to multiple matching point pairs corresponding to the same 3D point and pre-calibrated extrinsic parameters between multiple cameras, a current pose of the current frame relative to the historical image is estimated.
2. The method according to claim 1, It is characterized in that The estimating the current pose of the current frame relative to the historical image according to the multiple matching point pairs corresponding to the same 3D point and the pre-calibrated extrinsic parameters between the multiple cameras includes: For multiple matching point pairs corresponding to the same 3D point, normalize the image coordinates of the 2D points in the matching point pairs to obtain a scale coefficient of the image coordinates; Selecting a reference camera from the multiple cameras, and establishing a linear equation for solving the current pose according to pre-calibrated rotation matrices and translation vectors between other cameras in the multiple cameras and the reference camera, and the obtained scale coefficients; According to the linear equation, the world coordinates of the 3D point and the image coordinates of the 2D point in multiple matching point pairs corresponding to the same 3D point are used to solve the rotation matrix and translation vector of the reference camera in the world coordinate system to obtain the current posture.
3. The method according to claim 2, It is characterized in that The step of solving the rotation matrix and translation vector of the reference camera in the world coordinate system by using the world coordinates of the 3D point and the image coordinates of the 2D point in the multiple matching point pairs corresponding to the same 3D point according to the linear equation to obtain the current posture comprises: Selecting some candidate matching point pairs from multiple matching point pairs corresponding to the same 3D point, and substituting the world coordinates of the 3D points and the image coordinates of the 2D points in the candidate matching point pairs into the linear equation to obtain an initial pose; For each other matching point pair in the multiple matching point pairs, projecting the world coordinates of the 3D point in the other matching point pair into the environment image corresponding to the matching 2D point according to the initial pose to obtain the projection coordinates, and calculating the coordinate error between the projection coordinates and the image coordinates of the 2D point; If the coordinate error is greater than the preset error threshold, deleting the other matching point pairs; Take half of the sum of the squares of the coordinate errors of the remaining matching point pairs as the reprojection error; The current pose is obtained by minimizing the reprojection error.
4. The method according to claim 1, It is characterized in that After obtaining the current relative pose of the current frame relative to the historical image, the method further includes: Acquire the relative posture between the current frame and the adjacent frame, and obtain the relative posture between the adjacent frame and the historical image by combining the current posture of the current frame relative to the historical image; The posture trajectory in the historical map is optimized according to the relative posture between the adjacent frame and the historical image, and the current posture of the current frame relative to the historical image.
5. The method according to any one of claims 1 to 4, It is characterized in that Each posture is a 3*4 matrix, the first three columns of the matrix are rotation matrices, the fourth column of the matrix is a translation vector, and the rotation matrix is obtained by singular value decomposition.
6. A head mounted display device, It is characterized in that The device comprises a processor, a memory, a display and a plurality of cameras, wherein the plurality of cameras have different orientations and positions, and the plurality of cameras, the display, the memory and the processor are connected via a bus; The multiple cameras are used to collect surrounding environment images; The display is used to display a screen containing a virtual object; The memory stores a computer program, and the processor performs the following operations according to the computer program: Acquire a current frame sequence captured by the multiple cameras, wherein the current frame sequence includes an environment image captured synchronously by the multiple cameras; Searching for a historical image similar to the environment image in the current frame sequence from the generated historical map, and obtaining a historical 3D point cloud corresponding to the historical image; For each environment image in the current frame sequence, matching the 2D point set in the environment image with the historical 3D point cloud to obtain a matching point pair; According to multiple matching point pairs corresponding to the same 3D point and pre-calibrated extrinsic parameters between multiple cameras, a current pose of the current frame relative to the historical image is estimated.
7. The head mounted display device according to claim 6, It is characterized in that The processor estimates the current pose of the current frame relative to the historical image based on multiple matching point pairs corresponding to the same 3D point and pre-calibrated extrinsic parameters between multiple cameras. The specific operation is: For multiple matching point pairs corresponding to the same 3D point, normalize the image coordinates of the 2D points in the matching point pairs to obtain a scale coefficient of the image coordinates; Selecting a reference camera from the multiple cameras, and establishing a linear equation for solving the current pose according to pre-calibrated rotation matrices and translation vectors between other cameras in the multiple cameras and the reference camera, and the obtained scale coefficients; According to the linear equation, the world coordinates of the 3D point and the image coordinates of the 2D point in multiple matching point pairs corresponding to the same 3D point are used to solve the rotation matrix and translation vector of the reference camera in the world coordinate system to obtain the current posture.
8. The head mounted display device according to claim 7, It is characterized in that The processor solves the rotation matrix and translation vector of the reference camera in the world coordinate system according to the linear equation and uses the world coordinates of the 3D point and the image coordinates of the 2D point in the multiple matching point pairs corresponding to the same 3D point to obtain the current posture. The specific operation is: Selecting some candidate matching point pairs from multiple matching point pairs corresponding to the same 3D point, and substituting the world coordinates of the 3D points and the image coordinates of the 2D points in the candidate matching point pairs into the linear equation to obtain an initial pose; For each other matching point pair in the multiple matching point pairs, projecting the world coordinates of the 3D point in the other matching point pair into the environment image corresponding to the matching 2D point according to the initial pose to obtain the projection coordinates, and calculating the coordinate error between the projection coordinates and the image coordinates of the 2D point; If the coordinate error is greater than the preset error threshold, deleting the other matching point pairs; Take half of the sum of the squares of the coordinate errors of the remaining matching point pairs as the reprojection error; The current pose is obtained by minimizing the reprojection error.
9. The head mounted display device according to claim 6, It is characterized in that After obtaining the current relative pose of the current frame relative to the historical image, the processor further executes: Acquire the relative posture between the current frame and the adjacent frame, and obtain the relative posture between the adjacent frame and the historical image by combining the current posture of the current frame relative to the historical image; The posture trajectory in the historical map is optimized according to the relative posture between the adjacent frame and the historical image, and the current posture of the current frame relative to the historical image.
10. The head mounted display device according to any one of claims 6 to 9, It is characterized in that Each posture is a 3*4 matrix, the first three columns of the matrix are rotation matrices, the fourth column of the matrix is a translation vector, and the rotation matrix is obtained by singular value decomposition.