A method of estimating a handle pose and a virtual display device

CN116430986BActive Publication Date: 2026-08-07HISENSE ELECTRONICS TECH SHENZHEN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HISENSE ELECTRONICS TECH SHENZHEN CO LTD
Filing Date
2022-09-27
Publication Date
2026-08-07

AI Technical Summary

Benefits of technology

[0013]本申请提供的估计手柄位姿的方法及虚拟显示设备中,手柄上安装有IMU和多个发光器,虚拟显示设备上安装有多目相机,且相机的类型与发光器类型相匹配,通过估计手柄与虚拟显示设备间的相对位姿,实现手柄对控制虚拟显示设备显示的画面的控制,完成与虚拟世界的交互。在估计手柄与虚拟显示设备间相对位姿前,从不同位置、角度采集多帧初始手柄图像,保证获取到手柄上完整数量的发光器,从而基于多帧初始手柄图像中的发光器来优化发光器的3D空间结构,提高后续相对位姿计算的准确性;进一步的,基于优化后的3D空间结构以及相机采集的首帧目标手柄图像,初始化手柄与虚拟显示设备间的相对位姿,初始化完成后,针对相机采集的非首帧目标手柄图像,根据历史目标手柄图像对应的手柄与虚拟显示设备间的相对位姿,预测当前目标手柄图像对应的手柄与虚拟显示设备间的相对位姿,再结合IMU的观测数据,实现视觉惯导对相对位姿的联合优化,从而得到平稳、准确的当前手柄与虚拟显示设备间的目标相对位姿。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116430986B_ABST
    Figure CN116430986B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of virtual reality interaction, and provides a method for estimating the pose of a handle and a virtual display device, which utilizes an IMU on the handle and a multi-view camera on the virtual display device to realize visual inertial navigation joint optimization of the relative pose between the handle and the virtual display device. Before pose estimation, the 3D space structure of each light emitter on the handle is optimized according to the labeling results of the light emitters in multiple initial handle images collected at different position angles, so that the accuracy of relative pose calculation is improved; during the pose estimation process, if initialization has not been performed, the relative pose between the handle and the virtual display device is initialized based on the optimized 3D space structure and a target handle image collected by the camera, if initialization has been performed, the relative pose between the current handle and the virtual display device is predicted, the predicted relative pose is optimized in combination with the observation data of the IMU, and thus the target relative pose between the current stable and accurate handle and the virtual display device is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual reality interaction technology, and provides a method for estimating the pose of a controller and a virtual display device. Background Technology

[0002] For virtual display devices such as Virtual Reality (VR) and Augmented Reality (AR), gamepads are typically used for regular interaction, much like the control relationship between a personal computer (PC) and a mouse.

[0003] However, achieving interaction with the virtual world via a gamepad requires obtaining the 6DOF pose between the gamepad and the virtual display device, thus enabling the gamepad to control the displayed image on the virtual display device based on the 6DOF pose. Therefore, the gamepad's pose relative to the virtual display device directly affects the user's immersive experience and has significant research value. Summary of the Invention

[0004] This application provides a method for estimating the pose of a controller and a virtual display device, which improves the accuracy of controller pose estimation relative to the virtual display device.

[0005] On one hand, this application provides a method for estimating the pose of a controller, wherein the controller is used to control the screen displayed by a virtual display device, the controller is equipped with an IMU and multiple emitters, and the virtual display device is equipped with a multi-view camera matching the type of emitters. The method includes:

[0006] For the first frame of the target handle image captured by the camera, the relative pose between the handle and the virtual display device is initialized based on the target handle image, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the handle; wherein, the 3D spatial structure is optimized based on the annotation results of each emitter in multiple initial handle images acquired from different positions and angles;

[0007] For non-first frame target handle images captured by the camera, the relative pose between the current handle and the virtual display device is predicted based on the relative pose corresponding to historical target handle images. Combined with the observation data continuously acquired by the IMU, the target relative pose between the current handle and the virtual display device is determined.

[0008] On the other hand, this application provides a virtual display device, including a processor, a memory, a display screen, a communication interface, and a multi-view camera. The display screen is used to display images, and the virtual display device communicates with a controller through the communication interface. The controller is used to control the images displayed on the display screen, and the type of the multi-view camera matches the light emission type of multiple emitters on the controller.

[0009] The communication interface, the multi-view camera, the display screen, the memory, and the processor are connected via a bus. The memory stores a computer program, and the processor performs the following operations according to the computer program:

[0010] For the first frame of the target handle image captured by the camera, the relative pose between the handle and the virtual display device is initialized based on the target handle image, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the handle; wherein, the 3D spatial structure is optimized based on the annotation results of each emitter in multiple initial handle images acquired from different positions and angles;

[0011] For non-first frame target handle images captured by the camera, the relative pose between the current handle and the virtual display device is predicted based on the relative pose corresponding to historical target handle images. Combined with the observation data continuously acquired by the IMU, the target relative pose between the current handle and the virtual display device is determined.

[0012] On the other hand, this application provides a computer-readable storage medium storing computer-executable instructions for causing a computer device to perform the method for estimating handle pose provided in the embodiments of this application.

[0013] The method for estimating the pose of the controller and the virtual display device provided in this application include an IMU and multiple emitters installed on the controller, and a multi-view camera installed on the virtual display device, with the camera type matching the emitter type. By estimating the relative pose between the controller and the virtual display device, the controller can control the screen displayed on the virtual display device, thus completing the interaction with the virtual world. Before estimating the relative pose between the controller and the virtual display device, multiple frames of initial controller images are acquired from different positions and angles to ensure that all emitters on the controller are obtained. Based on these initial frames, the 3D spatial structure of the emitters is optimized, improving the accuracy of subsequent relative pose calculations. Furthermore, based on the optimized 3D spatial structure and the first frame of the target controller image acquired by the camera, the relative pose between the controller and the virtual display device is initialized. After initialization, for non-first frame target controller images acquired by the camera, the relative pose between the controller and the virtual display device corresponding to the historical target controller images is used to predict the relative pose between the controller and the virtual display device corresponding to the current target controller image. Combined with IMU observation data, joint optimization of the relative pose by visual-inertial navigation is achieved, resulting in a stable and accurate target relative pose between the current controller and the virtual display device. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic diagram illustrating the application scenario of the VR device and controller provided in the embodiments of this application;

[0016] Figure 2A A schematic diagram of a virtual display device including a multi-view camera, provided for an embodiment of this application;

[0017] Figure 2B A schematic diagram of a 6DOF handle containing multiple LED white lights provided for an embodiment of this application;

[0018] Figure 2C A schematic diagram of a 6DOF handle containing multiple LED infrared lights is provided for an embodiment of this application;

[0019] Figure 3 This is an overall architecture diagram of the handle pose estimation method provided in the embodiments of this application;

[0020] Figure 4 A flowchart illustrating the method for first optimizing the 3D spatial structure of each light emitter on the handle, as provided in this application embodiment;

[0021] Figure 5A The handle image captured by the binocular infrared camera before annotation is provided in the embodiments of this application;

[0022] Figure 5B The labeled binocular infrared camera captures the handle image in the embodiment of this application;

[0023] Figure 6 This is a schematic diagram of the PnP principle provided in the embodiments of this application;

[0024] Figure 7 A flowchart illustrating the method for second optimization of the 3D spatial structure of each light emitter on the handle, as provided in this application embodiment;

[0025] Figure 8 An architecture diagram for the joint visual-inertial navigation optimization estimation of handle pose provided in an embodiment of this application;

[0026] Figure 9 A flowchart of a method for jointly estimating handle pose using visual-inertial navigation provided in an embodiment of this application;

[0027] Figure 10 A flowchart illustrating a method for initializing the relative pose between a controller and a virtual display device, provided in an embodiment of this application.

[0028] Figure 11 A flowchart illustrating a method for real-time estimation of the relative pose between a controller and a virtual display device, provided in an embodiment of this application.

[0029] Figure 12 A flowchart illustrating a method for real-time determination of the correspondence between 3D and 2D points of each judge device on the handle, as provided in an embodiment of this application.

[0030] Figure 13 This is a structural diagram of a virtual display device provided in an embodiment of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0032] AR and VR virtual display devices generally refer to head-mounted display devices (simply called head-mounted displays or helmets, such as VR glasses and AR glasses) with independent processors, possessing independent computing, input, and output functions. Virtual display devices can be connected to external controllers, allowing users to control the virtual images displayed on the device and achieve regular interaction.

[0033] For example, in a game scenario, see Figure 1 This is a schematic diagram illustrating an application scenario of the virtual display device and controller provided in an embodiment of this application. Figure 1 In the game scenario shown, players interact with the virtual world using a gamepad. By utilizing the relative position of the gamepad and the virtual display device, they control the game screen and react physically to changes in the game environment, thus experiencing an immersive and engaging experience, enhancing the game's enjoyment. In particular, leveraging the large screen of a television to project the virtual game screen onto the TV further amplifies the entertainment value.

[0034] Generally, depending on the output pose, commonly used game controllers include 3DOF controllers and 6DOF controllers. 3DOF controllers output 3D rotational pose, while 6DOF controllers output 3D translation position and 3D rotational pose. Compared to 3DOF controllers, 6DOF controllers can perform more complex game actions and are more fun.

[0035] Currently, commonly used 6DOF controllers are equipped with multiple emitters (such as LEDs), which can emit different types of light (such as infrared light, white light, etc.), and the virtual display device has multiple cameras (in... Figure 2A The type (circled in the middle) should be compatible with the type of light emission.

[0036] For example, see Figure 2B This is a schematic diagram of a 6DOF handle provided in an embodiment of this application, as shown below. Figure 2B As shown, the LEDs on the 6DOF controller emit white light, and the white dots represent the positions of each LED. Therefore, to estimate the pose between the controller and the virtual display device using the positions of the LEDs on the controller, the multi-view camera on the virtual display device should be an RGB camera.

[0037] For example, see Figure 2C This is a schematic diagram of another 6DOF handle provided in an embodiment of this application, as shown below. Figure 2C As shown, the LED lights on the 6DOF controller emit infrared light (invisible to the human eye). In order to estimate the pose between the controller and the virtual display device by the position of the LED lights on the controller, the multi-view camera on the virtual display device should be an infrared camera.

[0038] In practical applications, the prerequisite for using a gamepad to interact with the virtual world is to obtain the gamepad's pose in the virtual world, so as to control the virtual display device's screen based on the 6DOF pose.

[0039] Currently, in most products on the market, the method for locating the handheld's pose is as follows: using an infrared camera on a virtual display device to capture infrared images of the emitters on the handheld, and then performing operations such as image recognition, image tracking of these infrared emitters, and matching and calculating 3D coordinates of the emitters in combination with the 3D spatial structure of the emitters on the handheld to finally obtain the relative pose between the handheld and the virtual display device.

[0040] However, in the above method, the 3D spatial structure of the emitter on the handle is measured based on the handle's design drawings, resulting in low accuracy and a large pose estimation error. At the same time, the pose of the handle in the current frame can be calculated by using the 3D spatial structure of the emitter on the handle and the corresponding 2D projection points of the emitter. However, on the one hand, the number of emitters in a single frame image captured by the camera is limited, resulting in low pose estimation accuracy. On the other hand, the observations of emitters in multiple consecutive frames captured by the camera are not correlated, resulting in poor pose smoothness during interaction and affecting the visual experience.

[0041] Generally, such as Figure 2B and Figure 2C The handle shown also contains an inertial measurement unit (IMU) to measure the handle's movement speed, including acceleration and angular velocity. The handle's movement speed also affects the relative pose between the handle and the virtual display device.

[0042] In view of this, embodiments of this application provide a method for estimating the pose of a controller and a virtual display device. Based on the annotation results of the emitters in the controller images acquired by the multi-view camera of the virtual display device at different positions and angles, the 3D spatial structure of the emitters on the controller is optimized, thereby improving the accuracy of controller pose estimation. Furthermore, by using the observation data acquired by the IMU on the controller and the controller images acquired by the camera on the virtual display device, a pose estimation method with joint optimization of visual and inertial navigation is adopted to obtain a smoother and more accurate controller pose.

[0043] See Figure 3This diagram illustrates the overall architecture of the handle pose estimation method provided in this application, primarily comprising two parts: preprocessing and relative pose estimation. The preprocessing part utilizes the annotation results of each emitter in multiple frames of initial handle images acquired by a multi-view camera on the virtual display device at different positions and angles to optimize the 3D spatial structure of the emitters on the handle, obtaining more accurate 3D coordinates and thus improving the accuracy of handle pose estimation. The relative pose estimation part primarily uses the target handle image acquired by the camera and the observation data acquired by the IMU, employing a visual-inertial navigation joint optimization method to estimate the relative pose between the handle and the virtual display device in real time.

[0044] Considering that the 3D spatial structure of each emitter can be obtained from the controller's design drawings before it leaves the factory, including the position of each emitter (represented by 3D coordinates) and the second identifier (represented by a numerically encoded ID), but due to differences in manufacturing processes, the actual 3D spatial structure of each emitter may differ from the design drawings. If the 3D spatial structure of each emitter on the controller corresponding to the design drawings is used directly for pose estimation, it may cause estimation errors and affect the user's immersive experience.

[0045] Therefore, in this embodiment of the application, before estimating the relative pose between the controller and the virtual display device, the 3D spatial structure of each emitter is optimized based on multiple frames of different initial controller images. The optimization process can use controller images captured by at least two pre-calibrated cameras on the virtual display device, or it can use controller images captured by multiple pre-calibrated independent cameras. Regardless of the type of camera used, the camera type must match the emission type of the emitters on the controller.

[0046] For a detailed optimization process of the 3D spatial structure of each emitter on the handle during implementation, please refer to [link / reference needed]. Figure 4 It mainly includes the following steps:

[0047] S401: Based on the pre-marked emitters on the multiple frames of initial handle images acquired from different positions and angles, obtain the 2D coordinates and first identifier of each emitter on the corresponding initial handle image.

[0048] In the embodiments of this application, with all the emitters on the handle lit up, a multi-view camera matching the emission type of the emitters is used to capture multiple frames of initial handle images from different positions and angles, ensuring that all emitters on the handle are captured. After obtaining multiple frames of initial handle images, the position of the center point of each emitter in each frame of the initial handle image is manually pre-marked (represented by 2D coordinates), as well as the first identifier of each emitter (represented by a numerically encoded ID). The first identifier of each emitter is consistent with the 3D spatial structure of each emitter.

[0049] Taking an LED infrared light as the emitter on the controller and a binocular infrared camera on a virtual display device as an example, the initial controller image is an infrared image. Figure 5A The image shown is an infrared handle image captured by a binocular infrared camera before labeling. After manual labeling, the binocular infrared handle image is as follows: Figure 5B As shown.

[0050] Because the position and angle of the binocular infrared cameras relative to the same handle differ, the position and number of the handle's emitters vary in the synchronously acquired single-frame infrared handle images. For example, as... Figure 5A and Figure 5B As shown, the infrared image of the handle captured by one infrared camera contains five LED infrared lights with first identifiers of 2, 3, 4, 5, and 7, while the infrared image of the handle captured by another infrared camera contains eight LED infrared lights with first identifiers of 2, 3, 4, 5, 6, 7, 8, and 9.

[0051] After annotating all the initial handle images captured by the multi-view camera at different positions and angles, in S401, the 2D coordinates and first identifier of each emitter on the corresponding initial handle image can be obtained based on the annotation results of each initial handle image.

[0052] Furthermore, based on the 2D coordinates and first identifier of each emitter, the 3D coordinates of each emitter are optimized using the Structure from Motion (SFM) approach to obtain the optimized 3D spatial structure of each emitter, as detailed in S402-S404.

[0053] S402: Based on the 3D spatial structure of each emitter before optimization, obtain the 3D coordinates and second identifier of each emitter.

[0054] In S402, the 3D spatial structure of each emitter before optimization is determined by the design drawings of the handle. By measuring the design drawings of the handle, the 3D coordinates of each emitter on the handle in the 3D spatial structure before optimization, as well as the second identifier of each emitter, can be obtained.

[0055] S403: For each frame of the initial handle image, determine the relative pose between the handle and the acquisition camera based on the 2D and 3D coordinates of the emitter with the same first and second identifiers, as well as the observation data of the IMU corresponding to the frame.

[0056] In S403, for each frame of the initial handle image, the following operations are performed: based on the 2D coordinates and 3D coordinates of the emitters that are identical to the first identifier in the 2D image and the second identifier in the 3D space, the PnP (Perspective-n-Points) algorithm is used to determine the first relative pose between the handle and the acquisition camera corresponding to the frame, and the second relative pose between the handle and the acquisition camera is obtained by integrating the observation data of the IMU corresponding to the frame. The relative pose between the handle and the acquisition camera is obtained by fusing the first relative pose and the second relative pose.

[0057] The PnP algorithm refers to the algorithm for solving the problem of object motion localization based on 3D and 2D point pairs. Its principle is as follows: Figure 6 As shown, O represents the optical center of the camera. Several 3D points (such as A, B, C, D) of an object in 3D space are projected onto the image plane by the camera, resulting in corresponding 2D points (such as a, b, c, d). Given the coordinates of the 3D points and the projection relationship between the 3D points and the 2D points, the pose between the camera and the object can be estimated. In this embodiment, the projection relationship between the 3D points and the 2D points can be reflected by the first and second identifiers of the emitter.

[0058] S404: Construct the reprojection error equation, and simultaneously optimize each relative pose and 3D coordinate based on the reprojection error equation to obtain the first optimized 3D spatial structure.

[0059] Since each camera was calibrated before use, the projection parameters (also called intrinsic parameters) of each camera and the relative poses between the cameras are known. Therefore, in S404, based on the projection parameters of each camera, the relative poses between the cameras, the 3D coordinates of each emitter on the handle, and the 2D coordinates of each emitter in the initial handle image acquired by each camera, a reprojection error equation is constructed, expressed as follows:

[0060]

[0061] In Formula 1, K n This represents the projection parameters of the nth camera. Let these represent the rotation matrix and translation vector between the handle and camera 0, respectively. Let n and 0 represent the rotation matrix and translation vector respectively between the nth camera and the 0th camera. p represents the 3D coordinates of the first emitter, identified as m, on the handle. m,n This represents the 2D coordinates of the second emitter, identified as m, projected onto the initial handle image captured by the nth camera.

[0062] in, This indicates the relative pose between the handle and camera number 0. This represents the relative pose between camera n and camera 0.

[0063] Optional. Camera 0 can be the camera with the most emitters on the acquisition handle, also known as the main camera. For example, with Figure 5B For example, if the number of emitters on the handle captured by the right infrared camera is greater than the number captured by the left infrared camera, then the right infrared camera is camera number 0 (the main camera).

[0064] Furthermore, in S404, by minimizing the reprojection error, the relative pose between the handle and the acquisition camera corresponding to the initial handle image of each frame, as well as the 3D coordinates of each emitter on the handle, are simultaneously optimized to obtain the first optimized 3D spatial structure.

[0065] After the first 3D spatial structure optimization, relatively accurate 3D coordinates of each emitter can be obtained. However, there will be some drift between the origin of the optimized 3D spatial structure and the origin of the unoptimized 3D spatial structure. In some embodiments, to further improve the accuracy of the 3D coordinates of each emitter, a 3-pair similarity transformation (SIM3) method is used to align the coordinate systems of the handle before and after optimization, thereby achieving a second optimization of the 3D spatial structure of each emitter. See [link to details] for further information. Figure 7 It mainly includes the following steps:

[0066] S405: Based on the first 3D point cloud composed of each emitter on the handle corresponding to the optimized 3D spatial structure, and the second 3D point cloud composed of each emitter on the handle corresponding to the unoptimized 3D spatial structure, determine the pose transition between the first 3D point cloud and the second 3D point cloud before and after optimization.

[0067] In S405, after the first optimization of the 3D spatial structure of each emitter on the handle, the 3D points of each emitter form the first 3D point cloud. Before the first optimization, the 3D points of each emitter form the second 3D point cloud. The coordinates of the 3D points of each emitter before and after optimization are known in both the first and second 3D point clouds. By minimizing the drift error between the 3D coordinates of each emitter before and after optimization, the pose transformation between the first and second 3D point clouds is obtained. The formula for calculating the pose transformation is as follows:

[0068]

[0069] in, This represents the 3D coordinates of the emitter, identified as m, in the handle coordinate system after the first optimization. Let m represent the 3D coordinates of the emitter labeled m in the handle coordinate system before the first optimization, s represent the scale transformation coefficient of the first 3D point cloud and the second 3D point cloud, and (R, t) represent the transformation pose between the first 3D point cloud and the second 3D point cloud. Here, R represents the rotation matrix between the handle coordinate system before and after optimization, and t represents the translation vector between the handle coordinate system before and after optimization.

[0070] S406: Based on the transformed pose, redetermine the 3D coordinates of each emitter on the handle to obtain the second optimized 3D spatial structure.

[0071] In S406, based on the quasi-transformation pose between the first and second 3D point clouds of each emitter before and after the first optimization of the 3D spatial structure, the final 3D coordinates of each emitter on the handle are calculated and denoted as follows. The calculation formula is as follows:

[0072]

[0073] Based on the final 3D coordinates of each emitter, a second optimized 3D spatial structure can be obtained. By optimizing the 3D spatial structure of each emitter on the handle, the accuracy of pose estimation can be improved.

[0074] It should be noted that the controllers in the same batch are produced based on the same design drawings. Therefore, only one optimization is needed for controllers in the same batch.

[0075] It should be noted that the above method for optimizing the 3D spatial structure of each light emitter on the handle can be executed by a virtual display device or by other devices, such as laptops or desktop computers.

[0076] After optimizing the 3D spatial structure of each emitter on the controller, more accurate 3D coordinates of each emitter can be obtained. Based on the optimized 3D coordinates of each emitter, the relative pose between the controller and the virtual display device can be estimated in real time.

[0077] See Figure 8 This is an architecture diagram of the visual-inertial joint optimization estimation of handle pose provided in the embodiments of this application. Figure 8 middle, These represent the relative poses of the controller between the IMU coordinate system and the world coordinate system, the relative poses of the controller coordinate system and the world coordinate system, and the relative poses of the camera (i.e., virtual display device) coordinate system and the world coordinate system, respectively, for the j-th (j = 1, 2, ... n) frame. This indicates the relative pose between the handle coordinate system and the IMU coordinate system.

[0078] like Figure 8As shown, by using pre-integration constraints between multiple frames of observation data continuously acquired by the IMU, and reprojection constraints between the same frame of data acquired by the IMU and the camera (i.e., the timestamps of the observation data and the target handle image are the same), the visual inertial navigation system can jointly optimize the relative pose between the handle and the virtual display device.

[0079] See Figure 9 The following is a flowchart of a method for jointly estimating handle pose using visual-inertial navigation according to an embodiment of this application. The process mainly includes the following steps:

[0080] S901: Determine whether the relative pose between the controller and the virtual display device has been initialized. If not, execute S902; if yes, execute S903.

[0081] In the process of real-time estimation of the relative pose between the controller and the virtual display device, the relative pose between the controller and the virtual display device can be predicted. The prediction process requires an initial value for the relative pose between the controller and the virtual display device. Therefore, in the pose estimation process, it is first determined whether the relative pose between the controller and the virtual display device has been initialized. If it has not been initialized, the relative pose between the controller and the virtual display device is initialized. If it has been initialized, the relative pose between the controller and the virtual display device is predicted and optimized.

[0082] S902: Based on the first frame image of the target handle captured by the camera, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the handle, initialize the relative pose between the handle and the virtual display device.

[0083] During real-time estimation of the relative pose between the controller and the virtual display device, if the relative pose between the controller and the virtual display device is not initialized, the initialization operation can be performed using the first frame image of the target controller captured by the camera on the virtual display device, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the controller. The initialization operation process is as follows: Figure 10 As shown, it mainly includes the following steps:

[0084] S9021: Extract 2D points of each emitter within the global scope of the target handle image.

[0085] During initialization, since the relative pose between the controller and the virtual display device is unknown, the positions of the 3D points of each emitter on the controller in 3D space, projected onto the 2D points in the target controller image captured by the camera on the virtual display device, are also unknown. Therefore, in S9021, it is necessary to detect each emitter on the controller within the global range of the target controller image, and use the center of each detected emitter as the 2D point of each emitter in the image.

[0086] In this application, the method for detecting 2D points of the emitter in the target handle image is not limited. For example, contour extraction algorithms in image processing (such as HOG, Canny, etc.) can be used, as well as deep learning models (such as CNN, YOLO, etc.).

[0087] S9022: Using a brute-force matching method, determine the one-to-one correspondence between the 3D points of each emitter on the 3D spatial structure and the 2D points of each emitter on the target handle image.

[0088] The projection of which 3D point of each emitter in the optimized 3D spatial structure corresponds to a 2D point extracted from the target handle image is unknown; that is, the correspondence between 2D and 3D points is unknown. Therefore, in S9022, a brute-force matching method is used to match the first identifier of each 2D point extracted from the target handle image with the second identifier of each 3D point in the optimized 3D spatial structure, establishing a one-to-one correspondence between 2D points with the same first and second identifiers and 3D points.

[0089] S9023: Based on the coordinates of corresponding 2D and 3D points, and the observation data synchronously acquired by the IMU, initialize the relative pose between the handle and the virtual display device.

[0090] In S9023, based on the pixel coordinates of the 2D points of each emitter extracted from the target handle image, and the 3D coordinates of the corresponding 3D points of the emitters after 3D spatial structure optimization, the PnP algorithm is used to initialize the relative pose between the handle and the virtual display device using vision. Simultaneously, by integrating the observation data synchronously acquired by the IMU, the initialization result of the relative pose between the handle and the virtual display device using inertial navigation is obtained. Through the fusion of vision and inertial navigation, the final initialized relative pose between the handle and the virtual display device can be obtained. The observation parameters acquired by the IMU include, but are not limited to, the acceleration and angular velocity of the handle. By integrating the acceleration once, the motion velocity of the handle can be obtained.

[0091] Generally, the acquisition frequencies of the IMU and the camera may be different. The pose estimation process needs to ensure that the observation data acquired by the IMU is synchronized with the target handle image acquired by the camera. The synchronization relationship between the observation data and the target handle image can be determined based on the timestamp.

[0092] S903: For non-first frame target handle images captured by the camera, predict the current relative pose between the handle and the virtual display device based on the relative pose between the handle and the virtual display device corresponding to the historical target handle images, and determine the target relative pose between the current handle and the virtual display device by combining the observation data continuously acquired by the IMU.

[0093] In the process of real-time estimation of the relative pose between the controller and the virtual display device, when the relative pose between the controller and the virtual display device has been initialized, the relative pose between the controller and the virtual display device is predicted based on the initialization results for the non-first frame target controller image captured by the camera.

[0094] In practice, based on the relative pose between the controller and the virtual display device corresponding to the first frame target controller image, the relative pose between the controller and the virtual display device corresponding to the second frame target controller image is predicted. Then, based on the relative pose between the controller and the virtual display device corresponding to the first frame target controller image and the second frame target controller image, the relative pose between the controller and the virtual display device corresponding to the third frame target controller image is predicted, and so on.

[0095] In this embodiment of the application, during the pose estimation process, the relative pose between the controller and the virtual display device corresponding to the historical target controller image is predicted, which ensures the smoothness of the relative pose between consecutive frames of target controller images. In this way, when using the controller to control the screen displayed on the virtual display device during actual interaction, the smoothness of the virtual display screen is guaranteed, and the user's immersive experience is improved.

[0096] After obtaining the relative pose between the current controller and the virtual display device, in order to further improve the accuracy of the relative pose, in S903, the target relative pose between the current controller and the virtual display device is determined based on the predicted current relative pose and the observation data continuously collected by the IMU.

[0097] For the process of determining the target's relative pose, please refer to [link / reference]. Figure 11 It mainly includes the following steps:

[0098] S9031: Based on the relative pose between the current controller and the virtual display device, determine the one-to-one correspondence between the 3D points of each emitter on the 3D spatial structure and the 2D points of each emitter on the target controller image.

[0099] During pose estimation, the relative pose between the controller and the virtual display device is predicted. Based on this relative pose, the approximate positions of the 3D points of each emitter on the controller in 3D space can be determined and projected onto the 2D points in the target controller image captured by the camera on the virtual display device. This establishes a one-to-one correspondence between the 2D and 3D points of each emitter. For details, please refer to [link to details]. Figure 12 It mainly includes the following steps:

[0100] S9031_1: Based on the 3D coordinates of each emitter on the handle in the optimized 3D spatial structure, and the predicted relative pose between the current handle and the virtual display device, determine the local range of each emitter in the target handle image.

[0101] S9031_2: Extract 2D points of each emitter within a local area of ​​the target handle image.

[0102] S9031_3: The nearest neighbor matching method is used to determine the one-to-one correspondence between the 3D points of each emitter on the optimized 3D spatial structure and the 2D points of each emitter on the target handle image.

[0103] Since the relative pose between the current controller and the virtual display device is known, the 3D points of each emitter on the controller in the optimized 3D spatial structure can be predicted and projected onto the 2D coordinates of the current target controller image. Therefore, in S9031_3, for the 3D point of each emitter, the nearest neighbor matching method can be used to select the 2D point that is closest to the projection point among the 2D points of each emitter extracted in the target controller image as the 2D point corresponding to that 3D point.

[0104] S9032: Based on the coordinates of corresponding 3D and 2D points, and the poses of the IMU and camera when the observation data and the current target handle image are synchronized, establish the reprojection constraint equation.

[0105] In S9032, the reprojection constraint equations are as follows:

[0106]

[0107] In Formula 4, Let represent the rotation matrix and translation vector of the IMU in the world coordinate system corresponding to the j-th frame of observation data acquired by the IMU. Let represent the rotation matrix and translation vector of the camera on the virtual display device in the world coordinate system corresponding to the j-th frame of observation data acquired by the IMU. Let these represent the rotation matrix and translation vector of the IMU in the handle coordinate system, respectively. p represents the 3D coordinates of the second emitter on the handle, labeled m. m Let represent the 2D coordinates of the first emitter, identified as m, on the current target handle image, and proj(·) represent the camera's projection equation. This refers to the pose of the IMU in the world coordinate system when the IMU is synchronized with the camera. This represents the camera's pose in the world coordinate system when the IMU is synchronized with the camera. The relative pose between the IMU and the handle when the IMU is synchronized with the camera.

[0108] S9033: Establish pre-integral constraint equations based on the pose of the IMU and the motion speed of the handle corresponding to two consecutive frames of observation data.

[0109] In S9033, the pre-integration constraint equations are as follows:

[0110]

[0111] In Formula 5, This represents the translation vector of the IMU in the world coordinate system corresponding to the (j+1)th frame of observation data acquired by the IMU. Let g represent the motion velocity of the IMU in the world coordinate system corresponding to the observation data of frame j and frame j+1, respectively. This velocity can be obtained by integrating the acceleration in the observation data of frame j and frame j+1, respectively. W Let Δt represent the acceleration due to gravity, Δt represent the time interval between the j-th and (j+1)-th frames of observation data acquired by the IMU, and LOG(·) represent the logarithmic function on the Lie group (Special Orthometri, SO3) corresponding to the quaternion array. These represent the pre-integration variables of the IMU's translation vector, motion velocity, and rotation matrix, respectively.

[0112] S9034: By combining the pre-integration constraint equation and the reprojection constraint equation, the pose of the IMU, the pose of the camera, and the relative pose of the IMU and the handle corresponding to the current target handle image are solved.

[0113] The combined formula of the pre-integral constraint equation and the reprojection constraint equation is expressed as follows:

[0114]

[0115] Where j represents the number of frames of observation data acquired by the IMU, f j Represent the pre-integral constraint equation, g j This represents the reprojection constraint equation.

[0116] By solving Equation 6, the pose of the IMU corresponding to the current target handle image in the world coordinate system can be obtained. The pose of the camera (i.e., the virtual display device) in the world coordinate system and the relative pose of the IMU and the controller

[0117] S9035: Based on the relative pose of the IMU and the controller, as well as the current pose of the IMU and the camera, obtain the target relative pose between the current controller and the virtual display device.

[0118] In S9035, based on the relative pose of the IMU and the controller, and the current pose of the IMU, the pose of the controller in the world coordinate system after joint visual-inertial navigation optimization is obtained, as expressed by the following formula:

[0119]

[0120] in, This indicates the current pose of the controller in the world coordinate system. This indicates the relative pose of the IMU and the controller.

[0121] because and Both are in the same world coordinate system, which can obtain the target relative pose between the current controller and the virtual display device, thereby controlling the screen displayed by the virtual display device by operating the controller.

[0122] It should be noted that since the camera is located on the virtual display device, the camera's pose can represent the pose of the virtual display device. A virtual display device typically has multiple cameras, each acquiring data synchronously. In this embodiment, the target handle image acquired by one camera can be used for pose estimation.

[0123] The method for estimating the handheld device pose provided in this application utilizes multiple emitters from the IMU on the handheld device and a multi-view camera on the virtual display device to achieve joint optimization of the relative pose between the handheld device and the virtual display device using visual-inertial navigation. Before pose estimation, the emitters are labeled on multiple initial handheld device images acquired at different positions and angles. The 3D spatial structure of the emitters is optimized based on the labeling results, improving the accuracy of subsequent relative pose calculations. During pose estimation, the relative pose between the handheld device and the virtual display device is initialized based on the optimized 3D spatial structure and the first frame target handheld device image acquired by the camera. After initialization, for non-first frame target handheld device images acquired by the camera, the current relative pose between the handheld device and the virtual display device is predicted based on the relative poses of the handheld device and the virtual display device corresponding to historical target handheld device images. Combined with the IMU's observation data, joint optimization of the relative pose using visual-inertial navigation is achieved, resulting in a stable and accurate target relative pose between the current handheld device and the virtual display device.

[0124] Based on the same technical concept, this application provides a virtual display device that can perform the above-described method for detecting the light emitter on the handle and achieve the same technical effect.

[0125] See Figure 13 The virtual display device includes a processor 1301, a memory 1302, a display screen 1303, a communication interface 1304, and a multi-view camera 1305. The display screen 1303 is used to display images. The virtual display device communicates with a gamepad through the communication interface 1304. The gamepad is used to control the images displayed on the display screen 1303. The type of the multi-view camera 1305 matches the light emission type of multiple emitters on the gamepad.

[0126] The communication interface 1304, the multi-view camera 1305, the display screen 1303, the memory 1302, and the processor 1301 are connected via a bus 1306. The memory 1302 stores a computer program, and the processor 1301 performs the following operations according to the computer program:

[0127] For the first frame of the target handle image captured by the camera, the relative pose between the handle and the virtual display device is initialized based on the target handle image, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the handle; wherein, the 3D spatial structure is optimized based on the annotation results of each emitter in multiple initial handle images acquired from different positions and angles;

[0128] For non-first frame target handle images captured by the camera, the relative pose between the current handle and the virtual display device is predicted based on the relative pose corresponding to historical target handle images. Combined with the observation data c continuously acquired by the IMU, the target relative pose between the current handle and the virtual display device is determined.

[0129] Optionally, the processor 1301 optimizes the 3D spatial structure of each light emitter on the handle in the following manner:

[0130] Based on the pre-marked emitters on multiple frames of initial handle images acquired from different positions and angles, obtain the 2D coordinates and first identifier of each emitter on the corresponding initial handle image;

[0131] Based on the 3D spatial structure of each emitter as described before optimization, obtain the 3D coordinates and second identifier of each emitter.

[0132] For each frame of the initial handle image, the relative pose between the handle and the acquisition camera is determined based on the 2D and 3D coordinates of the emitter with the same first and second identifiers, and the observation data of the IMU corresponding to the frame.

[0133] A reprojection error equation is constructed, and the 3D coordinates of each relative pose and each emitter are simultaneously optimized based on the reprojection error equation to obtain the first optimized 3D spatial structure.

[0134] Optionally, after obtaining the first optimized 3D spatial structure, the processor 1301 further executes:

[0135] Based on the first 3D point cloud composed of each light emitter on the handle corresponding to the optimized 3D spatial structure, and the second 3D point cloud composed of each light emitter on the handle corresponding to the unoptimized 3D spatial structure, the transformation pose between the first 3D point cloud and the second 3D point cloud before and after optimization is determined.

[0136] Based on the transformed pose, the 3D coordinates of each light emitter on the handle are re-determined to obtain the second optimized 3D spatial structure.

[0137] Optionally, the reprojection error equation is:

[0138]

[0139] Among them, K n This represents the projection parameters of the nth camera. Let these represent the rotation matrix and translation vector between the handle and the 0th camera, respectively. Let these represent the rotation matrix and translation vector between the nth camera and the 0th camera, respectively. p represents the 3D coordinates of the first emitter, identified as m, on the handle. m,n The 2D coordinates of the emitter with the second identifier m projected onto the initial handle image captured by the nth camera.

[0140] Optionally, the processor 1301 initializes the relative pose between the controller and the virtual display device based on the target controller image, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the controller. Specifically, the operation is as follows:

[0141] Extract 2D points of each emitter within the global scope of the target handle image;

[0142] A brute-force matching method is used to determine the one-to-one correspondence between the 3D points of each emitter on the 3D spatial structure and the 2D points of each emitter on the target handle image;

[0143] Based on the coordinates of the 2D and 3D points that have the aforementioned correspondence, and the observation data synchronously acquired by the IMU, the relative pose between the handle and the virtual display device is initialized.

[0144] Optionally, the processor 1301 determines the target relative pose between the current controller and the virtual display device based on the predicted relative pose between the current controller and the virtual display device, and the observation data continuously collected by the IMU. The specific operation is as follows:

[0145] Based on the relative pose between the current controller and the virtual display device, determine the one-to-one correspondence between the 3D points of each emitter on the 3D spatial structure and the 2D points of each emitter on the target controller image;

[0146] Based on the coordinates of the 2D and 3D points with the aforementioned correspondence, and the poses of the IMU and the camera when the observation data and the current target handle image are synchronized, a reprojection constraint equation is established.

[0147] Based on the pose of the IMU and the motion velocity of the handle corresponding to two consecutive frames of observation data, a pre-integral constraint equation is established;

[0148] By combining the pre-integration constraint equation and the reprojection constraint equation, the pose of the IMU, the pose of the camera, and the relative pose of the IMU and the handle corresponding to the current target handle image are solved.

[0149] Based on the relative pose of the IMU and the controller, the pose of the IMU, and the pose of the camera, the current target relative pose between the controller and the virtual display device is obtained.

[0150] Optionally, the processor 1301 determines a one-to-one correspondence between the 3D points of each emitter on the 3D spatial structure and the 2D points of each emitter on the target handheld image based on the relative pose between the current handheld controller and the virtual display device. Specifically, the operation is as follows:

[0151] Based on the 3D coordinates of each emitter on the handle in the optimized 3D spatial structure, and the predicted relative pose between the handle and the virtual display device, the local range of each emitter in the target handle image is determined.

[0152] Extract 2D points of each emitter within a local area of ​​the target handle image;

[0153] The nearest neighbor matching method is used to determine the one-to-one correspondence between the 3D points of each emitter on the optimized 3D spatial structure and the 2D points of each emitter on the target handle image.

[0154] Optionally, the pre-integration constraint equation is:

[0155]

[0156] The reprojection constraint equation is:

[0157]

[0158] in, These represent the rotation matrix and translation vector of the IMU in the world coordinate system corresponding to the j-th frame of observation data acquired by the IMU. This represents the translation vector of the IMU in the world coordinate system corresponding to the (j+1)th frame of observation data acquired by the IMU. g represents the motion velocity of the IMU in the world coordinate system corresponding to the observation data of frame j and frame j+1, respectively. WLet Δt represent the acceleration due to gravity, Δt represent the time interval between the observation data of the j-th frame and the (j+1)-th frame acquired by the IMU, and LOG(·) represent the logarithmic function on the Lie group SO3 corresponding to the quaternion array. Let these represent the pre-integration variables of the translation vector, motion velocity, and rotation matrix of the IMU, respectively. Let represent the rotation matrix and translation vector of the camera on the virtual display device in the world coordinate system corresponding to the j-th frame of observation data acquired by the IMU. Let these represent the rotation matrix and translation vector of the IMU in the handle coordinate system, respectively. p represents the 3D coordinates of the light emitter marked m on the handle. m Let m represent the 2D coordinates of the emitter marked m on the handle projected onto the current target handle image, and proj(·) represent the projection equation of the camera.

[0159] Optionally, the result of combining the pre-integration constraint equation and the reprojection constraint equation is:

[0160]

[0161] in, Let f represent the rotation matrix and translation vector of the IMU in the world coordinate system corresponding to the j-th frame of observation data acquired by the IMU, where j represents the frame number of the observation data acquired by the IMU, and f represents the translation vector of the IMU in the world coordinate system. j Let g represent the pre-integral constraint equation. j This represents the reprojection constraint equation.

[0162] It should be noted that, Figure 13 This is merely an example illustrating the hardware necessary for a virtual display device to implement the method steps for estimating the handle pose provided in the embodiments of this application. Not shown, the virtual display device also includes conventional hardware such as speakers, earpieces, lenses, and power interfaces.

[0163] Examples of this application Figure 13 The processor involved can be a central processing unit (CPU), a general-purpose processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.

[0164] This application also provides a computer-readable storage medium for storing instructions that, when executed, can perform the method for estimating the handle pose described in the foregoing embodiments.

[0165] This application also provides a computer program product for storing a computer program for executing the method for estimating handle pose in the foregoing embodiments.

[0166] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0167] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0168] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0169] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0170] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for estimating the pose of a handle, characterized in that, The handle is used to control the screen displayed by the virtual display device. The handle is equipped with an IMU and multiple emitters. The virtual display device is equipped with a multi-view camera matching the type of emitters. The method includes: For the first frame of the target handle image captured by the camera, the relative pose between the handle and the virtual display device is initialized based on the target handle image, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the handle; wherein, the 3D spatial structure is optimized based on the annotation results of each emitter in multiple initial handle images acquired from different positions and angles; For non-first frame target handle images captured by the camera, the relative pose between the current handle and the virtual display device is predicted based on the relative pose corresponding to historical target handle images. Combined with the observation data continuously acquired by the IMU, the target relative pose between the current handle and the virtual display device is determined. The 3D spatial structure of each light emitter on the handle is optimized in the following ways: Based on the pre-marked emitters on multiple frames of initial handle images acquired from different positions and angles, obtain the 2D coordinates and first identifier of each emitter on the corresponding initial handle image; Based on the 3D spatial structure of each emitter as described before optimization, obtain the 3D coordinates and second identifier of each emitter. For each frame of the initial handle image, the relative pose between the handle and the acquisition camera is determined based on the 2D and 3D coordinates of the emitter with the same first and second identifiers, and the observation data of the IMU corresponding to the frame. A reprojection error equation is constructed, and the 3D coordinates of each relative pose and each emitter are simultaneously optimized based on the reprojection error equation to obtain the first optimized 3D spatial structure.

2. The method as described in claim 1, characterized in that, The methods for optimizing the 3D spatial structure of each light emitter on the handle also include: After obtaining the first optimized 3D spatial structure, the transformation pose between the first 3D point cloud and the second 3D point cloud is determined based on the first 3D point cloud composed of each light emitter on the handle corresponding to the optimized 3D spatial structure and the second 3D point cloud composed of each light emitter on the handle corresponding to the unoptimized 3D spatial structure. Based on the transformed pose, the 3D coordinates of each light emitter on the handle are re-determined to obtain the second optimized 3D spatial structure.

3. The method as described in claim 1 or 2, characterized in that, The reprojection error equation is: in, This represents the projection parameters of the nth camera. , Let these represent the rotation matrix and translation vector between the handle and the 0th camera, respectively. , Let these represent the rotation matrix and translation vector between the nth camera and the 0th camera, respectively. This indicates the 3D coordinates of the first emitter, identified as m, on the handle. The 2D coordinates of the emitter with the second identifier m projected onto the initial handle image captured by the nth camera.

4. The method as described in claim 1, characterized in that, The step of initializing the relative pose between the controller and the virtual display device based on the target controller image, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the controller includes: Extract 2D points of each emitter within the global scope of the target handle image; A brute-force matching method is used to determine the one-to-one correspondence between the 3D points of each emitter on the 3D spatial structure and the 2D points of each emitter on the target handle image; Based on the coordinates of the 2D and 3D points that have the aforementioned correspondence, and the observation data synchronously acquired by the IMU, the relative pose between the handle and the virtual display device is initialized.

5. The method as described in claim 1, characterized in that, Based on the predicted relative pose between the current controller and the virtual display device, and the observation data continuously acquired by the IMU, the target relative pose between the current controller and the virtual display device is determined, including: Based on the relative pose between the current controller and the virtual display device, determine the one-to-one correspondence between the 3D points of each emitter on the 3D spatial structure and the 2D points of each emitter on the target controller image; Based on the coordinates of the 2D and 3D points with the aforementioned correspondence, and the poses of the IMU and the camera when the observation data and the current target handle image are synchronized, a reprojection constraint equation is established. Based on the pose of the IMU and the motion velocity of the handle corresponding to two consecutive frames of observation data, a pre-integral constraint equation is established; By combining the pre-integration constraint equation and the reprojection constraint equation, the pose of the IMU, the pose of the camera, and the relative pose of the IMU and the handle corresponding to the current target handle image are solved. Based on the relative pose of the IMU and the controller, the pose of the IMU, and the pose of the camera, the current target relative pose between the controller and the virtual display device is obtained.

6. The method as described in claim 5, characterized in that, The step of determining the one-to-one correspondence between the 3D points of each emitter on the 3D spatial structure and the 2D points of each emitter on the target handheld image based on the relative pose between the current handheld device and the virtual display device includes: Based on the 3D coordinates of each emitter on the handle in the optimized 3D spatial structure, and the predicted relative pose between the handle and the virtual display device, the local range of each emitter in the target handle image is determined. Extract 2D points of each emitter within a local area of ​​the target handle image; The nearest neighbor matching method is used to determine the one-to-one correspondence between the 3D points of each emitter on the optimized 3D spatial structure and the 2D points of each emitter on the target handle image.

7. The method as described in claim 5 or 6, characterized in that, The pre-integral constraint equation is: =0 The reprojection constraint equation is: in, , These represent the rotation matrix and translation vector of the IMU in the world coordinate system corresponding to the j-th frame of observation data acquired by the IMU. This represents the translation vector of the IMU in the world coordinate system corresponding to the (j+1)th frame of observation data acquired by the IMU. , These represent the motion velocities of the IMU in the world coordinate system corresponding to the observation data of the j-th frame and the (j+1)-th frame, respectively. Represents gravitational acceleration. LOG(j) represents the time interval between the observation data of the j-th frame and the (j+1)-th frame acquired by the IMU. () represents the logarithmic function on the Lie group SO3 corresponding to the quaternion array. , , Let these represent the pre-integration variables of the translation vector, motion velocity, and rotation matrix of the IMU, respectively. , Let represent the rotation matrix and translation vector of the camera on the virtual display device in the world coordinate system corresponding to the j-th frame of observation data acquired by the IMU. , Let these represent the rotation matrix and translation vector of the IMU in the handle coordinate system, respectively. This indicates the 3D coordinates of the light emitter marked 'm' on the handle. This indicates the 2D coordinates of the emitter marked 'm' on the handle projected onto the current target handle image. This represents the projection equation of the camera.

8. The method as described in claim 7, characterized in that, The result of combining the pre-integration constraint equation and the reprojection constraint equation is: in, , Let represent the rotation matrix and translation vector of the IMU in the world coordinate system corresponding to the j-th frame of observation data acquired by the IMU, where j represents the frame number of the observation data acquired by the IMU. This represents the pre-integral constraint equation. This represents the reprojection constraint equation.

9. A virtual display device, characterized in that, It includes a processor, memory, display screen, communication interface, and multi-view camera. The display screen is used to display images. The virtual display device communicates with the controller through the communication interface. The controller is used to control the images displayed on the display screen. The type of the multi-view camera matches the light emission type of multiple emitters on the controller. The communication interface, the multi-view camera, the display screen, the memory, and the processor are connected via a bus. The memory stores a computer program, and the processor performs the following operations according to the computer program: For the first frame of the target handle image captured by the camera, the relative pose between the handle and the virtual display device is initialized based on the target handle image, the observation data synchronously acquired by the IMU, and the optimized 3D spatial structure of each emitter on the handle; wherein, the 3D spatial structure is optimized based on the annotation results of each emitter in multiple initial handle images acquired from different positions and angles; For non-first frame target handle images captured by the camera, the relative pose between the handle and the virtual display device corresponding to historical target handle images is used to predict the current relative pose between the handle and the virtual display device. Combined with the observation data continuously acquired by the IMU, the current target relative pose between the handle and the virtual display device is determined. The 3D spatial structure of each light emitter on the handle is optimized in the following ways: Based on the pre-marked emitters on multiple frames of initial handle images acquired from different positions and angles, obtain the 2D coordinates and first identifier of each emitter on the corresponding initial handle image; Based on the 3D spatial structure of each emitter as described before optimization, obtain the 3D coordinates and second identifier of each emitter. For each frame of the initial handle image, the relative pose between the handle and the acquisition camera is determined based on the 2D and 3D coordinates of the emitter with the same first and second identifiers, and the observation data of the IMU corresponding to the frame. A reprojection error equation is constructed, and the 3D coordinates of each relative pose and each emitter are simultaneously optimized based on the reprojection error equation to obtain the first optimized 3D spatial structure.

Citation Information

Patent Citations

  • Hand-held three-dimensional surface information extraction method and extractor thereof

    CN101853528A

  • Method and system for locating object, and head-mounted display equipment

    CN107153369A