Virtual Reality Handheld Controller Tracking Method, Terminal, and Computer-Readable Storage Medium

By acquiring image and feature point matching at the camera device at the headset and handle ends, and using SLAM map and PNP algorithm to calculate the handle position, the problem of susceptibility to interference by electromagnetic and ultrasonic handle tracking is solved, achieving a more stable virtual reality experience.

CN114332423BActive Publication Date: 2025-07-11SHENZHEN SKYWORTH NEW WORLD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111654472.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-07-11
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Existing electromagnetic and ultrasonic handle tracking methods are susceptible to interference from complex electromagnetic signals and ultrasonic signals in the environment, resulting in large limitations in use and poor user experience.

Method used

Multi-frame images are acquired by the imaging device located at the headset, an SLAM map is established, feature points are extracted and their spatial coordinate information is calculated, and the image camera device at the handle end is acquired and matched with the feature points. The PNP algorithm is used to calculate the position information of the handle to avoid environmental interference.

Benefits of technology

Improves the robustness of handle tracking, reduces dependence on environmental factors, and provides a more stable virtual reality experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332423B_ABST
    Figure CN114332423B_ABST
Patent Text Reader

Abstract

The present invention discloses a virtual reality handle tracking method, a terminal and a computer-readable storage medium, comprising the following steps: obtaining multiple frames of first images through a first imaging device located at the headset end; establishing a SLAM map based on the multiple frames of first images; extracting first feature points in the first images; obtaining spatial coordinate information of the first feature points in the SLAM map; obtaining second images through a second imaging device located at the handle end; extracting second feature points in the second images and obtaining spatial coordinate information of the second feature points corresponding in the SLAM map; obtaining second feature plane coordinate information of the second feature points in the second images; and calculating the current handle pose information of the handle end through a PNP algorithm by combining the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points. By means of the present invention, effective tracking and positioning of the handle can be achieved to optimize the virtual reality experience of users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual reality, and particularly to a method for tracking a virtual reality handle, a terminal, and a computer-readable storage medium. Background Art

[0002] Virtual Reality (VR) is an "interactive computer simulation environment that can sense the state and behavior of a user, replace or enhance the sensory feedback information of one or more sensory systems, so that the user can obtain a feeling of being immersed in a virtual environment of the simulation environment". The characteristics of virtual reality technology are high immersion. When a user is in a virtual environment, it is like being on the spot; and when the user changes the angle, the virtual environment will also make corresponding changes.

[0003] With the development of virtual reality technology, the role of the handle is becoming more and more important. After a user wears a VR headset (head-mounted display device), the handle can be used to interact with the virtual reality scene. At present, the general methods for tracking the handle are electromagnetic tracking and ultrasonic tracking.

[0004] Among them, electromagnetic tracking is to embed an electromagnetic transmitter inside the handle, and at the same time embed an electromagnetic receiver inside the VR headset, and use the electromagnetic tracking principle to calculate the position and attitude information of the handle in three-dimensional space in real time. Ultrasonic tracking is to embed an ultrasonic transmitter inside the handle, and at the same time embed an ultrasonic receiver inside the VR headset, and use the ultrasonic tracking principle to calculate the position and attitude information of the handle in three-dimensional space in real time.

[0005] However, the electromagnetic sensor of the handle is sensitive to electromagnetic signals in the environment and is easily interfered by complex electromagnetic signals in the environment, resulting in incorrect electromagnetic tracking data of the handle by the electromagnetic sensor. For example, when the electromagnetic sensor of the handle is relatively close to the computer host, or in an environment relatively close to a stereo, a TV, a refrigerator, etc., affected by other electromagnetic signals, the tracking performance of the handle is poor, resulting in a poor virtual reality experience for the user. Therefore, the use of a handle with an electromagnetic sensor has great limitations. Similarly, the use of a handle with an ultrasonic sensor also has great limitations. Summary of the Invention

[0006] The main object of the present invention is to propose a method for tracking a virtual reality handle, a terminal, and a computer-readable storage medium, aiming to effectively track and locate the handle to optimize the virtual reality experience of the user.

[0007] To achieve the above object, the present invention proposes a method for tracking a virtual reality handle, including the following steps:

[0008] Obtain multiple first images through at least one first imaging device located at the headset end;

[0009] Build a SLAM map based on multiple frames of first images; extract first feature points in the first images; obtain the spatial coordinate information of the first feature points in the SLAM map;

[0010] Obtain second images through at least one second imaging device located at the handle end;

[0011] Extract second feature points in the second images; match the second feature points with the first feature points to obtain the spatial coordinate information corresponding to the second feature points in the SLAM map; obtain the second feature plane coordinate information of the second feature points in the second images;

[0012] Combine the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points, and calculate the current handle pose information of the handle end through the PNP algorithm.

[0013] Optionally, in the step of matching the second feature points with the first feature points, the following steps are included:

[0014] Generate corresponding first feature descriptors for the first feature points;

[0015] Generate corresponding second feature descriptors for the second feature points;

[0016] Match the second feature descriptors with the first feature descriptors.

[0017] Optionally, in the step of obtaining the spatial coordinate information of the first feature points in the SLAM map, the following steps are included:

[0018] Perform BA optimization processing on the first feature points in multiple frames of first images;

[0019] Obtain the spatial coordinate information of the first feature points after BA optimization processing in the SLAM map.

[0020] Optionally, after the step of obtaining the spatial coordinate information of the first feature points in the SLAM map, the following steps are further included:

[0021] Obtain the first feature plane coordinate information of the first feature points in the first images;

[0022] Combine the first feature plane coordinate information and the spatial coordinate information of the first feature points in the SLAM map, and calculate the current headset pose information of the headset end through the PNP algorithm.

[0023] Optionally, in the step of extracting the second feature points in the second images; matching the second feature points with the first feature points to obtain the spatial coordinate information corresponding to the second feature points in the SLAM map, the following steps are included:

[0024] Perform BA optimization processing on the second feature points in multiple frames of second images;

[0025] Match the second feature points after BA optimization processing with the first feature points to obtain the spatial coordinate information of the second feature points corresponding in the SLAM map.

[0026] Optionally, after the step of combining the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points, and calculating the current handle pose information of the handle end through the PNP algorithm, the following steps are further included:

[0027] Obtain second IMU information through a second inertial measurement unit located at the handle end;

[0028] Perform a fusion operation on the handle pose information and the second IMU information.

[0029] Optionally, in the step of extracting the first feature points in the first image and the step of extracting the second feature points in the second image, the following steps are included:

[0030] Obtain first timestamp information through a first timestamp unit located at the head-mounted display end;

[0031] Obtain second timestamp information through a second timestamp unit located at the handle end;

[0032] Obtain the first image and the second image at the same moment in the first timestamp information and the second timestamp information;

[0033] Extract the first feature points and the second feature points in the first image and the second image at the same moment.

[0034] Optionally, in the step of extracting the first feature points and the second feature points in the first image and the second image at the same moment, the following steps are included:

[0035] Match the second feature points in the second image according to the first feature points in the first image by using the optical flow tracking algorithm.

[0036] To achieve the above object, the present invention also proposes a virtual reality handle tracking device for executing the above virtual reality handle tracking method, including:

[0037] A first image acquisition module, which is used to process and obtain multiple frames of first images through a first imaging device located at the head-mounted display end;

[0038] A first feature point processing module, which is used to process and establish a SLAM map according to multiple frames of first images; extract the first feature points in the first image; obtain the spatial coordinate information of the first feature points in the SLAM map;

[0039] A second image acquisition module, which is used to process and acquire a second image through a second camera device located at the handle end;

[0040] A second feature point processing module, which is used to process and extract second feature points in the second image; obtain second feature plane coordinate information of the second feature points in the second image; match the second feature points with the first feature points to obtain spatial coordinate information corresponding to the second feature points;

[0041] A handle pose information module, which is used to process and combine the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points, and calculate the current handle pose information of the handle end through the PNP algorithm.

[0042] To achieve the above object, the present invention also provides a terminal, which includes: a processor, a memory, and a virtual reality handle tracking program stored on the memory and executable on the processor. When the virtual reality handle tracking program is executed by the processor, the steps of the above virtual reality handle tracking method are implemented.

[0043] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a virtual reality handle tracking program is stored. When the virtual reality handle tracking program is executed by a processor, the steps of the above virtual reality handle tracking method are implemented.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] Different from the electromagnetic tracking and ultrasonic tracking in the background technology, which are easily interfered by other electromagnetic signals and ultrasonic signals in the environment, resulting in greater limitations in use and poor user experience. The present invention first acquires multiple first images through a first camera device located at the head-mounted display end; then establishes a SLAM map based on the multiple first images; extracts first feature points in the first images; obtains spatial coordinate information of the first feature points in the SLAM map; then acquires a second image through a second camera device located at the handle end; subsequently extracts second feature points in the second image; matches the second feature points with the first feature points to obtain spatial coordinate information corresponding to the second feature points in the SLAM map; obtains second feature plane coordinate information of the second feature points in the second image; and finally combines the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points, and calculates the current handle pose information of the handle end through the PNP algorithm. By using the image acquisition technology of the camera device to directly detect feature points in the environment, and then through sharing the SLAM map, the current pose information of the handle is calculated, which is not easily interfered by other factors in the environment and has higher robustness. Brief Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1 Schematic diagram of the hardware structure of an embodiment of a mobile terminal;

[0048] Figure 2 Schematic flowchart of the first embodiment of the virtual reality handle tracking method of the present invention;

[0049] Figure 3 Schematic flowchart of the second embodiment of the virtual reality handle tracking method of the present invention;

[0050] Figure 4 Schematic flowchart of the third embodiment of the virtual reality handle tracking method of the present invention;

[0051] Figure 5 Schematic flowchart of the fourth embodiment of the virtual reality handle tracking method of the present invention;

[0052] Figure 6 Schematic flowchart of the fifth embodiment of the virtual reality handle tracking method of the present invention;

[0053] Figure 7 Schematic flowchart of the sixth embodiment of the virtual reality handle tracking method of the present invention. Detailed Embodiments

[0054] The following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, what is described is only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0055] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0056] In the following description, suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of description of the present invention, and have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.

[0057] The terminal can be implemented in various forms. For example, the terminal described in the present invention may include mobile terminals such as mobile phones, tablet computers, laptop computers, palmtop computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs, desktop computers, etc.

[0058] In the following description, a mobile terminal will be taken as an example for illustration. Those skilled in the art will understand that, except for components specifically for mobile purposes, the structure according to the embodiments of the present invention can also be applied to fixed-type terminals.

[0059] Please refer to Figure 1 , which is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present invention. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (audio / video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art can understand that Figure 1 the mobile terminal structure shown in

[0060] does not constitute a limitation to the mobile terminal. The mobile terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. Figure 1 The following will specifically introduce each component of the mobile terminal:

[0061] The radio frequency unit 101 can be used for receiving and transmitting information or signals during communication. Specifically, after receiving the downlink information from the base station, it is sent to the processor 110 for processing. Additionally, it sends the uplink data to the base station. Generally, the radio frequency unit 101 includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc. Moreover, the radio frequency unit 101 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), and TDD-LTE (Time Division Duplexing-Long Term Evolution), etc.

[0062] WiFi belongs to short-distance wireless transmission technology. The mobile terminal can help users send and receive emails, browse the web, and access streaming media through the WiFi module 102, which provides users with wireless broadband Internet access. Although Figure 1 the WiFi module 102 is shown, it can be understood that it is not an essential component of the mobile terminal and can be omitted entirely within the scope of not changing the essence of the invention according to needs.

[0063] The audio output unit 103 can convert the audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in call signal reception mode, call mode, recording mode, voice recognition mode, broadcast reception mode, etc. Moreover, the audio output unit 103 can also provide audio output related to specific functions executed by the mobile terminal 100 (such as call signal reception sound, message reception sound, etc.). The audio output unit 103 can include a speaker, a buzzer, etc.

[0064] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a Graphics Processing Unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes the image data of a still picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage media) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) via the microphone 1042 in operating modes such as a phone call mode, a recording mode, a voice recognition mode, etc., and can process such sound into audio data. The processed audio (voice) data can be output in a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in the case of the phone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to cancel (or suppress) the noise or interference generated during the reception and transmission of audio signals.

[0065] The mobile terminal 100 further includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0066] The display unit 106 is used to display the information input by the user or the information provided to the user. The display unit 106 may include a display panel 1061, and the display panel 1061 can be configured in the form of a Liquid Crystal Display (LCD), an Organic Light-Emitting Diode (OLED), etc.

[0067] The user input unit 107 can be used to receive input numerical or character information and generate key signal inputs related to the user settings and function control of the mobile terminal. Specifically, the user input unit 107 can include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel 1071), and drive corresponding connection devices according to a preset program. The touch panel 1071 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 110, and can receive and execute the commands sent by the processor 110. In addition, the touch panel 1071 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 can also include other input devices 1072. Specifically, the other input devices 1072 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc., and specific details are not limited here.

[0068] Further, the touch panel 1071 can cover the display panel 1061. After the touch panel 1071 detects a touch operation on or near it, it transmits the operation to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides a corresponding visual output on the display panel 1061 according to the type of touch event. Although in Figure 1 the touch panel 1071 and the display panel 1061 are implemented as two independent components to realize the input and output functions of the mobile terminal, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal, and specific details are not limited here.

[0069] The interface unit 108 serves as an interface through which at least one external device can be connected to the mobile terminal 100. For example, the external device can include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headset port, and so on. The interface unit 108 can be used to receive inputs from external devices (such as data information, power, etc.) and transmit the received inputs to one or more components within the mobile terminal 100 or can be used to transmit data between the mobile terminal 100 and external devices.

[0070] The memory 109 can be used to store software programs and various data. The memory 109 can be a computer storage medium, and the memory 109 stores the message reminder program of the present invention. The memory 109 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.), etc. In addition, the memory 109 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0071] The processor 110 is the control center of the mobile terminal, connects various parts of the entire mobile terminal through various interfaces and lines, and executes various functions of the mobile terminal and processes data by running or executing software programs and / or modules stored in the memory 109, and calling data stored in the memory 109, so as to perform overall monitoring of the mobile terminal. For example, the processor 110 executes the message reminder program in the memory 109 to implement the steps of each embodiment of the message reminder method of the present invention.

[0072] The processor 110 may include one or more processing units; optionally, the processor 110 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 110 either.

[0073] The mobile terminal 100 may further include a power supply 111 (such as a battery) for supplying power to each component. Optionally, the power supply 111 can be logically connected to the processor 110 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.

[0074] Although Figure 1 not shown, the mobile terminal 100 may further include a Bluetooth module, etc., which will not be elaborated here.

[0075] Based on the above mobile terminal hardware structure, each embodiment of the method of the present invention is proposed.

[0076] The present invention proposes a virtual reality handle tracking method. In the first embodiment of the virtual reality handle tracking method, refer to Figure 2 , including the following steps:

[0077] Step S10: Obtain multiple frames of first images through at least one first imaging device located at the headset end;

[0078] Step S20: Establish a SLAM map based on multiple frames of first images; extract first feature points in the first images, and generate corresponding first feature descriptors for the first feature points; obtain the spatial coordinate information of the first feature points in the SLAM map;

[0079] Step S30: Obtain second images through at least one second imaging device located at the handle end;

[0080] Step S40: Extract second feature points in the second images, and generate corresponding second feature descriptors for the second feature points; match the second feature descriptors with the first feature descriptors to obtain the spatial coordinate information of the second feature points corresponding in the SLAM map; obtain the second feature plane coordinate information of the second feature points in the second images;

[0081] Step S50: Combine the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points, and calculate the current handle pose information of the handle end through the PNP algorithm.

[0082] Regarding the electromagnetic sensor of the handle in the background technology, it is sensitive to electromagnetic signals in the environment and is easily interfered by complex electromagnetic signals in the environment, resulting in incorrect electromagnetic tracking data of the handle by the electromagnetic sensor. For example, when the electromagnetic sensor of the handle is relatively close to the computer host, or in an environment relatively close to speakers, TVs, refrigerators, etc., affected by other electromagnetic signals, the tracking performance of the handle is poor, leading to a poor virtual reality experience for users. Therefore, the use of the handle with an electromagnetic sensor has relatively large limitations. Similarly, the use of the handle with an ultrasonic sensor also has relatively large limitations.

[0083] To solve the above technical problems, in this embodiment, first, a first imaging device located at the headset end is used to obtain multiple frames of first images; then, a SLAM (simultaneous localization and mapping) map is established based on the multiple frames of first images; first feature points in the first images are extracted, and corresponding first feature descriptors are generated for the first feature points; spatial coordinate information of the first feature points in the SLAM map is obtained; then, a second imaging device located at the handle end is used to obtain a second image; subsequently, second feature points in the second image are extracted, and corresponding second feature descriptors are generated for the second feature points; the second feature descriptors are matched with the first feature descriptors to obtain the spatial coordinate information of the corresponding second feature points in the SLAM map; second feature plane coordinate information of the second feature points in the second image is obtained; finally, combining the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points, the current handle pose information of the handle end is calculated through the PNP algorithm. By using the imaging device image acquisition technology to directly detect feature points in the environment and then sharing the SLAM map, the current pose information of the handle is calculated, which is not easily interfered by other factors in the environment and has higher robustness.

[0084] Among them, the PNP (Perspective-n-Point) algorithm is a method for solving the motion from 3D to 2D point pairs. It describes how to estimate the pose of the imaging device when the spatial coordinates of N feature points and their projection positions are known. Generally speaking, the PNP algorithm is to calculate the pose of the imaging device when the spatial coordinate information of N feature points in the known world coordinate system and the projections (plane coordinate information) of these feature points on the image are known. Among them, the spatial coordinate information of the feature points can be determined by triangulation or the depth map of an RGB-D camera.

[0085] Among them, the information transmission between the headset end and the handle end can be carried out in a wired or wireless manner. Specifically, after the second imaging device at the handle end obtains the second image, the second image is compressed and transmitted to the headset end, and subsequent operations are performed by the host module at the headset end. In addition, after the second imaging device at the handle end obtains the second image, the extraction of the second feature points and the generation of the second feature descriptors can be performed at the handle end, and the extracted second feature points and second feature descriptors are sent to the headset end, and subsequent operations are performed by the host module at the headset end, which can reduce the bandwidth occupancy of the transmission. Furthermore, when the computing power at the handle end is sufficient, the SLAN map can also be shared between the handle end and the headset end, and the handle pose information is calculated by the handle end itself, and then the handle pose information is sent to the headset end, so as to further reduce the bandwidth occupancy of the transmission.

[0086] Further, in order to improve the accuracy and speed of SLAM map construction, multiple first camera devices can be set at the headset end. When multiple first camera devices are set, the shooting directions of the multiple first camera devices face different angles respectively. Synchronously acquiring first images by using the multiple first camera devices and sharing image information, and establishing a SLAM map through the multiple first images can help improve the accuracy and speed of SLAM map construction.

[0087] Further, in order to improve the accuracy of the handle pose information, the number of second camera devices can be set at the handle end according to the actual situation. When the detection situation is good, a single second camera device can be used. When the features detected by a single second camera device are not obvious, multiple second camera devices can be used to assist in the detection, which can optimize the problem of the handle segment failing in some scenarios where the features are not obvious. And the positions of the multiple second camera devices should be specifically distinguished, such as being set at the front, back, left, and right positions of the handle end respectively.

[0088] Further, in the step of matching the second feature points with the first feature points in step S40, the following steps are included: generating corresponding first feature descriptors for the first feature points; generating corresponding second feature descriptors for the second feature points; and matching the second feature descriptors with the first feature descriptors. With such a setting, by converting the first feature points and the second feature points into the first feature descriptors and the second feature descriptors, and then using the matching between the first feature descriptors and the second feature descriptors, since the transmission bandwidth occupancy can be reduced during transmission after being converted into feature descriptors due to the data transmission.

[0089] Further, a second embodiment of the virtual reality handle tracking method proposed based on the first embodiment, refer to Figure 3 , in step S20, the following steps are included:

[0090] Step S21: Performing BA optimization processing on the first feature points in multiple frames of first images;

[0091] Step S22: Obtaining the spatial coordinate information of the first feature points after BA optimization processing in the SLAM map.

[0092] When this embodiment is applied, first perform BA optimization processing on the first feature points in multiple frames of first images: then obtain the spatial coordinate information of the first feature points after BA optimization processing in the SLAM map. Using the BA optimization technology to optimize the first feature points to obtain more accurate spatial coordinate information of the first feature points in the SLAM map, and then obtain more accurate handle pose information.

[0093] Among them, the BA optimization processing of the first feature point is to first calculate the normalized spatial point coordinates corresponding to the plane coordinates on the first image of the previous frame according to the matched plane coordinates in the first camera device and the first and second frames, and then calculate the plane coordinates reprojected to the first image of the next frame according to the coordinates of the spatial point. The reprojected plane coordinates (estimated values) will not completely coincide with the plane coordinates (measured values) on the matched first image of the next frame. The purpose of the BA optimization processing is to establish an equation for each matched first feature point, and then combine them to form an overdetermined equation to solve the optimal spatial point coordinates of the first feature point. Among them, the so-called BA (Bundle Adjustment) refers to extracting the optimal 3D model and camera device parameters (intrinsic parameters and extrinsic parameters) from visual reconstruction. The process of several bundles of light rays reflected from each feature point, after we make the optimal adjustment (adjustment) of the camera device posture and the spatial position of the feature point, finally converging to the optical center of the camera device is referred to as BA.

[0094] Further, based on the third embodiment of the virtual reality handle tracking method proposed in the first embodiment, refer to Figure 4 , after step S20, the following steps are included:

[0095] Step S23: Acquire first feature plane coordinate information of the first feature point in the first image;

[0096] Step S24: combining the first feature plane coordinate information and the spatial coordinate information of the first feature point in the SLAM map, and calculating the current head display posture information of the head display end through the PNP algorithm.

[0097] When this embodiment is applied, the first feature plane coordinate information of the first feature point in the first image is first obtained; then the first feature plane coordinate information and the spatial coordinate information of the first feature point in the SLAM map are combined to calculate the current head display posture information of the head display end through the PNP algorithm. Considering that in some virtual reality, not only the posture of the handle end needs to be detected, but also the posture of the head display end needs to be detected, by simultaneously locating the posture information of the handle end and the head display end, a more realistic virtual experience can be achieved.

[0098] Further, based on the fourth embodiment of the virtual reality handle tracking method proposed in the first embodiment, refer to Figure 5 In step S40, the following steps are included:

[0099] Step S41: performing BA optimization processing on the second feature points in multiple frames of second images;

[0100] Step S42: Match the second feature points after BA optimization processing with the first feature points to obtain the spatial coordinate information of the second feature points corresponding in the SLAM map.

[0101] When this embodiment is applied, first perform BA optimization processing on the second feature points in multiple frames of second images; then match the second feature points after BA optimization processing with the first feature points to obtain the spatial coordinate information of the second feature points corresponding in the SLAM map. Use the BA optimization technology to optimize the second feature points to obtain more accurate second feature points, and then obtain more accurate handle pose information. Among them, the principle of BA optimization processing has been described above, so it will not be repeated here.

[0102] Further, a fifth embodiment of the virtual reality handle tracking method proposed based on the first embodiment, refer to Figure 6 , after step S50, includes the following steps:

[0103] Step S51: Obtain second IMU information through a second inertial measurement unit located at the handle end;

[0104] Step S52: Perform fusion operation processing on the handle pose information and the second IMU information.

[0105] When this embodiment is applied, first obtain second IMU information through a second inertial measurement unit located at the handle end; then perform fusion operation processing on the handle pose information and the second IMU information. Among them, the full name of IMU is inertial navigation system, and the main components in the second inertial measurement unit are gyroscopes, accelerometers, and magnetometers. Among them, the gyroscope can obtain the acceleration of each axis, the accelerometer can obtain the acceleration in the x, y, and z directions, and the magnetometer can obtain the information of the surrounding magnetic field. The main work is to fuse the data of the three sensors to obtain more accurate second IMU information. Then, more accurate handle pose information is obtained by fusing the handle pose information and the second IMU information.

[0106] Among them, the fusion methods between the handle pose information and the second IMU information include two types: loose coupling and tight coupling. Loose coupling treats the visual sensor and the IMU as two separate modules, and both modules can calculate the pose information, and then generally fuse through an EKF. Tight coupling means that the intermediate data obtained by the visual sensor and the IMU are processed through an optimization filter. Tight coupling requires adding image features to the feature vector, and finally obtaining the pose information. For this reason, the final dimension of the system state vector will also be very high, and the calculation amount is also very large.

[0107] Further, inspired by the technologies of the above embodiments, those skilled in the art can also perform the following steps after step S24:

[0108] Step S25: Obtain first IMU information through a first inertial measurement unit located at the headset end;

[0109] Step S26: Perform fusion operation processing on the headset pose information and the first IMU information.

[0110] Through the above steps, more accurate pose information of the headset can be obtained.

[0111] Further, a sixth embodiment of the virtual reality handle tracking method proposed based on the first embodiment, refer to Figure 7 , in steps S20 and step 40, the following steps are included:

[0112] Step S61: Obtain first timestamp information through a first timestamp unit located at the headset end;

[0113] Step S62: Obtain second timestamp information through a second timestamp unit located at the handle end;

[0114] Step S63: Obtain a first image and a second image at the same moment in the first timestamp information and the second timestamp information;

[0115] Step S64: Extract a first feature point and a second feature point from the first image and the second image at the same moment.

[0116] When this embodiment is applied, first obtain first timestamp information through a first timestamp unit located at the headset end and obtain second timestamp information through a second timestamp unit located at the handle end; then obtain a first image and a second image at the same moment in the first timestamp information and the second timestamp information; finally, extract a first feature point and a second feature point from the first image and the second image at the same moment. In order to improve the accuracy of the handle pose information and avoid deviation of the information shared by the headset end and the handle end caused by delay during signal transmission. In this embodiment, the first timestamp unit and the second timestamp unit are used to match the first image at the headset end and the second image at the handle end at the same moment to ensure time synchronization between the headset end and the handle end. At the same time, setting timestamps can simplify the complexity of algorithm matching for network delay.

[0117] Further, in step S64, the following steps are further included:

[0118] Step S641: According to the first feature point in the first image, use the optical flow tracking algorithm to match the second feature point in the second image.

[0119] Compared with first generating corresponding feature descriptors for feature points and using an optical flow tracking algorithm to match the feature points co-visible at the headset end and the handle end, it is more efficient and faster, and can calculate the pose information of the handle more accurately. Of course, if the optical flow tracking algorithm cannot be used for matching, the feature descriptor matching algorithm is used to find the second feature point in the SLAM map.

[0120] Among them, the principle of the optical flow tracking algorithm is to process a continuous sequence of video frames; for each video sequence, a certain object detection method is used to detect possible foreground objects; if a foreground object appears in a certain frame, its representative key feature points are found; for any two adjacent videos later, find the best position of the key feature points that appeared in the previous frame in the current frame, so as to obtain the position coordinates of the foreground object in the current frame; iterate in this way to achieve object tracking. Applying the principle of the above optical flow tracking algorithm to this step means using the first image at the headset end as the foreground object, using the first feature points in the first image as the key feature points, and then finding the position of the second feature points corresponding to the first feature points in the second image at the handle end, so as to obtain the position coordinates of the second feature points, thus achieving feature point tracking.

[0121] In addition, an embodiment of the present invention also proposes a terminal, which includes: a processor, a memory, and a virtual reality handle tracking program stored on the memory and executable on the processor. When the virtual reality handle tracking program is executed by the processor, it implements the steps of the virtual reality handle tracking method as described in the above embodiment.

[0122] In addition, an embodiment of the present invention also proposes a computer-readable storage medium, on which a virtual reality handle tracking program is stored. When the virtual reality handle tracking program is executed by a processor, it implements the steps of the virtual reality handle tracking method as described in the above embodiment.

[0123] It should be noted that the other contents of the virtual reality handle tracking method, terminal, and computer-readable storage medium disclosed in the present invention are prior art and will not be elaborated here.

[0124] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, then the directional indications are only used to explain the relative position relationship and movement conditions between components in a certain specific posture (as shown in the drawings). If this specific posture changes, the directional indications will also change accordingly.

[0125] In addition, it should be noted that in the present invention, the descriptions involving "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0126] The above are only alternative embodiments of the present invention and do not limit the patent scope of the present invention. All applications directly or indirectly applying the present invention in other related technical fields are included within the patent protection scope of the present invention.

Claims

1. A virtual reality handle tracking method, characterized in that, Including the following steps: Obtain multiple frames of first images through at least one first imaging device located at the headset end; Build a SLAM map based on the multiple frames of first images; extract first feature points in the first images; obtain the spatial coordinate information of the first feature points in the SLAM map; set multiple first imaging devices at the headset end, and when there are multiple first imaging devices, the shooting directions of the multiple first imaging devices face different angles respectively, use the multiple first imaging devices to synchronously obtain first images and perform image information sharing, and build a SLAM map through the multiple first images; Obtain second images through at least one second imaging device located at the handle end; Extract second feature points in the second images; match the second feature points with the first feature points to obtain the spatial coordinate information of the second feature points corresponding in the SLAM map; obtain the second feature plane coordinate information of the second feature points in the second images; Combine the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points, and calculate the current handle pose information of the handle end through the PNP algorithm; After the step of combining the second feature plane coordinate information and the spatial coordinate information corresponding to the second feature points and calculating the current handle pose information of the handle end through the PNP algorithm, the following steps are further included: Obtain second IMU information through a second inertial measurement unit located at the handle end; Perform fusion operation processing on the handle pose information and the second IMU information. Among them, the fusion method between the handle pose information and the second IMU information includes loose coupling and tight coupling. Loose coupling treats the visual sensor and the IMU as two separate modules, and both modules can calculate the pose information, and then perform fusion tracking through the EKF to obtain the motion pose of the intermediate data; tight coupling means that the intermediate data obtained by the visual sensor and the IMU is processed through an optimization filter. Tight coupling requires adding image features to the feature vector, and finally obtaining the pose information; In the step of extracting the first feature points in the first images and the step of extracting the second feature points in the second images, the following steps are included: Obtain first timestamp information through a first timestamp unit located at the headset end; Obtain second timestamp information through a second timestamp unit located at the handle end; Obtain the first images and the second images at the same moment in the first timestamp information and the second timestamp information; Extract the first feature points and the second feature points in the first images and the second images at the same moment.

2. The virtual reality handle tracking method according to claim 1, characterized in that: In the step of matching the second feature points with the first feature points, the following steps are included: Generate corresponding first feature descriptors for the first feature points; Generate corresponding second feature descriptors for the second feature points; Match the second feature descriptors with the first feature descriptors.

3. The virtual reality handle tracking method according to claim 1, characterized in that: In the step of obtaining the spatial coordinate information of the first feature points in the SLAM map, the following steps are included: Perform BA optimization processing on the first feature points in multiple frames of first images; Obtain the spatial coordinate information of the first feature points in the SLAM map after BA optimization processing.

4. The virtual reality handle tracking method according to claim 1, characterized in that: After the step of obtaining the spatial coordinate information of the first feature point in the SLAM map, the following steps are further included: Obtain the first feature plane coordinate information of the first feature point in the first image; Combine the first feature plane coordinate information and the spatial coordinate information of the first feature point in the SLAM map, and calculate the current headset pose information of the headset end through the PNP algorithm.

5. The virtual reality handle tracking method according to claim 1, wherein: In the step of extracting the second feature point in the second image; matching the second feature point with the first feature point to obtain the spatial coordinate information corresponding to the second feature point in the SLAM map, the following steps are included: Perform BA optimization processing on the second feature points in multiple frames of second images; Match the second feature points after BA optimization processing with the first feature points to obtain the spatial coordinate information corresponding to the second feature points in the SLAM map.

6. The virtual reality handle tracking method according to claim 1, characterized in that: In the step of extracting the first feature point and the second feature point in the first image and the second image at the same moment, the following steps are included: According to the first feature point in the first image, use the optical flow tracking algorithm to match the second feature point in the second image.

7. A terminal, characterized in that: The terminal includes: a processor, a memory, and a virtual reality handle tracking program stored on the memory and executable on the processor. When the virtual reality handle tracking program is executed by the processor, the steps of the virtual reality handle tracking method according to any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium, characterized in that: A virtual reality handle tracking program is stored on the computer-readable storage medium. When the virtual reality handle tracking program is executed by the processor, the steps of the virtual reality handle tracking method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Head-mounted display, cloud server, VR system and data processing method

    CN109375764A

  • Method and device for spatial positioning

    CN113658278A