Handle positioning method and device and head-mounted display equipment

The number of positioning spots is detected by multi-eye camera image flow, and combined with human-hand visual positioning algorithm and IMU data, the stability and accuracy problems when the handle positioning lamp is blocked are solved, and efficient handle positioning is achieved.

CN120014037APending Publication Date: 2025-05-16HISENSE ELECTRONICS TECH SHENZHEN CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411936984.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the handle positioning method is prone to failure when the number of positioning spots is insufficient, and because the positioning lamp is blocked, the handle cannot stabilize and high-precision positioning.

Method used

By acquiring the image stream collected by the multi-eye camera, the number of positioning spots is detected. If the threshold is reached, the positioning spot is directly used to calculate the handle positioning position; if it is insufficient, a visual positioning algorithm for the human hand is introduced, and the hand positioning positioning is converted into the handle initial positioning through the preset conversion relationship, and the second observation positioning is generated by combining the IMU rotation data.

Benefits of technology

The handle is stable and high-precision positioning when the positioning lamp is blocked, which improves the stability and accuracy of positioning, and avoids high computing power consumption in pure visual positioning technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014037A_ABST
    Figure CN120014037A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, provides a handle positioning method and device and head-mounted display equipment, and is used for stably positioning a handle at high precision. For image streams collected at intervals of long and short exposure durations, as the positioning precision of the positioning light spots is high, if the number of the positioning light spots detected in the current short exposure image reaches a preset number threshold value, the positioning light spots are used for calculating the current pose of the handle, otherwise, the pose of the handle is calculated according to the preset conversion relation between the handle and the corresponding hand. The posture of the hand in the next long-exposure image is converted into the current posture of the corresponding handle, so that the problem that the handle cannot be positioned when the positioning lamp is shielded is solved, the positioning stability of the handle is improved, and in addition, the IMU rotation data of the handle is a true value which can be normally transmitted, so that the positioning accuracy is improved. Therefore, the converted handle pose is optimized by using the IMU rotation data of the current handle, so that the pose conversion error can be reduced, and the positioning precision of the handle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] The 6-degree-of-freedom (Dof) handle is the most common interaction method in head-mounted display (HMD), and its positioning accuracy directly affects the user's interaction experience.

[0003] At present, the positioning method of the handle is mainly based on identifying the positioning light spots in the short-exposure infrared image. Due to the limitation of the positioning algorithm, 4 or more positioning light spots must be detected in an infrared image to meet the positioning requirements. When the positioning light spots are not collected enough, the handle positioning will fail. However, installing too many infrared positioning lights on the handle will cause the 2D points in the image to be too dense. The feature point matching during positioning is prone to errors and takes a long time, which has a negative effect on the positioning algorithm.

[0004] Therefore, achieving long-term, stable, high-precision positioning of the handle is an important issue that needs to be urgently solved in the interactive scenarios of head-mounted display devices. Summary of the invention

[0005] The embodiments of the present application provide a handle positioning method, device and head-mounted display device for continuously and stably performing high-precision positioning of the handle.

[0006] In a first aspect, an embodiment of the present application provides a method for positioning a handle, comprising:

[0007] Acquire an image stream of the hand grip continuously acquired by the multi-eye camera, wherein the image stream includes a first image and a second image, wherein the first image and the second image are acquired at a set interval;

[0008] Detecting the positioning light spots contained in each handle in the current first image in the image stream, and determining whether the number of the positioning light spots reaches a preset number threshold;

[0009] If reached, the first observation posture of the corresponding handle in the current first image is calculated and output according to the positioning spot detected by each handle;

[0010] If not reached, then based on the posture of the corresponding human hand of each handle in the adjacent second image of the current first image and the preset conversion relationship, the initial posture of the corresponding handle in the adjacent second image is generated, and the second observation posture of the corresponding handle in the current first image is generated using the IMU rotation data of the corresponding handle corresponding to the current first image and the initial posture.

[0011] The beneficial effect of the above technical solution is as follows: in order to realize the detection of the positioning lights while acquiring the external image, the head mounted display device usually shoots the image stream at a set short time interval with a long and short exposure time interval. Therefore, the state change of the hand holding the handle in the adjacent short exposure first image and the long exposure second image is negligible, that is, the position error of the handle between two adjacent frames is very small. Therefore, when the number of positioning lights meets the threshold, the positioning spot is directly used to calculate the handle position to achieve high-precision positioning of the handle. When the positioning lights on the handle are blocked to a number lower than the number required for positioning the handle, a visual positioning algorithm of the human hand is introduced. Through the preset conversion relationship corresponding to the human hand and the handle position, the position of the human hand in the adjacent second image with long exposure is converted to the initial position of the corresponding handle in the current first image with short exposure, thereby solving the problem that the handle cannot be positioned when the positioning lights are blocked when the short exposure image is used for positioning the handle, and improving the stability of the handle positioning. At the same time, because the visual positioning algorithm of the human hand is introduced when the positioning lights are blocked to a number lower than the number required for positioning the handle, compared with the pure visual handle positioning technology, there will not be too much computing power consumption.

[0012] In addition, since the IMU rotation data of the handle is a real value collected in real time, and the IMU rotation data of the handle can generally be transmitted normally during the interaction process, the second observed posture of the corresponding handle in the current first image of short exposure is determined according to the IMU rotation data of the handle corresponding to the current first image of short exposure and the converted initial posture, which can reduce the error in posture conversion and thus improve the positioning accuracy of the handle.

[0013] Optionally, generating the initial posture of the corresponding handle in the adjacent second image of the current first image according to the posture of the human hand corresponding to each handle in the adjacent second image and the preset conversion relationship includes:

[0014] For each handle, do the following:

[0015] According to the type of the handle, obtaining a preset conversion relationship between the handle and a corresponding human hand;

[0016] According to the preset conversion relationship, the first predicted posture of the handle in the current first image is converted into the second predicted posture of the corresponding human hand in the next second image; wherein the first predicted posture is determined according to the previous historical posture of the handle and the first IMU data segment between the image corresponding to the previous historical posture and the current first image;

[0017] Performing hand detection on the hand region in the adjacent second image according to the second predicted posture to output key points of the hand;

[0018] The posture of the corresponding human hand in the adjacent second image is determined according to the key points of the human hand, and the posture of the corresponding human hand is converted into the initial posture of the handle in the adjacent second image by using the preset conversion relationship.

[0019] The beneficial effects of the above technical solution are: in order to solve the problem of positioning failure caused by the inability to sense the sudden change of hand posture during visual positioning, since the IMU data of the handle can generally be transmitted normally during the interaction process, the handle's previous historical posture and the handle's IMU data between the corresponding image of the previous historical posture and the current first image are used to predict the handle's posture in the current first image, and then the predicted posture of the handle is applied to the hand in the adjacent image through the preset conversion relationship between the handle and the corresponding hand posture to perform local detection of the hand, thereby realizing the use of the IMU data of the handle to provide motion support for the visual positioning of the hand, ensuring that the key points of the hand can still be accurately detected in the adjacent second image when the hand movement suddenly changes, thereby improving the stability of hand tracking, and then, according to the preset conversion relationship, the hand posture is applied to the handle under the long-exposure adjacent second image, thereby ensuring a continuous and stable positioning process of the handle.

[0020] Optionally, the using the IMU rotation data of the corresponding handle corresponding to the current first image and the initial pose to generate a second observed pose of the corresponding handle under the current first image includes:

[0021] For each handle, do the following:

[0022] Generating a target pose of the handle in the current first image according to the initial pose of the handle in the adjacent second image;

[0023] Determining a first rotation parameter of the handle under the current first image according to the gravity acceleration in the current IMU data of the handle under the current first image;

[0024] The second rotation parameter in the target pose is replaced by the first rotation parameter to generate a second observed pose of the corresponding handle in the current frame image.

[0025] The beneficial effects of the above technical solution are as follows: since the first image and the second image are collected with long and short exposures at a set time interval, after the initial posture of the handle in the adjacent second image is known, the target posture of the handle in the current first image can be calculated. Since the gravity acceleration of the earth is vertically downward and is a constant, and the target posture determined according to the hand positioning result is an estimated value, there are hand positioning errors and posture conversion errors. Therefore, the accuracy of the rotation parameters calculated based on the gravity acceleration is higher than the accuracy of the rotation parameters in the target posture. In this way, after replacing the rotation parameters in the target posture with the rotation parameters determined by the gravity acceleration, the positioning accuracy of the handle can be improved.

[0026] Optionally, the detecting the positioning light spot in the current frame image of the image stream includes:

[0027] For each handle, do the following:

[0028] Acquire the previous historical posture of the handle, and the first IMU data segment between the image corresponding to the previous historical posture and the current first image;

[0029] Performing motion estimation on the first IMU data segment and the previous historical pose to output a first predicted pose of the handle under the current first image;

[0030] According to the first predicted position and posture, a handle area in the current first image is determined, and a positioning spot detection is performed on the handle area.

[0031] The beneficial effect of the above technical solution is: posture prediction is performed through the IMU data and historical posture of the handle, so as to perform positioning spot detection on the local area of ​​the current first image according to the predicted posture, thereby improving the detection efficiency of the positioning spot compared to the global detection.

[0032] Optionally, the process of determining the preset conversion relationship between the handle and the corresponding human hand includes:

[0033] Acquire multiple gripping images of the handle in motion, wherein two adjacent images of the multiple gripping images are taken as a group of image pairs, each group of image pairs includes a first gripping image and a second gripping image, the first gripping image and the second gripping image are acquired at a set time interval, and the exposure time of the first gripping image is shorter than that of the second gripping image;

[0034] For any set of image pairs among multiple sets of image pairs, execute:

[0035] identifying a first category of a handle in the first grasped image and a second category of a human hand in the second grasped image;

[0036] Matching handles and human hands that are the same as the first category and the second category to obtain a matching pair;

[0037] Calculate a first transformation relationship between the handle and the corresponding human hand according to the 6DOf pose of the human hand and the 6DOf pose of the handle in the same matching pair;

[0038] The mean of the first transformation relationship between the human hands and the handles of the same category in the multiple grasping images is calculated to obtain a preset transformation relationship between the handles and the human hands of the corresponding category; wherein the preset transformation relationship is used to bind the human hands and the handles of the same category in position and posture.

[0039] The beneficial effect of the above technical solution is: since human hands are divided into left and right categories, and handles are also divided into left and right categories, and users can correctly grasp the handles according to the instructions, a matching pair when the hand correctly grasps the handle can be obtained according to the first category of the human hand and the second category of the handle, that is, the first category of the hand and the second category of the handle are the same, and when the human hand correctly grasps the handle, the grasping posture of the human hand and the handle are basically consistent. Therefore, according to the 6DOf posture of the human hand and the handle of the category matching, the conversion relationship for binding the posture of the human hand and the handle of the same category can be calculated. On the other hand, by averaging the conversion relationships corresponding to multiple grasping images, the conversion relationship between the human hand and the handle of the same category can be smoothed, thereby avoiding the influence of the sudden change of the grasping state of the hand and the handle in the grasping image on the binding relationship, and improving the calculation accuracy of the conversion relationship.

[0040] Optionally, after obtaining matching pairs of hands and handles of the same category, the method further includes:

[0041] If the data of the human hand and the handle in the matching pair is incomplete, the matching pair is eliminated;

[0042] According to the 6Dof postures of the human hand and the handle in the matching pair, the distance between the human hand and the handle of the matching pair is calculated. If the distance is greater than a preset distance threshold, the matching pair is eliminated.

[0043] The beneficial effect of the above technical solution is: by eliminating matching pairs with incomplete data and matching pairs with a spacing greater than a preset distance threshold, it is ensured that the matching pairs used to calculate the posture transformation matrix are the correct posture of the hand holding the handle, thereby improving the accuracy of the posture transformation matrix calculation.

[0044] Optionally, the process of determining whether the data of the human hand and the handle in the matching pair are complete includes:

[0045] If the number of positioning light spots detected in the first grasping image is greater than a first preset threshold, it is determined that the data of the handle in the first grasping image is complete;

[0046] If all of the rectangular frames of the human hand in the second grasping image are located within the second grasping image, it is determined that the data of the human hand in the second grasping image is complete.

[0047] The beneficial effect of the above technical solution is: by ensuring the integrity of the data of the human hand and the handle, the 6DOf posture of the handle and the 6DOf posture of the human hand can be accurately calculated.

[0048] Optionally, the first conversion relationship includes rotation and translation, and the average of the first conversion relationships corresponding to the same category of human hands and handles in the plurality of grasping images is calculated to obtain a preset conversion relationship between the corresponding category of handles and human hands;

[0049] Set a sliding window list of preset length for each category;

[0050] For each first conversion relationship between a human hand and a handle of the same category, determining whether a rotation in the first conversion relationship is within a first value interval, and whether a translation in the first conversion relationship is within a second value interval;

[0051] If the rotation and the translation are both within the corresponding value range, determining whether the sliding window list of the corresponding category corresponding to the first conversion relationship is full of data, if not, adding the first conversion relationship to the sliding window list of the corresponding category, and if so, removing the first N first conversion relationships in the sliding window list of the corresponding category;

[0052] The average of the first conversion relationship in the sliding window list of the corresponding category is calculated to obtain the preset conversion relationship between the handle and the human hand of the corresponding category.

[0053] The beneficial effects of the above technical solution are: by setting the rotation and translation thresholds in the first conversion relationship, possible erroneous grasping states are eliminated, thereby improving the accuracy of the calculation of the preset conversion relationship; further, by updating the first conversion relationship of the same category of hands and handles in the sliding window list, it is ensured that the calculated preset conversion relationship is compatible with the latest grasping state of the human hand and handle, thereby improving the accuracy of the calculation of the preset conversion relationship.

[0054] In a second aspect, an embodiment of the present application provides a positioning device for a handle, comprising:

[0055] An acquisition module, used to acquire an image stream of the hand grip continuously acquired by the multi-eye camera, wherein the image stream includes a first image and a second image, and the first image and the second image are acquired at a set interval;

[0056] A detection module, used to detect the positioning light spots contained in each handle in the current first image in the image stream, and determine whether the number of the positioning light spots reaches a preset number threshold;

[0057] A positioning module, which is used to calculate and output the first observation posture of the corresponding handle in the current first image based on the positioning spots detected by each handle when the number of positioning spots reaches a preset number threshold; and when the number of positioning spots does not reach the preset number threshold, generate the initial posture of the corresponding handle in the adjacent second image of the current first image based on the posture of the corresponding human hand of each handle in the adjacent second image of the current first image and a preset conversion relationship, and use the IMU rotation data of the corresponding handle corresponding to the current first image and the initial posture to generate the second observation posture of the corresponding handle in the current first image.

[0058] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, a multi-eye camera, and a communication interface, wherein the communication interface, the multi-eye camera, the memory, and the processor are connected via a bus;

[0059] The communication interface is used to communicate with the handle;

[0060] The multi-eye camera is used to collect the image stream of the hand grip handle;

[0061] The memory stores a computer program, and the processor executes the steps of any one of the handle positioning methods provided in the first aspect according to the computer program.

[0062] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed, the steps of any handle positioning method provided in the first aspect can be implemented.

[0063] The technical effects brought about by any one of the implementation methods in the second to fourth aspects can refer to the technical effects brought about by the corresponding implementation method in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0065] Figure 1A A schematic diagram of a handle with a light ring provided in an embodiment of the present application;

[0066] Figure 1B A schematic diagram of a handle without a light ring provided in an embodiment of the present application;

[0067] Figure 2A flow chart of a method for determining a preset conversion relationship between a handle and a corresponding human hand provided in an embodiment of the present application;

[0068] Figure 3A A schematic diagram of a matching pair of a left handle and a left hand provided in an embodiment of the present application;

[0069] Figure 3B A schematic diagram of a matching pair of a right handle and a right hand provided in an embodiment of the present application;

[0070] Figure 3C A flow chart for determining the conversion relationship between handles of the same category and human hands provided in an embodiment of the present application;

[0071] Figure 4 A flowchart of a handle positioning method provided in an embodiment of the present application;

[0072] Figure 5 A flow chart of a handle positioning light spot detection method provided in an embodiment of the present application;

[0073] Figure 6 A flow chart of an auxiliary positioning method for 3D gesture detection provided in an embodiment of the present application;

[0074] Figure 7 A schematic diagram of key point detection of a human hand provided in an embodiment of the present application;

[0075] Figure 8 A flow chart of the optimization method for visual positioning provided in an embodiment of the present application;

[0076] Fig. 9 A schematic diagram of a complete handle positioning process provided in an embodiment of the present application;

[0077] Fig.10 A structural diagram of a positioning device for a handle provided in an embodiment of the present application;

[0078] Fig.11 A structural diagram of a head-mounted display device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0079] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solution of the present application, rather than all of the embodiments. Based on the embodiments recorded in the application documents, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the technical solution of the present application.

[0080] Based on the exemplary embodiments shown in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application. In addition, although the disclosure in this application is introduced according to one or several exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete technical solution separately.

[0081] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0082] The terms "first", "second", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, for example, they can be implemented in an order other than those given in the diagrams or descriptions of the embodiments of this application.

[0083] In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to those components expressly listed but may include other components not expressly listed or inherent to such products or devices.

[0084] The term "module" as used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0085] The following is an overview of the design concept of the embodiments of the present application in conjunction with application scenarios.

[0086] To solve the problem of handle positioning failure caused by insufficient number of positioning light spots, one method is to install a light ring on the top of the handle that is not easily blocked by human hands, and install multiple infrared positioning lights on the light ring to ensure that there are enough visible positioning light spots for handle positioning. This type of handle is generally called a light ring handle, such as Figure 1A However, due to the installation of the light ring, the handle with the light ring is heavy and bulky, inconvenient to carry, and long-term interaction will cause hand fatigue. Therefore, the related technology designs a handle with a hidden light ring, cancels the light ring hardware, and directly installs multiple infrared positioning lights on the handle surface, such as Figure 1BAs shown in the figure, compared with the handle with light ring, the handle with hidden light ring reduces weight and size, but the negative impact of this design is that the number of positioning light spots of the handle with hidden light ring can be easily blocked to below the minimum threshold required for the positioning algorithm to perform 6Dof pose calculation. Even with data provided by the handle IMU, the lack of 3D positioning from visual observation information will cause position deviation and poor positioning accuracy.

[0087] In view of this, an embodiment of the present application provides a method for positioning a handle, which uses the image stream of the hand holding the handle during the interaction process captured by a multi-eye camera for visual positioning. For the current first image in the image stream, when the number of positioning spots of each handle detected in the first image meets the requirements of visual positioning, since the positioning accuracy of the positioning spot is higher than the estimated accuracy, the positioning spot can be directly used to calculate the observed position and posture of the handle in the current first image to ensure high-precision positioning of the handle; when the number of positioning spots of each handle detected in the current first image does not meet the requirements of visual positioning, considering that the image stream is collected with a small and set long and short exposure time interval, and the state change of the hand holding the handle corresponding to the first image and the second image in two adjacent frames of images is negligible, that is, the position of the handle between the two adjacent frames of images The posture error is very small and the human eye can hardly feel the change. Therefore, a visual positioning algorithm of the human hand is introduced to address the situation where the visual positioning of the handle fails. Through the preset conversion relationship between the posture of the human hand and the handle, the posture of the human hand in the adjacent second image is converted into the initial posture of the corresponding handle in the adjacent second image, thereby solving the problem that the handle cannot be positioned when the positioning light is blocked and improving the stability of the handle positioning. In addition, since the IMU rotation data of the handle is a real value collected in real time, and the IMU rotation data of the handle can generally be transmitted normally during the interaction, the second observed posture of the corresponding handle in the current first image with short exposure is generated according to the IMU rotation data of the handle corresponding to the current first image and the converted initial posture, which can reduce the error existing in the posture conversion and thus improve the positioning accuracy of the handle.

[0088] In some embodiments, the short exposure is used to reduce ambient light interference to obtain an image of the positioning light on the detection handle. The long exposure is used to obtain a standard image that can be used to image recognize the hand and / or handle. The order of long and short exposures can be adjusted according to needs.

[0089] In some embodiments, the exposure time of the first image is shorter than the exposure time of the second image.

[0090] In some embodiments, the positioning light can be an LED light, or an infrared light, or an ultraviolet light, etc., which can generate light of a specific frequency band. In the embodiments of the present application, an infrared LED light is taken as an example.

[0091] It should be noted that the embodiments of the present application are not only applicable to the positioning process of a handle with a hidden light ring, but also applicable to the positioning process of a handle with a light ring.

[0092] When using the handle for interaction, the human hand needs to hold the handle. If the human hand and the handle are a fixed rigid body, then obviously no matter how they move, the relative relationship between the two will not change. In actual situations, the human hand and the handle are not rigid bodies, but because the human hand basically has a similar gripping posture when it grasps the handle correctly, it can be roughly assumed that when the same user grasps the handle correctly and moves, the relative spatial transformation relationship between the human hand and the handle is a close value. In this ideal observation situation, the transformation relationship between the relative posture of the human hand and the handle is pre-calculated.

[0093] For details on the calculation process of the conversion relationship, see Figure 2 , mainly includes the following steps:

[0094] S201: Acquire multiple grasping images of the handle in motion.

[0095] When the handle is in a static state (such as placed on a desktop), it indicates that the human hand may not be grasping the handle, and there is no need to calculate the conversion relationship between the handle and the corresponding relative posture of the human hand. Therefore, multiple grasping images of the handle in motion can be collected. Due to the limitations of the human body, even if you subjectively want to keep the handle in motion, there is still an inevitable slight tremor. The tremor of the hand holding the handle and the tremor of the handle when it is truly static will have an obvious magnitude difference in the parameter fluctuation range of the IMU acceleration sensor. Therefore, the acceleration parameters of the handle IMU can be used to judge the vibration amplitude of the handle, thereby judging whether the handle is in a motion state held by a human hand. If so, the multi-eye camera is controlled to continuously collect grasping images.

[0096] During the image acquisition process, the infrared LED light on the handle flashes at a fixed frequency, and the multi-eye camera uses two different exposure time intervals to acquire grasping images at a set time interval. Therefore, two adjacent images in the multiple grasping images can be used as a group of image pairs. Each image pair contains a first grasping image and a second grasping image. The exposure time of the first grasping image is shorter than that of the second grasping image, that is, the first grasping image is a short exposure image, and the second grasping image is a long exposure image. The LED light spot is bright in the short exposure image.

[0097] In some embodiments, the set time interval is a time interval that satisfies the position error of the handle between two frames.

[0098] In some embodiments, the set time interval is fixed. In some embodiments, the set time interval varies with the movement posture of the handle. For example, the time interval when the handle moves slowly is greater than the time interval when the handle moves quickly.

[0099] S202: For any one of the multiple image pairs, identify a first category of a handle in a first grasping image and a second category of a human hand in a second grasping image, match handles and hands of the same first category and second category to obtain matching pairs, and calculate a first conversion relationship between the handle and the corresponding hand based on the 6DOf pose of the human hand and the 6DOf pose of the handle in the same matching pair.

[0100] In actual use, there may be two hands (i.e., the left hand and the right hand) and two handles (i.e., the left handle and the right handle) in the grasping image at the same time. When calculating the conversion relationship between the relative postures of the human hands and the handles, the human hands are required to grasp the handles correctly, that is, the left hand holds the left handle and the right hand holds the right handle, without considering the special extreme cases such as the left hand holding the right handle and the right hand holding the left handle. Therefore, it is necessary to match the categories of the human hands and the handles.

[0101] Since the IMU data transmitted by the handle to the head-mounted display device carries category information, that is, whether the handle is left or right, when it is left, it indicates that the handle is a left handle, and when it is right, it indicates that the handle is a right handle. Therefore, when the LED light spot (that is, visual observation data) of the handle in the first grasping image captured by the multi-eye camera is matched with the IMU data of the handle, the handle in the first grasping image is given category information, that is, the first category of whether the handle detected in the first grasping image belongs to the left or right is obtained.

[0102] The usual 3D gesture detection method uses visual observation information to locate the coordinates of the joint points on the hand, and uses the 2D coordinates of the same hand key points observed by multiple cameras to perform positioning in 3D space. In general, 3D gesture detection algorithms usually use artificial intelligence models built with deep learning (such as OpenPose, HandNet, etc.). When the model detects the hand in the second grasp image, it also outputs the category information of the hand, that is, the second category of whether the hand detected in the second grasp image belongs to the left or right is obtained. Furthermore, after obtaining the first category of the handle and the second category of the hand, the handle and hand with the same first category and second category are treated as a matching pair, that is, the left hand and the left handle constitute a matching pair, such as Figure 3A As shown, the right hand and the right handle form a matching pair, such as Figure 3B shown.

[0103] It should be noted that when there is only one hand and handle of the same category, that is, only the left hand is missing the left handle, or only the left handle is missing the left hand, or only the right hand is missing the right handle, or only the right handle is missing the left hand, it is impossible to form a complete matching pair, so this type of data is not considered when calculating the conversion relationship.

[0104] For matching pairs of human hands and handles of the same type, when using the 6Dof handle positioning algorithm and the 3D gesture detection algorithm to calculate the position and posture of the handle and the human hand respectively, it is necessary to ensure that 4 or more LED spots on the handle appear in the first grasping image, and that 4 or more key points of the hand appear in the second grasping image. Therefore, it is necessary to judge the data integrity of the handle and hand in the matching pair. Specifically, when the number of LED spots detected in the first grasping image is greater than the first preset threshold, it indicates that the handle appears in the common viewing area of ​​the multi-eye camera and a sufficient number of LED spots are visible in the first grasping image. A stable and accurate 6Dof pose can be generated through the 6Dof handle positioning algorithm. Therefore, it can be determined that the handle corresponding to these LED spots in the first grasping image is complete in data; when the rectangular frame of the human hand in the second grasping image is all located in the second grasping image, it indicates that the human hand in the rectangular frame is in a complete form in the second grasping image, then it is determined that the human hand corresponding to the rectangular frame in the second grasping image is complete in data, and the 3D gesture detection algorithm can output the 3D pose of the human hand, that is, the 6Dof pose including the wrist root joint point. Furthermore, when the data of the human hand and the handle in the matching pair are complete, it indicates that the human hand holding the handle appears in the common viewing area of ​​the multi-eye camera, and a sufficient number of LED spots are visible in the short-exposure first grasping image. This is an ideal state that can simultaneously generate the 6Dof pose of the handle and the human hand. Therefore, the matching pair is retained. On the contrary, when the data of the human hand and the handle in the matching pair are incomplete, the state of holding the handle cannot simultaneously generate the 6Dof pose of the handle and the human hand, so the matching pair is eliminated.

[0105] In the embodiment of the present application, by ensuring that the data of the human hand and the handle are complete, the 6DOf posture of the handle and the 6DOf posture of the human hand can be accurately calculated.

[0106] Optionally, for a handle and a hand in a matching pair, after calculating their respective 6Dof poses, the translation of the handle and the hand can be extracted respectively, that is, the position coordinates of the two in three-dimensional space, and then further screening can be performed based on their position coordinates. Specifically, considering that the fingers may be blocked by the handle, resulting in inaccurate recognition, and for the human hand as a whole, the stability of the joint point positioning at the root of the wrist will be more accurate and stable than the flexible and changeable fingers at the end, so the wrist key point is used to represent the entire hand, and the Euclidean distance calculated using the position coordinates of the handle and the position coordinates of the wrist key point is used as the distance between the handle and the hand. If the distance is greater than a preset distance threshold (such as 5cm), it indicates that the hand must not be holding the corresponding handle correctly, or holding another handle of the wrong type, so the matching pair is eliminated, otherwise, it is considered that the hand and the corresponding handle maintain the correct grasping posture, so the matching pair is retained.

[0107] In an embodiment of the present application, by eliminating matching pairs with incomplete data and matching pairs with a spacing greater than a preset distance threshold, it is ensured that the matching pairs used to calculate the pose transformation matrix are the correct posture of the hand holding the handle, thereby improving the accuracy of the pose transformation matrix calculation.

[0108] Through the above screening, for the first grasping image and the second grasping image with a good observation angle, each pair of human hands and handles for which the conversion relationship needs to be calculated is obtained.

[0109] Since the time interval between the adjacent first grasping image and the second grasping image is very short, the gripping posture of the human hand and the corresponding handle remains basically unchanged during this time period. Therefore, for the matching pair of the handle and the human hand that maintain the correct gripping posture in the first grasping image and the second grasping image, the 6Dof pose of the handle and the 6Dof pose of the wrist key point are calculated based on the SLAM spatial positioning system of the head-mounted display device, and therefore belong to the same spatial coordinate system. After knowing the 6Dof pose of two objects in the same spatial coordinate system, the transformation matrix required to transform the pose of one object to the pose of another object can be directly calculated, and the matrix includes a rotation and a translation.

[0110] S203: Calculate the average of the first conversion relationship between the human hands and the handles of the same category in the plurality of grasping images to obtain a preset conversion relationship between the handles and the human hands of the corresponding category.

[0111] In the embodiment of the present application, the left hand and the left handle and the right hand and the right handle correspond to the conversion relationship between a hand and a handle respectively, and the two types of conversion relationships should not be used interchangeably. Therefore, a conversion relationship should be calculated for the left hand and the left handle and the right hand and the right handle respectively.

[0112] Specifically, the process of determining the conversion relationship between each type of human hand and the handle is as follows: Figure 3CAs shown, it mainly includes the following steps:

[0113] S2031: Setting a sliding window list of a preset length for each category.

[0114] Among them, the matching pair of the left hand and the left handle corresponds to a sliding window list, and the matching pair of the right hand and the right handle corresponds to a sliding window list, and the preset lengths of the two sliding window lists can be the same.

[0115] S2032: For each first transformation relationship between a human hand and a handle of the same category, determine whether the rotation in the first transformation relationship is within a first value range, and whether the translation in the first transformation relationship is within a second value range. If both are within the first value range, execute S2033; otherwise, execute S2037.

[0116] Among them, the rotation and translation in the first conversion relationship can characterize the posture of the hand holding the handle. For example, it is basically impossible for a human hand to rotate 120 degrees to face the handle back to grasp the corresponding handle under normal circumstances. Only when the human hand and the handle are at a good observation angle, a first conversion relationship will be calculated based on the 6Dof posture of the human hand and the corresponding handle. Therefore, the first value interval of rotation and the second value interval of translation can be set according to some extreme grasping actions. If at least one of the rotation and translation in the first conversion relationship exceeds the corresponding safe value interval, the first conversion relationship is considered to be an invalid abnormal value.

[0117] S2033: Determine whether the sliding window list of the corresponding category corresponding to the first conversion relationship is full, if so, execute S2034, if not, execute S2035.

[0118] Since the left and right categories each maintain a dynamic sliding window list, and the length of the sliding window list of each category is fixed, in order to ensure that the data does not overflow, it is necessary to determine whether the sliding window list of the corresponding category is full.

[0119] S2034: Eliminate the first N first conversion relations in the sliding window list of the corresponding category.

[0120] For each category of sliding window categories, when the sliding window list data of this category is full, in order to ensure that the first conversion relationship between the human hand and the handle of this category can be stored normally in the future, the earlier first N first conversion relationships can be removed from the sliding window list of the corresponding category, thereby ensuring that the latest first conversion relationship can be stored.

[0121] S2035: Add the first conversion relationship to the sliding window list of the corresponding category.

[0122] When the sliding window list data of the corresponding category is not full, if the rotation and translation in the first conversion relationship of the corresponding category are both within the preset value range, it indicates that the grip represented by the first conversion relationship is reasonable, so it is added to the sliding window list of the corresponding category.

[0123] S2036: Calculate the average of the first conversion relationships in the sliding window list of the corresponding category to obtain a preset conversion relationship between the handle of the corresponding category and the human hand.

[0124] In order to avoid the fluctuations in the calculation of the grasping images of the last few frames, which may cause frequent jumps in the subjective experience of the hands and handles of the same category, the first transformation relationship of the same category can be smoothed, that is, the average rotation and average translation of all the first transformation relationships in the sliding window list of the corresponding category are calculated to obtain the preset transformation relationship between the final handle and the corresponding hand corresponding to the category. The preset transformation relationship is used to bind the posture of the hands and handles of the same category.

[0125] S2037: Eliminate the first conversion relationship.

[0126] When at least one of the rotation and translation in the first conversion relationship between the same category of handles and human hands exceeds the corresponding safe value range, it indicates that the grip represented by the first conversion relationship is unreasonable and is an invalid outlier, so the first conversion relationship can be proposed.

[0127] In an embodiment of the present application, by setting the thresholds of rotation and translation in the first conversion relationship, possible erroneous grasping states are eliminated, thereby improving the accuracy of the calculation of the preset conversion relationship. Furthermore, by updating the first conversion relationship of the same category of hands and handles in the sliding window list, it is ensured that the calculated preset conversion relationship is compatible with the latest grasping state of the human hand and handle, thereby improving the accuracy of the calculation of the preset conversion relationship.

[0128] After obtaining the preset conversion relationship between the left hand and the left handle, and the preset conversion relationship between the right hand and the right handle, the handle can be positioned at 6Dof based on these two preset conversion relationships. Figure 4 , is a flow chart of a handle positioning method provided in an embodiment of the present application, and the process mainly includes the following steps:

[0129] S401: Acquire the image stream of the hand grip continuously captured by the multi-eye camera.

[0130] Generally, during the interaction, the positioning light on the handle flashes at a set frequency, so the multi-eye camera on the head-mounted display device collects the image stream of the hand holding the handle at a set frame rate. The image stream includes a first image and a second image, the first image and the second image are collected at a set interval, and the exposure time of the first image is shorter than the exposure time of the second image, that is, the first image is a short-exposure image, and the second image is a long-exposure image. The short-exposure image is an image collected when the positioning light is on.

[0131] S402: Detect the positioning light spots contained in each handle in the current first image in the image stream, and determine whether the number of positioning light spots reaches a preset number threshold. If so, execute S403; if not, execute S404.

[0132] Since the number of positioning spots determines the success or failure of the visual positioning of the handle, it is necessary to determine whether the number of positioning spots of each handle meets the definition requirements. Optionally, the preset number threshold is an integer greater than or equal to 4.

[0133] For details, please refer to the detection process of the positioning spot. Figure 5 , mainly includes the following steps:

[0134] S4021: For each handle, obtain the previous historical posture of the handle and the first IMU data segment between the corresponding image of the previous historical posture and the current first image.

[0135] Usually, the acquisition frame rate of the built-in IMU in the handle is higher than that of the multi-eye camera. In the period between the acquisition of two images, the IMU will collect multiple IMU data that represent the movement trend of the handle. Since the IMU data is transmitted to the head-mounted display device via wireless signals, it is not affected by the handle being blocked. The IMU data signal can always be sent normally, and the handle positioning is a continuous process. When the current first image is acquired, the historical posture of the handle in the previous image has been calculated.

[0136] In some embodiments, when the previous historical posture of the handle is the posture of the previous first image with short exposure, the first IMU data segment is the IMU data between the previous first image and the current first image.

[0137] In some embodiments, when the previous historical posture of the handle is the posture of the previous second image with a long exposure, the first IMU data segment is the IMU data between the previous second image and the current first image.

[0138] S4022: Perform motion estimation on the first IMU data segment and the previous historical pose to output a first predicted pose of the handle in the current first image.

[0139] Since the first IMU data segment can characterize the movement trend of the handle between the corresponding image of the previous historical posture and the current first image, kinematic estimation can be performed in combination with the previous historical posture, thereby outputting the first predicted posture of the handle corresponding to the current first image.

[0140] S4023: Determine a handle area in the current first image according to the first predicted posture, and perform LED spot detection on the handle area.

[0141] The current first image is the imaging result of the handle by the multi-eye camera. By using the pre-calibrated parameters of the multi-eye camera and the first predicted position of the handle in the three-dimensional space corresponding to the current first image, the rectangular area of ​​the 2D handle in the current first image can be determined, and the local detection of the positioning spot can be performed on the area.

[0142] In an embodiment of the present application, posture prediction is performed through the IMU data and historical posture of the handle, so that positioning spot detection is performed on the local area of ​​the current first image according to the predicted posture. Compared with global detection, the amount of calculation and calculation time are reduced, and the detection efficiency of the positioning spot is effectively improved.

[0143] S403: Calculate and output a first observation posture of the corresponding handle in the current first image according to the positioning spot detected by each handle.

[0144] Since the positioning accuracy of the handle LED spot is higher than that of the 3D gesture detection, when the number of LED spots detected on the handle meets the requirements of the 6Dof positioning algorithm, the handle is positioned directly according to the positioning spot, and the first observed posture of the corresponding handle in the three-dimensional camera coordinate system of the current first image is output.

[0145] S404: Generate an initial posture of the corresponding handle in the adjacent second image according to the posture of the human hand corresponding to each handle in the adjacent second image of the current first image and a preset conversion relationship.

[0146] When the number of positioning spots detected on the handle does not meet the requirements of the 6Dof positioning algorithm, the positioning spots cannot be used to track the handle. At this time, a 3D gesture detection algorithm can be introduced for assistance.

[0147] For detailed process, see Figure 6 , mainly includes the following steps:

[0148] S4041: For each handle in the current first image, according to the category of the handle, obtain a preset conversion relationship between the handle and the corresponding human hand.

[0149] Among them, there is a preset conversion relationship between the left handle and the left hand, and there is also a preset conversion relationship between the right handle and the right hand. The two types of preset conversion relationships are not universal. Therefore, for each handle in the current short exposure first image, the preset conversion relationship of the corresponding category is obtained.

[0150] S4042: According to a preset conversion relationship, convert the first predicted posture of the handle in the current first image into the second predicted posture of the corresponding human hand in the adjacent second image.

[0151] The first predicted posture is determined based on the previous historical posture of the handle and the first IMU data segment between the corresponding image of the previous historical posture and the current first image.

[0152] When performing posture prediction, the movement of the human hand is usually assumed to be uniform linear motion or uniformly accelerated linear motion. However, in real use scenarios, it is impossible for the human hand to always maintain a uniform or uniformly accelerated motion state. Therefore, during the interaction process, the human hand and the handle may suddenly change the direction and magnitude of the motion speed at any time. For the handle, it has the acceleration sensor signal of the IMU, and the data frame rate of the IMU is much higher than that of the multi-eye camera. Therefore, the sudden speed change will be reflected in the reading of the acceleration sensor. The 6Dof positioning algorithm of the handle can use the IMU data to correct the posture prediction result corresponding to the current first image. However, the human hand is different. The human hand without peripherals does not have an IMU that can read the acceleration parameters. Therefore, the traditional 3D gesture detection algorithm can only track the human hand through the observed posture of the previous few frames. It is impossible to quickly adjust the predicted posture when the human hand movement speed suddenly changes. The direct consequence is that when detecting the human hand, the tracking frame may be greatly deviated or even completely wrong according to the posture prediction result, which leads to the failure of the recognition of the key points of the human hand and the tracking state of the human hand is forced to be interrupted.

[0153] In order to solve the problem of positioning failure caused by the inability to sense sudden changes in hand posture during visual positioning, the handles of the same category have been bound to the hands through a preset conversion relationship, which is equivalent to providing IMU information for the corresponding hands. In this way, the first predicted posture of the handle can be used to accurately predict the posture of the corresponding hand, thereby improving the effect of tracking hand movements.

[0154] In some embodiments, the adjacent second image may be the previous second image of the current first image, or may be the next second image of the current first image.

[0155] In some embodiments, when the adjacent second image is the previous second image, since the time interval between the adjacent first image and the second image is very short, the motion state of the hand holding the handle remains basically unchanged. Therefore, for each handle, the first predicted pose of the handle in the three-dimensional camera coordinate system of the current first image can be directly used as the first predicted pose of the handle in the three-dimensional camera coordinate system of the previous second image.

[0156] In some embodiments, when the adjacent second image is the next second image, since the IMU data of the handle can be transmitted normally, the first predicted position of the handle in the three-dimensional camera coordinate system of the current first image can be motion estimated again based on the IMU data transmitted by the handle between the current first image and the next second image to obtain the first predicted position of the handle in the three-dimensional camera coordinate system of the next second image.

[0157] Furthermore, for each handle, after obtaining the first predicted pose of the handle in the three-dimensional camera coordinate system of the adjacent second image, the preset conversion relationship of the corresponding category is used to convert the first predicted pose of the handle in the adjacent second image into the second predicted pose of the corresponding hand in the three-dimensional camera coordinate system of the adjacent second image, thereby obtaining the 3D coordinates of the hand. In this way, in the case of frequent mutations in the 3D gesture motion trajectory, the IMU information built into the handle can be used to enhance the stability of hand tracking.

[0158] S4043: Perform hand detection on the hand area in the adjacent second image according to the second predicted posture to output hand key points.

[0159] Usually, 3D gesture detection will not run the full-image detection with large computational workload in real time. In order to reduce the computational workload and computational time of 3D gesture detection, local detection can be performed through prediction results. Specifically, the 3D coordinates of the hand are projected into the adjacent second image using the pre-calibrated parameters of the multi-camera, and then appropriately enlarged (for example, the length and width of the rectangle corresponding to the maximum and minimum values ​​of the projection point on the X-axis and Y-axis are enlarged by 1.2 times), and the local area of ​​the hand in the adjacent second image is obtained. Then, this small area is intercepted for hand detection, and the key points of the hand are output, such as Figure 7 As shown by the black dots in .

[0160] S4044: Determine the posture of the corresponding human hand in the adjacent second image according to the key points of the human hand, and convert the posture of the corresponding human hand into the initial posture of the handle in the adjacent second image by using a preset conversion relationship.

[0161] The key points of the human hand are used as the visual observation information of the common viewing area of ​​the multi-camera. Combined with the pre-calibrated camera parameters and the time difference between the multi-cameras, the 6DOf pose of the corresponding hand in the three-dimensional camera coordinate system corresponding to the adjacent second image can be obtained. The preset conversion relationship is used again to convert the 6Dof pose of the corresponding hand into the initial pose of the handle in the three-dimensional camera coordinate system corresponding to the adjacent second image. In this way, when it is impossible to observe sufficient positioning spots, the pose of the handle can still be located in combination with the observation data of the 3D gesture.

[0162] In the embodiments of the present application, in order to solve the problem of positioning failure caused by the inability to sense sudden changes in the hand posture during visual positioning, since the IMU data of the handle can generally be transmitted normally during the interaction process, the handle posture is predicted through the previous historical posture of the handle and the IMU data of the handle between the corresponding image of the previous historical posture and the current first image, and then the predicted posture of the handle is applied to the hand in the adjacent second image through the preset conversion relationship between the handle and the corresponding hand posture to perform local detection of the hand, thereby realizing the use of the IMU data of the handle to provide motion support for the visual positioning of the hand, ensuring that the key points of the hand can still be accurately detected in the adjacent second image when the hand movement suddenly changes, thereby improving the stability of hand tracking, and then, according to the preset conversion relationship, the hand posture is applied to the handle in the adjacent second image, thereby ensuring a continuous and stable positioning process of the handle.

[0163] S405: Generate a second observed posture of the corresponding handle in the current first image using the IMU rotation data and the initial posture of the corresponding handle corresponding to the current first image.

[0164] Considering the errors in 3D gesture estimation and posture conversion, the initial posture after conversion can be optimized through the IMU data transmitted in real time by the handle. For the specific optimization process, see Figure 8 , mainly includes the following steps:

[0165] S4051: For each handle in the current first image, determine a target position of the handle in the current first image according to the initial position of the handle in the adjacent second image.

[0166] In some embodiments, when the adjacent second image is the previous second image, the IMU data of the handle can be transmitted normally. Therefore, the initial posture can be motion estimated based on the handle IMU data between the previous second image and the current first image to obtain the target posture under the handle under the current first image.

[0167] In some embodiments, when the adjacent second image is the next second image, since the movement of the hand holding the handle is continuous, the previous historical posture of the handle (which may be the historical posture in the previous second image or the historical posture in the previous first image) and the initial posture of the handle in the next second image may be averaged to smoothly output the target posture of the handle in the current first image.

[0168] S4052: Determine a first rotation parameter of the handle in the current first image according to the gravity acceleration in the current IMU data of the handle in the current first image.

[0169] For each handle, its IMU data is transmitted to the head-mounted display device via wireless signals. Even if the positioning light of the handle is completely blocked, the IMU data signal can be sent normally. Therefore, the head-mounted display device can obtain the IMU data of the handle in real time and is not affected by occlusion. Therefore, the first rotation parameter of the current IMU can be calculated based on the gravity acceleration read in real time by the acceleration sensor in the IMU of the handle.

[0170] S4053: Replace the second rotation parameter in the target posture with the first rotation parameter to generate a second observed posture of the corresponding handle in the current frame image.

[0171] Since the earth's gravitational acceleration is a known and fixed value pointing straight down, the accuracy of the first rotation parameter calculated based on the gravitational acceleration in the IMU is higher than the second rotation parameter in the converted target pose. Therefore, the translation parameter in the target pose can be kept unchanged, and the second rotation parameter can be replaced by the first rotation parameter to accurately generate the second observed pose of the handle in the current frame image.

[0172] In an embodiment of the present application, since the first image and the second image are continuous motion images captured at set long and short exposure time intervals, after the initial posture of the handle in the adjacent second image is known, the target posture of the handle in the current first image can be determined, and since the earth's gravitational acceleration is vertically downward and is a constant, and the target posture determined according to the hand positioning result is an estimated value, there are hand positioning errors and posture conversion errors, so the accuracy of the rotation parameters calculated based on the gravitational acceleration is higher than the accuracy of the rotation parameters in the target posture. In this way, after replacing the rotation parameters in the target posture with the rotation parameters determined by the gravitational acceleration, the positioning accuracy of the handle can be improved.

[0173] For any one of the left handle or the right handle, the complete process of determining the preset conversion relationship and real-time positioning provided by the embodiment of the present application is as follows: Fig. 9 As shown, it mainly includes the following steps:

[0174] S1: Determine whether the handle is grasped by the corresponding human hand and the type of the handle based on the IMU data of the handle. If so, execute S2, otherwise do not perform the conversion relationship calculation.

[0175] Among them, whether the handle is in a moving state of being held by a human hand can be judged based on the acceleration parameters of the IMU.

[0176] S2: Detect the LED light spots in the current first grasping image and determine whether the number of LED light spots meets the positioning requirements. If so, execute S3; otherwise, do not perform the conversion relationship calculation.

[0177] Among them, the current first grasping image is a short exposure image. Since the infrared LED light on the handle may be blocked, resulting in positioning failure, it is necessary to determine whether the number of LED spots meets the positioning requirement of the 6Dof handle.

[0178] S3: Detect the human hand in the next second grasping image, obtain the category of the human hand and determine whether the human hand is complete. If so, execute S4, otherwise do not perform the conversion relationship calculation.

[0179] The next second grasped image is the next long exposure image of the current first grasped image, and the time interval between the two is fixed.

[0180] S4: Determine whether the category of the handle is the same as the category of the human hand. If so, execute S5; otherwise, do not calculate the conversion relationship.

[0181] To ensure that the human hand grasps the handle correctly, the category of the handle should be the same as the category of the human hand. For example, when the category of the handle is a left handle, the category of the human hand should be a left hand, and when the category of the handle is a right handle, the category of the human hand should be a right hand.

[0182] S5: Calculate the 6DOf pose of the handle according to the detected LED light spot, and calculate the 6DOf pose of the hand according to the detected key points of the hand.

[0183] S6: Calculate a first conversion relationship between the handle and the corresponding human hand according to the 6Dof posture of the handle and the 6Dof posture of the corresponding human hand.

[0184] S7: Determine whether the translation and rotation in the first conversion relationship are both within a preset value range. If so, execute S8; otherwise, do not perform conversion relationship calculation.

[0185] S8: Add the first conversion relationship to the sliding window list of the corresponding category.

[0186] Among them, the left handle and the left hand correspond to a first sliding window list, and the right handle and the right hand correspond to a second sliding window list. When the first conversion relationship is the posture conversion relationship between the left handle and the left hand, the first conversion relationship is added to the first sliding window list; when the first conversion relationship is the posture conversion relationship between the right handle and the right hand, the first conversion relationship is added to the second sliding window list.

[0187] S9: averaging the translation and rotation of the first conversion relationship in the sliding window list of the category, and obtaining a preset conversion relationship between the handle of the category and the corresponding human hand.

[0188] S10: Acquire the IMU data segment of the handle between the previous second image and the current first image and the historical position and posture of the handle in the previous second image.

[0189] Among them, the first image and the second image are images of the hand holding the handle captured at fixed time intervals, and the exposure time of the first image is shorter than the exposure time of the second image, that is, the current first image is a short exposure image, and the next second image is the next frame of the long exposure image of the current first image.

[0190] S11: Perform motion estimation on the IMU data segment and the historical poses, and output the first predicted pose of the handle in the current first image.

[0191] S12: Project the first predicted pose using pre-calibrated multi-camera parameters to obtain a handle area in the current first image.

[0192] S13: Perform LED spot detection on the handle area and determine whether the number of LED spots meets the positioning requirement. If so, execute S14; otherwise, execute S15.

[0193] S14: Calculate the first observation posture of the handle according to the detected LED light spot.

[0194] S15: Obtain a preset conversion relationship between the handle and the corresponding human hand, and convert the first predicted posture into a second predicted posture of the corresponding human hand in the next second image through the preset conversion relationship.

[0195] S16: Project the second predicted pose using pre-calibrated multi-camera parameters to obtain a human hand region in a next second image.

[0196] S17: Perform key point detection on the hand area and determine whether the number of key points meets the positioning requirements. If so, execute S18, otherwise global detection is used for the next frame of hands.

[0197] When the number of key points is insufficient, it indicates that hand tracking has failed, and a global detection method is used for hand detection in the next frame.

[0198] S18: Calculate the 6Dof pose of the corresponding human hand in the next second image according to the key points, and obtain the initial pose of the handle in the next second image by using a preset conversion relationship.

[0199] S19: Determine a target posture of the handle in the current first image according to the initial posture of the handle in the next second image and the historical posture of the handle in the previous second image.

[0200] S20: Calculate a first rotation parameter of the handle according to the gravity acceleration in the IMU data of the handle corresponding to the current first image.

[0201] S21: Replace the second rotation parameter in the target posture with the first rotation parameter of the handle to obtain the second observed posture of the handle in the current first image.

[0202] In the embodiment of the present application, in order to realize the detection of the positioning lights while acquiring the external image, the head mounted display device usually shoots the image stream at a set short time interval with a long and short exposure time interval. Therefore, the state change of the hand holding the handle in the adjacent short exposure first image and the long exposure second image is negligible, that is, the position error of the handle between two adjacent frames is very small. Therefore, when the number of positioning lights meets the threshold, the positioning spot is directly used to calculate the handle posture to achieve high-precision positioning of the handle. When the positioning lights on the handle are blocked to a number lower than the number required for positioning the handle, a visual positioning algorithm of the human hand is introduced. Through the preset conversion relationship corresponding to the human hand and the handle posture, the posture of the human hand in the adjacent second image with long exposure is converted to the initial posture of the corresponding handle in the current first image with short exposure, thereby solving the problem that the handle cannot be positioned when the positioning lights are blocked when the short exposure image is used for positioning the handle, and improving the stability of the handle positioning. At the same time, because the visual positioning algorithm of the human hand is introduced when the positioning lights are blocked to a number lower than the number required for positioning the handle, compared with the pure visual handle positioning technology, there will not be too much computing power consumption.

[0203] In addition, since the IMU rotation data of the handle is a real value collected in real time, and the IMU rotation data of the handle can generally be transmitted normally during the interaction process, the second observed posture of the corresponding handle in the current first image of short exposure is determined according to the IMU rotation data of the handle corresponding to the current first image of short exposure and the converted initial posture, which can reduce the error in posture conversion and thus improve the positioning accuracy of the handle.

[0204] Based on the same technical concept, an embodiment of the present application provides a handle positioning device that can implement the steps of the above-mentioned two-hand positioning method and achieve the same technical effect.

[0205] See also Fig.10The positioning device includes an acquisition module 1001, a detection module 1002 and a positioning module 1003, wherein:

[0206] An acquisition module 1001 is used to acquire an image stream of the hand grip continuously acquired by a multi-eye camera, wherein the image stream includes a first image and a second image, and the first image and the second image are acquired at a set interval;

[0207] The detection module 1002 is used to detect the positioning light spot contained in each handle in the current first image in the image stream, and determine whether the number of LED light spots reaches a preset number threshold;

[0208] The positioning module 1003 is used to calculate and output the first observation posture of the corresponding handle in the current first image based on the positioning spots detected by each handle when the number of positioning spots reaches a preset number threshold; and when the number of positioning spots does not reach the preset number threshold, generate the initial posture of the corresponding handle in the adjacent second image of the current first image based on the posture of the corresponding human hand of each handle in the adjacent second image of the current first image and a preset conversion relationship, and use the IMU rotation data and initial posture of the corresponding handle corresponding to the current first image to generate the second observation posture of the corresponding handle in the current first image.

[0209] Optionally, the positioning module 1003 is specifically used for:

[0210] For each handle, do the following:

[0211] According to the type of the handle, obtain the preset conversion relationship between the handle and the corresponding human hand;

[0212] According to a preset conversion relationship, the first predicted posture of the handle in the current first image is converted into the second predicted posture of the corresponding human hand in the adjacent second image; wherein the first predicted posture is determined according to the previous historical posture of the handle and the first IMU data segment between the corresponding image of the previous historical posture and the current first image;

[0213] Performing hand detection on the hand region in the adjacent second image according to the second predicted posture to output key points of the hand;

[0214] The posture of the corresponding human hand in the adjacent second image is determined according to the key points of the human hand, and the posture of the corresponding human hand is converted into the initial posture of the handle in the adjacent second image by using a preset conversion relationship.

[0215] Optionally, the positioning module 1003 is specifically used for:

[0216] For each handle, do the following:

[0217] Generate a target pose of the handle in the current first image according to the initial pose of the handle in the adjacent second image;

[0218] Determine a first rotation parameter of the handle in the current first image according to the gravity acceleration in the current IMU data of the handle in the current first image;

[0219] The second rotation parameter in the target pose is replaced by the first rotation parameter to generate a second observed pose of the corresponding handle in the current frame image.

[0220] Optionally, the detection module 1002 is specifically used for:

[0221] For each handle, do the following:

[0222] Obtain the last historical posture of the handle and the first IMU data segment between the corresponding image of the last historical posture and the current first image;

[0223] Perform motion estimation on the first IMU data segment and the previous historical pose to output a first predicted pose of the handle in the current first image;

[0224] According to the first predicted position and posture, a handle area in the current first image is determined, and a positioning spot detection is performed on the handle area.

[0225] Optionally, the positioning module 1003 is further used for:

[0226] Acquire multiple gripping images of the handle in motion, wherein two adjacent images in the multiple gripping images are taken as a group of image pairs, each group of image pairs includes a first gripping image and a second gripping image, the first gripping image and the second gripping image are acquired at a set time interval, and the exposure time of the first gripping image is shorter than that of the second gripping image;

[0227] For any set of image pairs among multiple sets of image pairs, execute:

[0228] identifying a first category of a handle in the first grasped image and a second category of a human hand in the second grasped image;

[0229] Matching the handles and hands that are the same in the first category and the second category to obtain matching pairs;

[0230] Calculate a first transformation relationship between the handle and the corresponding human hand according to the 6DOf pose of the human hand and the 6DOf pose of the handle in the same matching pair;

[0231] The average of the first transformation relationship between the same category of human hands and handles in multiple grasping images is calculated to obtain the preset transformation relationship between the corresponding category of handles and human hands; wherein the preset transformation relationship is used to bind the same category of human hands and handles in position and posture.

[0232] Optionally, the detection module 1002 is further configured to:

[0233] If the data of the human hand and the handle in a matching pair is incomplete, the matching pair will be removed;

[0234] According to the 6Dof poses of the human hand and the handle in the matching pair, the distance between the human hand and the handle is calculated. If the distance is greater than the preset distance threshold, the matching pair is eliminated.

[0235] Optionally, the detection module 1002 is further used for:

[0236] If the number of positioning light spots detected in the first grasping image is greater than a first preset threshold, it is determined that the data of the handle in the first grasping image is complete;

[0237] If the rectangular frame of the human hand in the second grasped image is entirely located within the second grasped image, it is determined that the data of the human hand in the second grasped image is complete.

[0238] Optionally, the positioning module 1003 is specifically used to: set a sliding window list of a preset length for each category;

[0239] The detection module 1002 is also used to: for each first conversion relationship between a human hand and a handle of the same category, determine whether the rotation in the first conversion relationship is within a first value interval, and whether the translation in the first conversion relationship is within a second value interval; determine whether both the rotation and the translation are within the corresponding value interval, and determine whether the sliding window list of the corresponding category corresponding to the first conversion relationship is full of data;

[0240] The positioning module 1003 is specifically used for: when the sliding window list is not full, adding the first conversion relationship to the sliding window list of the corresponding category; when the sliding window list is full, removing the first N first conversion relationships in the sliding window list of the corresponding category; calculating the mean of the first conversion relationships in the sliding window list of the corresponding category to obtain the preset conversion relationship between the handle and the human hand of the corresponding category.

[0241] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0242] After introducing the handle positioning method and device according to the exemplary embodiment of the present application, next, a head-mounted display device according to another exemplary embodiment of the present application is introduced.

[0243] Those skilled in the art will appreciate that various aspects of the present application may be implemented as a system, method or program product. Therefore, various aspects of the present application may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to as "circuit", "module" or "system" herein.

[0244] Based on the same inventive concept as the above method embodiment, the structure of the head mounted display device provided in the embodiment of the present application can be as follows: Fig.11 As shown, it includes a processor 1101, a memory 1102, a communication interface 1103 and a multi-eye camera 1104;

[0245] The multi-eye camera 1104 is used to collect the image stream of the hand grip handle;

[0246] The communication interface 1103 is used to communicate with the handle;

[0247] The memory 1102 stores a computer program, and the processor 1101 executes the steps of any one of the handle positioning methods in the above embodiments according to the computer program.

[0248] In the embodiment of the present application, the memory 1102 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, and programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc. The memory 1102 may be a volatile memory (volatile memory), such as a random-access memory (RAM); the memory may also be a non-volatile memory (non-volatile memory), such as a read-only memory, a flash memory (flash memory), a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 1102 may be any other medium that can be used to carry or store a desired computer program in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 1102 may be a combination of the above memories.

[0249] The processor 1101 may include one or more central processing units (CPU), GPU or a digital processing unit, etc.

[0250] In the embodiment of the present application, the specific connection medium between the multi-eye camera 1104, the communication interface 1103, the memory 1102 and the processor 1101 is not limited. In the embodiment of the present application, the bus 1105 between the multi-eye camera 1104, the communication interface 1103, the memory 1102 and the processor 1101 is Fig.11 The connections between the other components are only for illustration and are not intended to be limiting. The bus 1104 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Fig.11 The diagram shows that only one thick line is used, but this does not mean that there is only one bus or only one type of bus.

[0251] It should be noted that Fig.11 It is only the equipment necessary for the head-mounted display device to implement the positioning method of the handle in the embodiment of the present application. Not shown, the head-mounted display device may also include hardware of conventional positioning devices such as IMU, speakers, microphones, power supplies, buttons, etc.

[0252] The embodiment of the present application also provides a computer-readable storage medium for storing some instructions. When these instructions are executed, the steps of any handle positioning method in the aforementioned embodiments can be completed.

[0253] An embodiment of the present application also provides a computer program product for storing a computer program, wherein the computer program is used to execute the steps of any one of the handle positioning methods in the aforementioned embodiments.

[0254] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0255] The present application is described with reference to the flowchart and / or block diagram of the method, device (system), and computer program product according to the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram, and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart or multiple flows and / or one box or multiple boxes in the block diagram.

[0256] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0257] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0258] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for positioning a handle, characterized in that: include: Acquire an image stream of the hand grip continuously acquired by the multi-eye camera, wherein the image stream includes a first image and a second image, wherein the first image and the second image are acquired at a set interval; Detecting the positioning light spots contained in each handle in the current first image in the image stream, and determining whether the number of the positioning light spots reaches a preset number threshold; If reached, the first observation posture of the corresponding handle in the current first image is calculated and output according to the positioning spot detected by each handle; If not reached, then based on the posture of the corresponding human hand of each handle in the adjacent second image of the current first image and the preset conversion relationship, the initial posture of the corresponding handle in the adjacent second image is generated, and the second observation posture of the corresponding handle in the current first image is generated using the IMU rotation data of the corresponding handle corresponding to the current first image and the initial posture.

2. The method according to claim 1, characterized in that The step of generating an initial posture of the corresponding handle in the adjacent second image according to the posture of the human hand corresponding to each handle in the adjacent second image of the current first image and a preset conversion relationship comprises: For each handle, do the following: According to the type of the handle, obtaining a preset conversion relationship between the handle and a corresponding human hand; According to the preset conversion relationship, the first predicted posture of the handle in the current first image is converted into the second predicted posture of the corresponding human hand in the adjacent second image; wherein the first predicted posture is determined according to the previous historical posture of the handle and the first IMU data segment between the image corresponding to the previous historical posture and the current first image; Performing hand detection on the hand region in the adjacent second image according to the second predicted posture to output key points of the hand; The posture of the corresponding human hand in the adjacent second image is determined according to the key points of the human hand, and the posture of the corresponding human hand is converted into the initial posture of the handle in the adjacent second image by using the preset conversion relationship.

3. The method according to claim 1, characterized in that The step of using the IMU rotation data of the corresponding handle corresponding to the current first image and the initial pose to generate a second observed pose of the corresponding handle under the current first image includes: For each handle, do the following: Generating a target pose of the handle in the current first image according to the initial pose of the handle in the adjacent second image; Determining a first rotation parameter of the handle under the current first image according to the gravity acceleration in the current IMU data of the handle under the current first image; The second rotation parameter in the target pose is replaced by the first rotation parameter to generate a second observed pose of the corresponding handle in the current frame image.

4. The method according to claim 1, characterized in that The detecting of the positioning light spot in the current frame image of the image stream includes: For each handle, do the following: Acquire the last historical posture of the handle, and the first IMU data segment between the image corresponding to the last historical posture and the current first image; Performing motion estimation on the first IMU data segment and the previous historical pose to output a first predicted pose of the handle under the current first image; According to the first predicted position and posture, a handle area in the current first image is determined, and a positioning spot detection is performed on the handle area.

5. The method according to any one of claims 1 to 4, characterized in that The process of determining the preset conversion relationship between the handle and the corresponding human hand includes: Acquire multiple gripping images of the handle in motion, wherein two adjacent images of the multiple gripping images are taken as a group of image pairs, each group of image pairs includes a first gripping image and a second gripping image, the first gripping image and the second gripping image are acquired at a set time interval, and the exposure time of the first gripping image is shorter than that of the second gripping image; For any set of image pairs among multiple sets of image pairs, execute: identifying a first category of a handle in the first grasped image and a second category of a human hand in the second grasped image; Matching handles and human hands that are the same as the first category and the second category to obtain a matching pair; Calculate a first transformation relationship between the handle and the corresponding human hand according to the 6DOf pose of the human hand and the 6DOf pose of the handle in the same matching pair; The mean of the first transformation relationship between the human hands and the handles of the same category in the multiple grasping images is calculated to obtain a preset transformation relationship between the handles and the human hands of the corresponding category; wherein the preset transformation relationship is used to bind the human hands and the handles of the same category in position and posture.

6. The method according to claim 5, characterized in that After obtaining matching pairs of hands and handles of the same category, the method further includes: If the data of the human hand and the handle in the matching pair is incomplete, the matching pair is eliminated; According to the 6Dof postures of the human hand and the handle in the matching pair, the distance between the human hand and the handle of the matching pair is calculated. If the distance is greater than a preset distance threshold, the matching pair is eliminated.

7. The method according to claim 6, characterized in that The process of determining whether the data of the human hand and the handle in the matching pair are complete includes: If the number of positioning light spots detected in the first grasping image is greater than a first preset threshold, it is determined that the data of the handle in the first grasping image is complete; If all of the rectangular frames of the human hand in the second grasping image are located within the second grasping image, it is determined that the data of the human hand in the second grasping image is complete.

8. The method according to claim 5, characterized in that The first conversion relationship includes rotation and translation, and the average of the first conversion relationships corresponding to the human hands and handles of the same category in the plurality of grasping images is calculated to obtain a preset conversion relationship between the handles and human hands of the corresponding category; Set a sliding window list of preset length for each category; For each first conversion relationship between a human hand and a handle of the same category, determining whether a rotation in the first conversion relationship is within a first value interval, and whether a translation in the first conversion relationship is within a second value interval; If the rotation and the translation are both within the corresponding value range, determining whether the sliding window list of the corresponding category corresponding to the first conversion relationship is full of data, if not, adding the first conversion relationship to the sliding window list of the corresponding category, and if so, removing the first N first conversion relationships in the sliding window list of the corresponding category; The average of the first conversion relationship in the sliding window list of the corresponding category is calculated to obtain the preset conversion relationship between the handle and the human hand of the corresponding category.

9. A handle positioning device, characterized in that: include: An acquisition module, used to acquire an image stream of the hand grip continuously acquired by the multi-eye camera, wherein the image stream includes a first image and a second image, and the first image and the second image are acquired at a set interval; A detection module, used to detect the positioning light spots contained in each handle in the current first image in the image stream, and determine whether the number of the positioning light spots reaches a preset number threshold; A positioning module, configured to calculate and output a first observation posture of the corresponding handle under the current first image according to the positioning spots detected by each handle when the number of positioning spots reaches a preset number threshold; Furthermore, when the number of positioning spots does not reach a preset threshold, the initial posture of the corresponding handle in the adjacent second image of the current first image is generated according to the posture of the corresponding human hand of each handle in the adjacent second image of the current first image and the preset conversion relationship, and the IMU rotation data of the corresponding handle corresponding to the current first image and the initial posture are used to generate the second observation posture of the corresponding handle in the current first image.

10. A head mounted display device, characterized in that: It includes a processor, a memory, a multi-eye camera and a communication interface, wherein the communication interface, the multi-eye camera, the memory and the processor are connected via a bus; The communication interface is used to communicate with the handle; The multi-eye camera is used to collect the image stream of the hand holding the handle; The memory stores a computer program, and the processor executes the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Full-motion flight simulator positioning device based on time deviation automatic calibration

    CN121459665A