Methods for tracking extended reality input gestures and systems using them
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2026-08-14
AI Technical Summary
然而,以上提到的操作方法对于用户并不简单且可使用户筋疲力尽
[0016]基于以上描述,本公开通过利用手持装置与用户的手部之间的关系来识别用于与扩展现实进行交互的输入手势。因此,输入手势的追踪精确度可显著提高。
Smart Images

Figure CN116246335B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a method for tracking input gestures and a system using the method, and more particularly, to a method for tracking input gestures in extended reality (XR) and a system using the method. Background Technology
[0002] With technological advancements, extended reality (e.g., augmented reality (AR), virtual reality (VR), or mixed reality (MR)) headsets are becoming increasingly popular. To interact with users, these headsets can generate virtual scenes and display virtual objects (e.g., virtual buttons) within them. Users can operate the headset by pressing or pulling on these virtual objects. However, these methods are not user-friendly and can be exhausting. Summary of the Invention
[0003] This disclosure relates to a method for tracking input gestures in extended reality and a system using said method.
[0004] This disclosure relates to a system for tracking extended reality input gestures, wherein the system includes an output device, an image capture device, and a processor. The image capture device acquires an image. The processor is coupled to the output device and the image capture device, wherein the processor is configured to: detect a hand-held device and a hand in the image; detect at least one joint of the hand from the image in response to detecting a first bounding box of the hand and a second bounding box of the hand-held device; perform data fusion of the first and second bounding boxes based on the at least one joint to obtain an input gesture; and output a command corresponding to the input gesture via the output device.
[0005] In one embodiment, the processor is further configured to detect at least one joint of the hand from the image in response to the overlap of the first bounding box and the second bounding box.
[0006] In one embodiment, the processor is further configured to: perform data fusion according to a first weight of a first bounding box in response to the number of at least one joint being greater than a threshold; and perform data fusion according to a second weight of the first bounding box in response to the number of at least one joint being less than or equal to the threshold, wherein the second weight is less than the first weight.
[0007] In one embodiment, the processor is further configured to: obtain an input gesture based on a second bounding box in response to the absence of a first bounding box; and obtain an input gesture based on a first bounding box in response to the absence of a second bounding box.
[0008] In one embodiment, the system further includes a handheld device, wherein the handheld device includes a touchscreen.
[0009] In one embodiment, the processor is further configured to detect a handheld device based on positioning markers displayed by the touchscreen.
[0010] In one embodiment, the handheld device is communicatively connected to the processor, and the processor is further configured to: receive signals from the handheld device; and perform data fusion of the first bounding box, the second bounding box, and the signals to obtain an input gesture.
[0011] In one embodiment, the signal corresponds to user input received from the touchscreen of a handheld device.
[0012] In one embodiment, the handheld device further includes an inertial measurement unit, wherein the signal corresponds to data generated by the inertial measurement unit.
[0013] In one embodiment, the output device includes a display, wherein the display outputs an expanded real-world scene according to a command.
[0014] In one embodiment, the output device includes a transceiver communicatively connected to a handheld device, wherein the process is further configured to output commands to the handheld device via the transceiver.
[0015] This disclosure relates to a method for tracking extended reality input gestures, comprising: acquiring an image; detecting a handheld device and a hand in the image; detecting at least one joint of the hand from the image in response to detecting a first bounding box of the hand and a second bounding box of the handheld device; performing data fusion of the first and second bounding boxes based on the at least one joint to obtain an input gesture; and outputting a command corresponding to the input gesture via an output device.
[0016] Based on the above description, this disclosure identifies input gestures for interacting with augmented reality by utilizing the relationship between the handheld device and the user's hand. Therefore, the tracking accuracy of input gestures can be significantly improved.
[0017] To better understand the foregoing, several embodiments with accompanying drawings are described in detail below. Attached Figure Description
[0018] The accompanying drawings are included to provide a further understanding of this disclosure and are incorporated in and form a part of this specification. The drawings illustrate exemplary embodiments of this disclosure and, together with the implementation methods, serve to explain the principles of this disclosure.
[0019] Figure 1 A schematic diagram of a system for tracking extended reality input gestures according to an embodiment of the present invention is shown.
[0020] Figure 2A and Figure 2B A schematic diagram illustrating an input gesture presented by a handheld device and a user's hand according to an embodiment of the present invention.
[0021] Figure 3A and Figure 3B A schematic diagram of a touchscreen of a handheld device according to an embodiment of the present invention is shown.
[0022] Figure 4 A schematic diagram illustrating an interaction with extended reality using a handheld device with an inertial measurement unit, according to an embodiment of the present invention.
[0023] Figure 5 A flowchart illustrating a method for tracking extended reality input gestures according to an embodiment of the present invention is shown.
[0024] Explanation of icon numbers
[0025] 10: System;
[0026] 20, 30: Boundary frames;
[0027] 100: Head-mounted device;
[0028] 110, 210: Processor;
[0029] 120, 220: Storage medium;
[0030] 130: Image capturing device;
[0031] 140: Output device;
[0032] 141, 230: Transceiver;
[0033] 142: Monitor;
[0034] 200: Handheld device;
[0035] 240: Touchscreen;
[0036] 241: Location marker / touch area;
[0037] 242: Button;
[0038] 250: Inertial Measurement Unit;
[0039] 300: Hands;
[0040] 310: Joint;
[0041] 600: Extended Reality Scenarios;
[0042] 610: cursor;
[0043] S501, S502, S503, S504, S505: Steps. Detailed Implementation
[0044] Figure 1 A schematic diagram of a system 10 for tracking extended reality input gestures according to an embodiment of the present invention is shown, wherein the input gestures can be used to interact with a virtual scene generated based on extended reality technology. System 10 may include a head-mounted device 100. In one embodiment, system 10 may further include a handheld device 200 communicatively connected to the head-mounted device 100.
[0045] The head-mounted device 100 can be worn by a user to explore extended reality scenarios. The head-mounted device 100 may include a processor 110, a storage medium 120, an image capture device 130, and an output device 140.
[0046] Processor 110 may be, for example, a central processing unit (CPU) or other programmable general-purpose or special-purpose microcontroller (MCU), microprocessor, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), graphics processing unit (GPU), arithmetic logic unit (ALU), complex programmable logic device (CPLD), field programmable gate array (FPGA), or other similar device, or a combination of the above. Processor 110 may be coupled to storage medium 120, image capture device 130, and output device 140.
[0047] Storage medium 120 may be, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar element, or a combination thereof, configured to record multiple modules or various applications that can be executed by processor 110.
[0048] The image capture device 130 may be a camera or photographic device for capturing images. The image capture device 130 may include an image sensor, such as a complementary metal oxide semiconductor (CMOS) sensor or a charge coupled device (CCD) sensor.
[0049] Output device 140 may include, but is not limited to, transceiver 141 and display 142. Transceiver 141 may be configured to transmit or receive wired / wireless signals. Transceiver 141 may also perform operations such as low-noise amplification, impedance matching, mixing, up-conversion or down-conversion, filtering, amplification, etc. Headset 100 can communicate with handheld device 200 via transceiver 141.
[0050] Display 142 may include, but is not limited to, a liquid-crystal display (LCD) or an organic light-emitting diode (OLED) display. Display 142 can provide an image beam to the user's eyes to form an image on the user's retina, allowing the user to see a virtual scene generated by head-mounted device 100.
[0051] The handheld device 200 may include, but is not limited to, a smartphone or a joystick. The handheld device 200 may include a processor 210, a storage medium 220, a transceiver 230, and a touchscreen 240. In one embodiment, the handheld device 200 may further include an inertial measurement unit (IMU) 250.
[0052] Processor 210 is, for example, a CPU, or other programmable general-purpose or special-purpose MCU, microprocessor, DSP, programmable controller, ASIC, GPU, ALU, CPLD, FPGA, or other similar device, or a combination of the above. Processor 210 may be coupled to storage medium 220, transceiver 230, touch screen 240, and IMU 250.
[0053] Storage medium 220 may be, for example, any type of fixed or removable RAM, ROM, flash memory, HDD, SSD or similar element, or a combination thereof, configured to record multiple modules or various applications that can be executed by processor 210.
[0054] Transceiver 230 can be configured to transmit or receive wired / wireless signals. Transceiver 230 can also perform operations such as low-noise amplification, impedance matching, mixing, up-conversion or down-conversion, filtering, and amplification. Handheld device 200 can communicate with head-mounted device 100 via transceiver 230.
[0055] Touchscreen 240 may include, but is not limited to, a capacitive touchscreen or a resistive touchscreen. IMU 250 may include, but is not limited to, an accelerometer, a gyroscope, or a magnetometer.
[0056] Image capture device 130 can acquire images. Processor 110 can detect the acquired images to determine the handheld device 200 or the user's hand (e.g., as shown in the image). Figure 2A or Figure 2B The processor 110 determines whether the hand (300) shown in the image is in the image. Specifically, the processor 110 may detect the image based on, for example, an object detection algorithm to determine whether an object is in the image. If the object is in the image, the object detection algorithm may generate a bounding box for the object. The processor 110 may then perform an image recognition algorithm on the bounding box to identify the object within the bounding box.
[0057] Figure 2A and Figure 2B A schematic diagram illustrating an input gesture presented by a handheld device 200 and a user's hand 300 according to an embodiment of the present invention is shown. If the handheld device 200 or the hand 300 is in an image captured by the image capture device 130, then the processor 110 may generate a bounding box 20 of the handheld device 200 or a bounding box 30 of the hand 300 on the image.
[0058] In one embodiment, the positioning marker 241 may be displayed by the touchscreen 240 of the handheld device 200. The processor 110 may locate and detect the handheld device 200 and / or the hand 300 attempting to operate the handheld device 200 based on the positioning marker 241.
[0059] In response to the detection of bounding boxes 20 and 30 on the image, processor 110 may detect one or more joints 310 of hand 300 from the image. Processor 110 may detect joints 310 based on a hand tracking algorithm.
[0060] In one embodiment, if bounding box 20 and bounding box 30 overlap, then processor 110 can detect joint 310 of hand 300 from the image. If bounding box 20 and bounding box 30 overlap, then processor 110 can determine that an input gesture is obtained by handheld device 200 and hand 300. However, if bounding box 20 and bounding box 30 do not overlap, then processor 110 can determine that an input gesture is obtained based on either bounding box 20 or bounding box 30. That is, it is possible that processor 110 will use only handheld device 200 and hand 300 to obtain an input gesture. For example, if bounding box 20 of handheld device 140 is detected from the image but bounding box 30 of hand 300 is not detected from the image, then processor 110 can determine that an input gesture is obtained based solely on handheld device 140. If the bounding box 30 of hand 300 is detected from the image but the bounding box 20 of handheld device 200 is not detected from the image, then the processor 110 can determine the input gesture based solely on hand 300.
[0061] The processor 110 can perform data fusion of bounding boxes 20 and 30 based on the detected joint 310 to obtain or recognize an input gesture, wherein the input gesture can be associated with the six degrees of freedom (6DOF) pose of the hand 300. The weights of the bounding boxes 20 or 30 used for performing the data fusion can be dynamically adjusted. In some cases, the weight of bounding box 20 can be greater than the weight of bounding box 30. That is, the result of data fusion is more influenced by the handheld device 200 than by the hand 300. In some cases, the weight of bounding box 30 can be greater than the weight of bounding box 20. That is, the result of data fusion is more influenced by the hand 300 than by the handheld device 200.
[0062] In one embodiment, processor 110 may perform data fusion of bounding boxes 20 and 30 according to a first weight of bounding boxes 30 in response to the number of joints 310 being greater than a threshold (e.g., 3), and processor 110 may perform data fusion of bounding boxes 20 and 30 according to a second weight of bounding boxes 30 in response to the number of joints 310 being less than or equal to the threshold (e.g., 3), wherein the second weight is less than the first weight. In other words, if the number of joints 310 detected by processor 110 is greater than the threshold, then the weight of the bounding boxes 30 used for performing data fusion may be increased due to the image clearly displaying the hand 300. Therefore, the weight of the bounding boxes 20 used for performing data fusion may be decreased. On the other hand, if the number of joints 310 detected by processor 110 is less than or equal to the threshold, then the weight of the bounding boxes 30 used for performing data fusion may be decreased due to the fact that a large area of the hand 300 can be covered by the handheld device 200 (e.g., ...). Figure 2B(As shown in the diagram) and thus decrease. Therefore, the weights of the bounding box 20 used to perform data fusion can be increased.
[0063] In one embodiment, processor 110 may receive signals from handheld device 200 via transceiver 141. Processor 110 may perform data fusion of bounding boxes 20 and 30 and the signals to obtain or recognize input gestures.
[0064] In one embodiment, a signal from the handheld device 200 may correspond to user input received by the touchscreen 240 of the handheld device 200. Figure 3A and Figure 3B A schematic diagram of a touchscreen 240 of a handheld device 200 according to an embodiment of the present invention is shown. The touchscreen 240 provides a user interface for receiving user input, wherein the user interface may include a touch area 241 for receiving drag or swipe operations or one or more buttons 242 for receiving click operations. The user interface may be as follows: Figure 3A The image shown is presented in a vertical mode, or as follows: Figure 3B The image shown is presented in a horizontal mode.
[0065] In one embodiment, the signal from the handheld device 200 may correspond to data generated by the IMU 250. For example, the signal from the handheld device 200 may include acceleration information of the handheld device 200. Therefore, the input gesture obtained by the processor 110 may be influenced by the data generated by the IMU 250.
[0066] After performing data fusion of bounding boxes 20 and 30, processor 110 can obtain or recognize input gestures based on the results of the data fusion. Therefore, processor 110 can operate head-mounted device 100 based on input gestures. Processor 110 can output commands corresponding to input gestures via output device 140.
[0067] In one embodiment, processor 110 can transmit commands corresponding to input gestures to transceiver 141. Transceiver 141 can output the received commands to an external electronic device, such as handheld device 200. That is, head-mounted device 100 can feed back information corresponding to input gestures to handheld device 200.
[0068] In one embodiment, processor 110 can transmit a command corresponding to an input gesture to display 142. Display 142 can output an extended reality scene based on the received command. For example, suppose an input gesture obtained by processor 110 is associated with data generated by IMU 250. Then processor 110 can transmit a command corresponding to the input gesture to display 200, wherein the command can move cursor 610 in the extended reality scene 600 displayed by display 142, such as... Figure 4 shown in.
[0069] Figure 5 A flowchart illustrating a method for tracking extended reality input gestures according to an embodiment of the present invention is shown, wherein the method may be performed by, for example Figure 1 The system 10 illustrated herein is implemented as follows: In step S501, an image is acquired. In step S502, a handheld device and a hand are detected in the image. In step S503, in response to the detection of a first bounding box of the hand and a second bounding box of the handheld device, at least one joint of the hand is detected from the image. In step S504, data fusion of the first and second bounding boxes is performed based on at least one joint to obtain an input gesture. In step S505, a command corresponding to the input gesture is output via an output device.
[0070] In summary, the system of the present invention can recognize input gestures presented by the handheld device and the user's gestures based on bounding box data fusion. Users can interact with extended reality with minimal physical force. The weights used to calculate the results of data fusion can be adjusted based on the relative position between the handheld device and the user's hand, resulting in the most accurate recognition of the input gestures. The input gestures can also be correlated with data generated by the inertial measurement unit of the handheld device. Based on the above description, this disclosure provides a convenient method for users to interact with extended reality.
Claims
1. A system for tracking input gestures in extended reality, characterized in that, include: Output device; Handheld devices, including touchscreens; An image capture device, used to acquire images; as well as A processor, coupled to the output device and the image capture device, is communicatively connected to the handheld device, wherein the processor is configured to: Detect the handheld device and hand in the image; In response to detecting a first bounding box of the hand and a second bounding box of the handheld device, at least one joint of the hand is detected from the image; Performing data fusion of the first bounding box and the second bounding box based on the at least one joint to obtain the input gesture includes: receiving a signal from the handheld device, wherein the signal corresponds to user input received by the touchscreen of the handheld device; and performing data fusion of the first bounding box, the second bounding box, and the signal to identify the input gesture; and The output device outputs a command corresponding to the input gesture; The processor is further configured as follows: Detecting at least one joint of the hand from the image in response to the overlap of the first bounding box and the second bounding box; In response to the number of at least one joint being greater than a threshold, the data fusion is performed according to the first weight of the first bounding box; and In response to the number of the at least one joint being less than or equal to the threshold, the data fusion is performed according to a second weight of the first bounding box, wherein the second weight is less than the first weight.
2. The system for tracking extended reality input gestures according to claim 1, wherein the processor is further configured to: In response to the absence of the first bounding box, the input gesture is obtained based on the second bounding box; and In response to the absence of the second bounding box, the input gesture is obtained based on the first bounding box.
3. The system for tracking extended reality input gestures according to claim 1, wherein the processor is further configured to: The handheld device is detected based on the positioning markers displayed by the touchscreen.
4. The system for tracking extended reality input gestures according to claim 1, wherein the handheld device further comprises: An inertial measurement unit, wherein the signal corresponds to data generated by the inertial measurement unit.
5. The system for tracking extended reality input gestures according to claim 1, wherein the output device includes a display, wherein the display outputs an extended reality scene according to the command.
6. The system for tracking extended reality input gestures according to claim 1, wherein the output device includes a transceiver communicatively connected to the handheld device, wherein the process is further configured to: The command is output to the handheld device via the transceiver.
7. A method for tracking input gestures in extended reality, characterized in that, include: Obtain the image; Detect the handheld device and hand in the image; In response to detecting a first bounding box of the hand and a second bounding box of the handheld device, at least one joint of the hand is detected from the image; Performing data fusion of the first bounding box and the second bounding box based on the at least one joint to obtain the input gesture includes: receiving a signal from a handheld device including a touchscreen, wherein the signal corresponds to user input received by the touchscreen of the handheld device; and performing data fusion of the first bounding box, the second bounding box, and the signal to identify the input gesture; and The output device outputs a command corresponding to the input gesture; The method is characterized in that it further includes: Detecting at least one joint of the hand from the image in response to the overlap of the first bounding box and the second bounding box; In response to the number of at least one joint being greater than a threshold, the data fusion is performed according to the first weight of the first bounding box; and In response to the number of the at least one joint being less than or equal to the threshold, the data fusion is performed according to a second weight of the first bounding box, wherein the second weight is less than the first weight.
Citation Information
Patent Citations
Object matching method and device, storage medium and electronic device
CN110866532A
Keyboard input system and keyboard input method using finger gesture recognition
US20200117282A1