Gesture-interaction-enabled ar smart helmet device

By combining a multi-camera module and a SoC computing board with the RetinaHand network, high-precision gesture recognition for AR helmets was achieved, solving the problem of the single interaction method of existing AR helmets and realizing real-time interaction capabilities with short latency and accurate hand movement recognition.

CN115167666BActive Publication Date: 2026-02-10HENAN COSTAR GRP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210723856.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2026-02-10
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

Existing AR headsets primarily rely on buttons, voice control, or smartphone connections for interaction, lacking high-precision and high-accuracy gesture interaction capabilities, which limits communication between people, between people and machines, and even between humanoid intelligent machines.

Method used

Employing a multi-camera module, a SoC computing board, and a gesture recognition module, the system acquires gesture images through RGB and IR cameras, performs gesture recognition processing using an ARM+NPU-based SoC computing board, and combines RetinaHand's hand detection network and static gesture classification network to achieve high-precision gesture recognition and interactive control.

Benefits of technology

It achieves real-time interaction with short latency and accurate hand gesture recognition, enhancing the natural interaction capabilities of AR smart helmets and supporting operations such as function selection, clicking, and page switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115167666B_ABST
    Figure CN115167666B_ABST
Patent Text Reader

Abstract

The application discloses a gesture-interaction AR smart helmet device, which comprises an AR helmet and a gesture recognition module; wherein the AR helmet is composed of a multi-camera module, a SoC computing board, a binocular micro display screen and an optical-mechanical module; the gesture image is collected by the multi-camera module, and gesture action control instructions are recognized by gesture recognition processing on the SoC computing board based on an ARM+NPU architecture, so that functions are selected, clicked, exited and page-changed, and the functions are displayed on the binocular micro display screen and the optical-mechanical module simultaneously; the gesture recognition module runs on the SoC computing board, and hand detection and static gesture recognition are performed on the gesture image collected by the multi-camera module, wherein the hand detection adopts a hand detection network based on RetinaHand. Compared with the prior art, the application can realize gesture interaction recognition, has the characteristics of short delay, accurate hand action recognition and support for real-time interaction, and has important significance for communication and exchange between people, between people and machines, and even between human-like intelligent machines and machines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of VR / AR and natural interaction technology, and specifically to an AR smart helmet device with gesture interaction capability. Background Technology

[0002] Gestures, as an innate and natural form of interaction, serve as a crucial bridge for communication between people, between people and machines, and even between humanoid intelligent machines. There is an urgent need for them in many fields, such as communication for the deaf and mute, smart homes, robotics, healthcare, and national defense. Achieving high-precision and high-accuracy gesture recognition has become a key focus of gesture interaction research. Large-field-of-view, highly immersive AR headsets, as portable large-screen display devices, can meet the display requirements of "large field of view, high immersion, and high resolution," and are currently developing towards multi-sensor integration, virtual-real image fusion, and digital information overlay. Existing AR headsets generally use buttons, voice control, or smartphone connections for interaction. If AR headsets could also adopt gesture interaction, it would be of great significance for communication between people, between people and machines, and even between humanoid intelligent machines. Summary of the Invention

[0003] To address the aforementioned technical deficiencies, the present invention aims to provide an AR smart helmet device with gesture interaction capability, which can realize gesture interaction recognition, has the characteristics of short latency, accurate hand movement recognition, and support for real-time interaction, and is of great significance for communication between people, between people and machines, and even between humanoid intelligent machines.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is: an AR smart helmet device with gesture interaction, comprising an AR helmet and a gesture recognition module; wherein the AR helmet is composed of a multi-camera module, a SoC computing board, a binocular micro-display screen and an optical engine module, the multi-camera module captures gesture images, which are then processed by the gesture recognition on the SoC computing board based on the ARM+NPU architecture to identify gesture action control commands for function selection, clicking, exiting, and page changing actions, and are simultaneously displayed on the binocular micro-display screen and the optical engine module;

[0005] The gesture recognition module runs on the SoC computing board and performs hand detection and static gesture recognition on gesture images captured by the multi-camera module. The hand detection uses a RetinaHand-based hand detection network, and the static gesture recognition uses a static gesture classification network.

[0006] Furthermore, the multi-camera module includes an RGB high-definition camera and an IR detection camera, and transmits the video images taken by the RGB high-definition camera and the IR detection camera to the SoC computing board through the MIPI interface.

[0007] The multi-camera module includes two low-light high-definition cameras and one IR detection camera. The two low-light high-definition cameras combined have the function of expanding the field of view (FoV). The video images captured by the cameras are transmitted to the SoC computing board through the MIPI interface.

[0008] The SoC computing board features an SoC chip with an ARM core and an NPU core architecture. It inputs multiple video signals from a multi-camera module via a MIPI interface. After ISP processing, the signals undergo RGB+IR or two low-light+IR image fusion processing. The goal of this fusion is to make targets more visible and easier to identify under different lighting conditions. The NPU core runs target detection and gesture recognition algorithms. The fused image, target detection results, and gesture recognition results are simultaneously output via MIPI to the binocular micro-display and optomechanical module for display.

[0009] The binocular microdisplay uses an OLED microdisplay or an LCoS microdisplay; the optomechanical module is a near-field optical system or an optical waveguide diffraction device for near-field AR enhancement display.

[0010] The hand detection and static gesture recognition are mainly based on the NPU core of the SoC on the SoC computing board. The RetinaHand-based hand detection network includes a backbone network for feature extraction, a feature processing fusion module FPN, and a regression head module. The regression head module is used to regress the specific category and coordinate information of the target from the features processed by the feature processing fusion module FPN.

[0011] The results of static gesture recognition will be used to control functions and menu selection on the SoC computing board, APP clicks, exits, and page transitions. The static gesture classification network that implements static gesture recognition includes a feature extraction module and a normalized exponential function. The feature extraction module includes a fully connected layer, a batch normalization layer, and a nonlinear activation layer. The detected hand region of the static gesture classification network outputs C-dimensional features, representing the probability that the static gesture belongs to C categories. The normalized exponential function normalizes the probability to the range [0, 1].

[0012] The above technical solution enables an AR smart helmet device with gesture interaction, based on an AR helmet and a gesture recognition module. This invention features low latency, accurate hand gesture recognition, and support for real-time interaction, thereby improving the natural interaction capabilities of AR smart helmets. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the principle of the block diagram of the present invention.

[0014] Figure 2 This is a schematic diagram of the hand detection network structure based on RetinaHand in the gesture recognition module of this invention.

[0015] Figure 3 This is a flowchart illustrating a method for implementing visual Attention-FPN in this invention.

[0016] Figure 4 This is a schematic diagram of a static gesture classification network structure in this invention. Detailed Implementation

[0017] See appendix Figure 1 This is one embodiment of the present invention, which discloses an AR smart helmet device with gesture interaction, including: an AR helmet and a gesture recognition module; wherein, the AR helmet is composed of a multi-camera module, a SoC computing board, a binocular micro-display and an optical engine module, the multi-camera module captures gesture images, and the gesture recognition processing is performed on the SoC computing board based on the ARM+NPU architecture to identify gesture action control commands for function selection, clicking, exiting, and page changing actions, and simultaneously display them on the binocular micro-display and the optical engine module.

[0018] The multi-camera module includes an RGB high-definition camera and an IR detection camera, which transmit the captured video images to the SoC computing board via a MIP interface.

[0019] The multi-camera module may also include two low-light high-definition cameras and one IR detection camera. The two low-light high-definition cameras can be combined to extend the field of view (FoV) and send the captured video images to the SoC computing board through the MIPI interface.

[0020] The aforementioned SoC computing board typically contains a single SoC chip, which employs an ARM core + NPU core architecture. It inputs multiple video signals from a multi-camera module via a MIPI interface. After ISP processing, it performs RGB+IR or two-channel low-light+IR image fusion processing. The goal of this fusion is to make targets more visible and easier to identify under different lighting conditions. Its NPU core is used to run target detection and gesture recognition algorithms. The fused image, target detection results, and gesture recognition results are synchronously output via MIPI to the binocular microdisplay and optomechanical module for display.

[0021] The aforementioned binocular microdisplay is generally an OLED microdisplay or an LCoS microdisplay.

[0022] The aforementioned optomechanical module is a near-field optical system, which can be an optical waveguide diffraction device used for near-field AR enhancement displays.

[0023] The gesture recognition module, running on the SoC computing board, performs hand detection and static gesture action classification on gesture images captured by the multi-camera module. Hand detection uses a RetinaHand-based hand detection network, while gesture recognition uses a static gesture classification network. The results of static gesture recognition will be used to control functions and menu selection on the SoC computing board, as well as actions such as app clicks, exits, and page transitions.

[0024] The hand detection and classification network is mainly based on the NPU core of the SoC on the SoC computing board for processing.

[0025] The RetinaHand-based hand detection network is a single-stage object detection network. While its structure is based on the RetinaFace framework, many of its modules have been improved and upgraded. Specifically, more lightweight networks have been introduced as backbone networks, the Feature Pyramid Networks (FPN) have been improved, the positive and negative sample generation strategy has been changed, the Neck part has been simplified, and different loss functions have been tried.

[0026] The RetinaHand-based hand detection network follows the classic Backbone, Neck, and Head design flow in object detection algorithms. See the appendix for its network structure. Figure 2 As shown. Its network structure mainly consists of three main parts:

[0027] 1) The backbone network used for feature extraction is usually called the backbone.

[0028] 2) Feature Processing Fusion Module (FPN), also known as the network's neck module.

[0029] 3) The regression head part, usually called the head module, is used to regress the target's specific category, coordinates, and other information from the features processed by the neck module.

[0030] The RetinaHand-based hand detection network processes hand detection in three steps:

[0031] Step 1: Generation of prior anchor boxes and matching of anchor boxes with ground truth (GT). The basic principle of all single-stage object detection algorithms based on prior anchor boxes can be summarized as classification and regression after dense sampling of the original image. Therefore, generating anchor boxes is an essential step. Although the geometric meaning of the anchor box is relative to the original image, its specific generation needs to be combined with feature maps. For Retina-hand, three layers of feature maps in the network are retained, with downsampling ratios of 1 / 8, 1 / 16, and 1 / 32 relative to the original image, respectively.

[0032] Considering the characteristics of the infrared gesture image dataset of this invention and speed requirements, in one instance, the original size of the input infrared image is limited to 224x224. Therefore, the scales of the three feature maps are 28x28, 14x14, and 7x7, respectively. Each pixel in each feature map corresponds to an 8x8, 16x16, or 32x32 region in the original image, respectively. Traditional algorithms such as Faster R-CNN, SSD, and RetinaNet generate k anchor boxes with different scales and aspect ratios based on each pixel in the feature map. Typically, k=9, representing three different scales and three different aspect ratios. Furthermore, considering the near-square nature of the infrared gesture image data of this invention, only the scale needs to be considered, while the aspect ratio can be ignored, thus simplifying the anchor box design. Simultaneously, when processing the dataset, the shorter sides can be padded to force all annotations to be square.

[0033] After generating the anchor boxes, only the dense sampling of the original image is completed. Further work is needed to construct a target for supervised learning for each sample. This specifically represents the position of the target box relative to the anchor box and the category of each anchor box. That is, to determine whether the anchor box belongs to the foreground or background. If it belongs to the foreground, its specific position needs to be determined. This position is represented by the offset of the anchor box relative to the target box. This offset has two parts: the offset of the target box center point relative to the anchor box center point, and the transformation of the target box's width and height relative to the anchor box's width and height. This transformation specifically represents the scale ratio of the target box and the anchor box after logarithmic transformation.

[0034] It's important to note that to eliminate the influence of the anchor box's scale and treat all anchor boxes equally, it's necessary to normalize the target box's width and height relative to the anchor box's center point using these dimensions. Without normalization, large anchor boxes can tolerate greater deviations, while small anchor boxes become highly sensitive to deviations, which is detrimental to model training. Converting the regression to absolute scale to relative scale solves this problem. Another crucial step is transforming the target box's width and height relative to the anchor box's width and height to logarithmic space. Without this transformation, the model's output width and height can only be positive values, increasing the demands on the model and making optimization more difficult. Transforming to logarithmic space resolves this issue.

[0035] The second step is the mapping process from input to output of the entire network. The input image of 3x224x224 first passes through a backbone network composed of stacked convolutional layers for feature extraction. The features of each layer in the middle of the network are extracted and sent to the next FPN for processing. Here, the features of the last three layers of the entire backbone network are extracted. For MobileNetV1x0.25 as the backbone network, the scales of the three feature maps are 64x28x28, 128x14x14, and 256x7x7, respectively.

[0036] After FPN feature fusion, three layers of features are obtained, each with a large number of prior anchor boxes. To improve the expressive power of the features, the feature maps at this point will undergo further feature extraction through a feature refinement module composed of large convolutional kernels, expanding the receptive field of the feature maps.

[0037] The FPN described is an Attention-FPN. Feature pyramids, as an essential component in current mainstream object detection models, effectively improve the algorithm's ability to locate objects of different scales. For hand detection tasks, the size of the hand varies drastically due to the different distances and orientations of the object relative to the camera in real-world scenarios. Targets close to the camera can have a maximum pixel size of 400x400, while the furthest targets are only 20x20, demonstrating a significant scale variation. This necessitates that the object detection network possess excellent detection capabilities for both large and small targets. Traditional FPNs achieve this by upsampling high-level features and directly adding low-level features. This invention designs and implements an improved FPN that incorporates the Attention concept.

[0038] Inspired by MobileViT, this invention extends the self-attention mechanism and introduces it into the FPN module. Here, Query, Key, and Value no longer come from the same input. Query comes from a non-linear transformation of the shallow feature map, while Key and Value both come from a linear transformation of the deep feature map after upsampling. The element-wise addition operation used in the original FPN is replaced by a fusion using an attention mechanism. From the perspective of the principle of the attention mechanism, this operation can be understood as representing each pixel in the shallow feature map using a weighted sum of all pixels in the deep feature map. The advantage of this is that using a deep attention mechanism to represent the shallow layer can effectively introduce global information into each pixel in the shallow feature map, while convolution focuses more on local information. Therefore, the fused feature map retains both global and local information, which is more conducive to model learning. Finally, after obtaining a new feature map fused from the shallow and deep features using the attention mechanism, the self-attention mechanism is used again to further transform the feature map, improving the expressive power of the features.

[0039] The specific operation is as follows: The feature map of a relatively deep layer is upsampled, from 7x7 to 14x14. Then, a 1x1 convolution is used to align the number of channels with the previous layer, mapping 256 to 128, resulting in a 128x14x14 map. To perform attention operations on the obtained feature map, a approach borrowed from MobileViT is adopted: the feature map is first sliced, and self-attention is performed on all pixels within each slice. The final result is then inversely transformed to obtain the same shape as the original input feature map, thus completing one attention calculation process. (Appendix) Figure 3 The complete implementation process of Attention-FPN is demonstrated.

[0040] Step 3: These feature maps will be processed by the target bounding box regression branch and the confidence classification branch to regress the final coordinates and the probabilities of foreground and background. For this invention, if the total number of anchor boxes is represented by N, then the final output of the classification branch of the network model will be 2N, and the final output of the coordinate box regression branch will be 4N, representing the probability that each anchor box belongs to the foreground or background, and if it belongs to the foreground, the offset of the target's center point relative to the anchor box and the logarithmic transformation value of the target's width and height relative to the anchor box's width and height, respectively.

[0041] To improve localization accuracy, the loss function for regressing the target bounding box coordinates was replaced with the Intersection over Union (IoU) loss. When using absolute error to measure the distance between the output and the target, the regressed geometric quantities are independent of each other, lacking inherent geometric constraints. However, by directly optimizing the IoU between the predicted and ground truth bounding boxes, this geometric relationship can be modeled, which can also be seen as a direct optimization of the evaluation metric.

[0042] In another embodiment, the static gesture classification network includes a feature extraction module and a normalized exponential function; wherein, the feature extraction module includes a fully connected layer, a batch normalization layer, and a nonlinear activation layer; the detected hand region of the static gesture classification network outputs a C-dimensional feature, representing the probability that the static gesture belongs to C categories respectively; the normalized exponential function normalizes the probability to the range [0, 1].

[0043] In this embodiment, the static gesture classification network, such as Figure 4 As shown, the network mainly consists of a feature extraction module and a normalization exponential function. The feature extraction module is mainly composed of stacked fully connected layers, batch normalization layers, and nonlinear activation layers. The input of the dynamic gesture classification network is a sequence of K key point positions, and the output is a C-dimensional feature (C represents the number of categories), representing the probability that the gesture belongs to each of the C categories. In order to facilitate the comparison between the maximum output probability and the set threshold, the probability needs to be normalized to the range [0, 1] using an exponential normalization function.

[0044] The embodiments described above are only some embodiments of the present invention, and the concept and scope of the present invention are not limited to the details of the above exemplary embodiments. Therefore, various modifications and improvements made by other people skilled in the art based on the technical solutions of the present invention without departing from the design concept of the present invention should fall within the protection scope of the present invention, and all the contents of the claims are described in the claims.

Claims

1. An AR smart helmet device with gesture interaction capability, characterized in that: Includes AR helmets and gesture recognition modules; The AR headset comprises a multi-camera module, a SoC computing board, a binocular micro-display, and an optical engine module. The multi-camera module captures gesture images, which are then processed by the gesture recognition module on the SoC computing board based on an ARM+NPU architecture. The system identifies gesture control commands for function selection, clicking, exiting, and page changing, and displays these commands simultaneously on the binocular micro-display and the optical engine module. The gesture recognition module, running on the SoC computing board, performs hand detection and static gesture recognition on the gesture images captured by the multi-camera module. Hand detection uses a RetinaHand-based hand detection network, while static gesture recognition uses a static gesture classification network. The SoC computing board has an SoC chip with an ARM core + NPU core architecture. It inputs multiple video signals from the multi-camera module via a MIPI interface. After ISP processing, the signals undergo RGB+IR or two-channel low-light+IR image fusion processing. The goal of fusion is to make targets more visible and easier to identify under different lighting conditions. The NPU core is used to run target detection and gesture recognition algorithms. The fused image, target detection results, and gesture recognition results are synchronously output to the binocular micro-display and optomechanical module via MIPI for display. The hand detection and static gesture recognition are primarily processed by the NPU core of the SoC on the SoC computing board. The RetinaHand-based hand detection network includes a backbone network for feature extraction, a feature processing fusion module (FPN), and a regression head module. The regression head module regresses the specific category and coordinate information of the target from the features processed by the FPN. The results of static gesture recognition are used to control functions and menu selections, APP clicks, exits, and page transitions on the SoC computing board. The static gesture classification network for static gesture recognition includes a feature extraction module and a normalized exponential function. The feature extraction module includes a fully connected layer, a batch normalization layer, and a non-linear activation layer. The detected hand region of the static gesture classification network outputs C-dimensional features, representing the probability that the static gesture belongs to one of C categories. The normalized exponential function normalizes the probability to the range [0, 1]. The RetinaHand-based hand detection network processes hand detection in three steps: 1) generating prior anchor boxes and matching anchor boxes with ground truth boxes; 2) the entire network's mapping process from input to output. The input image first passes through a backbone network composed of stacked convolutional layers for feature extraction, and the features of each layer in the middle of the network are extracted and sent to the subsequent FPN for processing. 3) The feature map will be regressed through the target bounding box regression branch and the confidence classification branch to obtain the final coordinates and the probability of foreground and background.

2. The AR smart helmet device with gesture interaction according to claim 1, characterized in that: The multi-camera module includes an RGB high-definition camera and an IR detection camera, and transmits the video images taken by the RGB high-definition camera and the IR detection camera to the SoC computing board through the MIP interface.

3. The AR smart helmet device with gesture interaction according to claim 1, characterized in that: The multi-camera module includes two low-light high-definition cameras and one IR detection camera. The two low-light high-definition cameras combined have the function of expanding the field of view (FoV). The video images captured by the cameras are transmitted to the SoC computing board through the MIPI interface.

4. The AR smart helmet device with gesture interaction according to claim 1, characterized in that: The binocular microdisplay uses an OLED microdisplay or an LCoS microdisplay; the optomechanical module is a near-field optical system or an optical waveguide diffraction device for near-field AR enhancement display.

Citation Information

Patent Citations

  • Holographic helmet display with gesture recognition function

    CN104570366A

  • AR glasses control method and device, equipment and storage medium

    CN111736709A

  • Target tracking method and device, electronic equipment and storage medium

    CN112102364A

  • Augmented reality intelligence helmet based on two mesh demonstration functions

    CN207096572U