Gesture interaction control method, system, device and medium for aerial imaging device
By reconstructing a 3D skeletal model using infrared sensors and combining it with coordinate system mapping, the spatial deviation between gesture operations and virtual images was solved, achieving high-precision gesture recognition and interactive control, and providing a natural and intuitive air-based touch interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 北京中嘉空间展示设计有限公司
- Filing Date
- 2026-05-28
- Publication Date
- 2026-07-21
AI Technical Summary
Existing computer vision-based gesture recognition methods only consider single-dimensional feature information in gesture recognition feature extraction, which results in the inability to establish accurate spatial mapping and reliable semantic recognition between user gesture operations and virtual images in the air, leading to low precision in interactive control.
By acquiring images of the user's hand using an infrared sensor, a three-dimensional skeletal model is reconstructed, the three-dimensional spatial coordinates of the hand feature points are determined, and combined with a pre-calibrated coordinate system mapping relationship, the hand feature points are accurately transformed to the display coordinate system, generating mapping coordinates that precisely correspond to the aerial imaging content, and extracting comprehensive gesture recognition features from static posture spatial geometric features and dynamic motion trajectory features.
It significantly improves the accuracy and robustness of gesture recognition, reliably recognizes complex gesture types and generates precise interactive commands, achieving a high-precision and high-reliability contactless air touch interaction experience, and providing a natural and intuitive human-computer interaction method.
Smart Images

Figure CN122431536A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, specifically to a gesture interaction control method, system, device, and medium for an aerial imaging device. Background Technology
[0002] With the rapid development of human-computer interaction technology, traditional physical contact-based interaction methods can no longer meet the higher requirements of modern display devices in terms of hygiene, convenience and immersion. Especially in application scenarios with strict hygiene requirements or requiring remote control, such as medical care, industrial control, and public information display, there is an urgent need for a human-computer interaction technology that can achieve non-contact precise control.
[0003] Currently, computer vision-based gesture recognition methods are commonly used to address the aforementioned technical problems. These methods acquire images of the user's hand using cameras or depth sensors, extract hand features using image processing algorithms, and recognize hand gestures. The recognition results are then converted into corresponding interactive commands to control the display device. However, these methods often only consider single-dimensional feature information in gesture recognition feature extraction, lacking a comprehensive analysis of hand posture and movement processes. This results in an inability to establish an accurate spatial mapping and reliable semantic recognition between user gesture operations and virtual images in the air, leading to low precision in interactive control. Summary of the Invention
[0004] This application provides a gesture interaction control method, system, device, and medium for aerial imaging equipment to improve the accuracy of interactive control.
[0005] In a first aspect, this application provides a gesture interaction control method for an aerial imaging device. The method includes: acquiring an infrared image of a user's hand captured by a sensor below the imaging area of the aerial imaging device; reconstructing a three-dimensional skeletal model of the user's hand based on the infrared image; determining the three-dimensional spatial coordinates of multiple hand feature points in the sensor coordinate system based on the three-dimensional skeletal model; transforming the three-dimensional spatial coordinates of each hand feature point to the display coordinate system according to a pre-calibrated mapping relationship between the sensor coordinate system and the display coordinate system of the aerial imaging device, generating mapped coordinates of each hand feature point; generating gesture recognition features by combining the three-dimensional spatial coordinates of each hand feature point and the temporal changes of each hand feature point over multiple consecutive frames, wherein the gesture recognition features include spatial geometric features representing the static posture of the hand and / or dynamic trajectory features representing the movement process of the hand; identifying the type of gesture performed by the user's hand based on the gesture recognition features; and generating an interaction command by combining the gesture type and the mapped coordinates.
[0006] By employing the aforementioned technical solution, infrared sensors are used to acquire images of the user's hand and reconstruct a 3D skeletal model. This accurately determines the 3D spatial coordinates of hand feature points and, combined with a pre-calibrated coordinate system mapping relationship, precisely transforms the hand feature points from the sensor coordinate system to the display coordinate system, generating mapped coordinates that precisely correspond to the aerial imaging content. This effectively solves the spatial deviation problem between the gesture operation position and the virtual image display position. Simultaneously, by extracting comprehensive gesture recognition features that include static posture spatial geometric features and dynamic motion trajectory features, and combining this with temporal changes in hand feature points across multiple consecutive frames, the accuracy and robustness of gesture recognition are significantly improved. This enables reliable recognition of complex gesture types and the generation of precise interactive commands, achieving a high-precision, high-reliability non-contact aerial touch interaction experience. This provides users with a natural and intuitive human-computer interaction method and improves the precision of interactive control.
[0007] Secondly, this application provides a gesture interaction control system for an aerial imaging device, the system comprising: an acquisition module, a reconstruction module, a mapping module, a combination module, and an output module; wherein, The acquisition module is used to acquire an infrared image of the user's hand collected by a sensor below the imaging area of the aerial imaging device; the reconstruction module is used to reconstruct a three-dimensional skeletal model of the user's hand based on the infrared image, and determine the three-dimensional spatial coordinates of multiple hand feature points in the sensor coordinate system based on the three-dimensional skeletal model; the mapping module is used to transform the three-dimensional spatial coordinates of each hand feature point to the display coordinate system according to the pre-calibrated mapping relationship between the sensor coordinate system and the display coordinate system of the aerial imaging device, generating the mapped coordinates of each hand feature point; the combining module is used to combine the three-dimensional spatial coordinates of each hand feature point and the temporal changes of each hand feature point in consecutive frames to generate gesture recognition features, the gesture recognition features including spatial geometric features representing the static posture of the hand and / or dynamic trajectory features representing the movement process of the hand; the output module is used to identify the type of gesture performed by the user's hand based on the gesture recognition features; and generate an interaction command by combining the gesture type and the mapped coordinates.
[0008] Thirdly, this application provides an electronic device that adopts the following technical solution: it includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to enable the electronic device to execute a computer program of a gesture interaction control method for any of the above-mentioned aerial imaging devices.
[0009] Fourthly, this application provides a computer-readable storage medium that stores a computer program capable of being loaded by a processor and executing the gesture interaction control method of any of the above-mentioned aerial imaging devices.
[0010] In summary, this application includes at least one of the following beneficial technical effects: By acquiring images of the user's hand using an infrared sensor and reconstructing a 3D skeletal model, the 3D spatial coordinates of hand feature points are accurately determined. Combined with a pre-calibrated coordinate system mapping relationship, the hand feature points in the sensor coordinate system are accurately transformed to the display coordinate system, generating mapped coordinates that precisely correspond to the aerial imaging content. This effectively solves the spatial deviation problem between the gesture operation position and the virtual image display position. Simultaneously, by extracting comprehensive gesture recognition features including static posture spatial geometric features and dynamic motion trajectory features, and combining this with temporal changes in hand feature points across multiple consecutive frames, the accuracy and robustness of gesture recognition are significantly improved. This enables reliable recognition of complex gesture types and generation of precise interactive commands, achieving a high-precision, high-reliability non-contact aerial touch interaction experience. This provides users with a natural and intuitive human-computer interaction method and improves the precision of interactive control. Attached Figure Description
[0011] Figure 1 This application provides an aerial imaging gesture interaction system. Figure 2 This is a flowchart illustrating a gesture interaction control method for an aerial imaging device provided in an embodiment of this application; Figure 3 This is a diagram illustrating the architecture of a gesture recognition and interaction control system based on a Leapmotion device, as provided in an embodiment of this application. Figure 4 This is a schematic diagram of the structure of a gesture interaction control system for an aerial imaging device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0012] Explanation of reference numerals in the attached figures: 1000, electronic device; 1001, processor; 1002, communication bus; 1003, user interface; 1004, network interface; 1005, memory. Detailed Implementation
[0013] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0014] In the description of the embodiments in this application, words such as "illustrative," "for example," or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "illustrative," "for example," or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "illustrative," "for example," or "for example" is intended to present the relevant concepts in a specific manner.
[0015] Figure 1 This application provides an aerial imaging gesture interaction system, such as... Figure 1 As shown, the aerial imaging gesture interaction system of the present invention includes two core components: an aerial imaging device and a gesture recognition sensor. The gray cube represents the main body of the aerial imaging device, which integrates optical imaging elements (such as an isosceles right-angled prism array or a reflective imaging lens group) to project the image on the display screen into the air to form a suspended real image through special optical principles. The red frustum-shaped area is the three-dimensional imaging area formed by the aerial imaging device in space, where the user can see a clear suspended image with the naked eye. The white device marked with a blue circle represents the gesture recognition sensor (such as a Leap Motion sensor) installed at the bottom of the aerial imaging device. This sensor uses an active infrared illumination method combining a binocular infrared camera with three infrared LEDs, with its field of view facing upwards. The inverted frustum-shaped area marked by the white frame is the effective recognition space range of the sensor. When the user's hand enters the area within the white frame, the sensor can collect the infrared image of the hand and calculate the depth information. The black rhombus-shaped plane represents the ground or installation platform, used as a reference benchmark for establishing the world coordinate system.
[0016] The core technology of this invention lies in establishing a precise mapping relationship between the coordinate system of the gesture recognition sensor and the display coordinate system of the aerial imaging device. Specifically, firstly, multiple calibration reference points with known three-dimensional positions (such as the four vertices and the center point of the imaging area) are pre-set in the red imaging area of the aerial imaging device. Then, the user is guided to touch these calibration points that are suspended in the air with their fingers in sequence. The gesture recognition sensor captures the three-dimensional spatial coordinates of the user's finger touch in real time within its white recognition area. By recording the corresponding coordinate pairs of each calibration point in the sensor coordinate system (a three-dimensional coordinate system with the sensor center as the origin) and the display coordinate system (a three-dimensional coordinate system with the imaging area as the reference), the coordinate transformation matrix between the two coordinate systems is calculated by fitting using the least squares method or affine transformation algorithm. This matrix contains rotation, translation, and scaling parameters, which can accurately transform the coordinates of the hand feature points recognized by the sensor to the display space of the aerial imaging, thereby achieving precise alignment between the gesture operation and the suspended image.
[0017] During actual interaction, when a user's hand performs a gesture within the white recognition area, the gesture recognition sensor acquires infrared images of the hand in real time and reconstructs a 3D skeletal model. It extracts the 3D spatial coordinates of 27 hand feature points, including five fingertips, knuckles, and the center of the palm. Then, using a pre-calibrated coordinate transformation matrix, these feature point coordinates are transformed from the sensor coordinate system to the display coordinate system, generating mapped coordinates for each feature point within the red imaging area. The system uses these mapped coordinates to determine whether the user's finger has touched a specific virtual button or control element in the floating screen. Simultaneously, it considers the spatial geometric relationships of the hand feature points (such as finger bending angle and distance between fingertips) and temporal variations. The system identifies the specific type of gesture performed by the user (such as click, swipe, pinch, rotate, etc.) based on characteristics (such as motion trajectory and speed), and finally generates corresponding interaction commands to send to the aerial imaging device. This allows the content displayed in the red imaging area to respond to the user's gestures in real time and be updated and adjusted synchronously. For example, when a right swipe gesture is detected, the imaging content flips to the right; when a two-finger pinch gesture is detected, the imaging content is scaled. This achieves a natural and smooth contactless aerial interaction experience. The entire system eliminates visual deviations between two independent subsystems through precise spatial registration technology, ensuring that the user's perceived hand position completely coincides with the interaction point of the suspended image, achieving sub-millimeter level interaction accuracy.
[0018] Figure 2 This is a flowchart illustrating a gesture interaction control method for an aerial imaging device provided in an embodiment of this application. Figure 2 As shown, the method includes S101-S105: S101, acquire the infrared image of the user's hand collected by the sensor below the imaging area of the aerial imaging device.
[0019] The system acquires infrared images of the user's hand using an infrared binocular sensor positioned below the imaging area of the aerial imaging device. This arrangement addresses the technical problems of low accuracy and limited dimensionality in traditional infrared touch recognition. Traditional infrared matrix touch technology can only form a two-dimensional light curtain on the imaging plane. When the user's finger blocks the infrared beam, only planar coordinates can be detected, and depth information cannot be obtained, resulting in recognition accuracy of only millimeters and an inability to support complex three-dimensional gesture interactions.
[0020] The infrared binocular sensor used in this invention is a depth sensing device based on the principle of stereo vision. It includes two infrared cameras and three infrared LED emitters. The infrared LED emitters emit near-infrared light into the imaging area in a specific timing pattern, forming infrared illumination with a characteristic pattern on the user's hand surface. This allows for clear hand images even in low-light conditions or when visible light interference is present. The two infrared cameras simultaneously acquire infrared images of the hand from different angles, with an acquisition frequency of 120-200Hz, ensuring the capture of rapid and fine hand movements.
[0021] The sensor is mounted below the imaging area of the aerial imaging device, with an upward viewing angle. This "looking up" installation method offers several technological advantages. First, this mounting position effectively covers the user's hand movement range above the aerial imaging area, and its large 140° × 120° field of view ensures ample interaction space. Second, the downward viewing angle allows for better observation of the shape and features of the palm and fingers, avoiding the problem of fingers obscuring each other and improving the accuracy of gesture recognition. Furthermore, the sensor is concealed within the imaging device, not interfering with the user's visual experience, and also has better resistance to environmental interference.
[0022] During image acquisition, the binocular infrared camera utilizes the principle of stereo vision. By calculating the parallax of the same hand feature point in the images from two cameras, and combining the camera's intrinsic and extrinsic parameters, it can accurately calculate the three-dimensional spatial coordinates of that feature point. Infrared images have better stability than visible light images, unaffected by changes in ambient lighting, ensuring consistent image quality in various usage environments. Furthermore, infrared light is invisible to the human eye, so users will not experience lighting interference during interaction, providing a more natural interactive experience.
[0023] The acquired infrared images contain rich information about the hand, including not only its contours and textures, but more importantly, precise depth information. This depth information provides a reliable data foundation for subsequent 3D skeletal model reconstruction, enabling the system to achieve sub-millimeter-level hand tracking accuracy, a significant improvement over the millimeter-level accuracy of traditional infrared touch.
[0024] S102, based on the infrared image, reconstruct a three-dimensional skeletal model of the user's hand, and based on the three-dimensional skeletal model, determine the three-dimensional spatial coordinates of multiple hand feature points in the sensor coordinate system.
[0025] The system first performs hand detection and region segmentation on the acquired infrared images. The hand detection algorithm, based on a deep learning-based object detection network, can accurately locate the hand region in the infrared image, maintaining high detection accuracy even in complex backgrounds or with multiple object interference. The region segmentation process combines depth information and image features to precisely separate the hand from the background, generating a hand region image. This segmentation is not only based on pixel-level image features but also fully utilizes the depth information provided by the infrared binocular sensor. Through depth thresholding and connectivity analysis, background interference can be effectively eliminated, ensuring that subsequent processing targets only the actual hand region.
[0026] Next, the system extracts features from the hand region image to obtain the hand's appearance and depth features. Appearance features include two-dimensional visual features such as hand contour, texture patterns, and grayscale variations at joints; these features are extracted and encoded using a convolutional neural network. Depth features are derived from three-dimensional point cloud data obtained through stereo vision calculations, including the geometric shape of the hand surface, the relative depth relationships of different parts, and the degree of finger curvature. The combined use of appearance and depth features provides a richer and more accurate description of the hand, offering sufficient information for subsequent pose estimation.
[0027] The system inputs extracted appearance and depth features into a pre-trained hand pose estimation model for 3D skeletal model reconstruction. The hand pose estimation model is a regression network based on a deep learning architecture, trained on a large amount of labeled hand pose data, learning a complex mapping relationship from image features to 3D joint positions. The model outputs the 3D coordinates of each joint in the user's hand, including the wrist joint, the root joints of the five fingers, the middle joints, and the fingertip joints, totaling 27 key joints. The output for each joint includes precise 3D coordinates in the sensor coordinate system, with sub-millimeter accuracy.
[0028] Based on the output 3D positions of each joint and the preset topological connections of the hand skeleton, the system constructs a complete 3D skeletal model. The topological connections define the connection methods and hierarchical structure between each joint point; for example, the center of the palm connects to the base of each finger, and the internal connections of each finger are sequentially arranged from proximal to mid-distal. The 3D skeletal model not only contains the spatial position information of each joint point but also preserves the kinematic constraints between joints, such as the bending angle range of finger joints and the kinematic correlation between adjacent joints, ensuring that the reconstructed skeletal model conforms to the actual movement patterns of the human hand.
[0029] When determining the 3D spatial coordinates of hand feature points, the system extracts key hand feature points from the reconstructed 3D skeletal model. These hand feature points include the fingertips of the five fingers, the interphalangeal joints, and the calculated palm center point. The fingertips directly correspond to the distal joints of each finger in the skeletal model, while the interphalangeal joints include the proximal and distal interphalangeal joints of each finger. The palm center point is calculated based on the positions of the wrist joints and the root joints of each finger, using a weighted average to determine the geometric center of the palm. This center point serves as an important reference for hand posture.
[0030] The three-dimensional spatial coordinates of all hand feature points are represented and stored in the sensor coordinate system. The sensor coordinate system is a three-dimensional Cartesian coordinate system established with the geometric center of the infrared binocular sensor as the origin. The X and Y axes are parallel to the sensor's imaging plane, and the Z axis points vertically upwards towards the imaging area. This coordinate system ensures the consistency and accuracy of the hand feature point positions, providing a reliable data foundation for subsequent coordinate transformations and gesture recognition.
[0031] Based on the above embodiments, as an optional implementation, in S102, reconstructing the three-dimensional skeletal model of the user's hand based on the infrared image specifically includes S21-S23: S21, perform hand detection and region segmentation on the infrared image to generate a hand region image.
[0032] The system requires hand detection and region segmentation of infrared images to generate hand region images. This preprocessing step is crucial for ensuring the accuracy of subsequent 3D reconstruction. Since the raw images acquired by the infrared binocular sensor contain information about all objects within the entire detection range, such as the user's face, body, and background environment, directly processing the complete image would not only significantly increase the computational burden but also reduce the accuracy of hand pose estimation due to background noise and interference from irrelevant objects. Through precise hand detection and region segmentation, the system can accurately locate the hand position and extract a clean hand region image from complex infrared images, providing high-quality data input for subsequent feature extraction and 3D reconstruction.
[0033] The system first employs a deep learning-based object detection algorithm to coarsely locate the hand region in the infrared image. Object detection algorithms are typically based on mature architectures such as YOLO or Faster R-CNN. These algorithms, trained on large amounts of hand image data, can quickly and accurately detect the approximate location and bounding box of the hand in the image. During detection, the algorithm outputs the bounding box coordinates of the hand region, a confidence score, and information on the possible number of hands. To handle cases where both hands appear simultaneously, the system supports multi-instance detection, enabling the simultaneous identification and localization of multiple hand regions in the image.
[0034] Based on the coarse localization results, the system further employs a semantic segmentation algorithm for fine-grained pixel-level segmentation of the hand region. This semantic segmentation algorithm accurately distinguishes hand pixels from background pixels, generating a high-precision hand mask image. In the mask image, hand pixels are marked as foreground regions, and other pixels are marked as background regions. This precise pixel-level segmentation effectively removes background noise from the hand edges, ensuring that the extracted hand region image has clear boundaries and a complete hand structure. The system then crops a rectangular region containing the complete hand based on the segmentation mask, while retaining appropriate edge margins to avoid truncation effects, ultimately generating a clean and accurate hand region image for subsequent processing.
[0035] S22, perform feature extraction on the hand region image to obtain the appearance and depth features of the hand.
[0036] The system needs to extract features from hand region images to obtain appearance and depth features. This feature extraction process provides a rich visual information foundation for subsequent hand pose estimation. Hand pose estimation is a complex computer vision task. Relying solely on raw pixel information is insufficient to accurately infer 3D hand pose. However, by extracting multi-level and multi-dimensional feature representations, the system can capture key visual information such as hand shape, texture, contour, and depth, providing richer and more discriminative input features for the pose estimation model, thereby significantly improving the accuracy and robustness of 3D hand reconstruction.
[0037] The extraction of appearance features is primarily based on processing infrared intensity information from hand region images. The system employs a multi-scale convolutional neural network to extract hand appearance features, with the network architecture typically based on classic backbone networks such as ResNet or MobileNet. The convolutional neural network extracts visual features from low to high levels through multiple layers of convolution, pooling, and activation operations. Low-level features include local image patterns such as edges, corners, and textures, which describe the basic shape and contour information of the hand; high-level features include semantic information such as finger shape, palm structure, and joint position, which reflect the overall structure and posture configuration of the hand. Through a multi-scale feature fusion mechanism, the system effectively combines features from different levels to generate appearance feature vectors containing rich semantic information.
[0038] Depth feature extraction utilizes the stereo vision capabilities of an infrared binocular sensor, obtaining depth information of the hand region through disparity calculation. The system performs stereo matching on hand region images acquired by the left and right infrared cameras, calculates the disparity values between corresponding pixels, and converts the disparity into actual depth distance based on binocular calibration parameters. Acquiring depth information requires precise stereo correction and disparity calculation to ensure the accuracy and consistency of the depth data. Based on the obtained depth images, the system extracts geometric features such as depth gradient, curvature, and normal vectors. These depth features directly reflect the three-dimensional geometric structure of the hand surface, providing crucial spatial constraint information for subsequent three-dimensional pose estimation. The organic combination of appearance and depth features forms a complete multimodal feature representation, providing comprehensive and accurate visual input for the hand pose estimation model.
[0039] S23. Input appearance features and depth features into the pre-trained hand pose estimation model to generate the three-dimensional position of each joint of the user's hand. Combine the three-dimensional position of each joint with the preset topological connection relationship of the hand skeleton to construct a three-dimensional skeleton model.
[0040] The system needs to input the extracted appearance and depth features into a pre-trained hand pose estimation model to generate the 3D positions of each hand joint and construct a complete 3D skeletal model. This step is the core of achieving accurate 3D hand reconstruction. Traditional 2D hand detection methods can only obtain planar position information and cannot provide the true pose of the hand in 3D space. 3D hand pose estimation requires inferring complex 3D joint configurations from limited visual observation, which is a challenging problem in the field of computer vision. By using a specially trained deep learning model, the system can accurately infer the 3D spatial positions of each hand joint from multimodal features, and then construct a complete hand skeletal model that conforms to the human anatomical structure.
[0041] The hand pose estimation model employs an end-to-end deep neural network architecture, pre-trained on a large-scale 3D hand dataset containing hundreds of thousands of labeled samples. The model's input layer receives the appearance and depth features extracted in step S22, effectively integrating multimodal information through a feature fusion module. The backbone of the network uses a multi-layer fully connected network or Transformer architecture, capable of learning the complex nonlinear mapping relationship between features and 3D joint positions. The model's output layer directly regresses the coordinate positions of 21 key hand joints in 3D space, including the wrist joint, the root joints, interphalangeal joints, and fingertip joints of each finger.
[0042] The predefined hand skeletal topology defines the anatomical connections between the joints of the hand, based on accurate anatomical knowledge of the human hand. The hand skeletal topology includes the wrist as the root node, with each of the five fingers extending from the wrist. The joints within each finger connect sequentially from the root to the fingertip. Specifically, the thumb has two joints, and the other four fingers each have three joints, forming a tree-like hierarchical structure. This topological connection not only defines the geometric structure of the skeleton but also provides physical constraints on joint movement, ensuring that the reconstructed 3D skeletal model conforms to the actual movement patterns of the human hand.
[0043] The system constructs a complete 3D skeletal model based on the 21 joint 3D position coordinates output by the pose estimation model and the preset hand skeleton topology connections. The construction process first establishes the spatial distribution of joint nodes, and then creates bone connection segments between the corresponding joints according to the topology connections. Each bone connection includes not only the position information of the starting and ending joints, but also geometric attributes such as bone length, direction, and relative angle. To ensure the physical rationality of the skeletal model, the system performs bone length constraint checks and joint angle constraint verification, correcting or marking any abnormal configurations that exceed the normal range. The final constructed 3D skeletal model accurately reflects the user's true 3D hand pose, providing a reliable geometric foundation for subsequent hand feature point localization and gesture recognition. The model's joint localization accuracy reaches the sub-centimeter level, supporting fine gesture analysis and interactive control.
[0044] Based on the above embodiments, as an optional implementation, in S102, determining the three-dimensional spatial coordinates of multiple hand feature points of the user's hand in the sensor coordinate system according to the three-dimensional skeletal model specifically includes S24-S26: S24. In the sensor coordinate system, based on the three-dimensional skeletal model, extract the three-dimensional spatial coordinates of the fingertips and interphalangeal joints of each finger of the user's hand.
[0045] The system needs to extract the 3D spatial coordinates of the fingertips and interphalangeal joints of each finger based on a 3D skeletal model in the sensor coordinate system. This coordinate extraction process is the foundation for establishing a set of hand feature points. Since subsequent gesture recognition and interactive control require analysis based on specific key hand positions, a complete skeletal model alone is insufficient to directly support gesture understanding. It is essential to accurately extract the coordinates of feature points with clear geometric and interactive meanings. Fingertips and interphalangeal joints are core elements of hand movement and gesture expression. Fingertips represent the end position of the finger and are the primary tool for precise pointing and touch operations; interphalangeal joints determine the bending shape of the finger and are a key basis for determining the gesture type. By accurately extracting the 3D coordinates of these feature points, the system can provide a precise spatial positioning basis for gesture analysis.
[0046] The system first traverses all joint nodes in the 3D skeletal model, determining the fingertip positions of each finger based on preset hand anatomical markers. In the skeletal model, each fingertip corresponds to the terminal joint at the end of each finger, including five key points: thumb tip, index finger tip, middle finger tip, ring finger tip, and little finger tip. The system directly obtains the 3D coordinate information of these fingertip joints using joint identifiers or topological location indexes. These coordinates have already been accurately located in the sensor coordinate system. The sensor coordinate system is a 3D Cartesian coordinate system established with the optical center of the infrared binocular sensor as the origin. The X and Y axes correspond to the horizontal and vertical directions of the sensor, respectively, and the Z axis corresponds to the depth direction. All coordinate values are precisely expressed in millimeters or centimeters.
[0047] Next, the system extracts the coordinates of the interphalangeal joints of each finger. Interphalangeal joints are the joint positions connecting different bone segments of the finger, determining the finger's bending angle and posture. According to human anatomy, each of the four fingers (excluding the thumb) contains two interphalangeal joints: a proximal interphalangeal joint and a distal interphalangeal joint; the thumb, due to its unique structure, contains only one interphalangeal joint. The system identifies all intermediate joint nodes located within the fingers by traversing the topological connection structure of the skeletal model and extracts their precise three-dimensional coordinates in the sensor coordinate system. These interphalangeal joint coordinates not only record the spatial position of the joints but also implicitly contain geometric information about the finger's bending state, providing an important data foundation for subsequent gesture morphology analysis. Through this processing step, the system successfully obtains a set of key coordinates describing the hand's end-effector manipulation capabilities and finger morphological characteristics.
[0048] S25. Calculate the three-dimensional spatial coordinates of the center point of the palm based on the positions of the wrist joints and the root joints of each finger in the three-dimensional skeletal model.
[0049] The system needs to calculate the 3D spatial coordinates of the palm center point based on the positions of the wrist joints and the base joints of each finger in the 3D skeletal model. This calculation process is a crucial step in determining the overall geometric reference of the hand. The palm center point, as the geometric centroid of the hand, is of particular importance in gesture recognition and interactive control. It is not only a reference point for measuring the relative positions of the fingers and calculating hand dimensions, but also a key reference for judging the overall movement trend and spatial position of the hand. Since the palm center point does not correspond to a specific joint structure anatomically, it cannot be directly extracted from the skeletal model and must be accurately determined through geometric calculations using the positional information of relevant joint points. By calculating the weighted average based on the wrist joints and the base joints of each finger, the system can obtain palm center point coordinates that are both anatomically sound and computationally stable.
[0050] The system first identifies and extracts the 3D coordinates of wrist joints from the 3D skeletal model. Wrist joints are the root nodes of the entire hand skeletal structure, located at the connection between the palm and forearm, representing the anatomical starting point of the hand. The coordinates of the wrist joints are denoted as W(wx, wy, wz), where wx, wy, and wz represent the coordinate components of the joint along the X, Y, and Z axes of the sensor coordinate system, respectively. Subsequently, the system extracts the coordinates of the root joints of each finger. Root joints are the starting joints connecting each finger to the palm, including the root joints of the thumb, index finger, middle finger, ring finger, and little finger. The coordinates of these joints are denoted as T1(t1x, t1y, t1z), F1(f1x, f1y, f1z), M1(m1x, m1y, m1z), R1(r1x, r1y, r1z), and L1(l1x, l1y, l1z), respectively.
[0051] Based on the extracted coordinates of wrist joints and finger root joints, the system uses a weighted average method to calculate the three-dimensional spatial coordinates of the palm center point. The calculation formula considers the contribution weight of different joints to the palm center position. Wrist joints are typically assigned a higher weight due to their importance in the hand structure, while finger root joints are assigned corresponding weights based on their anatomical importance and geometric distribution. The coordinates P(px, py, pz) of the palm center point are calculated using the following weighted average formula: px = (α·wx + β·t1x + β·f1x + β·m1x + β·r1x + β·l1x) / (α + 5β). The calculation methods for py and pz coordinates are similar, where α is the weight coefficient of the wrist joints and β is the weight coefficient of each finger root joint. The weight coefficients are set based on extensive anatomical data and experimental verification to ensure that the calculated palm center point position conforms to human anatomy principles and possesses good numerical stability and applicability for gesture recognition.
[0052] S26, the fingertips, interphalangeal joints, and the center of the palm are identified as multiple hand feature points.
[0053] S103, based on the pre-calibrated mapping relationship between the sensor coordinate system and the display coordinate system of the aerial imaging device, transforms the three-dimensional spatial coordinates of each hand feature point to the display coordinate system, generating the mapped coordinates of each hand feature point.
[0054] The system needs to transform hand feature points from the sensor coordinate system to the display coordinate system and generate mapped coordinates. This coordinate transformation process is a crucial step in achieving accurate aerial interaction. Since the infrared binocular sensor and the aerial imaging device are two independent systems, they each establish different coordinate reference systems. The sensor coordinate system uses the sensor's geometric center as its origin, while the display coordinate system uses the imaging space of the aerial imaging device as its reference. Without accurate coordinate transformation, the user's gesture position will deviate from the actual position of the aerial image content, resulting in a poor interactive experience or even rendering the device unusable. This is precisely the spatial registration inaccuracy problem commonly found in traditional aerial imaging equipment.
[0055] The system first acquires a pre-calibrated coordinate transformation matrix, which is a mathematical mapping relationship between two coordinate systems established through a precise spatial calibration process. The coordinate transformation matrix is a 4×4 homogeneous transformation matrix containing complete information on rotation and translation transformations, capable of describing the complete geometric transformation relationship from the sensor coordinate system to the display coordinate system. Acquiring this transformation matrix requires precise calibration during the system installation and commissioning phase. During calibration, multiple calibration reference points at known locations are set in the imaging space of the aerial imaging device. These reference points are typically selected from key locations in the imaging space, such as the four corner points, the center point, and the midpoints of each side.
[0056] The selection and measurement process of calibration reference points are crucial to transformation accuracy. Each calibration reference point needs to obtain its precise coordinates in both the sensor and display coordinate systems. The coordinates in the sensor coordinate system are obtained by guiding the operator to precisely point a calibration tool (such as a thin indicator rod) to the calibration point, with the sensor detecting the three-dimensional position of the tool's tip. The coordinates in the display coordinate system are calculated using the internal parameters and geometric dimensions of the aerial imaging device, or directly measured using a dedicated coordinate measuring machine. Using the corresponding coordinate data from multiple calibration reference points, the system employs the least squares method or other optimization algorithms to fit the optimal coordinate transformation matrix, which minimizes the transformation error at all calibration points.
[0057] Using the coordinate transformation matrix obtained from calibration, the system calculates the coordinate transformation of each hand feature point in the sensor coordinate system. The coordinate transformation process employs homogeneous coordinate representation, expanding the three-dimensional coordinates of each hand feature point into a four-dimensional homogeneous coordinate vector, which is then multiplied by a 4×4 transformation matrix. The transformation calculation formula is: Display coordinates = Transformation matrix × Sensor coordinates, where the transformation matrix includes the rotation matrix R, the translation vector T, and the standard form of homogeneous coordinates. Through this mathematical transformation, each hand feature point can obtain its precise three-dimensional display coordinates in the display coordinate system, and the transformation accuracy directly inherits the sub-millimeter measurement accuracy of the sensor and the calibration accuracy of the calibration process.
[0058] The display coordinate system is a three-dimensional coordinate system established based on the imaging space of the aerial imaging device. Its origin is usually set at the geometric center of the imaging space. The X-axis and Y-axis correspond to the horizontal and vertical directions of the imaging plane, respectively, and the Z-axis is perpendicular to the imaging plane and points in the direction of user operation. In the display coordinate system, the three-dimensional display coordinates of each hand feature point completely describe the real spatial relationship of the feature point relative to the imaging content, providing an accurate spatial positioning basis for subsequent interactive operations.
[0059] To generate mapping coordinates suitable for 2D interface interaction, the system needs to project 3D display coordinates onto the imaging plane of the aerial imaging device. The projection process employs mathematical models of perspective or orthographic projection, calculating the 2D projection position of each hand feature point on the imaging plane based on its 3D display coordinates. The projection calculation must consider the intrinsic parameters of the aerial imaging device, such as focal length, principal point position, and distortion parameters, to ensure the projection result matches the actual imaging effect. For 3D spatial interaction applications, the system retains the complete 3D display coordinates; while for traditional 2D interface interaction, the projected 2D coordinates are primarily used as the mapping coordinates.
[0060] The generation of mapped coordinates also needs to consider the interaction priority and accuracy requirements of different hand feature points. Fingertip points typically have the highest interaction accuracy requirements because they are the primary tools for users to make precise selections and operations; palm center points are more used for coarse gesture judgment and overall hand position tracking; interphalangeal joint points are mainly used for posture analysis of complex gestures. The system will adopt corresponding coordinate accuracy and update frequency according to the purpose of different feature points, optimizing system performance while ensuring interaction accuracy.
[0061] Based on the above embodiments, as an optional implementation, in S103, transforming the three-dimensional spatial coordinates of each hand feature point to the display coordinate system and generating the mapped coordinates of each hand feature point specifically includes S31-S33: S31, obtain the pre-calibrated coordinate transformation matrix, which is obtained by fitting the corresponding coordinates of multiple calibration reference points in the sensor coordinate system and the display coordinate system respectively.
[0062] The system needs to acquire a pre-calibrated coordinate transformation matrix, which is obtained by fitting the corresponding coordinates of multiple calibration reference points in different coordinate systems. This calibration process is fundamental to ensuring the accuracy of gesture interaction. Since the infrared binocular sensor and the aerial imaging device are two independent optical systems, they each establish different coordinate reference systems. The sensor coordinate system has its origin at the optical center of the infrared sensor, while the display coordinate system is established based on the imaging optical system of the aerial imaging device. Without an accurate coordinate transformation relationship, the user's gestures within the sensor's detection range cannot accurately correspond to the correct position on the imaging plane, leading to deviations or even complete failure of the interaction. Through precise coordinate system calibration, the system can establish an accurate mapping relationship between the two coordinate systems, ensuring that the user's gestures are precisely applied to the intended display content.
[0063] The calibration process of the coordinate transformation matrix adopts a multi-point calibration method. First, multiple calibration reference points with known locations are set within the operating space of the aerial imaging device. These reference points are usually selected from key locations in the imaging space, such as the four corner points or the center point. The number of calibration reference points is generally no less than eight, and they are well distributed in three-dimensional space, covering the entire working area where the user may perform gesture operations. The system uses both the infrared binocular sensor and the coordinate measurement system of the aerial imaging device to accurately measure the three-dimensional coordinates of each calibration reference point, obtaining the coordinate set {Ps1, Ps2, ..., Psn} of the calibration reference points in the sensor coordinate system and the coordinate set {Pd1, Pd2, ..., Pdn} in the display coordinate system, where n is the total number of calibration reference points.
[0064] Based on the obtained corresponding coordinate pairs, the system uses a least-squares fitting algorithm to calculate the optimal coordinate transformation matrix T. The coordinate transformation matrix T is a 4×4 homogeneous transformation matrix that can simultaneously describe rotation, translation, and scaling transformations. Its mathematical form is a composite transformation containing a rotation matrix R, a translation vector t, and a scaling factor s. The fitting process determines the matrix parameters by minimizing the transformation errors of all calibration points. The error function is defined as the sum of squared Euclidean distances between the transformed coordinates and the true coordinates. To improve calibration accuracy and robustness, the system performs multiple rounds of calibration verification, eliminating outlier calibration points and optimizing the matrix parameters. The final coordinate transformation matrix typically achieves sub-millimeter level calibration accuracy, meeting the requirements for precise gesture interaction.
[0065] S32 uses a coordinate transformation matrix to transform the three-dimensional spatial coordinates of each hand feature point in the sensor coordinate system, generating the three-dimensional display coordinates of each hand feature point in the display coordinate system.
[0066] The system needs to use a coordinate transformation matrix to transform the 3D spatial coordinates of each hand feature point in the sensor coordinate system, generating the 3D display coordinates of each hand feature point in the display coordinate system. This transformation process is the core computational step for achieving accurate spatial mapping. The original coordinates of the hand feature points are obtained in the sensor coordinate system. Although these coordinates accurately reflect the positional relationship of the hand in the sensor's perception space, they cannot be directly used for the display control of the aerial imaging device. They must be transformed into the display coordinate system to establish a correct correspondence with the display content on the imaging plane. Through precise coordinate transformation calculations, the system can accurately map the user's true 3D position of the hand into the display space, providing a correct spatial positioning basis for subsequent interactive operations.
[0067] The system first converts the three-dimensional spatial coordinates of each hand feature point in the sensor coordinate system into homogeneous coordinates for matrix transformation calculations. For the feature point coordinates Ps(xs, ys, zs) in the sensor coordinate system, its homogeneous coordinates are represented as a four-dimensional vector [xs, ys, zs, 1]T, where the fourth component 1 is the standard form of homogeneous coordinates. The advantage of the homogeneous coordinate system is that it can simultaneously perform multiple geometric transformations such as rotation, translation, and scaling using a single matrix multiplication, greatly simplifying the coordinate transformation calculation process and improving computational efficiency.
[0068] Next, the system performs matrix transformation calculations on each hand feature point, multiplying the homogeneous coordinate vector of the feature point by the 4×4 coordinate transformation matrix T obtained in step S31 to obtain the homogeneous coordinates of the feature point in the display coordinate system. The transformation formula is Pd_homo = T × Ps_homo, where Ps_homo is the homogeneous coordinate in the sensor coordinate system, Pd_homo is the transformed homogeneous coordinate, and T is the coordinate transformation matrix. After normalization, the first three components of the calculated homogeneous coordinates are extracted as the three-dimensional display coordinates Pd(xd, yd, zd) of the feature point in the display coordinate system. This transformation process is performed simultaneously on all hand feature points to ensure that the entire hand structure maintains the correct relative position and geometry after the coordinate transformation. The accuracy of the transformation calculation directly affects the accuracy of subsequent interactive operations. The system uses double-precision floating-point arithmetic to ensure calculation accuracy, and the transformation error is controlled within 0.1 mm.
[0069] Although the 3D display coordinates of hand feature points are represented in the correct display coordinate system, the user interface of the aerial imaging device is essentially a two-dimensional imaging plane. User interactions ultimately need to act on specific pixel positions on this two-dimensional plane. Through accurate projection calculations, the system can convert the hand's position information in 3D space into 2D coordinates on the imaging plane, establishing a precise correspondence between gesture operations and interface elements, thus achieving true aerial touch interaction.
[0070] The system employs a perspective projection model to calculate the projected positions of each hand feature point on the imaging plane. This model accurately simulates the optical imaging process of an aerial imaging device. The projection calculation is based on the intrinsic parameter matrix of the aerial imaging device, which includes key optical parameters such as the focal length, principal point position, and pixel size of the imaging system. For the three-dimensional feature point coordinates Pd(xd, yd, zd) in the display coordinate system, its projected coordinates on the imaging plane are calculated using the perspective projection formulas: u = fx × xd / zd + cx, v = fy × yd / zd + cy, where (u, v) are the pixel coordinates on the imaging plane, fx and fy are the focal length parameters in the X and Y directions, cx and cy are the principal point coordinates, and zd is the depth coordinate of the feature point. This projection calculation considers the true optical characteristics of the imaging system, ensuring the accuracy of the projection results.
[0071] To adapt to different display resolutions and screen sizes, the system also performs coordinate normalization and scaling. The calculated pixel coordinates are first normalized to the standard range of [0,1], and then scaled according to the actual display resolution to obtain the final mapped coordinates. The mapped coordinates are represented in pixels and can be directly used for the positioning of interface elements and the execution of interactive operations. For feature points that exceed the boundary of the imaging plane, the system performs boundary clipping to ensure that all mapped coordinates are within the effective display range.
[0072] The generated set of mapped coordinates provides precise positional information for each hand feature point on the imaging plane. These mapped coordinates not only maintain the relative positional relationships between hand feature points but also establish an accurate spatial correspondence with the displayed content on the imaging plane. Through this complete coordinate transformation chain, the system successfully achieves precise mapping from the sensor detection space to the display interaction space, with mapping accuracy reaching the pixel level, supporting fine gesture operations and accurate interactive control. When the user performs gesture operations, the system can accurately translate their hand movements into corresponding interface operations, realizing a natural and intuitive air touch interaction experience.
[0073] S33, based on the projection position of the three-dimensional display coordinates of each hand feature point on the imaging plane of the aerial imaging device, generate the mapped coordinates of each hand feature point.
[0074] Although the 3D display coordinates of hand feature points are represented in the correct display coordinate system, the user interface of the aerial imaging device is essentially a two-dimensional imaging plane. User interactions ultimately need to act on specific pixel positions on this two-dimensional plane. Through accurate projection calculations, the system can convert the hand's position information in 3D space into 2D coordinates on the imaging plane, establishing a precise correspondence between gesture operations and interface elements, thus achieving true aerial touch interaction.
[0075] The system employs a perspective projection model to calculate the projected positions of each hand feature point on the imaging plane. This model accurately simulates the optical imaging process of an aerial imaging device. The projection calculation is based on the intrinsic parameter matrix of the aerial imaging device, which includes key optical parameters such as the focal length, principal point position, and pixel size of the imaging system. For the three-dimensional feature point coordinates Pd(xd, yd, zd) in the display coordinate system, its projected coordinates on the imaging plane are calculated using the perspective projection formulas: u = fx × xd / zd + cx, v = fy × yd / zd + cy, where (u, v) are the pixel coordinates on the imaging plane, fx and fy are the focal length parameters in the X and Y directions, cx and cy are the principal point coordinates, and zd is the depth coordinate of the feature point. This projection calculation considers the true optical characteristics of the imaging system, ensuring the accuracy of the projection results.
[0076] To adapt to different display resolutions and screen sizes, the system also performs coordinate normalization and scaling. The calculated pixel coordinates are first normalized to the standard range of [0,1], and then scaled according to the actual display resolution to obtain the final mapped coordinates. The mapped coordinates are represented in pixels and can be directly used for the positioning of interface elements and the execution of interactive operations. For feature points that exceed the boundary of the imaging plane, the system performs boundary clipping to ensure that all mapped coordinates are within the effective display range.
[0077] The generated set of mapped coordinates provides precise positional information for each hand feature point on the imaging plane. These mapped coordinates not only maintain the relative positional relationships between hand feature points but also establish an accurate spatial correspondence with the displayed content on the imaging plane. Through this complete coordinate transformation chain, the system successfully achieves precise mapping from the sensor detection space to the display interaction space, with mapping accuracy reaching the pixel level, supporting fine gesture operations and accurate interactive control. When the user performs gesture operations, the system can accurately translate their hand movements into corresponding interface operations, realizing a natural and intuitive air touch interaction experience.
[0078] S104, combining the three-dimensional spatial coordinates of each hand feature point and the temporal changes of each hand feature point over multiple consecutive frames, generates gesture recognition features, which include spatial geometric features representing the static posture of the hand and / or dynamic trajectory features representing the hand movement process.
[0079] The system needs to extract gesture recognition features from the spatial coordinates and temporal changes of hand feature points. This feature extraction process is the core of achieving intelligent gesture recognition. Traditional infrared touch technology can only detect simple positional information and cannot understand the complex meaning and intention of user gestures, resulting in interaction methods limited to basic click and drag operations. However, by extracting spatial geometric features and dynamic trajectory features, the system can deeply understand the spatial structure and motion patterns of hand postures, providing rich feature descriptions for recognizing complex 3D gestures, multi-finger collaborative operations, and continuous action sequences, thereby achieving a technological leap from simple touch control to intelligent gesture interaction.
[0080] The system first calculates spatial geometric features based on the three-dimensional spatial coordinates of each hand feature point in the same frame. These features are used to characterize the static posture information of the hand. The calculation of spatial geometric features begins with the relative distance between the center point of the palm and each fingertip. The center point of the palm serves as the geometric reference of the hand, and its distance from the five fingertips reflects the extension degree of each finger. When the fingers are fully extended, the distance between the fingertips and the center of the palm reaches its maximum value; when the fingers are bent or clenched into a fist, this distance decreases significantly. By monitoring these changes in distance, the system can determine whether the user is performing basic hand gestures such as clenching a fist, opening the palm, or extending a single finger.
[0081] Next, the system calculates the bending angle of each finger to obtain more refined hand posture information. The bending angle is calculated based on the direction vector of the line connecting adjacent joints within the same finger, quantifying the degree of finger bending by calculating the angle between the lines connecting adjacent bones. Each finger contains three bone segments: the proximal phalanx, the middle phalanx, and the distal phalanx, corresponding to two key bending angles: the proximal interphalangeal joint angle and the distal interphalangeal joint angle. This angle information accurately describes the bending state of each finger, providing detailed feature descriptions for recognizing digit gestures, letter gestures, and complex finger combinations.
[0082] The calculation of the opening angle between adjacent fingers further enriches the descriptive power of spatial geometry. The system calculates the spatial angle between adjacent fingers based on the direction vector of the line connecting their bones; this angle reflects the degree of finger separation. For example, when a user makes an "OK" gesture, a specific angle is formed between the thumb and index finger, and the other fingers also open accordingly; when making a "victory" gesture, the angle between the index and middle fingers increases significantly, while the ring and little fingers move closer together. This opening angle information provides crucial geometric constraints for distinguishing different static gestures.
[0083] To ensure consistent gesture recognition across different users and distances, the system incorporates a hand reference scale normalization mechanism. The hand reference scale is determined by calculating the distance between the center point of the palm and the wrist joint; this distance reflects, to some extent, the overall size of the user's hand. Using this reference scale, the system normalizes all calculated relative distances, generating normalized distance features. This normalization process eliminates the impact of individual user differences and variations in operating distance on gesture recognition, enabling the system to provide consistent recognition accuracy for users with different hand sizes.
[0084] The system combines normalized distance features, the bending angles of each finger, and the spread angles between adjacent fingers to form a complete spatial geometric feature vector. This feature vector typically contains geometric parameters in dozens of dimensions, comprehensively describing the static pose configuration of the hand in three-dimensional space. The extraction frequency of spatial geometric features is synchronized with the image acquisition frequency to ensure that changes in the user's hand pose can be reflected in real time.
[0085] In terms of generating dynamic trajectory features, the system acquires the three-dimensional spatial coordinates of each hand feature point in multiple consecutive frames, constructing a temporal data sequence for motion analysis. The temporal data sequence records the positional change history of each hand feature point over time, providing a data foundation for analyzing hand movement patterns. The system calculates the displacement vector of each hand feature point between adjacent frames. This displacement vector describes the spatial movement of the feature point from the previous frame to the current frame, including the direction and distance of movement. By performing a temporal derivative on the displacement vector, the system further calculates the velocity of each feature point; this velocity information reflects the speed and direction of hand movement.
[0086] Based on the temporal sequence of displacement vectors, the system employs a curve fitting algorithm to reconstruct the motion trajectory of each hand feature point. The motion trajectory fitting process considers the smoothness constraints and physical feasibility of hand movements, connecting discrete position points into a continuous and smooth spatial trajectory using mathematical models such as spline curves or Bézier curves. The fitted trajectory not only contains path information but also preserves the temporal characteristics of the motion, providing an accurate trajectory description for recognizing dynamic gestures such as waving, circling, and linear swiping.
[0087] Dynamic trajectory features also include statistical analysis of motion patterns, such as high-level motion descriptors like total trajectory length, average velocity, acceleration variation, number of direction changes, and periodicity. These statistical characteristics capture the global features of gesture movements, helping the system distinguish between different types of dynamic gestures. For example, circular trajectories exhibit obvious periodicity and relatively constant curvature, while zigzag trajectories show frequent direction changes and varying velocity patterns.
[0088] The system ultimately combines the extracted spatial geometric features and dynamic trajectory features into a comprehensive gesture recognition feature vector. The combination strategy is adjusted according to the specific application scenario and gesture type. For static gestures, it primarily relies on spatial geometric features; for dynamic gestures, it focuses on trajectory features; and for composite gestures, it needs to consider information from both types of features. The feature combination process also includes adaptive adjustment of feature weights. The system dynamically adjusts the importance weights of different features based on the currently detected gesture type and confidence level.
[0089] Based on the above embodiments, as an optional implementation, in S104, combining the three-dimensional spatial coordinates of each hand feature point and the temporal changes of each hand feature point across multiple consecutive frames to generate gesture recognition features specifically includes S41-S43: S41, based on the three-dimensional spatial coordinates of each hand feature point in the same frame, calculate the relative distance between preset hand feature point pairs and the angle between the connecting lines of each finger bone, and generate spatial geometric features representing the static posture of the hand.
[0090] The system needs to calculate the relative distances between preset pairs of hand feature points and the angles between the lines connecting the finger bones based on the three-dimensional spatial coordinates of each hand feature point in the same frame. This generates spatial geometric features representing the static posture of the hand, and this feature extraction process is a key step in building the foundation for gesture recognition. The core of gesture recognition lies in understanding the shape and spatial relationships of the hand. However, simple feature point coordinate information lacks the ability to describe the overall shape of the hand and cannot effectively distinguish different gesture types. By calculating the relative distances between feature points and the angles between the bone lines, the system can obtain geometric descriptors directly related to hand posture. These descriptors have good scale invariance and rotation invariance, and can effectively cope with the interference caused by differences in user hand size and changes in operating angles on gesture recognition.
[0091] The system first pre-defines multiple pairs of geometrically significant hand feature points based on hand anatomy and gesture recognition requirements. These feature point pairs include the distance between adjacent fingertips (e.g., the fingertips of the thumb and index finger, index finger and middle finger), the distance between fingertips and the center of the palm, and the distance between adjacent joints within the same finger. The system calculates the Euclidean distance between each pair of feature points. For feature points Pi(xi, yi, zi) and Pj(xj, yj, zj), the relative distance is calculated as d = √[(xi-xj)² + (yi-yj)² + (zi-zj)²]. These distance features directly reflect the degree of hand opening, finger extension, and overall scale information, providing crucial information for distinguishing basic postures such as clenched fists, open palms, and digital gestures.
[0092] Next, the system calculates the angle features between the bone lines connecting each finger. A bone line is a straight line segment connecting adjacent joints within the same finger. The angle features describe the degree of bending and relative position of the fingers. The system calculates the angles between bone lines through vector operations. For two bone lines AB and BC, formed by three consecutive joints A, B, and C, the angle θ is calculated using the dot product formula of vectors AB and BC: cosθ = (AB·BC) / (|AB||BC|). These angle features accurately describe the bending state of each finger, providing detailed morphological information for recognizing complex gestures. The generated spatial geometric feature vector contains approximately 30-50 geometric descriptors, forming a complete mathematical description of the static posture of the hand, providing rich and stable input features for subsequent gesture classification algorithms.
[0093] Based on the above embodiments, as an optional implementation, in S41, the relative distance between preset pairs of hand feature points and the angle between the lines connecting the finger bones are calculated according to the three-dimensional spatial coordinates of each hand feature point in the same frame, and spatial geometric features representing the static posture of the hand are generated, specifically including S411-S415: S411, Calculate the relative distance between each fingertip and the center point of the palm based on the three-dimensional spatial coordinates of the palm center point and the three-dimensional spatial coordinates of each fingertip point.
[0094] The system calculates the relative distances between the palm center point and each fingertip using the three-dimensional spatial coordinates of hand feature points. This distance calculation is a fundamental step in constructing static hand posture features. The distance from the palm center point to each fingertip directly reflects the extension state of the fingers and the overall shape of the hand, and is a key geometric parameter for distinguishing basic gesture types such as clenched fist, open palm, and digital gestures. By accurately calculating these distance values, the system can quantify the basic morphological features of the hand, providing a stable and reliable geometric basis for subsequent gesture recognition. The system calculates the Euclidean distances from the palm center point P(px, py, pz) to the fingertips T (thumb), F (index finger), M (m middle finger), R (ring finger), and L (little finger), respectively, using the formula d = √[(x2-x1)² + (y2-y1)² + (z2-z1)²], resulting in five distance values denoted as dT, dF, dM, dR, and dL. These distance values can effectively distinguish between finger extension and bending states, providing important discriminative information for gesture classification.
[0095] S412, calculate the angle between adjacent bone lines based on the direction vector of the line connecting adjacent joints of the same finger, and use it as the bending angle of each finger.
[0096] The system calculates the angle between adjacent bone segments based on the direction vector of the line connecting adjacent joints of the same finger. This angle is used as the bending angle of each finger, precisely quantifying the degree of finger bending. Finger bending angle is a core feature in gesture recognition, accurately reflecting the joint activity state of the fingers and serving as a crucial basis for distinguishing similar gestures. The system first calculates the direction vector between adjacent joints within the same finger. For example, for the index finger, it calculates the vector V1 from the base joint to the proximal interphalangeal joint and the vector V2 from the proximal interphalangeal joint to the distal interphalangeal joint. Then, it calculates the angle between the two bone segments using the vector angle formula θ = arccos(V1·V2 / (|V1|×|V2|)). This angle directly reflects the degree of bending of the corresponding joint; a smaller angle indicates greater finger bending, while an angle close to 180 degrees indicates the finger is essentially straight. By performing similar calculations on each joint of all fingers, the system obtains a complete set of finger bending angles, accurately describing the bending state of each finger.
[0097] S413, calculate the opening angle between adjacent fingers based on the direction vector of the line connecting the bones of adjacent fingers.
[0098] The system calculates the opening angle between adjacent fingers based on the direction vectors of the lines connecting their bones. This angle reflects the relative positional relationship between the fingers. The degree of finger opening is a crucial feature for gesture recognition, especially in digital and symbolic gestures, where the opening angle often determines the specific meaning of the gesture. The system calculates the direction vectors of the principal bones of adjacent fingers, defined as the straight line from the base of the finger joint to the fingertip. Then, it calculates the angle between the principal bones of adjacent fingers using the vector angle formula. For example, the opening angle between the index and middle fingers is obtained by calculating the angle between the principal bone vector VF of the index finger and the principal bone vector VM of the middle finger, using the formula θFM = arccos(VF·VM / (|VF|×|VM|)). The system sequentially calculates the opening angles between the thumb and index finger, index and middle finger, middle and ring finger, and ring and little finger, obtaining four opening angle values. These angle values effectively distinguish between different finger configurations, such as fingers tightly closed, naturally open, and deliberately separated.
[0099] S414. Determine the hand reference scale based on the distance between the center point of the palm and the wrist joint. Normalize the relative distances using the hand reference scale to generate normalized distance features.
[0100] A hand reference scale is determined based on the distance between the center point of the palm and the wrist joint. This reference scale is then used to normalize the relative distances, generating normalized distance features. Since hand sizes vary significantly among users, the original distance values can change dramatically due to individual differences. Directly using the original distances for gesture recognition can lead to insufficient generalization ability of the algorithm. The introduction of the hand reference scale eliminates the influence of individual differences, making the gesture recognition algorithm more adaptable to users. The system first calculates the distance dref = √[(px-wx)² + (py-wy)² + (pz-wz)²] between the center point P of the palm and the wrist joint W. This distance is relatively stable anatomically and can well represent the overall size of the user's hand. Then, the system normalizes the fingertip distances calculated in step S411 using a reference scale. The normalization formula is dnorm = d / dref, resulting in a normalized distance feature set {dT_norm, dF_norm, dM_norm, dR_norm, dL_norm}. These normalized features eliminate the influence of hand size, exhibit good individual invariance, and significantly improve the accuracy and robustness of gesture recognition.
[0101] S415 combines the normalized distance features, the bending angle of each finger, and the opening angle between adjacent fingers into spatial geometric features.
[0102] S42, obtain the three-dimensional spatial coordinates of each hand feature point in multiple consecutive frames, calculate the displacement vector and motion velocity of each hand feature point between adjacent frames, fit the motion trajectory of each hand feature point according to the temporal sequence of each displacement vector, and generate dynamic trajectory features that characterize the hand movement process.
[0103] The system needs to acquire the three-dimensional spatial coordinates of each hand feature point in multiple consecutive frames, calculate the displacement vector and motion velocity of each hand feature point between adjacent frames, and fit the motion trajectory of each hand feature point based on the temporal sequence of each displacement vector to generate dynamic trajectory features representing the hand movement process. While static spatial geometric features can describe the instantaneous posture of the hand, they cannot capture the dynamic changes in gestures. Many distinctive gestures are reflected in their motion patterns and temporal changes. For example, waving gestures, swiping gestures, and rotating gestures all require analysis of the hand's motion trajectory for accurate recognition. By extracting dynamic trajectory features, the system can understand the temporal evolution of gestures, greatly improving the accuracy and richness of gesture recognition, enabling the system to support more diverse types of interactive gestures.
[0104] The system maintains a sliding time window, continuously acquiring and storing the coordinate data of hand feature points from the most recent N frames (typically N = 10-20 frames). The length of the time window is optimized based on the typical duration of the gesture and the system's frame rate. For each hand feature point Pi, the system calculates its displacement vector between adjacent frames. The displacement vector is defined as the difference between the coordinates of the current frame and the coordinates of the previous frame: ΔPi(t) = Pi(t) - Pi(t-1), where t represents the frame number. Based on the displacement vector and the inter-frame time interval Δt, the system further calculates the instantaneous velocity of each feature point: Vi(t) = ΔPi(t) / Δt. This velocity information reflects the speed changes of hand movements, providing crucial information for recognizing gestures of varying intensity and rhythm.
[0105] The system employs cubic spline interpolation or Kalman filtering algorithms to fit the time-series sequences of each displacement vector, generating smooth and continuous motion trajectory curves. The trajectory fitting process not only eliminates trajectory jitter caused by sensor noise and detection errors but also extracts geometric features of the trajectory through curve parameters, such as trajectory length, curvature variation, and principal direction. The system extracts multidimensional dynamic features from the fitted trajectory, including statistical features such as total displacement, average velocity, maximum velocity, rate of change of velocity, and trajectory direction consistency. These dynamic trajectory features form feature vectors describing hand movement patterns, effectively distinguishing different hand gesture patterns such as linear motion, circular motion, and reciprocating motion, providing a reliable data foundation for the accurate recognition of complex dynamic gestures.
[0106] S43, combining spatial geometric features and / or dynamic trajectory features into gesture recognition features.
[0107] S105: Based on gesture recognition features, identify the type of gesture performed by the user's hand; combine the gesture type and mapped coordinates to generate interactive instructions.
[0108] The system inputs the gesture recognition features generated in step S104 into a pre-trained gesture classification model for intelligent recognition processing. The gesture classification model is a multi-class classifier built on a deep learning architecture. This model is trained on a large-scale dataset containing tens of thousands of labeled gesture samples, learning a complex mapping relationship from gesture recognition features to gesture types. The model architecture employs a hybrid structure combining a multilayer perceptron and a recurrent neural network. The multilayer perceptron handles spatial geometric features, extracting feature representations of static hand postures; the recurrent neural network specifically handles dynamic trajectory features, capturing the temporal patterns and trajectory characteristics of hand movements. This hybrid architecture fully utilizes the complementary information of different types of features, significantly improving the accuracy and robustness of gesture recognition.
[0109] The gesture classification model operates based on feature matching and pattern recognition principles. Internally, the model pre-stores multiple gesture templates, each containing a standard feature description and variation range for the corresponding gesture type. When a gesture recognition feature is input, the model calculates similarity and performs matching analysis against all pre-defined gesture templates. The matching process is not a simple Euclidean distance calculation, but rather employs advanced similarity metrics such as weighted Mahalanobis distance or cosine similarity, which considers the differences in importance across different feature dimensions and the correlation between features. The model also integrates an uncertainty estimation mechanism, outputting the confidence probability for each gesture type. When the highest confidence level falls below a preset threshold, the system either refuses to recognize the gesture or requires the user to re-perform it.
[0110] The preset gesture templates cover a rich library of gesture types, including basic static gestures such as clenched fist, open palm, OK, victory sign, and thumbs-up; dynamic gestures such as waving, circling, swiping in a straight line, pinching to zoom, and rotating; and compound gestures such as two-handed operation and multi-finger fine control. Each gesture type corresponds to specific interaction semantics. For example, a clenched fist usually indicates selection or confirmation, an open palm indicates cancellation or reset, a wave indicates switching or page turning, and a pinch indicates zooming. The design of the gesture templates fully considers users' natural habits and cultural background, ensuring the intuitiveness and consistency of gesture semantics.
[0111] The gesture type information output by the model includes the semantic identifier and confidence score of the gesture. The semantic identifier is the specific name or code of the gesture, such as "pinch_zoom", "swipe_left", and "palm_open". These identifiers provide clear guidance for subsequent interaction mapping. The confidence score reflects the model's degree of confidence in the recognition result. High-confidence recognition results will directly execute the corresponding interaction operation, while low-confidence results may require additional confirmation mechanisms or be ignored by the system to avoid accidental operation.
[0112] Based on the recognized gesture type, the system determines the corresponding interaction operation category. The interaction operation category serves as an intermediate mapping layer between gesture semantics and specific device functions, transforming abstract gesture meanings into executable operation commands. For example, the "pinch_zoom" gesture type corresponds to the "zoom operation" category, "swipe_left" corresponds to the "switch left" category, and "tap" corresponds to the "selection operation" category. The advantage of this hierarchical mapping mechanism lies in its excellent scalability and adaptability; the same gesture type can be mapped to different specific operations according to different application scenarios, and the same operation category can be triggered by multiple different gestures.
[0113] The system simultaneously determines the effective position and area of the interactive operation on the aerial imaging device's display interface based on the mapped coordinates of hand feature points. The effective position is typically determined based on the mapped coordinates of the primary interaction point (such as the fingertip of the index finger), indicating the specific interface element or screen area the user intends to manipulate. For gestures requiring regional manipulation, such as zooming or rotating, the system calculates the effective area based on the mapped coordinates of multiple hand feature points. This effective area may be a rectangle, a circle, or an irregular polygon. The precise calculation of the effective position and area ensures that the user's gestures are accurately applied to the intended interface content.
[0114] Determining the target location and area also requires considering the hierarchical structure and interaction priority of interface elements. The display interface of an aerial imaging device typically contains multiple levels of interactive elements, such as windows, buttons, menus, and content areas. Different elements have different interaction priorities and response ranges. The system determines the target element for the current operation based on the hierarchical affiliation of the mapped coordinates and adjusts the specific method of the operation response according to the element's interaction attributes. For example, a click in the button area triggers a button event, a swipe in the content area scrolls the content, and a drag in the border area moves the window.
[0115] The system ultimately combines the determined interactive operation category with the precise location and area of effect to generate a complete interactive instruction. An interactive instruction is a structured command object containing complete information such as operation type, coordinates, parameter settings, and a timestamp. The operation type field specifies the specific interactive action, such as "click," "drag," "zoom," and "rotate." The coordinates field records the spatial location information of the operation, including two-dimensional screen coordinates and three-dimensional space coordinates. The parameter settings field contains the specific parameters of the operation, such as zoom level, rotation angle, and movement distance. The timestamp field records the generation time of the instruction, used for subsequent timing control and operation merging.
[0116] The generation of interactive commands also includes command optimization and conflict detection mechanisms. Command optimization analyzes continuous gesture sequences, merging multiple related simple operations into compound operations to improve the smoothness and efficiency of the interaction. For example, continuous small movements are merged into smooth drag trajectories, and rapid repeated clicks are identified as double-click operations. The conflict detection mechanism handles multiple gestures or contradictory operation commands detected simultaneously, ensuring the clarity and consistency of the final output interactive commands through priority rules and time window mechanisms.
[0117] The generated interactive commands are then sent to the aerial imaging device for execution. The aerial imaging device drives corresponding interface response operations based on the received commands. Interface response operations may include visual feedback, content updates, state changes, and function execution, among other things. The timely and accurate execution of these responses provides users with immediate operational feedback, forming a complete interactive loop. The execution latency of interactive commands is typically controlled within tens of milliseconds, ensuring users experience a natural and smooth interactive experience.
[0118] Based on gesture recognition features, the system identifies the type of gesture performed by the user's hand. Combining the gesture type and mapped coordinates, it generates interactive instructions, including: inputting the gesture recognition features into a pre-trained gesture classification model, which matches the spatial geometric features and / or dynamic trajectory features in the gesture recognition features with multiple preset gesture templates to output the type of gesture performed by the user's hand; determining the corresponding interactive operation category based on the gesture type; determining the position and / or area of the interactive operation category on the display interface of the aerial imaging device based on the mapped coordinates; generating interactive instructions by combining the interactive operation category and its position and / or area, and sending the interactive instructions to the aerial imaging device to execute the corresponding interface response operation.
[0119] In the gesture recognition stage, the system inputs the extracted gesture recognition features into a pre-trained gesture classification model for pattern matching and classification. The gesture classification model is a deep learning model trained on a large amount of labeled gesture data, internally storing multiple preset gesture templates, each corresponding to a specific gesture type such as click, swipe, zoom, and rotate. The model analyzes the static posture of the hand based on the input spatial geometric features and the hand's movement pattern through dynamic trajectory features. These features are then compared with the built-in gesture templates to calculate similarity, ultimately outputting the gesture type identifier with the highest confidence. Through this template-matching-based classification method, the system can accurately identify the specific gestures performed by the user.
[0120] During the interaction command generation phase, the system first determines the corresponding interaction operation category based on the recognized gesture type. Different gesture types map to different interface operation functions; for example, a tap with the index finger corresponds to a selection operation, a grasp with the palm corresponds to a drag operation, and a pinch-to-zoom operation corresponds to a pinch-to-zoom operation. Simultaneously, the system determines the specific location or area of the interaction operation on the display interface based on the mapped coordinates of hand feature points. These coordinates directly indicate the interface area or element targeted by the user's gesture. Finally, the system combines the interaction operation category and location information to generate a complete interaction command. This command includes complete information such as the operation type, target location, and operation parameters, and is sent to the aerial imaging device to execute the corresponding interface response, achieving precise aerial touch interaction control.
[0121] Figure 3This application provides an embodiment of a gesture recognition and interaction control system architecture based on the Leapmotion device, illustrating the complete data flow processing chain from depth information acquisition to final application response. The entire system uses the Leapmotion small depth acquisition device as the data source, which is responsible for acquiring and imaging depth information within its range. This raw depth data is then transmitted to the host program for further processing. In the host software data processing module, the system performs a series of complex data processing operations on the received depth information, including key steps such as data filtering, calibration, and preprocessing. These processing operations aim to improve data quality and prepare for subsequent recognition and analysis. The preprocessed data enters the recognition and analysis module, the core component of the entire system, responsible for in-depth analysis and processing of the depth data. Specifically, this includes static gesture posture recognition, dynamic gesture trajectory analysis, and recognition of complex interaction logic such as multi-finger operations. Advanced algorithm models are used to convert the raw depth information into understandable gesture semantics. The results of the recognition and analysis are ultimately passed to the application layer. The application layer transforms the recognized gesture information into standardized interactive operations and API calls, and controls the imaging content to respond to the user's gestures according to predefined mapping rules. This achieves a seamless transition from physical gesture actions to digital interface operations, forming a complete non-contact gesture interaction control system.
[0122] Based on the above method, this application also discloses a gesture interaction control system for aerial imaging equipment, such as... Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of a gesture interaction control system for an aerial imaging device provided in an embodiment of this application. The system includes: an acquisition module, a reconstruction module, a mapping module, a combination module, and an output module; wherein, The system comprises the following modules: an acquisition module for acquiring infrared images of the user's hand captured by sensors below the imaging area of the aerial imaging device; a reconstruction module for reconstructing a 3D skeletal model of the user's hand based on the infrared images, and determining the 3D spatial coordinates of multiple hand feature points in the sensor coordinate system based on the 3D skeletal model; a mapping module for transforming the 3D spatial coordinates of each hand feature point to the display coordinate system according to a pre-calibrated mapping relationship between the sensor coordinate system and the display coordinate system of the aerial imaging device, generating mapped coordinates for each hand feature point; a combination module for combining the 3D spatial coordinates of each hand feature point with the temporal changes of each hand feature point over multiple consecutive frames to generate gesture recognition features, including spatial geometric features representing the static posture of the hand and / or dynamic trajectory features representing the hand movement process; and an output module for identifying the type of gesture performed by the user's hand based on the gesture recognition features and generating interactive commands by combining the gesture type and the mapped coordinates.
[0123] It should be noted that the system provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0124] Please see Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.
[0125] The communication bus 1002 is used to realize the connection and communication between these components.
[0126] The user interface 1003 may include a display screen and a camera. Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface.
[0127] The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0128] The processor 1001 may include one or more processing cores. The processor 1001 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005. Optionally, the processor 1001 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 1001 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 1001 and may be implemented as a separate chip.
[0129] The memory 1005 may include random access memory (RAM) or read-only memory. Optionally, the memory 1005 may include a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 5 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a gesture interaction control method for an aerial imaging device.
[0130] exist Figure 5In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 1001 can be used to call an application program stored in the memory 1005 for a gesture interaction control method of an aerial imaging device. When executed by one or more processors, the electronic device performs one or more of the methods described in the above embodiments.
[0131] An electronic device readable storage medium stores instructions that, when executed by one or more processors, cause the electronic device to perform one or more of the methods described in the above embodiments.
[0132] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0133] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0134] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some service interfaces; indirect couplings or communication connections between devices or units may be electrical or other forms.
[0135] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0136] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0137] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0138] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A gesture-based interactive control method for an aerial imaging device, characterized in that, The method includes: Acquire infrared images of the user's hand captured by sensors below the imaging area of the aerial imaging device; Based on the infrared image, a three-dimensional skeletal model of the user's hand is reconstructed, and based on the three-dimensional skeletal model, the three-dimensional spatial coordinates of multiple hand feature points in the sensor coordinate system are determined. Based on the pre-calibrated mapping relationship between the sensor coordinate system and the display coordinate system of the aerial imaging device, the three-dimensional spatial coordinates of each hand feature point are transformed to the display coordinate system, generating the mapped coordinates of each hand feature point; By combining the three-dimensional spatial coordinates of each hand feature point and the temporal changes of each hand feature point across multiple consecutive frames, gesture recognition features are generated. The gesture recognition features include spatial geometric features representing the static posture of the hand and / or dynamic trajectory features representing the hand movement process. Based on the gesture recognition features, the type of gesture performed by the user's hand is identified; and by combining the gesture type and the mapped coordinates, an interaction command is generated.
2. The gesture interaction control method for an aerial imaging device according to claim 1, characterized in that, The step of reconstructing a three-dimensional skeletal model of the user's hand based on the infrared image includes: The infrared image is subjected to hand detection and region segmentation to generate a hand region image; Feature extraction is performed on the hand region image to obtain the appearance and depth features of the hand; The appearance features and depth features are input into a pre-trained hand pose estimation model to generate the three-dimensional positions of each joint of the user's hand. The three-dimensional positions of each joint and the preset topological connection relationship of the hand skeleton are combined to construct a three-dimensional skeleton model.
3. The gesture interaction control method for an aerial imaging device according to claim 1, characterized in that, The step of determining the three-dimensional spatial coordinates of multiple hand feature points in the sensor coordinate system based on the three-dimensional skeletal model includes: In the sensor coordinate system, based on the three-dimensional skeletal model, the three-dimensional spatial coordinates of the fingertips and interphalangeal joints of each finger of the user's hand are extracted; Based on the positions of the wrist joints and the root joints of each finger in the three-dimensional skeletal model, calculate the three-dimensional spatial coordinates of the center point of the palm. The fingertip, each of the interphalangeal joints, and the center point of the palm are identified as multiple hand feature points.
4. The gesture interaction control method for an aerial imaging device according to claim 1, characterized in that, The step of transforming the three-dimensional spatial coordinates of each of the hand feature points to the display coordinate system, and generating the mapped coordinates of each of the hand feature points, includes: A pre-calibrated coordinate transformation matrix is obtained, which is obtained by fitting the corresponding coordinates of multiple calibration reference points in the sensor coordinate system and the display coordinate system respectively; Using the coordinate transformation matrix, the three-dimensional spatial coordinates of each hand feature point in the sensor coordinate system are transformed to generate the three-dimensional display coordinates of each hand feature point in the display coordinate system. Based on the projection position of the three-dimensional display coordinates of each hand feature point on the imaging plane of the aerial imaging device, the mapped coordinates of each hand feature point are generated.
5. The gesture interaction control method for an aerial imaging device according to claim 1, characterized in that, The step of generating gesture recognition features by combining the three-dimensional spatial coordinates of each hand feature point and the temporal changes of each hand feature point across multiple consecutive frames includes: Based on the three-dimensional spatial coordinates of each hand feature point in the same frame, calculate the relative distance between preset pairs of hand feature points and the angle between the lines connecting each finger bone to generate spatial geometric features representing the static posture of the hand. The three-dimensional spatial coordinates of each hand feature point in multiple consecutive frames are obtained, the displacement vector and motion velocity of each hand feature point between adjacent frames are calculated, and the motion trajectory of each hand feature point is fitted according to the temporal sequence of each displacement vector to generate dynamic trajectory features characterizing the hand movement process. The spatial geometric features and / or the dynamic trajectory features are combined into gesture recognition features.
6. The gesture interaction control method for an aerial imaging device according to claim 5, characterized in that, The step involves calculating the relative distance between preset pairs of hand feature points and the angle between the lines connecting the finger bones based on the three-dimensional spatial coordinates of each hand feature point in the same frame, thereby generating spatial geometric features representing the static posture of the hand, including: Based on the three-dimensional spatial coordinates of the center point of the palm and the three-dimensional spatial coordinates of each fingertip, calculate the relative distance between each fingertip and the center point of the palm. Calculate the angle between adjacent bone lines based on the direction vector of the line connecting adjacent joints of the same finger, and use it as the bending angle of each finger. Calculate the opening angle between adjacent fingers based on the direction vector of the line connecting the bones of adjacent fingers; The hand reference scale is determined based on the distance between the center point of the palm and the wrist joint. The relative distances are then normalized using the hand reference scale to generate normalized distance features. The normalized distance feature, the bending angle of each finger, and the opening angle between each adjacent finger are combined to form the spatial geometric feature.
7. The gesture interaction control method for an aerial imaging device according to claim 1, characterized in that, The gesture type performed by the user's hand is identified based on the gesture recognition features. Combining the gesture type and the mapped coordinates, an interaction command is generated, including: The gesture recognition features are input into a pre-trained gesture classification model, which then matches the spatial geometric features and / or dynamic trajectory features in the gesture recognition features with multiple preset gesture templates to output the type of gesture performed by the user's hand. Based on the gesture type, determine the corresponding interactive operation category; based on the mapping coordinates, determine the position and / or area of action of the interactive operation category in the display interface of the aerial imaging device. Based on the interaction operation category, the location of action, and / or the area of action, an interaction command is generated and sent to the aerial imaging device to execute the corresponding interface response operation.
8. A gesture interaction control system for an aerial imaging device, characterized in that, The system includes: an acquisition module, a reconstruction module, a mapping module, a combination module, and an output module; wherein... The acquisition module is used to acquire infrared images of the user's hand collected by sensors below the imaging area of the aerial imaging device; The reconstruction module is used to reconstruct a three-dimensional skeletal model of the user's hand based on the infrared image, and to determine the three-dimensional spatial coordinates of multiple hand feature points in the sensor coordinate system based on the three-dimensional skeletal model. The mapping module is used to transform the three-dimensional spatial coordinates of each hand feature point to the display coordinate system according to the pre-calibrated mapping relationship between the sensor coordinate system and the display coordinate system of the aerial imaging device, and generate the mapped coordinates of each hand feature point. The combining module is used to combine the three-dimensional spatial coordinates of each of the hand feature points and the temporal changes of each of the hand feature points in consecutive frames to generate gesture recognition features. The gesture recognition features include spatial geometric features representing the static posture of the hand and / or dynamic trajectory features representing the hand movement process. The output module is used to identify the type of gesture performed by the user's hand based on the gesture recognition features; and to generate an interaction command by combining the gesture type and the mapped coordinates.
9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1-7.