Interactive positioning system and method based on real-time high-precision detection of hand key points

By using an interactive positioning system that works in collaboration with multiple devices, combined with real-time flow limiting and super-resolution optimization, the problem of insufficient recognition efficiency and accuracy under a single media stream mode is solved, and real-time high-precision detection and three-dimensional positioning of key hand points are achieved.

CN120656243BActive Publication Date: 2026-03-31NINGBO THREDIM OPTOELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies for real-time high-precision detection of key points on the hand, the single media stream method limits the recognition efficiency and the accuracy of the recognition results depends on the resolution of the image acquisition device, resulting in excessively high requirements for the equipment when high precision is needed.

Method used

Multiple image acquisition devices are used, combined with a real-time flow limiting module, a palm detection module, a super-resolution module, a key point detection module, and a coordinate calculation module. Data is transmitted through different numbers of transmission channels. Global detection and region detection are combined, a lightweight multi-path feature calibration network is used for resolution optimization, and three-dimensional coordinates are calculated through a binocular camera.

Benefits of technology

It improves detection efficiency and accuracy, ensures real-time and high-precision interaction, and meets the real-time and high-precision requirements of gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656243B_ABST
    Figure CN120656243B_ABST
Patent Text Reader

Abstract

The application provides an interactive positioning system and method based on real-time high-precision detection of hand key points, and the system comprises an image acquisition device, a real-time flow limiting module, a palm detection module, a super-resolution module, a key point detection module and a coordinate calculation module. The application uses different numbers of transmission channels for data transmission at different stages through the real-time flow limiting module, thereby improving the real-time performance of data transmission. The palm detection module uses a combination of global detection and regional detection to improve the detection efficiency. The super-resolution module is used to optimize the resolution of the image, thereby improving the high precision of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of interactive positioning technology, and more specifically, to an interactive positioning system and method based on real-time high-precision detection of hand key points. Background Technology

[0002] In computer science, hand key detection is the foundation of gesture recognition; real-time high-precision positioning is the basis for XR (Extended Reality) to achieve natural interaction between the real and virtual worlds. Interactive positioning is a topic that uses mathematical algorithms to recognize human gestures. Users can use simple gestures to control or interact with devices, allowing computers to understand human behavior.

[0003] Current technologies for interactive positioning often employ a single media stream approach. While this reduces resource consumption, it limits recognition efficiency. Furthermore, the accuracy of this method depends entirely on the resolution of the image acquisition device, placing high demands on such equipment when high-precision results are required. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide an interactive positioning system and method based on real-time high-precision detection of hand key points, so as to overcome the problems in the prior art.

[0005] In a first aspect, embodiments of this application provide an interactive positioning system based on real-time high-precision detection of hand key points, including multiple image acquisition devices deployed at preset positions for acquiring initial images of the shooting area;

[0006] A real-time flow limiting module is connected to the image acquisition device and initially sends the initial image to the palm detection module through a first number of transmission channels.

[0007] A palm detection module, connected to the real-time rate limiting module, detects the initial image based on the palm detection method to obtain a detection result. When the detection result meets a preset detection condition, a second number of transmission channels are activated to receive the initial image; the second number is greater than the first number.

[0008] The hand detection method is used to detect the region to be detected in the initial image in the first transmission channel to obtain a hand detection box; wherein, the region to be detected is predicted based on the actual position of the hand in the detection result; the first transmission channel is a portion of the second number of transmission channels;

[0009] A super-resolution module, connected to the palm detection module, is used to optimize the resolution based on a first cropped image and a second cropped image to obtain an optimized target image; wherein, the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection box, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection box; the second transmission channel is a channel other than the first transmission channel in the second number of transmission channels;

[0010] A key point detection module, connected to the super-resolution module, is used to extract the two-dimensional coordinates of key points of the hand skeleton from the target image;

[0011] The coordinate calculation module, connected to the key point detection module, is used to convert the two-dimensional coordinates into three-dimensional coordinates, which are used for virtual interaction.

[0012] In some technical solutions of this application, the above-mentioned hand detection method is used to detect the initial image to obtain detection results, including:

[0013] The initial image is detected using a preset target detection model to obtain the number of hands contained in the initial image and the actual location of each hand.

[0014] In some technical solutions of this application, the above-mentioned method is used to determine whether the detection result meets the preset detection conditions:

[0015] The number of hands included in the detection results is greater than or equal to a preset threshold, and the detection results meet the preset detection conditions.

[0016] The number of hands included in the detection results is less than a preset threshold, and the detection results meet the non-preset detection conditions.

[0017] In some technical solutions of this application, the above-mentioned palm detection method is used to detect the region to be detected in the initial image in the first transmission channel based on the detection result, to obtain a palm detection box, including:

[0018] Based on the positional relationship between the various image acquisition devices and the actual position, the movement of the hand is predicted to obtain the area to be detected;

[0019] The palm detection method is used to detect the area to be detected, and a palm detection box containing the palm is obtained in the area to be detected.

[0020] In some technical solutions of this application, the above-mentioned resolution optimization based on the first cropped image and the second cropped image to obtain the optimized target image includes:

[0021] The first cropped image and the second cropped image are input into a preset lightweight multi-path feature calibration network to obtain the target image output by the lightweight multi-path feature calibration network.

[0022] In some technical solutions of this application, the target image is two images, and the conversion of the two-dimensional coordinates into three-dimensional coordinates includes:

[0023] Based on the optical axis distance and focal length between the image acquisition devices, the two-dimensional coordinates are transformed into three-dimensional coordinates.

[0024] In some technical solutions of this application, the extraction of two-dimensional coordinates of key hand bone points from the target image includes:

[0025] The target image is input into a high-resolution network to obtain the two-dimensional coordinates of the key points of the hand skeleton output by the high-resolution network.

[0026] Secondly, embodiments of this application provide an interactive positioning method based on real-time high-precision detection of hand key points, the method comprising:

[0027] Initial images of the shooting area are acquired using image acquisition devices at multiple preset locations;

[0028] Initially, the real-time flow limiting module sends the initial image to the palm detection module through a first number of transmission channels;

[0029] The initial image is detected using a palm detection module based on a palm detection method to obtain detection results. When the detection results meet preset detection conditions, a second number of transmission channels are activated to receive the initial image; the second number is greater than the first number.

[0030] The hand detection method is used to detect the region to be detected in the initial image in the first transmission channel to obtain a hand detection box; wherein, the region to be detected is predicted based on the actual position of the hand in the detection result; the first transmission channel is a portion of the second number of transmission channels;

[0031] The super-resolution module is used to optimize the resolution based on the first cropped image and the second cropped image to obtain the optimized target image; wherein, the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection box, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection box; the second transmission channel is a channel other than the first transmission channel in the second number of transmission channels.

[0032] The key point detection module extracts the two-dimensional coordinates of key points of the hand bones from the target image;

[0033] The coordinate calculation module converts the two-dimensional coordinates into three-dimensional coordinates, which are then used for virtual interaction.

[0034] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the interactive positioning method based on real-time high-precision detection of hand key points described above are performed.

[0035] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described interactive positioning method based on real-time high-precision detection of hand key points.

[0036] The technical solutions provided by the embodiments of this application may include the following beneficial effects:

[0037] This system includes multiple image acquisition devices deployed at preset locations to acquire initial images of the shooting area; a real-time rate limiting module connected to the image acquisition devices, which initially sends the initial images to the hand detection module through a first number of transmission channels; the hand detection module connected to the real-time rate limiting module, which detects the initial images based on a hand detection method, obtains detection results, and activates a second number of transmission channels to receive the initial images when the detection results meet preset detection conditions; the second number is greater than the first number; the system uses the hand detection method to detect the area to be detected in the initial images in the first transmission channel, obtaining a hand detection bounding box; wherein, the area to be detected is predicted based on the actual position of the hand in the detection results; The first transmission channel is a portion of the second number of transmission channels; the super-resolution module, connected to the palm detection module, is used to optimize the resolution based on the first cropped image and the second cropped image to obtain an optimized target image; wherein, the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection box, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection box; the second transmission channel is a channel other than the first channel in the second number of transmission channels; the key point detection module, connected to the super-resolution module, is used to extract the two-dimensional coordinates of key points of the hand skeleton from the target image; the coordinate calculation module is connected to the key point detection module.

[0038] This application effectively saves transmission resources by using a real-time rate limiting module to transmit data through different numbers of transmission channels at different stages; the hand detection module uses a combination of global detection and regional detection to improve detection efficiency; the super-resolution module optimizes the image resolution to improve the accuracy of interaction; and the combination of all modules ensures the real-time nature of the interaction.

[0039] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This illustration shows a schematic diagram of an interactive positioning system based on real-time high-precision detection of hand key points provided in an embodiment of this application;

[0042] Figure 2 This illustration shows a schematic diagram of skeletal key points provided in an embodiment of this application;

[0043] Figure 3 This illustration shows a schematic diagram of the relationship between pixel coordinates and world coordinates of a binocular camera according to an embodiment of this application;

[0044] Figure 4 This illustration shows a schematic diagram of a specific implementation of an interactive positioning system based on real-time high-precision detection of hand key points provided in this application.

[0045] Figure 5 A flowchart illustrating an interactive positioning method based on real-time high-precision detection of hand key points provided in an embodiment of this application is shown.

[0046] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0048] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0049] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0050] In computer science, interactive localization based on real-time, high-precision detection of hand key points is a topic that uses mathematical algorithms to recognize human gestures. Users can use simple gestures to control or interact with devices, allowing computers to understand human behavior.

[0051] Existing technologies for interactive positioning based on real-time, high-precision hand landmark detection often employ a single media stream approach. While this reduces resource consumption, it limits recognition efficiency. Furthermore, the accuracy of this method depends entirely on the resolution of the image acquisition device, placing high demands on such equipment when high-precision results are required.

[0052] Based on this, this application provides an interactive positioning system and method based on real-time high-precision detection of hand key points, which is described below through embodiments. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0053] Figure 1This illustration shows a schematic diagram of an interactive positioning system based on real-time high-precision detection of hand key points, according to an embodiment of this application. The interactive positioning system includes an image acquisition device, a real-time rate limiting module, a palm detection module, a super-resolution module, a key point detection module, and a coordinate calculation module. The image acquisition device is connected to the real-time rate limiting module, which is connected to the palm detection module, which is connected to the super-resolution module, which is connected to the key point detection module, which is connected to the coordinate calculation module, and the coordinate calculation module is connected to an external interactive device.

[0054] In this embodiment, the number of image acquisition devices (cameras, mobile phones, webcams, etc.) is multiple, such as four, nine, fifteen, or twenty-two. They are deployed at preset locations, and their relative positions remain unchanged after deployment. For example, four image acquisition devices (A1, A2, A3, A4) are fixed in pairs, with A1 adjacent to A2 and A3 adjacent to A4. The distance between A1 and A3 is x1, and the distance between A2 and A4 is also x1. The acquisition devices capture images towards a preset shooting area (the area where the hand appears) (based on a preset acquisition frequency; the specific acquisition frequency can be set according to requirements, and this embodiment does not limit the specific acquisition frequency) to obtain an initial image. The initial image includes a background or a background and a hand. After acquiring the initial image, the image acquisition devices send it to the real-time rate limiting module.

[0055] The real-time rate limiting module transmits the initial image to the hand detection module through a first number of transmission channels. In this embodiment, multiple transmission channels (e.g., represented by X) are set between the real-time rate limiting module and the hand detection module, but the number of transmission channels opened varies at different stages. For example, initially (when the initial image has not been detected or the detection result of the initial image does not meet the preset detection conditions), the real-time rate limiting module and the hand detection module open a first number of transmission channels (here, the first number is less than X). For example, if the first number is two, the initial image can only be transmitted through these two transmission channels. Or, if the first number is five, the initial image can only be transmitted through these five transmission channels. Or, if the first number is nine, the initial image can only be transmitted through these nine transmission channels.

[0056] After the real-time rate limiting module sends the initial image to the palm detection module, the palm detection module performs detection on the initial image. To improve detection efficiency, this embodiment uses a palm detection method, rather than a hand detection method, for the palm detection module. The hand is located distal to the wrist and is the distal structure of the entire upper limb. The hand has a complex and delicate structure, including the wrist, palm, and fingers. The wrist connects to the forearm and is composed of many small bones, connected to the forearm by ligaments and muscles; the palm supports the fingers and contains many muscles and nerves; the fingers are composed of phalanges, joints, muscles, and skin, enabling grasping and pinching actions. The palm is part of the hand.

[0057] The initial image is then processed using a hand detection method to obtain the detection results. It's important to emphasize that this detection process checks all content in the initial image, including background detection and detection of both the background and hands. The detection results include the number of hands and their corresponding actual locations. For example, the initial image may show no hands (zero count). Alternatively, it may show one hand and its location.

[0058] In practical implementation, the hand detection method uses the SSD (Single Shot Multibox Detector) object detection model. SSD is a single-stage object detection algorithm that extracts features through a convolutional neural network and outputs detection at different feature layers to achieve multi-scale detection. It employs an anchor (bounding box) strategy, pre-setting anchors with different aspect ratios and predicting their positions (represented by hand detection boxes) on each output feature layer. The SSD framework includes multi-scale detection methods, with shallow layers for detecting small objects and deep layers for detecting large objects. Hand detection is a very complex task because hands come in many sizes and are subject to occlusion and self-occlusion. Unlike face detection, which is a high-contrast pattern (mouth, eyes, etc., can assist in face detection), the hand is a dynamic and complex pattern, making it difficult to predict the hand based solely on visual features. To address this issue, this embodiment adopts the following strategy: first, a detector for the palm is trained, instead of a detector for the entire hand. Because the palm is a square, rigid-body-like object, it doesn't move as flexibly as the fingers, making it a relatively stable feature. Furthermore, the palm is smaller, so non-maximum suppression algorithms perform better when two hands are shaking hands or when there is self-occlusion. Additionally, the square shape of the palm eliminates the need to consider other aspect ratios, reducing unnecessary anchor boxes. Next, an encoder-decoder is used for feature extraction.

[0059] After obtaining the detection results, the hand detection module needs to compare these results with preset detection conditions to determine whether the results meet the conditions. These detection conditions can be preset quantity thresholds: if the number of hands in the detection results is greater than or equal to the preset quantity threshold, the detection results meet the preset detection conditions; if the number of hands in the detection results is less than the preset quantity threshold, the detection results do not meet the preset detection conditions. For example, if the quantity threshold is set to three, if the initial image does not contain any hands, then the initial image does not meet the preset detection conditions; if the initial image contains four hands, then the initial image meets the preset detection conditions.

[0060] When the preset detection conditions are met, this embodiment of the application considers that subsequent initial images need to be prioritized for detection. Therefore, to improve detection efficiency, this embodiment of the application expands the original first number of transmission channels to a second number. Then, the real-time rate limiting module transmits images to the hand detection module based on the second number of transmission channels.

[0061] On the other hand, since the initial image has already been identified as containing a preset threshold number of hands during the aforementioned detection process, and due to the interactive nature of the hand being in motion, subsequent initial images will also contain hands of a number greater than or equal to the threshold. Furthermore, based on the actual position of the detected hands, their motion attributes, and the relative position between the image acquisition devices, the position of the hands in subsequent initial images can be predicted, thus obtaining the positions where the hands will appear in each subsequent initial image. The specific prediction method can use existing prediction methods in the prior art; this application embodiment does not limit the specific prediction method. After obtaining the position where the hands will appear, this application embodiment determines a region to be detected based on the predicted position. To improve detection efficiency, the region to be detected can be detected directly, rather than performing global detection on the initial image.

[0062] In specific implementation, this application adopts a combination of global detection and ROI (Region of Interest) detection. If no hand is detected or the number of detected hands is less than a preset threshold (e.g., 2), the global detection mode is used. When a hand is detected or the preset threshold of hands is detected, the detection area to be detected is predicted based on the position of the hand obtained from global detection (which can be represented by a bounding box or ROI, and the detected hand detection box is passed to other transmission channels), and then ROI detection is performed on the area to be detected.

[0063] Furthermore, after acquiring the initial images from the second number of transmission channels, the hand detection module, considering image resolution issues, does not detect all initial images in all of the second number of transmission channels. Instead, it divides the second number of transmission channels into a first transmission channel and a second transmission channel. The first transmission channel comprises a portion of the second number of transmission channels; the second transmission channel comprises all transmission channels except the first transmission channel. That is, in this embodiment, the hand detection module only detects the region to be detected in the initial images from the first transmission channel, obtains a hand detection bounding box, and sends this bounding box to the super-resolution module, while the initial images from the second transmission channel are directly transmitted to the super-resolution module.

[0064] To improve processing efficiency and accuracy, the super-resolution module does not process the entire image, but first performs image cropping: a first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection box, and a second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection box. Then, the resolution of the first cropped image and the second cropped image is optimized to obtain the optimized target image.

[0065] In optimizing resolution, this application employs a lightweight multi-path feature calibration network (MPFCN). The MPFCN aims to achieve a good balance between parameters and performance, realizing excellent super-resolution while maintaining real-time detection capabilities. It models context dependencies across multiple scales and channels, fully mining the spatial information and channel features of the image to obtain more accurate results in image reconstruction.

[0066] After obtaining the target image, this embodiment of the application extracts coordinates using a key point detection module, thereby obtaining the two-dimensional coordinates of key points of the hand bones. For example... Figure 2 As shown, after deep learning algorithm, the information of 21 skeletal key points of the hand is given (2D / 2.5D): the 21 key points of the hand correspond to the 21 joints of the human hand, one joint at the base of the palm, four joints of the thumb, four joints of the index finger, four joints of the middle finger, four joints of the ring finger, and four joints of the little finger. The base of the palm is defined as joint number 0, and the joint numbers of each finger increase upwards from the base of the palm, with the joint numbers of the thumb to the little finger being 1~4, 5~8, 9~12, 13~16, and 17~20 respectively. Figure 2The joint numbers given are abbreviated, only showing the numbers for the palm root and fingertip joints. The z-axis data is based on the rule of larger values ​​for farther objects and smaller values ​​for closer ones, with the palm root node (node ​​0) as the base point. For example, hand keypoint detection is an improved implementation based on the open-source HRNet (High Resolution Network), constructing a complete training and testing implementation for hand skeletal keypoint detection. The network connects convolutional streams in parallel from high to low resolution. It maintains high-resolution representation throughout the process and generates reliable high-resolution representations with strong position sensitivity by repeatedly fusing representations from multi-resolution streams. To ensure the real-time performance and high accuracy of the detection system, improvements are made in two aspects: first, the input to HRNet is not the entire image, but rather the predicted detection boxes output from the palm detection module, which greatly improves HRNet's detection speed; second, detection boxes are cropped from images from multiple channels, and super-resolution processing is performed on the cropped distributed images to improve the resolution of the input image, thereby improving the localization accuracy of keypoints. It should be noted that the hand detection box ROI is adjusted based on the detected keypoint coordinates (2D) output to achieve accurate tracking.

[0067] After obtaining the two-dimensional coordinates, the coordinate calculation module transforms them into three-dimensional coordinates. The main task of the coordinate calculation module is to realize the three-dimensional coordinate calculation of the hand skeleton keypoints detected in two target images. Although the estimation of the 3D z-axis of the hand skeleton keypoints can be achieved through learning, this accuracy cannot meet the requirements of high-precision positioning. Therefore, this embodiment combines the principle of binocular camera positioning to realize the calculation of the 3D coordinates of the keypoints.

[0068] The process of converting 2D keypoint coordinates (pixel coordinates) to 3D keypoint coordinates (world coordinates) using a stereo camera. For example... Figure 3 As shown (Z) L Z R The Z-axis coordinates of the left and right hand keypoints are respectively (the Z-coordinates are achieved through binocular vision localization). Ideally, with the y-axis of the two sets of cameras perfectly aligned, the optical axis distance between the two cameras is b, and the focal lengths of the two cameras are equal, both being f. Given the pixel coordinates of the hand keypoints (PL(uXL, vYL), PR(uXR, vYR)), the world coordinates P(xW, yW, zW) of the corresponding keypoints are calculated as follows:

[0069]

[0070] Where d is the phase difference, i.e., d = ux - uR.

[0071]

[0072]

[0073] Where u0 and v0 are the origin of the camera image. Let XW=X / W, YW=Y / W, ZW=Z / W, the 2D coordinates (pixel coordinates) to 3D coordinates (world coordinates) of the keypoints of the stereo camera can be calculated using the following formula:

[0074]

[0075] Achieving 3D coordinate localization of key points on the hand skeleton provides a crucial element for skeleton-based motion recognition. By identifying the currently executed action from a series of temporally continuous skeletal key points (2D / 3D), gesture interaction functionality is achieved. Through the detection of 21 key points on the hand skeleton and the real-time accurate localization of their spatial coordinates (world coordinates), a fusion interaction between the virtual and real worlds is realized.

[0076] The embodiments of this application effectively solve the problem of simultaneously meeting the requirements of real-time performance, high precision, and three-dimensional positioning for hand key point detection.

[0077] In an optional implementation, in specific implementation, the system in the embodiments of this application can be implemented according to... Figure 4 The process described is as follows: The image acquisition device uses cameras, specifically four cameras: Camera 1, Camera 2, Camera 3, and Camera 4. Initial images captured by the four cameras (Image 1 from Camera 1, Image 2 from Camera 2, Image 3 from Camera 3, and Image 4 from Camera 4) are sent to the real-time rate limiting module. Initially, the real-time rate limiting module sends the initial images from the four cameras to the hand detection module through a single transmission channel. The hand detection module performs global detection on the initial images. When the hand detection module detects two hands in the initial image, the real-time rate limiting module transmits the initial image to the hand detection module through four transmission channels (one channel for each camera: the first, second, third, and fourth transmission channels). Based on the hand detection bounding boxes detected during global detection, the hand detection module predicts the hand regions in subsequent initial images to obtain the regions to be detected. Then, the hand detection module detects the regions to be detected in the initial images transmitted through the second and third transmission channels, obtaining hand detection bounding boxes, and simultaneously sends these bounding boxes to the first and fourth transmission channels. Then, the super-resolution module performs image cropping based on the hand detection bounding box, obtaining cropped images for each transmission channel. The cropped images from the first and second transmission channels are then optimized for resolution to obtain a single target image; similarly, the cropped images from the third and fourth transmission channels are optimized for resolution to obtain another target image. After obtaining the coordinates of the two target images, a coordinate transformation is performed based on the principle of a stereo camera. This yields three-dimensional coordinates, enabling virtual interaction.

[0078] Figure 5 The diagram illustrates a flowchart of an interactive positioning method based on real-time high-precision detection of hand key points according to an embodiment of this application. The method includes steps S101-S107; specifically:

[0079] S101. Acquire initial images of the shooting area using image acquisition devices at multiple preset locations;

[0080] S102. Initially, the real-time flow limiting module sends the initial image to the palm detection module through a first number of transmission channels.

[0081] S103. The initial image is detected using a palm detection module based on a palm detection method to obtain a detection result. When the detection result meets a preset detection condition, a second number of transmission channels are activated to receive the initial image; the second number is greater than the first number.

[0082] S104. The hand detection method is used to detect the region to be detected in the initial image in the first transmission channel to obtain a hand detection box; wherein, the region to be detected is predicted based on the actual position of the hand in the detection result; the first transmission channel is a portion of the second number of transmission channels;

[0083] S105. Using a super-resolution module, the resolution is optimized based on the first cropped image and the second cropped image to obtain an optimized target image; wherein, the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection box, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection box; the second transmission channel is a channel other than the first transmission channel in the second number of transmission channels.

[0084] S106. Extract the two-dimensional coordinates of key points of the hand bones from the target image using the key point detection module;

[0085] S107. The two-dimensional coordinates are converted into three-dimensional coordinates through the coordinate calculation module, and the three-dimensional coordinates are used for virtual interaction.

[0086] The initial image is then analyzed using a hand detection method to obtain detection results, including:

[0087] The initial image is detected using a preset target detection model to obtain the number of hands contained in the initial image and the actual location of each hand.

[0088] The method determines whether the detection result meets the preset detection conditions in the following manner:

[0089] The number of hands included in the detection results is greater than or equal to a preset threshold, and the detection results meet the preset detection conditions.

[0090] The number of hands included in the detection results is less than a preset threshold, and the detection results meet the non-preset detection conditions.

[0091] The step of using a palm detection method to detect the region to be detected in the initial image in the first transmission channel based on the detection results to obtain a palm detection box includes:

[0092] Based on the positional relationship between the various image acquisition devices and the actual position, the movement of the hand is predicted to obtain the area to be detected;

[0093] The palm detection method is used to detect the area to be detected, and a palm detection box containing the palm is obtained in the area to be detected.

[0094] The resolution optimization based on the first cropped image and the second cropped image to obtain the optimized target image includes:

[0095] The first cropped image and the second cropped image are input into a preset lightweight multi-path feature calibration network to obtain the target image output by the lightweight multi-path feature calibration network.

[0096] The target image consists of two images, and the process of converting the two-dimensional coordinates into three-dimensional coordinates includes:

[0097] Based on the optical axis distance and focal length between the image acquisition devices, the two-dimensional coordinates are transformed into three-dimensional coordinates.

[0098] The extraction of two-dimensional coordinates of key hand bone points from the target image includes:

[0099] The target image is input into a high-resolution network to obtain the two-dimensional coordinates of the key points of the hand skeleton output by the high-resolution network.

[0100] like Figure 6 As shown, this application provides an electronic device for executing the interactive positioning method based on real-time high-precision detection of hand key points in this application. The device includes a memory, a processor, a bus, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the interactive positioning method based on real-time high-precision detection of hand key points.

[0101] Specifically, the aforementioned memory and processor can be general-purpose memory and processor, without any specific limitations. When the processor runs the computer program stored in the memory, it can execute the aforementioned interactive positioning method based on real-time high-precision detection of hand key points.

[0102] Corresponding to the interactive positioning method based on real-time high-precision detection of hand key points in this application, this application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the steps of the above-described interactive positioning method based on real-time high-precision detection of hand key points.

[0103] Specifically, the storage medium can be a general-purpose storage medium, such as a portable disk or hard disk. When the computer program on the storage medium is run, it can execute the above-mentioned interactive positioning method based on real-time high-precision detection of hand key points.

[0104] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0105] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0106] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0107] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0108] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0109] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. An interactive positioning system based on real-time high-precision detection of hand key points, characterized in that, The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points.

2. The interactive positioning system based on hand key point real-time high-precision detection according to claim 1, characterized in that, The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points.

3. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1, characterized in that, The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points.

4. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1, characterized in that, The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points.

5. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1, characterized in that, The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real-time high-precision detection of hand key points. The application relates to an interactive positioning system based on real Input the first cropped image and the second cropped image into a preset lightweight multi-path feature calibration network to obtain a target image output by the lightweight multi-path feature calibration network.

6. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 5, characterized in that, The target image is two, the two-dimensional coordinates are converted into three-dimensional coordinates, including: According to the optical axis distance and focal length between the image acquisition devices, the two-dimensional coordinates are converted into three-dimensional coordinates.

7. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1, characterized in that, The two-dimensional coordinates of the hand skeleton key points are extracted from the target image, including: Input the target image into a high-resolution network to obtain the two-dimensional coordinates of the hand skeleton key points output by the high-resolution network.

8. An interactive positioning method based on real-time high-precision detection of hand key points, characterized in that, The interactive positioning system based on real-time high-precision detection of hand key points according to any one of claims 1-7, the interactive positioning method based on real-time high-precision detection of hand key points comprises: Collecting initial images of a shooting area through a plurality of image acquisition devices at preset positions; Initially, control the real-time current limiting module to send the initial image to the palm detection module through a first number of transmission channels; Using the palm detection module to detect the initial image based on the palm detection method to obtain a detection result, when the detection result meets the preset detection condition, activating a second number of transmission channels to receive the initial image; the second number is greater than the first number; Using the palm detection method to detect the to-be-detected region of the initial image in the first transmission channel to obtain a palm detection frame; wherein the to-be-detected region is predicted based on the real position of the palm in the detection result; the first transmission channel is part of the transmission channels in the second number of transmission channels; Using the super-resolution module to perform resolution optimization based on the first cropped image and the second cropped image to obtain an optimized target image; wherein the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection frame, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection frame; the second transmission channel is a channel other than the first transmission channel in the second number of transmission channels; Extracting the two-dimensional coordinates of the hand skeleton key points from the target image through the key point detection module; Converting the two-dimensional coordinates into three-dimensional coordinates through the coordinate calculation module, the three-dimensional coordinates are used for virtual interaction.

9. An electronic device, comprising: Comprise: A processor, a memory and a bus, the memory stores machine readable instructions executable by the processor, when the electronic device is running, the processor and the memory communicate through the bus, the machine readable instructions are executed by the processor to execute the steps of the interactive positioning method based on real-time high-precision detection of hand key points according to claim 8.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which is executed by the processor to execute the steps of the interactive positioning method based on real-time high-precision detection of hand key points according to claim 8.

Citation Information

Patent Citations

  • Gesture recognition system in three-dimensional space and recognition method thereof

    CN103440035A

  • Global hand gesture detecting method based on depth data

    CN105759967A