Interactive positioning system and method based on real-time high-precision detection of hand key points
Through a multi-device interactive positioning system, using a combination of global and regional detection, and using a lightweight multi-channel feature calibration network and a high-resolution network, the problems of insufficient efficiency and accuracy in hand key point detection in existing technologies are solved, and real-time and high-precision interactive positioning of hand key points is achieved.
Patent Information
- Application Number
- CN202511116828.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-11
AI Technical Summary
In the existing technology of real-time high-precision detection of key points of the hand, the single media stream method limits the recognition efficiency and the accuracy of the recognition results depends on the resolution of the image acquisition device, resulting in excessively high requirements for the equipment when high precision is required.
It uses multiple image acquisition devices, combined with a real-time current limiting module, a palm detection module, a super-resolution module, a key point detection module, and a coordinate calculation module. By combining global detection and regional detection, it uses a lightweight multi-channel feature calibration network and a high-resolution network to achieve image resolution optimization and three-dimensional coordinate transformation.
It improves detection efficiency and accuracy, meets the needs of real-time and high-precision interactive positioning of hand key points, and realizes the real-time and high-precision of virtual interaction.
Smart Images

Figure CN120656243A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of interactive positioning technology, and in particular to an interactive positioning system and method based on real-time and high-precision detection of key points of the hand. Background Art
[0002] In computer science, hand gesture detection is the foundation of gesture recognition; real-time, high-precision positioning is the foundation of XR (Extended Reality) for enabling natural interaction between the real and virtual worlds. Interactive positioning is a topic that uses mathematical algorithms to recognize human gestures. Users can use simple gestures to control or interact with devices, allowing computers to understand human behavior.
[0003] Existing technologies for interactive positioning often use a single media stream. While this approach reduces resource usage, it limits recognition efficiency. Furthermore, the accuracy of the recognition results depends entirely on the resolution of the image acquisition device itself, placing high demands on the image acquisition device when high-precision results are required. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide an interactive positioning system and method based on real-time and high-precision detection of hand key points to overcome the problems in the prior art.
[0005] In a first aspect, an embodiment of the present application provides an interactive positioning system based on real-time and high-precision detection of key points of a hand, comprising a plurality of image acquisition devices deployed at preset positions for acquiring an initial image of a shooting area; a real-time current limiting module, connected to the image acquisition device, and initially sending the initial image to the palm detection module through a first number of transmission channels; a palm detection module, connected to the real-time current limiting module, detecting the initial image based on a palm detection method to obtain a detection result, and activating a second number of transmission channels to receive the initial image when the detection result meets a preset detection condition; the second number is greater than the first number; Using a palm detection method, the to-be-detected area of the initial image in the first transmission channel is detected to obtain a palm detection frame; wherein the to-be-detected area is predicted based on the actual position of the palm in the detection result; and the first transmission channel is a portion of the second number of transmission channels; a super-resolution module, connected to the palm detection module, configured to perform resolution optimization based on the first cropped image and the second cropped image to obtain an optimized target image; wherein the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection frame, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection frame; and the second transmission channel is a channel of the second number of transmission channels excluding the first transmission channel; A key point detection module, connected to the super-resolution module, for extracting the two-dimensional coordinates of the key points of the hand skeleton from the target image; A coordinate calculation module is connected to the key point detection module and is used to convert the two-dimensional coordinates into three-dimensional coordinates, and the three-dimensional coordinates are used for virtual interaction.
[0006] In some technical solutions of the present application, the palm detection method is used to detect the initial image to obtain a detection result, including: The initial image is detected as a whole using a preset target detection model to obtain the number of palms contained in the initial image and the real position corresponding to each palm.
[0007] In some technical solutions of the present application, whether the test result meets the preset test conditions is determined in the following manner: The number of palms included in the detection result is greater than or equal to a preset number threshold, and the detection result meets the preset detection condition; The number of palms included in the detection result is less than a preset number threshold, and the detection result meets the non-preset detection condition.
[0008] In some technical solutions of the present application, the palm detection method is used to detect the to-be-detected area of the initial image in the first transmission channel based on the detection result to obtain a palm detection frame, including: Predicting the movement of the palm according to the positional relationship between the image acquisition devices and the real position to obtain the area to be detected; The area to be detected is detected using the palm detection method to obtain a palm detection frame containing a palm in the area to be detected.
[0009] In some technical solutions of the present application, the above-mentioned resolution optimization based on the first cropped image and the second cropped image to obtain the optimized target image includes: The first cropped image and the second cropped image are input into a preset lightweight multi-channel feature calibration network to obtain a target image output by the lightweight multi-channel feature calibration network.
[0010] In some technical solutions of the present application, there are two target images, and converting the two-dimensional coordinates into three-dimensional coordinates includes: The two-dimensional coordinates are converted into three-dimensional coordinates according to the optical axis distance and focal length between the image acquisition devices.
[0011] In some technical solutions of the present application, the above-mentioned step of extracting the two-dimensional coordinates of the key points of the hand skeleton from the target image includes: The target image is input into a high-resolution network to obtain the two-dimensional coordinates of the hand skeleton key points output by the high-resolution network.
[0012] In a second aspect, an embodiment of the present application provides an interactive positioning method based on real-time and high-precision detection of hand key points, the method comprising: Capturing an initial image of the shooting area by using image acquisition devices at multiple preset positions; Initially, controlling the real-time current limiting module to send the initial image to the palm detection module through a first number of transmission channels; Using a palm detection module to detect the initial image based on a palm detection method to obtain a detection result, and when the detection result meets a preset detection condition, activating a second number of transmission channels to receive the initial image; the second number is greater than the first number; Using a palm detection method, the to-be-detected area of the initial image in the first transmission channel is detected to obtain a palm detection frame; wherein the to-be-detected area is predicted based on the actual position of the palm in the detection result; and the first transmission channel is a portion of the second number of transmission channels; Using a super-resolution module to optimize the resolution of the first cropped image and the second cropped image to obtain an optimized target image; wherein the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection frame, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection frame; the second transmission channel is a channel in the second number of transmission channels excluding the first transmission channel; Extracting two-dimensional coordinates of key points of the hand skeleton from the target image through a key point detection module; The two-dimensional coordinates are converted into three-dimensional coordinates by a coordinate calculation module, and the three-dimensional coordinates are used for virtual interaction.
[0013] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of the above-mentioned interactive positioning method based on real-time and high-precision detection of hand key points are performed.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned interactive positioning method based on real-time and high-precision detection of hand key points are executed.
[0015] The technical solutions provided by the embodiments of the present application may have the following beneficial effects: The system includes a plurality of image acquisition devices, which are deployed at a preset position and are used to acquire an initial image of a shooting area; a real-time current limiting module, which is connected to the image acquisition device and initially sends the initial image to the palm detection module through a first number of transmission channels; a palm detection module, which is connected to the real-time current limiting module and detects the initial image based on a palm detection method to obtain a detection result, and when the detection result meets a preset detection condition, activates a second number of transmission channels to receive the initial image; the second number is greater than the first number; the palm detection method is used to detect the area to be detected in the initial image in the first transmission channel to obtain a palm detection frame; wherein, the area to be detected is predicted based on the actual position of the palm in the detection result; The first transmission channel is part of the second number of transmission channels; the super-resolution module is connected to the palm detection module, and is used to optimize the resolution based on the first cropped image and the second cropped image to obtain an optimized target image; wherein, the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection frame, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection frame; the second transmission channel is the channel of the second number of transmission channels other than the first transmission channel; the key point detection module is connected to the super-resolution module, and is used to extract the two-dimensional coordinates of the key points of the hand bones from the target image; the coordinate calculation module is connected to the key point detection module.
[0016] This application uses different numbers of transmission channels for data transmission at different stages through the real-time current limiting module, effectively saving transmission resources; the palm detection module uses a combination of global detection and regional detection to improve detection efficiency, and optimizes the image resolution through the super-resolution module to improve the accuracy of interaction. The combination of all modules also ensures the real-time nature of the interaction.
[0017] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 A schematic diagram of an interactive positioning system based on real-time and high-precision detection of hand key points provided by an embodiment of the present application is shown; Figure 2 A schematic diagram of skeleton key points provided in an embodiment of the present application is shown; Figure 3 A schematic diagram showing the relationship between pixel coordinates and world coordinates of a binocular camera provided in an embodiment of the present application is shown; Figure 4 A schematic diagram showing a specific implementation of an interactive positioning system based on real-time high-precision detection of hand key points provided by an embodiment of the present application is shown; Figure 5 A flow chart of an interactive positioning method based on real-time and high-precision detection of hand key points provided by an embodiment of the present application is shown; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.
[0021] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0022] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0023] In computer science, interactive localization based on real-time, high-precision detection of hand keypoints is a topic that uses mathematical algorithms to recognize human gestures. Users can use simple gestures to control or interact with devices, allowing computers to understand human behavior.
[0024] Existing technologies for interactive positioning based on real-time, high-precision detection of hand key points often rely on a single media stream approach. While this approach reduces resource usage, it limits recognition efficiency. Furthermore, the accuracy of this recognition result depends entirely on the resolution of the image acquisition device itself, placing high demands on the image acquisition device when high-precision results are required.
[0025] Based on this, the embodiments of the present application provide an interactive positioning system and method based on real-time high-precision detection of hand key points, which are described below through embodiments. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0026] Figure 1 A schematic diagram of an interactive positioning system based on real-time, high-precision detection of hand key points, provided by an embodiment of the present application, is shown. The interactive positioning system based on real-time, high-precision detection of hand key points includes an image acquisition device, a real-time current limiting module, a palm detection module, a super-resolution module, a key point detection module, and a coordinate calculation module. The image acquisition device is connected to the real-time current limiting module, which is connected to the palm detection module, which is connected to the super-resolution module, which is connected to the key point detection module, which is connected to the coordinate calculation module, and the coordinate calculation module is connected to an external interactive device.
[0027] In the embodiment of the present application, there are multiple image acquisition devices (cameras, mobile phones, webcams, etc.), for example, four, nine, fifteen, twenty-two, etc. They are deployed at preset positions, and their positions relative to each other do not change after deployment. For example, four image acquisition devices (A1, A2, A3, A4) are fixed in pairs, with A1 adjacent to A2, A3 adjacent to A4, the distance between A1 and A3 is x1, and the distance between A2 and A4 is also x1. The acquisition device captures images toward the preset shooting area (the area where the palm appears) (based on a preset acquisition frequency, the specific acquisition frequency can be set according to needs, and the embodiment of the present application does not limit the specific acquisition frequency) to obtain an initial image. The initial image includes the background or the background and the hand. After capturing the initial image, the image acquisition device sends the initial image to the real-time current limiting module.
[0028] The real-time current limiting module transmits the initial image to the palm detection module via a first number of transmission channels. In the embodiment of the present application, multiple transmission channels (e.g., represented by X) are set between the real-time current limiting module and the palm detection module, but the number of transmission channels opened varies at different stages. For example, initially (when the initial image has not been detected or the detection result of the initial image does not meet the preset detection conditions), a first number of transmission channels (where the first number is less than X) is opened between the real-time current limiting module and the palm detection module. For example, if the first number is two, the initial image can only be transmitted via these two transmission channels. Alternatively, if the first number is five, the initial image can only be transmitted via these five transmission channels. Alternatively, if the first number is nine, the initial image can only be transmitted via these nine transmission channels.
[0029] After the real-time current limiting module sends the initial image to the palm detection module, the palm detection module will detect the initial image. When performing the detection, in order to improve the detection efficiency, the palm detection module of the embodiment of the present application uses a palm detection method rather than a hand detection method. The hand is located at the distal end of the wrist and is the terminal structure of the entire upper limb. The structure of the hand is complex and delicate, including the wrist, palm and fingers. The wrist is connected to the forearm and is composed of many small bones. It is connected to the forearm through ligaments and muscles; the palm supports the fingers and contains many muscles and nerves; the fingers are composed of phalanges, joints, muscles and skin, and can perform actions such as grasping and pinching. The palm is part of the hand.
[0030] The initial image is detected using a palm detection method to obtain a detection result. It is important to emphasize that the detection process at this point is performed on all content in the initial image, including background detection or background and hand detection. The detection result includes the number of palms contained and the actual location of each palm. For example, the detection of the initial image may determine that there are no palms in the initial image (the number is zero), or that there is a palm in the initial image and the location of the palm is determined.
[0031] In specific implementations, the palm detection method is implemented using the SSD (Single Shot Multibox Detector) target detection model. SSD is a single-stage target detection algorithm that extracts features through a convolutional neural network and outputs detection at different feature layers to achieve multi-scale detection. It adopts an anchor (bounding box) strategy, presetting anchors of different aspect ratios and predicting their positions (represented by a palm detection box) on each output feature layer. The SSD framework includes a multi-scale detection method, with shallow layers used to detect small targets and deep layers used to detect large targets. Detecting hands is a very complex task because hands come in many sizes and are subject to occlusion and self-occlusion. Unlike face detection, a face is a high-contrast pattern (the mouth and eyes can assist in detecting the face), while a hand is a dynamic and complex pattern, making it difficult to predict a hand based solely on visual features. To address this issue, the embodiment of the present application adopts the following strategy: first, a palm detector is trained, rather than a detector for the entire hand. Because the palm is a square, rigid object that doesn't move around like fingers, it's a relatively stable feature. Furthermore, palms are smaller, so the non-maximum suppression algorithm performs better in situations like handshakes or self-occlusion. Furthermore, square palms don't need to consider different aspect ratios, reducing unnecessary anchor boxes. Next, an encoder-decoder is used for feature extraction.
[0032] After obtaining the detection results, the palm detection module needs to use the detection results to compare them with the preset detection conditions to determine whether the detection results meet the detection conditions. The detection conditions here can be a preset number threshold: if the number of palms contained in the detection results is greater than or equal to the preset number threshold, the detection results meet the preset detection conditions; if the number of palms contained in the detection results is less than the preset number threshold, the detection results meet the non-preset detection conditions. For example, if the number threshold is set to three, if the initial image does not contain any palms, then the initial image does not meet the preset detection conditions; if the initial image contains four palms, then the initial image meets the preset detection conditions.
[0033] When the preset detection conditions are met, the embodiment of the present application determines that subsequent initial images require focused detection. To improve detection efficiency, the embodiment of the present application expands the original first number of transmission channels to a second number of transmission channels. The real-time current limiting module then transmits the image to the palm detection module based on the second number of transmission channels.
[0034] On the other hand, since the above detection process has determined that the initial image contains a preset number threshold of palms, due to the interactive characteristics, the hands are in motion, and the subsequent initial images also contain palms greater than or equal to the number threshold. Moreover, based on the actual position of the detected palm, the motion properties of the palm and the relative position between the image acquisition devices, the position where the palm appears in the subsequent initial images can be predicted to obtain the position where the palm will appear in each subsequent initial image. The specific prediction method can use the prediction method already available in the prior art, and the embodiment of the present application does not limit the specific prediction method. After obtaining the position where the subsequent palm will appear, the embodiment of the present application determines a region to be detected based on the predicted position. In order to improve the detection efficiency, the region to be detected can be directly detected instead of performing a global detection on the initial image.
[0035] In specific implementation, the embodiment of the present application adopts a combination of global detection and ROI (Region of Interest) detection. If no palm is detected or the number of detected palms is less than a preset threshold (for example, 2), the global detection mode is adopted; when the palm is detected or the preset threshold number of palms is detected, the area to be detected is predicted based on the position of the palm obtained by global detection (which can be represented by a bounding box or ROI, and the detected palm detection frame is transmitted to other transmission channels), and then ROI detection is performed on the area to be detected.
[0036] Furthermore, after the palm detection module acquires the initial image from the second number of transmission channels, considering image resolution, the embodiment of the present application does not detect all of the initial images from the second number of transmission channels. Instead, the second number of transmission channels is divided into a first transmission channel and a second transmission channel. The first transmission channel is a portion of the second number of transmission channels; the second transmission channel is a transmission channel from the second number of transmission channels excluding the first transmission channel. That is, in the embodiment of the present application, the palm detection module only detects the to-be-detected area of the initial image from the first transmission, obtains a palm detection frame, and sends the palm detection frame to the super-resolution module, while the initial image from the second transmission channel is directly transmitted to the super-resolution module.
[0037] To improve processing efficiency and accuracy, the super-resolution module does not process the entire image, but first crops the image: the initial image in the first transmission channel is cropped based on the palm detection frame to obtain a first cropped image, and the initial image in the second transmission channel is cropped based on the palm detection frame to obtain a second cropped image. The resolution of the first cropped image and the second cropped image are then optimized to obtain an optimized target image.
[0038] When optimizing resolution, the present embodiment utilizes a lightweight multi-path feature calibration network (MPFCN). This lightweight multi-path feature calibration network (MPFCN) strives to achieve a good balance between parameters and performance, achieving excellent super-resolution while also ensuring real-time system detection. By modeling contextual dependencies across multiple spatial scales and channel dimensions, it fully exploits the spatial information and channel characteristics of the image, resulting in more accurate image reconstruction results.
[0039] After obtaining the target image, the embodiment of the present application performs coordinate extraction through the key point detection module, and obtains the two-dimensional coordinates of the key points of the hand skeleton by extracting the coordinates of the target image. Figure 2 As shown in the figure, a deep learning algorithm generates 21 skeletal keypoints (2D / 2.5D) of the hand. These 21 keypoints correspond to the 21 joints of the human hand: one at the base of the palm, four at the thumb, four at the index finger, four at the middle finger, four at the ring finger, and four at the pinky finger. Joint 0 is defined as the base of the palm, and the joints of each finger are numbered incrementally upwards, from the base of the palm. The joints from thumb to pinky are numbered 1-4, 5-8, 9-12, 13-16, and 17-20, respectively. Figure 2The joint numbers given in the figure are abbreviations; only the joint numbers for the base of the palm and the tips of each finger are given. The z-axis data are coordinates based on the palm base node (node 0), according to the rule that distance is greater than distance and distance is smaller than distance. For example, hand keypoint detection is based on an improved implementation of the open-source HRNet (High Resolution Network), which constructs a complete training and testing implementation for hand skeletal keypoint detection. The network connects high- and low-resolution convolutional streams in parallel. It maintains a high-resolution representation throughout the entire process and repeatedly fuses representations from multiple resolution streams to generate a reliable, position-sensitive, high-resolution representation. To ensure real-time and high-accuracy detection, two improvements were made. First, the HRNet input is not the entire image, but the predicted bounding box output by the palm detection module, which significantly improves HRNet detection speed. Second, the bounding box is cropped from images from multiple channels and super-resolution processing is performed on the cropped distributed image to increase the resolution of the input image, thereby improving the localization accuracy of keypoints. It is important to note that the hand bounding box ROI is adjusted based on the detected keypoint coordinates (2D) output to achieve accurate tracking.
[0040] After obtaining the two-dimensional coordinates, the two-dimensional coordinates are converted into three-dimensional coordinates by the coordinate calculation module. The main task of the coordinate calculation module is to realize the three-dimensional coordinate calculation of the hand skeleton key point detection coordinates of the two target images. Although the estimation of the 3D coordinate z-axis of the hand skeleton key points can be achieved through learning, this accuracy cannot meet the requirements of high-precision positioning applications. Therefore, the embodiment of the present application combines the positioning principle of binocular cameras to realize the calculation of the 3D coordinates of the key points.
[0041] The process of obtaining the key point 2D coordinates (pixel coordinates) of the binocular camera to the key point 3D coordinates (world coordinates). Figure 3 As shown (Z L 、Z R The Z coordinates of the left and right hand keypoints are respectively corresponding. The Z coordinate is achieved through binocular vision positioning. Ideally, when the y-axes of the left and right cameras are perfectly aligned, the optical axis distance between the two cameras is b, and the focal lengths of the two cameras are equal, both f. Given the pixel coordinates of the hand keypoints (PL(uXL, vYL), PR(uXR, vYR)), the world coordinates P (xW, yW, zW) of the corresponding keypoints are solved as follows:
[0042] Among them, d is the phase difference, that is, d=ux-uR.
[0043]
[0044]
[0045] Where u0 and v0 are the camera image origins. Let XW = X / W, YW = Y / W, and ZW = Z / W. The 2D coordinates (pixel coordinates) of the key points of the binocular camera can be calculated as follows:
[0046] The 3D coordinate positioning of key points on the hand skeleton provides a key element for skeleton-based action recognition. By identifying the action being performed from a series of time-continuous skeleton key points (2D / 3D), gesture interaction is realized. By detecting 21 key points on the hand skeleton and accurately positioning their spatial coordinates (world coordinates) in real time, the fusion of virtual and real-world interaction is achieved.
[0047] The embodiments of the present application effectively solve the problem of hand key point detection while meeting the requirements of real-time, high-precision, and three-dimensional positioning.
[0048] In an optional embodiment, in specific implementation, the system in the embodiment of the present application can be Figure 4 The process shown here operates as follows: The image acquisition device uses cameras, specifically four cameras: Camera 1, Camera 2, Camera 3, and Camera 4. The initial images captured by the four cameras (Image 1 from Camera 1, Image 2 from Camera 2, Image 3 from Camera 3, and Image 4 from Camera 4) are sent to the real-time current limiting module. Initially, the real-time current limiting module transmits the initial images captured by the four cameras via a single transmission channel to the palm detection module. The palm detection module performs global detection on the initial images. When it detects two palms in the initial images, it transmits the initial images to the palm detection module via four transmission channels (one for each camera: the first, second, third, and fourth channels). Based on the palm detection frames detected during global detection, the palm detection module predicts palms in subsequent initial images to obtain the detection area. The palm detection module then detects the detection area in the initial images transmitted via the second and third transmission channels, obtaining palm detection frames and simultaneously transmitting these frames to the first and fourth transmission channels. The super-resolution module then crops the images based on the palm detection frame, generating cropped images for each transmission channel. The cropped images from the first and second transmission channels are then resolution-optimized to create a single target image. The cropped images from the third and fourth transmission channels are then resolution-optimized to create a single target image. After obtaining the coordinates of the two target images, they are transformed based on the principles of a binocular camera. This results in three-dimensional coordinates, enabling virtual interaction.
[0049] Figure 5A flow chart of an interactive positioning method based on real-time high-precision detection of hand key points provided by an embodiment of the present application is shown, wherein the method includes steps S101-S107; specifically: S101, capturing an initial image of a shooting area using image capture devices at multiple preset positions; S102: Initially, controlling the real-time current limiting module to send the initial image to the palm detection module through a first number of transmission channels; S103: Detecting the initial image using a palm detection module based on a palm detection method to obtain a detection result, and activating a second number of transmission channels to receive the initial image when the detection result meets a preset detection condition; the second number of transmission channels is greater than the first number; S104: Detecting a to-be-detected area of the initial image in the first transmission channel using a palm detection method to obtain a palm detection frame; wherein the to-be-detected area is predicted based on the actual position of the palm in the detection result; and the first transmission channel is a portion of the second number of transmission channels; S105. Using a super-resolution module, perform resolution optimization based on the first cropped image and the second cropped image to obtain an optimized target image; wherein the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection frame, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection frame; and the second transmission channel is a channel in the second number of transmission channels excluding the first transmission channel; S106, extracting the two-dimensional coordinates of the hand skeleton key points from the target image through a key point detection module; S107: Convert the two-dimensional coordinates into three-dimensional coordinates through a coordinate calculation module, and use the three-dimensional coordinates for virtual interaction.
[0050] The initial image is detected based on a palm detection method to obtain a detection result, including: The initial image is detected as a whole using a preset target detection model to obtain the number of palms contained in the initial image and the real position corresponding to each palm.
[0051] The method determines whether the detection result meets the preset detection conditions in the following manner: The number of palms included in the detection result is greater than or equal to a preset number threshold, and the detection result meets the preset detection condition; The number of palms included in the detection result is less than a preset number threshold, and the detection result meets the non-preset detection condition.
[0052] The detecting the to-be-detected area of the initial image in the first transmission channel using a palm detection method based on the detection result to obtain a palm detection frame includes: Predicting the movement of the palm according to the positional relationship between the image acquisition devices and the real position to obtain the area to be detected; The area to be detected is detected using the palm detection method to obtain a palm detection frame containing a palm in the area to be detected.
[0053] The step of performing resolution optimization based on the first cropped image and the second cropped image to obtain an optimized target image includes: The first cropped image and the second cropped image are input into a preset lightweight multi-channel feature calibration network to obtain a target image output by the lightweight multi-channel feature calibration network.
[0054] There are two target images, and converting the two-dimensional coordinates into three-dimensional coordinates includes: The two-dimensional coordinates are converted into three-dimensional coordinates according to the optical axis distance and focal length between the image acquisition devices.
[0055] The step of extracting the two-dimensional coordinates of the hand skeleton key points from the target image includes: The target image is input into a high-resolution network to obtain the two-dimensional coordinates of the hand skeleton key points output by the high-resolution network.
[0056] like Figure 6 As shown, an embodiment of the present application provides an electronic device for executing the interactive positioning method based on real-time and high-precision detection of hand key points in the present application. The device includes a memory, a processor, a bus, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the interactive positioning method based on real-time and high-precision detection of hand key points are implemented.
[0057] Specifically, the above-mentioned memory and processor can be general-purpose memory and processor, which are not specifically limited here. When the processor runs the computer program stored in the memory, it can execute the above-mentioned interactive positioning method based on real-time and high-precision detection of hand key points.
[0058] Corresponding to the interactive positioning method based on real-time and high-precision detection of hand key points in the present application, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by the processor, the steps of the above-mentioned interactive positioning method based on real-time and high-precision detection of hand key points are executed.
[0059] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, it can execute the above-mentioned interactive positioning method based on real-time and high-precision detection of hand key points.
[0060] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of the system or unit, which can be electrical, mechanical or other forms.
[0061] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.
[0062] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0063] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0064] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0065] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. However, these modifications, changes, or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims.
Claims
1. An interactive positioning system based on real-time high-precision detection of hand key points, characterized in that: include: Multiple image acquisition devices are deployed at preset positions to acquire initial images of the shooting area; a real-time current limiting module, connected to the image acquisition device, and initially sending the initial image to the palm detection module through a first number of transmission channels; a palm detection module, connected to the real-time current limiting module, detecting the initial image based on a palm detection method to obtain a detection result, and activating a second number of transmission channels to receive the initial image when the detection result meets a preset detection condition; the second amount is greater than the first amount; Using a palm detection method, the to-be-detected area of the initial image in the first transmission channel is detected to obtain a palm detection frame; wherein the to-be-detected area is predicted based on the actual position of the palm in the detection result; and the first transmission channel is a portion of the second number of transmission channels; a super-resolution module, connected to the palm detection module, configured to perform resolution optimization based on the first cropped image and the second cropped image to obtain an optimized target image; wherein the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection frame, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection frame; and the second transmission channel is a channel of the second number of transmission channels excluding the first transmission channel; A key point detection module, connected to the super-resolution module, for extracting the two-dimensional coordinates of the key points of the hand skeleton from the target image; A coordinate calculation module is connected to the key point detection module and is used to convert the two-dimensional coordinates into three-dimensional coordinates, and the three-dimensional coordinates are used for virtual interaction.
2. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1 is characterized in that: The detecting the initial image based on the palm detection method to obtain a detection result includes: The initial image is detected as a whole using a preset target detection model to obtain the number of palms contained in the initial image and the real position corresponding to each palm.
3. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1, characterized in that: The interactive positioning system based on real-time high-precision detection of hand key points determines whether the detection result meets the preset detection conditions in the following manner: The number of palms included in the detection result is greater than or equal to a preset number threshold, and the detection result meets the preset detection condition; The number of palms included in the detection result is less than a preset number threshold, and the detection result meets the non-preset detection condition.
4. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1, characterized in that: The detecting the to-be-detected area of the initial image in the first transmission channel using a palm detection method based on the detection result to obtain a palm detection frame includes: Predicting the movement of the palm according to the positional relationship between the image acquisition devices and the real position to obtain the area to be detected; The area to be detected is detected using the palm detection method to obtain a palm detection frame containing a palm in the area to be detected.
5. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1, characterized in that: The step of performing resolution optimization based on the first cropped image and the second cropped image to obtain an optimized target image includes: The first cropped image and the second cropped image are input into a preset lightweight multi-channel feature calibration network to obtain a target image output by the lightweight multi-channel feature calibration network.
6. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 5, characterized in that: There are two target images, and converting the two-dimensional coordinates into three-dimensional coordinates includes: The two-dimensional coordinates are converted into three-dimensional coordinates according to the optical axis distance and focal length between the image acquisition devices.
7. The interactive positioning system based on real-time high-precision detection of hand key points according to claim 1, characterized in that: The step of extracting the two-dimensional coordinates of the hand skeleton key points from the target image includes: The target image is input into a high-resolution network to obtain the two-dimensional coordinates of the hand skeleton key points output by the high-resolution network.
8. An interactive positioning method based on real-time high-precision detection of hand key points, characterized in that: The interactive positioning system based on real-time high-precision detection of hand key points according to any one of claims 1 to 7, wherein the interactive positioning method based on real-time high-precision detection of hand key points comprises: Capturing an initial image of the shooting area by using image acquisition devices at multiple preset positions; Initially, controlling the real-time current limiting module to send the initial image to the palm detection module through a first number of transmission channels; Using a palm detection module to detect the initial image based on a palm detection method to obtain a detection result, and when the detection result meets a preset detection condition, activating a second number of transmission channels to receive the initial image; the second number is greater than the first number; Using a palm detection method, the to-be-detected area of the initial image in the first transmission channel is detected to obtain a palm detection frame; wherein the to-be-detected area is predicted based on the actual position of the palm in the detection result; and the first transmission channel is a portion of the second number of transmission channels; Using a super-resolution module to optimize the resolution of the first cropped image and the second cropped image to obtain an optimized target image; wherein the first cropped image is obtained by cropping the initial image in the first transmission channel based on the palm detection frame, and the second cropped image is obtained by cropping the initial image in the second transmission channel based on the palm detection frame; the second transmission channel is a channel in the second number of transmission channels excluding the first transmission channel; Extracting two-dimensional coordinates of key points of the hand skeleton from the target image through a key point detection module; The two-dimensional coordinates are converted into three-dimensional coordinates by a coordinate calculation module, and the three-dimensional coordinates are used for virtual interaction.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the steps of the interactive positioning method based on real-time and high-precision detection of hand key points are performed as described in claim 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the interactive positioning method based on real-time and high-precision detection of hand key points as claimed in claim 8.
Citation Information
Patent Citations
Gesture recognition system in three-dimensional space and recognition method thereof
CN103440035A
Global hand gesture detecting method based on depth data
CN105759967A
Smart home control system based on hand gesture recognition
CN108021880A
Image detection method and apparatus, electronic device and storage medium
CN108230294A
Method for automatically generating annotation data of hand and method for calculating skeleton length
CN112767300A