Hand rotation angle regression method based on RGB image
The global visual and local geometric features are extracted through the dual-stream convolutional structure based on RGB cameras, and high-precision hand rotation angle regression is achieved, solving the problems of high equipment costs and environmental sensitivity in the prior art, and is applied to remote operation of robotic arm.
Patent Information
- Application Number
- CN202510281921.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-22
AI Technical Summary
When realizing high-precision gesture regression, the prior art relies on high-cost depth cameras or depth sensors, and is sensitive to light and background, has high equipment cost, complex calculations, poor scene generalization, and lacks research on directly implementing hand rotation angle estimates based on RGB cameras.
The image is acquired based on the RGB camera, and the global visual features and local geometric features are extracted using the dual-stream convolutional structure. The hand rotation angle regression is realized through weighted fusion, which is applied to remote operation of the robotic arm.
It realizes high-precision hand rotation angle regression, improves the naturalness and accuracy of remote operation robot arms, reduces equipment costs, and enhances application flexibility in complex environments.
Smart Images

Figure CN120356259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gesture recognition, and particularly relates to a method for hand rotation angle regression based on RGB images. Background Art
[0002] In the fields of robotics and intelligent human-computer interaction, gesture recognition technology is one of the core technologies, which enables users to directly interact with machines or computing systems through natural hand movements.
[0003] Currently, when achieving high-precision gesture pose regression, it usually relies on high-cost depth cameras or depth sensors. The spatial regions for these depth devices to achieve high-precision regression are limited, and they are extremely sensitive to light changes and background complexity in the environment. These factors limit the practicality and flexibility of the technology. In addition, when using multiple RGB cameras for gesture pose regression, multi-view gesture data is usually required to achieve high precision, which not only increases the costs of devices, deployment, and computing, but also has poor scene generalization. The current gesture recognition technology based on a single RGB camera mainly focuses on accurately measuring the positions of finger joints and the overall shape of the hand, and the datasets it relies on are mostly computer-generated 3D models, which limits its application in the real world. The hand rotation angle is also part of the gesture pose. Currently, there is still a lack of research on estimating the hand rotation angle directly based on RGB images collected by a single RGB camera. Summary of the Invention
[0004] In order to overcome the defects and deficiencies of the prior art, the present invention provides a method for hand rotation angle regression based on RGB images. The present invention makes full use of the global visual features and local geometric features of the RGB images collected by the RGB camera to achieve end-to-end high-precision hand rotation angle regression.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a method for hand rotation angle regression based on RGB images, including the following steps:
[0007] Collect hand rotation images based on an RGB camera and label hand rotation angle tags to construct a hand rotation angle dataset;
[0008] Divide the hand rotation angle dataset into a training set and a test set;
[0009] Extract features from the hand rotation images of the hand rotation angle dataset based on a two-stream convolutional structure to obtain global visual features and local geometric features;
[0010] Perform weighted fusion on the global visual features and local geometric features;
[0011] Obtain a training set, train a hand rotation angle regression model based on the difference between the hand rotation angle obtained by regression and the hand rotation angle of the label value, and update the parameters of each layer of the network in the hand rotation angle regression model;
[0012] Verify the accuracy of the updated hand rotation angle regression model based on the test set. When the mean absolute error is the smallest, use the updated hand rotation angle regression model as the trained hand rotation angle regression model;
[0013] Apply the trained hand rotation angle regression model to the end of the robotic arm, and change the posture by rotating the end of the robotic arm remotely with the human hand.
[0014] As a preferred technical solution, collect hand rotation images based on an RGB camera and label hand rotation angle labels to construct a hand rotation angle data set, specifically including:
[0015] Use a Leap Motion sensor to collect the palm normal vector of the hand in the same hand rotation image and label the hand rotation angle label;
[0016] Organize the collected image data and hand rotation angle labels into a hand rotation angle data set. The input data included in each sample in the hand rotation angle data set is the hand rotation image at the initial moment when the hand rotation angle is -0° and the current hand rotation image of the hand rotation angle to be regressed, and the output label value is the hand rotation angle.
[0017] As a preferred technical solution, the hand rotation angle label is expressed as:
[0018] θ = degrees(arcsin(n x ))
[0019] where n x represents the component of the palm normal vector on the X-axis of the right-hand coordinate system of the Leap Motion sensor, and degrees(·) represents converting the obtained angle from radians to degrees.
[0020] As a preferred technical solution, perform feature extraction on the hand rotation images of the hand rotation angle data set based on a two-stream convolutional structure to obtain global visual features and local geometric features, specifically including:
[0021] Extract the global visual features of the hand rotation images based on a two-stream ResNet18 network. The two ResNet18s in the two-stream ResNet18 network independently extract the global visual feature tensors of the hand rotation image at the initial moment and the hand rotation image at the current moment in parallel;
[0022] Use the Mediapipe framework to detect hand key points in the hand rotation images and extract the coordinates of the hand key points.
[0023] As a preferred technical solution, the global visual feature and the local geometric feature are weighted and fused, specifically including:
[0024] Concatenate the global visual feature tensors of the hand rotation image at the initial moment and the hand rotation image at the current moment on dimension 1 to obtain the concatenated global visual feature vector;
[0025] Input the ratio of the obtained initial image distance feature vector to the distance feature vector of the current image into two fully connected layers and the ReLU activation function to obtain a new local feature vector;
[0026] Weight and add the concatenated global visual feature vector and the new local feature vector according to weights α and β respectively, and input the fused vector into a single fully connected layer to regress the hand rotation angle.
[0027] As a preferred technical solution, the construction steps of the image distance feature vector include:
[0028] Select the middle finger as the reference line l mid , select the metacarpophalangeal joint and the distal interphalangeal joint of the middle finger as the reference points, define the reference line as the straight line connecting the metacarpophalangeal joint and the distal interphalangeal joint, and the perpendicular distance d from each key point of the hand other than the middle finger to the reference line mid,i is:
[0029]
[0030] where and represent the coordinates of the metacarpophalangeal joint and the distal interphalangeal joint respectively, is the two-dimensional coordinate vector representation of the key point of the hand other than the middle finger;
[0031] Based on the perpendicular distance d mid,i Obtain an image distance feature vector containing multiple distance features.
[0032] As a preferred technical solution, apply the trained hand rotation angle regression model to the end of the robotic arm, and change the posture by remotely operating the end of the robotic arm through the rotation of the human hand. The specific steps include:
[0033] Input the hand rotation image at the initial moment and the hand rotation image at the current moment into the trained hand rotation angle regression model to obtain the hand rotation angle;
[0034] Map the hand rotation angle to the end posture in the base coordinate system of the robotic arm.
[0035] The present invention also provides a hand rotation angle regression system based on RGB images, including: a hand rotation angle dataset construction module, a dataset division module, a feature extraction module, a weighted fusion module, a regression model training module, an optimal regression model output module, and a pose control module;
[0036] The hand rotation angle dataset construction module is used to collect hand rotation images based on an RGB camera and label hand rotation angle tags, and construct a hand rotation angle dataset;
[0037] The dataset division module is used to divide the hand rotation angle dataset into a training set and a test set;
[0038] The feature extraction module is used to extract features from the hand rotation images of the hand rotation angle dataset based on a two-stream convolutional structure, and obtain global visual features and local geometric features;
[0039] The weighted fusion module is used to perform weighted fusion on the global visual features and local geometric features;
[0040] The regression model training module is used to obtain a training set, and train a hand rotation angle regression model based on the difference between the hand rotation angle obtained by regression and the labeled hand rotation angle, and update the parameters of each layer of the network in the hand rotation angle regression model;
[0041] The optimal regression model output module is used to verify the accuracy of the updated hand rotation angle regression model based on the test set, and when the mean absolute error is the smallest, use the updated hand rotation angle regression model as the trained hand rotation angle regression model;
[0042] The pose control module is used to apply the trained hand rotation angle regression model to the end of the robotic arm, and change the pose by remotely operating the end of the robotic arm through hand rotation.
[0043] The present invention also provides a computer-readable storage medium storing a program, and when the program is executed by a processor, it implements the above-mentioned hand rotation angle regression method based on RGB images.
[0044] The present invention also provides a computer device, including a processor and a memory for storing programs executable by the processor. When the processor executes the programs stored in the memory, it implements the above-mentioned hand rotation angle regression method based on RGB images.
[0045] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0046] (1) The present invention constructs a hand rotation angle regression model based on RGB images. This model extracts global visual features through a two-stream ResNet18 network architecture and fuses global visual features and local geometric features through a weighted fusion network, thereby achieving high-precision hand rotation angle regression.
[0047] (2) The present invention extracts the distance from the hand key points other than the middle finger to the middle finger reference line as local geometric features through a local geometric feature extraction method, effectively improving the accuracy of the hand rotation angle regression model.
[0048] (3) The present invention applies the trained hand rotation angle regression model to the end of a teleoperated robotic arm, making the teleoperation trajectory more natural, thereby enhancing the naturalness and accuracy of human-machine skill learning. Description of the Drawings
[0049] Figure 1 is a schematic flowchart of the hand rotation angle regression method based on RGB images of the present invention;
[0050] Figure 2 is a schematic diagram of the network architecture of the hand rotation angle regression model of the present invention. Detailed Embodiments
[0051] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] Embodiment 1
[0053] As Figure 1 shown, this embodiment provides a hand rotation angle regression method based on RGB images, including the following steps:
[0054] S1: Collect a hand rotation angle data set, use an RGB camera to collect hand rotation angle data, and at the same time use a LeapMotion sensor to label the image data;
[0055] S11: Continuously collect hand rotation images using an RGB camera as input sample data for the model;
[0056] Specifically, in this embodiment, 8 collection objects are used, each person collects 5 groups of data, each person continuously collects 90 frames of hand rotation images in each group, the palm is facing down at the initial moment of each group, and the hand rotation angle corresponding to the image at the initial moment is set to -0°, and the clockwise direction is the direction in which the angle increases, and hand rotation images with a hand rotation angle range of -0° to -90° are collected;
[0057] S12: Use the Leap Motion sensor to collect the palm normal vector N of the hand in the same hand rotation image, and obtain the labeled hand rotation angle θ using the following formula:
[0058] θ = degrees(arcsin(n x ))
[0059] where n x represents the component of the palm normal vector N on the X-axis of the right-hand coordinate system of the Leap Motion sensor, and degrees(·) represents converting the obtained angle from radians to degrees;
[0060] S13: Organize the collected image data and hand rotation angle labels into a hand rotation angle dataset. Each sample in the hand rotation angle dataset contains the input data which is the hand rotation image at the initial moment when the hand rotation angle is -0° and the current hand rotation image of the hand rotation angle to be regressed, and the output label value is the hand rotation angle;
[0061] S14: Divide the obtained hand rotation angle dataset into a training dataset and a test dataset.
[0062] Specifically, in this embodiment, the obtained hand rotation angle dataset is divided into a training dataset and a test dataset according to a ratio of 8:2;
[0063] S2: Extract features from the hand rotation images in the hand rotation angle dataset to obtain global visual features and local geometric features;
[0064] S21: Use a two-stream convolutional structure to extract the global visual features of the hand rotation images. In the two-stream convolutional structure, each single-stream convolutional network independently extracts the global visual feature tensors of the hand rotation image at the initial moment and the hand rotation image at the current moment in parallel;
[0065] Specifically, use a two-stream ResNet18 network to extract the global visual features of the hand rotation images. The two ResNet18s in the two-stream ResNet18 network independently extract the global visual feature tensors of the hand rotation image at the initial moment and the hand rotation image at the current moment in parallel;
[0066] S22: Use the Mediapipe framework to perform hand keypoint detection on the hand rotation images and extract the coordinates Pi of 21 hand keypoints, where i = 0, 1, 2,..., 20;
[0067] Select the middle finger as the reference line l mid , select the metacarpophalangeal joint (MCP, P9) and the distal interphalangeal joint (TIP, P12) of the middle finger as the reference points, and define the reference line l mid as the straight line connecting MCP and TIP. For each keypoint of the hand other than the middle finger, the distance to the reference line lmid The vertical distance d mid,i is:
[0068]
[0069] Wherein, and respectively represent the coordinates of the key points MCP and TIP of the middle finger, i = 0, 1, …, 8, 13, …, 20, is the two-dimensional coordinate vector representation of the key points of the hand other than the middle finger. In this way, a distance feature vector D containing 17 distance features can be obtained mid :
[0070] D mid ={d mid,0 , d mid,1 , … d mid,8 , d mid,13 , …, d mid,20}
[0071] The initial image distance feature vector D init,f and the distance feature vector D cur,f of the current image, the ratio R mid is:
[0072]
[0073] S3: Perform weighted fusion on the obtained global visual features and local geometric features, and perform regression prediction;
[0074] S31: Concatenate the global visual feature tensors of the hand rotation image at the initial moment and the hand rotation image at the current moment on dimension 1 to obtain the concatenated global visual feature vector;
[0075] S32: Input the ratio R init,f of the obtained initial image distance feature vector D cur,f and the distance feature vector D mid of the current image into two fully connected layers and the ReLU activation function to obtain a new local feature vector;
[0076] S33: Weight and add the concatenated global visual feature vector and the new local feature vector by weights α and β respectively, and input the fused vector into a fully connected layer to regress the hand rotation angle;
[0077] Specifically, the weights α and β in this embodiment are preferably 1 and 0.0025;
[0078] S4: As Figure 2As shown, the collected training set is input into the hand rotation angle regression model composed of feature extraction and weighted fusion. The regression training is carried out using the difference between the regression hand rotation angle and the hand rotation angle of the label value, and the parameters of each layer of the network in the hand rotation angle regression model are updated using an optimization algorithm.
[0079] Specifically, the L1 loss function is used to perform regression training on the difference between the regression hand rotation angle and the hand rotation angle of the label value. During the training process, the Adam optimizer is used to update the parameters of each layer of the network in the hand rotation angle regression model.
[0080] S5: Use the collected test set to verify the accuracy of the updated hand rotation angle regression model. When the minimum mean absolute error is achieved, the updated hand rotation angle regression model is used as the current optimal hand rotation angle regression model.
[0081] S6: Apply the trained hand rotation angle regression model with the minimum mean absolute error to the end of the Elite robotic arm, and change the posture by remotely operating the end of the Elite by hand rotation.
[0082] S61: Use the RGB camera to collect hand rotation images in real time. The hand rotation images at the initial moment and the current moment are input into the trained hand rotation angle regression model in real time to obtain the hand rotation angle θ.
[0083] S62: Map the hand rotation angle θ to the end posture rx in the base coordinate system of the robotic arm:
[0084] rx = (180° + θ) × π / 180°
[0085] where rx represents the rotation angle of the end of the Elite robotic arm around the X axis of the base coordinate system.
[0086] Embodiment 2
[0087] This embodiment provides a hand rotation angle regression system based on RGB images for implementing the hand rotation angle regression method based on RGB images in the above Embodiment 1. The system includes: a hand rotation angle dataset construction module, a dataset division module, a feature extraction module, a weighted fusion module, a regression model training module, an optimal regression model output module, and a posture control module.
[0088] In this embodiment, the hand rotation angle dataset construction module is used to collect hand rotation images based on the RGB camera and label the hand rotation angle labels to construct a hand rotation angle dataset.
[0089] In this embodiment, the dataset division module is used to divide the hand rotation angle dataset into a training set and a test set.
[0090] In this embodiment, the feature extraction module is used to extract features from the hand rotation images of the hand rotation angle dataset based on a two-stream convolutional structure to obtain global visual features and local geometric features;
[0091] In this embodiment, the weighted fusion module is used to perform weighted fusion on the global visual features and local geometric features;
[0092] In this embodiment, the regression model training module is used to obtain a training set, train the hand rotation angle regression model based on the difference between the hand rotation angle obtained by regression and the hand rotation angle of the label value, and update the parameters of each layer of the network in the hand rotation angle regression model;
[0093] In this embodiment, the optimal regression model output module is used to verify the accuracy of the updated hand rotation angle regression model based on the test set. When the mean absolute error is the smallest, the updated hand rotation angle regression model is used as the trained hand rotation angle regression model;
[0094] In this embodiment, the attitude control module is used to apply the trained hand rotation angle regression model to the end of the robotic arm, and change the attitude by remotely operating the end of the robotic arm through hand rotation.
[0095] Embodiment 3
[0096] This embodiment provides a storage medium, which can be a storage medium such as ROM, RAM, disk, or optical disc. The storage medium stores one or more programs. When the programs are executed by a processor, the method for regressing the hand rotation angle based on RGB images in Embodiment 1 is implemented.
[0097] Embodiment 4
[0098] This embodiment provides a computing device, which can be a desktop computer, a laptop computer, a smart phone, a PDA handheld terminal, a tablet computer, or other terminal devices with a display function. The computing device includes a processor and a memory. The memory stores one or more programs. When the processor executes the programs stored in the memory, the method for regressing the hand rotation angle based on RGB images in Embodiment 1 is implemented.
[0099] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for regressing the rotation angle of a hand based on an RGB image, characterized in that, Including the following steps: Collect hand rotation images based on an RGB camera and label hand rotation angle tags to construct a hand rotation angle dataset; Divide the hand rotation angle dataset into a training set and a test set; Extract features from the hand rotation images in the hand rotation angle dataset based on a two-stream convolutional structure to obtain global visual features and local geometric features; Perform weighted fusion on the global visual features and local geometric features; Obtain the training set, train the hand rotation angle regression model based on the difference between the hand rotation angle obtained by regression and the hand rotation angle of the label value, and update the parameters of each layer network in the hand rotation angle regression model; Verify the accuracy of the updated hand rotation angle regression model based on the test set. When the mean absolute error is the smallest, use the updated hand rotation angle regression model as the trained hand rotation angle regression model; Apply the trained hand rotation angle regression model to the end of the robotic arm, and change the posture by remotely operating the end of the robotic arm through hand rotation; 2. The method for regressing the hand rotation angle based on the RGB image according to claim 1, wherein Collect hand rotation images based on an RGB camera and label hand rotation angle tags to construct a hand rotation angle dataset, specifically including: Use a Leap Motion sensor to collect the palm normal vector of the hand in the same hand rotation image and label the hand rotation angle tag; Organize the collected image data and hand rotation angle tags into a hand rotation angle dataset. The input data included in each sample in the hand rotation angle dataset is the hand rotation image at the initial moment when the hand rotation angle is -0° and the current hand rotation image of the hand rotation angle to be regressed, and the output label value is the hand rotation angle; 3. The method for regressing the hand rotation angle based on RGB images according to claim 2, wherein The hand rotation angle tag is expressed as: θ = degrees(arcsin(n x )) where n x represents the component of the palm normal vector on the X-axis of the right-handed coordinate system of the Leap Motion sensor, and degrees(·) represents converting the obtained angle from radians to degrees.
4. The method for regressing the hand rotation angle based on RGB images according to claim 1, characterized in that Extract features from the hand rotation images in the hand rotation angle dataset based on a two-stream convolutional structure to obtain global visual features and local geometric features, specifically including: Extract the global visual features of the hand rotation image based on the two-stream ResNet18 network. The two ResNet18s in the two-stream ResNet18 network independently extract the global visual feature tensors of the hand rotation image at the initial moment and the hand rotation image at the current moment in parallel; Use the Mediapipe framework to detect the hand key points in the hand rotation image and extract the coordinates of the hand key points; 5. The method for regressing the hand rotation angle based on RGB images according to claim 1, characterized in that, Perform weighted fusion on the global visual features and local geometric features, specifically including: Concatenate the global visual feature tensors of the hand rotation image at the initial moment and the hand rotation image at the current moment obtained on dimension 1 to obtain a concatenated global visual feature vector; Input the ratio of the obtained initial image distance feature vector to the distance feature vector of the current image into two fully connected layers and the ReLU activation function to obtain a new local feature vector; Weight and add the concatenated global visual feature vector and the new local feature vector according to weights α and weight β respectively, and input the fused vector into a fully connected layer to regress the hand rotation angle; 6. The method for regressing the hand rotation angle based on RGB images according to claim 5, characterized in that, The construction steps of the image distance feature vector include: Select the middle finger as the reference line l mid , select the metacarpophalangeal joint and the distal interphalangeal joint of the middle finger as the reference points, define the reference line as the straight line connecting the metacarpophalangeal joint and the distal interphalangeal joint, and the perpendicular distance d mid,i from each key point of the hand other than the middle finger to the reference line is defined as: Among them, and represent the coordinates of the metacarpophalangeal joint and the distal interphalangeal joint respectively, is the two-dimensional coordinate vector representation of the key points of the hand other than the middle finger; Based on the vertical distance d mid,i An image distance feature vector containing multiple distance features is obtained.
7. The method for regressing the hand rotation angle based on the RGB image according to claim 1, wherein Apply the trained hand rotation angle regression model to the end of the robotic arm, and change the posture by remotely operating the end of the robotic arm through hand rotation. The specific steps include: Input the hand rotation image at the initial moment and the hand rotation image at the current moment into the trained hand rotation angle regression model to obtain the hand rotation angle; Map the hand rotation angle to the end pose in the base coordinate system of the robotic arm.
8. A hand rotation angle regression system based on RGB images, characterized in that, It includes: Hand rotation angle dataset construction module, dataset division module, feature extraction module, weighted fusion module, regression model training module, optimal regression model output module, pose control module; The hand rotation angle dataset construction module is used to collect hand rotation images based on an RGB camera and label hand rotation angle tags to construct a hand rotation angle dataset; The dataset division module is used to divide the hand rotation angle dataset into a training set and a test set; The feature extraction module is used to extract features from the hand rotation images of the hand rotation angle dataset based on a two-stream convolutional structure to obtain global visual features and local geometric features; The weighted fusion module is used to perform weighted fusion on the global visual features and local geometric features; The regression model training module is used to obtain a training set, train the hand rotation angle regression model based on the difference between the hand rotation angle obtained by regression and the hand rotation angle of the label value, and update the parameters of each layer of the network in the hand rotation angle regression model; The optimal regression model output module is used to verify the accuracy of the updated hand rotation angle regression model based on the test set. When the mean absolute error is the smallest, the updated hand rotation angle regression model is used as the trained hand rotation angle regression model; The pose control module is used to apply the trained hand rotation angle regression model to the end of the robotic arm, and change the pose by remotely operating the end of the robotic arm through hand rotation.
9. A computer-readable storage medium storing a program, characterized in that, When the program is executed by a processor, it implements the RGB image-based hand rotation angle regression method according to any one of claims 1-7.
10. A computer device, comprising a processor and a memory for storing processor-executable programs, characterized in that, When the processor executes the program stored in the memory, it implements the RGB image-based hand rotation angle regression method according to any one of claims 1-7.