A fingertip detection method based on visual features
The finger tip detection method employs rotated HOG features and AdaBoost learning to enhance hand gesture detection accuracy in complex environments, addressing lighting and background interference issues.
Patent Information
- Application Number
- CN202210782400.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-05
AI Technical Summary
Traditional fingertip detection methods have poor results in complex backgrounds and rely heavily on gesture detection results, resulting in inaccurate fingertip detection.
Using the fingertip detection method based on HOG features, the rotation-invariant HOG features and AdaBoost integrated learning is used, and the fingertip points are accurately detected in complex environments by combining the Mean Shift algorithm.
It realizes accurate detection of one-finger fingertip points in complex environments, with grayscale, lighting and rotation invariance, improving the robustness and accuracy of detection.
Smart Images

Figure CN115171158B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robots and the related technology of human-computer interaction, and particularly to a fingertip point detection method based on visual features applicable to complex environments. Background Art
[0002] With the development of society, human-computer interaction has gradually become an essential part of people's lives. The current human-computer interaction mainly includes contact type and non-contact type. Among them, contact human-computer interaction refers to realizing interaction by operating tools such as a mouse, a keyboard, and a touch screen. However, this method is easily limited by hardware devices and increasingly fails to meet people's needs. Therefore, non-contact human-computer interaction has emerged as the times require. Gesture, as the most natural human-computer interaction method, has become one of the research hotspots. In order to realize human-computer interaction based on gestures, gesture detection needs to be carried out first.
[0003] Currently, the commonly used methods for gesture detection are data glove-based detection and computer vision-based detection. The data glove-based gesture detection method needs to rely on a data glove. This kind of glove generally needs to install an inertial sensor, and the gesture posture information is obtained through the attitude inclination angle of the inertial sensor. However, it is very difficult to obtain the position information of the gesture by this method. The computer vision-based method does not require wearing a glove on the hand. Only a camera needs to be configured at the position of the operator's head or glasses, and the gesture can be detected by capturing the hand image through the camera. However, gestures based on the computer vision method are usually based on skin color detection or the detection of other color markers. Such simple methods are very easily interfered by light or similar colors in the background and are only suitable for simple background scenes. For complex scenes with similar color interference or even drastic light changes, the gesture detection effect based on color will be greatly reduced or even unusable. Fingertip detection is an important part of gesture detection. It can accurately describe the position of the operator's gesture, and natural human-computer interaction, such as touch control in a virtual space, can be realized through the fingertip position. The fingertip detection method usually needs to first obtain the gesture detection result, and then further realize the fingertip point detection through methods such as the center of gravity distance method or morphological features on the basis of the gesture detection result. If the gesture detection result is inaccurate, the fingertip point detection result will also be inaccurate. That is to say, the traditional fingertip point detection method highly depends on the gesture detection result. In a complex background, the visual-based gesture detection effect is very poor, which also leads to a very poor traditional fingertip point detection effect. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a fingertip detection method based on HOG (Histogram of Oriented Gradient) features in a complex environment. The rotation-invariant HOG features of the edge image are used, and effective features are classified through AdaBoost ensemble learning training. A rotating offset vector is used to locate the fingertip, realizing the accurate detection of the single-finger fingertip point in a complex environment.
[0005] The present invention provides the following technical solutions: A fingertip detection method based on visual features includes the following steps: mainly including a model training stage and a testing stage, where steps 1-11 are the first-stage model training stage, and steps 12-20 are the second-stage testing stage;
[0006] Step 1, Take two sets of images, one is a hand image with a background having a color difference from the hand color, that is, a hand image with a clean background, and the other is a background image in a complex environment without a hand;
[0007] Step 2, For the hand image obtained in step 1, use a simple background hand segmentation method to segment the hand image, and sequentially perform grayscale conversion and binarization on the segmented image;
[0008] Step 3, For the hand image obtained in step 2, detect the fingertip point coordinates of each image and save them;
[0009] Step 4, For the hand image and the background image obtained in step 1, use the HED method for edge detection and binarization to obtain a hand edge image and a background edge image;
[0010] Step 5, Extract the rotation-invariant HOG features of each edge point in the hand edge image, and save the edge point HOG features, the position of the features, the offset vector of the features from the fingertip point, and the main direction of the features;
[0011] Step 6, Select effective features from the rotation-invariant HOG features of the hand according to the effective feature criterion, and save the HOG features, the position of the features, the offset vector of the features from the fingertip point, and the main direction of the features as a feature dictionary;
[0012] Step 7, Perform agglomerative hierarchical clustering on the feature dictionary for subsequent searching;
[0013] Step 8, Extract the rotation-invariant HOG features of the edge points in the background edge image;
[0014] Step 9, Make a data set, which includes a training set, a validation set, and a test set;
[0015] Step 10: Design a feature classifier; use AdaBoost ensemble learning as the feature classifier, which consists of M multi-layer perceptrons (MLPs) as weak classifiers to form a strong classifier;
[0016] Step 11: Train the feature classifier with the training set, adjust the parameters through the validation set, select the optimal model, and retain the model information;
[0017] Step 12: Collect hand images in a complex environment, and the hand posture should be consistent with the requirements of the training set;
[0018] Step 13: Use the HED edge detection method to perform edge detection and binary processing on the images obtained in Step 12;
[0019] Step 14: Extract rotation-invariant HOG features from the edge detection images obtained in Step 13;
[0020] Step 15: Use the trained AdaBoost feature classifier model to classify the predicted valid features from the HOG features obtained in Step 14;
[0021] Step 16: Compare the predicted valid features obtained in Step 15 with the valid features in the feature dictionary after rotating them by the same angle to the same main direction;
[0022] Step 17: Use the KNN algorithm to find the K features that best match each valid feature from the predicted valid features, and define them as quasi-valid features;
[0023] Step 18: Rotate the offset vectors of the valid features matching each quasi-valid feature by the same angle;
[0024] Step 19: Use each quasi-valid feature and the corresponding rotated offset vector to predict the fingertip coordinates, forming a fingertip point space;
[0025] Step 20: Use the Mean Shift algorithm to find the point with the maximum predicted density of fingertip points in the fingertip point space to obtain the finally predicted fingertip points.
[0026] In Step 1, the background with a color difference from the hand color ensures that all the features extracted subsequently come from the hand. The hand posture in the hand image should be a rotational transformation with any four fingers closed and the remaining one finger extended, with diverse distance transformations;
[0027] During the segmentation in Step 2, convert the original RGB image to a YCbCr image, use the Gaussian model of skin color to calculate the similarity between the input image and the skin color image, set a threshold to segment the hand, and finally denoise and smooth the image through median filtering.
[0028] In step 3, the simple fingertip detection algorithm is used to detect the fingertip point coordinates of each image. The specific simple fingertip detection algorithm is as follows:
[0029] Step 31: Obtain the binary hand segmentation map segmented in step 2;
[0030] Step 32: Use the Canny operator edge detection method to find the gesture contour from the image;
[0031] Step 33: Calculate the centroid of the gesture contour by finding the zeroth moment M 00 of the gesture contour, the first moment M 01 , M 10 ; the centroid of the gesture contour is the centroid of the hand;
[0032] Step 34: Find the point farthest from the centroid among the gesture contour points, and this point is the fingertip point;
[0033] Step 35: Record and save the fingertip point coordinates.
[0034] The specific HED (Holistically-Nested Edge Detection) edge detection method in step 4 is as follows:
[0035] Step 41: Construct the HED network structure; the HED network is improved based on VGG16. Compared with the VGG16 network, first, the output of each layer of convolution in the VGG16 network is connected to the output layer, that is, the weighted fusion layer. Second, the 5th pooling layer and all fully connected layers in the VGG16 network are removed, enabling training and prediction for images of any size.
[0036] Step 42: Determine the loss function;
[0037] The loss function of the HED network consists of two parts: the side output layer loss L side and the fusion weight loss L fuse .
[0038] Let the input image be |X n | represent the number of pixels contained in the nth image, and its corresponding label is The HED network has 5 side output layers, and each side output layer is associated with a classifier. Then, the weights of each layer are defined as w = (w (1) , …, w (5) ), and the remaining parameter values in the network are all W.
[0039]
[0040]
[0041] Among them, α m is the weight of each side output layer loss function and can be set according to the training log or set to 1 / 5. β is used to solve the problem of unbalanced numbers of edge pixels and non-edge pixels. |Y - | and |Y + | represent the numbers of edge pixel points and non-edge pixel points respectively. Pr represents the prediction result. Therefore, P r (y j =1|X; W, w (m) ) is the predicted value output by the m-th side output layer after being calculated by the Sigmoid function σ, that is
[0042]
[0043]
[0044]
[0045] Among them, L fuse (W, w, h) represents the fusion loss, which is the weighted sum of each side output layer and is the fused prediction result. h m is the fusion weight of the m-th side output layer. The Dist function is used to calculate the distance between the label Y and the fused prediction result therebetween.
[0046] Step 43: Perform training and testing in sequence and retain the trained model.
[0047] (W, w, h) * = arg min(L side (W, w) + L fuse (W, w, h)) (6)
[0048] Among them, (W, w, h) * represents the model that minimizes the sum of L side and L fuse .
[0049] In Step 5, the rotation-invariant HOG features of each edge point in the hand edge image are extracted. The core idea of the HOG features is to describe the shape of the object of interest in the image through the gradient or edge direction density distribution, and its essence is to statistically analyze the gradient information. The rotation-invariant HOG feature extraction only adds the rotation-invariant property to it on the basis of HOG. When extracting features:
[0050] Step 51: Read the edge image and perform normalization processing.
[0051] Step 52: Let H(x, y) be the image after being processed in Step 51, and calculate the gradient magnitude m(x, y) and direction θ(x, y) of the edge image.
[0052]
[0053]
[0054] Step 53: Calculate the gradient histogram.
[0055] When calculating the gradient histogram of HOG features, a block is used as the sampling window form. One block contains n×n cells, and one cell contains a×a pixels. The gradient direction is used to weight and project each pixel in the cell, and the gradient magnitude therein is used as the weight value of the projection.
[0056] Step 54: Perform normalization processing on the block.
[0057] Taking the block as a unit, perform L2 regularization on the gradient intensity. This can effectively reduce the influence of illumination on HOG features. The normalization formula is as follows:
[0058]
[0059] Among them, s n is the normalization result, x n is the corresponding block vector, and ξ is a very small positive number.
[0060] Step 55: Taking a certain point as the center, find the HOG feature histogram of the block where it is located as the HOG feature of this point. The angle with the largest gradient amplitude in the block histogram is used as the main direction of the block, that is, the main direction of the HOG feature of this point. Record the HOG features, position coordinates, and main directions of each point.
[0061] Step 56: Connect the HOG features of all points in the edge image to form the HOG feature of the entire image.
[0062] The specific effective feature extraction criterion described in Step 6 is as follows:
[0063] Step 61: The L2 between the fingertip position predicted by the effective feature and the true fingertip position.
[0064] Distance j (F i ) = ||P j (F i ) - Fingertip(j)||2 <d (10)
[0065] P j (F i )=Q j (n)-Offset(F i ) (11)
[0066]
[0067] Among them, F i represents the i-th feature, j represents the j-th hand image, and P j (F i ) represents the fingertip coordinates predicted by the feature most similar to F i in the j-th image. Fingertip(j) represents the true fingertip coordinates in the j-th image. Therefore, Distance j (F i ) represents the L2 norm of the predicted fingertip coordinates and the true fingertip coordinates in the j-th image. d is a threshold. If it is less than the threshold d, it is considered to satisfy "very small". n represents the label of the feature most similar to the feature F i in the j-th image. Q j (n) represents the coordinates of the feature most similar to F i in the j-th image. Offset(F i ) represents the offset vector between the feature F i and the true fingertip coordinates.
[0068] Step 62: On the basis of the condition in 61, the effective features should also satisfy that they have appeared in most hand images. In other words, the frequency of appearance of the effective features in the hand images exceeds the set threshold.
[0069]
[0070] Among them, α represents the indicator function. If Distance j (F i ) < d, the value is 1, otherwise it is 0. T is the set threshold, and P N is the number of hand images.
[0071] In Step 7, the agglomerative hierarchical clustering method is a bottom-up clustering method. The steps of agglomerative hierarchical clustering are as follows:
[0072] Step 71: Initialize each sample in the training sample as a cluster.
[0073] Step 72: Calculate the distance between any two clusters, and merge the two clusters with the closest distance.
[0074] Step 73. Repeat Step 72 until 90% of the clusters are merged.
[0075] The data set production process described in Step 9 is as follows:
[0076] Step 91. The data set includes a valid feature set and a background feature set. Among them, the data of the valid hand feature set is the set of valid HOG features selected in Step 6, and the data of the background feature set is the background HOG feature obtained in Step 8. Step 92. Assign labels of 1 and 0 to the valid feature set and the background feature set respectively. Step 93. Combine the valid feature set and the background feature set into a data set, shuffle the order, and divide it into a training set of 70%, a validation set of 10%, and a test set of 20%.
[0077] The steps of the AdaBoost feature classifier classification described in Step 10 are specifically as follows:
[0078] Step 101. Initialize the weights of each sample and assign them the same initial value.
[0079] Let the total number of samples be N, and D1(i) be the weight of the initial i-th sample, where i = 1, 2, 3,..., N.
[0080]
[0081] Step 102. Calculate the error rate of each weak classifier, and select the weak classifier with the smallest error rate as the base classifier C for this iteration.
[0082]
[0083] Among them, e t is the error rate, t is the number of iterations, T is the set maximum number of iterations, D t (i) is the weight of each sample at the t-th iteration, C t (x i ) is the predicted value of the i-th sample at the t-th iteration, y i is the true label value of the i-th sample, y i = {-1, 1}. If the i-th sample is a valid feature, then y i = 1; otherwise, y i = -1.
[0084] Step 103. Calculate the weight α t ,
[0085]
[0086] that the base classifier C occupies in the final strong classifier at the t-th iteration.
[0087]
[0088] Step 105. Repeat steps 102 - 104 until the maximum number of iterations T is satisfied.
[0089] Step 106. Combine according to the weights of each base classifier to obtain the strong classifier model C. final ,
[0090]
[0091] where sign is the sign function.
[0092] The method of rotating to the same angle described in steps 16 and 18 is as follows: Let (x, y) be the original coordinates, θ be the rotation angle of the coordinate axis, and (x′, y′) be the rotated coordinates. The formula is as follows:
[0093]
[0094] The specific KNN method described in step 17 is as follows: Given a training sample set with labels and a test sample, the KNN classifier compares the data features of the test sample with the corresponding data features in the training set, and obtains the classification labels of the K nearest training samples in the training set according to the set distance metric, and selects the class label that appears the most among the K samples as the classification of the test sample.
[0095] The specific Mean Shift method described in step 20 is as follows: Mean Shift is a non - parametric density gradient estimation method. The steps are as follows: Given n sample points x in the d - dimensional space R d , i = 1, 2, 3, …, n, the Mean Shift vector at the sample point x can be expressed as: i
[0096]
[0097] where k represents the number of samples falling within the high - dimensional ball region, S h h T 2 represents the high - dimensional ball region with radius h, satisfying S T (x) ≡ {y: (y - x) 2 (y - x) ≤ h i}. x i - x represents the offset vector of the sample point x h relative to the point x, and M h is obtained by summing the offset vectors of the k sample points falling within the high - dimensional ball region S
[0098] Let h be the bandwidth, d be the space dimension, K(x) be the kernel function, ck,d is a constant is the unit density, then the Gaussian kernel multivariate kernel density estimation function at the sample point x can be expressed as:
[0099]
[0100] The formula for iterative update of the sample point x is as follows:
[0101]
[0102] The termination condition for convergence is ||x - y i || ≤ P, where P is the threshold, that is, the allowable error.
[0103] From the above description, it can be seen that the beneficial effects of this solution compared with the prior art are as follows: (1) The present invention uses rotation-invariant HOG features to locate the effective features of the hand, which has gray-scale invariance, illumination invariance, and rotation invariance, enabling fingertip detection to be unaffected by illumination, hand skin color, and hand rotation transformation. Even if the operator wears gloves, it will not affect the detection result. (2) The method of AdaBoost ensemble learning is used to classify effective features and background features, which can accurately detect the fingertip position in a complex background. (3) The present invention uses the rotation-invariant HOG features of the hand and the rotation offset vector of the fingertip point to locate the fingertip. Compared with the modular combined detection method, it has the advantages of fast detection speed, strong robustness, and high accuracy. (4) In the formation process of human brain vision, the first information obtained is edge and direction feature information, and human vision is very sensitive to edges. The HOG features of edge points used in the present invention have bionic characteristics and greatly reduce the amount of calculation. (5) By using the method of predicting the fingertip based on hand parts, the anti-interference ability can be improved, and accurate prediction can still be achieved when there is partial occlusion of the hand. Description of the Drawings
[0104] Figure 1 is the flowchart of the first-stage model training phase of the specific implementation manner of the present invention.
[0105] Figure 2 is the flowchart of the second-stage testing phase of the specific implementation manner of the present invention.
[0106] Figure 3 is the overall block diagram of the specific implementation manner of the present invention.
[0107] Figure 4 is the HED edge detection structure diagram of the specific implementation manner of the present invention. Specific Implementation Manner
[0108] The technical solutions in the specific embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the specific embodiments of the present invention. Obviously, the described specific embodiments are only specific embodiments of the present invention, rather than all specific embodiments. All other specific embodiments obtained by those of ordinary skill in the art based on the specific embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0109] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof;
[0110] As can be seen from the accompanying drawings, the detection method of this solution of the present invention is divided into two stages. Steps 1-11 are the first-stage model training stage, and steps 12-20 are the second-stage testing stage.
[0111] Step 1, take two sets of images.
[0112] One set is the hand image P with a clean background j , j = 1, 2, 3,..., and one set is the background image B of complex and variable backgrounds without hands m , m = 1, 2, 3,.... The background is clean to ensure that all the features extracted subsequently come from the hand. The hand posture should be a rotational transformation with any four fingers closed and the remaining one finger extended, with diverse distance transformations.
[0113] Step 2, for the hand image P obtained in Step 1 j , use a simple background hand segmentation method to segment the hand, and perform grayscale conversion and binarization processing in sequence to obtain PB j .
[0114] Step 3, for the hand image PB obtained in Step 2 j , use a simple fingertip detection algorithm to detect the fingertip point coordinates Fingertip(j) of each image and save them.
[0115] Step 4, for the hand image P j and the background image B m obtained in Step 1, use the HED method for edge detection and binarization processing to obtain the hand edge map PE j and the background edge map BE m .
[0116] Step 5, Extract the hand edge image PE j The rotation invariant HOG feature F of each edge point in i , i = 1, 2, 3, …, and save the HOG feature PE of the edge point j (F i ), the position Position(i) of the feature, the offset vector Offset(F i ) from the feature to the fingertip point, and the main direction D(i) of the feature.
[0117] Step 6, Select the effective features PE from the hand rotation invariant HOG features using the effective feature criterion j (F c ), c ∈ i, and save the position Position(c) of the corresponding feature, the offset vector Offset(F c ) from the feature to the fingertip point, and the main direction D(c) of the feature as the feature dictionary Dictionary{PE j (F c ), Position(c), Offset(F c ), D(c)}.
[0118] Step 7, Perform agglomerative hierarchical clustering on the feature dictionary for subsequent searching.
[0119] Step 8, Extract the background edge image BE m The rotation invariant HOG feature BE of the edge points in m (F g ), g = 1, 2, 3, ….
[0120] Step 9, Make a dataset.
[0121] Step 10, Design a feature classifier.
[0122] Use AdaBoost ensemble learning as the feature classifier, which consists of 10 multi-layer perceptrons MLP as weak classifiers to form a strong classifier.
[0123] Step 11, Train the feature classifier with the training set, adjust the parameters through the validation set, select the optimal model, and retain the model Model information.
[0124] So far, the model training is completed, and the flowchart is as Figure 1 shown. Next, enter the testing phase.
[0125] Step 12, Collect hand images C in a complex environment d , d = 1, 2, 3, …, and the hand posture should be consistent with the requirements of the training set.
[0126] Step 13: Use the HED edge detection method to perform edge detection and binary processing on the image C obtained in Step 12 d to obtain the image CE d .
[0127] Step 14: Extract the rotation-invariant HOG features F h , h = 1, 2, 3, … and the main direction D(h) of the features
[0128] Step 15: Use the trained AdaBoost feature classifier model model to classify the predicted valid features Predict(s) from the HOG features obtained in Step 14, where s ∈ h
[0129] Step 16: Compare the predicted valid features Predict(s) obtained in Step 15 with the valid features PE j (F c ) rotated by θ a , a = 1, 2, 3, … to the same main direction
[0130] Step 17: Use the KNN algorithm to find the 3 features that best match each valid feature from the predicted valid features, defined as the quasi-valid features PE j (F z ), z ∈ c
[0131] Step 18: Rotate the offset vector Offset(F j (F z ) of the valid feature PE j (F c ) that matches each quasi-valid feature PE c ) by the same angle θ a to become NewOffset(F c )
[0132] Step 19: Use each quasi-valid feature PE j (F z ) and the corresponding, rotated offset vector NewOffset(F c ) to predict the fingertip coordinates, forming a fingertip space F
[0133] Step 20: Use the Mean Shift algorithm to find the point with the highest predicted density of fingertip points from the fingertip space F to obtain the finally predicted fingertip Fingertip final .
[0134] So far, the test phase ends. The flowchart of the test phase is as shown in Figure 2 and the overall block diagram is as shown in Figure 3as shown
[0135] The simple background hand segmentation described in step 2 is specifically as follows:
[0136] Step 21: Convert the original RGB image to a YCbCr image, use the Gaussian model of skin color to calculate the similarity between the input image and the skin color image, set a reasonable threshold to segment the hand, and finally denoise through median filtering to smooth the image.
[0137] Specifically, the simple fingertip detection algorithm described in step 3 is specifically as follows:
[0138] Step 31: Obtain the binary hand segmentation map obtained by segmentation in step 2; Step 32: Use the Canny operator edge detection method to find the gesture contour from the image; Step 33: Calculate the centroid of the gesture contour by finding the zero-order moment M 00 and the first-order moment M 01 and M 10 to calculate the centroid of the gesture contour which is the centroid of the hand; Step 34: Find the point farthest from the centroid from the gesture contour points, and this point is the fingertip point; Step 35: Record the fingertip point coordinates and save them.
[0139] Specifically, the HED (Holistically-Nested Edge Detection) edge detection method described in step 4 is specifically as follows:
[0140] Step 41: Construct the HED network structure.
[0141] The HED network is improved based on VGG16. Compared with the VGG16 network, first, the HED network connects the output of each layer of convolution in the VGG16 network to the output layer, that is, the weighted fusion layer. Second, the 5th pooling layer and all fully connected layers in the VGG16 network are removed, realizing training and prediction for images of any size.
[0142] Step 42: Determine the loss function.
[0143] The loss function of the HED network includes two parts: the side output layer loss L side and the fusion weight loss L fuse .
[0144] Let the input image be |X n | represents the number of pixels contained in the nth image, and its corresponding label is The HED network has 5 side output layers, and each side output layer is associated with a classifier, then the weights of each layer are defined as w = (w (1) ,…,w (5) ), and the other parameter values in the network are all W.
[0145]
[0146]
[0147] Among them, α m is the weight of the loss function of each side output layer and can be set according to the training log or set to 1 / 5. β is used to solve the problem of unbalanced numbers of edge pixels and non-edge pixels. |Y - | and |Y + | represent the numbers of edge pixel points and non-edge pixel points respectively. Pr represents the prediction result. Therefore, P r (y j = 1|X; W, w (m) ) is the predicted value output by the m-th side output layer after being calculated by the Sigmoid function σ, that is
[0148]
[0149]
[0150]
[0151] Among them, L fuse (W, w, h) represents the fusion loss, which is the weighted sum of each side output layer and is the fused prediction result. h m is the fusion weight of the m-th side output layer. The Dist function is used to calculate the distance between the label Y and the fused prediction result therebetween.
[0152] Step 43: Train and test sequentially and retain the trained model. The HED network structure diagram is as Figure 4 shown.
[0153] (W, w, h) * = arg min(L side (W, w) + L fuse (W, w, h)) (6)
[0154] Among them, (W, w, h) * represents the model that minimizes the sum of L side and L fuse .
[0155] Specifically, the rotation-invariant HOG feature extraction described in step 5 is specifically as follows:
[0156] The core idea of HOG features is to describe the shape of the object of interest in the image through the gradient or edge direction density distribution, and its essence is to statistically analyze the gradient information. The rotation-invariant HOG feature extraction only adds the rotation-invariant property to HOG on the basis of HOG. The process of rotation-invariant HOG feature extraction is as follows:
[0157] Step 51: Read the edge image and perform normalization processing.
[0158] Step 52: Let H(x, y) be the image processed in Step 51, and calculate the gradient magnitude m(x, y) and direction θ(x, y) of the edge image.
[0159]
[0160]
[0161] Step 53: Calculate the gradient histogram.
[0162] When calculating the gradient histogram of HOG features, a block is used as the sampling window form. One block contains n×n cells, and one cell contains a×a pixels. The gradient direction is used to weight and project each pixel in the cell, and the gradient magnitude therein is used as the weight of the projection.
[0163] Step 54: Perform normalization processing on the block.
[0164] Taking the block as a unit, perform L2 regularization on the gradient intensity, which can effectively reduce the influence of illumination on HOG features. The normalization formula is as follows:
[0165]
[0166] where s n is the normalization result, x n is the corresponding block vector, and ξ is a very small positive number.
[0167] Step 55: Taking a certain point as the center, find the HOG feature histogram of the block where it is located as the HOG feature of this point. The angle with the largest gradient amplitude in the block histogram is used as the main direction of the block, that is, the main direction of the HOG feature of this point. Record the HOG features, position coordinates, and main directions of each point.
[0168] Step 56: Connect the HOG features of all points in the edge image to form the HOG feature of the entire image.
[0169] Specifically, the effective feature extraction criterion described in Step 6 is specifically as follows:
[0170] Step 61. The L2 norm between the fingertip position predicted by the effective feature and the true fingertip position is very small.
[0171] Distance j (F i ) = ||P j (F i ) - Fingertip(j)|| 2 < d (10)
[0172] P j (F i ) = Q j (n) - Offset(F i ) (11)
[0173]
[0174] Wherein, F i represents the i-th feature, j represents the j-th hand image, P j (F i ) represents the fingertip coordinates predicted by the feature most similar to F i in the j-th image, and Fingertip(j) represents the true fingertip coordinates in the j-th image. Therefore, Distance j (F i ) represents the L2 norm between the predicted fingertip coordinates and the true fingertip coordinates in the j-th image, d is the threshold, and if it is less than the threshold d, it is considered to satisfy "very small". n represents the label of the feature most similar to the feature F i in the j-th image, and Q j (n) represents the coordinates of the feature most similar to F i in the j-th image, and Offset(F i ) represents the offset vector between the feature F i and the true fingertip coordinates.
[0175] Step 62. On the basis of the condition in 61, the effective feature should also satisfy that it appears in most hand images. In other words, the frequency of the effective feature appearing in the hand images exceeds the set threshold.
[0176]
[0177] Wherein, α represents the indicator function. If Distance j (F i ) < d, the value is 1, otherwise it is 0. T is the set threshold, and P N is the number of hand images.
[0178] Specifically, the agglomerative hierarchical clustering method described in step 7 is specifically as follows:
[0179] The agglomerative hierarchical clustering method is a bottom-up clustering method, and the clustering steps are as follows:
[0180] Step 71: Initialize each sample in the training samples as a cluster.
[0181] Step 72: Calculate the Euclidean distance between any two clusters, and merge the two closest clusters.
[0182] Step 73: Repeat Step 72 until 90% of the clusters are merged.
[0183] Specifically, the data set production process described in Step 9 is as follows:
[0184] Step 91: The data set includes a valid feature set and a background feature set. Among them, the data of the hand valid feature set is the valid HOG feature set selected in Step 6, and the data of the background feature set is the background HOG feature obtained in Step 8. Step 92: Assign labels of 1 and 0 to the valid feature set and the background feature set respectively. Step 93: Combine the valid feature set and the background feature set into a data set, and randomly shuffle and divide it into a training set of 70%, a validation set of 10%, and a test set of 20%.
[0185] Specifically, the steps of the AdaBoost feature classifier classification described in Step 10 are as follows:
[0186] Step 101: Initialize the weights of each sample and assign them the same initial value.
[0187] Let the total number of samples be N, and D1(i) be the weight of the initial i-th sample, where i = 1, 2, 3,..., N.
[0188]
[0189] Step 102: Calculate the error rate of each weak classifier, and select the weak classifier with the minimum error rate as the base classifier C for this iteration.
[0190]
[0191] Among them, e t is the error rate, t is the number of iterations, and T is the set maximum number of iterations. D t (i) is the weight of each sample at the t-th iteration, and C t (x i ) is the predicted value of the i-th sample at the t-th iteration, y i is the true label value of the i-th sample, y i = {-1, 1}, if the i-th sample is a valid feature, then y i = 1; otherwise, yi = -1.
[0192] Step 103: Calculate the weight α of the base classifier C in the final strong classifier at the t-th iteration. t .
[0193]
[0194] Step 104: Update the weights of the training samples.
[0195]
[0196] Step 105: Repeat Steps 102 - 104 until the maximum number of iterations T is satisfied.
[0197] Step 106: Combine according to the weights of each base classifier to obtain the strong classifier model C. final .
[0198]
[0199] where sign is the sign function.
[0200] Specifically, the method of rotating to the same angle described in Steps 16 and 18 is as follows:
[0201] Let (x, y) be the original coordinates, θ be the coordinate axis rotation angle, and (x′, y′) be the rotated coordinates. The formula is as follows:
[0202]
[0203] Specifically, the KNN method described in Step 17 is as follows:
[0204] Given a training sample set with labels and test samples, the KNN classifier compares the test sample data features with the corresponding data features in the training set, and obtains the classification labels of the K nearest training samples in the training set according to a specific distance metric, and selects the category label that appears the most among the K samples as the classification of the test sample. This method is also known as the "voting method".
[0205] Specifically, the Mean Shift method described in Step 20 is as follows:
[0206] Mean Shift is a non-parametric density gradient estimation method. Its basic principle is: Given n sample points x in the d-dimensional space R d i, i = 1, 2, 3,..., n, the Mean Shift vector at the sample point x can be expressed as: i , i = 1, 2, 3,..., n, the Mean Shift vector at the sample point x can be expressed as:
[0207]
[0208] Among them, k represents the number of samples falling within the high-dimensional sphere region, and S h represents a high-dimensional sphere region with a radius of h, satisfying S h (x) ≡ {y: (y - x) T (y - x) ≤ h 2}}. x i -x represents the offset vector of the sample point x i relative to the point x, and M h (x) is obtained by summing the offset vectors of the k sample points falling within the high-dimensional sphere region S h relative to the point x and then taking the average value.
[0209] Let h be the bandwidth, d be the spatial dimension, K(x) be the kernel function, and c k,d be a constant, be the unit density, then the Gaussian kernel multivariate kernel density estimation function at the sample point x can be expressed as:
[0210]
[0211] The formula for iterative updating of the sample point x is as follows:
[0212]
[0213] The termination condition for convergence is ||x - y i || ≤ P, where P is the threshold, that is, the allowable error.
[0214] The above are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, the present disclosure can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A fingertip detection method based on visual features, characterized in that It includes the following steps: Step 1, capture two sets of images, one is a hand image with a background having a color difference from the hand color, and the other is a background image in a complex environment without a hand; Step 2, for the hand image obtained in Step 1, use a simple background hand segmentation method to segment the hand image, and sequentially perform grayscale conversion and binaryzation on the segmented image; Step 3, for the hand image obtained in Step 2, detect and save the fingertip point coordinates of each image; Step 4, for the hand image and the background image obtained in Step 1, perform edge detection and binaryzation using the HED method to obtain a hand edge image and a background edge image; Step 5, extract the rotation-invariant HOG features of each edge point in the hand edge image, and save the edge point HOG features, the positions of the features, the offset vectors of the features from the fingertip points, and the main directions of the features; Step 6, select effective features from the rotation-invariant HOG features of the hand using an effective feature criterion, and save the HOG features, the positions of the features, the offset vectors of the features from the fingertip points, and the main directions of the features as a feature dictionary; Step 7, perform agglomerative hierarchical clustering on the feature dictionary; Step 8, extract the rotation-invariant HOG features of the edge points in the background edge image; Step 9, make a data set, which includes a training set, a validation set, and a test set; Step 10, design a feature classifier; use AdaBoost ensemble learning as the feature classifier, and consist of M multi-layer perceptrons MLP as weak classifiers to form a strong classifier; Step 11, train the feature classifier with the training set, adjust the parameters through the validation set, select the optimal model, and retain the model information; Step 12, collect hand images in a complex environment, and the hand postures should be consistent with the requirements of the training set; Step 13, perform edge detection and binaryzation on the image obtained in Step 12 using the HED edge detection method; Step 14, extract the rotation-invariant HOG features from the edge detection image obtained in Step 13; Step 15, use the trained AdaBoost feature classifier model to classify the predicted effective features from the HOG features obtained in Step 14; Step 16, compare the predicted effective features obtained in Step 15 with the effective features in the feature dictionary after rotating the same angle to the same main direction; Step 17, use the KNN algorithm to find the K features that best match each effective feature from the predicted effective features, and define them as quasi-effective features; Step 18, rotate the offset vectors of the effective features matching each quasi-effective feature by the same angle; Step 19, use each quasi-effective feature and the corresponding rotated offset vector to predict the fingertip point coordinates to form a fingertip point space; Step 20, use the Mean Shift algorithm to find the point with the maximum predicted density of fingertip points from the fingertip point space to obtain the finally predicted fingertip points.
2. The fingertip detection method based on visual features according to claim 1, wherein, In step 1, the background with a color difference from the hand color ensures that all the features extracted subsequently come from the hand. In the hand image, the hand posture should be a rotation transformation with any four fingers closed and the remaining one finger extended, and the distance transformation is diverse; During the segmentation in step 2, the original RGB image is converted into a YCbCr image. Using the Gaussian model of skin color, the similarity between the input image and the skin color image is calculated, a threshold is set, and the hand is segmented. Finally, median filtering is used to denoise and smooth the image.
3. The fingertip detection method based on visual features according to claim 2, wherein, In step 3, a simple fingertip detection algorithm is used to detect the fingertip point coordinates of each image. The simple fingertip detection algorithm is specifically as follows: Step 31: Obtain the binary hand segmentation map segmented in step 2; Step 32: Use the method of Canny operator edge detection to find the gesture contour from the image; Step 33, calculate the centroid of the gesture contour by finding the zeroth moment M 00 of the gesture contour, the first moment M 01 , M 10 ; the centroid of the gesture contour is calculated, which is the centroid of the hand ; that is, the centroid of the hand Step 34: Find the point farthest from the centroid from the gesture contour points, and this point is the fingertip point; Step 35: Record the fingertip point coordinates and save them.
4. The fingertip detection method based on visual features according to claim 3, wherein, The HED edge detection method in step 4 is specifically as follows: Step 41: Construct the HED network structure; Step 42: Determine the loss function; The loss function of the HED network consists of two parts: the side output layer loss L side and the fusion weight loss L fuse , Let the input image be and |X n | represent the number of pixels contained in the nth image, and its corresponding label is , . Since the HED network has 5 side output layers and each side output layer is associated with a classifier, the weights of each layer are defined as , and the values of the remaining parameters in the network are all W. (1) (2) Among them, α m is the weight of the loss function of each side output layer and is set according to the training log or set to 1 / 5. β is used to solve the problem of imbalance in the number of edge pixels and non-edge pixels. , , |Y - | and |Y + | represent the number of edge pixel points and non-edge pixel points respectively. Pr represents the prediction result. Therefore, the predicted value output by the m-th side output layer after being calculated by the Sigmoid function σ, that is , (3) (4) (5) Among them, represents the fusion loss, is the weighted sum of each side output layer, that is, the predicted result after fusion, ℎ m is the fusion weight of the m-th side output layer, and the Dist function is used to calculate the distance between the label Y and the predicted result after fusion therebetween. Step 43: Train and test in sequence, and retain the trained model, (6) where (W, w, h) * represents the model that minimizes the sum of side L fuse and L 5. The fingertip detection method based on visual features according to claim 4, wherein, When extracting the rotation-invariant HOG features of each edge point in the hand edge image in step 5, Step 51: Read the edge image and perform normalization processing; Step 52: Let H(x, y) be the image processed in step 51, and calculate the gradient magnitude m(x, y) and direction θ(x, y) of the edge image; (7) (8) Step 53: Calculate the gradient histogram; When calculating the gradient histogram of the HOG feature, it adopts the form of a block as the sampling window. One block contains n×n cells, and one cell contains a×a pixels. The gradient direction is used to weight and project each pixel in the cell, and the gradient magnitude therein is used as the weight of the projection. Step 54: Perform block normalization processing; Taking the block as the unit, perform L2 regularization on the gradient intensity, which can effectively reduce the influence of illumination on the HOG feature. The normalization formula is as follows: (9) where s n is the normalization result, x n is the corresponding block vector, is a very small positive number Step 55: Taking a certain point as the center, find the HOG feature histogram of the block where it is located as the HOG feature of this point. The angle with the largest gradient amplitude in the block histogram is used as the main direction of the block, that is, the main direction of the HOG feature of this point. Record the HOG features, position coordinates, and main directions of each point. Step 56: Connect the HOG features of all points in the edge image to form the HOG feature of the entire image.
6. The fingertip detection method based on visual features according to claim 5, wherein, The effective feature extraction criterion in step 6 is specifically as follows: Step 61: Through the L2 of the fingertip position predicted by the effective feature and the true fingertip position, (10) (11) (12) Among them, F i represents the i-th feature, j represents the j-th hand image, and P j (F j ) represents the fingertip coordinates predicted by the feature that is most similar to F i in the j-th image. Fingertip(j) represents the true fingertip coordinates in the j-th image. Therefore, Distance j (F i ) represents the L2 norm between the predicted fingertip coordinates and the true fingertip coordinates in the j-th image. d is the threshold, and n represents the label of the feature that is most similar to F i in the j-th image. Q j (n) represents the coordinates of the feature that is most similar to F i in the j-th image. Offset(F i ) represents the offset vector between the feature F i and the true fingertip coordinates. Step 62: The frequency of the effective features appearing in the hand image exceeds the set threshold. (13) Among them, α represents an indicator function. If Distance j (F i ) < d, the value is 1; otherwise, it is 0. T is the set threshold, and P N is the number of hand images.
7. The fingertip detection method based on visual features according to claim 6, wherein: The agglomerative hierarchical clustering step in step 7 is as follows: Step 71: Initialize each sample in the training samples as a cluster. Step 72: Calculate the distance between any two clusters, and merge the two closest clusters. Step 73: Repeat step 72 until 90% of the clusters are merged.
8. The fingertip detection method based on visual features according to claim 7, wherein: The data set production process in step 9 is as follows: Step 91: The data set includes an effective feature set and a background feature set. Among them, the hand effective feature set data is the effective HOG feature set selected in step 6, and the background feature set data is the background HOG feature obtained in step 8. Step 92: Assign labels of 1 and 0 to the effective feature set and the background feature set respectively. Step 93: Combine the effective feature set and the background feature set into a data set, and randomly shuffle and divide it into a training set of 70%, a validation set of 10%, and a test set of 20%.
9. The fingertip detection method based on visual features according to claim 8, wherein: The steps of the AdaBoost feature classifier classification in step 10 are specifically as follows: Step 101: Initialize the weights of each sample and assign them the same initial value. Let the total number of samples be N, and D1(i) be the weight of the i-th sample in the initial stage, where i = 1, 2, 3, …, N. (14) Step 102: Calculate the error rates of each weak classifier, and select the weak classifier with the minimum error rate as the base classifier C for this iteration. (15) Among them, e t is the error rate, t is the number of iterations, T is the set maximum number of iterations, D t (i) is the weight of each sample at the t-th iteration, C t (x i ) is the predicted value of the i-th sample at the t-th iteration, y i is the true label value of the i-th sample, y i ={-1, 1}, if the i-th sample is a valid feature, then y i = 1; otherwise, y i = -1, Step 103: Calculate the weight α of the base classifier C in the final strong classifier at the t-th iteration. t , (16) Step 104: Update the weights of the training samples. (17) Step 105: Repeat steps 102 - 104 until the maximum number of iterations T is satisfied. Step 106: Combine according to the weights of each base classifier to obtain a strong classifier model C final , (18) Among them, sign is the sign function.
10. The fingertip detection method based on visual features according to claim 9, wherein: The method of rotating to the same angle described in Steps 16 and 18 is as follows. Let (x, y) be the original coordinates and θ be the rotation angle of the coordinate axes. are the coordinates after rotation, and the formula is as follows: (19) The KNN method in step 17 is specifically as follows: Given a training sample set with labels and a test sample, the KNN classifier compares the data features of the test sample with the corresponding data features in the training set, and obtains the classification labels of the K closest training samples in the training set according to the set distance metric, and selects the category label that appears the most among the K samples as the classification of the test sample. The Mean Shift method described in step 20 is specifically as follows: Given n sample points x d in the d-dimensional space R i , i = 1, 2, 3, …, n, the Mean Shift vector at the sample point x is expressed as: (20) Among them, k represents the number of samples falling within the high-dimensional sphere region, and S h represents the high-dimensional sphere region with a radius of h, satisfying , x i -x represents the offset vector of the sample point x i relative to the point x, and M h (x) is obtained by summing the offset vectors of the k sample points falling within the high-dimensional sphere region S h relative to the point x and then taking the average. Let \(h\) be the bandwidth, \(d\) be the spatial dimension, \(K(x)\) be the kernel function, and \(c\) k,d be a constant, be the unit density. Then the Gaussian kernel multivariate kernel density estimation function at the sample point \(x\) is expressed as: (21) The formula for iterative update of the sample point x is as follows: (22) The termination condition for convergence is ||x - y i || ≤ P, where P is the threshold value, i.e., the allowable error.
Citation Information
Patent Citations
Snow pressing vehicle appearance defect detection method based on vision under complex background
CN113160192A
Posture-adjusted calculation of physiological signals
US20190313915A1