Improved YOLO11 hand key point detection method
By improving the YOLO11 network, introducing the spatial attention mechanism and Wensheng graph technology, optimizing the key point loss weight and pruning model, the robustness and real-time performance issues of hand key point detection in complex scenarios are solved, and high-precision and fast detection effects are achieved.
Patent Information
- Application Number
- CN202510555818.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-16
AI Technical Summary
Existing hand key point detection methods have poor robustness when lighting and hand posture changes, are computationally intensive, lack real-time performance, and are difficult to adapt to complex scenes.
An improved YOLO11 network is adopted, a feature fusion module with spatial attention mechanism is introduced, and Wensheng graph technology is combined to simulate hand postures in different environments. The key point loss weights are adjusted, and the model structure is optimized through pruning to improve detection accuracy and real-time performance.
It achieves high-precision and fast detection of hand key points, can adapt to changes in complex scenes, and improves the robustness and flexibility of detection.
Smart Images

Figure CN120656230A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and deep learning technology, specifically to the application of hand key point detection in the fields of human-computer interaction, virtual reality, posture estimation, medical rehabilitation, etc., and more specifically to an improved YOLO11 hand key point detection method. Background Art
[0002] Hand keypoint detection is an important task in gesture recognition and human-computer interaction. Existing hand keypoint detection methods generally rely on traditional computer vision techniques and deep learning models. Early hand keypoint detection was mainly based on image processing and geometric feature analysis. The hand contour was extracted through edge detection algorithms or color segmentation techniques. Then, morphological features such as the convex hull and concave points of the fingers were used to identify keypoints, or predefined hand shape templates were used to match the input image to infer the positions of the finger joints. Traditional methods are very sensitive to changes in lighting, background, and hand posture. Once the environment changes, the robustness of the algorithm is greatly reduced. Moreover, because the algorithm relies on geometric shapes, it is difficult to accurately handle complex gestures and diverse hand postures.
[0003] With the development of deep learning technology, hand keypoint detection methods based on convolutional neural networks (CNNs) have gradually become a research hotspot. These methods can effectively improve detection accuracy and robustness by automatically learning features from images. However, existing deep learning-based hand keypoint detection methods still have some problems, such as high computational complexity, poor real-time performance, and insufficient adaptability to complex scenes. Summary of the Invention
[0004] The embodiment of the present invention provides a hand key point detection method based on an improved YOLO11, which improves the accuracy, real-time performance and adaptability of hand key point detection to complex scenes, thereby providing a better user experience and higher efficiency in applications such as human-computer interaction and hand posture estimation.
[0005] In a first aspect, the present invention provides a hand key point detection method based on an improved YOLO11, comprising:
[0006] Collect hand images under different hand postures and lighting conditions, annotate each hand image, and obtain key point information corresponding to each hand image;
[0007] YOLO11 is used as the basic network architecture, and a feature fusion module based on the spatial attention mechanism is introduced to perform convolution operations on feature maps of different dimensions in the spatial dimension, thereby obtaining an improved YOLO11 network model;
[0008] The annotated hand image is input into the improved YOLO11 network model for training to obtain a trained target YOLO11 network model, which is used to detect the hand image to be detected and obtain the position information of the hand key points.
[0009] In some examples, collecting hand images under different hand postures and lighting conditions, annotating each hand image, and obtaining key point information corresponding to each hand image includes:
[0010] Collect hand gesture images with specific requirements in various scenarios and generate hand images using text-based graph technology to simulate various natural hand movements in different environments;
[0011] Each hand image is preprocessed, and then annotated using an annotation tool to obtain the coordinates of the hand key points corresponding to each hand image.
[0012] In some instances, the feature fusion module based on the spatial attention mechanism is used to replace the connection module in the YOLO11 network.
[0013] In some instances, the feature fusion module based on the spatial attention mechanism is used to obtain a shallow spatial attention feature map and a deep spatial attention feature map by performing a convolution operation on feature maps of different dimensions in the spatial dimension, calculate the spatial attention weights of the shallow spatial attention feature map and the deep spatial attention feature map, multiply the shallow spatial attention feature map and the deep spatial attention feature map by their respective spatial attention weights, and then perform addition or splicing operations to achieve feature fusion.
[0014] In some instances, during the training of the improved YOLO11 network model, by changing the degree of finger bending, the model can learn the characteristics of the key points of the hand under different postures, and by adjusting the key point loss weights, different loss weights are set for different key points.
[0015] In some instances, and Update the joint point position, where the bending degree parameter is k and the finger base joint point is A(x A ,y A ), the fingertip joint point is C(x C ,y C ), the middle knuckle point is B(x B ,y B ).
[0016] In some examples, adjusting the key point loss weights to set different loss weights for different key points includes:
[0017] Assume that there are N key points in the hand key point detection task, where the index set of the fingertip key points is I, and the predicted key point coordinates are The real key point coordinates are y=(y1,y2,...,y N );
[0018] Define the weight vector w=(w1,w2,...,w N ), for the fingertip key point i∈I, set w i = l, for other key points Set w i =m, l>m;
[0019] Depend on Get the weighted loss function.
[0020] In some instances, a decoupled head design is used for each prediction branch in YOLO11 to handle classification and regression tasks separately.
[0021] In some examples, the method further comprises:
[0022] After the model is trained, it is compressed and accelerated through pruning.
[0023] In some examples, compressing and accelerating the model by pruning includes:
[0024] The improved YOLO11 model is trained and its performance is evaluated on the validation set to obtain the benchmark performance index P0;
[0025] Try to cut off each layer in turn, and retrain and evaluate the performance of the improved YOLO11 model on the validation set P i , P i The performance of the model on the validation set after pruning the i-th layer;
[0026] If |P i If -P0| is less than a set tolerance ΔP, the i-th layer will be cut off.
[0027] In some examples, compressing and accelerating the model by pruning includes:
[0028] For the entire improved YOLO11 model, according to the selected pruning criteria, the connections that need to be pruned are marked, and then the marked connections are pruned.
[0029] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:
[0030] (1) By improving and optimizing the original YOLO11 network, and using a large amount of training data and effective training methods, this method can achieve high-precision hand key point detection. The improved YOLO11 network has a faster inference speed, and this method can complete hand key point detection in a shorter time, meeting real-time requirements.
[0031] (2) Using the Wensheng graph technique, we simulate various natural hand movements in different environments, such as clenching, stretching, bending, and other posture changes, to increase the diversity of training data. This method can effectively handle the problem of hand key point detection in complex scenes such as lighting changes, posture changes, and occlusion, and has strong adaptability.
[0032] (3) This method can be deployed on different hardware platforms, such as CPU, GPU, NPU, etc., with high flexibility and scalability.
[0033] (4) This method includes: data collection and preprocessing, network structure adjustment and model training and fine-tuning, hand key point detection based on the training model, and post-processing. The first is the collection and preprocessing of hand posture data. According to specific business needs, hand images of different scenes are collected, where these data contain different hand postures and lighting conditions. Using the Wensheng graph technology, various natural hand movements in different environments are simulated, such as fisting, stretching, bending and other posture changes, to increase the diversity of training data. All these data are annotated to obtain the key point information of the hand. The second is model construction. YOLO11 is used as the basic network architecture. According to the characteristics of the hand key point detection task, the YOLO11 network is improved and optimized, such as adding a hand key point detection branch and adding a feature fusion layer. Then the model is trained. The preprocessed image data is input into the constructed model for training. The appropriate loss function and optimization algorithm are used. Considering the difficulty of later deployment, the model will be pruned and quantized. On this basis, the model is fine-tuned to maintain high performance while reducing computational costs. The last step is hand key point detection. The image to be detected is input into the final trained model to obtain the key point position information of the hand. The detection results are post-processed to improve the accuracy and stability of the detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0035] Figure 12 is a schematic diagram of a hand key point detection method based on improved YOLO11 provided in an embodiment of the present invention;
[0036] Figure 2 This is a flowchart of a specific embodiment of the improved YOLO11 hand key point detection provided by an embodiment of the present invention;
[0037] Figure 3 This is the original YOLO11 network structure diagram provided by an embodiment of the present invention;
[0038] Figure 4 1 is a diagram of the improved YOLO11 network structure provided by an embodiment of the present invention;
[0039] Figure 5 This is a schematic diagram of the YOLO layer decoupling head design provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0041] In the following description, specific embodiments of the present invention will be described with reference to steps and symbols performed by one or more computers, unless otherwise specified. Therefore, these steps and operations will be mentioned several times as being performed by a computer, and computer execution as referred to herein includes operations by a computer processing unit that represents electronic signals of data in a structured form. This operation converts the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise change the operation of the computer in a manner familiar to testers in the field. The data structure in which the data is maintained is a physical location in the memory, which has specific characteristics defined by the data format. However, the principles of the present invention are described in the above text, which does not represent a limitation, and testers in the field will understand that the various steps and operations below can also be implemented in hardware.
[0042] As used herein, the terms "module" or "unit" may be considered software objects executed on the computing system. The various components, modules, engines, and services herein may be considered implementation objects on the computing system. While the devices and methods herein are preferably implemented in software, they may also be implemented in hardware and remain within the scope of protection of the present invention.
[0043] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0044] In the first embodiment of the present invention, a hand key point detection method based on improved YOLO11 is provided. Figure 1 As shown, the following steps are included:
[0045] S101: Collect hand images under different hand postures and lighting conditions, annotate each hand image, and obtain key point information corresponding to each hand image;
[0046] S102: Using YOLO11 as the basic network architecture, a feature fusion module based on the spatial attention mechanism is introduced to perform convolution operations on feature maps of different dimensions in the spatial dimension, thereby obtaining an improved YOLO11 network model;
[0047] S103: Input the annotated hand image into the improved YOLO11 network model for training to obtain a trained target YOLO11 network model, and detect the hand image to be detected through the target YOLO11 network model to obtain the position information of the hand key points.
[0048] In an embodiment of the present invention, the public data set can provide a wealth of hand posture samples, but there may be situations that do not conform to the actual application scenario. The data actually captured can better reflect the changes in hand posture under specific application scenarios. Secondly, the currently more successful Wenshengtu technology is used to simulate various natural movements of the hands in different environments, such as changes in postures such as clenching a fist, stretching, and bending, to increase the diversity of training data. The images generated later and the self-collected images are used as training data sets and test sets. The positions of the key points of the hands are annotated by combining manual annotation and automatic annotation. Preprocessing operations can remove noise and interference in the image and improve the quality and consistency of the data. For example, image cropping can remove irrelevant information in the background, and normalization can adjust the pixel value range of the image to a certain range.
[0049] In the embodiment of the present invention, the YOLO11 network is a target detection algorithm based on deep learning, which has high detection speed and good detection accuracy. In the hand key point detection task, the hand can be regarded as a target and detected and located using the YOLO11 network. In order to improve the accuracy of hand key point detection, the YOLO11 network can be improved and optimized. Figure 3 The figure shows the original YOLO11 network structure, which introduces a new feature fusion module, such as Figure 4 This layer represents a new feature fusion layer, which combines shallow, high-resolution features with deep, semantic features for a more refined integration. This allows for the detection of hand keypoints of varying sizes and postures. For subtle hand movements or finger joint keypoints, shallow features provide more accurate location information, while deep features help identify complex hand postures and background information.
[0050] In an embodiment of the present invention, a feature fusion based on a spatial attention mechanism is introduced. The advantage of this fusion method is that it can highlight the important spatial areas in the feature map for detecting key points on the hand. The principle of spatial attention fusion is to first perform a convolution operation on feature maps of different dimensions in the spatial dimension to obtain a spatial attention map, which represents the importance of different spatial positions, and then calculate the spatial attention weights of the shallow and deep feature maps respectively. Then, the shallow and deep feature maps are multiplied by their respective spatial attention weights, and finally, operations such as addition or splicing are performed to realize feature fusion. It is specifically implemented in the following way:
[0051] (1) Calculating spatial attention weights
[0052] For the shallow feature map F s ∈R C×H×W (C is the number of channels, H is the height, and W is the width), through the convolution operation Conv s Calculate the spatial attention weight A s :
[0053] A s =σ(Conv s (F s ))
[0054] Where is the σ activation function, here we assume it is a sigmoid function, that is Convolution operation s (F s ) can be expressed as:
[0055]
[0056] Here we assume that the convolution kernel size is k×k, W ijis the convolution kernel weight, and x,y are the spatial positions on the feature map. The calculation method for spatial attention weights of deep feature maps is similar to that of shallow feature maps and will not be repeated here.
[0057] (2) Feature Fusion
[0058] Weighting shallow feature maps: ( represents element-wise multiplication)
[0059] Weighting deep feature maps:
[0060] Fused feature map (Here we take addition as an example).
[0061] Through the above feature fusion method, the model's detection accuracy of key point positions is improved. According to the characteristics of the hand key points, the key point prediction branch in the prediction head is specially designed. The details are as follows Figure 4 shown.
[0062] In an embodiment of the present invention, in order to optimize the training strategy, in addition to using common data enhancement methods (such as rotation, flipping, cropping, etc.) during the model training process, a specific posture transformation enhancement method is designed based on the characteristics of the hand, so that the model can learn the characteristics of the key points of the hand under different postures and improve the generalization ability of the model. It mainly includes finger joint angle transformation, finger bending degree transformation, finger overall rotation and translation, etc. The main method used is finger bending degree transformation, which generates hand images in different postures by changing the bending degree of the finger. The specific implementation is as follows:
[0063] Define a bending degree parameter k (k∈[0,1]), where 0 means the finger is fully extended and 1 means the finger is maximally bent. For each finger, update the joint point position according to its current joint point position and bending degree parameter k. For example, let the base joint of the finger be A(x A ,y A ), the fingertip joint point is C(x C ,y C ), the middle knuckle point is B(x B ,y B ). When the finger bends, point B moves closer to point C. The new position of point B can be calculated using the following formula:
[0064]
[0065] Keypoint loss weight adjustment: Different loss weights are set for different keypoints based on their importance and detection difficulty. For example, the keypoints of the fingertips are very important for applications such as gesture recognition, but because the fingertips occupy a small proportion in the image and are difficult to detect, a higher loss weight can be set for the keypoints of the fingertips to encourage the model to pay more attention to the detection accuracy of these keypoints. The specific implementation is as follows:
[0066] Assume that there are N key points in the hand key point detection task, where the index set of fingertip key points is I (I = {i1, i2, i3, i4, i5}, which means there are 5 fingertip key points), and the predicted key point coordinates are The real key point coordinates are y=(y1,y2,...,y N ).
[0067] Define a basic loss function, such as the mean squared error loss function to measure the error between the predicted keypoints and the true keypoints:
[0068]
[0069] In the embodiment of the present invention, weights are introduced to adjust the importance of different key points. Define a weight vector w=(w1,w2,...,w N ), for the fingertip key point i∈I, set w i =k (k>1, for example 2), for other key points Set w j =1.
[0070] The final loss function is the weighted loss function: Expand to get:
[0071] In the embodiment of the present invention, fine-tuning mainly compresses and accelerates the model. First, mixed precision training is adopted in training to reduce memory usage and computing time and improve training efficiency. Pruning is mainly carried out from the depth and breadth of the network structure to remove redundant connections and weights, and further reduce the number of model parameters.
[0072] Furthermore, the implementation of deep pruning (layer pruning) from the network structure is as follows:
[0073] First, the original model is trained and its performance on the validation set is evaluated to obtain the benchmark performance index P0;
[0074] Then, try to cut off each layer in turn (you can start from the end of the network and try it forward), and retrain and evaluate the performance of the model on the validation set. For example, after cutting off the i-th layer, the performance of the model on the validation set is P i ;
[0075] If |P i If -P0| is less than a set tolerance ΔP, it means that the i-th layer has little impact on the model performance and can be pruned. Repeat this process until the appropriate number of layers is found for pruning.
[0076] Furthermore, we prune the network structure from its breadth (connection pruning):
[0077] For the entire network, based on the selected pruning criterion, mark the connections that need to be pruned. For example, if the weight magnitude criterion is used, mark all connections whose weight magnitude is less than the threshold T.
[0078] Then, actually pruning these marked connections can be achieved by setting the corresponding weights to or directly deleting the relevant calculation paths.
[0079] In practical applications, the image to be detected is fed into a trained model, which then outputs information about the key points of the hand. This information can include coordinates and confidence levels of the key points. To improve the accuracy and stability of the detection results, post-processing optimization can be performed on the results.
[0080] In the second embodiment of the present invention, Figure 2 The figure shows an improved YOLO11 hand key point detection flow chart provided by the present invention, which is composed of Figure 2 It can be seen that the method includes:
[0081] Step 1: Collect photos of hand postures with specific requirements in various scenarios and generate hand photos using the Wensheng graph technology to simulate various natural hand movements in different environments, such as fist clenching, stretching, bending, and other posture changes. All acquired data is pre-processed, and then these sample images are annotated using annotation tools to obtain the coordinates of the corresponding hand key points.
[0082] Step 2: Model construction, which is modified based on the original YOLO11 network structure;
[0083] Step 3: After the model is built, use the sample images and labels generated in step 1 to train and fine-tune the model, and then obtain the optimal model based on the training process;
[0084] Step 4: Use the optimal model obtained in step 3 for testing, and finally output the hand key point information through post-processing optimization.
[0085] The hand key point detection method provided by the embodiment of the present invention greatly improves the accuracy, real-time performance and adaptability of hand key point detection to complex scenes.
[0086] In the embodiment of the present invention, in step 3, specifically as follows Figure 3 As shown in , a feature fusion module based on spatial attention mechanism is added to the feature extraction part, which enables the model to focus on key areas, such as Figure 4 As shown in the figure, the model employs multiple parallel prediction branches, each with a decoupled head design to handle classification and regression tasks separately. This improves the model's efficiency in independently predicting different types of hand keypoints. The keypoint loss function at the output layer is adjusted, with different loss weights assigned to different hand keypoints based on their importance and detection difficulty.
[0087] Furthermore, the YOLO layer in step 3 is decoupled from the head design, as follows Figure 5 As shown. For the cross entropy loss used in the classification task, it is specifically expressed as:
[0088]
[0089] where y class ∈{1,2,...,N}, the predicted category probability distribution is y class (i, j) is the true category label of position (i, j). If the true category of position (i, j) is k, then y class (i,j,k)=1, and other categories are 0.
[0090] For the bounding box regression task, the SmoothL1 loss is used, which is expressed as follows:
[0091]
[0092] in, is the true bounding box parameter, To predict the bounding box regression results,
[0093] In an embodiment of the present invention, in step 4, all training samples obtained in step 2 are input into the network structure constructed in step 3 for model training. Mixed precision is used in the training process, which can greatly reduce memory usage and computing time and improve training efficiency.
[0094] In this embodiment of the present invention, to meet application requirements, the model needs to be fine-tuned and further optimized to improve inference speed. Based on step 4, the model is pruned. If the detection accuracy after fine-tuning is lower than expected, the model returns to step 3, modifies specific parameter settings, and retrains until satisfactory results are achieved. Ultimately, an optimal model with an accuracy of 98.5% is obtained.
[0095] In this embodiment of the present invention, keypoint detection begins by acquiring an image to be predicted and performing appropriate processing. Furthermore, the image to be predicted is processed using the normalization applied to the validation set input images during model training. This process converts the image to the format expected by the model. The trained model is then used to predict the image to be detected, and further post-processing is performed to output hand keypoint information.
[0096] Based on the characteristic information of hand key points, the present invention innovatively proposes a model structure based on the improved YOLO11, deeply fuses deep features and shallow features during feature extraction, and sets a special detection head in combination with the attention mechanism; in order to accurately locate the hand key point information, the key point loss weight is adjusted, and different loss weights are set for different key points according to the importance and detection difficulty of the hand key points, prompting the model to pay more attention to those key points that are easy to be ignored, thereby improving the overall effect of the model; utilizing the currently more successful Wenshengtu technology, various natural movements of the hands in different environments are simulated to increase the diversity of training data, and a specific posture transformation enhancement method is designed according to the characteristics of the hands, so that the model can learn the characteristics of the hand key points under different postures. With the support of these two unique data enhancement methods, the model finally trained can more effectively handle the hand key point detection problem in complex scenes such as illumination changes, posture changes, and occlusion; after training, the model is further fine-tuned and optimized, and the model is compressed by pruning and quantization operations to improve the model's reasoning speed, which reduces the difficulty of the next step of deployment of the method.
[0097] The above is a detailed introduction to an improved YOLO11 hand key point detection method provided by an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A hand key point detection method based on improved YOLO11, characterized in that: include: Collect hand images under different hand postures and lighting conditions, annotate each hand image, and obtain key point information corresponding to each hand image; YOLO11 is used as the basic network architecture, and a feature fusion module based on the spatial attention mechanism is introduced to perform convolution operations on feature maps of different dimensions in the spatial dimension, thereby obtaining an improved YOLO11 network model; The annotated hand image is input into the improved YOLO11 network model for training to obtain a trained target YOLO11 network model, which is used to detect the hand image to be detected and obtain the position information of the hand key points.
2. The method according to claim 1, characterized in that The collecting of hand images under different hand postures and lighting conditions, labeling of each hand image, and obtaining key point information corresponding to each hand image include: Collect hand gesture images with specific requirements in various scenarios and generate hand images using text-based graph technology to simulate various natural hand movements in different environments; Each hand image is preprocessed, and then annotated using an annotation tool to obtain the coordinates of the hand key points corresponding to each hand image.
3. The method according to claim 1 or 2, characterized in that The feature fusion module based on the spatial attention mechanism is used to replace the connection module in the YOLO11 network.
4. The method according to claim 3, characterized in that The feature fusion module based on the spatial attention mechanism is used to obtain a shallow spatial attention feature map and a deep spatial attention feature map by performing convolution operations on feature maps of different dimensions in the spatial dimension, calculate the spatial attention weights of the shallow spatial attention feature map and the deep spatial attention feature map, multiply the shallow spatial attention feature map and the deep spatial attention feature map by their respective spatial attention weights, and then perform addition or splicing operations to achieve feature fusion.
5. The method according to claim 4, characterized in that During the training process of the improved YOLO11 network model, by changing the degree of finger bending, the model can learn the characteristics of the key points of the hand under different postures, and by adjusting the key point loss weights, different loss weights are set for different key points.
6. The method according to claim 5, characterized in that Depend on and Update the joint point position, where the bending degree parameter is k and the finger base joint point is A(x A ,y A ), the fingertip joint point is C(x C ,y C ), the middle knuckle point is B(x B ,y B ).
7. The method according to claim 6, characterized in that The method of adjusting the loss weight of key points and setting different loss weights for different key points includes: Assume that there are N key points in the hand key point detection task, where the index set of the fingertip key points is I, and the predicted key point coordinates are The real key point coordinates are y=(y1,y2,...,y N ); Define the weight vector w=(w1,w2,...,w N ), for the fingertip key point i∈I, set w i = l, for other key points Set w i =m, l>m; Depend on Get the weighted loss function.
8. The method according to claim 7, characterized in that A decoupled head design is used for each prediction branch in YOLO11 to handle classification and regression tasks separately.
9. The method according to claim 8, characterized in that The method further comprises: After the model is trained, it is compressed and accelerated through pruning.
10. The method according to claim 9, characterized in that The compression and acceleration of the model through pruning includes: The improved YOLO11 model is trained and its performance is evaluated on the validation set to obtain the benchmark performance index P0; Try to cut off each layer in turn, and retrain and evaluate the performance of the improved YOLO11 model on the validation set P i , P i The performance of the model on the validation set after pruning the i-th layer; If |P i If -P0| is less than a set tolerance ΔP, the i-th layer will be cut off.
Citation Information
Cited By
Improved YOLOv11-pose-based automobile crane gesture key point detection system and method
CN121330768A