Gesture recognition method, system and device based on fine-grained clustering and medium

Through fine-grained clustering and knowledge distillation technology, the accuracy and resource limitation problems of gesture recognition methods in complex environments are solved, efficient gesture recognition on resource-constrained devices is achieved, and disassembly efficiency and security are improved.

CN120340137APending Publication Date: 2025-07-18QINGDAO UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510606911.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing gesture recognition methods are insufficient in complex and changeable disassembly environments, and the high computing and storage costs limit real-time deployment on resource-constrained devices, making it difficult to meet the needs of disassembly of used household appliances.

Method used

The gesture recognition method based on fine-grained clustering is adopted, and the knowledge of complex models is transferred to the lightweight model through fine-grained clustering feature extraction and knowledge distillation technology, combining dynamic weight regulation and layered feature distillation mechanism to improve the recognition accuracy and efficiency.

Benefits of technology

Implement efficient gesture recognition on resource-constrained devices, improve disassembly efficiency and security, reduce labor intensity, and adapt to real-time deployment of embedded systems and mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340137A_ABST
    Figure CN120340137A_ABST
Patent Text Reader

Abstract

The invention discloses a gesture recognition method, system and device based on fine-grained clustering and a medium, and the method comprises the steps: inputting an obtained gesture image into a teacher model, extracting a high-dimensional feature embedding vector of a specific layer, decomposing the high-dimensional feature embedding vector into a plurality of fine-grained sub-blocks, carrying out the clustering of the sub-blocks through a distance matrix between the sub-blocks, and carrying out the recognition of the fine-grained sub-blocks. According to the similarity between each clustering set and the original gesture image, distributing a weight for each clustering set, and carrying out dot multiplication on the weights and the label matrix to generate a weight matrix for subsequent knowledge distillation; the student model performs weighted distillation learning by utilizing the teacher features extracted by the teacher model and the weight matrix through a dynamic weight regulation and control mechanism; through a hierarchical feature distillation mechanism, automatically extracting fine-grained features of gestures from deep features and shallow features of the teacher model and the student model through clustering, and accurately describing the features of the gestures; knowledge of a complex model is migrated to a lightweight model, and efficient operation is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing and computer vision technologies, and particularly relates to a gesture recognition method, system, device, and medium based on fine-grained clustering. Background Art

[0002] With the rapid development of the economy and the continuous improvement of people's living standards, the replacement speed of household appliances has increased significantly, resulting in a sharp increase in the number of waste household appliances. However, despite the large number of waste household appliances, the current situation of their recycling and dismantling is not optimistic. During the dismantling process of waste household appliances, workers need to perform complex gestures and body movements, and the feature extraction of these movements is crucial for optimizing the dismantling process and reducing risks. For example, by extracting the gesture movement features and body movement features during the worker's operation, real-time monitoring and guidance of the worker's operation can be achieved to ensure the dismantling quality; in addition, the feature extraction of risk information during the household appliance dismantling process is also the key to ensuring the safety of workers. By real-time monitoring the gesture and body movement conditions and analyzing features such as the stability, speed, and acceleration of body movements, potential dangerous movements can be detected in a timely manner to ensure the safety of workers.

[0003] Traditional gesture recognition methods usually rely on manually designed feature extraction and simple machine learning algorithms, such as those based on skin color segmentation, contour detection, or optical flow methods. However, these methods often have certain limitations when facing complex and changing dismantling environments. For example, changes in lighting may cause skin color segmentation to fail, background interference will affect the accuracy of contour detection, and the diversity and complexity of gestures make it difficult for manually designed features to comprehensively describe the features of gestures.

[0004] With the development of deep learning technologies, gesture recognition methods based on neural networks have made significant progress. For example, the Mediapipe library, based on deep learning algorithms, can capture hand key points in real time through a camera and achieve gesture recognition. In addition, technologies such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have also been widely applied to gesture recognition tasks. These deep learning-based methods can more accurately recognize gesture images or video sequences and respond to the worker's operation instructions in real time. Although deep learning-based gesture recognition methods perform well in terms of accuracy and robustness, these methods usually rely on large-scale and high-complexity models, and their high computational and storage costs limit their real-time deployment on mobile devices and edge devices. For example, complex deep learning models require a large amount of computing resources to process real-time video streams, which is difficult to achieve in resource-constrained dismantling scenarios. Summary of the Invention

[0005] The technical problem to be solved by this application is to overcome the deficiencies of the prior art and provide a gesture recognition method, system, device and medium based on fine-grained clustering. By clustering, the fine-grained features of gestures can be automatically extracted, which can more accurately describe the features of gestures; using the knowledge distillation technology, the knowledge of complex models is transferred to lightweight models to ensure efficient operation on resource-constrained devices while maintaining high recognition accuracy; it can not only improve the accuracy of disassembly efficiency, but also reduce labor intensity and safety risks, providing an innovative solution for the intelligent upgrade of the waste household appliance disassembly industry.

[0006] To achieve the above object, the first aspect of this application provides a gesture recognition method based on fine-grained clustering, including the following steps:

[0007] Step S1, obtaining a gesture image and performing preprocessing;

[0008] Step S2, extracting fine-grained clustering features;

[0009] Input the gesture image obtained in step S1 into the teacher model, and extract the teacher features of the teacher model through the method of fine-grained clustering feature extraction, including specific deep features and shallow features. Among them, the method of fine-grained clustering feature extraction is: extracting the high-dimensional feature embedding vector of a specific layer, decomposing the high-dimensional feature embedding vector into multiple fine-grained sub-blocks, clustering the sub-blocks through the distance matrix between the sub-blocks, and according to the similarity between each clustering set and the original gesture image, assigning weights to each clustering set, and generating a weight matrix for subsequent knowledge distillation by multiplying the weights and the label matrix;

[0010] Step S3, knowledge distillation and training;

[0011] The student model divides the extracted student features into sub-blocks with the same number and dimension as the teacher features. The student model uses the teacher features and the weight matrix extracted by the teacher model through the dynamic weight regulation mechanism for weighted distillation learning; through the hierarchical feature distillation mechanism, the deep features and shallow features of the extracted teacher model and student model are used to sum the distillation losses of the deep features and shallow features by weighting, and then the total hierarchical distillation loss is obtained to train the student model.

[0012] Optionally, the fine-grained clustering feature extraction in step S2 includes:

[0013] Extracting the image embedding vector E: Extracting the image embedding vector from the input gesture image through a pre-trained convolutional neural network model, expressed as:

[0014] E = f(I);

[0015] Among them, f represents the feature extraction function of the convolutional neural network model, I represents the input image, and E represents the image embedding vector;

[0016] Image embedding vector division: Divide the image embedding vector E into c image patches, and the image embedding vector of each image patch is e n , which is expressed as:

[0017] e = {e1, e2,..., e c};

[0018] where e represents the set of c image patches, and e c represents the c-th image patch;

[0019] Calculate the distance matrix: For each image patch vector e n , calculate the Euclidean distance between the current image patch vector e i and other image patch vectors e j , which is expressed as:

[0020]

[0021] where k represents the dimension of the image patch e n , e i represents the i-th image patch, e j represents the j-th image patch, e i,k represents the image patch e i with dimension k, e j,k represents the image patch e j with dimension k. The differences are calculated separately in k dimensions. dist(*) represents the Euclidean distance. The rows of the obtained distance matrix are sorted in ascending order of distance to form the distance matrix C n ;

[0022] Clustering: Merge the distance matrix C n , and use the union-find algorithm to group the semantically similar image patches e n together. Specifically, if the distance dist(e i , e j ) between two image patches e i , e j exceeds the threshold β, then the two image patches are assigned to different sets; if the distance dist(e i , e j ) between two image patches e i , e j is less than or equal to the threshold β, then the two image patches are merged into the same set; finally, a total of m clustering sets V m containing similar semantic image patches are obtained;

[0023] Evaluate the similarity between each clustering set and the original image: Calculate the image embedding vector E and the m clustering sets V mThe cosine similarity of the remaining similarity, and the set with a higher similarity can better represent the features of the original image. For the clustering set V m For the i-th set V i in, calculate the cosine similarity between E and all the image patches e i in V i , and take the average value of the |V i | image patches, which is expressed as:

[0024]

[0025] where E represents the image embedding vector, and e i represents the image patch in the set V i , |V i | represents the number of image patches in the set V i , and Sim i represents the average cosine similarity between the image embedding vector E and the i-th clustering set V i .

[0026] Optionally, assigning weights to each clustering set in step S2 includes:

[0027] Assign a weight to each clustering set according to the similarity between each clustering set and the original image, which is expressed as:

[0028]

[0029] where weight j is the weight of the j-th clustering set, representing the weight of all |V j | sub-blocks in the j-th clustering set V j in distillation; ∈ is set to a small constant to avoid a zero denominator; θ controls the dynamic scaling of the weight.

[0030] Optionally, in step S3, weighted distillation learning is performed through a dynamic weight regulation mechanism using the teacher features extracted by the teacher model and the weight matrix, specifically including:

[0031] The teacher model extracts features at a specific layer to obtain teacher features:

[0032] E t,layer = f t,layer (I);

[0033] where E t,layer represents the teacher features obtained at the specific layer, I is the input image, and f t,layer (·) is the feature extraction function based on the fine-grained clustering feature extraction module;

[0034] The student model extracts features in the corresponding layer and divides them into sub-blocks with the same number and dimension as the teacher features:

[0035] E s,layer = g s,layer (I) = div s=t (t s,layer (I));

[0036] Among them, g s,layer (·) is the feature extraction function of the student model, which is used to extract the features of the corresponding layer of the student model and divide the student features into image patches with the same number and dimension as the teacher features; t s,layer (·) extracts the features of the corresponding layer of the student model, and div s=t (·) is a function used to divide the student features into image patches with the same number as the teacher features; E s,layer represents the feature vector extracted by the student model at the corresponding layer;

[0037] Perform a dot product operation with the weight weight j and the label matrix Y to generate the weight matrix M layer , and use the weight matrix M layer generated by the teacher model to perform weighted distillation on the features of the student model:

[0038]

[0039] Among them, L layer is a loss function used to measure the difference between the student features and the teacher features at a specific layer, and M layer represents the weight matrix generated by using the teacher model. The rows of the weight matrix M layer correspond to the image sub-patches, and the columns correspond to the weights of the image sub-patches; E t,layer represents the teacher features obtained at a specific layer, represents the square of the Euclidean norm.

[0040] Optionally, the deep and shallow features of the teacher model and the student model are expressed as:

[0041] E t,deep = f t,deep (I)E t,shallow = f t,shallow (I);

[0042] E s,deep = g s,deep (I)E s,shallow = g s,shallow (I);

[0043] Among them, f t,deep (I) and f t,shallow (I) are the deep and shallow feature extraction functions of the teacher model; f t,·(·) represents the process of extracting image embeddings, dividing into sub - blocks, clustering, and weight assignment; g s,deep (I) and g s,shallow (I) are the deep and shallow feature extraction functions of the student model; I is the input image, E t,deep and E t,shallow represents that the teacher model uses the feature extraction function f t,deep (I) and f t,shallow (I) to extract the feature vectors obtained respectively in the deep layer and the shallow layer, E s,deep and E s,shallow represents that the student model uses the feature extraction function g s,deep (I) and g s,shallow (I) to extract the feature vectors obtained respectively in the deep layer and the shallow layer.

[0044] Optionally, the hierarchical feature distillation mechanism in step S3 includes:

[0045] Calculating the hierarchical distillation loss: Calculate the corresponding deep - layer feature distillation loss and shallow - layer feature distillation loss between the student model and the teacher model respectively, including:

[0046] Calculating the deep - layer feature distillation loss, expressed as:

[0047]

[0048] where L deep represents the deep - layer feature distillation loss, E s,deep represents the feature vector extracted by the student in the deep layer, E t,deep represents the feature vector extracted by the teacher in the deep layer, represents the square of the Euclidean norm. Calculating the shallow - layer feature distillation loss, expressed as:

[0049]

[0050] where L shallow represents the shallow - layer feature distillation loss, E s,shallow represents the feature vector extracted by the student in the shallow layer, E t,shallow represents the feature vector extracted by the teacher in the shallow layer, represents the square of the Euclidean norm.

[0051] Calculating the total hierarchical distillation loss, which is the weighted sum of the deep - layer feature distillation loss and the shallow - layer feature distillation loss, expressed as:

[0052] L d =α·L deep +β·L shallow ;

[0053] where L dDenote the total hierarchical distillation loss, and α and β denote hyperparameters used to balance the importance of deep and shallow features.

[0054] Optionally, the total distillation loss includes the standard logits distillation loss and the total hierarchical distillation loss, which is expressed as:

[0055] L total =γ·L logits +(1 - γ)·L d ;

[0056] where L total denotes the total distillation loss, L logits is the standard logits distillation loss, L d is the total hierarchical distillation loss, and γ is a hyperparameter used to balance the two parts of the loss.

[0057] To achieve the above object, the second aspect of the present application provides a gesture recognition system based on fine-grained clustering, and the gesture recognition system includes:

[0058] An acquisition module that acquires a gesture image and performs preprocessing;

[0059] A fine-grained clustering feature extraction module, which is used to input the acquired gesture image into a teacher model, and extract teacher features of the teacher model through a fine-grained clustering feature extraction method, including specific deep features and shallow features. Among them, the fine-grained clustering feature extraction method is: extracting high-dimensional feature embedding vectors of a specific layer, decomposing the high-dimensional feature embedding vectors into multiple fine-grained sub-blocks, clustering the sub-blocks through the distance matrix between the sub-blocks, and according to the similarity between each clustering set and the original gesture image, assigning weights to each clustering set, and generating a weight matrix for subsequent knowledge distillation through the dot product of the weights and the label matrix;

[0060] A knowledge distillation module; including a teacher-student model, the student model divides the extracted student features into sub-blocks with the same number and dimension as the teacher features, and the student model uses the teacher features and the weight matrix extracted by the teacher model through a dynamic weight regulation mechanism for weighted distillation learning; through a hierarchical feature distillation mechanism, the deep features and shallow features of the extracted teacher model and student model are weighted and summed for the distillation loss of the deep features and shallow features, and then the total hierarchical distillation loss is obtained to train the student model.

[0061] To achieve the above object, the third aspect of the present application provides a gesture recognition device based on fine-grained clustering, including a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the gesture recognition method based on fine-grained clustering as described above is implemented.

[0062] To achieve the above object, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the gesture recognition method based on fine-grained clustering as described above.

[0063] After adopting the above technical solutions, the present application has the following beneficial effects compared with the prior art:

[0064] In the present application, the retention of detailed information is enhanced through fine-grained clustering feature extraction. In existing gesture recognition methods, as the network depth increases, the feature maps gradually fuse more context information, resulting in the loss of detailed information (such as finger bending, palm contour, etc.). In the present application, the feature maps are decomposed into multiple fine-grained sub-blocks, and a clustering algorithm is used to group similar features. This process not only effectively retains various detailed information of the gesture, but also significantly improves the recognition accuracy of the model for complex gestures, enabling the model to capture the subtle features of the gesture more comprehensively and accurately.

[0065] In the present application, by introducing a dynamic weight regulation mechanism, the teacher model and the student model can perform more accurate knowledge transfer at the feature level, significantly optimizing the knowledge transfer process between the teacher model and the student model. By enhancing the knowledge transfer between the teacher-student models through dynamic weight regulation, the most valuable knowledge can be intelligently identified and transferred, and valuable knowledge can be selectively transferred to the student model. This accurate knowledge transfer method not only significantly improves the performance of the student model in gesture recognition, but also enables it to capture and understand the detailed features of the gesture more accurately, thus showing higher robustness and accuracy in complex scenarios.

[0066] In the present application, a hierarchical feature distillation mechanism is set for the differences in semantic information of deep and shallow features. Different weights are set for features with different deep and shallow semantic information. The student model can simultaneously learn the global semantic information and local detailed information of the teacher model. By measuring and selecting important semantic information through the hierarchical feature distillation mechanism, this method not only improves the learning efficiency of the student model, but also enhances its expression ability for different levels of features, enabling it to more accurately identify and understand gestures in complex tasks, thus significantly improving the overall performance.

[0067] Efficient Deployment of Deep Learning Models on Resource-Constrained Devices: By combining knowledge distillation and model compression strategies, a resource-efficient lightweight gesture recognition network is designed. The supervision signals in the distillation process are concentrated on key features, reducing the transmission of redundant signals, thereby effectively compressing the number of parameters of the student model. In addition, the lightweight model achieves multi-branch parallel processing during the inference process, significantly improving the speed of gesture recognition and enabling it to adapt to resource-constrained devices such as embedded systems and mobile devices. This innovation not only improves the practicality of the model but also expands its application scope in actual scenarios.

[0068] The following further describes the specific embodiments of the present application in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The accompanying drawings, as part of the present application, are used to provide a further understanding of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application but do not constitute an improper limitation of the present application. Obviously, the accompanying drawings in the following description are only some embodiments, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0070] In the accompanying drawings of the specification:

[0071] Figure 1 is a schematic diagram of the overall steps of the gesture recognition method based on fine-grained clustering in this specific embodiment;

[0072] Figure 2 is a framework diagram of the gesture recognition method based on fine-grained clustering in this specific embodiment;

[0073] Figure 3 is a flowchart of the gesture recognition method based on fine-grained clustering in this specific embodiment. SPECIFIC EMBODIMENTS

[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not used to limit the scope of the present application.

[0075] Regarding the gesture recognition technology mentioned in the background art, the following elaboration is provided:

[0076] The application of gesture recognition technology in the field of industrial automation is one of the key directions to promote the intelligent upgrading of modern industry. Its core task is to improve production efficiency and operation safety through efficient human-machine interaction. Applying gesture recognition in the process of waste household appliance disassembly is of great significance for improving disassembly efficiency, reducing labor intensity and ensuring worker safety. Specifically, gesture recognition technology aims to achieve real-time control of industrial equipment, operation guidance, and monitoring and early warning of abnormal behaviors by accurately recognizing the gesture actions of workers or operators. This task usually involves the following processes: starting from data collection, through preprocessing, feature extraction, classification and recognition, post-processing, and finally outputting the recognition results and applying them to the actual scenario. Among them, feature extraction and classification recognition models are the two most important parts of the gesture recognition task.

[0077] Among them, feature extraction is the process of extracting key information that can effectively describe gestures from the original data (such as images or videos). The original data (such as high-resolution images or video frames) usually contains a large amount of redundant information. Directly using it for classification will lead to waste of computing resources and degradation of model performance. Feature extraction can convert high-dimensional data into low-dimensional feature vectors, remove irrelevant information, and retain key information, thereby improving the efficiency and performance of the system.

[0078] Among them, the classification model is the core of the gesture recognition system. Its task is to classify gestures into predefined categories according to the extracted features. The performance of the classification model directly affects the accuracy and generalization ability of gesture recognition. An excellent classification model can maintain high accuracy under different scenarios, different users, and different lighting conditions. In practical applications, the gesture recognition system needs to respond to users' gesture commands in real time. An efficient classification model can quickly process the input features and output results to meet the requirements of real-time interaction.

[0079] However, traditional gesture recognition methods face many challenges in industrial environments. For example, complex backgrounds and lighting conditions may lead to a decrease in the accuracy of gesture recognition; the diversity and deformation of gestures make it difficult to comprehensively capture features with single-modal data; in addition, the high dependence of existing methods on computing resources makes it difficult to adapt to resource-constrained mobile devices or embedded systems, restricting the wide application of gesture recognition technology in actual industrial scenarios.

[0080] Based on this, the present application provides a gesture recognition method based on fine-grained clustering for intelligent interaction and operation monitoring in industrial automation scenarios. By using the fine-grained clustering method to extract features, the robustness of feature expression is enhanced. At the same time, knowledge distillation technology is used to transfer the knowledge of the complex teacher model to the lightweight student model, significantly reducing the computational overhead. Combining with the dynamic distillation strategy, through the hierarchical learning of features and the task weighting mechanism, the performance of gesture recognition is dynamically optimized. Further, lightweight design is adopted to ensure the efficient deployment of the model in embedded devices. In addition, by improving the accuracy of gesture recognition and the model deployment efficiency, reliable technical support is provided for intelligent interaction and operation monitoring in industrial scenarios.

[0081] Please refer to Figure 1 , Figure 2 and Figure 3 , the gesture recognition method based on fine-grained clustering provided by the present application includes the following steps:

[0082] Step S1, obtain a gesture image and perform preprocessing;

[0083] Step S2, extract fine-grained clustering features;

[0084] Input the gesture image obtained in step S1 into the teacher model, and extract the teacher features of the teacher model through the method of fine-grained clustering feature extraction, including specific deep features and shallow features. Among them, the method of fine-grained clustering feature extraction is: extract the high-dimensional feature embedding vector of a specific layer, decompose the high-dimensional feature embedding vector into multiple fine-grained sub-blocks, perform clustering of the sub-blocks through the distance matrix between the sub-blocks, and assign weights to each clustering set according to the similarity between each clustering set and the original gesture image. Generate a weight matrix for subsequent knowledge distillation by performing a dot product of the weights and the label matrix;

[0085] Step S3, perform knowledge distillation and training;

[0086] The student model divides the extracted student features into sub-blocks with the same number and dimension as the teacher features. The student model uses the teacher features and the weight matrix extracted by the teacher model through the dynamic weight regulation mechanism for weighted distillation learning; through the hierarchical feature distillation mechanism, the deep features and shallow features of the extracted teacher model and student model are weighted and summed to obtain the total hierarchical distillation loss for training the student model.

[0087] It should be noted that the execution subject of the gesture recognition method in this embodiment is a gesture recognition device based on fine-grained clustering knowledge distillation, and this device can be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, etc., and the non-mobile electronic device can be a server and a personal computer, etc., which are not specifically limited in this application. Hereinafter, taking the server as the execution subject as an example, the gesture recognition method based on fine-grained clustering knowledge distillation in this embodiment will be described.

[0088] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, "a plurality of" means two or more, unless otherwise specifically defined.

[0089] In an implementable embodiment, obtaining the input image and preprocessing includes the steps of data acquisition, data preprocessing, and data loading.

[0090] Data acquisition: In the intelligent industrial equipment monitoring task, the acquisition of image data is the basis of the entire process. High-quality image data can significantly improve the accuracy of subsequent processing and analysis, which requires using a high-resolution RGB camera as the image acquisition device to ensure that the details of the worker's gestures can be clearly captured; at the same time, according to the actual scene requirements, appropriate camera parameters also need to be selected, such as frame rate (30fps or higher), resolution (at least 1920×1080), etc.

[0091] In addition, to ensure the diversity of the data set, the image data acquisition should cover multiple scenarios. For example, images are acquired under different lighting environments, including bright environments (such as natural light, strong lights), dim environments (such as weak lights, shadow areas), and mixed lighting conditions, which helps the model learn the gesture features under various lighting conditions; simulate the actual working scenario and acquire the gesture images of workers in different device placement positions. For example, the devices may be partially blocked, overlapped, or completely exposed, which can enhance the model's adaptability to complex scenarios; acquire the operation gesture images of workers near different device types (such as robotic arms, control consoles, tool cabinets, etc.). These devices may affect light reflection and shadows, so these scenarios need to be covered to improve the robustness of the model.

[0092] Data Preprocessing: The purpose of data preprocessing is to convert the collected raw image data into a format suitable for model training, while removing noise, enhancing data diversity, and improving the training efficiency and stability of the model. For all collected RGB images, first adjust them to a unified size, such as 224×224 pixels or 256×256 pixels, which can ensure that the images input into the model have consistent dimensions. Secondly, normalize the pixel values of the images to the range of [0,1]. Normalization can accelerate the convergence speed of the model and improve the stability of training. For RGB images, normalization can be performed through the following formula:

[0093]

[0094] where I min and I max are the minimum and maximum values of the image pixel values respectively. In addition, through data augmentation techniques (such as random rotation, cropping, flipping, brightness adjustment, etc.), the data diversity is increased, enabling the model to be exposed to more diverse data during the training phase, thereby improving the model's adaptability to actual scenarios.

[0095] Data Loading: The preprocessed image data needs to be stored in the training data pipeline so that the model can load and use it efficiently. The shape of the image data loaded into the computer should be B×W×H×C, where B is the batch size, W and H are the width and height of the image respectively, and C is the number of channels of the image. For RGB images, C = 3.

[0096] It should be noted that in the scenario of waste household appliance disassembly, in order to achieve accurate recognition and risk assessment of workers' gestures, first, image data of workers' gestures are collected through a high-resolution RGB camera to ensure coverage of a variety of actual working scenarios, including different lighting conditions, equipment placement positions, and equipment types. These images are preprocessed, adjusted to a unified size and normalized pixel values, and at the same time, the generalization ability of the model is improved through data augmentation techniques (such as rotation, cropping, brightness adjustment). The preprocessed image data is loaded into the training data pipeline for efficient use by the model.

[0097] In an implementable embodiment, the teacher model includes a teacher network, and the teacher network includes N modules, representing each layer of the teacher network; the student model includes a student network, and the student network includes N modules, representing each layer of the student network. The modules between the teacher network and the student network correspond to each other one by one.

[0098] In an implementable embodiment, fine-grained clustering features are extracted.

[0099] Specifically, the core objective of fine-grained clustering feature extraction is to extract feature vectors from gesture images that can accurately describe gesture details. Through deep learning and clustering algorithms, redundant information can be effectively reduced, and the robustness and representativeness of features can be enhanced, thus significantly improving the accuracy of gesture recognition.

[0100] The following are the detailed steps:

[0101] Extract image embeddings: Select a pre-trained convolutional neural network (CNN) model as the teacher model for extracting image features. The pre-trained model has been trained on large-scale datasets (such as ImageNet) and can extract features with strong semantic information, which is suitable for gesture recognition tasks. Commonly used models include ResNet, VGG, or MobileNet, etc.

[0102] Input the input image I into the selected teacher model to extract the feature embedding vectors E of specific layers (deep and shallow layers). t = f t (I), where f t (·) represents the feature extraction function of the teacher model, and E t is a compact representation of the image that can capture the key information in the image. The dimension of E t becomes B×D. D is determined by the teacher model and represents the dimension of the feature. B represents the batchsize, which is determined by the specific experimental details. If ResNet-50 is used as the feature extractor, the size of the input image I is B×224×224×3. After feature extraction by ResNet-50, the dimension of the obtained feature vector E t is B×2048.

[0103] Image embedding vector partitioning: Image embedding vector partitioning is a basic step in the fine-grained clustering feature extraction module. Its purpose is to decompose the high-dimensional feature embedding vector into multiple fine-grained image patches so that subsequent clustering operations can more accurately capture the local details of the gesture. At the same time, each image patch can be organized and distributed to different nodes for parallel computing to improve computing efficiency. Divide the image embedding vector E t into c image patches e t,i ∈{e t,1 , e t,2 ,..., e t,c}, where e t,i represents the i-th image patch after the division of the image embedding vector E t , and its dimension is:

[0104]

[0105] where B represents the batchsize, D represents the dimension of the feature, and c represents the number of image patches;

[0106] Clustering: In the fine-grained clustering feature extraction module, the clustering of image patches is performed through the distance matrix between image patches. For each image patch e t,i , calculate the Euclidean distance between this image patch and other patches to measure the similarity between image patches. For any image patch e t,i and e t,j , the Euclidean distance between the two is:

[0107]

[0108] where e t,i represents the i-th image patch, e t,j represents the j-th image patch, and dist(*) represents the Euclidean distance. Calculate the distances between all image patches to form a c×c global distance matrix C n , and the global distance matrix contains the similarity information between all image patches.

[0109] Use the Union-Find algorithm to quickly merge similar elements. Its core idea is to manage the grouping of elements through a dynamic set structure. Each element is initially an independent set, and similar elements are combined together through merge operations. First, traverse the global distance matrix C n , check the distance dist(e t,i ,e t,j ) between each pair of image patches (e t,i ,e t,j ). If the distance dist(e t,i ,e t,j ) is less than the set threshold β, then merge the sets where these two image patches are located. When merging, use the path compression technique to find the root nodes of the two image patches (i.e., the representative elements of the sets they belong to). If the root nodes are different, point the root node of one set to the root node of the other set, thereby merging the two sets, and use the root node of each set as the set number. Finally, obtain m clustering sets V i , and it is known the numbers of each set and the image patches in the sets. Assign class labels y j ∈[y1,y2,...,y c to each clustering set. Traverse each set to generate a label matrix Y with a size of c×m. The rows of the matrix correspond to each image patch, and the columns correspond to the various classes of the clustering sets.

[0110] Generating weights: Evaluate the similarity between each clustering set and the original image, and select the feature set that is closer to the semantic information of the original image from multiple clustering sets, thereby enhancing the robustness and semantic consistency of the features. After clustering, each clustering set Vi Contains a set of similar image patches. To evaluate the similarity between these sets and the original image, it is necessary to calculate the feature embedding vector E of the original image t The average cosine similarity with all image patches in each clustering set:

[0111]

[0112] where E represents the image embedding vector, and e i represents the image patch in set V i |V i | represents the number of image patches in set V i , and Sim i represents the average cosine similarity between the image embedding vector E and the i-th clustering set V i . According to the similarity between each clustering set and the original image, a weight is assigned to each set. The higher the similarity (i.e., the smaller the distance), the greater the weight. The calculation formula for the weight is as follows:

[0113]

[0114] where weight j is the weight of the j-th clustering set, representing the weights of all |V j | image patches in the j-th clustering set V j | during distillation; ∈ is set to a small constant to avoid a zero denominator; θ controls the dynamic scaling of the weight. The dot product operation is performed on the weight weight j and the label matrix Y to generate the weight matrix M, where the rows correspond to the image patches and the columns correspond to the weights of the image patches.

[0115] It should be noted that in the feature extraction stage, a pre-trained convolutional neural network (such as ResNet-50) is used to extract the high-dimensional feature embedding vectors of the images, and these feature vectors are subdivided into multiple fine-grained image patches. By calculating the Euclidean distance between the image patches and using the union-find algorithm for clustering, the similar image patches are grouped to form multiple clustering sets. The similarity between each clustering set and the original image is evaluated, and based on this, weights are assigned to each set to ensure that the higher the similarity, the greater the weight. These weights are combined with the label matrix to generate the weight matrix for subsequent distillation.

[0116] In an achievable implementation, weighted distillation learning is performed using the teacher features extracted by the teacher model and the weight matrix through a dynamic weight regulation mechanism, specifically including:

[0117] Through the dynamic weight regulation mechanism, the features of the teacher model are finely grained clustered and weighted, and the features of the student model learn these weighted features through weighted distillation, so as to more efficiently absorb the knowledge of the teacher model. The teacher model extracts features in a specific layer and obtains the teacher features and the weight matrix for distillation:

[0118] E t,layer =f t,layer (I);

[0119] Among them, E t,layer represents the teacher features obtained in a specific layer, I is the input image, and f t,layer (·) is the feature extraction function based on the fine-grained clustering feature extraction module, and the finally generated weight matrix for distillation is M layer .

[0120] The student model extracts features in the corresponding layer and divides them into sub-blocks with the same number and dimension as the teacher features:

[0121] E s,layer =g s,layer (I)=div s=t (t s,layer (I));

[0122] Among them, g s,layer (·) is the feature extraction function of the student model, which completes the function of extracting the features of a certain layer of the student model and dividing the student features into sub-blocks with the same number and dimension as the teacher features; t s,layer (·) extracts the features of a certain layer of the student model, and div s=t (·) divides the student features into sub-blocks with the same number as the teacher features, and E s,layer represents the feature vector extracted by the student model in the corresponding layer.

[0123] Use the weight matrix M layer generated by the teacher model to perform weighted distillation on the features of the student model:

[0124]

[0125] Among them, L layer is a loss function used to measure the difference between the student features and the teacher features in a specific layer, M layer represents the weight matrix generated by the teacher model, and the rows of the weight matrix M layer correspond to the image sub-blocks, and the columns correspond to the weights of the image sub-blocks; E t,layer represents the teacher features obtained in a specific layer, and E s,layer represents the student features obtained in a specific layer, represents the square of the Euclidean norm.

[0126] In an implementable embodiment, the hierarchical feature distillation in step S3 is described below: The goal of hierarchical feature distillation is to simultaneously learn the deep and shallow features of the teacher model, and set different weights according to the difference in their semantic information, so that the student model can more comprehensively absorb the knowledge of the teacher model and enhance its expression ability for features at different levels.

[0127] In an implementable embodiment, first, specific deep and shallow features of the teacher model and the student model are extracted respectively to provide a basis for the subsequent distillation process. The deep and shallow features of the teacher model and the student model are expressed as:

[0128] E t,deep = f t,deep (I)E t,shallow = f t,shallow (I);

[0129] E s,deep = g s,deep (I)E s,shallow = g s,shallow (I);

[0130] where f t,deep (I) and f t,shallow (I) are the deep and shallow feature extraction functions of the teacher model; f t,· (·) represents the processes of extracting image embeddings, dividing sub-blocks, clustering, and weight assignment; g s,deep (I) and g s,shallow (I) are the deep and shallow feature extraction functions of the student model; I is the input image, E t,deep and E t,shallow represent the feature vectors extracted by the teacher model using the feature extraction functions f t,deep (I) and f t,shallow (I) at the deep and shallow layers respectively, and E s,deep and E s,shallow represent the feature vectors extracted by the student model using the feature extraction functions g s,deep (I) and g s,shallow (I) at the deep and shallow layers respectively.

[0131] Deep features are usually located in the latter part of the network and have stronger semantic information. Shallow features are usually located in the former part of the network and contain more texture and edge information. In this embodiment, taking ResNet50 as the teacher network, the first layer is selected as the shallow layer and the third layer is selected as the deep layer.

[0132] In an implementable embodiment, the hierarchical feature distillation mechanism in step S3 includes: calculating the distillation losses of the corresponding deep and shallow features between the student model and the teacher model respectively;

[0133] Calculate the deep feature distillation loss, denoted as:

[0134]

[0135] where L deep denotes the deep feature distillation loss, M deep denotes the deep weight matrix, E s,deep denotes the feature vector extracted by the student at the deep layer, E t,deep denotes the feature vector extracted by the teacher at the deep layer, denotes the square of the Euclidean norm;

[0136] Calculate the shallow feature distillation loss, denoted as:

[0137]

[0138] where L shallow denotes the shallow feature distillation loss, M shallow denotes the shallow weight matrix, E s,shallow denotes the feature vector extracted by the student at the shallow layer, E t,shallow denotes the feature vector extracted by the teacher at the shallow layer, denotes the square of the Euclidean norm;

[0139] Weightedly sum the distillation losses of the deep and shallow features to obtain the total hierarchical distillation loss:

[0140] L d = α · L deep + β · L shallow ;

[0141] where α and β are hyperparameters used to balance the importance of deep and shallow features. L d denotes the total hierarchical distillation loss. Deep features usually contain higher-level semantic information, so a higher weight α is required; shallow features contain more texture and edge information, which helps the student model learn details, so an appropriate weight β is required;

[0142] Calculate the total distillation loss: Combine the total hierarchical distillation loss with the standard logits distillation loss to obtain the total distillation loss, thereby optimizing the training process of the student model.

[0143] Calculate the standard logits distillation loss between the student model and the teacher model:

[0144]

[0145] where K is the number of classes, F s,k and F t,kThey are the logits outputs of the student model and the teacher model respectively. Logits represent the predicted values output by the model, and softmax is a function that converts logits into a probability distribution;

[0146] Combine the hierarchical distillation loss and the standard logits distillation loss to obtain the total distillation loss:

[0147] L total = γ · L logits +(1 - γ) · L d

[0148] Where L total represents the total distillation loss, and γ is a hyperparameter used to control the relative importance between the logits distillation loss and the hierarchical distillation loss.

[0149] It should be noted that during the knowledge distillation stage, the features of the teacher model are passed to the student model after fine-grained clustering and weight assignment. The features of the student model are divided into sub-blocks with the same number and dimension as those of the teacher model, and the knowledge of the teacher model is learned through weighted distillation. In addition, a hierarchical distillation mechanism is introduced to extract the deep and shallow features of the teacher model and the student model respectively, and different weights are set according to the semantic information differences. Deep features usually contain stronger semantic information, while shallow features contain more texture and edge information. By weighted summing the distillation losses of the deep and shallow features, the total hierarchical distillation loss is obtained. Finally, the hierarchical distillation loss is combined with the standard logits distillation loss to form the total distillation loss, thereby optimizing the training process of the student model.

[0150] In an implementable embodiment, in terms of hyperparameter settings, the initial learning rate is set to 0.001, and a decay strategy of reducing the learning rate by 10% every certain number of rounds is adopted; to prevent overfitting, L2 regularization is used, and the decay value is set to 0.0001; during the clustering process, the distance threshold β is set to 0.07 to determine whether an image sub-block belongs to the same set; in addition, during the distillation process, the hyperparameters α and β are set to 1 and 1.5 respectively to balance the hierarchical distillation loss, and the hyperparameter γ is set to 0.5 to balance the relative importance between the logits distillation loss and the hierarchical distillation loss.

[0151] Based on the same inventive concept, the present application also provides a gesture recognition system based on fine-grained clustering. The gesture recognition system includes:

[0152] An acquisition module that acquires a gesture image and performs preprocessing;

[0153] The fine-grained clustering feature extraction module is used to input the obtained gesture image into the teacher model, extract the high-dimensional feature embedding vectors of a specific layer, decompose the high-dimensional feature embedding vectors into multiple fine-grained sub-blocks, perform clustering of the sub-blocks through the distance matrix between the sub-blocks, assign weights to each clustering set according to the similarity between each clustering set and the original gesture image, and generate a weight matrix for subsequent knowledge distillation by performing a dot product of the weights and the label matrix;

[0154] The knowledge distillation module is used to perform knowledge distillation. The student model divides the extracted student features into sub-blocks with the same number and dimension as the teacher features. The student model uses the teacher features and the weight matrix extracted by the teacher model through a dynamic weight regulation mechanism for weighted distillation learning; the deep features and shallow features of the teacher model and the student model are respectively extracted through a hierarchical feature distillation mechanism, and the distillation losses of the deep features and shallow features are weighted and summed to obtain the total hierarchical distillation loss for training the student model.

[0155] In an implementable embodiment, the trained student model is deployed on the waste household appliance disassembly production line. During the waste household appliance disassembly process, the obtained gesture image of the worker is input into the trained student model for gesture recognition, so as to realize real-time monitoring and guidance of the worker's operation and ensure the disassembly quality; analyze features such as the stability, speed, and acceleration of the limb movements, and timely discover potential dangerous actions to ensure the safety of the worker.

[0156] Specifically, the obtained gesture image of the worker is input into the teacher model, and the teacher features of the teacher model, including specific deep features and shallow features, are extracted through the method of fine-grained clustering feature extraction. Among them, the method of fine-grained clustering feature extraction is: extract the high-dimensional feature embedding vectors of a specific layer, decompose the high-dimensional feature embedding vectors into multiple fine-grained sub-blocks, perform clustering of the sub-blocks through the distance matrix between the sub-blocks, assign weights to each clustering set according to the similarity between each clustering set and the original gesture image, and generate a weight matrix for subsequent knowledge distillation by performing a dot product of the weights and the label matrix; the student model divides the extracted student features into sub-blocks with the same number and dimension as the teacher features. The student model uses the teacher features and the weight matrix extracted by the teacher model through a dynamic weight regulation mechanism for weighted distillation learning; the deep features and shallow features of the teacher model and the student model extracted are weighted and summed through the hierarchical feature distillation mechanism to obtain the total hierarchical distillation loss for training the student model, and gesture recognition is performed on the input gesture image through the trained student model.

[0157] Based on the same inventive concept, the present application also provides a gesture recognition device based on fine-grained clustering, including a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, it implements the gesture recognition method based on fine-grained clustering knowledge distillation as described above.

[0158] Based on the same inventive concept, the present application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the gesture recognition method based on fine-grained clustering as described above.

[0159] The program product for implementing the above method in the present application may be a portable compact disc read-only memory and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In the present application, the readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0160] It should be noted that the computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0161] The above are only the preferred embodiments of the present application, and do not impose any formal limitations on the present application. Although the present application has been disclosed above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art of the present application can make some changes or modifications to the above-disclosed technical content within the scope of the technical solution of the present application to obtain equivalent embodiments of equivalent changes. The implementation schemes in the above embodiments can also be further combined or replaced. However, as long as the content does not depart from the technical solution of the present application, any simple modification, equivalent change, and modification made to the above embodiments according to the technical essence of the present application still fall within the scope of the present application solution.

Claims

1. A gesture recognition method based on fine-grained clustering, characterized in that It includes the following steps: Step S1: Obtain a gesture image and perform preprocessing; Step S2: Extract fine-grained clustering features; Input the gesture image obtained in Step S1 into the teacher model, and extract the teacher features of the teacher model through the method of fine-grained clustering feature extraction, including specific deep features and shallow features. Among them, the method of fine-grained clustering feature extraction is: extract the high-dimensional feature embedding vector of a specific layer, decompose the high-dimensional feature embedding vector into multiple fine-grained sub-blocks, perform clustering of the sub-blocks through the distance matrix between the sub-blocks, and assign weights to each clustering set according to the similarity between each clustering set and the original gesture image. Generate a weight matrix for subsequent knowledge distillation by multiplying the weights and the label matrix; Step S3: Perform knowledge distillation and training; The student model divides the extracted student features into sub-blocks with the same number and dimension as the teacher features. The student model uses the teacher features and the weight matrix extracted by the teacher model through the dynamic weight regulation mechanism for weighted distillation learning; through the hierarchical feature distillation mechanism, the deep features and shallow features of the extracted teacher model and student model are used to calculate the distillation loss of the weighted sum of the deep features and shallow features, and then the total hierarchical distillation loss is obtained to train the student model.

2. The method according to claim 1, characterized in that, The fine-grained clustering feature extraction in Step S2 includes: Extract the image embedding vector: Extract the image embedding vector from the input gesture image through a pre-trained convolutional neural network model, expressed as: E = f(I); where f represents the feature extraction function of the convolutional neural network model, I represents the input image, and E represents the image embedding vector; Image embedding vector division: Divide the image embedding vector E into c image patches, and the vector of each image patch is e n , which is expressed as: e = {e1, e2,..., e c}; Among them, e represents a set of c image patches, and e c represents the c-th image patch; Calculate the distance matrix: For each image patch vector e n , calculate the Euclidean distance between the current image patch vector e i and other image patch vectors e j , which is expressed as: where k represents the dimension of the image patch e n ; e i represents the i-th image patch, e j represents the j-th image patch, e i,k represents the image patch e with dimension k i ; e j,k represents the image patch e with dimension k j ; the two are calculated separately in k dimensions, dist(*) represents the Euclidean distance, and the rows of the obtained distance matrix are sorted in ascending order of distance to form the distance matrix C n ; Clustering: Merge distance matrix C n , use the union-find algorithm to group semantically similar image blocks n Grouped together, specifically including: if two image blocks e i 、e j The distance between i ,e j ) exceeds the threshold β, the two image blocks are assigned to different sets; if the two image blocks e i 、e j The distance between i ,e j ) is less than or equal to the threshold β, the two image blocks are merged into the same set; finally, a total of m clustering sets V containing image blocks with similar semantics are obtained m ; Evaluate the similarity between each cluster set and the original image: Calculate the cosine similarity between the image embedding vector E and the m cluster sets V m respectively. The set with a higher similarity can better represent the features of the original image. For the i-th set V m in the cluster set V i , calculate the cosine similarity between E and all the image patches e i in V i , and take the average of |V i | image patches, which is expressed as: Among them, E represents the image embedding vector, e i represents the image patch in the set V i and |V i | represents the number of image patches in the set V i , and Sim i represents the average cosine similarity between the image embedding vector E and the i-th clustering set V i .

3. The method according to claim 1, wherein Assigning weights to each clustering set in Step S2 includes: Assign a weight to each clustering set according to the similarity between each clustering set and the original image, expressed as: Among them, weight j is the weight of the j-th clustering set, representing all |V j | image patches in the j-th clustering set V j in the distillation; ∈ is set to a constant to avoid a zero denominator; θ controls the dynamic scaling of the weight.

4. The method according to claim 1, characterized in that, In Step S3, the weighted distillation learning is performed using the teacher features and the weight matrix extracted by the teacher model through the dynamic weight regulation mechanism, specifically including: The teacher model extracts features at a specific layer to obtain teacher features, expressed as: E t,layer = f t,layer (I); Among them, E t,layer represents the teacher feature vector obtained from a specific layer, I is the input image, and f t,layer (·) is a feature extraction function based on the fine-grained clustering feature extraction module; The student model extracts features in the corresponding layer and divides them into image blocks with the same number and dimension as the teacher features, expressed as: E s,layer = g s,layer (I) = div s=t (t s,layer (I)); Among them, g s,layer (·) is the feature extraction function of the student model, which is used to extract the features of the corresponding layer of the student model and divide the student features into image patches with the same number and dimension as the teacher features; t s,layer (·) extracts the features of the corresponding layer of the student model, and div s=t (·) is a function used to divide the student features into the same number of image patches as the teacher features; E s,layer represents the feature vector extracted by the student model at the corresponding layer; Dot multiply with the weight j and the label matrix Y to generate the weight matrix M layer , and use the weight matrix M generated by the teacher model layer to perform weighted distillation on the features of the student model, expressed as: Among them, L layer is a loss function used to measure the difference between the features of the student model and the features of the teacher model, and M layer represents the weight matrix generated using the teacher model. The rows of the weight matrix M layer correspond to the image patches, and the columns correspond to the weights of the image patches; E t,layer represents the teacher features obtained from a specific layer, represents the square of the Euclidean norm, which is used to calculate the sum of the squares of the vectors.

5. The method according to claim 1, characterized in that The deep and shallow features of the teacher model and the student model, expressed as: E t,deep = f t,deep (I)E t,shallow = f t,shallow (I); E s,deep = g s,deep (I) E s,shallow = g s,shallow (I); Among them, f t,deep (I) and f t,shallow (I) are the deep and shallow feature extraction functions of the teacher model; f t,· (·) represents the processes of extracting image embeddings, dividing into sub-blocks, clustering, and weight assignment; g s,deep (I) and g s,shallow (I) are the deep and shallow feature extraction functions of the student model; I is the input image, E t,deep and E t,shallow represent the feature vectors extracted by the teacher model using the feature extraction functions f t,deep (I) and f t,shallow (I) at the deep and shallow layers respectively, and E s,deep and E s,shallow represent the feature vectors extracted by the student model using the feature extraction functions g s,deep (I) and g s,shallow (I) at the deep and shallow layers respectively.

6. The method according to claim 5, characterized in that, The hierarchical feature distillation mechanism in Step S3 includes: Calculate the hierarchical distillation loss: Calculate the corresponding deep feature distillation loss and shallow feature distillation loss between the student model and the teacher model respectively, including: Calculate the deep feature distillation loss, expressed as: Among them, L deep represents the deep feature distillation loss, E s,deep represents the feature vector extracted by the student at the deep layer, E t,deep represents the feature vector extracted by the teacher at the deep layer, represents the square of the Euclidean norm; Calculate the shallow feature distillation loss, expressed as: Among them, L shallow represents the shallow feature distillation loss, and E s,shallow represents the feature vector extracted by the student at the shallow layer, and E t,shallow represents the feature vector extracted by the teacher at the shallow layer; Calculate the total hierarchical distillation loss, and perform a weighted sum of the deep feature distillation loss and the shallow feature distillation loss, expressed as: L d = α·L deep + β·L shallow ; Among them, L d represents the total hierarchical distillation loss, and α and β represent hyperparameters used to balance the importance of deep and shallow features.

7. The method according to claim 5, wherein The total distillation loss includes the standard logits distillation loss and the total hierarchical distillation loss, expressed as: L total = γ·L logits + (1 - γ)·L d ; Among them, L total represents the total distillation loss, L logits is the standard logits distillation loss, L d is the total hierarchical distillation loss, and γ is a hyperparameter used to balance the two parts of the loss.

8. A gesture recognition system based on fine-grained clustering, characterized in that, The gesture recognition system includes: An acquisition module that acquires a gesture image and performs preprocessing; A fine-grained clustering feature extraction module is used to input the obtained gesture image into the teacher model, and extract the teacher features of the teacher model through the method of fine-grained clustering feature extraction, including specific deep features and shallow features. Among them, the method of fine-grained clustering feature extraction is: extract the high-dimensional feature embedding vector of a specific layer, decompose the high-dimensional feature embedding vector into multiple fine-grained sub-blocks, perform clustering on the sub-blocks through the distance matrix between the sub-blocks, assign weights to each clustering set according to the similarity between each clustering set and the original gesture image, and generate a weight matrix for subsequent knowledge distillation by performing a dot product of the weights and the label matrix; A knowledge distillation module; includes a teacher-student model. The student model divides the extracted student features into sub-blocks with the same number and dimension as the teacher features. The student model uses the teacher features and the weight matrix extracted by the teacher model through a dynamic weight regulation mechanism for weighted distillation learning; through a hierarchical feature distillation mechanism, the deep features and shallow features of the extracted teacher model and student model are weighted and summed to obtain the total hierarchical distillation loss, which is used to train the student model.

9. A gesture recognition device based on fine-grained clustering, characterized in that, It includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, it implements the gesture recognition method based on fine-grained clustering as described in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the gesture recognition method based on fine-grained clustering as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Hierarchical identification method, device and system for electronic data trust intensity authentication and medium

    CN121071434A