Human body acupuncture point accurate recognition method based on DWPose algorithm
Through the two-stage knowledge distillation method and self-attention fusion module of the DWPose algorithm, the problems of low accuracy and efficiency in human acupoint recognition are solved, and accurate recognition and lightweight models are achieved, which are suitable for clinical and research use in traditional Chinese medicine.
Patent Information
- Application Number
- CN202510719614.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
AI Technical Summary
The existing technology of human acupoint identification has problems of low accuracy and low efficiency, especially in complex parts and large-scale acupoint research, which affects the treatment effect and increases medical costs.
A two-stage knowledge distillation method based on the DWPose algorithm was adopted. By improving the RTMPose-l network model and combining the self-attention and variance attention fusion modules, a DWPose network model was built to perform feature extraction and recognition of human acupoint images.
It achieves precise identification of human acupuncture points, improves recognition accuracy and real-time performance, reduces the number of model parameters, and improves detection efficiency. It is suitable for deployment on resource-constrained devices.
Smart Images

Figure CN120635203A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of health care acupoint identification, and in particular to a method for accurately identifying human acupoints based on a DWPose algorithm. Background Art
[0002] Accurate identification of acupuncture points on the human body is crucial in Traditional Chinese Medicine (TCM) clinical practice and research. Accurately locating acupuncture points plays a crucial role in the effectiveness of massage and acupuncture treatments and the accuracy of TCM diagnoses. Traditional acupuncture point identification methods rely primarily on manual location by professional physicians, relying on experience and body landmarks. However, this approach has numerous limitations.
[0003] On the one hand, doctors' experience levels vary, and different doctors may have certain differences in their judgment of acupoint locations. This leads to a high degree of subjectivity in acupoint positioning and a lack of unified standards and quantitative indicators. For example, in some complex areas of the body, such as the ear and wrist, the positioning of acupoints by different doctors may have slight deviations, which in turn affects the treatment effect. On the other hand, manual positioning is inefficient, especially when facing a large number of patients or when large-scale acupoint research is required. Doctors need to spend a lot of time and energy to make judgments one by one, which not only increases medical costs but also limits the promotion and development of Traditional Chinese Medicine. Therefore, developing a method that can accurately identify acupoints on the human body is of great practical significance. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for accurately identifying human acupoints based on the DWPose algorithm. This method can solve the problem of low accuracy in human acupoint identification, improve the accuracy and real-time performance of acupoint identification, and provide more reliable technical support for clinical practice and research in Traditional Chinese Medicine.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is: a method for accurately identifying human acupoints based on the DWPose algorithm, comprising the following steps:
[0006] Step 1: Obtain a dataset of human acupoint images;
[0007] Step 2: Preprocess the collected human acupoint image dataset and divide it into training set, test set and validation set;
[0008] Step 3: Build RTMPose-s and RTMPose-l network models, improve the RTMPose-l model, and obtain an improved RTMPose-l network model;
[0009] Step 4: Use the training set to train the RTMPose-s network model and the improved RTMPose-l network model respectively, and test the RTMPose-s network model using the test set;
[0010] Step 5: Build the DWPose network model and train the RTMPose-s network model through two-stage distillation, and test it;
[0011] Step 6: Use the RTMPose-s network model trained by distillation to perform acupoint recognition.
[0012] A further improvement of the technical solution of the present invention is that: step 1 specifically uses a depth camera to collect human acupoint images to obtain a human acupoint image dataset, and the collected human acupoint images include acupoint images under different lighting conditions and covering different body shapes.
[0013] The further improvement of the technical solution of the present invention is that the specific steps of step 2 are as follows:
[0014] Step 2.1: Use filtering algorithm to remove image noise;
[0015] Step 2.2: Use the image annotation tool Labelme to annotate the human acupoint image dataset after removing noise;
[0016] Step 2.3: Divide the human acupoint image dataset into training set, test set and validation set in a ratio of 8:1:1, and convert them into coco format.
[0017] A further improvement of the technical solution of the present invention is that in step 3, CSPNeXt-s is used with a depth factor of 0.33 and a width factor of 0.5 as the backbone network of RTMPose-s to build it, and CSPNeXt-l is used with a depth factor of 1.0 and a width factor of 1.0 as the backbone network of RTMPose-l to build it.
[0018] A further improvement of the technical solution of the present invention is that the RTMPose-1 network model is improved in step 3 by replacing the second ConvModule module in the first layer StemLayer in the original network structure with a "ConvModule->SVAFM->ConvModule(1×1)" structure, and adding an SVAFM module after the CSPLayer module in the second layer StageLayer1 and the 7×7Conv module in the Head part.
[0019] A further improvement of the technical solution of the present invention is that the SVAFM module is a self-attention and variance attention fusion module, which adopts a dual-branch parallel structure: the left branch captures global context dependency through self-attention and adjusts the channel through 1×1 convolution; the right branch first performs variance pooling on the input features to extract local feature dynamics, and then generates local attention weights through 1×1 convolution and Hardsigmoid activation; finally, the left and right branch features are fused by element-by-element addition to combine the global context information with the local dynamic salient features.
[0020] A further improvement of the technical solution of the present invention is that the improved RTMPose-1 network model in step 3 is used as the teacher model and RTMPose-s is used as the student model.
[0021] A further improvement of the technical solution of the present invention is that: in step 4, the learning rate is set to 0.001, the momentum is set to 0.9, the number of training rounds is 100, and the image size is set to 256×256.
[0022] A further improvement of the technical solution of the present invention is that: in step 5, the pre-trained weight file of the teacher model and the configuration files of the improved RTMPose-l teacher model and the RTMPose-s student model are used to train them through the two-stage distillation of DWPose. The improved RTMPose-l teacher model guides the student model RTMPose-s to learn from scratch through knowledge distillation. The student model inherits the teacher model's ability to understand acupoints through feature distillation, and imitates the teacher model's output distribution of acupoint coordinates through logical distillation. Then, it enters the second stage of self-training to complete the model distillation process based on DWPose, and finally realizes the lightweight model.
[0023] By adopting the above technical solution, the present invention achieves the following technological advancements: through a two-stage knowledge distillation process, a lightweight model is achieved, providing a feasible solution for future model deployment. By adding the Self-Attention and Variance Attention Fusion Module (SVAFM) to the Backbone network structure and the Head structure, the model's ability to focus on global information and local feature changes when processing complex features and noise is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0025] Figure 1 This is a flow chart of a method for accurately identifying human acupoints based on the DWPose algorithm provided in the present invention;
[0026] Figure 2 It is a principle block diagram of the DWPose algorithm provided in the present invention;
[0027] Figure 3 It is a principle block diagram of the RTMPose network model provided in the present invention;
[0028] Figure 4 It is the SVAFM module proposed in the present invention;
[0029] Figure 5 This is a principle block diagram of the improved RTMPose network model in the present invention. DETAILED DESCRIPTION
[0030] The present invention is described in further detail below in conjunction with the embodiments:
[0031] like Figure 1 FIG. 1 is a flow chart of a method for accurately identifying human acupoints based on the DWPose algorithm, including the following steps:
[0032] Step 1: Obtain a dataset of human acupoint images; use a depth camera to capture human acupoint images. The captured human acupoint images include acupoint images under different lighting conditions and covering different body shapes. In this embodiment, the acupoint images of the human back are taken as an example.
[0033] Step 2: Preprocess the collected human acupoint image dataset and divide it into training set, test set and validation set;
[0034] Step 2.1: Use filtering algorithm to remove image noise;
[0035] Step 2.2: Use the image annotation tool Labelme to annotate the human acupoint image dataset after removing noise;
[0036] Step 2.3: Divide the human acupoint image dataset into training set, test set and validation set in a ratio of 8:1:1, and convert them into coco format.
[0037] Step 3: Build RTMPose-s and RTMPose-l network models, improve the RTMPose-l model, and obtain an improved RTMPose-l network model;
[0038] like Figure 3As shown in the figure, the RTMPose network architecture model consists of a backbone, a 7×7 convolution layer, a fully connected FC layer, a global attention unit (GAU), and X- and Y-axis coordinate classifiers. The CSPNeXt-s (depth factor 0.33, width factor 0.5) network is used as the backbone network for RTMPose-s, while the CSPNeXt-l (depth factor 1.0, width factor 1.0) network is used as the backbone network for RTMPose-l. The backbone extracts features from the input human acupoint image, obtaining basic features that include human structure and potential acupoint locations. The 7×7 convolution layer further extracts spatial features from these basic features to enhance the feature representation of local regions. The fully connected FC layer integrates feature dimensions, and the global attention unit (GAU) captures long-range dependencies between different parts of the human body, focusing on acupoint-related areas. The X- and Y-axis coordinate classifiers are responsible for regressing and classifying the acquired acupoint points and outputting their coordinate positions in the image.
[0039] like Figure 4 Figure 2 shows the structure of the proposed Self-Variance Attention Fusion Module (SVAFM). The SVAFM module employs a two-branch parallel architecture: the left branch captures global contextual dependencies through self-attention and adjusts the channels via 1×1 convolution. The right branch first performs variance pooling on the input features to extract local feature dynamics, then generates local attention weights via 1×1 convolution and Hardsigmoid activation. Finally, the left and right branch features are fused through element-by-element addition, combining global contextual information with local dynamic salient features. This achieves efficient feature enhancement and improves the model's ability to express complex scenes.
[0040] like Figure 5As shown in the figure, the improved RTMPose-1 network model is mainly improved by replacing the second ConvModule module in the first StemLayer layer of the original network structure with the "ConvModule->SVAFM->ConvModule(1×1)" structure. At the same time, the SVAFM module is added after the CSPLayer module in the second StageLayer1 and the 7×7Conv module in the Head part. Adding the SVAFM module to the first StemLayer layer can effectively improve the model's ability to understand local details and global features. The CSPLayer module can enhance feature diversity through cross-stage connections. After adding the SVAFM module, the global scale and local scale can be further integrated. Adding the SVAFM module after the 7×7Conv module in the Head part can better learn the relative position constraints of multiple acupuncture points, improve the anti-interference ability during coordinate regression, and reduce the coordinate regression error.
[0041] The improved RTMPose-l teacher model has more backbone network layers and wider channels, and can extract high-precision, multi-scale feature information from the input human acupoint images; RTMPose-s, as a student model, has a similar structural framework to the teacher model, but its backbone network has fewer layers, which reduces the amount of calculation by compressing the number of channels. It has a lower computational cost and can achieve fast and accurate recognition of human acupoints on resource-constrained devices.
[0042] Step 4: Use the training set to train the RTMPose-s network model and the improved RTMPose-l network model, and test the RTMPose-s network model using the test set. When training the RTMPose network model using the training set, use AdamW as the network optimizer, set the learning rate to 0.001, the momentum to 0.9, the number of training rounds to 100, the batch size to 8, and the image size to 256×256. Use the training set to train the improved RTMPose-l and RTMPose-s network models to obtain the trained weight files. Use the test dataset to test the RTMPose-s network model.
[0043] Step 5: Build the DWPose network model and train the RTMPose-s network model through two-stage distillation, and test it;
[0044] like Figure 2The figure shows a two-stage distillation training using the pre-trained weight file of the teacher model, as well as the configuration files for the teacher model (the improved RTMPose-l model) and the student model (the RTMPose-s model). In the first stage, the improved RTMPose-l teacher model guides the student model (RTMPose-s) from scratch through knowledge distillation. Because acupuncture points are often small, such as the Fengmen acupoint, which is approximately 2 mm in diameter, and rely on surrounding anatomical landmarks such as bones and joints, the teacher model analyzes and classifies the overall structure of human acupoint images. It first extracts shallow features to identify more obvious structures such as the spine and joints, and then extracts deep features to identify the subtle features of acupoints, resulting in more accurate acupoint identification. The student model (RTMPose-s) aligns the spatial dimensions of the backbone network features of the improved RTMPose-l teacher model through 1×1 convolutional layers. Feature distillation inherits the teacher model's understanding of acupoints, and logical distillation allows the student model to mimic the teacher model's output distribution of acupoint coordinates. In the second stage of self-training, the backbone structure of RTMPose-s is fixed, and the head network of RTMPose-s is continuously adjusted according to the teacher model, ultimately improving the detection efficiency and accuracy of human acupoints.
[0045] Step 6: Use the RTMPose-s network model trained by distillation to identify acupoints, that is, the RTMPose-s network model after the two-stage distillation of DWPose is used to identify the acupoints on the back of the human body.
[0046] The method of the present invention is verified using the back acupuncture point embodiment, as follows:
[0047] Table 1 Performance comparison of DWPose model before and after distillation
[0048]
[0049] Through the two-stage distillation based on DWPose, it can be seen that the parameter size of the RTMPose-s model is maintained at 5.5M while the model accuracy is increased from the original 84.6% to 99%, which greatly reduces the parameter size of the teacher model. At the same time, the FPS is increased from 26 to 99, which improves the detection speed of the cloud detection model and better meets the real-time requirements. In summary, the model is lightweight, the detection efficiency of the model is improved, and a feasible solution is provided for model deployment.
[0050] Table 2 lists the prediction error values of acupoints before and after distillation of the student model.
[0051] Table 2 Human acupoint error table
[0052]
[0053] Table 2 continued
[0054]
[0055] Note: BL11 is Dazhui acupoint, BL12 is Fengmen acupoint, BL13 is Feishu acupoint, BL14 is Jueyinshu acupoint, BL15 is Xinshu acupoint, BL16 is Dushu acupoint, BL17 is Geshu acupoint, BL18 is Ganshu acupoint, BL19 is Gallbladdershu acupoint, BL20 is Pishu acupoint, BL21 is Weishu acupoint, BL22 is Sanjiaoshu acupoint, BL23 is Shenshu acupoint, BL43 is Gaomang acupoint, BL44 is Shentang acupoint, BL46 is Geguan acupoint, BL49 is Yishe acupoint, BL50 is Weicang acupoint, DU14 is Dazhui acupoint, DU12 is Shenzhu acupoint, DU9 is Zhiyang acupoint, DU6 is Jizhong acupoint, and DU4 is Mingmen acupoint.
[0056] The units of the data in Table 2 are all cm. The first row of the table is the names of the acupuncture points on the back of the human body, the second row is the prediction error values of the student model for the acupuncture points before distillation, and the third row is the prediction error values of the student model for the acupuncture points after distillation.
[0057] Analysis of the experimental data above shows that the student model's error for back acupuncture points before distillation was a maximum of 1.107 cm and a minimum of 0.156 cm. However, after the lightweight student model underwent two-stage distillation using the DWPose algorithm, its error for back acupuncture point prediction was significantly improved, with a maximum error of 0.140 cm and a minimum error of 0.039 cm. Therefore, the DWPose algorithm employed in this invention enables accurate acupuncture point identification.
[0058] In summary, the present invention provides a method for accurately identifying human acupoints based on DWPose, which can accurately detect human acupoints, realize the lightweight of the model, and improve the detection efficiency.
[0059] The embodiments described above are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should fall within the scope of protection determined by the claims of the present invention.
Claims
1. A method for accurately identifying human acupoints based on the DWPose algorithm, characterized by: The steps are as follows: Step 1: Obtain a dataset of human acupoint images; Step 2: Preprocess the collected human acupoint image dataset and divide it into training set, test set and validation set; Step 3: Build RTMPose-s and RTMPose-l network models, improve the RTMPose-l model, and obtain an improved RTMPose-l network model; Step 4: Use the training set to train the RTMPose-s network model and the improved RTMPose-l network model respectively, and test the RTMPose-s network model using the test set; Step 5: Build the DWPose network model and train the RTMPose-s network model through two-stage distillation, and test it; Step 6: Use the RTMPose-s network model trained by distillation to perform acupoint recognition.
2. The method for accurately identifying human acupoints based on the DWPose algorithm according to claim 1, characterized in that: Step 1 specifically involves using a depth camera to capture human acupoint images to obtain a human acupoint image dataset. The captured human acupoint images include acupoint images under different lighting conditions and covering different body shapes.
3. The method for accurately identifying human acupoints based on the DWPose algorithm according to claim 1, characterized in that: Step 2: Step 2.1: Use filtering algorithm to remove image noise; Step 2.2: Use the image annotation tool Labelme to annotate the human acupoint image dataset after removing noise; Step 2.3: Divide the human acupoint image dataset into training set, test set and validation set in a ratio of 8:1:1, and convert them into coco format.
4. The method for accurately identifying human acupoints based on the DWPose algorithm according to claim 1, characterized in that: In step 3, CSPNeXt-s is used with a depth factor of 0.33 and a width factor of 0.5 as the backbone network of RTMPose-s to build it, and CSPNeXt-l is used with a depth factor of 1.0 and a width factor of 1.0 as the backbone network of RTMPose-l to build it.
5. The method for accurately identifying human acupoints based on the DWPose algorithm according to claim 4, characterized in that: The improvement method for the RTMPose-1 network model in step 3 is to replace the second ConvModule module in the first layer StemLayer in the original network structure with the "ConvModule->SVAFM->ConvModule(1×1)" structure, and add the SVAFM module after the CSPLayer module in the second layer StageLayer1 and the 7×7Conv module in the Head part.
6. The method for accurately identifying human acupoints based on the DWPose algorithm according to claim 4, characterized in that: The SVAFM module is a self-attention and variance attention fusion module with a dual-branch parallel structure: the left branch captures global contextual dependencies through self-attention and adjusts the channels through 1×1 convolution; the right branch first performs variance pooling on the input features to extract local feature dynamics, and then generates local attention weights through 1×1 convolution and Hardsigmoid activation; finally, the left and right branch features are fused through element-by-element addition to combine global contextual information with local dynamic salient features.
7. The method for accurately identifying human acupoints based on the DWPose algorithm according to claim 6, characterized in that: The improved RTMPose-l network model in step 3 is used as the teacher model, and RTMPose-s is used as the student model.
8. The method for accurately identifying human acupoints based on the DWPose algorithm according to claim 1, characterized in that: In step 4, the learning rate is set to 0.001, the momentum is set to 0.9, the number of training rounds is 100, and the image size is set to 256×256.
9. The method for accurately identifying human acupoints based on the DWPose algorithm according to claim 7, characterized in that: In step 5, the pre-trained weight file of the teacher model and the configuration files of the improved RTMPose-l teacher model and RTMPose-s student model are used to train them through the two-stage distillation of DWPose. The improved RTMPose-l teacher model guides the student model RTMPose-s to learn from scratch through knowledge distillation. The student model inherits the teacher model's ability to understand acupoints through feature distillation, and imitates the teacher model's output distribution of acupoint coordinates through logical distillation. Then, the second stage of self-training is entered to complete the model distillation process based on DWPose, and finally the model is lightweight.
Citation Information
Cited By
A hand acupoint recognition method based on multi-scale feature fusion and spatial constraint regression
CN122715210A