Human body back acupoint automatic identification method and system based on key point detection

By introducing the Focus module and the C3STR module into the target detection network, the problem of poor back acupuncture point detection effect in the prior art is solved, and more efficient feature extraction and acupuncture point positioning are achieved.

CN120198677AActive Publication Date: 2025-06-24GUANGDONG EMBOSSED STORM ROBOT CO LTD

Patent Information

Application Number
CN202510270133.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The prior art has poor detection results when dealing with back acupoints with high density distribution and small target characteristics, and lacks public data sets specifically for back acupoints in humans, resulting in insufficient model adaptability.

Method used

The Focus module and C3STR module are introduced into the target detection network. The Focus module enhances the attention of local areas through slice recombination and feature extraction, and the C3STR module improves the efficiency of feature extraction through the hierarchical window self-attention mechanism.

Benefits of technology

The detection effect of the target detection model when detecting back acupuncture points with high density distribution and small target characteristics is improved, and the model's ability to capture key features and accurate positioning of acupuncture points is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198677A_ABST
    Figure CN120198677A_ABST
Patent Text Reader

Abstract

The invention discloses a human body back acupoint automatic identification method and system based on key point detection. And inputting a to-be-analyzed human body back image into the trained target detection network to obtain acupoint identification information. The trained target detection network comprises a Focus module and a plurality of C3STR modules. The Focus module is located on the first layer and carries out slice recombination and feature extraction on an input image, local areas can be processed in a refined mode, attention to the back and acupuncture point areas is enhanced, and a target detection network can better capture key features. The C3STR module performs feature extraction on the output image of the previous layer through a hierarchical window self-attention mechanism, so that the feature extraction efficiency is effectively improved, and the acupuncture point can be more accurately positioned. Therefore, after the Focus module and the C3STR module are introduced, a good detection effect can be achieved when the target detection network detects back acupuncture points with high-density distribution and small target characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image analysis, and particularly to an automatic recognition method and system for human back acupoints based on key point detection. Background Art

[0002] The positioning of human acupoints is very important in the fields of traditional Chinese medicine physiotherapy and the like. The back acupoints are closely related to the human meridians and are of great significance to therapies such as massage. With the rapid development of artificial intelligence technology, automatic acupoint recognition has become an important direction for improving the intelligent level of physiotherapy.

[0003] Deep learning has made remarkable progress in object detection and key point detection tasks. The YOLO series of models are widely used, but there are still technical problems such as difficulty in adapting to the characteristics of acupoint distribution and inferring unknown acupoints when directly applied to the recognition of human back acupoints, and it is necessary to specifically optimize the algorithm structure and data characteristics. Object detection technology can achieve the positioning of specific parts of the human body, but the detection effect of back acupoints with high-density distribution and small target characteristics is not good, especially when the background is complex or the human body posture changes greatly. Key point detection technology is mainly used for motion capture or pose estimation, but it is easily restricted by data distribution and annotation accuracy when dealing with acupoints with high-density distribution on the back, and does not combine the specific anatomical characteristics of acupoint distribution, making it difficult to model and infer the positional relationship between acupoints.

[0004] In addition, there is currently little research on the human back acupoint dataset, and there is a lack of a publicly available dataset specifically for human back acupoints. The lack of the dataset leads to insufficient adaptability of existing models in detecting back acupoints.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an automatic recognition method and system for human back acupoints based on key point detection in view of the above-mentioned defects of the existing technology, aiming to solve the problem of poor detection effect of the existing detection algorithms when dealing with back acupoints with high-density distribution and small target characteristics.

[0007] The technical solution adopted by the present invention to solve the problem is as follows:

[0008] In a first aspect, an embodiment of the present invention provides an automatic recognition method for human back acupoints based on key point detection, and the method includes:

[0009] Obtain a human back image to be analyzed;

[0010] Input the human back image into the trained object detection network; wherein, the trained object detection network includes a Focus module and several C3STR modules; the Focus module is located at the first layer and is used for slicing, reorganizing, and feature extraction of the input image; the C3STR module is used for feature extraction of the output image of the previous layer through a hierarchical window self-attention mechanism;

[0011] Obtain the acupoint recognition information corresponding to the human back image through the trained object detection network; wherein, the acupoint recognition information includes the back detection frame, the positions and categories of key acupoint points.

[0012] In a second aspect, an embodiment of the present invention further provides a human back acupoint automatic recognition system based on key point detection, and the system includes:

[0013] An image acquisition module, configured to acquire a human back image to be analyzed;

[0014] An image input module, configured to input the human back image into the trained object detection network; wherein, the trained object detection network includes a Focus module and several C3STR modules; the Focus module is located at the first layer and is used for slicing, reorganizing, and feature extraction of the input image; the C3STR module is used for feature extraction of the output image of the previous layer through a hierarchical window self-attention mechanism;

[0015] An image detection module, configured to obtain the acupoint recognition information corresponding to the human back image through the trained object detection network; wherein, the acupoint recognition information includes the back detection frame, the positions and categories of key acupoint points.

[0016] In a third aspect, an embodiment of the present invention further provides a terminal, and the terminal includes a memory and more than one processor; the memory stores more than one program; the program includes instructions for executing the human back acupoint automatic recognition method based on key point detection as described in any one of the above; the processor is configured to execute the program.

[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor to implement the steps of the human back acupoint automatic recognition method based on key point detection as described in any one of the above.

[0018] Advantages of the present invention: In the embodiments of the present invention, a Focus module and a C3STR module are introduced into the target detection network. The Focus module is used to refine the processing of local areas, enhancing the attention to the back and acupoint areas, enabling the target detection network to better capture key features. The C3STR module effectively improves the efficiency of feature extraction by introducing a spatial attention mechanism, and can more accurately locate acupoints. After introducing the Focus module and the C3STR module, the target detection model can achieve better detection results when detecting back acupoints with high-density distribution and small target characteristics. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of the automatic recognition method for human back acupoints based on key point detection provided by the embodiments of the present invention.

[0021] Figure 2 It is a flowchart of the model training process provided by the embodiments of the present invention.

[0022] Figure 3 It is a schematic diagram of the data augmentation combination provided by the embodiments of the present invention.

[0023] Figure 4 It is a schematic diagram of Mosaic augmentation provided by the embodiments of the present invention.

[0024] Figure 5 It is a structural diagram of the improved target detection network provided by the embodiments of the present invention.

[0025] Figure 6 It is a structural diagram of the Focus network provided by the embodiments of the present invention.

[0026] Figure 7 It is a structural diagram of the C3STR module network provided by the embodiments of the present invention.

[0027] Figure 8 It is a structural diagram of the Swin Transformer Layer network provided by the embodiments of the present invention.

[0028] Figure 9 It is a diagram showing the model prediction effect provided by the embodiments of the present invention.

[0029] Figure 10It is a schematic diagram of the modules of the human back acupoint automatic recognition system based on key point detection provided by an embodiment of the present invention.

[0030] Figure 11 It is a schematic block diagram of the terminal provided by an embodiment of the present invention. Specific embodiments

[0031] The present invention discloses a method and system for automatically recognizing human back acupoints based on key point detection. To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. The "including" used herein means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.

[0032] Aiming at the above-mentioned defects of the prior art, the present invention provides a method for automatically recognizing human back acupoints based on key point detection, as Figure 1 shown, the method specifically includes the following steps:

[0033] Step S100, obtaining a human back image to be analyzed.

[0034] Specifically, first, an image of the human back needs to be obtained. The image can be collected from directly above the human body by an RGBD depth camera or other image acquisition devices.

[0035] Step S200, inputting the human back image into a trained target detection network; wherein, the trained target detection network includes a Focus module and several C3STR modules; the Focus module is located at the first layer and is used for slicing and reorganizing the input image and feature extraction; the C3STR module is used for feature extraction of the output image of the previous layer through a hierarchical window self-attention mechanism.

[0036] Specifically, the human back image is input into the trained object detection network (or object detection model). To improve the model performance, the Focus and C3STR modules are further introduced, enhancing the attention to local regions and the ability to extract spatial features respectively. The Focus module is located at the first layer. By slicing the input image in the height and width dimensions and concatenating them in the channel dimension, it realizes the conversion from spatial information to channel information while retaining all pixel information, which helps to improve the efficiency and accuracy of subsequent feature extraction. The C3STR module combines a convolutional neural network and a self-attention mechanism. By using a hierarchical window self-attention mechanism to extract features from the output image of the previous layer, it can effectively capture key points and global information in the image, enhancing the model's detection ability for targets of different scales. In this embodiment, the Focus and C3STR modules are added to the object detection network. The Focus module achieves efficient downsampling while maintaining information integrity, and the C3STR module well integrates the advantages of traditional convolution and modern Transformer, enabling the object detection network to significantly reduce the number of model parameters while maintaining the model's computational efficiency, ensuring the balance between performance and computational resources.

[0037] For example, the basic architecture of the object detection network can choose YOLOv11. YOLOv11 is an efficient version of the YOLO series, which redefines the performance boundary of real-time object detection through architecture optimization and function extension. Based on the optimization of YOLOv9 and YOLOv10, an improved feature extraction module is introduced to enhance the ability to capture image details, especially performing better in complex scenarios. Adopting a lightweight design, it achieves higher computational efficiency by reducing the number of parameters while maintaining the accuracy. More post-processing optimization methods are added to improve the inference speed of the model. Through multi-dimensional optimization, YOLOv11 has set a new benchmark in terms of speed, accuracy, and functional diversity, becoming a general vision model widely adopted in the industrial and academic fields.

[0038] Step S300: Obtain the acupoint recognition information corresponding to the human back image through the trained object detection network; wherein, the acupoint recognition information includes the back detection frame, the positions and categories of key acupoints.

[0039] Specifically, under the action of the Focus and C3STR modules, the network gradually extracts the features of the image, abstracting richer feature representations from low-level to high-level. By predicting the extracted feature map, the acupoint recognition information can be obtained, including three types of information: the back detection frame, the positions of key acupoints, and the categories. The back detection frame is used to locate the back region, the positions of key acupoints are used to accurately locate the acupoints, and the categories are used to distinguish different acupoints.

[0040] In one implementation, the method for constructing the training dataset of the trained object detection network includes:

[0041] Obtain a number of original human back images through a depth camera;

[0042] For each of the original human back images, obtain the enhanced image corresponding to the original human back image through data augmentation operations, and generate the annotation data corresponding to the enhanced image;

[0043] Generate a dataset according to all the enhanced images and the annotation data, and divide the training dataset from the dataset.

[0044] Specifically, a depth camera can capture images and simultaneously obtain depth information, which can provide the depth information of each pixel point in the image and is very important for object detection tasks. In practical applications, a depth camera is used to collect a number of original human back images as the basic data for subsequent processing. The original human back images are raw data without any processing and may contain problems such as noise, illumination changes, background interference, etc. Data augmentation brings significant advantages to the object detection model, especially in enhancing the generalization ability of the model, improving the detection ability of small objects, enhancing the adaptability to complex backgrounds, enhancing the robustness to target position and scale changes, and reducing overfitting. By reasonably using different data augmentation methods, the performance of the object detection model can be effectively improved, especially when facing a changing real environment and limited training data. For each original human back image, one or more data augmentation operations are applied to generate one or more enhanced images. The annotation data includes, but is not limited to, the back detection frame, the positions and categories of key acupoints. In practical applications, if the annotation data is generated before the data augmentation operation, the annotation data needs to be synchronously transformed during the data augmentation operation to ensure the consistency between the enhanced image and the annotation data. The dataset is a set composed of enhanced images and corresponding annotation data. The dataset is divided into a training set, a validation set, and a test set to evaluate the performance of the model at different stages.

[0045] For example, such as Figure 2As shown, the data collection used a Deptrum Aurora 930 series RGBD depth camera, which was collected directly above the human body to ensure that the images included a full-back view and captured the details of the human back. A total of 1500 images were collected, covering the diversity of different body shapes, skin colors, postures, and lighting conditions. During the annotation process, in addition to marking the overall detection frame of the back, 27 key acupoint points were accurately marked to ensure a high degree of accuracy and consistency in the position of each acupoint. To improve the annotation accuracy, depth information was combined to ensure the correct identification of the three-dimensional spatial position of the back. The dataset was divided into a training set, a validation set, and a test set, which were allocated in a ratio of 8:1:1 to ensure the effectiveness of model training and the comprehensiveness of validation. In actual application, the model training adopted an architecture based on a deep convolutional neural network and used the SGD optimizer to adjust the parameters. For the acupoint recognition task, a multi-task learning strategy was designed to ensure that the model could not only detect the back area but also accurately locate the 27 acupoints. After the training was completed, through the evaluation of the validation set and the test set, the model showed excellent performance in indicators such as accuracy, recall rate, and mAP, achieving the design goal.

[0046] In one implementation, the data augmentation operations include at least one of random cropping, random rotation, horizontal flipping, brightness adjustment, image merging, and depth map enhancement.

[0047] In the actual scenario, the target is often intertwined with the complex background. To enhance the generalization ability of the model, help the model better adapt to this complex background, and reduce the interference of the background on the detection results. In this embodiment, a variety of data augmentation techniques are adopted during the training process, enabling the model to be exposed to diverse training data. As a result, when the model encounters unseen images, it can better perform inference and improve its generalization ability in different scenarios. As Figure 3 shown, in this embodiment, a variety of data augmentations are combined. The main data augmentation operations include random cropping, rotation, flipping, brightness adjustment, Mosaic technology, etc. The depth map enhancement technology is also combined, such as adding noise and depth distortion to the depth map, to enhance the generalization ability of the model and enable the model to perform stably in complex scenarios. It is also possible to simulate back images in different environments to further improve the robustness of the model to target spatial transformation and its performance in complex scenarios. In the case of limited training data, data augmentation effectively expands the training set by generating diverse samples and reduces the risk of model overfitting.

[0048] For example, for random cropping: Random cropping is to randomly crop a small area from the original image to simulate objects of different sizes and positions. This data augmentation operation can increase the robustness of the model to changes in the object's position. Suppose the size of the original image is W×H, and a sub-image of size w×h is cropped from it. The mathematical formula for cropping can be expressed as:

[0049] I crop (x,y) = I(x + Δx,y + Δy);

[0050] where I(x,y) is the value of the pixel (x,y) in the original image; Δx and Δy are the cropping offsets, which are randomly selected.

[0051] For random rotation: Random rotation is to rotate the image within a certain range of angles. Through rotation, the objects in the image can be trained at different angles, enhancing the rotational invariance of the model. Suppose the rotation angle is θ, and the pixel values of the rotated image I rot can be calculated through the following transformation:

[0052]

[0053] where (x′,y′) are the coordinates after rotation, (x,y) are the coordinates before rotation, and θ is the rotation angle.

[0054] For horizontal flipping: Horizontal flipping is to reverse the image left and right, enhancing the model's ability to learn the left - right symmetry of objects. The expression for horizontal flipping is:

[0055] I flip (x,y) = I(W - x,y);

[0056] where W is the width of the image, x is the abscissa of the current pixel, and I(x,y) is the pixel value of the original image.

[0057] For brightness adjustment: Brightness adjustment is to perform a weighted operation on the pixel values of the image to simulate the image performance under different lighting conditions. The brightness adjustment formula is as follows:

[0058] I bright (x,y) = I(x,y)×α + β;

[0059] where α and β are randomly selected coefficients, controlling the scaling factor and offset of brightness respectively, and I(x,y) is the pixel value of the original image.

[0060] For Mosaic enhancement: The basic idea of the Mosaic technique is to splice multiple images together at random positions to form a new large image, which not only increases the diversity of training samples but also makes the target objects appear in different proportions and positions in the image, thereby enhancing the model's object detection ability under various different backgrounds and scales (such as Figure 4 shown). The process of Mosaic specifically includes: Random cropping: Randomly crop a region from each original image, and the size of the cropped region generally has a certain random variation (the cropping size can be adjusted according to requirements); Image splicing: Splice the cropped image patches into a new large image in a certain way. Assuming that 4 images are spliced according to a 2x2 grid, the size of the spliced image is:

[0061] W′ = W1 + W2, H′ = H1 + H2;

[0062] where W1, W2, H1, and H2 are the widths and heights of each spliced image.

[0063] Target annotation update: Since the position of the image has changed, the coordinates of the target object need to be adjusted accordingly. The specific adjustment process can be achieved by calculating the relative position of each target in the spliced image.

[0064] In one implementation, to further improve the efficiency and accuracy of model training, this embodiment provides several more efficient training methods / training strategies:

[0065] By adopting an adaptive learning rate (such as optimization algorithms like AdamW, RAdam, etc.), automatically adjust the learning rate according to the performance of the model during training. This helps to converge quickly in the initial stage of training and avoid overfitting in the later stage of training, thereby improving the stability and generalization ability of the model;

[0066] Or adopt a multi-stage training method. First, use a larger dataset and stronger data augmentation strategies for rough training, and then perform refined training on a smaller target dataset. Such a training strategy can make the model gradually adapt to different training tasks, avoid falling into local optimal solutions prematurely, and thus improve the overall training effect.

[0067] In one implementation, the object detection network includes: a backbone network, an intermediate network, and a detection head module. The structures and functions of each part are as follows:

[0068] The backbone network includes a Focus module, several feature extraction modules, a Spatial Pyramid Pooling Fast module, and a Cross-Stage Partial Self-Attention module cascaded in sequence; each of the feature extraction modules includes a standard convolution module and a C3STR module; the backbone network is used to extract features from the input image to obtain an initial feature map.

[0069] Specifically, for the Focus module: Considering computational efficiency, in this embodiment, the Focus module is introduced at the first layer of the object detection network. Traditional downsampling methods (such as max pooling, average pooling) will lose image information. The Focus module retains all pixel information through a reorganization operation, while achieving downsampling, reducing the computational load of subsequent network layers, improving the overall efficiency, and converting spatial information into channel information, which is beneficial for feature extraction. For the feature extraction module: Each such module includes a standard convolution module (Conv module) and a C3STR module. The standard convolution module contains a convolutional layer, batch normalization, and an activation function, and is the most basic feature extraction unit in the network. The C3STR module effectively improves the efficiency of feature extraction by introducing a spatial attention mechanism. Especially in complex backgrounds and pose changes, it can more accurately locate acupoints. For the Spatial Pyramid Pooling Fast module (SPPF module): This module extracts multi-scale features by using pooling kernels of multiple different sizes, enhancing the model's detection ability for objects of different scales. For the Cross-Stage Partial Self-Attention module (C2PSA module): This module combines the advantages of the cross-stage local network and the self-attention mechanism, enabling the model to more effectively capture context information across multiple layers, thereby improving the object detection accuracy.

[0070] For example, as Figure 5 shown, the specific structure of the backbone network includes: a Focus module located at the first layer; four feature extraction modules are cascaded in sequence after the Focus module, and each feature extraction module includes a standard convolution module (Conv module) and a C3STR module; a SPPF module is connected after the fourth feature extraction module; a C2PSA module is connected after the SPPF module.

[0071] The intermediate network (Neck) includes a number of first feature extraction modules and a number of second feature extraction modules cascaded in sequence; each of the first feature extraction modules includes an upsampling module, a splicing module, and a C3STR module cascaded in sequence; the input image of the splicing module in the first feature extraction module includes the output image of the previous layer and the output image of a specific C3STR module in the backbone network; each of the second feature extraction modules includes a standard convolution module and a splicing module cascaded in sequence, and adjacent two second feature extraction modules are connected by a C3STR module; the input image of the splicing module in the second feature extraction module includes the output image of the previous layer and the output image of a specific C3STR module in the first feature extraction module, or the output image of the previous layer and the output image of the cross-stage local self-attention module; the intermediate network is used to perform feature enhancement and fusion on the initial feature map output by the backbone network to obtain the target feature map.

[0072] Specifically, as Figure 5 shown, in this embodiment, the intermediate network is mainly divided into two types of modules, the first feature extraction module and the second feature extraction module, and adjacent two second feature extraction modules are connected by a C3STR module.

[0073] Regarding the first feature extraction module: Each of the first feature extraction modules has a cascaded connection relationship. Each first feature extraction module includes an upsampling module, a splicing module, and a C3STR module cascaded in sequence. Among them, the upsampling module is used to upsample (enlarge) the input feature map to increase the resolution of the feature map. The splicing module in the first feature extraction module is used to splice the output image of the previous layer and the output image of a specific C3STR module in the backbone network, and a specific C3STR module in the backbone network can be selected based on the size of the output image of the previous layer, so as to fuse feature information at different levels to enhance the expression ability of features. The C3STR module in the first feature extraction module is used to perform further feature extraction on the spliced feature map output by the splicing module. Regarding the second feature extraction module: Each of the second feature extraction modules is located after the first feature extraction module and also has a cascaded connection relationship. Each second feature extraction module includes a standard convolution module and a splicing module cascaded in sequence. Among them, the standard convolution module is used to perform feature extraction on the input feature map. The splicing module in the second feature extraction module is used to splice the output image of the previous layer and the output image of a specific C3STR module in the first feature extraction module, or splice the output image of the previous layer and the output image of the cross-stage local self-attention module, so as to enhance the diversity and expression ability of features. The C3STR module in the second feature extraction module is used to connect adjacent two second feature extraction modules to further enhance feature fusion.

[0074] For example, as Figure 5 shown, the specific structure of the intermediate network includes: cascading two first feature extraction modules in sequence. Each first feature extraction module includes a standard convolution module and a splicing module cascaded in sequence. A second feature extraction module is connected after the second first feature extraction module, a C3STR module is connected after the second feature extraction module, and another second feature extraction module is connected after the C3STR module. Each second feature extraction module includes a standard convolution module and a splicing module cascaded in sequence.

[0075] The detection head module (head) is used to predict the target feature map output by the intermediate network to obtain the acupoint recognition information.

[0076] Specifically, the detection head module is the last layer of the object detection network, which is used to predict the target feature map output by the intermediate network and generate the final prediction results, including the category, location, and key points of the object.

[0077] For example, as Figure 5 shown, the detection head module fuses information from three different feature levels for bounding box detection (target location, target class confidence) and obtaining key point information (key point coordinates, key point visibility). All required information is output in a single forward pass.

[0078] In one implementation, the Focus module includes:

[0079] A slicing module for slicing the input image in the length and width dimensions to obtain sub-sliced images;

[0080] A splicing module for splicing the sub-sliced images in the channel dimension to obtain a spliced image;

[0081] A standard convolution module for extracting features from the spliced image to obtain the feature map output by the Focus module.

[0082] Specifically, the specific structure of the Focus module includes: a slicing module, a splicing module, and a standard convolution module. The data flow in the Focus module is as follows: the slicing module receives the feature map input to the Focus module and performs a slicing operation on this map in two dimensions, namely the length (height direction) and width (width direction), thereby obtaining multiple sub-sliced images. Through the slicing operation, a larger image can be segmented into multiple smaller images, reducing the complexity of subsequent processing and improving the computational efficiency. The splicing module receives the multiple sub-sliced images output by the slicing module and is responsible for splicing and combining each sub-sliced image along the channel dimension to form a spliced image, thereby further enriching the feature representation of the image. The standard convolution module receives the spliced image output by the splicing module and uses convolution operations to extract features from it, thereby generating the feature map finally output by the Focus module.

[0083] For example, as Figure 6 shown, in the Focus module, first perform a slicing (Slice) operation on each input feature map in the H and W dimensions, then splice in the channel (Channel) dimension, and then pass through the standard convolution module (Conv), and finally output the feature image.

[0084] Among them, the process of input tensor dimension transformation is as follows:

[0085] Assume the input is: x(B, C, H, W); after slicing: x(B, 4C, H / 2, W / 2).

[0086] The results of the slicing operation are as follows: slice 1: x1 = x[..., ::2, ::2]; slice 2: x2 = x[..., 1::2, ::2]; slice 3: x3 = x[..., ::2, 1::2]; slice 4: x4 = x[..., 1::2, 1::2]. Among them,... means keeping the first two dimensions (B and C) unchanged; ::2 means taking one value every 2 elements in this dimension (the step size is 2); x[..., ::2, ::2] means taking elements with even indices in both height and width; x[..., 1::2, ::2] means taking elements with odd indices in height and even indices in width; x[..., ::2, 1::2] means taking elements with even indices in height and odd indices in width; x[..., 1::2, 1::2] means taking elements with odd indices in both height and width.

[0087] The calculation formula for the splicing operation is as follows: x cat = Concat([1, xx2, x3, x4], dim = 1);

[0088] The calculation formula for the convolution operation is as follows: y = Conv(x cat).

[0089] The Focus module in this embodiment achieves downsampling without information loss through clever slicing and reorganization operations, converts information in the spatial dimension into the channel dimension, increases the richness of features, has high computational efficiency, and reduces the amount of computation for subsequent processing.

[0090] In one implementation, the C3STR module includes:

[0091] A first standard convolution module, used to extract features from an input image to obtain a first feature map;

[0092] A second standard convolution module is used to extract features from the input image to obtain a second feature map;

[0093] A Swin Transformer module, used for extracting features from the first feature map through a layered window self-attention mechanism to obtain a third feature map;

[0094] A splicing module, used for splicing the second feature map and the third feature map to obtain a spliced ​​feature map;

[0095] The third standard convolution module is used to extract features from the concatenated feature map to obtain the feature map output by the C3STR module.

[0096] Specifically, acupuncture points on the back (such as lung points, heart points, etc.) usually present small (about 1-3 cm in diameter) and similar texture features, and the distribution of acupuncture points on the back follows the direction of the meridians (such as the bladder meridian running through the back), and it is difficult for traditional CNNs to capture such cross-regional associations. In order to achieve fine-grained feature extraction and improve anti-interference capabilities, this embodiment introduces the SwinTransformer module in the C3STR module. The window self-attention mechanism of the Swin Transformer module can focus on local high-response areas, for example: by calculating the attention weights of the skin surface micro-convexities and color differences (such as redness and swelling), the sensitivity to the boundaries of acupuncture points is enhanced. Secondly, the global context modeling of the Swin Transformer module can distinguish acupuncture point features from non-target areas, and effectively deal with interference in actual scenes (such as scars, birthmarks, and clothing occlusions): for example, by comparing the difference in smoothness of the surrounding skin to suppress interference from occluders.

[0097] like Figure 7As shown in the figure, the specific structure of the C3STR module includes: three standard convolutional modules, a Swin Transformer module, and a splicing module. The data flow in the C3STR module is as follows: The first standard convolutional module receives the feature map input to the C3STR module for further feature extraction and outputs the first feature map; the second standard convolutional module and the Swin Transformer module each receive the first feature map for further feature extraction. The second standard convolutional module outputs the second feature map, and the Swin Transformer module outputs the third feature map. The splicing module receives the second feature map and the third feature map for feature splicing / fusion and outputs the spliced feature map; the third standard convolutional module receives the spliced feature map for further feature extraction to obtain the final output feature map of the C3STR module. Among them, the three standard convolutional modules perform basic feature extraction work, while the Swin Transformer Block calculates the attention weights within the local window through the window self-attention mechanism (Window Multi-Head Self Attention), and at the same time realizes cross-window information interaction through window shifting (Shifted Window), taking into account both global and local features, significantly improving the model's ability to capture complex textures and long-range dependencies. Generally speaking, the C3STR in this embodiment captures local textures in the shallow network through hierarchical window interaction (such as the pore arrangement around a single acupoint); and establishes the topological relationship between acupoint groups in the deep network (such as the regular distance between adjacent acupoints), which can assist in correcting the positioning deviation caused by individual body type differences.

[0098] In one implementation, the Swin Transformer module includes: a number of Swin Transformer layers cascaded in sequence.

[0099] Specifically, as Figure 7 shown, the Swin Transformer Block is composed of multiple Swin Transformer Layers cascaded in sequence. As Figure 8As shown, the structure of each Swin Transformer Layer includes: the first channel dimension normalization layer (Layer Norm), the multi-stage window partitioning module (Window Partition), the position embedding module (PositionEmbedding), the windowed multi-head self-attention module, the window fusion module (Window Reverse), the first fusion module, the second channel dimension normalization layer, the multi-layer perceptron (MLP), and the second fusion module. The data flow in each Swin Transformer Layer is as follows: the first channel dimension normalization layer receives the initial feature map of the input Swin Transformer Layer and normalizes it in the channel dimension to obtain the first normalized data. The multi-stage window partitioning module receives the first normalized data and performs multi-stage window partitioning on it to gradually expand the receptive field and adapt to the multi-scale characteristics of the acupoint area. The position embedding module receives the partitioned image windows output by the multi-stage window partitioning module and explicitly models the spatial relative position relationship between acupoints to solve the problem that traditional CNNs are sensitive to geometric transformations. The window fusion module receives the image windows with the modeled spatial relative position relationship and fuses each image window. The first fusion module receives the initial feature map of the input SwinTransformer Layer and the image window fusion data output by the window fusion module and outputs the first fusion data. The second channel dimension normalization layer receives the first fusion data and normalizes it in the channel dimension to obtain the second normalized data. The second normalized data enters the second fusion module after passing through the multi-layer perceptron, and at the same time, the first fusion data output by the first fusion module also enters the second fusion module. After the two types of data are merged, the output data of this Swin TransformerLayer is obtained.

[0100] Furthermore, the Window Multi-Head Self Attention (W-MSA) in Swin Transformer can solve the problem of high computational complexity of traditional Transformers in high-resolution image processing. Among them, the specific calculation process of W-MSA includes:

[0101] Query, Key, Value calculation: Q = XW q , K = XW k , V = XW v ;

[0102] Among them, Query (Q): represents the information to be queried at the current position; Key (K): represents the key information at each position, used to match with the Query; Value (V): represents the information content actually contained at each position; Wq represents the query weight matrix; Wk represents the key weight matrix; Wv represents the value weight matrix.

[0103] Attention weight calculation:

[0104] Among them, T represents the transpose of the matrix; d k represents the dimension of the K matrix; B represents the relative position encoding.

[0105] Relative position encoding: B i,j = B[(i - j) x , (i - j) y ;

[0106] Among them, B i,j represents the relative position relationship between position i and position j in the feature map, which is a learnable bias term used to enhance the attention mechanism's perception of position information.

[0107] The final MSA output: MSA(X) = [head1; head2;...; head h W O ;

[0108] Among them, head h represents the h-th attention head; W O represents the corresponding weight matrix.

[0109] The overall calculation process of the C3STR module: Y1 = Conv1(X); Y2 = Conv2(X); Y STR = SwinTransformer(Y1); Output = Conv3(Concat[Y STR , Y2]);

[0110] Among them, Conv1 represents the first standard convolution module, Y1 represents the first feature map; Conv2 represents the second standard convolution module, Y2 represents the second feature map; Conv3 represents the third standard convolution module, Y STR represents the third feature map.

[0111] In this embodiment, the W-MSA reduces the computational complexity while retaining the global modeling potential of the Transformer through local window partitioning and relative position encoding. Combining the window shifting strategy of the SW-MSA, the Swin Transformer achieves efficient long-range dependence capture and becomes the mainstream architecture to replace CNNs in vision tasks. Its design concept has also been further extended by subsequent works (such as frequency domain enhancement and cross-modal attention), promoting the continuous evolution of vision Transformers.

[0112] To demonstrate the technical effects of the present invention, relevant experiments were conducted using the technical solution of the present invention. The experimental results show that the method proposed in the present invention has high accuracy and robustness in the task of human back acupoint recognition, can accurately identify 27 acupoints in a variety of actual environments, and can be widely applied to the fields of intelligent medicine, rehabilitation assistance, and health management.

[0113] In the acupoint recognition task, choosing appropriate evaluation metrics is crucial for model optimization. Commonly used evaluation metrics include accuracy, precision, recall, F1-score, IoU, localization error, mAP, etc., which help comprehensively evaluate the performance of the model. According to the focus of different tasks, these metrics can be comprehensively selected for evaluation according to requirements such as precision, recall, and localization accuracy, so as to better optimize and improve the performance of the model. The present invention evaluates the model performance from three metrics: precision, recall, and mean average precision.

[0114] Precision: Precision is used to measure the proportion of acupoints predicted by the model that are actually correct. High precision indicates that the model will not give too many incorrect predictions when predicting acupoints.

[0115]

[0116] Among them, TP (True Positive) is the number of acupoints correctly identified, and FP (False Positive) is the number of non-acupoints wrongly predicted as acupoints.

[0117] Recall: Recall measures the proportion of actual acupoints that are identified by the model. High recall means that the model can identify more actual acupoints and reduce the situation of missed detections.

[0118]

[0119] Among them, FN (False Negative) is the number of missed acupoints, that is, the actual acupoints that are not identified by the model.

[0120] Mean Average Precision (mAP): Mean Average Precision (mAP) is a commonly used evaluation metric in object detection tasks and is applicable to multi-class recognition tasks. For acupoint recognition, the precision can be calculated for each acupoint category, and finally the average precision of all acupoint categories is obtained. The higher the mAP value, the better the recognition effect of the model.

[0121]

[0122] Among them, N is the total number of acupoint categories, and AP i is the average precision of each acupoint category.

[0123] Experimental results: The evaluation of key points (KeyPoints) usually focuses on the accuracy and precision of their predictions. Bounding box evaluation not only involves correctly identifying the location of the target acupoint but also needs to consider the overlap between the predicted box and the actual target area.

[0124] Table 1. Comparison of multi-evaluation metric parameters

[0125]

[0126] Analysis of the evaluation results of the bounding box shows that: The model of the present invention has strong accuracy and robustness in acupoint recognition tasks, can efficiently identify most target acupoints, and the overlap between the prediction results and the true values is relatively high. The high performance of mAP and other evaluation metrics also indicates that the overall performance of the model is excellent, with a 0.5% improvement in precision, a 0.6% improvement in recall, and a 0.4% improvement in mAP-50, making it suitable for actual acupoint recognition applications.

[0127] In the present invention, the model is optimized by introducing the Focus module and the C3STR module to improve the performance of acupoint recognition tasks. When comparing the model parameter quantities, the Focus module and the C3STR module can effectively reduce the computational amount while maintaining the high precision of the model. The advantages of this technical solution are not only reflected in the recognition accuracy of the model but also significantly reduce the demand for computing resources, making it suitable for deployment on edge devices and low-resource environments.

[0128] Table 2. Comparison of parameter quantities and effects

[0129]

[0130] Through the innovative technology of the present invention, especially the introduction of the Focus and C3STR modules, the model performance is retained or even improved while the number of parameters is significantly reduced. This makes the model more efficient and scalable, and also suitable for deployment on resource-constrained devices. Specifically, the number of model parameters is reduced by 59%, and the recognition accuracy is almost not lost. This feature gives the present invention strong advantages in practical applications, especially for scenarios that require real-time processing and efficient computing. The comparison results of the prediction effects between the object detection model in the present invention and the traditional model are as Figure 9 shown.

[0131] In summary, the innovation points / advantages of the present invention mainly lie in:

[0132] 1. Combining the automatic recognition and positioning technology of human back acupoints, and improving the accuracy, robustness and efficiency of acupoint recognition through a series of processes such as data collection and annotation, data augmentation, model improvement and training.

[0133] 2. Based on the traditional deep learning model, combining the Focus module and the C3STR module, and significantly improving the recognition accuracy of the model for key points and reducing the consumption of computing resources by enhancing the feature extraction ability of the model and reducing the model parameters. The Focus module focuses on enhancing the features of important regions, and the C3STR module improves the feature learning method of the network, enabling the model to better focus on the target acupoints when dealing with complex human backgrounds.

[0134] 3. Providing a dedicated data augmentation technology for the acupoint recognition task. High-quality annotation of human back acupoints is carried out, and a variety of innovative data augmentation methods are adopted, such as rotation, translation, scaling, mirroring, Mosaic, etc., combined with depth map information to enhance the generalization ability and robustness of the model. Especially by using the Mosaic technology, it is possible to process image data from multiple different perspectives at the same time, enhancing the adaptability of the model to acupoint detection in complex environments.

[0135] The possible deformation directions of the present invention are:

[0136] 1. Using a lighter deep learning model: In the future, it may be possible to reduce the computing resource requirements by designing lightweight neural networks (such as MobileNet, EfficientNet). This network structure is more compact and suitable for embedded devices and mobile applications, and has advantages for scenarios with high requirements for real-time processing and low power consumption.

[0137] 2. Multimodal Fusion: To improve the accuracy and robustness of acupoint recognition, it may be considered to fuse data from multiple sensors in the future, such as combining infrared imaging or thermal imaging data. These sensors can help identify the thermal distribution characteristics of the human back and further refine the recognition ability of the model by combining RGBD information. Through multimodal data fusion, the model can comprehensively understand the geometric structure and temperature distribution of the human body, thereby improving the overall recognition effect.

[0138] 3. Adaptive Enhancement Strategy: It may be considered to introduce an adaptive data enhancement strategy in the future, such as dynamically adjusting the method and intensity of data enhancement according to the current training stage or the error distribution of the model. Use a stronger enhancement strategy in the initial training stage of the model, and reduce the interference of enhancement during fine-tuning in the later stage to improve the generalization ability of the final model.

[0139] 4. Optimization Methods to Replace Focus and C3STR Modules: The Focus and C3STR modules reduce the number of parameters of the model by streamlining the network structure and feature extraction. In the future, it may be considered to introduce other convolutional optimization methods, such as the Squeeze-and-Excitation module (SE module) or attention mechanisms (such as SE-Net, CBAM, etc.). These methods weight the features, enabling the network to pay more attention to the regions that are important for the final task, further improving the performance and efficiency of the model.

[0140] 5. Combination of Traditional Computer Vision Methods and Deep Learning: It may be combined with the advantages of traditional computer vision techniques and deep learning techniques in the future. For example, features can be extracted through traditional methods such as HOG+SVM and image pyramids, and then combined with a deep learning model for feature classification and regression. This can reduce the dependence of the model on big data and still work effectively in the case of scarce data.

[0141] 6. Real-time Performance Optimization: Considering that the application scenario of the present invention requires real-time processing, the following suggestions for real-time performance optimization are proposed: Edge computing and model compression. To meet the real-time requirements, the model can be deployed to edge computing devices. By using compression techniques such as quantization, pruning, and distillation of the model, the storage space and computational requirements of the model are reduced, while maintaining its accuracy. This is very important for improving the response speed and reducing latency in real-time applications. Low-latency inference framework. Use a low-latency inference framework (such as TensorRT or ONNX Runtime, etc.) for inference acceleration. By optimizing the computational graph and parallelizing the inference process, the running efficiency of the model on edge devices is maximally improved. This can significantly reduce the latency from data input to model output, ensuring that the system can complete tasks within a limited time.

[0142] 7. Application scenario integration: Closely integrate the automatic detection of human acupoints with actual application scenarios. For example, conduct customized design for actual applications such as massage robots, enhance the linkage with automated devices, enable direct application in scenarios such as intelligent massage, and expand the practical value and promotion potential of the technology.

[0143] Based on the above embodiments, the present invention also provides an automatic human back acupoint recognition system based on key point detection, as Figure 10 shown. The system includes:

[0144] An image acquisition module 01, configured to acquire a human back image to be analyzed;

[0145] An image input module 02, configured to input the human back image into a trained target detection network; wherein, the trained target detection network includes a Focus module and several C3STR modules; the Focus module is located at the first layer and is used for slicing, reorganizing, and feature extraction of the input image; the C3STR module is used for feature extraction of the output image of the previous layer through a hierarchical window self-attention mechanism;

[0146] An image detection module 03, configured to obtain acupoint recognition information corresponding to the human back image through the trained target detection network; wherein, the acupoint recognition information includes a back detection frame, the positions and categories of key acupoint points.

[0147] Based on the above embodiments, the present invention also provides a terminal, and its principle block diagram can be as Figure 11 shown. The terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes an automatic human back acupoint recognition method based on key point detection. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.

[0148] Those skilled in the art can understand that Figure 11 the principle block diagram shown in

[0149] In one implementation, more than one program is stored in the memory of the terminal, and the more than one program configured to be executed by more than one processor includes instructions for performing an automatic recognition method of human back acupoints based on key point detection.

[0150] In summary, the present invention discloses an automatic recognition method and system for human back acupoints based on key point detection. By introducing a Focus module and a C3STR module into the target detection network. The Focus module is used to refine the processing of local regions, enhancing the attention to the back and acupoint regions, enabling the target detection network to better capture key features. The C3STR module effectively improves the efficiency of feature extraction by introducing a spatial attention mechanism and can more accurately locate acupoints. After introducing the Focus module and the C3STR module, the target detection model can achieve better detection results when detecting back acupoints with high-density distribution and small target characteristics.

[0151] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for automatic recognition of acupuncture points on the back of the human body based on key point detection, characterized in that: The method comprises: Acquire a human back image to be analyzed; Input the human back image into a trained target detection network; wherein the trained target detection network includes a Focus module and a plurality of C3STR modules; the Focus module is located at the first layer and is used for slicing and reorganizing the input image and extracting features; the C3STR module is used for extracting features from the output image of the previous layer through a layered window self-attention mechanism; The acupoint recognition information corresponding to the human back image is obtained through the trained target detection network; the acupoint recognition information includes a back detection frame, and the positions and categories of key acupoints.

2. The method for automatic recognition of acupuncture points on the back of the human body based on key point detection according to claim 1, characterized in that: The method for constructing the training data set of the trained target detection network includes: Acquire a plurality of original human back images through a depth camera; for each of the original human back images, acquire an enhanced image corresponding to the original human back image through a data enhancement operation, and generate annotation data corresponding to the enhanced image; A data set is generated according to all the enhanced images and the annotated data, and a training data set is obtained by dividing the data set.

3. The method for automatic recognition of acupuncture points on the back of the human body based on key point detection according to claim 2, characterized in that: The data enhancement operation includes at least one of random cropping, random rotation, horizontal flipping, brightness adjustment, image merging, and depth map enhancement.

4. The method for automatic recognition of acupuncture points on the back of the human body based on key point detection according to claim 1, characterized in that: The target detection network includes: The backbone network includes a Focus module, several feature extraction modules, a spatial pyramid pooling fast module and a cross-stage local self-attention module which are cascaded in sequence; each of the feature extraction modules includes a standard convolution module and a C3STR module; the backbone network is used to extract features from the input image to obtain an initial feature map; The intermediate network includes a plurality of first feature extraction modules and a plurality of second feature extraction modules cascaded in sequence; each of the first feature extraction modules includes an upsampling module, a splicing module and a C3STR module cascaded in sequence; the input image of the splicing module in the first feature extraction module includes the output image of the previous layer and the output image of the specific C3STR module in the backbone network; each of the second feature extraction modules includes a standard convolution module and a splicing module cascaded in sequence, and two adjacent second feature extraction modules are connected through a C3STR module; the input image of the splicing module in the second feature extraction module includes the output image of the previous layer and the output image of the specific C3STR module in the first feature extraction module, or the output image of the previous layer and the output image of the cross-stage local self-attention module; the intermediate network is used to perform feature enhancement and fusion on the initial feature map output by the backbone network to obtain a target feature map; The detection head module is used to predict the target feature map output by the intermediate network to obtain the acupoint identification information.

5. The method for automatic recognition of acupuncture points on the back of the human body based on key point detection according to claim 1, characterized in that: The Focus module includes: A slicing module is used to slice the input image in length and width dimensions to obtain sub-slice images; A stitching module, used for stitching the sub-slice images in the channel dimension to obtain a stitched image; The standard convolution module is used to extract features from the stitched image to obtain a feature map output by the Focus module.

6. The method for automatic recognition of acupuncture points on the back of the human body based on key point detection according to claim 1, characterized in that: The C3STR module includes: A first standard convolution module, used to extract features from an input image to obtain a first feature map; A second standard convolution module is used to extract features from the input image to obtain a second feature map; A Swin Transformer module, used for extracting features from the first feature map through a layered window self-attention mechanism to obtain a third feature map; A splicing module, used for splicing the second feature map and the third feature map to obtain a spliced ​​feature map; The third standard convolution module is used to extract features from the concatenated feature map to obtain the feature map output by the C3STR module.

7. The method for automatic recognition of acupuncture points on the back of the human body based on key point detection according to claim 6, characterized in that: The Swin Transformer module includes: a plurality of Swin Transformer layers cascaded in sequence.

8. An automatic recognition system for acupuncture points on the back of the human body based on key point detection, characterized in that: The system comprises: An image acquisition module, used to acquire a human back image to be analyzed; An image input module is used to input the human back image into a trained target detection network; wherein the trained target detection network includes a Focus module and a plurality of C3STR modules; the Focus module is located at the first layer and is used to perform slice reorganization and feature extraction on the input image; the C3STR module is used to perform feature extraction on the output image of the previous layer through a layered window self-attention mechanism; The image detection module is used to obtain the acupoint recognition information corresponding to the human back image through the trained target detection network; wherein the acupoint recognition information includes a back detection frame, and the location and category of key acupoints.

9. A terminal, characterized in that: The terminal includes a memory and one or more processors; the memory stores one or more programs; the program contains instructions for executing the method for automatic identification of acupuncture points on the back of the human body based on key point detection as described in any one of claims 1-7; and the processor is used to execute the program.

10. A computer-readable storage medium having a plurality of instructions stored thereon, characterized in that: The instructions are suitable for being loaded and executed by a processor to implement the steps of the method for automatic identification of acupuncture points on the back of the human body based on key point detection as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Human body back acupoint recognition method and device and computer storage medium

    CN114882526A

  • Model training method and device, infrared small target detection method and device and electronic equipment

    CN116152591A

  • ACU-YOLO deep learning method for rapid human body acupoint recognition

    CN119181111A

Cited By

  • Human body meridian point detection method and device, electronic equipment and storage medium

    CN122115434A