Acupuncture method and robot fusing multi-modal visual data
By fusing multimodal visual data and utilizing improved YOLOv8s and OpenPose algorithms to accurately identify acupoints, the problems of inaccurate positioning and insufficient interaction in existing acupuncture robots have been solved, enabling efficient and personalized acupuncture treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN UNIV
- Filing Date
- 2026-02-26
- Publication Date
- 2026-06-02
AI Technical Summary
Existing acupuncture robots suffer from insufficient acupoint positioning accuracy, lack of personalized interactive functions, inability to adapt to the treatment needs of different patients, and reliance on the experience of acupuncturists.
By employing a method that integrates multimodal visual data, the system acquires three-dimensional human contour and depth image data of patients through a head depth camera and LiDAR. Combined with an improved YOLOv8s target detection model and OpenPose pose recognition algorithm, it identifies key skeletal joints and uses a dual-stream fusion neural network to accurately locate acupoints, thereby achieving intelligent acupuncture treatment.
It improves the accuracy and personalization of acupoint location, lowers the operational threshold, enhances treatment efficiency and quality, and reduces reliance on the experience of acupuncturists.
Smart Images

Figure CN122123874A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the medical field, and in particular relates to an acupuncture method and robot that integrates multimodal visual data. Background Technology
[0002] Acupuncture, as a traditional Chinese medicine therapy, has significant effects in disease treatment and health maintenance. However, traditional acupuncture relies on the experience and techniques of professional acupuncturists, resulting in high labor costs, difficulty in ensuring consistency of operation, and significant influence from the individual condition of the acupuncturist. With the increasing aging population and growing demand for acupuncture, traditional acupuncture methods are struggling to meet market needs.
[0003] Currently, although some acupuncture assistive devices and simple acupuncture robots have appeared on the market, these devices generally suffer from insufficient acupoint positioning accuracy and a lack of effective interaction with patients. Most existing acupuncture robots rely solely on a single sensor for acupoint positioning, failing to comprehensively consider the complex shape of the human body and individual differences, resulting in significant positioning errors and affecting the effectiveness of acupuncture treatment. Furthermore, existing devices lack intelligent interactive functions, preventing the formulation and adjustment of personalized acupuncture plans based on the patient's specific condition. Therefore, developing a humanoid acupuncture robot system capable of accurately locating acupoints and achieving intelligent interaction has significant practical significance and value. Summary of the Invention
[0004] In view of this, this application provides an acupuncture method and robot that integrates multimodal visual data, aiming to accurately locate acupoints on the human body according to the acupuncture plan.
[0005] Firstly, this application provides an acupuncture method that integrates multimodal visual data, including: Collect patient information and determine the acupuncture plan based on the collected patient information; Acquire the patient's body contour data and depth image data; Based on the depth image data, the patient's human body region is detected using a pre-trained human target detection model to obtain the region of interest of the patient's human body; Based on the region of interest, and combined with a human posture recognition algorithm, the key skeletal joints of the human body in the region of interest are identified. Based on key skeletal joints in the region of interest and human contour data, the location information of human acupoints corresponding to the acupuncture plan is determined through a pre-trained acupoint prediction model. Acupuncture treatment is performed on the patient based on the location information of the acupuncture points corresponding to the acupuncture plan.
[0006] Optionally, the human body contour data includes: a three-dimensional point cloud of the human body.
[0007] Optionally, the depth image data includes: a depth map of the human body.
[0008] Optionally, the human target detection model includes an improved YOLOv8s, which includes a backbone network, a neck network, and a task head; The backbone network adopts an improved CSPDarNet, which is obtained by introducing the CBAM attention mechanism and residual connections into the C2f module; the backbone network is used to extract large-scale, medium-scale and small-scale feature maps from the input image respectively. The neck network adopts a bidirectional feature flow PAFPN structure or a PAN-OS structure; the neck network is used to transmit the fused large-scale, medium-scale, and small-scale feature maps. The task head adopts a task adaptive decoupling head, which includes a large target decoupling head, a medium target decoupling head, and a small target decoupling head, respectively used to perform classification and localization operations on the fused large-scale, medium-scale, and small-scale feature maps.
[0009] Optionally, the human pose recognition algorithm includes an improved OpenPose algorithm, which is obtained by employing an HRNet backbone network and an improved PAF algorithm.
[0010] Optionally, the acupoint prediction model employs a dual-stream fusion neural network, comprising: Image processing networks are used to extract texture features from images of regions of interest. Point cloud processing network, used to extract spatial geometric features from 3D point clouds; The fusion network employs mid-term fusion feature splicing and attention mechanism-guided feature fusion to fuse texture features and spatial geometric features, thereby obtaining fused features; The acupoint classification head is used to determine the acupoint type based on fusion characteristics; and the acupoint location head is used to determine the acupoint location based on fusion characteristics.
[0011] Optionally, the fusion network employs an attention-weighted approach to fuse texture features and spatial geometric features.
[0012] Secondly, this application provides an acupuncture robot that integrates multimodal visual data, comprising: The determination module is used to collect patient information and determine the acupuncture plan based on the collected patient information. The acquisition module is used to acquire the patient's human body contour data and depth image data; The detection module is used to detect the patient's human body region based on the depth image data using a pre-trained human target detection model, and obtain the region of interest of the patient's human body. The recognition module is used to identify key skeletal joints of the human body in the region of interest, based on the region of interest and in conjunction with a human posture recognition algorithm. The localization module is used to determine the location information of the acupuncture points corresponding to the acupuncture plan based on the key skeletal joints in the region of interest and human contour data, using a pre-trained acupuncture point prediction model. The acupuncture module is used to perform acupuncture treatment on patients based on the location information of the acupoints corresponding to the acupuncture plan.
[0013] Thirdly, this application provides an electronic device, including the acupuncture robot that fuses multimodal visual data as described above.
[0014] Fourthly, this application provides a computer-readable storage medium storing at least one piece of program code, which is executed by a processor to implement the acupuncture method for fusing multimodal visual data as described in any of the preceding claims.
[0015] The beneficial effects of the technical solution provided in this application include: (1) Precise and personalized acupoint positioning: By combining the human-computer interaction module with the large model, acupuncture points are recommended based on the patient's personalized information. Then, the image acquisition unit is used to accurately identify and locate the points. Compared with traditional methods and existing equipment, this greatly improves the accuracy and personalization of acupoint positioning and can better meet the treatment needs of different patients.
[0016] (2) Efficient human-computer interaction: Rich interaction methods (touchscreen, voice, body language) make communication between doctors and patients and the device more convenient and natural, reduce the operation threshold, and improve treatment efficiency. The large model responds quickly and recommends acupoints, saving time for manual judgment and acupoint search.
[0017] (3) Intelligent operation: The fusion of multimodal data and large models enables the device to make intelligent decisions, automatically generate and optimize acupuncture treatment plans based on patient information, reduce reliance on the experience of acupuncturists, and improve the overall quality of treatment. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1This is a schematic diagram of the structure of an acupuncture robot that integrates multimodal visual data, provided in an embodiment of this application. Figure 2 This is a structural block diagram of an acupuncture robot that integrates multimodal visual data, provided in one embodiment of this application. Figure 3 This is a control flowchart of an acupuncture robot that integrates multimodal visual data, provided in one embodiment of this application. Figure 4 A flowchart illustrating an acupuncture method that integrates multimodal visual data, as provided in an embodiment of this application. Figure 5 A schematic diagram of a human target detection model provided in an embodiment of this application; Figure 6 A schematic diagram of a human pose recognition algorithm provided in an embodiment of this application; Figure 7 A schematic diagram of an acupoint prediction model provided in an embodiment of this application; Figure 8 A structural block diagram of an acupuncture robot that integrates multimodal visual data, provided in another embodiment of this application; Figure 9 This is a structural block diagram of an electronic device provided in an embodiment of this application.
[0020] The attached figures are labeled as follows: 1: Head depth camera; 2: LiDAR; 3: Hand depth camera; 4: Five-finger dexterity hand; 5: Shoulder joint motor; 6: Elbow joint motor; 7: Microphone; 8: Speaker; 9: Touch screen; 10: High-performance edge computing platform; 11: Control microcontroller; 12: Lithium battery; 13: Differential wheel.
[0021] 21: Determination Module; 22: Acquisition Module; 23: Detection Module; 24: Recognition Module; 25: Positioning Module; 26: Acupuncture Module; 31: Processor; 32: Memory. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] See Figures 1 to 3 This application provides an acupuncture robot that integrates multimodal visual data. The robot includes: The main framework includes a multimodal data acquisition module, a data processing and control module, an acupuncture execution module, and an enhanced human-computer interaction module.
[0024] In some examples, the main frame adopts a humanoid design, with a structure adapted to the human body shape. The upper body of the main frame is designed in a humanoid manner, including a head, chest, and arms, facilitating the mimicry of human movements and interaction with the patient. The lower body of the main frame uses a wheeled chassis, powered by a high-capacity lithium battery. The four moving wheels under the wheeled chassis are differential wheels, driven by DC motors to control the movement of the main frame. This design provides excellent flexibility and mobility, allowing for free movement within the treatment area.
[0025] In some examples, the main frame is made of a high-strength, lightweight alloy material to ensure the stability and flexibility of the device.
[0026] In some examples, the multimodal data acquisition module includes: a head depth camera 1, a LiDAR 2, and a hand depth camera 3.
[0027] In some examples, head depth camera 1 may include a Realsense D435i depth camera.
[0028] In some examples, the lidar may include the Livox Mid-360 lidar.
[0029] In some examples, a hand depth camera may include: a Realsense D435i depth camera.
[0030] By integrating a head depth camera 1 and a LiDAR 2 into the head of the humanoid robot's main frame, the 3D LiDAR can quickly scan and acquire high-precision three-dimensional geometric information of the patient, while the head depth camera can capture dense depth and RGB image data of the patient's body surface. High-precision alignment of three-dimensional geometry and visual color texture information is achieved through the intrinsic and extrinsic parameter calibration technology of the radar and camera. Simultaneously, a hand depth camera 3 is installed above the wrist. When the robotic arm moves above the acupoint, the hand depth camera 3 can capture images and depth details of the acupoint area at close range with high precision. Combined with the image processing algorithms of the data processing and control module, precise positioning of the acupoint is achieved.
[0031] In some examples, the data processing and control module employs a high-performance Jetson Orin NX processor to receive RGB images, dense depth data, and point cloud data transmitted from the multimodal data acquisition module. Subsequently, a lightweight YOLOv8 object detection method based on RGB images and dense depth data achieves accurate perception of human contours; and an Openpose human pose recognition algorithm enables precise perception of key human body parts. Key body part data is obtained based on aligned image and point cloud data, and an acupoint recognition method fusing image and point cloud data is trained for acupoint identification and localization. Simultaneously, the data processing and control module connects to a large model system to process data commands during human-computer interaction.
[0032] In some examples, the acupuncture execution module is the upper limb of the acupuncture robot, which includes a five-finger dexterous hand 4, a shoulder joint motor 5, and an elbow joint motor 6. The five-finger dexterous hand 4 can autonomously retrieve acupuncture needles and insert them. The shoulder and elbow joint motors are both high-precision servo motors. Through the high-precision driving of the shoulder and elbow joints, multi-degree-of-freedom movement of the upper limb can be achieved, thereby accurately controlling the upper limb to the designated acupuncture point. This allows the robot to perform functions such as needle retrieval, insertion, and withdrawal according to the acupuncture plan.
[0033] In some examples, the human-computer interaction module includes: a high-definition touchscreen 9, a voice interaction device (microphone 7 and speaker 8), and a camera.
[0034] The human-computer interaction module is located on the front of the main frame. Doctors or patients can input symptom descriptions, medical history and other information through the touch screen, or they can talk to the robot through voice. The robot's built-in large model system quickly recommends suitable acupuncture points based on the input information of doctors or patients (including touch input or voice input) and combined with the theoretical knowledge of traditional Chinese medicine acupuncture.
[0035] See you again Figure 1 The robot also includes: a high-performance edge computing platform 10 and a control microcontroller 11 for intelligent control of the robot and its operator; a lithium battery 12 for powering the robot; and differential wheels 13 for driving the robot's movement.
[0036] The humanoid acupuncture robot system of the present invention has the following advantages: First, precise and personalized acupoint positioning: By combining the human-computer interaction module with a large model, acupuncture points are recommended based on the patient's personalized information, and then the image acquisition unit is used to accurately identify and locate them. Compared with existing acupuncture robots, this greatly improves the accuracy and personalization of acupoint positioning, and can better adapt to the treatment needs of different patients.
[0037] Secondly, efficient human-computer interaction: Rich interaction methods (touchscreen and voice) make communication between doctors and patients and the acupuncture robot faster and more convenient. Connecting the robot to a large model allows for rapid acquisition of acupoint plans based on the model's built-in functions, thus improving treatment efficiency.
[0038] Third, intelligent operation: The integration of multimodal data and large models enables acupuncture robots to make intelligent decisions, automatically generate and optimize acupuncture treatment plans based on patient information, reduce reliance on the experience of acupuncturists, and improve the overall quality of treatment.
[0039] To further describe the specific implementation process of the acupuncture method in this application, this application additionally provides... Figure 4 , Figure 4 A flowchart of an acupuncture method that integrates multimodal visual data, provided in an embodiment of this application, is shown. The method includes: S101. Collect patient information and determine the acupuncture plan based on the collected patient information using a large model.
[0040] In some examples, the specific implementation process of step S101 is as follows: In practical applications, the acupuncture robot device of this application is first placed in a suitable location and activated to enter working mode. The acupuncture robot uses a microphone and speaker to communicate with the patient through the Qwen3-VL visual language model, inquiring about the patient's symptoms, medical history, and other information. Based on traditional Chinese medicine theory and acupoint knowledge, the model initially determines the areas or acupoints requiring acupuncture. The Qwen3-VL visual language model first communicates with the patient through the microphone and speaker, collecting information such as symptoms and medical history. Simultaneously, it combines visual capabilities (if there is image information such as facial color and tongue coating) to convert the patient's spoken language into standardized traditional Chinese medicine terminology. Then, relying on the built-in traditional Chinese medicine theoretical knowledge base and syndrome differentiation rules, it completes the syndrome differentiation and locates the affected meridians and organs. Following the principles of local, distal, and syndrome differentiation-based acupoint selection, it matches acupoints and excludes risky acupoints based on contraindications. Subsequently, it refines parameters such as acupuncture techniques, depth, and treatment course to generate a preliminary plan. Finally, by confirming contraindications and tolerance with the patient a second time and referring to similar past cases, it optimizes the plan to form a personalized and safe acupuncture plan.
[0041] After receiving information, the robot's built-in large model system quickly generates acupuncture point recommendations based on traditional Chinese medicine acupuncture theory and a large amount of case data, and displays them to doctors or patients through a touch screen.
[0042] S102. Obtain depth image data and human body contour data of the patient's body.
[0043] In some examples, the patient's body is scanned using a head-mounted LiDAR (Mid-360 LiDAR) and a head-mounted depth camera (Realsense D435i depth camera) to acquire human contour data and depth image data, which are then transmitted to the data processing and control module.
[0044] S103. Based on the depth image data and the pre-trained human target detection model, the patient's body is detected to obtain the region of interest of the patient's body.
[0045] In some examples, the specific implementation process of step S103 is as follows: The data processing and control module first detects the human body contour based on depth image data (RGB image and depth map formed by dense depth), then identifies key parts, and finally integrates human body contour data (i.e. three-dimensional point cloud) to identify and locate acupoints.
[0046] In some examples, the human target detection model uses an improved YOLOv8s based on depth image data (RGB images and depth maps formed by dense depth). By optimizing the architecture of YOLOv8s, an efficient and accurate human target detector is built.
[0047] See Figure 5 First, we selected YOLOv8s, a lightweight variant of YOLOv8, as the basic detection model, which has a good balance between inference speed and detection accuracy.
[0048] The overall network adopts a modular fully convolutional architecture consisting of a backbone network, a neck feature fusion FusionNet, and a task head. The backbone network is based on an optimized version of CSPDarknet, which achieves efficient extraction of multi-scale features through multi-layer downsampling and cross-stage feature fusion. The C2f module introduces residual connections in its structure to enhance gradient propagation capability and alleviate the gradient vanishing problem in deep networks.
[0049] Based on this, in order to improve the model's ability to represent key target features such as human body contours in complex scenes, the CBAM attention mechanism is introduced into the YOLOv8s backbone network and embedded into the residual branch inside each C2f module, specifically after the output of the Bottleneck convolutional layer.
[0050] CBAM first performs global average pooling and global max pooling on the feature map through the channel attention branch to aggregate spatial dimensional information, and generates channel weights through a shared multilayer perceptron to adaptively enhance important channel features. Then, the channel-weighted features are fed into the spatial attention branch, where channel dimensional information is aggregated through pooling operations, and spatial weight maps are generated using convolution to highlight the target region and suppress background interference. Finally, the enhanced features are seamlessly integrated into the downsampling process of the backbone network.
[0051] This embedding method retains the original residual structure and gradient propagation advantages of the C2f module, while enabling the network to adaptively focus on key features in both channel and spatial dimensions, effectively improving the target representation ability under complex backgrounds and occlusion conditions.
[0052] The neck feature fusion part (i.e., the neck network) adopts a bidirectional feature flow PAFPN (or PAN-OS) structure. It constructs a closed-loop feature transfer mechanism through "top-down semantic sinking" and "bottom-up detail upflow" to achieve full complementarity between deep semantic information and shallow detail information, thereby significantly enhancing the model's ability to detect small targets and occluded targets.
[0053] In terms of detection head design, YOLOv8s adopts a task-adaptive decoupled head and abandons the traditional anchor box mechanism. Based on the Anchor-Free idea, it directly regresses the target center offset, width and height parameters and class probability at each position of the feature map. By separating the classification and regression branches, it reduces task conflicts and improves training stability.
[0054] Regarding the loss function, the model focuses on optimizing the bounding box regression performance, and uses CIoU Loss to comprehensively constrain the overlap between the predicted box and the ground truth box, the distance between the center points, and the consistency of the aspect ratio, thereby further improving the target localization accuracy and overall detection performance.
[0055] Specifically, an image containing a human body for Given a matrix where each element is an integer between 0 and 255, normalize it to... Once within the specified range, the data is input into the backbone network to generate small-scale, medium-scale, and large-scale feature maps. Then, the fused features are output through the neck network. The data are input into the detection head to detect human target bounding boxes of different scales. and mask ,in These are the coordinates of the top left corner of the box. It refers to the width and height of the frame. yes A boolean matrix, where areas covered by human figures are 1 and areas not covered by human figures are 0.
[0056] To further improve the detection accuracy of human contours and adapt to embedded or mobile deployments, model compression technology is adopted, including structured pruning to remove redundant channels and layers, and quantizing the weights of the perceptual training model from FP32 to INT8, which significantly reduces the model size and computational cost while maintaining accuracy as much as possible.
[0057] Model training was conducted on a dataset containing accurately annotated human bounding boxes under various scenes, poses, and lighting conditions, with multi-scale training and validation. The training strategy employed a pre-trained model that was fine-tuned using data annotated in a hospital acupuncture setting. The hospital acupuncture setting primarily included a dataset of 5000 annotated human bounding box images in standing, sitting, and lying positions. During network training, the parameters of a YOLOv8 human bounding box detection model pre-trained on ImageNet were loaded, and then fine-tuned on the 5000 annotated images to improve its accuracy and robustness in human detection within the hospital acupuncture setting.
[0058] The final model will be deployed on the Jetson Nx hardware platform, using the TensorRT inference engine for hardware acceleration to achieve real-time, high recall and high IoU human region detection, providing accurate regions of interest (ROI) for subsequent steps.
[0059] S104. Based on the region of interest, and combined with a human posture recognition algorithm, the key skeletal joints of the human body are identified.
[0060] Step S104 The goal of this stage is to accurately identify key skeletal joints of the human body (such as head, neck, shoulder, elbow, wrist, hip, knee, ankle, etc.) within the human body ROI detected by YOLOv8s.
[0061] See Figure 5 In some examples, the human pose recognition algorithm is based on the OpenPose human pose recognition algorithm. The input is the original image. Human body mask input in YOLOv8s The output is the coordinates of the acupoints. and corresponding acupoint types .
[0062]
[0063] The technical approach adopts the classic OpenPose lightweight algorithm, which uses a bottom-up paradigm. This involves first detecting all possible key points in the human mask in the image, and then associating and grouping them through part affinity fields to form an individual skeleton.
[0064] To improve the accuracy of keypoint localization, especially when dealing with occlusion, complex poses, and small human bodies, several improvements were made to the standard OpenPose: (1) Backbone network replacement / enhancement. The more powerful feature extractor HRNet was used to replace the original VGG network. (2) Improved PAF (Part Affinity Fields) representation algorithm, optimized the connection logic between keypoints, and improved the grouping accuracy in multi-person scenes. When annotating limbs, PAF is a 2D vector of each limb of the body, and the positional and orientation information between limb regions must be maintained.
[0065] Specifically, let the input depth image be: .in These represent the height and width of the depth image, respectively, with 3 representing the RGB channels. In the improved model, HRNet is used as the backbone network. Its key feature is maintaining a high-resolution feature flow and obtaining highly expressive keypoint features through multi-scale fusion. HRNet outputs a series of feature maps at different resolutions. After fusion, the final high-resolution feature map is obtained: Fusion refers to cross-resolution feature fusion (including operations such as upsampling, element-wise weighting, or channel concatenation). The final feature map dimension is: This is the downsampling ratio (usually 4 or 8). This represents the number of feature channels.
[0066] For each person's K key points (e.g., head, shoulder, elbow, knee, foot, etc.), the network outputs a K-channel key point heatmap: ,in This represents the heatmap prediction function generated by the convolutional layer. The heatmap value at each pixel location. This indicates that the coordinates represent a key point. Confidence level. Hotspots of truth-annotated key points. Figure 1 Generally generated using a two-dimensional Gaussian function:
[0067] During training, the loss function is typically MSE or focal loss. .
[0068] For L limb connections in the human body (e.g., shoulder-elbow, elbow-wrist), define their corresponding two-dimensional vector fields. : .in Indicated in pixels The direction vector of the limb (a unit vector from one end of the connection to the other). The optimized PAF generation process is as follows: ,in It is a convolutional prediction head specifically designed for learning limb connections. The ground truth labeled PAF vector is defined as:
[0069] in , representing the connection between two key points and The direction. The loss function uses vector field regression:
[0070] This module outputs precise 2D human body keypoint coordinates and their connections, providing an accurate location reference framework for acupoint recognition. Model Output ,in Heatmap of key points; This is a limb connection vector field (PAF). The complete human skeletal structure can be reconstructed by grouping key points.
[0071] The training process uses a human pose dataset with precise keypoint annotations and implements data augmentation for pose recognition (such as rotation, scaling, and elastic deformation).
[0072] S105. Based on the region of interest and human body contour data, determine the location information of the human acupoints corresponding to the acupuncture plan through a pre-trained acupoint prediction model.
[0073] The ultimate goal of this stage is to use the visual information (human outline, key points) provided in the first two steps, combined with depth information, to accurately identify and locate specific acupoints in three-dimensional space.
[0074] See Figure 7 In some examples, the specific implementation process of step S105 is as follows: First, ensure that the depth image data (depth map) and human contour data (3D point cloud) are strictly registered and aligned in time and space, achieved using a RealSense D435i RGB-D camera and mid360 LiDAR with calibrated intrinsic and extrinsic parameters. The model's input includes RGB images and depth images. Let the input RGB image be denoted as... Depth image is denoted as By calibrating the camera and mapping the point cloud, a 3D point cloud with color information is generated through registration. .in Let the coordinates be the points. This corresponds to the color value.
[0075] In the human detection stage, the YOLOv8 model is used to detect human regions in the input image, obtaining the Region of Interest (ROI): Subsequently, OpenPose was used to extract a set of key points on the human body. Based on this, local interest regions are defined around each key point. Extracting image patches and point cloud subsets corresponding to local regions of interest from 2D images and depth data. , .
[0076] After entering the two-stream network, the model extracts features from both image and point cloud modalities. RGB image patches are encoded using a ResNet network to obtain texture features. The point cloud data is processed by a Point Transformer network to extract spatial geometric features. The feature interactions within the Point Transformer are described by an attention mechanism:
[0077] The feature reweighting of the point neighborhood is achieved through an attention mechanism.
[0078] Subsequently, the cross-modal feature synthesis representation is achieved in the fusion layer through attention-weighted concatenation. The fusion function is:
[0079] The weight coefficients are hyperparameters. The fused feature vector... It also includes image texture and spatial layout information.
[0080] In the output phase, the network includes two task branches: acupoint classification and acupoint spatial localization.
[0081] The classification branch uses fully connected layers and activation functions for probability prediction. , in This indicates the probability of the existence of acupoints corresponding to each category.
[0082] The localization branch directly regresses the position coordinates of the acupoint in three-dimensional space: ,in .
[0083] Accordingly, the classification task uses cross-entropy loss:
[0084] Localization tasks use L1 distance or smoothed L1 loss.
[0085] The entire model achieves a total loss function through joint optimization.
[0086] in is a hyperparameter used to balance the weights of classification and regression tasks.
[0087] Finally, the model outputs a prediction result for each detected acupoint. ,in Represents class probability, This represents the predicted coordinates of the acupoint in three-dimensional space.
[0088] This structure enables the processing of input images. With depth map The end-to-end learning process for acupoint recognition and localization enables high-precision three-dimensional human acupoint recognition and localization under complex human postures, occlusion, and small-sized target conditions.
[0089] S106. Acupuncture treatment is performed on the patient based on the location information of the acupoints.
[0090] The three algorithms executed in the aforementioned steps S103, S104, and S105 work together to form a complete human acupoint perception and positioning technology system, from human contour recognition to key part recognition and then to acupoint recognition, providing strong technical support for the execution of robotic acupuncture.
[0091] If doctors or patients have questions about the recommended treatment plan from the large model, they can interact with the device further via voice or touchscreen to make adjustment suggestions; the Realsense D435i camera on the head captures the body language of doctors or patients to help understand their needs, and the large model system optimizes the acupoint treatment plan based on the feedback.
[0092] Once the acupuncture point information corresponding to the acupuncture plan is determined, the robotic arm moves the RealSense D435i depth camera above the wrist to the recommended acupuncture point to capture images of the acupuncture point area at close range. The data processing and control module uses image processing algorithms to analyze the images and, combined with the acupuncture point information recommended by the large model, accurately determines the acupuncture point position under the robotic arm, making it easier to perform acupuncture.
[0093] The robot's lower body features a wheeled chassis design for easy and flexible movement. A high-capacity lithium battery is integrated within the chassis, providing power for extended operation. Driven by differential wheels, the robot can freely move within the treatment area, quickly accessing different parts of the patient's body for acupuncture treatment.
[0094] The data processing and control module generates control commands based on the determined acupuncture point locations, controlling the right arm and dexterity hand movements of the acupuncture execution module to accurately grasp and insert the acupuncture needles into the designated acupuncture points, performing treatment according to the set needling angle, depth, and frequency. During treatment, the doctor can view the treatment progress in real time through the human-computer interaction module and, if necessary, interact with the device again to adjust the treatment plan. After treatment, the doctor can also view the relevant data records of this treatment through the human-computer interaction module, providing a reference for subsequent treatments.
[0095] The robot interacts with the patient through a large-scale model. The system combines physiological data, historical case records, and a big data model to automatically identify acupuncture points and perform acupuncture treatment. The acupuncture needle insertion mechanism in the right arm employs a precision control system, accurately inserting the needles into the acupuncture points based on data from a depth camera and sensors, and operating according to pre-set acupuncture techniques and parameters. During acupuncture, the robot can monitor the patient's reactions and the state of the acupuncture points in real time, adjusting the acupuncture plan as needed. After the acupuncture is completed, the robot retracts the needles, moves them to a suitable position, and awaits its next task.
[0096] Figure 8 This is a structural block diagram of an acupuncture robot that integrates multimodal visual data, provided in one embodiment of this application. See also... Figure 8 ,include: Module 21 is used to collect patient information and determine the acupuncture plan based on the collected patient information. The acquisition module 22 is used to acquire human body contour data and depth image data of the patient's body; The detection module 23 is used to detect the patient's human body region based on the depth image data using a pre-trained human target detection model, and obtain the region of interest of the patient's human body. The recognition module 24 is used to identify key skeletal joints of the human body in the region of interest based on the region of interest and in combination with a human posture recognition algorithm. The positioning module 25 is used to determine the location information of the human acupoints corresponding to the acupuncture plan based on the key skeletal joints in the region of interest and human contour data, through a pre-trained acupoint prediction model. The acupuncture module 26 is used to perform acupuncture treatment on the patient based on the location information of the acupoints corresponding to the acupuncture plan.
[0097] Figure 9 This is a structural block diagram of an electronic device provided according to an embodiment of this application. See also... Figure 9 Electronic devices may include Figure 8The acupuncture robot that integrates multimodal visual data. Typically, the electronic device includes a processor 31 and a memory 32. The processor 31 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 31 can be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 31 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. The memory 32 may include one or more computer-readable storage media, which may be non-transitory. The memory 32 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage medium in memory 32 is used to store at least one instruction for execution by processor 31 to implement the acupuncture method for fusing multimodal visual data performed by an electronic device, as provided in the method embodiments of this application.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An acupuncture method that integrates multimodal visual data, characterized in that, include: Collect patient information and determine the acupuncture plan based on the collected patient information; Acquire the patient's body contour data and depth image data; Based on the depth image data, the patient's human body region is detected using a pre-trained human target detection model to obtain the region of interest of the patient's human body; Based on the region of interest, and combined with a human posture recognition algorithm, the key skeletal joints of the human body in the region of interest are identified. Based on key skeletal joints in the region of interest and human contour data, the location information of human acupoints corresponding to the acupuncture plan is determined through a pre-trained acupoint prediction model. Acupuncture treatment is performed on the patient based on the location information of the acupuncture points corresponding to the acupuncture plan.
2. The acupuncture method for fusing multimodal visual data according to claim 1, characterized in that, The human body contour data includes: a three-dimensional point cloud of the human body.
3. The acupuncture method for fusing multimodal visual data according to claim 1, characterized in that, The depth image data includes: a depth map of the human body.
4. The acupuncture method for fusing multimodal visual data according to claim 1, characterized in that, The human target detection model includes an improved YOLOv8s, which includes a backbone network, a neck network, and a task head. The backbone network adopts an improved CSPDarNet, which is obtained by introducing the CBAM attention mechanism and residual connections into the C2f module; the backbone network is used to extract large-scale, medium-scale and small-scale feature maps from the input image respectively. The neck network adopts a bidirectional feature flow PAFPN structure or a PAN-OS structure; the neck network is used to transmit the fused large-scale, medium-scale, and small-scale feature maps. The task head adopts a task adaptive decoupling head, which includes a large target decoupling head, a medium target decoupling head, and a small target decoupling head, respectively used to perform classification and localization operations on the fused large-scale, medium-scale, and small-scale feature maps.
5. The acupuncture method for fusing multimodal visual data according to claim 1, characterized in that, The human pose recognition algorithm includes an improved OpenPose algorithm, which is obtained by using the HRNet backbone network and an improved PAF algorithm.
6. The acupuncture method for fusing multimodal visual data according to claim 1, characterized in that, The acupoint prediction model employs a dual-stream fusion neural network, including: Image processing networks are used to extract texture features from images of regions of interest. Point cloud processing network, used to extract spatial geometric features from 3D point clouds; The fusion network employs mid-term fusion feature splicing and attention mechanism-guided feature fusion to fuse texture features and spatial geometric features, thereby obtaining fused features; The acupoint classification head is used to determine the acupoint type based on fusion characteristics; and the acupoint location head is used to determine the acupoint location based on fusion characteristics.
7. The acupuncture method for fusing multimodal visual data according to claim 6, characterized in that, The fusion network uses an attention-weighted approach to fuse texture features and spatial geometric features.
8. An acupuncture robot that integrates multimodal visual data, characterized in that, include: The determination module is used to collect patient information and determine the acupuncture plan based on the collected patient information. The acquisition module is used to acquire the patient's human body contour data and depth image data; The detection module is used to detect the patient's human body region based on the depth image data using a pre-trained human target detection model, and obtain the region of interest of the patient's human body. The recognition module is used to identify key skeletal joints of the human body in the region of interest, based on the region of interest and in conjunction with a human posture recognition algorithm. The localization module is used to determine the location information of the acupuncture points corresponding to the acupuncture plan based on the key skeletal joints in the region of interest and human contour data, using a pre-trained acupuncture point prediction model. The acupuncture module is used to perform acupuncture treatment on patients based on the location information of the acupoints corresponding to the acupuncture plan.
9. An electronic device, characterized in that, This includes the acupuncture robot that integrates multimodal visual data as described in claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is executed by a processor to implement the acupuncture method for fusing multimodal visual data as described in any one of claims 1 to 7.