Massage robot multi-mode fusion treatment system and method based on 3D vision

Through a multimodal fusion treatment system based on 3D vision, combined with image recognition and hybrid ant colony algorithm to optimize path planning, real-time monitoring of skin pressure and dynamic adjustment of massage intensity, the problem that existing massage robots cannot be dynamically adjusted is solved, and the accuracy and safety of massage are improved.

CN120661365AInactive Publication Date: 2025-09-19CHANGSHA KANGMIN MEDICAL DEVICE TECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510762289.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing massage robots cannot dynamically adjust according to the user's body shape or massage area, resulting in collisions of the robotic arms or unstable massage effects. They are unable to perceive the user's pain or fatigue state in real time, which may lead to excessive or uncomfortable massage.

Method used

A multimodal fusion treatment system based on 3D vision is adopted, including image acquisition, area recognition, massage terminal, pain recognition and strength adjustment modules. It uses 3D vision to identify human body features, combines the hybrid ant colony algorithm to optimize path planning, monitors skin pressure in real time and dynamically adjusts massage strength, and uses multimodal data fusion analysis to identify pain.

Benefits of technology

It improves the accuracy and safety of massage operations, avoids excessive stimulation, significantly improves user comfort and safety, and adapts to individual differences among different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120661365A_ABST
    Figure CN120661365A_ABST
Patent Text Reader

Abstract

The invention provides a massage robot multi-mode fusion treatment system and method based on 3D vision, and relates to the technical field of massage robots. The region recognition module is used for establishing a human body recognition model and recognizing human body features in the processed image; the massage terminal is used for carrying out massage operation according to parts needing to be massaged and human body characteristics and monitoring skin pressure of the massaged parts in real time; the pain sense recognition module is used for human body pain sense recognition; the force adjusting module is used for adjusting skin pressure. According to the method, pertinence and safety of massage operation are remarkably improved through multi-modal image preprocessing and an initialized human body recognition model of a fusion segmentation model and a key point detection model; a global path strategy is generated and optimized through a hybrid ant colony algorithm, efficient and accurate motion control of the mechanical arm is achieved, through multi-modal data fusion analysis, the pain feeling of a user is sensed in real time, the massage strength is dynamically adjusted, and the comfort and the safety coefficient are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of massage robots, and in particular to a 3D vision-based multimodal fusion treatment system and method for massage robots. Background Art

[0002] Massage robots, as an emerging technology product, have garnered widespread attention in recent years. With the accelerating pace of life and increasing stress, the demand for massage has gradually increased, driving rapid development in this field. Massage robots utilize advanced mechanical structures and flexible materials to mimic traditional massage techniques, leveraging various sensors and algorithms to achieve precise massage effects. Technically, the core of massage robots lies in their motion control systems and human-machine interaction design. These motion control systems typically employ high-precision servo motors and intelligent algorithms, adjusting massage intensity and tempo through real-time feedback to suit the user's individual needs. Furthermore, integrated tactile sensors enable the robot to sense the user's physical condition, providing a more personalized service. Furthermore, with the continuous advancement of artificial intelligence and machine learning technologies, massage robots can analyze user data, optimize massage plans, and enhance the user experience. The application of deep learning models enables robots to quickly learn and adapt to different users, providing customized massage services for each user. In terms of design, modern massage robots emphasize ergonomics and aesthetics, striving to strike a balance between functionality and aesthetics to enhance user acceptance and user experience. At the same time, with the development of the Internet of Things (IoT), massage robots are increasingly being integrated with smart home systems, enabling remote control and data sharing, providing users with a more convenient lifestyle. As an interdisciplinary product, massage robots integrate advanced technologies such as mechanical engineering, electronics, and artificial intelligence, demonstrating promising development prospects and broad market potential.

[0003] Existing massage robots typically use preset paths or simple trajectory-following strategies, which are unable to dynamically adjust to the user's body shape or massage area, easily leading to collisions and unstable massage effects. They also typically use fixed force or simple force feedback control, which cannot sense the user's pain or fatigue in real time. This can lead to excessive massage, discomfort, or even injury.

[0004] In order to solve the above-mentioned defects in the existing technology, this technical solution proposes a 3D vision-based multimodal fusion treatment system and method for massage robots. Summary of the Invention

[0005] The present invention provides a 3D vision-based multimodal fusion treatment system and method for a massage robot to address the deficiencies in the prior art.

[0006] In one aspect, the present invention provides a 3D vision-based massage robot multimodal fusion treatment system, comprising:

[0007] An image acquisition module is used to acquire human body images, pre-process the human body images, and output the processed images;

[0008] A region recognition module is used to establish a human body recognition model based on a historical human body database, and to identify human body features in the processed image based on the human body recognition model;

[0009] Massage terminal, used to perform massage operations according to the required massage parts and human body characteristics, monitor the skin pressure of the massage parts in real time, and output massage terminal data;

[0010] The pain recognition module is used to perform human pain recognition based on multimodal feature analysis, human body characteristics and massage terminal data, and output the pain recognition results;

[0011] The force adjustment module is used to adjust skin pressure based on the pain recognition results.

[0012] According to the 3D vision-based multimodal fusion treatment system of the massage robot provided by the present invention, the image acquisition module includes an image acquisition unit and an image preprocessing unit; the image acquisition unit is used to acquire human body images, and the image preprocessing unit is used to preprocess the human body images and output the processed images.

[0013] According to the 3D vision-based multimodal fusion therapy system for massage robots provided by the present invention, the steps of preprocessing the human body image by the image preprocessing unit include:

[0014] Perform geometric correction and image correction on human body images and output standardized images;

[0015] Perform Gaussian filtering on the standardized image to output a low-noise image;

[0016] Perform image enhancement on low-noise images, including sharpening and edge detection, and output feature-enhanced images;

[0017] A cutout algorithm is used to separate the human body area from the background area in the feature-enhanced image, and the processed image is output.

[0018] According to the 3D vision-based multimodal fusion treatment system for massage robots provided by the present invention, the region recognition module includes a modeling unit and a region recognition unit; the modeling unit is used to establish a human body recognition model based on a historical human body database, and the region recognition unit is used to identify human body features in the processed image based on the human body recognition model.

[0019] According to the 3D vision-based multimodal fusion therapy system for massage robots provided by the present invention, the steps of establishing a human body recognition model by the modeling unit include:

[0020] Preprocess the historical human images in the historical human database and output standardized historical images;

[0021] Extract features from standardized historical images and output human body image features;

[0022] Divide the human body image features into data sets, including training sets and validation sets;

[0023] Establish an initialized human recognition model based on the fusion of segmentation model and key point detection model;

[0024] Use the classification cross entropy function as the loss function to initialize the human recognition model;

[0025] Use the training set to train the human body recognition model, and use the validation set to verify the loss of the loss function. When the loss tends to be stable, stop training and obtain the human body recognition model.

[0026] According to the 3D vision-based massage robot multimodal fusion treatment system provided by the present invention, the massage terminal includes a control terminal, a robotic arm and a massage claw; the control terminal is used to output control instructions for the position of the robotic arm based on the required massage part and according to the characteristics of the human body; the robotic arm is used to adjust its own position according to the control instructions; the massage claw is used to perform massage operations based on its own position and monitor the skin pressure of the massage part in real time.

[0027] According to the 3D vision-based multimodal fusion therapy system for massage robots provided by the present invention, the step of controlling the terminal to output control instructions for the position of the robotic arm according to human body characteristics includes:

[0028] Extract the required massage parts from the case database, match the human body features with the required massage parts, and output the matching results;

[0029] According to the matching results and human body characteristics, the robot arm position is planned based on the hybrid ant colony algorithm and combined with the human body environment information to obtain the optimal path strategy;

[0030] Convert the optimal path strategy into the control signal of the robot arm and output the robot arm control instruction.

[0031] According to the 3D vision-based multimodal fusion therapy system for massage robots provided by the present invention, the steps of controlling the terminal to plan the path of the robot arm position based on the hybrid ant colony algorithm include:

[0032] Discretize the workspace in the human environment information into a grid or construct a free space map, mark obstacles and feasible areas; define the initial position and target point of the robot arm;

[0033] Initialize algorithm parameters, including population size;

[0034] Each individual in the population starts from the initial position and selects the next node based on the current pheromone concentration and heuristic information. The formula is expressed as:

[0035]

[0036] Where, represents the probability of individual k moving from the current position i to the next position j, τ ij Indicates the pheromone concentration left by the individual on the path (i, j), t is the time window, allowed k represents the target point set of individual k, α is the pheromone influence constant, β is the heuristic function influence constant, is the heuristic information, d ij is the distance between the individual’s current position and the next position j, d jw is the distance between the next position j of the individual and the final target point w;

[0037] Calculate the total cost for each feasible path and update the pheromone; when the change in the optimal path cost for multiple generations is less than the threshold, stop updating and output the node sequence;

[0038] Connect each node in the node sequence according to the time window order to obtain the optimal path strategy.

[0039] According to the 3D vision-based multimodal fusion therapy system for massage robots provided by the present invention, the steps of the pain recognition module performing human pain recognition include:

[0040] Collect multimodal data of the human body through multiple sensors;

[0041] Perform feature extraction on multimodal data, including extracting human skin electrical response signals, speech features, and behavioral features;

[0042] A deep learning model is used to fuse skin electrical response signals, speech features, and behavioral features to output pain recognition results.

[0043] The multimodal fusion treatment method of a massage robot based on 3D vision provided by the present invention includes:

[0044] Collect human body images, pre-process the human body images, and output the processed images;

[0045] Establishing a human body recognition model based on a historical human body database, and identifying human body features in the processed image based on the human body recognition model;

[0046] Perform massage operations according to the required massage parts and human body characteristics, monitor the skin pressure of the massage parts in real time, and output massage terminal data;

[0047] Based on multimodal feature analysis, human pain recognition is performed according to human characteristics and massage terminal data, and the pain recognition results are output;

[0048] Adjust skin pressure based on pain recognition results.

[0049] The present invention provides a 3D vision-based massage robot multimodal fusion treatment system and method. By adopting a human body recognition model initialized by a multimodal image preprocessing and fusion segmentation model and a key point detection model, the system can accurately identify human body features and locate massage areas, significantly improving the pertinence and safety of massage operations. The control terminal combines human body features with environmental information, and uses a hybrid ant colony algorithm to generate and optimize the global path strategy to achieve efficient and precise motion control of the robotic arm. Through multimodal data fusion analysis, it can perceive the user's pain in real time and dynamically adjust the massage intensity to avoid excessive stimulation, significantly improving comfort and safety factors. The system integrates multi-source information such as images, pressure, and physiological signals, and combines deep learning models for feature analysis, effectively improving the accuracy of pain recognition and human feature recognition, and adapting to individual differences among different users. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0051] Figure 1 Schematic diagram of the structure of a 3D vision-based multimodal fusion therapy system for massage robots provided by an embodiment of the present invention;

[0052] Figure 2 This is a flow chart of the 3D vision-based multimodal fusion treatment method of a massage robot provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0053] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0054] Example 1:

[0055] The following combination Figure 1-Figure 2The present invention describes a 3D vision-based massage robot multimodal fusion treatment system and method.

[0056] like Figure 1 As shown, the 3D vision-based massage robot multimodal fusion treatment system provided by the embodiment of the present invention includes: an image acquisition module, a region recognition module, a massage terminal, a pain recognition module and a strength adjustment module.

[0057] The image acquisition module is used to acquire human body images, pre-process the human body images, and output the processed images. The image acquisition module includes an image acquisition unit and an image pre-processing unit. The image acquisition unit is used to acquire human body images, and the image pre-processing unit is used to pre-process the human body images and output the processed images. The steps include:

[0058] First, the human body image is subjected to geometric correction and image correction, including adjusting the perspective deformation, rotation angle and scale of the image to ensure that the posture and size of the human body image meet the standardization requirements, and then a standardized image is output.

[0059] Secondly, the standardized image is Gaussian filtered and multi-scale smoothing is performed using different Gaussian kernel sizes (such as 3×3, 5×5 or 7×7) to effectively suppress noise while retaining the main structural information of the image and output a low-noise image.

[0060] Furthermore, the low-noise image is subjected to image enhancement, sharpening, and edge detection. Sharpening enhances image details through the Laplacian operator or unsharp masking, and edge detection uses the Canny algorithm, Sobel operator, or Laplacian operator to extract human contours and key features, and outputs a feature-enhanced image.

[0061] Finally, a cutout algorithm is used to separate the human body area from the background area in the feature-enhanced image. The cutout algorithm can adopt a threshold segmentation based on color space (such as HSV or Lab space), an instance segmentation model based on deep learning (such as MaskR-CNN), or the GrabCut algorithm based on edge detection to ensure accurate segmentation of the human body area and output the processed image.

[0062] The color space-based threshold segmentation model uses the characteristics of the color space to distinguish the foreground (human body) from the background by setting a threshold range. HSV (hue, saturation, brightness) and Lab (brightness, a channel, b channel) color spaces are more suitable for segmentation than RGB because they are more robust to lighting changes. The specific steps include: converting the image from RGB to HSV or Lab color space. Setting a threshold based on the typical range of human body color (such as the skin color range in HSV). Binarizing each channel to generate a mask. Merging the masks of multiple channels to obtain the final segmentation result. This method is suitable for situations where the lighting is stable and the background color is significantly different from the human body. It has high computational efficiency, but may be affected by complex backgrounds or lighting changes.

[0063] The deep learning-based instance segmentation model combines object detection and semantic segmentation to simultaneously detect objects and generate pixel-level segmentation masks. The steps include: The input image is passed through a backbone network (such as ResNet) to extract features. The Region Proposal Network (RPN) generates candidate object regions. Each candidate region is then classified, subjected to bounding box regression, and instance segmentation (generating a mask). Ultimately, the model outputs an accurate human segmentation result.

[0064] The GrabCut algorithm, based on edge detection, combines graph cut optimization with energy minimization to rapidly segment foreground and background using a user-provided initial rectangle or simple marker. The specific steps include: The user specifies an initial rectangle (enclosing the human body) or a simple marker. A graph model is constructed, in which pixels are classified as foreground, background, possible foreground, and possible background. A Gaussian mixture model (GMM) is used to estimate the color distribution of the foreground and background. The segmentation results are iteratively adjusted through graph cut optimization until convergence.

[0065] It can automatically distinguish the complexity of the scene and automatically associate relevant algorithms to perform image preprocessing according to needs.

[0066] The region recognition module is used to establish a human body recognition model based on the historical human body database and, based on the human body recognition model, identify human body features in the processed image. The region recognition module includes a modeling unit and a region recognition unit. The modeling unit is used to establish a human body recognition model based on the historical human body database, and the region recognition unit is used to identify human body features in the processed image based on the human body recognition model. The steps of establishing the human body recognition model by the modeling unit include:

[0067] First, the historical human images in the historical human body database are preprocessed to output standardized historical images. The preprocessing steps include image denoising, grayscale processing, size normalization, and contrast enhancement to ensure the quality and consistency of the input images, laying the foundation for subsequent feature extraction.

[0068] Secondly, feature extraction is performed on the standardized historical images to output human image features. Feature extraction uses a multi-scale convolutional neural network (CNN) or a pre-trained deep learning model (such as ResNet, VGG, etc.) to capture global and local features of human images, including contour, texture, and semantic information.

[0069] Furthermore, the human image features were divided into a dataset consisting of training, validation, and test sets. This dataset division followed an 8:1:1 or 7:2:1 ratio to ensure the independence of model training, validation, and testing, and to prevent data leakage. Furthermore, data augmentation techniques (such as rotation, flipping, and cropping) were employed on small datasets to improve the model's generalization capabilities.

[0070] Furthermore, an initial human recognition model is established by fusing a segmentation model with a keypoint detection model. The segmentation model uses an architecture such as U-Net or Mask-R-CNN to locate human body regions. A keypoint detection model (such as OpenPose or HRNet) is used to extract the coordinates of human joints. These two models are combined through a feature fusion module (such as an attention mechanism or weighted fusion) to generate an initial human recognition model.

[0071] Furthermore, the classification cross entropy function is used as the loss function to initialize the human recognition model. The loss function also includes auxiliary loss terms (such as L2 loss for key point detection or Dice loss for segmentation) to optimize model performance in a multi-task learning manner.

[0072] Finally, the human recognition model is trained using the training set, and the loss function is verified using the validation set. When the loss stabilizes, training is stopped, resulting in the human recognition model. The training process employs an adaptive optimization algorithm and a dynamic learning rate scheduling strategy to accelerate convergence and improve model robustness. An early stopping mechanism is also introduced during training to prevent overfitting.

[0073] The massage terminal is used to perform massage operations based on the desired massage area and human body characteristics, monitor skin pressure at the massage area in real time, and output massage terminal data. The massage terminal includes a control terminal, a robotic arm, and a massage claw. The control terminal is used to output control instructions for the robotic arm's position based on the desired massage area and human body characteristics. The robotic arm is used to adjust its position according to the control instructions. The massage claw is used to perform massage operations based on its own position and monitor skin pressure at the massage area in real time. The steps for the control terminal to output control instructions for the robotic arm's position based on human body characteristics include:

[0074] First, the system extracts the desired massage area from the case database, matches the body's characteristics with the desired massage area, and outputs the matching results. During this process, the system dynamically adjusts the massage plan based on the user's historical health records and real-time physiological data (such as heart rate and muscle tension) to ensure the safety and personalization of the massage plan.

[0075] Secondly, based on the matching results and human characteristics, the system uses a hybrid ant colony algorithm and combines human environmental information to plan the path of the robotic arm to obtain the optimal path strategy. During the path planning process, the system comprehensively considers the human anatomy, the range of motion of the massage area, the kinematic limitations of the robotic arm, and the surrounding environment (such as obstacles, changes in human posture, etc.) to ensure that the massage claws can reach the target position accurately and safely. In addition, the system will adjust the path in real time to cope with slight movements or changes in posture of the human body to ensure the stability and comfort of the massage process. The steps for the control terminal to plan the path of the robotic arm based on the hybrid ant colony algorithm include:

[0076] Discretize the workspace in the human environment information into a grid or construct a free space map, mark obstacles and feasible areas, and define the initial position and target point of the robot arm.

[0077] Initialize algorithm parameters, including population size.

[0078] Each individual in the population starts from the initial position and selects the next node based on the current pheromone concentration and heuristic information. The formula is expressed as:

[0079]

[0080] Where, represents the probability of individual k moving from the current position i to the next position j, τ ij Indicates the pheromone concentration left by the individual on the path (i, j), t is the time window, allowed k represents the target point set of individual k, α is the pheromone influence constant, β is the heuristic function influence constant, is the heuristic information, d ij is the distance between the individual's current position and the next position j, d jwis the distance between the next position j of the individual and the final target point w. Among them, the amount of residual pheromone released by the individual will affect the size of the heuristic information. Therefore, after each individual completes a search path, it is necessary to update the residual information. The pheromone will evaporate over time, and its volatilization rate is the pheromone volatilization factor ρ, which has a value range of (0, 1). When the ρ value is too small, the residual pheromone will also increase, which will cause a large number of useless path searches during the search process, and ultimately reduce the convergence speed of the algorithm. However, when the ρ value is too large, although the residual pheromone is less, the individual's search range is also increased, causing the target search to take longer, thereby affecting the individual's convergence speed. Therefore, the pheromone volatilization factor ρ in the form of a constant will make the adaptability of the algorithm worse. Therefore, the calculation method of the pheromone volatilization factor ρ can be expressed as ρ(T)=ρ(1+log N T), so the improved global pheromone update formula is as follows:

[0081] τ ij (t+n)=(1-ρ(1+log N T))τ ij (t)+Δτ ij (t)

[0082]

[0083] Where ρ is the pheromone volatility factor, Δτ ij (t) represents the pheromone increment on the current path (i, j), L represents the residual information of the kth individual on the path (i, j), Q is the pheromone intensity, which represents the total amount of pheromone released by the individual. k Indicates the total length of the k-th individual to complete a search path, L best Indicates the shortest path taken by the individual in this cycle, L worst represents the longest path taken by an individual in this cycle, λ represents the number of best individuals during the path search cycle, ω represents the number of worst individuals during the path search cycle, G is the pheromone update formula, T is the number of iterations, and N is the maximum number of iterations. Due to the low number of iterations in the early stages, the pheromone level is also relatively low, which increases the algorithm's global search capability and reduces the unnecessary guidance of the ant colony's pheromone. However, as the number of iterations increases, the pheromone evaporates faster, allowing the algorithm to converge quickly in the later stages and more accurately select the optimal path.

[0084] The total cost is calculated for each feasible path, and the pheromone is updated. When the change in the optimal path cost for multiple generations is less than the threshold, the update is stopped and the node sequence is output. To solve the problems of unstable and non-smooth paths during path planning, a quadratic Bezier curve is mixed into the path fitting process to improve the stability of the motion trajectory. The formula is expressed as:

[0085] B(a)=(1-a) 2 P0+2a(1-a)P1

[0086] Where B(a) is the quadratic Bezier curve formula, and P0, P1, and P2 represent the pneumatic, control, and endpoint points of the motion trajectory, respectively. a is the optimization parameter of the quadratic Bezier curve, a∈[0,1]. The range of a is limited to [0,1] to ensure that points on the Bezier curve can be completely and uniquely determined. Within this interval, different values ​​of a correspond to different positions on the curve. The endpoint properties of the quadratic Bezier curve ensure the continuity of the derivatives at both ends of the curve, ensuring that the optimized trajectory always maintains displacement continuity, reducing the instability of the robot's operation and improving the smoothness and service life of the robot's path.

[0087] Connect each node in the node sequence according to the time window order to obtain the optimal path strategy.

[0088] Finally, the optimal path strategy is converted into a control signal for the robotic arm, and the robotic arm control instructions are output. The joint trajectories in the optimal path strategy are converted into the Cartesian space trajectory (position and attitude) of the end effector through the kinematic forward solution. In combination with motion control algorithms (such as PID control, feedforward compensation, or adaptive control), real-time control signals for each joint are generated. The control instructions must be calibrated and filtered according to the hardware characteristics of the robotic arm (including the response characteristics of the servo motor and the accuracy of the encoder) to ensure execution accuracy and stability. The control instructions are sent to the robotic arm's driver or controller via a communication interface (such as the CAN bus), completing the closed-loop execution of the path planning.

[0089] The pain recognition module is used to perform human pain recognition based on multimodal feature analysis, human body characteristics and massage terminal data, and output pain recognition results. The steps of the pain recognition module for human pain recognition include:

[0090] First, multimodal human body data is collected through multiple sensors, including galvanic skin response (GSR) sensors, high-precision microphones, and high-frame-rate cameras, to ensure comprehensive and real-time data collection.

[0091] Secondly, feature extraction is performed on multimodal data, including extraction of human skin electrical response signals (such as skin conductance level, skin conductance response), speech features (such as fundamental frequency, Mel-frequency cepstral coefficients MFCC, emotional speech features) and behavioral features (such as facial expressions, body movements, posture changes), and time series analysis methods are used to enhance feature stability.

[0092] Finally, a deep learning model is used to fuse the skin electrical response signals, speech features, and behavioral features, and the feature weights are optimized with the attention mechanism to output the recognition results.

[0093] The force adjustment module is used to adjust skin pressure based on the recognition results. It usually adopts PID control algorithm or reinforcement learning strategy to achieve accurate pressure feedback to ensure comfort and safety.

[0094] like Figure 2 As shown, the present invention also provides a 3D vision-based multimodal fusion treatment method for a massage robot, comprising:

[0095] Collect human body images, pre-process the human body images, and output the processed images. Pre-process the human body images and output the processed images, the steps include:

[0096] First, the human body image is subjected to geometric correction and image correction, including adjusting the perspective deformation, rotation angle and scale of the image to ensure that the posture and size of the human body image meet the standardization requirements, and then a standardized image is output.

[0097] Secondly, the standardized image is Gaussian filtered and multi-scale smoothing is performed using different Gaussian kernel sizes (such as 3×3, 5×5 or 7×7) to effectively suppress noise while retaining the main structural information of the image and output a low-noise image.

[0098] Furthermore, the low-noise image is subjected to image enhancement, sharpening, and edge detection. Sharpening enhances image details through the Laplacian operator or unsharp masking, and edge detection uses the Canny algorithm, Sobel operator, or Laplacian operator to extract human contours and key features, and outputs a feature-enhanced image.

[0099] Finally, the cutout algorithm is used to separate the human body area from the background area in the feature enhanced image, and the processed image is output.

[0100] A human body recognition model is established based on the historical human body database, and human body features in the processed image are identified based on the human body recognition model. The steps of the modeling unit to establish the human body recognition model include:

[0101] First, the historical human images in the historical human body database are preprocessed to output standardized historical images. The preprocessing steps include image denoising, grayscale processing, size normalization, and contrast enhancement to ensure the quality and consistency of the input images, laying the foundation for subsequent feature extraction.

[0102] Secondly, feature extraction is performed on the standardized historical images to output human body image features.

[0103] Furthermore, the human body image features are divided into data sets, including training sets, validation sets and test sets.

[0104] Furthermore, an initialized human recognition model based on the fusion of the segmentation model and the key point detection model is established.

[0105] Furthermore, the classification cross entropy function is used as the loss function to initialize the human recognition model. The loss function also includes auxiliary loss terms (such as L2 loss for key point detection or Dice loss for segmentation) to optimize model performance in a multi-task learning manner.

[0106] Finally, the training set is used to train the human recognition model, and the validation set is used to verify the loss of the loss function. When the loss tends to be stable, the training is stopped to obtain the human recognition model.

[0107] The massage operation is performed according to the required massage parts and human body characteristics, and the skin pressure of the massage parts is monitored in real time, and the massage terminal data is output.

[0108] Based on multimodal feature analysis, human pain recognition is performed according to human body characteristics and massage terminal data, and the pain recognition results are output. The steps for human pain recognition include:

[0109] First, multimodal human body data is collected through multiple sensors, including galvanic skin response (GSR) sensors, high-precision microphones, and high-frame-rate cameras, to ensure comprehensive and real-time data collection.

[0110] Secondly, feature extraction is performed on multimodal data, including extraction of human skin electrical response signals (such as skin conductance level, skin conductance response), speech features (such as fundamental frequency, Mel-frequency cepstral coefficients MFCC, emotional speech features) and behavioral features (such as facial expressions, body movements, posture changes), and time series analysis methods are used to enhance feature stability.

[0111] Finally, a deep learning model is used to fuse skin electrical response signals, speech features, and behavioral features, and the feature weights are optimized with the attention mechanism to output pain recognition results.

[0112] Based on the pain recognition results, skin pressure is adjusted. PID control algorithms or reinforcement learning strategies are usually used to achieve accurate pressure feedback to ensure comfort and safety.

[0113] In summary, the present invention provides a 3D vision-based massage robot multimodal fusion treatment system and method. By adopting a human body recognition model initialized by a multimodal image preprocessing and fusion segmentation model and a key point detection model, the system can accurately identify human body features and locate massage parts, significantly improving the pertinence and safety of massage operations. The control terminal combines human body features with environmental information, and adopts a hybrid ant colony algorithm to generate and optimize the global path strategy to achieve efficient and precise motion control of the robotic arm. Through multimodal data fusion analysis, it can perceive the user's pain in real time and dynamically adjust the massage intensity to avoid excessive stimulation, significantly improving comfort and safety factors. The system integrates multi-source information such as images, pressure, and physiological signals, and combines deep learning models for feature analysis, effectively improving the accuracy of pain recognition and human feature recognition, and adapting to individual differences among different users.

[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A 3D vision-based massage robot multimodal fusion treatment system, characterized in that: include: An image acquisition module is used to acquire human body images, pre-process the human body images, and output the processed images; A region recognition module, configured to establish a human body recognition model based on a historical human body database, and to recognize human body features in the processed image based on the human body recognition model; A massage terminal, configured to perform massage operations according to the desired massage area and the characteristics of the human body, monitor the skin pressure of the massage area in real time, and output massage terminal data; a pain recognition module, configured to perform human pain recognition based on multimodal feature analysis, according to the human body features and the massage terminal data, and output a pain recognition result; A force adjustment module is used to adjust the skin pressure according to the recognition result.

2. The 3D vision-based massage robot multimodal fusion treatment system according to claim 1 is characterized in that: The image acquisition module includes an image acquisition unit and an image preprocessing unit; the image acquisition unit is used to acquire the human body image, and the image preprocessing unit is used to preprocess the human body image and output the processed image.

3. The 3D vision-based massage robot multimodal fusion treatment system according to claim 2 is characterized in that: The step of preprocessing the human body image by the image preprocessing unit includes: Performing geometric correction and image correction on the human body image to output a standardized image; Performing Gaussian filtering on the standardized image to output a low-noise image; Performing image enhancement on the low-noise image, including sharpening and edge detection, and outputting a feature-enhanced image; A cutout algorithm is used to separate the human body region from the background region in the feature-enhanced image, and a processed image is output.

4. The 3D vision-based massage robot multimodal fusion treatment system according to claim 1 is characterized in that: The region recognition module includes a modeling unit and a region recognition unit; the modeling unit is used to establish a human body recognition model based on a historical human body database, and the region recognition unit is used to recognize human body features in the processed image based on the human body recognition model.

5. The 3D vision-based multimodal fusion therapy system for massage robots according to claim 4 is characterized in that: The step of the modeling unit establishing the human body recognition model comprises: Preprocessing the historical human images in the historical human database to output standardized historical images; Performing feature extraction on the standardized historical image and outputting human body image features; Dividing the human body image features into a data set including a training set and a validation set; Establish an initialized human recognition model based on the fusion of segmentation model and key point detection model; Using the classification cross entropy function as the loss function of the initialized human recognition model; The human body recognition model is trained using the training set, and the loss of the loss function is verified using the validation set. When the loss tends to be stable, the training is stopped to obtain the human body recognition model.

6. The 3D vision-based massage robot multimodal fusion treatment system according to claim 1 is characterized in that: The massage terminal includes a control terminal, a robotic arm and a massage claw; the control terminal is used to output a control instruction for the position of the robotic arm based on the desired massage area and the human body characteristics; the robotic arm is used to adjust its own position according to the control instruction; the massage claw is used to perform massage operations based on its own position and monitor the skin pressure of the massage area in real time.

7. The 3D vision-based massage robot multimodal fusion treatment system according to claim 6 is characterized in that: The step of the control terminal outputting a control instruction for the position of the robotic arm according to the human body characteristics comprises: Extracting the desired massage part from the case database, matching the human body features with the desired massage part, and outputting a matching result; According to the matching results and the human body features, based on the hybrid ant colony algorithm and combined with the human body environment information, the path of the robot arm position is planned to obtain the optimal path strategy; The optimal path strategy is converted into a control signal for the robotic arm, and a robotic arm control instruction is output.

8. The 3D vision-based multimodal fusion therapy system for massage robots according to claim 7 is characterized in that: The step of the control terminal performing path planning for the position of the robotic arm based on the hybrid ant colony algorithm includes: Discretize the workspace in the human body environment information into a grid or construct a free space map, mark obstacles and feasible areas; define the initial position and target point of the robot arm; Initialize algorithm parameters, including population size; Each individual in the population starts from the initial position and selects the next node based on the current pheromone concentration and heuristic information. The formula is expressed as: Where, represents the probability of individual k moving from the current position i to the next position j, τ ij Indicates the pheromone concentration left by the individual on the path (i, j), t is the time window, allowed k represents the target point set of individual k, α is the pheromone influence constant, β is the heuristic function influence constant, is the heuristic information, d ij is the distance between the individual's current position and the next position j, d jw is the distance between the next position j of the individual and the final target point w; Calculate the total cost for each feasible path and update the pheromone; when the change in the optimal path cost for multiple generations is less than the threshold, stop updating and output the node sequence; Each node in the node sequence is connected in a time window order to obtain the optimal path strategy.

9. The 3D vision-based massage robot multimodal fusion treatment system according to claim 1 is characterized in that: The steps of the pain recognition module performing human pain recognition include: Collect multimodal data of the human body through multiple sensors; Performing feature extraction on the multimodal data, including extracting human skin electrical response signals, speech features, and behavioral features; A deep learning model is used to perform feature fusion on the skin electrical response signal, speech features and behavioral features, and output the pain recognition result.

10. A 3D vision-based massage robot multimodal fusion treatment method, based on the 3D vision-based massage robot multimodal fusion treatment system according to any one of claims 1 to 9, characterized in that: include: Collecting human body images, preprocessing the human body images, and outputting processed images; Establishing a human body recognition model based on a historical human body database, and identifying human body features in the processed image based on the human body recognition model; Performing massage operations according to the desired massage area and the human body characteristics, and monitoring the skin pressure of the massage area in real time; Based on multimodal feature analysis, human pain recognition is performed and pain recognition results are output; The skin pressure is adjusted according to the recognition result.

Citation Information

Cited By

  • Magnetic therapy back massage instrument based on vision assistance

    CN121129620A

  • Visual aid-based magnetic therapy back massager

    CN121129620B

  • Massage robot, control method and control device thereof and storage medium

    CN121132704A

  • Self-adaptive massage equipment based on multi-mode perception

    CN121401107A

  • Microwave irradiation area matching method and system

    CN121489428A