Massage robot control method based on multi-modal feedback
By constructing a mapping model through multimodal physiological data acquisition and deep learning algorithms, and combining it with reinforcement learning to optimize the control strategy, the problems of personalization and closed-loop feedback in the control of massage robots in the existing technology have been solved, and precise adjustment of massage parameters and stable improvement of comfort have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YUEYANG INTEGRATED TRADITIONAL CHINESE & WESTERN MEDICINE HOSPITAL SHANGHAI UNIV OF CHINESE TRADITIONAL MEDICINE
- Filing Date
- 2025-11-28
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies lack the ability to accurately map multi-dimensional physiological characteristics to individualized treatment parameters in the control of massage robots, making it impossible to achieve precise control and personalized adjustment. Furthermore, the lack of a closed-loop feedback mechanism leads to unstable therapeutic effects and comfort.
By collecting multimodal physiological data (EEG, heart rate, EMG) and combining it with deep learning algorithms to construct a mapping model, massage parameters are adjusted in real time, and the control strategy is optimized using reinforcement learning adaptive algorithms to achieve individualized closed-loop feedback control.
It achieves precise mapping from multi-dimensional physiological characteristics to massage parameters, improves the stability of massage efficacy and comfort, reduces reliance on patient subjective feedback, and enhances the personalization and reliability of treatment.
Smart Images

Figure CN121928535A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent massage technology, specifically relating to a control method for a massage robot based on multimodal feedback. Background Technology
[0002] Tuina (Chinese massage) is a traditional Chinese medicine treatment method that uses techniques such as pressing, kneading, pushing, and grasping to act on specific parts or acupoints of the body to regulate bodily functions, relieve pain, and promote blood circulation. In modern clinical practice, the effectiveness of tuina highly depends on the experience of the practitioner and the patient's subjective feelings, resulting in problems such as significant individual differences, difficulty in standardization, and difficulty in quantifying efficacy.
[0003] With the development of robotics technology, massage robots are increasingly being applied in the field of rehabilitation medicine. In existing technology, CN116196196A discloses a massage chair control method that utilizes multiple sensors to collect human physiological data, generates a user emotion recognition model through convolutional neural networks and deep learning algorithms, and controls the massage intensity based on the user's emotional state to achieve feedback-based massage optimization. This method includes steps such as establishing a standard database, real-time collection of physiological data, emotion recognition, and adjusting massage parameters based on user feedback.
[0004] However, this method has the following limitations: 1. This method classifies emotions by comparing real-time physiological data with reference data in a standard database, and can only determine whether the user is relaxed or tense. It lacks the ability to accurately map multi-dimensional physiological characteristics to individualized treatment parameters. 2. Existing technologies can only correlate emotional states with massage intensity, and have not established a direct mapping relationship between multimodal physiological signals and multi-dimensional massage control parameters, making it difficult to achieve precise control and personalized adjustment. 3. Existing methods adjust massage intensity after determining the state through emotion recognition, but the feedback adjustment granularity is limited. It does not form a continuous closed-loop control based on real-time physiological signals and expected comfort, and cannot dynamically optimize the comfort and therapeutic effect of the entire massage process.
[0005] Therefore, there is an urgent need for an intelligent massage robot control method that can adjust massage parameters in real time based on multimodal physiological data, integrate subjective feelings and objective indicators, and has the ability to learn individually and provide closed-loop feedback, so as to improve the therapeutic effect, comfort and standardization of massage. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a control method for a massage robot based on multimodal feedback.
[0007] The objective of this invention can be achieved through the following technical solutions: This invention provides a control method for a massage robot based on multimodal feedback, comprising the following steps: The physiological data generated by multiple patients during the massage was collected using a multimodal acquisition device, and the corresponding Qi-de-point value and massage control parameters for each patient during each treatment were recorded. Based on the physiological data, Qi-de-score, and massage control parameters of multiple patients, a deep learning algorithm was used to train the mapping relationship between physiological data and Qi-de-score, as well as the mapping relationship between physiological data and massage control parameters, to obtain the first mapping model and the second mapping model respectively. Before the massage begins, the patient’s desired Qi score is obtained, and based on the desired Qi score, the ideal physiological data corresponding to the desired Qi score is derived through the first mapping model. During the massage, the patient's physiological data is collected in real time, and the current Qi attainment score is inferred using the first mapping model. The score is then compared with the expected Qi attainment score to calculate the difference between the two scores. Based on the differences between the Qi-de scores, the massage control parameters corresponding to the ideal physiological data are inferred through the second mapping model, and the massage control parameters are dynamically adjusted using a feedback mechanism. Based on the dynamically adjusted massage control parameters, the massage operation continues, and the massage control parameters are continuously optimized based on real-time feedback. During the long-term treatment process, historical treatment data and feedback information from each patient are continuously collected, and reinforcement learning adaptive algorithms are used to optimize the second mapping model.
[0008] Furthermore, the multimodal acquisition device includes: EEG acquisition equipment is used to monitor a patient's brain electrical activity and acquire brain wave signals related to emotions and relaxation levels; Heart rate sensor, used to monitor a patient's heart rate signal in real time; Electromyography (EMG) sensors are used to collect electrical signals of muscle activity from patients. The physiological data includes brain wave signals, heart rate signals, and electrical signals of muscle activity. The "Deqi" score refers to the patient's subjective rating of the comfort level during the massage.
[0009] Furthermore, the massage control parameters include: Massage intensity, controlling the amount of force applied to the patient's body by the robotic arm; Massage frequency refers to the frequency of massage movements; Massage depth refers to the depth of penetration of the robotic arm during the massage process; Massage speed, or the speed at which the robotic arm moves, controls the rhythm of the massage movements.
[0010] Furthermore, a deep learning algorithm is used to train the mapping relationship between physiological data and Qi-de-scores to obtain the first mapping model, which specifically includes: The collected physiological data are preprocessed, and a physiological data feature vector is constructed based on the preprocessed physiological data, represented as follows: in, Indicates the first i Physiological data feature vectors of each sample; , , These represent the first and second parts after preprocessing. i EEG characteristic data, heart rate characteristic data and electromyography characteristic data of each sample; Using the physiological data feature vector as input and the corresponding Qi score of the physiological data feature vector as [something], training sample pairs are constructed, represented as follows: in, For the training dataset; Indicates the first i The gas yield score corresponding to each sample; This represents the number of training samples; A deep neural network model is constructed by inputting the physiological data feature vector into the deep neural network, obtaining the predicted result of the Qi score through forward propagation, and training the deep neural network based on the training samples. A supervised learning method is used to train a deep neural network. A loss function is constructed based on the difference between the predicted and actual Qi scores. The network parameters are iteratively updated using a gradient descent optimization algorithm, which makes the loss function gradually converge. Training is completed when the loss function meets the preset convergence condition or reaches the preset number of training rounds, and the trained deep neural network is used as the first mapping model.
[0011] Furthermore, the deep neural network includes an input layer, multiple hidden layers, and an output layer. The input layer is used to receive physiological data feature vectors, and the output layer is used to output the predicted result of the Qi score.
[0012] Furthermore, the construction process of the second mapping model includes: The collected physiological data are preprocessed, and physiological data feature vectors are constructed based on the preprocessed physiological data. ,in, Indicates the first i Physiological data feature vectors of each sample; physiological data feature vectors As input, the corresponding massage control parameter vector As supervisory labels, training sample pairs are constructed. ; Construct a deep neural network, input training sample pairs into the deep neural network, train the deep neural network, and the output of the deep neural network is the prediction result of the massage control parameters; During training, a loss function is constructed based on the difference between the predicted massage control parameters and the actual massage control parameters, and an optimization algorithm based on gradient descent is used to update the parameters of the deep neural network. When the loss function meets the preset convergence condition or the number of training rounds reaches the preset threshold, training stops, and the trained deep neural network is used as the second mapping model.
[0013] Furthermore, the preprocessing of the collected physiological data specifically includes: The EEG signal was bandpass filtered to remove power frequency interference and high frequency noise. The filtered EEG signal was then segmented and its features were extracted to obtain EEG feature data. The heart rate signal is smoothed and baseline corrected, and the heart rate variability correlation characteristics are calculated to obtain heart rate characteristic data; The electromyographic signals are rectified and integrated to normalize the amplitude of muscle activity and obtain electromyographic characteristic data. The processed EEG, heart rate, and electromyography (EMG) feature data are combined in a preset order to form a physiological data feature vector. .
[0014] Furthermore, based on the differences in the Qi (vital energy) scores, the massage control parameters corresponding to the ideal physiological data are inferred through the second mapping model, and the massage control parameters are dynamically adjusted using a feedback mechanism, specifically including: During the massage, the patient's physiological data is collected in real time, and the current moment is obtained. t Physiological data feature vector ; Current moment t Physiological data feature vector Input the first mapping model Predict the current Qi score : in, This is the first mapping model; For the current moment t The predicted Qi score; According to the current time t The error in the predicted Qi attainment score is calculated by comparing it with the patient's expected Qi attainment score. in, The expected score for obtaining Qi set for the patient; Indicates the current time t The error in the Qi-gathering score; Using the second mapping model For the current moment t Physiological data feature vector With the feature vector of ideal physiological data By performing massage control parameter reasoning, the control parameter increments are obtained: in, Indicates the increment of the control parameter; The feature vector representing ideal physiological data is obtained based on the patient's desired Qi attainment score and the first mapping model. Based on the error between the control parameter increment and the Qi (vital energy) value, the dynamically adjusted massage control parameters are obtained: in, This is the feedback gain coefficient; For the current moment t The dynamically adjusted massage control parameters; Will The data is sent to the massage robotic arm, which then performs the actions according to the massage force, frequency, depth, and speed.
[0015] Furthermore, during the long-term treatment process, by continuously collecting historical treatment data and feedback information from each patient, the second mapping model is optimized using a reinforcement learning adaptive algorithm, specifically including: An individual treatment history dataset is created for each patient, represented as follows: in, Indicates the first Historical treatment dataset of 100 patients Indicates the first i Physiological data feature vectors of each sample This indicates the corresponding massage control parameters. This indicates the patient's Qi attainment score for that sample; Indicates the first Number of historical samples from patients; A state-action-reward model is constructed for each patient based on reinforcement learning, where: state For the first t The physiological data feature vector of the patient at any given time; action The vector of adjustment for control parameters; Using individual patients' historical data Personalized strategy network for each patient Training is performed, and the policy network input state is used. Output action And based on rewards Update network parameters : in, For policy network parameters; The learning rate; This represents the gradient operation applied to the policy network parameters; For the reward function; During real-time massage, the trained policy network is used to generate adjustment movements. And combined with the massage control parameters output by the second mapping model The final massage control parameters are obtained as follows: Status updated in real time during each treatment cycle Deqi score and final massage control parameters This forms an individualized closed-loop adaptive adjustment, enabling continuous optimization of the second mapping model.
[0016] Furthermore, the reward function is expressed as: in, The current air quality score is predicted using the first mapping model; The patient's desired Qi score; , Preset weighting coefficients; It is the Euclidean norm.
[0017] Compared with the prior art, the present invention has the following advantages: (1) Existing technologies rely on standard databases and single emotion recognition models, which can only determine whether the user is relaxed or tense, and lack the ability to accurately map multi-dimensional physiological characteristics to individualized treatment parameters. To solve this problem, this invention collects multimodal physiological data (EEG, heart rate, EMG, etc.) and combines it with the patient's Qi-de (de-qi) score, and uses deep learning algorithms to construct a first mapping model to achieve accurate mapping between physiological data and the patient's subjective comfort. This allows for dynamic assessment of the patient's comfort based on real-time physiological signals, achieving higher precision massage adjustment than single emotion classification, and improving the level of personalized and refined treatment.
[0018] (2) Existing technologies lack a closed-loop real-time adjustment mechanism and can only adjust the massage intensity simply after the emotional state is determined, with limited feedback adjustment granularity. This invention collects the patient's physiological data in real time, uses the first mapping model to predict the current Qi score, and then calculates the error by combining the expected Qi score. The second mapping model is used to dynamically adjust the massage control parameters to form a closed-loop adaptive control, enabling the massage robot to continuously optimize the intensity, frequency, depth, and speed throughout the treatment process. This achieves real-time adjustment and dynamic optimization of the massage movements, improving patient comfort and therapeutic stability.
[0019] (3) Existing technologies lack individualized optimization capabilities, fail to establish individualized models for different patients, and do not utilize historical treatment data for adaptive learning. This invention establishes an individual historical treatment dataset for each patient and trains an individualized policy network using reinforcement learning algorithms. This enables dynamic adjustment of massage control parameters based on the patient's historical data and real-time feedback, allowing each patient to obtain a personalized closed-loop adaptive massage plan, thereby optimizing long-term treatment effects and continuously improving patient comfort.
[0020] (4) Existing technologies rely heavily on subjective patient feedback data and cannot effectively adjust when patients cannot provide continuous feedback. This invention utilizes the mapping relationship between multimodal physiological data and Qi-de-score, as well as the mapping relationship between multimodal physiological data and massage control parameters. It can infer the ideal physiological state and obtain the corresponding control parameters based on real-time physiological data without the patient providing real-time scores. This enables automated adjustment and continuous optimization of massage operations, thereby reducing reliance on subjective patient feedback and improving the reliability and practicality of massage robot control. Attached Figure Description
[0021] Figure 1 This is a flowchart of the massage robot control method according to an embodiment of the present invention; Figure 2 This is a model diagram of the massage robot control system according to an embodiment of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] Example 1: This embodiment specifically provides a control method for a massage robot based on multimodal feedback, such as... Figure 1 As shown, it includes the following steps: Step S1: Collect physiological data generated by multiple patients during massage using a multimodal acquisition device, and record the corresponding Qi-de-point value and massage control parameters for each patient during each treatment. The multimodal acquisition device includes: EEG acquisition equipment is used to monitor a patient's brain electrical activity and acquire brain wave signals related to emotions and relaxation levels; Heart rate sensor, used to monitor a patient's heart rate signal in real time; Electromyography (EMG) sensors are used to collect electrical signals of muscle activity from patients. Physiological data include brain wave signals, heart rate signals, and electrical signals of muscle activity. The "Deqi" score refers to the patient's subjective rating of the comfort level during the massage.
[0024] Massage control parameters include: Massage intensity, controlling the amount of force applied to the patient's body by the robotic arm; Massage frequency refers to the frequency of massage movements; Massage depth refers to the depth of penetration of the robotic arm during the massage process; Massage speed, or the speed at which the robotic arm moves, controls the rhythm of the massage movements.
[0025] Step S2: Based on the physiological data, Qi-de-score, and massage control parameters of multiple patients, a deep learning algorithm is used to train the mapping relationship between physiological data and Qi-de-score, as well as the mapping relationship between physiological data and massage control parameters, to obtain the first mapping model and the second mapping model respectively. The mapping relationship between physiological data and Qi-de-qi scores is trained using deep learning algorithms to obtain the first mapping model, which specifically includes: The collected physiological data are preprocessed, and a physiological data feature vector is constructed based on the preprocessed physiological data, represented as follows: in, Indicates the first i Physiological data feature vectors of each sample; , , These represent the first and second parts after preprocessing. i EEG characteristic data, heart rate characteristic data and electromyography characteristic data of each sample; Using the physiological data feature vector as input and the corresponding Qi score of the physiological data feature vector as [something], training sample pairs are constructed, represented as follows: in, For the training dataset; Indicates the first i The gas yield score corresponding to each sample; This represents the number of training samples; A deep neural network model is constructed. Physiological data feature vectors are input into the deep neural network, and the predicted result of the Qi-de-qi score is obtained through forward propagation. The deep neural network is trained based on training samples. The deep neural network is a multi-layer feedforward neural network, including an input layer, multiple hidden layers and an output layer. The input layer is used to receive physiological data feature vectors, and the output layer is used to output the predicted result of the Qi-de-qi score.
[0026] A supervised learning method is used to train a deep neural network. A loss function is constructed based on the difference between the predicted and actual Qi scores. The network parameters are iteratively updated using a gradient descent optimization algorithm, which makes the loss function gradually converge. Training is completed when the loss function meets the preset convergence condition or reaches the preset number of training rounds, and the trained deep neural network is used as the first mapping model.
[0027] The second mapping model is also a deep neural network model. Its training process is similar to that of the first mapping model, specifically including the construction of an input layer, multiple hidden layers, and an output layer. The input layer receives physiological data feature vectors, but the difference is that the output layer of the second mapping model is used to output the predicted results of massage control parameters, rather than the predicted results of the Qi-de (de-qi) score. The training sample pairs are physiological data feature vectors and corresponding massage control parameter vectors. The loss function is constructed based on the difference between the predicted massage control parameters and the actual massage control parameters. The optimization objective is to minimize the massage parameter prediction error, thereby achieving an accurate mapping from physiological data to massage parameters, which can then be used for the dynamic control of massage robots.
[0028] During massage stimulation, the subjective experience of "deqi" (the sensation of qi) is usually accompanied by observable changes in physiological characteristics, such as changes in electroencephalogram (EEG) signals. Synchronous changes in rhythm, regulation of heart rate and heart rate variability, and the degree of relaxation or tension in electromyographic activity. These physiological changes reflect the response of the autonomic nervous system, muscle tone, and central nervous system processing to external stimuli, and show a consistent trend with the subjective perception of massage pressure, depth, and comfort. Therefore, although the Deqi score is a subjective evaluation, the underlying physiological response has quantifiable and statistically learnable objective laws. This invention utilizes this intrinsic correlation between subjective experience and physiological response, automatically extracting stable features from a large amount of multimodal physiological data through deep learning methods. This enables the model to capture the implicit mapping relationship between subjective Deqi sensation and physiological signals, thereby achieving effective prediction of the Deqi score.
[0029] In step S2, deep learning training is first performed on the physiological data, Qi-de scores, and corresponding massage control parameters collected from multiple patients during massage. A first mapping model and a second mapping model are constructed, which is a crucial foundation for the personalized, real-time adjustment of massage in this invention. By training the mapping relationship between physiological data and Qi-de scores, the obtained first mapping model can accurately predict the patient's current physiological state as their subjective comfort experience (Qi-de score), solving the problem in existing technologies of being unable to quantify the patient's comfort experience in real time and objectively. A physiological data feature vector is constructed, integrating multimodal signals such as EEG, heart rate, and EMG into a unified input. Forward propagation and supervised learning training are performed through a multi-layer feedforward neural network, enabling the model to learn complex nonlinear relationships and achieve high-precision mapping from physiological data to subjective Qi-de scores, thus providing a reliable basis for subsequent feedback adjustments.
[0030] Step S3: Before the massage begins, obtain the patient's desired Qi attainment score, and based on this desired score, derive the ideal physiological data corresponding to the desired Qi attainment score through the first mapping model, specifically including: Obtain the patient's desired Qi De score, which represents the level of comfort or Qi De sensation the patient hopes to achieve during the massage.
[0031] The desired Qi score is input into the first mapping model, and the model is used to predict the ideal physiological data feature vector corresponding to the Qi score, including multimodal physiological data such as EEG features, heart rate features, and electromyography features.
[0032] Step S4: During the massage, the patient's physiological data is collected in real time, and the current Qi De score is inferred using the first mapping model. The score is then compared with the expected Qi De score to calculate the difference between the two scores. Step S5: Based on the differences in the Qi attainment scores, the massage control parameters corresponding to the ideal physiological data are inferred through the second mapping model. The massage control parameters are then dynamically adjusted using a feedback mechanism, specifically including: During the massage, the patient's physiological data is collected in real time, and the current moment is obtained. t Physiological data feature vector ; Current moment t Physiological data feature vector Input the first mapping model Predict the current Qi score : in, This is the first mapping model; For the current moment t The predicted Qi score; According to the current time t The error in the predicted Qi attainment score is calculated by comparing it with the patient's expected Qi attainment score. in, The expected score for obtaining Qi set for the patient; Indicates the current time t The error in the Qi-gathering score; Using the second mapping model For the current moment t Physiological data feature vector With the feature vector of ideal physiological data By performing massage control parameter reasoning, the control parameter increments are obtained: in, Indicates the increment of the control parameter; The feature vector representing ideal physiological data is obtained based on the patient's desired Qi attainment score and the first mapping model. Based on the error between the control parameter increment and the Qi (vital energy) value, the dynamically adjusted massage control parameters are obtained: in, This is the feedback gain coefficient; For the current moment t The dynamically adjusted massage control parameters; Will The data is sent to the massage robotic arm, which then performs the actions according to the massage force, frequency, depth, and speed.
[0033] The main purpose of step S5 is to achieve real-time, personalized closed-loop control during massage to ensure optimal patient comfort while improving the response accuracy and safety of the massage robot. During massage, the patient's physiological state changes over time, and subjective feelings (the "deqi" score) are difficult to quantify continuously. Therefore, this step uses multimodal physiological data as real-time feedback, solving the problem of traditional massage control's inability to dynamically adjust parameters and its reliance on fixed preset schemes. Specifically, firstly, the patient's physiological data is collected in real-time and converted into feature vectors, which are then input into the first mapping model to predict the current "deqi" score. This step maps complex multimodal physiological signals into quantifiable comfort indicators, thereby achieving objective monitoring of the patient's subjective feelings. Subsequently, by comparing the predicted "deqi" score with the patient's desired "deqi" score, the error is calculated. This error quantifies the deviation between the current massage effect and the ideal treatment effect, providing a basis for subsequent adjustments.
[0034] Using a second mapping model, the massage control parameters corresponding to the current physiological characteristics and ideal physiological characteristics are inferred to obtain the incremental adjustment value of the control parameters. By combining the incremental control parameters with the error of the Qi (deqi) score, the dynamically adjusted massage control parameters are calculated and sent to the robotic arm for execution, realizing a closed-loop mapping from real-time physiological state to massage action. This dynamic adjustment mechanism based on feedback error not only solves the problems of lack of real-time personalized adjustment and comfort differences caused by fixed control parameters in existing technologies, but also continuously optimizes the massage force, frequency, depth, and speed according to the patient's actual response, improving the accuracy and safety of the massage effect.
[0035] Step S6: Based on the dynamically adjusted massage control parameters, continue performing the massage operation and continuously optimize the massage control parameters based on real-time feedback, specifically including: During the massage, the patient's real-time physiological data and the predicted results of the Qi attainment score are continuously collected as real-time feedback information.
[0036] The real-time feedback is compared with the expected Qi score to generate a new Qi score error.
[0037] Using the second mapping model and feedback mechanism, the massage control parameters are further adjusted based on the new Qi-de-point error.
[0038] By continuously iterating and updating the massage control parameters, continuous optimization is achieved during the massage process, so that the patient's sensation of Qi is as close as possible to the desired level, while ensuring the safety and comfort of the massage operation.
[0039] Step S7: During long-term treatment, by continuously collecting historical treatment data and feedback information from each patient, the second mapping model is optimized using a reinforcement learning adaptive algorithm, specifically including: An individual treatment history dataset is created for each patient, represented as follows: in, Indicates the first Historical treatment dataset of 100 patients Indicates the first i Physiological data feature vectors of each sample This indicates the corresponding massage control parameters. This indicates the patient's Qi attainment score for that sample; Indicates the first Number of historical samples from patients; A state-action-reward model is constructed for each patient based on reinforcement learning, where: state For the first t The physiological data feature vector of the patient at any given time; action The vector of adjustment for control parameters; The reward function is expressed as: in, The current air quality score is predicted using the first mapping model; The patient's desired Qi score; , Preset weighting coefficients; It is the Euclidean norm.
[0040] Using individual patients' historical data Personalized strategy network for each patient Training is performed, and the policy network input state is used. Output action And based on rewards Update network parameters : in, For policy network parameters; The learning rate; This represents the gradient operation applied to the policy network parameters; For the reward function; During real-time massage, the trained policy network is used to generate adjustment movements. And combined with the massage control parameters output by the second mapping model To obtain the final massage control parameters : Status updated in real time during each treatment cycle Deqi score and final massage control parameters This forms an individualized closed-loop adaptive adjustment, enabling continuous optimization of the second mapping model.
[0041] It should be noted that step S7 differs from the aforementioned real-time massage control process. It is mainly used for individualized model optimization in long-term treatment. Through reinforcement learning mechanism, the second mapping model is continuously adaptively adjusted to improve the accuracy of massage control parameter prediction and patient comfort, thereby achieving continuous optimization of the model in long-term treatment.
[0042] The core design of step S7 lies in achieving individualized adaptive optimization during long-term treatment, solving the problem that traditional massage robot control cannot continuously learn and dynamically optimize for different patients. In long-term treatment, each patient's physiological state and subjective experience of massage differ significantly, and a fixed control strategy can easily lead to inconsistent comfort or unsatisfactory treatment results. Therefore, this step establishes an individualized closed-loop adaptive mechanism through reinforcement learning to achieve continuous optimization of the second mapping model.
[0043] Specifically, firstly, an individual historical treatment dataset is established for each patient, including physiological characteristics, corresponding massage control parameters, and the patient's Deqi score. This dataset reflects the patient's physiological responses and subjective feelings under different treatment conditions, providing personalized samples for strategy learning. Subsequently, a state-action-reward model is constructed for each patient based on reinforcement learning. The state is represented by a feature vector of real-time physiological data, the action is the adjustment amount of the control parameters, and the reward function comprehensively considers the error between the Deqi score and the expected Deqi score, as well as the amplitude of the action. By penalizing cases where the control parameters deviate too much or the Deqi score fails to meet the target, the control strategy is effectively constrained.
[0044] During the training phase, the individualized policy network, based on historical data, takes the current state as input and adjusts its output actions. The network parameters are optimized through a reward function, enabling the policy to gradually learn the optimal control strategy under different physiological states. During real-time massage, the trained policy network is combined with the second mapping model to generate dynamically adjusted final massage control parameters, achieving personalized closed-loop control of the robotic arm's massage movements. By continuously updating the state, Qi (vital energy) score, and control parameters in real time, the system can continuously adapt to the patient's physiological changes and comfort needs, forming an individualized closed-loop adjustment mechanism. This significantly improves the accuracy and comfort of the massage effect and ensures intelligent adaptive optimization during long-term treatment.
[0045] Example 2: This embodiment provides a massage robot control system based on multimodal feedback, such as... Figure 2 As shown, it includes: Multimodal acquisition module: used to collect physiological data of patients in real time during massage, including electroencephalogram (EEG) signals, heart rate signals and electromyogram (EMG) signals, and preprocess the collected data to form physiological data feature vectors.
[0046] The first mapping model module uses a deep learning model to map the preprocessed physiological data feature vectors to the current predicted Qi score, which reflects the patient's current comfort level.
[0047] The second mapping model module uses a deep neural network to map physiological data feature vectors to ideal physiological data feature vectors as massage control parameters, which are used to generate control commands for massage force, frequency, depth and speed.
[0048] Dynamic adjustment module for control parameters: Based on the difference between the Qi-de-qi score predicted by the first mapping model and the patient's desired Qi-de-qi score, the module calculates the Qi-de-qi score error and, in conjunction with the adjustment amount of the massage control parameters output by the second mapping model, dynamically adjusts the massage control parameters in real time to achieve closed-loop feedback control.
[0049] Massage robotic arm execution module: Receives dynamically adjusted control parameters and drives the massage robotic arm to perform massage operations according to the adjusted force, frequency, depth and speed.
[0050] Historical data acquisition and reinforcement learning optimization module: During long-term treatment, historical treatment data and feedback information of each patient are continuously collected. The second mapping model is optimized individually using reinforcement learning adaptive algorithms, so that the system can continuously improve its adaptability to patients and the massage effect.
[0051] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0052] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A control method for a massage robot based on multimodal feedback, characterized in that, Includes the following steps: The physiological data generated by multiple patients during the massage was collected using a multimodal acquisition device, and the corresponding Qi-de-point value and massage control parameters for each patient during each treatment were recorded. Based on the physiological data, Qi-de-score, and massage control parameters of multiple patients, a deep learning algorithm was used to train the mapping relationship between physiological data and Qi-de-score, as well as the mapping relationship between physiological data and massage control parameters, to obtain the first mapping model and the second mapping model respectively. Before the massage begins, the patient’s desired Qi score is obtained, and based on the desired Qi score, the ideal physiological data corresponding to the desired Qi score is derived through the first mapping model. During the massage, the patient's physiological data is collected in real time, and the current Qi attainment score is inferred using the first mapping model. The score is then compared with the expected Qi attainment score to calculate the difference between the two scores. Based on the differences between the Qi-de scores, the massage control parameters corresponding to the ideal physiological data are inferred through the second mapping model, and the massage control parameters are dynamically adjusted using a feedback mechanism. Based on the dynamically adjusted massage control parameters, the massage operation continues, and the massage control parameters are continuously optimized based on real-time feedback. During the long-term treatment process, historical treatment data and feedback information from each patient are continuously collected, and reinforcement learning adaptive algorithms are used to optimize the second mapping model.
2. The control method for a massage robot based on multimodal feedback according to claim 1, characterized in that, The multimodal acquisition device includes: EEG acquisition equipment is used to monitor a patient's brain electrical activity and acquire brain wave signals related to emotions and relaxation levels; Heart rate sensor, used to monitor a patient's heart rate signal in real time; Electromyography (EMG) sensors are used to collect electrical signals of muscle activity from patients. The physiological data includes brain wave signals, heart rate signals, and electrical signals of muscle activity. The "Deqi" score refers to the patient's subjective rating of the comfort level during the massage.
3. The control method for a massage robot based on multimodal feedback according to claim 1, characterized in that, The massage control parameters include: Massage intensity, controlling the amount of force applied to the patient's body by the robotic arm; Massage frequency refers to the frequency of massage movements; Massage depth refers to the depth of penetration of the robotic arm during the massage process; Massage speed, or the speed at which the robotic arm moves, controls the rhythm of the massage movements.
4. The control method for a massage robot based on multimodal feedback according to claim 1, characterized in that, The mapping relationship between physiological data and Qi-de-qi scores is trained using deep learning algorithms to obtain the first mapping model, which specifically includes: The collected physiological data are preprocessed, and a physiological data feature vector is constructed based on the preprocessed physiological data, represented as follows: in, Indicates the first i Physiological data feature vectors of each sample; , , These represent the first and second parts after preprocessing. i EEG characteristic data, heart rate characteristic data and electromyography characteristic data of each sample; Using the physiological data feature vector as input and the corresponding Qi score of the physiological data feature vector as [something], training sample pairs are constructed, represented as follows: in, For the training dataset; Indicates the first i The gas yield score corresponding to each sample; This represents the number of training samples; A deep neural network model is constructed by inputting the physiological data feature vector into the deep neural network, obtaining the predicted result of the Qi score through forward propagation, and training the deep neural network based on the training samples. A supervised learning method is used to train a deep neural network. A loss function is constructed based on the difference between the predicted and actual Qi scores. The network parameters are iteratively updated using a gradient descent optimization algorithm, which makes the loss function gradually converge. Training is completed when the loss function meets the preset convergence condition or reaches the preset number of training rounds, and the trained deep neural network is used as the first mapping model.
5. The control method for a massage robot based on multimodal feedback according to claim 4, characterized in that, The deep neural network includes an input layer, multiple hidden layers, and an output layer. The input layer is used to receive physiological data feature vectors, and the output layer is used to output the predicted result of the Qi score.
6. The control method for a massage robot based on multimodal feedback according to claim 1, characterized in that, The second mapping model is constructed as follows: The collected physiological data are preprocessed, and physiological data feature vectors are constructed based on the preprocessed physiological data. ,in, Indicates the first i Physiological data feature vectors of each sample; physiological data feature vectors As input, the corresponding massage control parameter vector As supervisory labels, training sample pairs are constructed. ; Construct a deep neural network, input training sample pairs into the deep neural network, train the deep neural network, and the output of the deep neural network is the prediction result of the massage control parameters; During training, a loss function is constructed based on the difference between the predicted massage control parameters and the actual massage control parameters, and an optimization algorithm based on gradient descent is used to update the parameters of the deep neural network. When the loss function meets the preset convergence condition or the number of training rounds reaches the preset threshold, training stops, and the trained deep neural network is used as the second mapping model.
7. A control method for a massage robot based on multimodal feedback according to claim 4 or 6, characterized in that, The preprocessing of the collected physiological data specifically includes: The EEG signal was bandpass filtered to remove power frequency interference and high frequency noise. The filtered EEG signal was then segmented and its features were extracted to obtain EEG feature data. The heart rate signal is smoothed and baseline corrected, and the heart rate variability correlation characteristics are calculated to obtain heart rate characteristic data; The electromyographic signals are rectified and integrated to normalize the amplitude of muscle activity and obtain electromyographic characteristic data. The processed EEG, heart rate, and electromyography (EMG) feature data are combined in a preset order to form a physiological data feature vector. .
8. The control method for a massage robot based on multimodal feedback according to claim 1, characterized in that, The process involves inferring massage control parameters corresponding to ideal physiological data based on the differences in Qi (vital energy) scores using a second mapping model, and dynamically adjusting these parameters using a feedback mechanism. Specifically, this includes: During the massage, the patient's physiological data is collected in real time, and the current moment is obtained. t Physiological data feature vector ; Current moment t Physiological data feature vector Input the first mapping model Predict the current Qi score : in, This is the first mapping model; For the current moment t The predicted Qi score; According to the current time t The error in the predicted Qi attainment score is calculated by comparing it with the patient's expected Qi attainment score. in, The expected score for obtaining Qi set for the patient; Indicates the current time t The error in the Qi-gathering score; Using the second mapping model For the current moment t Physiological data feature vector With the feature vector of ideal physiological data By performing massage control parameter reasoning, the control parameter increments are obtained: in, Indicates the increment of the control parameter; The feature vector representing ideal physiological data is obtained based on the patient's desired Qi attainment score and the first mapping model. Based on the error between the control parameter increment and the Qi (vital energy) value, the dynamically adjusted massage control parameters are obtained: in, This is the feedback gain coefficient; For the current moment t The dynamically adjusted massage control parameters; Will The data is sent to the massage robotic arm, which then performs the actions according to the massage force, frequency, depth, and speed.
9. The control method for a massage robot based on multimodal feedback according to claim 1, characterized in that, During the long-term treatment process, by continuously collecting historical treatment data and feedback information from each patient, the second mapping model is optimized using a reinforcement learning adaptive algorithm, specifically including: An individual treatment history dataset is created for each patient, represented as follows: in, Indicates the first Historical treatment dataset of 100 patients Indicates the first i Physiological data feature vectors of each sample This indicates the corresponding massage control parameters. This indicates the patient's Qi attainment score for that sample; Indicates the first Number of historical samples from patients; A state-action-reward model is constructed for each patient based on reinforcement learning, where: state For the first t The physiological data feature vector of the patient at any given time; action The vector of adjustment for control parameters; Using individual patients' historical data Personalized strategy network for each patient Training is performed, and the policy network input state is used. Output action And based on rewards Update network parameters : in, For policy network parameters; The learning rate; This represents the gradient operation applied to the policy network parameters; For the reward function; During real-time massage, the trained policy network is used to generate adjustment movements. And combined with the massage control parameters output by the second mapping model To obtain the final massage control parameters : Status updated in real time during each treatment cycle Deqi score and final massage control parameters This forms an individualized closed-loop adaptive adjustment, enabling continuous optimization of the second mapping model.
10. A control method for a massage robot based on multimodal feedback according to claim 9, characterized in that, The reward function is expressed as: in, The current air quality score is predicted using the first mapping model; The patient's desired Qi score; , Preset weighting coefficients; It is the Euclidean norm.