Children burn dressing change pain management method and system based on tactile and auditory stimulation
By constructing personalized pain management models and reinforcement learning models, and combining multimodal data acquisition with real-time feedback optimization, the synergistic regulation of tactile and auditory stimuli was achieved, solving the problem of limited analgesic effect during dressing changes for children with burns, and improving analgesic efficiency and treatment safety.
Patent Information
- Application Number
- CN202511508646.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, tactile and auditory stimulation lacks multimodal collaborative design during dressing changes for children's burns. Stimulation parameters cannot be adjusted in real time, and behavioral responses cannot be effectively integrated, resulting in limited analgesic effects and difficulty in achieving personalized intervention.
A personalized pain management model is built based on the patient's basic physiological data. Multimodal data is collected for feature extraction and fusion. Tactile and auditory stimulation parameters are generated using a reinforcement learning model. These parameters are dynamically optimized through real-time monitoring and combined with behavioral feedback for closed-loop control.
It achieves synergistic regulation of tactile and auditory stimulation, significantly improves analgesic efficiency, lowers tolerance threshold, enhances treatment compliance and personalized adaptation, avoids excessive or insufficient stimulation, and enhances treatment safety and adherence.
Smart Images

Figure CN120977494A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of burn dressing change pain management, in particular to a children burn dressing change pain management method and system based on tactile and auditory stimulation. BACKGROUND
[0002] Currently, pain management during children burn dressing change is one of the difficulties in clinical treatment. Due to the high sensitivity of children to pain and low cooperation degree, the traditional drug analgesic method has side effect risk, and non-drug intervention technology has gradually become a research hotspot. In the prior art, tactile stimulation (such as vibration cooling device) and auditory stimulation (such as music therapy) have been applied alone for pain relief, and the mechanism is to achieve analgesic effect by interfering with pain conduction pathway and regulating emotional state respectively. In addition, some studies attempt to combine multi-modal physiological signals (such as heart rate variability, skin electric response) to monitor the degree of pain, and provide data support for personalized intervention.
[0003] However, the prior art has significant deficiencies: first, tactile and auditory stimulation are mostly applied in isolation, lacking multi-modal fusion and collaborative design, resulting in limited analgesic effect; second, stimulation parameters (such as vibration frequency, music rhythm) are usually set based on experience, and cannot be dynamically adjusted according to real-time physiological feedback of patients, making it difficult to adapt to individual differences; third, the behavioral responses (such as limb avoidance, crying) of children patients are not effectively integrated into the intervention model, making it difficult to achieve a balance between stimulation intensity and comfort; fourth, the existing method does not construct a patient-specific model, and cannot fuse basic physiological data (such as burn area, past pain history) with real-time multi-modal signals, limiting the precision of personalized intervention. SUMMARY
[0004] The present application aims to at least solve the technical problems in the prior art that tactile and auditory stimulation are applied in isolation and lack multi-modal collaborative design, and particularly innovatively proposes a children burn dressing change pain management method and system based on tactile and auditory stimulation.
[0005] In order to achieve the above-mentioned purpose of the present application, the present application provides a children burn dressing change pain management method based on tactile and auditory stimulation, the method comprising: S1, constructing a personalized pain management model based on patient basic physiological data; S2, collecting multi-modal data of the patient, the multi-modal data including pressure / temperature data of the wound area, biological signals and neural activity data; S3, performing feature extraction and data fusion on the multi-modal data based on the personalized pain management model to obtain fusion features; S4, generating tactile stimulation parameters and auditory stimulation parameters based on the fusion features using a reinforcement learning model; S5, applying the generated stimulation parameters to the haptic device and the auditory device, and synchronously applying the personalized stimulation during dressing change; S6, monitoring patient pain feedback data in real time, dynamically updating the reinforcement learning model parameters through the personalized pain management model, and iteratively optimizing the haptic stimulation parameters and the auditory stimulation parameters.
[0006] In another aspect, the present application also provides a child burn dressing change pain management system based on haptic and auditory stimulation, which is used to perform the child burn dressing change pain management method based on haptic and auditory stimulation; the system comprises: a signal acquisition module for acquiring voice signals and physiological signals of a child in real time, including EEG signals and oxygenated hemoglobin concentration data; a feature extraction module for extracting a multi-dimensional MFCC coefficient vector and a fundamental frequency trajectory from the voice signals to generate an emotional feature vector, and extracting a theta band power spectral density and an oxygenated hemoglobin concentration change from the physiological signals to generate an electroencephalogram feature; a feature fusion module for fusing the emotional feature vector and the electroencephalogram feature into a unified fusion feature; a reinforcement learning model module comprising an Actor network and a Critic network, for generating an action vector of haptic stimulation parameters and an action vector of auditory stimulation parameters based on the fusion feature; a stimulation output module for outputting customized haptic stimulation and auditory stimulation according to the action vectors; a feedback monitoring module for monitoring heart rate variability, pain expression indicators, comfort scores, and stimulus avoidance tendencies of a patient in real time; a model updating module for dynamically optimizing parameters of the reinforcement learning model through a time difference error and a policy gradient algorithm based on data of the feedback monitoring module; and a control interface module for receiving user inputted treatment requirements and boundary value settings.
[0007] The present application has the following beneficial effects: the present application can realize the synergistic regulation of haptic and auditory stimulation through multi-modal data fusion (fusion of pressure / temperature, biological signals, and neural activity data) and a reinforcement learning model. Specifically, time domain features, emotional feature vectors, and electroencephalogram features are deeply fused to generate a fusion feature containing pain level, emotional state, and neural activity; based on the fusion feature, an Actor network generates a joint action vector of haptic and auditory parameters (such as dynamic matching of vibration frequency and music rhythm). Through multi-modal information interaction, haptic stimulation (such as cooling vibration) can specifically inhibit wound pain signal conduction, while auditory stimulation (such as soothing music) can simultaneously regulate amygdala emotional response, and the synergistic effect of the two significantly improves analgesic efficiency while reducing the tolerance threshold of children to single stimulation.
[0008] The application also realizes personalized adaptive adjustment of stimulation parameters through real-time monitoring and dynamic optimization mechanism. Specifically, the Q value of the current action is evaluated using the Critic network, and the timing difference error is calculated in combination with the multi-objective reward function (including heart rate variability improvement, comfort score, etc.); the Actor network parameters are updated through the policy gradient algorithm, so that the tactile / auditory parameters (such as vibration amplitude, music volume) are optimized in the direction of maximizing cumulative rewards. Get rid of experience dependence, realize dynamic matching of parameters (such as automatically increasing vibration frequency or reducing volume according to real-time pain feedback of children), avoid overstimulation (such as secondary damage caused by pressure exceeding limit) or insufficient stimulation (such as weak analgesic effect).
[0009] The application also incorporates behavioral feedback (such as limb avoidance, crying) into the closed-loop control through multi-modal data acquisition (including biological signals) and reward function design. Specifically, the fundamental frequency trajectory of the child's crying sound (such as high-frequency screaming indicating strong pain) is extracted through voice emotion analysis, limb movements (such as wound avoidance movements) are monitored in real time, and a "stimulation avoidance tendency" penalty term (weight coefficient adjustable) is introduced into the reward function. Behavioral feedback directly drives parameter adjustment (such as automatically reducing tactile pressure when avoidance movements are detected), avoiding the single mapping defect of "stimulation parameters-physiological indicators" in traditional methods, improving the treatment cooperation degree of children and reducing the interruption rate of intervention.
[0010] The application also realizes deep integration of patient-specific features through personalized pain management model construction and feature fusion. Specifically, collect patient's basic physiological data (such as burn depth, past pain history), and train patient-specific features (such as children who are more sensitive to high-frequency vibration automatically adjust parameter range) through reinforcement learning; based on the model, weight the multi-modal data (such as children with a history of pain increase the weight of emotional features). The model can learn individual differences (such as differences in children's sensitivity to sound), avoid uniform intervention, reduce the use of analgesic drugs in personalized model groups, and improve long-term treatment compliance.
[0011] Additional aspects and advantages of the application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and / or additional aspects and advantages of the application will become apparent and be readily understood from the following description, taken in conjunction with the accompanying drawings, in which: Figure 1 is a flowchart of a children's burn dressing pain management method based on tactile and auditory stimulation according to the application. DETAILED DESCRIPTION
[0013] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0014] Example 1 like Figure 1 As shown, a method for pain management during burn dressing changes in children based on tactile and auditory stimulation includes: S1. Construct a personalized pain management model based on the patient's basic physiological data; The basic physiological data collected in this embodiment include burn depth (classified according to the International Burn Society grading standard), past pain history (including records of analgesic drug use and pain sensitivity scores), child's age (accurate to month), weight, baseline heart rate (measured at rest), skin sensitivity index (obtained by Von Frey fiber test), and burn wound area (calculated as a percentage of body surface area).
[0015] In this embodiment, the method for constructing a personalized pain management model is as follows: First, based on the collected basic physiological data of the patient (including burn depth, history of pain, age, weight, baseline heart rate, skin sensitivity index, and burn wound area), the parameters of the Actor network and Critic network in the reinforcement learning model are initialized. Specifically, the Actor network is constructed using a multilayer perceptron architecture, where the input layer receives fused feature vectors, the hidden layer uses the ReLU activation function, and the output layer generates action vector distributions for tactile and auditory stimulus parameters, respectively. The Critic network uses a fully connected structure to evaluate the Q-value of the state-action pair. During initialization, the weight matrix and bias vector of the Actor network are normalized and scaled according to the patient's age and weight. For example, younger patients correspond to a more conservative initial variance setting to reduce the risk of overstimulation. At the same time, the model parameters are pre-trained using the patient's basic physiological data, and the initial strategy is generated through an offline simulation environment. The simulation environment is constructed based on historical burn case data and includes typical pain feedback sequences.
[0016] S2. Collect multimodal data from the patient, including pressure / temperature data, biosignals, and neural activity data of the wound area; In step S2, the pressure / temperature data of the wound area is acquired in real time by a flexible sensor array attached to the inside of the wound dressing, which records the local pressure distribution and surface temperature changes at a sampling frequency of 10 Hz; the biological signals include real-time electrocardiogram (ECG), galvanic skin response (GSR), and facial electromyography (EMG) signals, which are collected synchronously by wireless wearable biosensors at a sampling rate not less than 256 Hz; the neural activity data are acquired by a portable near-infrared functional near-infrared spectroscopy (fNIRS) system, which focuses on monitoring the changes in oxygenated hemoglobin concentration in the anterior cingulate gyrus and the primary somatosensory cortex. All multi-modal data streams are synchronized at the millisecond level by time stamping and transmitted to the central processing unit via Bluetooth 5.0, forming a multi-dimensional pain response feature matrix with spatiotemporal alignment. For pediatric patients, the sensor interface is designed with a soft silicone package to reduce the risk of secondary damage, and all acquisition processes are started 5 minutes before the dressing change operation to establish an individualized pain baseline reference.
[0017] In this embodiment, the specific steps for forming a multi-dimensional pain response feature matrix with spatiotemporal alignment are as follows: First, the original acquisition data is time-stamped, and a Bluetooth 5.0 module based on the NTP protocol is used to achieve cross-device millisecond-level synchronization; then, the pressure / temperature data is spatially gridded, converting the original voltage signals from the flexible sensor array into a 32x32 pressure distribution matrix, while extracting the temperature gradient of each unit; finally, data standardization is performed, i.e., R-wave detection is performed on the ECG signal and the RR interval is calculated, the fNIRS data is calculated using the modified Beer-Lambert law to calculate the oxygenated hemoglobin concentration change value, and all feature dimensions are standardized by Z-score to eliminate individual baseline differences, forming a multi-dimensional pain response feature matrix containing 128 spatiotemporal features.
[0018] The central processing unit in this embodiment is deployed on a medical-grade embedded computing platform in the hospital cloud, which integrates a quad-core ARM Cortex-A53 processor (clock speed 1.8 GHz) and an FPGA acceleration module, and realizes high-speed data interaction through a PCIe bus. The hardware architecture adopts a dual-redundancy design, with the main processor responsible for running the Actor-Critic network inference of the reinforcement learning model (processing 32 frames of feature matrix per second), and the FPGA module performing parallel multi-modal data preprocessing (including spatial filtering of pressure matrix, QRS wave detection of ECG signal, and motion artifact correction of fNIRS data). The memory configuration includes 8 GB DDR4 running memory and 256 GB solid state storage, supporting real-time storage of 120 hours of multi-modal data stream. To meet the electromagnetic compatibility requirements of medical devices, the platform shell is made of aluminum-magnesium alloy by one-piece forming process, with independent analog signal acquisition area and digital processing area inside, and the radiation interference is controlled below -60 dBm through multi-layer shielding design. The operating system selects a real-time Linux kernel (PREEMPT-RT patch) with Xenomai hard real-time extension to ensure that the delay of the stimulation output module is stable within 5 ms.
[0019] S3, feature extraction and data fusion of multi-modal data based on personalized pain management model to obtain fused features; S4, generating haptic stimulation parameters and auditory stimulation parameters based on the fused features using a reinforcement learning model; S5, applying the generated stimulation parameters to the haptic device and the auditory device to synchronize the application of personalized stimulation during dressing change; In step S5, the need for detailed description is that the tactile device uses a flexible embedded micro-vibration motor array integrated inside the wound dressing, composed of multiple independent adjustable piezoelectric ceramic units, which can realize precise regulation according to the action vector of the tactile stimulation parameters (such as vibration frequency, amplitude and spatial distribution pattern). The vibration frequency dynamic range is 5-200Hz, and the amplitude control accuracy is ±0.05mm, to match the pain feedback of the child; at the same time, the auditory device is implemented through low-delay Bluetooth bone conduction earphones, which receive the action vector of the auditory stimulation parameters (including music rhythm, volume and spectral characteristics), real-time synthesis of personalized audio sequences, music rhythm and real-time heart rate variability of the child are synchronized, and the volume is self-adaptive according to the environmental noise level (noise suppression threshold is adjustable). During the dressing change operation, the control system ensures that the application of tactile and auditory stimulation is strictly synchronized with the dressing change action of medical staff through the time stamp mechanism, for example, enhancing high-frequency vibration (>100Hz) in the dressing removal stage to disperse pain attention, and playing soothing melodies with a fundamental frequency lower than 200Hz in the wound cleaning link to regulate emotional response; In addition, the device has a built-in safety monitoring module that continuously detects local skin temperature (threshold upper limit 40°C) and pressure distribution (maximum pressure limit 15kPa), and once the limit is exceeded, the stimulation output is automatically interrupted and an alarm is sent through the control interface module to avoid the risk of secondary injury. The application of stimulation parameters is also combined with a personalized pain management model, adjusting the initial parameter range for children with different burn depths, such as shallow burns prefer low-frequency vibration (<50Hz), while deep burns activate the temperature regulation mechanism (cooling mode activated temperature difference ±2°C), ensuring individualized adaptation of the intervention.
[0020] In the above-mentioned "During the dressing change operation, the control system ensures that the application of tactile and auditory stimuli is strictly synchronized with the dressing change action of the medical staff through the timestamp mechanism", the timestamp mechanism is implemented as follows: First, before the dressing change operation starts, the medical staff presets the dressing change action sequence (including the dressing removal, wound cleaning and new dressing application stages) through the control interface module (such as a touch screen operation terminal or a voice command device) and inputs the expected starting time of each stage; the system generates a global timestamp reference based on the real-time clock of the central processing unit (accuracy ±1ms, synchronized with the hospital network through NTP protocol) (such as recorded in UTC time format, formatted as "year-month-day hour: minute: second.millisecond"); during the dressing process, the action start signal is triggered by a foot switch or a voice command (such as the medical staff saying "start removing" instruction), the system immediately records the accurate timestamp of that moment, and transmits it to the stimulus output module through Bluetooth 5.0; at the same time, the flexible sensor array monitors the dressing pressure change in real time (sampling rate 10Hz), and automatically corrects the action start timestamp when a sudden pressure drop (threshold >5kPa / s) is detected, ensuring the robustness of action detection. The stimulus output module applies personalized stimulation (such as activating the vibration motor array at a frequency of 200Hz) within 5ms after the start of the action according to the timestamp and the preset action type (such as high-frequency vibration enhancement stage for dressing removal), and adjusts the delay compensation in real time through the FPGA acceleration module. In addition, the system has a built-in timeout protection mechanism (default timeout threshold 30s), if the completion signal (such as pressure recovery stable) is not detected at the expected end of action timestamp, the stimulus output is automatically interrupted and an alarm is triggered to prevent the stimulus from being out of sync with the action.
[0021] S6, real-time monitoring of patient pain feedback data, dynamic updating of reinforcement learning model parameters through personalized pain management model, iterative optimization of tactile stimulation parameters and auditory stimulation parameters.
[0022] The core principle of the method is to build a closed-loop adaptive control system, which realizes the precise optimization of pain intervention through the synergistic mechanism of multi-modal perception and reinforcement learning. Specifically, before the dressing change operation starts, the system initializes the individualized pain management model based on the patient's basic physiological data (such as burn depth, age, weight, and skin sensitivity). The Actor network generates the action vector distribution of tactile and auditory parameters (such as vibration frequency, amplitude, and music rhythm) through a multi-layer perception architecture, and the Critic network evaluates the Q value of the state-action pair to quantify the intervention effect. During the dressing change process, the real-time collected wound pressure / temperature data, biological signals (ECG, GSR, EMG), and neural activity data (fNIRS monitored oxygenated hemoglobin changes) are converted into emotional feature vectors (multi-dimensional MFCC coefficients and fundamental frequency trajectories) and electroencephalogram features (theta band power spectral density) by the feature extraction module, and are integrated into a unified fusion feature vector by the feature fusion module, which deeply encodes the pain level, emotional state, and neural activation level. The reinforcement learning model uses the Actor network to output tactile stimulation parameters (such as vibration frequency 5-200Hz, amplitude ±0.05mm) and auditory stimulation parameters (such as music rhythm synchronized with heart rate variability, and volume adaptive noise suppression), and applies customized intervention through the stimulation output module: tactile stimulation (flexible vibration motor array) directly inhibits the wound pain conduction pathway (such as Aδ fiber signal), auditory stimulation (bone conduction earphones) synchronously regulates the emotional response of the limbic system (such as reduced amygdala activation), and the two are dynamically matched (such as high-frequency vibration and soothing melody coordination) to distract attention and improve analgesic efficiency. At the same time, the feedback monitoring module real-time tracks heart rate variability, comfort score, pain performance indicators (such as high-frequency peaks in the crying fundamental frequency trajectory), and stimulus avoidance tendency (such as limb movement detection), and the model update module drives the policy gradient algorithm to optimize the Actor and Critic network parameters through the time difference error (calculated based on the multi-objective reward function, including heart rate improvement, comfort weight, and avoidance penalty term), and the Actor network parameter update is realized through gradient ascent to ensure that the stimulation parameters adaptively adjust towards the direction of maximizing cumulative rewards (such as automatically reducing pressure below the safety threshold of 15kPa when avoidance actions are detected). This principle not only realizes individual difference learning (such as increasing the emotional feature weight for those with a history of severe pain), but also avoids overstimulation (temperature exceeding 40°C interrupts output) or insufficient stimulation through closed-loop control, significantly improving treatment safety and patient compliance.
[0023] In the step of generating the haptic stimulation parameters in the present embodiment, first, the Actor network of the reinforcement learning model receives the 128-dimensional fusion feature vector output by the feature fusion module as input; the network is processed through a multi-layer perceptron architecture, the input layer maps the feature vector to the hidden layer, the hidden layer uses the ReLU activation function to perform nonlinear transformation (including three layers of fully connected structure, the number of neurons in each layer is 64, 32 and 16 respectively); the output layer generates the action vector distribution of the haptic stimulation parameters, specifically including three dimensions: the vibration frequency parameter (continuous value domain 5-200Hz), the amplitude parameter (precision ±0.05mm) and the spatial distribution mode parameter (discrete coding of the active unit index of the vibration motor array); the action vector is modeled by a Gaussian probability distribution, where the mean and variance are dynamically calculated by the network weights, for example, for the high-frequency pain feedback scene, the variance is automatically reduced to improve the stability of the parameters; the sampling process uses the importance sampling strategy to extract specific parameter values from the distribution, and combines the Q value evaluated by the Critic network to calibrate the confidence; at the same time, the parameter generation module integrates real-time safety constraint logic, such as when the fusion feature detects that the skin sensitivity index is over-standard, the upper limit of the amplitude is forcibly locked below 10kPa, or the frequency range is dynamically adjusted according to the depth of the burn (superficial burns prefer to limit to <50Hz), to ensure that the output parameters meet the individualized safety threshold.
[0024] The step of generating the auditory stimulation parameter is that: the Actor network of the reinforcement learning model receives the same 128-dimensional fusion feature vector input; the auditory parameter generation adopts a double-layer GRU network architecture to process the time sequence features, wherein the input layer divides the feature vector sequence by 10 ms frame length, the hidden layer extracts the music emotion features (such as the dynamic track of the MFCC coefficient) and the neural response features (the change of the theta band power spectrum density) through the gating mechanism; the output layer generates the action vector distribution of the auditory stimulation parameter, including three core dimensions: the music rhythm parameter (BPM range 60-180, with the real-time heart rate variability of the patient synchronized deviation controlled within ±5%), the volume parameter (dynamic range 40-85 dB, accuracy ±2 dB) and the frequency spectrum characteristic parameter (personalized equalizer curve is generated through 128-point FFT, focusing on enhancing the frequency band with a fundamental frequency lower than 200 Hz). The action vector is modeled by Beta distribution to adapt to the bounded parameter space, and the bias correction is performed by combining the advantage function output by the Critic network during sampling; in particular, an environmental noise adaptive mechanism is integrated in the parameter generation (the environmental sound pressure level is collected in real time through a microphone array), when the noise exceeds 55 dB, the volume is automatically increased for compensation (slope 3 dB / dB), and the spectrum subtraction algorithm is activated to suppress noise in specific frequency bands (such as 1 kHz notch filter for medical equipment alarm sound). The safety constraint logic includes: according to the age of the patient, the upper limit of the volume is set (children under 6 years old are locked below 75 dB), and when the EEG feature detects the tendency of restlessness (the theta / beta power ratio exceeds 1.5 times the baseline), the soothing spectrum mode is automatically switched (the fundamental frequency is reduced by 20%).
[0025] As an optional embodiment of the present application, optionally, the personalized pain management model is constructed in step S1, which includes: S101, collecting basic physiological data of the patient, the basic physiological data including age, gender, weight, height, total burn surface area percentage, burn depth and previous pain response history; S102, preprocessing the basic physiological data; It should be noted that the preprocessing includes data cleaning, missing value filling and outlier detection. Specifically, the data cleaning aims to eliminate invalid data points caused by sensor failure or collection error, and ensure data quality; the missing value filling adopts mean interpolation or linear interpolation strategy of previous and next values to maintain the integrity of the data sequence; the outlier detection identifies and labels extreme values outside the reasonable range, and manually reviews and corrects if necessary.
[0026] S103, based on the preprocessed data, applying a reinforcement learning model to train a personalized pain management model and integrating patient-specific features; In step S103, the training process of the reinforcement learning model uses the Deep Deterministic Policy Gradient (DDPG) algorithm framework, where the state space is defined as a multi-dimensional vector that integrates patient-specific features, including: normalized age factor (mapped to the [0, 1] interval), body mass index (BMI percentile), total burn surface area percentage (TBSA%), burn depth encoding (I degree = 0.2, II degree shallow = 0.4, II degree deep = 0.6, III degree = 0.8), past pain response sensitivity score (converted to a 0-10 scale based on historical analgesic drug dosage), and skin sensitivity index (log-transformed value of Von Frey test results). The action space is defined as a continuous value domain, including the joint vector of tactile stimulus parameters (vibration frequency 5-200 Hz, amplitude 0.1-1.0 mm) and auditory stimulus parameters (music fundamental frequency 100-400 Hz, volume 40-70 dB SPL). The training process is carried out in a simulated environment, which is based on a historical multi-modal dataset of burn children and simulates typical pain feedback patterns corresponding to different dressing operation stages (such as dressing removal, debridement, and bandaging). During training, the Critic network uses a double Q network structure to alleviate overestimation problems, its input is the concatenation of state vector and action vector, and its output is the Q value estimation of state-action pairs; the Actor network is updated through policy gradient, and its objective function is defined as maximizing the expected cumulative discounted reward. Patient-specific features are directly input into the network through the corresponding dimensions of the state vector and participate in gradient calculation during network weight update, thereby learning individual differences, for example: for children with high past pain sensitivity scores, the Critic network will automatically assign greater negative reward weights to heart rate variability (HRV) decreases; for children with high skin sensitivity index, the Actor network will constrain the amplitude action output to a lower range (<0.5 mm) during the initial exploration stage. The training uses a segmented strategy, the first stage performs 10000 offline iteration updates to stabilize the basic policy, the second stage introduces real-time biosignal feedback (simulated ECG, GSR fluctuations) for online fine-tuning, and finally the model parameters with the highest cumulative reward on the validation set are selected as the deployment version through cross-validation to ensure the model's generalization ability for specific children.
[0027] S104, verify the performance of the individualized pain management model through cross-validation or clinical historical data set, the evaluation indicators include accuracy, recall rate and F1 score, and dynamically adjust the model hyperparameters according to the verification results.
[0028] It is necessary to explain in step S104 that cross-validation adopts k-fold strategy, where k value is set to 10, and stratified sampling is used to ensure balanced distribution of different burn depths and age groups to avoid model overfitting; the clinical history data set is selected from a multi-center burn pediatric database, containing at least 500 anonymous case records, and the data set is divided into training set (70%), validation set (15%) and test set (15%) according to time stamp order, the training set is used for preliminary training of the model, the validation set is used for hyperparameter tuning, and the test set is used for final performance evaluation. The calculation of evaluation indexes is based on confusion matrix analysis: accuracy is defined as (true positive + true negative) / total sample number, which is used to measure the overall prediction correctness; recall rate (i.e. sensitivity) is true positive / (true positive + false negative), which reflects the comprehensiveness of identifying children with high pain state; F1 score is 2 x (accuracy x recall rate) / (accuracy + recall rate), which is a comprehensive index balancing accuracy and recall rate, and AUC-ROC curve area is introduced to evaluate the robustness of the model at different thresholds. According to the verification results, the process of dynamically adjusting the model hyperparameters adopts automatic grid search combined with Bayesian optimization framework: grid search covers learning rate (range 0.0001-0.01), batch size (16-128), Actor network hidden layer neuron number (64-256) and discount factor γ (0.9-0.99); Bayesian optimization models the hyperparameter space through Gaussian process, with the F1 score of the validation set as the objective function, and real-time monitoring of overfitting signs (such as training loss and validation loss difference exceeding 10%) during the iteration process, once overfitting is detected, immediately adjust the regularization parameter (such as increasing L2 regularization coefficient by 0.001) or stop training in advance; finally, the optimized hyperparameter configuration is updated to the deployment environment through the control interface module to ensure that the model maintains high generalization performance in real-time application.
[0029] In this embodiment, the process of cross-validation is as follows: the k-fold strategy is adopted, and the value of k is set to 10. Stratified sampling is used to ensure the balanced distribution of different burn depths (such as I degree, II degree shallow, II degree deep, III degree) and age groups (such as <3 years old, 3-6 years old, >6 years old), to avoid overfitting of the model to a specific subgroup. In specific implementation, the clinical history dataset (selected from the multi-center burn pediatric database, containing at least 500 anonymous cases) is randomly divided into 10 mutually exclusive subsets (folds), each fold contains the same proportion of burn depth and age samples; in each iteration, 9 folds are used as the training set for model parameter optimization, and the remaining 1 fold is used as the validation set for performance evaluation. The training process is based on the deep deterministic policy gradient (DDPG) algorithm framework, and the model update is performed on the training set: the Actor network maximizes the cumulative discounted reward through policy gradient, and the Critic network uses a double Q network structure to alleviate the overestimation problem, the input is the concatenation of the state vector (normalized age factor, BMI percentile, TBSA%, burn depth encoding, past pain response sensitivity score, skin sensitivity index) and the action vector (tactile stimulus parameters such as vibration frequency, amplitude, auditory stimulus parameters such as music rhythm, volume), and the output is the Q value estimate; during training, the Adam optimizer is used, and the batch size is dynamically adjusted according to grid search (range 16-128), and the learning rate is initially set to 0.001. In the validation phase, the confusion matrix is calculated on the validation set: based on the model's prediction of high pain state (threshold set to pain score ≥7) and actual label, the accuracy is defined as (true positives + true negatives) / total number of samples, the recall rate is true positives / (true positives + false negatives), and the F1 score is 2 * (accuracy * recall) / (accuracy + recall), and the AUC-ROC curve is simultaneously drawn to evaluate the robustness of the model at different decision thresholds. At the same time, based on the validation results, the hyperparameters are dynamically adjusted: using automated grid search to cover learning rate (0.0001-0.01), Actor network hidden layer neuron number (64-256), discount factor γ (0.9-0.99), and batch size (16-128), combined with the Bayesian optimization framework (Gaussian process modeling of hyperparameter space, target function is to maximize the F1 score of the validation set); during optimization, the training loss and validation loss gap are monitored in real time, and if the gap exceeds 10%, it is determined that there is a risk of overfitting, and the L2 regularization coefficient is immediately increased (step size 0.001) or the training is stopped in advance. This cross-validation cycle is executed 10 times (each fold is rotated as the validation set), and the final performance indicators are calculated by averaging the results of all folds (such as average F1 score and AUC-ROC value), the optimal hyperparameter configuration (such as learning rate 0.0005, batch size 64, Actor neuron number 128) is selected, and the model is updated to the deployment environment through the control interface module to ensure that the model maintains high generalization performance in real-time applications.
[0030] As an optional embodiment of the present application, the feature extraction and data fusion of the multi-modal data based on the personalized pain management model in step S3 comprises: S301, using time-frequency analysis method to extract time domain features and frequency domain features from the wound area pressure / temperature data; It should be noted that in step S301, the time-frequency analysis method such as short-time Fourier transform (STFT) or continuous wavelet transform (CWT) can effectively capture the dynamic characteristics of the wound pressure / temperature data changing with time, the time domain features include average pressure, maximum pressure change rate and temperature fluctuation standard deviation, and the frequency domain features focus on the main frequency band energy distribution and spectral entropy, which together reflect the local microcirculation state of the wound and the potential pain stimulation response.
[0031] In step S301, the data collection is a short-time collection during the preparation stage of the dressing change operation, and the specific collection time is 15-30 seconds, which can cover the typical dynamic change period of the wound pressure / temperature data (such as physiological response in the initial stage of dressing removal or debridement), and at the same time avoid prolonging the waiting time of the patient. The system of the present embodiment integrates a high-speed pressure / temperature sensor array (sampling rate ≥100Hz), quickly acquires data stream before dressing change, and efficiently extracts features by using real-time time-frequency analysis method (such as short-time Fourier transform STFT window length set to 1 second, overlap rate 50%); The process is seamlessly connected with the dressing change operation, ensuring that the feature extraction is completed within the clinically feasible time (total processing delay <2 seconds).
[0032] S302, performing speech emotion analysis on the biological signal to obtain an emotion feature vector; It is necessary to explain in step S302 that the voice emotion analysis is for the sound data in the collected biological signals (such as the crying or speech response of the child patient), the mel frequency cepstral coefficient (MFCC) algorithm is used to extract the frequency spectrum characteristics, and the pitch dynamic change is analyzed combined with the fundamental frequency (F0) track. Specifically, the original audio signal (the sampling rate is not less than 16 kHz) is processed through pre-emphasis, framing and windowing, the 39-dimensional MFCC coefficients (including static, first-order difference and second-order difference characteristics) of each frame are calculated, and the autocorrelation function or cepstrum method is used to track the fundamental frequency track to quantify the pitch fluctuation range (such as the peak value of the fundamental frequency > 500 Hz in the pain state) and the harmonic-to-noise ratio (HNR); the feature vector construction also includes short-time energy, spectral centroid and zero-crossing rate and other time domain indicators, these parameters are reduced to 15-20 dimensional emotion feature vectors through principal component analysis (PCA), encode the emotional states such as anxiety, pain or calm, and are time-aligned with ECG, GSR and EMG signals to enhance the robustness of pain emotion recognition; the analysis process integrates the open source tool library (such as LibROSA) to realize real-time processing, and the filter bank (mel scale range 80-5000 Hz) is optimized for children's pitch characteristics, to ensure that the emotional classification accuracy rate is maintained above 90% under environmental noise interference (signal-to-noise ratio > 15 dB).
[0033] S303, using electroencephalogram analysis method on neural activity data to extract electroencephalogram characteristics; It is necessary to explain in step S303 that the electroencephalogram analysis method is for the collected neural activity data (such as the electrophysiological signals recorded by the high-density EEG electrode array), and the power spectral density (PSD) analysis is used to extract key electroencephalogram characteristics; specifically, the signal preprocessing includes band-pass filtering (range 0.5-40 Hz) to suppress power frequency noise and motion artifacts, and the independent component analysis (ICA) is used to separate and remove eye movement, electromyographic interference components; in the feature extraction stage, the Welch method is applied to calculate the power spectral density, and the average power values of δ (0.5-4 Hz), θ (4-8 Hz), α (8-12 Hz), β (12-30 Hz) and γ (30-40 Hz) bands are quantified, wherein the power spectral density of the θ band is used as the core index of pain sensitivity (the threshold is set as ±20% change of the baseline value); the analysis process integrates the open source tool library (such as EEGLAB or MNE-Python) to realize real-time processing, the sampling rate is not less than 256 Hz to ensure the time-frequency resolution, the feature vector construction includes the power ratio between frequency bands (such as θ / β ratio) and event-related desynchronization (ERD) index, which enhances the coding robustness of the pain-induced neural response (such as prefrontal cortex activation), and is time-aligned with the fNIRS oxyhemoglobin data, and is reduced to 10-15 dimensional electroencephalogram feature vectors through principal component analysis (PCA).
[0034] S304, based on the personalized pain management model, the time domain features, the frequency domain features, the emotional feature vectors and the electroencephalogram features are fused to obtain the fusion features.
[0035] It should be noted that in step S304, the multi-modal data fusion adopts a feature-level fusion strategy. First, the time domain features, the frequency domain features, the emotional feature vectors and the electroencephalogram features are standardized (such as normalization based on Z-score, to ensure that the mean of each feature is 0 and the standard deviation is 1) to eliminate the dimensional difference. Then, through the feature fusion module integrated in the personalized pain management model (implemented based on a multi-layer perceptron architecture), the standardized feature vectors are spliced into a high-dimensional joint vector. The module includes a fully connected layer with 128-256 neurons, applies a ReLU activation function for nonlinear transformation, and outputs a fixed-dimension fusion feature vector (for example, 15-20 dimensions). The fusion process integrates open source libraries (such as TensorFlow or PyTorch) to realize real-time calculation. The fusion feature vector encodes the pain intensity, emotional fluctuations and neural response patterns (such as the synergistic changes of frontal θ wave power and heart rate variability) in depth, and is directly input into the Actor and Critic networks of the reinforcement learning model for dynamic optimization of stimulation parameters. At the same time, the fusion module further reduces the dimension through principal component analysis (PCA) or automatic encoder technology to minimize feature redundancy (such as a feature variance contribution rate of >95%), and participates in the time difference error calculation in the model update module to improve the robustness and real-time performance of the closed-loop control system.
[0036] As an optional embodiment of the present application, the expressions for extracting the time domain features and the frequency domain features in S301 are as follows: ; ; wherein, represents the mean pressure, represents the number of samples, represents the instantaneous pressure value at time , represents the peak pressure, represents the continuous wavelet transform, represents the wavelet basis function, represents the scale, represents the translation, represents the dominant frequency component, represents the fast Fourier transform; The expression for obtaining the emotional feature vector in step S302 is as follows: , ; wherein, representing a multi-dimensional MFCC coefficient vector, representing a mel-frequency cepstral coefficient, representing a sampling point of a speech signal, representing a fundamental frequency trajectory; The expression of the brain electrical feature extracted in step S303 is: , ; wherein, represents a wave band power spectral density, represents a sample number, represents a time point a wave band EEG signal, represents a change in oxygenated hemoglobin concentration, represents oxygenated hemoglobin concentration at a time point, represents oxygenated hemoglobin concentration at a time point.
[0037] As an optional embodiment of the present application, optionally, the generating of the tactile stimulation parameter and the auditory stimulation parameter based on the fusion feature by using the reinforcement learning model in step S4 comprises: S401, constructing a reinforcement learning model based on an Actor network and a Critic network; It needs to be explained in step S401 that the Actor network is responsible for generating the action policy of the tactile stimulation parameters (including amplitude, frequency) and the auditory stimulation parameters (including tone, rhythm), and its output is defined as the action probability distribution; the Critic network is responsible for evaluating the state value function and predicting the cumulative discounted reward to guide the policy optimization. Specifically, the Actor network architecture adopts a multi-layer perceptron (MLP) implementation: the input layer receives the fusion feature vector generated by step S304 (the dimension is fixed at 15-20 dimensions), the hidden layer contains 128-256 neurons, and the ReLU activation function is applied for nonlinear transformation; the output layer of the Actor network uses the Softmax function to generate discrete action probabilities (such as amplitude intervals of 0.1-0.5mm, 0.6-1.0mm, etc.), or Gaussian distribution parameters (mean and standard deviation) to generate continuous actions (such as frequency range 20-200Hz); the output layer of the Critic network is a single value Q function estimate; the Critic network loss function uses mean square error (MSE). Network initialization uses the Xavier method, weight update uses the Adam optimizer, learning rate is dynamically adjusted through Bayesian optimization (initial value 0.001), and batch size is set to 32-64 to balance training efficiency and stability. At the same time, the network integrates the open source framework (such as TensorFlow or PyTorch) to realize real-time inference, and supports GPU acceleration calculation to ensure that the delay on embedded hardware (such as RaspberryPi) is less than 50ms.
[0038] S402, map the fusion features to the state space representation of the reinforcement learning model, and extract key physiological indicators as state variables; It needs to be explained in step S402 that mapping the fusion features to the state space representation of the reinforcement learning model includes: first, extracting key physiological indicators as state variables from the fusion feature vector (dimension fixed at 15-20 dimensions); these indicators include heart rate variability (HRV, calculated as the standard deviation of RR intervals), skin conductance level (SCL, average skin conductance value), theta band power density (extracted from EEG signals), and principal components of the emotional feature vector (such as anxiety index); the mapping process is implemented through a lightweight state encoder, which uses linear projection or a two-layer fully connected network (32-64 hidden layer neurons) to apply the ReLU activation function to convert the input features into a low-dimensional state vector (e.g. 8-12 dimensions); this process integrates open source libraries (such as NumPy or Pandas) to implement real-time calculation with a sampling interval of no more than 100ms, ensuring a delay of less than 30ms on embedded hardware (such as RaspberryPi), and seamlessly connecting with the input layers of the Actor and Critic networks for dynamic optimization of the stimulation strategy.
[0039] S403, setting boundary values of the tactile stimulation parameters and the auditory stimulation parameters according to the physiological tolerance of the child and the treatment needs, forming a continuous action space; It needs to be explained in step S403 that the boundary value setting strictly follows the clinical safety guidelines and the research data of the child pain perception threshold: the tactile stimulation amplitude is limited to 0.1-2.0mm (resolution 0.05mm), the frequency range is controlled to 20-200Hz (resolution 1Hz), and the mechanical vibration energy density is ensured not to exceed 3.0J / cm² to avoid tissue damage; in the auditory stimulation parameters, the tone is defined as the fundamental frequency range of 80-500Hz (covering the comfortable hearing frequency band of children), the rhythm is set to 60-120BPM (beats per minute), and the dynamic range is compressed to 45-70dBSPL to prevent auditory discomfort. The action space is represented by a two-dimensional continuous vector, the first dimension encodes the tactile stimulation intensity (a weighted combination of amplitude and frequency, the weight coefficient is determined by Bayesian optimization), and the second dimension encodes the auditory parameters (linear mapping of tone and rhythm). Boundary constraints are achieved through hard clipping in the environmental state space, when the Actor network output exceeds the safety threshold, the system automatically corrects the parameters to the nearest valid boundary value, and introduces a boundary penalty term (such as the part exceeding the threshold multiplied by the penalty coefficient λ=0.3) in the value evaluation of the Critic network, to strengthen the safety learning of the strategy. This design integrates the open source reinforcement learning library (such as StableBaselines3) to realize real-time action constraints, ensuring that the stimulation parameters always meet the safety standards of service robots and the specifications of children's medical equipment.
[0040] S404, based on the real-time monitoring of the patient's pain feedback data and the patient's comfort score, a multi-objective reward function is constructed, when the pain feedback data indicates that the physiological indicators have improved or the patient has no avoidance action, the reward is given, when the patient shows signs of pain, the punishment is applied, and the adjustable coefficient is introduced to balance the analgesic effect and the stimulation tolerance; It is necessary to explain in step S404 that the multi-objective reward function design comprehensively considers the analgesic effect and stimulation tolerance, ensuring the maximum comfort of children during dressing change. Specifically, the pain feedback data includes heart rate variability (HRV), skin conductance level (SCL), and facial expression analysis (pain, anxiety, and other emotional signals are identified through computer vision technology), which are input into the reward function module in real time. When HRV increases (indicating parasympathetic activation and pain relief), SCL decreases (reflecting reduced stress), or facial expression turns calm, the system gives positive rewards; on the contrary, if there is a painful expression, HRV decreases, or SCL rises sharply, punishment is applied. In addition, patient comfort scores are obtained through an interactive interface (such as through voice instructions or touch screen selection) as another important input of the reward function. To balance the analgesic effect and stimulation tolerance, an adjustable coefficient w (range 0-1) is introduced, where w=0 focuses entirely on the analgesic effect, and w=1 focuses on the stimulation tolerance. The w value is dynamically adjusted through online learning algorithms (such as the policy gradient method) to adapt to the pain sensitivity and treatment needs of different children. The reward mechanism is implemented using open source libraries (such as Gym or OpenAIBaselines) to ensure efficient and safe pain management in complex and variable clinical environments.
[0041] S405, based on the state variable, the continuous action space and the multi-objective reward function, the action vector of the tactile stimulation parameter and the action vector of the auditory stimulation parameter are generated by using the Actor network.
[0042] It is necessary to explain in step S405 that the Actor network generates the corresponding action vector according to the real-time input state variable (dimension 8-12 dimensions) through its policy network. The action vector is a two-dimensional continuous vector: the first dimension represents the intensity of the tactile stimulus, defined as the weighted linear combination of amplitude and frequency, i.e. ; wherein is the normalized amplitude value, is the normalized frequency value, and the weight coefficient is dynamically learned and determined in the training phase through Bayesian optimization; the second dimension represents the auditory stimulus parameter, encoded as the joint output of the pitch base frequency and the rhythm, i.e., wherein is the normalized value of the pitch base frequency, and is the normalized value of the rhythm. The network output layer adopts a double-branch structure, each branch outputs the mean and standard deviation parameters of the corresponding dimension, and the final action value is obtained by sampling through the reparameterization trick, ensuring the differentiability of the policy. In the action execution stage, the action vector obtained by sampling is mapped to the physical stimulation device in real time through the environment interface module: the tactile component drives the servo motor to generate mechanical vibration with corresponding amplitude and frequency, and the auditory component generates a soothing sound effect with the specified base frequency and rhythm through a digital audio synthesizer. This process integrates the open-source reinforcement learning library (such as RayRLlib) to achieve millisecond-level reasoning, and the embedded system ensures that the action instructions are transmitted and executed within 50ms through the hardware abstraction layer (HAL), realizing the closed-loop real-time regulation of the pain management strategy.
[0043] The reinforcement learning model realizes the dynamic optimization of tactile and auditory stimulus parameters through the following key mechanisms: 1. Reinforcement learning model training process: the model is trained in the offline stage through the historical child burn dressing pain response data set. The data set contains a spatiotemporally aligned multi-dimensional pain response feature matrix (input), corresponding physiological indicators and behavior observation labels (state), preset stimulus parameter attempts within the safety boundary (action), and reward values calculated based on pain feedback. Training uses the Proximal Policy Optimization (PPO) algorithm, which is iterated in a simulated environment (usually 105-106 steps) until the reward function converges and the policy is stable. The training goal is to maximize the cumulative reward (i.e. optimize the analgesic effect and comfort) while strictly meeting the preset action boundary constraints.
[0044] 2. Input and output: Input: The input of the model when running in real time is the fusion feature vector generated in step S304 (fixed dimension 15-20 dimensions), which has been mapped to a low-dimensional state vector (8-12 dimensions) by the state encoder, representing the real-time pain physiological and emotional state of the child.
[0045] Output: The output of the model (Actor network) is a two-dimensional continuous action vector: First dimension: tactile stimulus intensity action value (composed of normalized amplitude and normalized frequency weighted combination).
[0046] Second dimension: auditory stimulation parameter action values (encoded normalized pitch fundamental frequency and normalized rhythm).
[0047] These action values are sampled through parameterized probability distributions (e.g. Gaussian distribution) at the output layer, and finally mapped to specific stimulation parameter values within pre-defined safety boundaries (amplitude 0.1-2.0mm, frequency 20-200Hz, pitch 80-500Hz, rhythm 60-120BPM).
[0048] 3. Real-time optimization and adaptation: when deployed online, the model fine-tunes the Actor network parameters through policy gradient method based on the multi-objective reward function (combining real-time HRV, SCL, facial expression analysis, patient comfort score) defined in step S404 and the value evaluation of Critic network. This allows the system to dynamically adjust the stimulation strategy (e.g. change the w coefficient balance between analgesia and comfort) according to the real-time pain response and treatment tolerance of individual children, achieving personalized pain management closed-loop regulation. The model adapts to different age groups of children (e.g. 2-5 years old, 6-12 years old) through transfer learning technology, ensuring the effectiveness and safety of the strategy.
[0049] Therefore, the reinforcement learning model of the present application can safely and efficiently generate and real-time optimize personalized tactile and auditory stimulation parameters through structured network design, strict action boundary constraints, and reward mechanisms combined with multi-source feedback, effectively relieving the pain of children during the burn dressing change process.
[0050] As an optional embodiment of the present application, optionally, the expression of the multi-objective reward function in step S404 is: ; wherein, represents the total reward value, represents the adjustable weight coefficient of the improvement amount of heart rate variability, represents the improvement amount of heart rate variability, represents the adjustable weight coefficient of the pain performance indicator, represents the pain performance indicator, represents the adjustable weight coefficient of the comfort score, represents the comfort score, represents the adjustable weight coefficient of the stimulation avoidance tendency, represents the stimulation avoidance tendency.
[0051] As an optional embodiment of the present application, optionally, the expression of the action vector of the tactile stimulation parameter generated in step S404 is: ; ; ; wherein, represents a haptic stimulus parameter action vector, represents a mean function of the haptic stimulus parameter, represents a standard deviation function of the haptic stimulus parameter, represents a random noise vector, represents a hyperbolic tangent function, represents an Actor network weight matrix corresponding to the mean function of the haptic stimulus parameter, represents a state variable, represents an Actor network bias vector corresponding to the mean function of the haptic stimulus parameter, represents a haptic parameter scaling factor vector, represents a Softplus function, represents an Actor network weight matrix corresponding to the standard deviation function of the haptic stimulus parameter, represents an Actor network bias vector corresponding to the standard deviation function of the haptic stimulus parameter; The expression of the action vector of the auditory stimulus parameter generated in step S404 is: ; ; ; wherein, represents an auditory stimulus parameter action vector, represents a mean function of the auditory stimulus parameter, represents a standard deviation function of the auditory stimulus parameter, represents an Actor network weight matrix corresponding to the auditory mean function, represents an Actor network bias vector corresponding to the auditory mean function, represents an auditory parameter scaling factor vector, represents an Actor network weight matrix corresponding to the auditory standard deviation function, represents an Actor network bias vector corresponding to the auditory standard deviation function.
[0052] As an optional embodiment of the present application, optionally, in step S6, the patient pain feedback data is monitored in real time, the reinforcement learning model parameters are dynamically updated by the personalized pain management model, and the haptic stimulus parameter and the auditory stimulus parameter are iteratively optimized, comprising: S601, using a Critic network to evaluate the Q value of the action vector in the state space of the reinforcement learning model, and calculating the time difference error according to a multi-objective reward function; It is necessary to explain in step S601 that the Critic network receives the current state variable (dimension 8-12 dimensions) and the action vector generated by the Actor network as input, and outputs a single value Q function estimate through its fully connected layer (hidden layer neuron number 64-128, ReLU activation function is applied), indicating the expected cumulative reward of the state-action pair; at the same time, the system calculates the immediate reward value based on the multi-objective reward function (combined with heart rate variability improvement, pain performance indicator, comfort score and stimulus avoidance tendency, weight coefficient dynamic adjustment). The time difference error is used for the loss function optimization of the Critic network, and the loss is defined as the mean square error (MSE). The optimization process integrates the open source reinforcement learning framework (such as StableBaselines3 or RayRLlib), uses the Adam optimizer to dynamically update the network weight (the learning rate is maintained at 0.0005-0.001 through Bayesian optimization), ensures that the calculation delay on the embedded hardware (such as RaspberryPi) is less than 20ms, and realizes the efficient online update of the Critic network, providing a stable value benchmark for the policy gradient of the Actor network.
[0053] S602, update the Actor network parameters based on the time difference error through the policy gradient algorithm, optimize the action vector to the direction of maximizing the cumulative reward, and obtain the optimized action vector of the tactile stimulation parameter and the auditory stimulation parameter.
[0054] It is necessary to explain in step S602 that the policy gradient algorithm calculates the policy gradient of the Actor network through the temporal difference error, and the Proximal Policy Optimization (PPO) algorithm is used to improve the training stability; the gradient calculation is based on the Q value estimation output by the Critic network, and the Actor network weight parameters (including the weight matrix and bias vector corresponding to the mean function and the standard deviation function) are updated through the back propagation algorithm combined with the temporal difference error. The weight update process uses the Adam optimizer, and the learning rate is dynamically adjusted in the range of 0.0005-0.001 (fine-tuned online through Bayesian optimization) to maximize the cumulative reward and ensure the convergence efficiency. During the update process, the system monitors the boundary constraints of the action vector in real time, and when the optimized tactile stimulation parameters (amplitude and frequency weighted combination) or auditory stimulation parameters (tone fundamental frequency and rhythm linear mapping) exceed the safety threshold, the hard truncation mechanism is automatically applied to correct to the nearest valid value, and a boundary penalty term (penalty coefficient λ=0.3) is introduced in the gradient calculation to strengthen the safety learning. The process is realized by integrating the open source reinforcement learning library (such as RayRLlib or StableBaselines3), and the inference delay on embedded hardware (such as RaspberryPi) is strictly controlled within 20ms, and the action command is executed efficiently through the hardware abstraction layer (HAL); the optimized action vector is directly output to the physical stimulation device to drive the servo motor and digital audio synthesizer, realize closed-loop real-time control, and support continuous iteration and update of personalized pain management models.
[0055] As an optional embodiment of the application, optionally, the expression for calculating the temporal difference error in step S601 is: ; ; wherein, represents the state-action pair value function output by the Critic network, represents the state variable, represents the current action vector, represents the parameter set of the Critic network, represents the weight matrix of the Critic network, represents the vector splicing operation, represents the bias vector of the Critic network, represents the temporal difference error, represents the multi-objective reward function, represents the discount factor, represents the next time state variable, represents the next time action vector; The expression of the policy gradient algorithm in step S602 is: ; ; wherein, denotes a parameter set of a haptic stimulus parameter Actor network, denotes a learning rate, denotes a gradient operator on haptic Actor network parameters, denotes a log probability density of a haptic stimulus parameter action vector, denotes a haptic stimulus parameter action vector, denotes a parameter set of an auditory stimulus parameter Actor network, denotes a gradient operator on auditory Actor network parameters, denotes a log probability density of an auditory stimulus parameter action vector, denotes an auditory stimulus parameter action vector.
[0056] Embodiment 2 A child burn dressing change pain management system based on haptic and auditory stimulation, the system is used to execute a child burn dressing change pain management method based on haptic and auditory stimulation; the system comprises: a signal acquisition module for acquiring real-time voice signals and physiological signals of children, including EEG signals and oxygenated hemoglobin concentration data; a feature extraction module for extracting a multi-dimensional MFCC coefficient vector and a fundamental frequency trajectory from the voice signals to generate an emotional feature vector, and extracting a theta band power spectral density and an oxygenated hemoglobin concentration change from the physiological signals to generate an electroencephalogram feature; a feature fusion module for fusing the emotional feature vector and the electroencephalogram feature into a unified fusion feature; a reinforcement learning model module including an Actor network and a Critic network, for generating a haptic stimulus parameter action vector and an auditory stimulus parameter action vector based on the fusion feature; a stimulus output module for outputting customized haptic stimulation and auditory stimulation according to the action vector; a feedback monitoring module for real-time monitoring of patient heart rate variability, pain performance indicators, comfort scores, and stimulus avoidance tendencies; a model updating module for dynamically optimizing the parameters of the reinforcement learning model through a time difference error and a policy gradient algorithm based on the data of the feedback monitoring module; and a control interface module for receiving user inputted treatment requirements and boundary value settings.
[0057] It should be noted that the child burn dressing change pain management system based on tactile and auditory stimulation of the present embodiment is used to execute the child burn dressing change pain management method based on tactile and auditory stimulation in embodiment 1, the signal acquisition module captures the child's voice and physiological response in real time through the microphone array and physiological signal sensor, ensures the accuracy and real-time of the data; the feature extraction module adopts signal processing algorithm, such as the mel frequency cepstral coefficient (MFCC) analysis in embodiment 1, extracts the emotional features in the voice, and simultaneously processes the physiological signals by using fast fourier transform (FFT) and power spectral density estimation method, generates the features reflecting the changes of brain activity and blood oxygen level; the feature fusion module adopts deep learning model, such as convolutional neural network (CNN) or long short-term memory network (LSTM), fuses multi-dimensional features, improves the accuracy of pain assessment; the reinforcement learning model module integrates Actor-Critic architecture, uses deep deterministic policy gradient (DDPG) or proximal policy optimization (PPO) algorithm, dynamically generates tactile and auditory stimulation parameters according to the fused features, realizes personalized pain management; the stimulation output module includes servo motor driven tactile stimulator and digital audio synthesizer, outputs customized tactile and auditory stimulation according to the generated stimulation parameters, such as vibration frequency and tone change; the feedback monitoring module monitors the patient's pain response in real time through the heart rate monitor, facial expression recognition system and comfort rating scale; the model updating module integrates the open source reinforcement learning framework, such as StableBaselines3, dynamically optimizes the parameters of the reinforcement learning model through the time difference error and policy gradient algorithm, realizes the continuous optimization of the pain management strategy; the control interface module provides a user-friendly graphical user interface (GUI), supports therapists to input treatment requirements and set the boundary values of stimulation parameters, ensures the safety and effectiveness of treatment.
[0058] Although the embodiments of the present application have been shown and described, those skilled in the art can understand that various changes, modifications, replacements and variations can be made to these embodiments without departing from the principles and purposes of the present application, and the scope of the present application is defined by the claims and their equivalents.
Claims
1. A method for managing pain in children with burn dressing change based on tactile and auditory stimulation, characterized by, The method comprises: S1, constructing a personalized pain management model based on patient basic physiological data; S2, collecting multi-modal data of the patient, including pressure / temperature data of the wound area, biological signals, and neural activity data; S3, extracting features from the multi-modal data based on the personalized pain management model and performing data fusion to obtain fusion features; S4, generating haptic stimulation parameters and auditory stimulation parameters using a reinforcement learning model based on the fusion features; S5, applying the generated stimulation parameters to the haptic device and the auditory device, and synchronously applying personalized stimulation during dressing change; S6, monitoring patient pain feedback data in real time, dynamically updating the reinforcement learning model parameters through the personalized pain management model, and iteratively optimizing the haptic stimulation parameters and the auditory stimulation parameters.
2. The method of claim 1, wherein the method further comprises: Constructing a personalized pain management model in step S1 comprises: S101, collecting patient basic physiological data, including age, gender, weight, height, total burn surface area percentage, burn depth, and past pain response history; S102, preprocessing the basic physiological data; S103, applying a reinforcement learning model to train a personalized pain management model based on the preprocessed data and incorporating patient-specific features; S104, verifying the performance of the personalized pain management model through cross-validation or a clinical historical data set, evaluating indicators including accuracy, recall rate, and F1 score, and dynamically adjusting model hyperparameters according to the verification results.
3. The method of claim 1, wherein the method further comprises: In step S3, extracting features from the multi-modal data based on the personalized pain management model and performing data fusion to obtain fusion features comprises: S301, using time-frequency analysis method to extract time domain features and frequency domain features from the pressure / temperature data of the wound area; S302, performing speech emotion analysis on the biological signals to obtain an emotion feature vector; S303, using electroencephalogram analysis method to extract electroencephalogram features from the neural activity data; S304, based on the personalized pain management model, performing multi-modal data fusion on the time domain features, frequency domain features, emotion feature vector, and electroencephalogram features to obtain the fusion features.
4. The method of claim 3, wherein the tactile and audible stimuli are applied to the child in a manner that is different from the manner in which the stimuli are applied to the child during the first dressing change. In S301, the expression of the extracted time domain features and frequency domain features is: ; ; wherein, represents the mean value of pressure, represents the number of samples, represents the instantaneous pressure value at time represents the peak value of pressure, represents the continuous wavelet transform, represents the wavelet basis function, represents the scale, represents the translation, represents the dominant frequency component, represents the fast Fourier transform; In step S302, the expression of the obtained emotion feature vector is: , ; in, Represents the multidimensional MFCC coefficient vector. Represents the Mel frequency cepstral coefficients. Indicates sampling point The voice signal, Indicates the fundamental frequency trajectory; In step S303, the expression of the extracted electroencephalogram features is: , ; wherein, denotes band power spectral density, denotes the number of samples, denotes time instant band EEG signal, denotes the change in oxyhemoglobin concentration, denotes oxyhemoglobin concentration at time instant, denotes oxyhemoglobin concentration at time instant.
5. The haptic and auditory stimulus based pain management method for children with burn dressing change according to claim 1, wherein, In step S4, generating haptic stimulation parameters and auditory stimulation parameters using a reinforcement learning model based on the fusion features comprises: S401, constructing the reinforcement learning model based on Actor network and Critic network; S402, mapping the fusion features to the state space representation of the reinforcement learning model, and extracting key physiological indicators as state variables; S403, setting boundary values for haptic stimulation parameters and auditory stimulation parameters according to child physiological tolerance and treatment needs, forming a continuous action space; S404, constructing a multi-objective reward function based on the real-time monitoring of the patient's pain feedback data and the patient's comfort score, giving a reward when the pain feedback data indicates an improvement in physiological indicators or the patient has no avoidance action, and imposing a penalty when the patient shows signs of distress, and introducing an adjustable coefficient to balance the analgesic effect and stimulation tolerance; S405, generating an action vector of the tactile stimulation parameters and an action vector of the auditory stimulation parameters using the Actor network based on the state variables, the continuous action space, and the multi-objective reward function.
6. The haptic and auditory stimulus-based pain management method for children with burn dressing change according to claim 5, wherein, In step S404, the expression of the multi-objective reward function is: ; wherein represents a total reward value, represents an adjustable weight coefficient of an improvement amount of heart rate variability, represents an improvement amount of heart rate variability, represents an adjustable weight coefficient of a pain expression marker, represents a pain expression marker, represents an adjustable weight coefficient of a comfort score, represents a comfort score, represents an adjustable weight coefficient of a stimulus avoidance tendency, represents a stimulus avoidance tendency.
7. The haptic and auditory stimulus based pain management method for children with burn dressing change as claimed in claim 5 wherein, In step S404, the expression of the action vector of the tactile stimulation parameters is: ; ; ; wherein, denotes a haptic stimulus parameter action vector, denotes a mean function of a haptic stimulus parameter, denotes a standard deviation function of a haptic stimulus parameter, denotes a random noise vector, denotes a hyperbolic tangent function, denotes an Actor network weight matrix corresponding to the mean function of the haptic stimulus parameter, denotes a state variable, denotes an Actor network bias vector corresponding to the mean function of the haptic stimulus parameter, denotes a haptic parameter scaling factor vector, denotes a Softplus function, denotes an Actor network weight matrix corresponding to the standard deviation function of the haptic stimulus parameter, denotes an Actor network bias vector corresponding to the standard deviation function of the haptic stimulus parameter; In step S404, the expression of the action vector of the auditory stimulation parameters is: ; ; ; wherein, an action vector representing an auditory stimulus parameter, a mean function representing an auditory stimulus parameter, a standard deviation function representing an auditory stimulus parameter, an Actor network weight matrix corresponding to the auditory mean function, an Actor network bias vector corresponding to the auditory mean function, an auditory parameter scaling factor vector, an Actor network weight matrix corresponding to the auditory standard deviation function, an Actor network bias vector corresponding to the auditory standard deviation function.
8. The haptic and auditory stimulus-based pain management method for children with burn dressing change according to claim 5, wherein, In step S6, real-time monitoring of the patient's pain feedback data, dynamic updating of the reinforcement learning model parameters through the personalized pain management model, and iterative optimization of the tactile stimulation parameters and the auditory stimulation parameters include: S601, evaluating the Q value of the action vector in the state space of the reinforcement learning model using the Critic network, and calculating the temporal difference error according to the multi-objective reward function; S602, updating the Actor network parameters based on the temporal difference error through the policy gradient algorithm, optimizing the action vector in the direction of maximizing the cumulative reward, and obtaining the optimized action vectors of the tactile stimulation parameters and the auditory stimulation parameters.
9. The method for managing pain in children with burns during dressing change based on tactile and auditory stimuli according to claim 8, wherein, In step S601, the expression of the temporal difference error is: ; ; wherein, represents the state-action pair value function output by the Critic network, represents the state variable, represents the current action vector, represents the set of parameters of the Critic network, represents the Critic network weight matrix, represents the vector concatenation operation, represents the Critic network bias vector, represents the time-difference error, represents the multi-objective reward function, represents the discount factor, represents the next time step state variable, represents the next time step action vector; In step S602, the expression of the policy gradient algorithm is: ; ; wherein, denotes a set of parameters of a tactile stimulus parameter actor network, denotes a learning rate, denotes a gradient operator on the tactile actor network parameters, denotes a log probability density of a tactile stimulus parameter action vector, denotes a tactile stimulus parameter action vector, denotes a set of parameters of an auditory stimulus parameter actor network, denotes a gradient operator on the auditory actor network parameters, denotes a log probability density of an auditory stimulus parameter action vector, denotes an auditory stimulus parameter action vector.
10. A child burn dressing change pain management system based on tactile and auditory stimulation, characterized in that, The system is used to perform the method of burn dressing pain management for children based on tactile and auditory stimulation as claimed in any one of claims 1 to 9; the system comprises: a signal acquisition module for real-time acquisition of the child's voice signal and physiological signal, including EEG signal and oxygenated hemoglobin concentration data; a feature extraction module for extracting a multi-dimensional MFCC coefficient vector and a fundamental frequency trajectory from the voice signal to generate an emotional feature vector, and extracting a theta band power spectral density and oxygenated hemoglobin concentration change from the physiological signal to generate an electroencephalogram feature; a feature fusion module for fusing the emotional feature vector and the electroencephalogram feature into a unified fusion feature; a reinforcement learning model module including an Actor network and a Critic network, for generating an action vector of tactile stimulation parameters and an action vector of auditory stimulation parameters based on the fusion feature; a stimulation output module for outputting customized tactile stimulation and auditory stimulation according to the action vectors; a feedback monitoring module for real-time monitoring of the patient's heart rate variability, signs of distress, comfort score, and stimulus avoidance tendency; a model updating module for dynamically optimizing the parameters of the reinforcement learning model through the temporal difference error and the policy gradient algorithm based on the data of the feedback monitoring module; and a control interface module for receiving user input of treatment requirements and boundary value settings.
Citation Information
Patent Citations
Burn patient rehabilitation training intelligent guidance system based on deep learning
CN118918638A
Method for optimizing dysphagia training based on multimode stimulation
CN119296721A
Analgesia feedback and control method and system based on multi-modal data, electronic equipment and storage medium
CN120514963A
Immersive music experience through synchronized auditory and haptic feedback tailored to cognitive and emotional states
WO2024180549A1
Cited By
Analysis method of multi-modal data in motion scene, electronic equipment and storage medium
CN121812195A