Pain assessment system and method based on multi-modal physiological signals

By combining multimodal physiological signal synchronous acquisition with deep learning fusion technology, the subjectivity and environmental interference problems of existing pain assessment methods have been solved, enabling real-time and individualized assessment of pain in infants and young children, and improving the accuracy and applicability of pain management.

CN120899167APending Publication Date: 2025-11-07NANJING CHILDRENS HOSPITAL
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510918433.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing pain assessment methods rely on subjective observation, making continuous monitoring difficult. They lack multimodal information fusion, cannot accurately distinguish pain levels, and do not consider individual differences and clinical environment interference, resulting in poor pain management effectiveness.

Method used

A multimodal physiological signal synchronous acquisition and analysis system is adopted, which combines deep learning and multimodal information fusion technology to achieve objective quantitative assessment of pain level through modular data acquisition, deep feature extraction, cross-modal fusion and individualized calibration.

Benefits of technology

It enables multi-dimensional, real-time, and individualized assessment of pain in infants and young children, improving the accuracy and timeliness of pain management, reducing pain-related complications, and enhancing the robustness and applicability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120899167A_ABST
    Figure CN120899167A_ABST
Patent Text Reader

Abstract

The invention provides a pain assessment system and method based on multi-modal physiological signals. The system comprises a multi-modal signal acquisition module, a signal preprocessing module, a multi-modal feature extraction module, a deep fusion analysis module, an individualized calibration module and a result output and early warning module. By synchronously collecting and analyzing multi-dimensional data such as facial expressions, sound features, physiological signs and behavior responses and combining deep learning and multi-modal information fusion technologies, objective quantitative evaluation and real-time monitoring of the pain degree are achieved, and accurate decision support is provided for clinical pain management.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of multi-modal physiological signal detection, and in particular to a pain assessment system and method based on multi-modal physiological signals. BACKGROUND

[0002] Targeted pain management has long been a major challenge in the medical community. Due to the inability of the target to effectively express pain in language, its pain experience is often underestimated or ignored by medical staff and caregivers. A large number of studies have confirmed that the target's pain experience not only leads to immediate physiological and behavioral reactions, but also has long-term negative effects on its nervous system development, pain sensitivity, stress response patterns, and cognitive behavioral development. According to the International Association for the Study of Pain, babies in the neonatal intensive care unit (NICU) experience an average of 11-14 painful procedures per day, and only about 30% of these pains are properly managed.

[0003] Traditional target pain assessment mainly relies on scale scoring systems, such as the Neonatal Facial Coding System (NFCS), the Neonatal Infant Pain Scale (NIPS), the CRIES scale, and the N-PASS scale. These scales are mainly based on the subjective scoring of observers on infant facial expressions, body posture, crying characteristics, changes in physiological indicators, etc. However, such methods have many limitations: first, the scoring process relies on the experience and subjective judgment of the observer, and there are significant differences between different evaluators; second, the evaluation process is time-consuming and cannot be continuously monitored, making it difficult to discover intermittent and chronic pain in a timely manner; third, existing scales do not fully consider the specific response patterns of targets at different developmental stages, as well as the regulatory effects of disease states on pain expression.

[0004] In recent years, with the rapid development of artificial intelligence and multi-modal signal processing technology, objective pain assessment methods based on physiological signals and behavioral characteristics have gradually emerged. Studies have shown that pain causes a series of changes in physiological parameters, including increased heart rate, elevated blood pressure, changes in respiratory rate, decreased blood oxygen saturation, and pupil response. At the same time, pain also triggers specific behavioral responses and facial expression changes, such as characteristic crying, furrowed brows, and clenched lips. International research teams have begun to explore computer vision-based infant facial pain recognition systems and infant crying pain classification methods based on speech analysis. For example, a deep learning-based infant facial pain expression recognition system developed by a research team at the University of Toronto in Canada achieved an accuracy of over 80% in an experimental environment; a frequency spectrum analysis-based infant crying pain classification algorithm developed by researchers at Stanford University in the United States can distinguish between painful crying and hunger crying.

[0005] However, the prior art still has obvious deficiencies. On the one hand, existing systems focus on single-modal signal analysis, such as only analyzing facial expressions or only analyzing crying features, and fail to fully utilize the complementarity of multi-modal information, resulting in insufficient robustness in complex clinical environments. On the other hand, existing algorithms generally lack precise quantification of pain intensity, making it difficult to distinguish between mild, moderate and severe pain. In addition, existing systems are mostly laboratory prototypes that fail to fully consider the complexity and specificity requirements of clinical application scenarios, such as medical device interference in NICU environments, infant body position changes, and various pipeline obstructions. Finally, existing methods lack individualized design and fail to consider the differences in target pain response patterns at different developmental stages and different disease states.

[0006] Therefore, there is an urgent need for an intelligent system that can capture target pain response characteristics in multiple dimensions and comprehensively, and objectively quantify and accurately evaluate pain intensity, to assist medical personnel and caregivers in achieving precise and individualized target pain management, improving clinical efficacy and reducing short-term and long-term negative effects of pain. SUMMARY

[0007] To solve the problems of the prior art, the present application provides a pain assessment system and method based on multi-modal physiological signals, which synchronously collects and analyzes multi-dimensional data such as facial expressions, sound features, physiological signs and behavioral responses, combines deep learning and multi-modal information fusion technology, and realizes objective quantitative evaluation and real-time monitoring of pain intensity, providing precise decision support for clinical pain management.

[0008] The present application aims to solve the following core technical problems:

[0009] Real-time synchronous acquisition and alignment of multi-modal medical signals. Infant pain assessment requires simultaneous analysis of multi-dimensional data such as facial expressions, sound features, physiological indicators and behavioral responses. These signals have different sampling rates, signal-to-noise characteristics and time resolution. How to accurately collect and time-align multi-channel data without interfering with the normal activities of infants is an important challenge in the field of multi-modal physiological signal processing. The present application designs a modular data acquisition system and a distributed sensing network, which solves the problem of accurate alignment of heterogeneous signals through global clock synchronization and data pipeline processing, achieving a synchronization accuracy of microseconds.

[0010] Pain feature representation and quantification in weakly labeled environment. Unlike adult pain assessment, infant pain response is more complex and individual difference is significant, and clinical annotation data is scarce and subjective bias exists. How to build a robust pain representation model under the condition of limited labeled data is a typical weakly supervised learning problem. The invention automatically discovers pain-related feature patterns from unstructured physiological signals by designing a deep feature extraction network and a self-supervised learning strategy, and alleviates the problem of insufficient labeled data through transfer learning and data augmentation techniques. In particular, the invention innovatively defines a set of pain-related deep feature space, which can realize knowledge transfer between different clinical scenarios.

[0011] Semantic alignment and fusion of multi-modal medical information. Pain is a subjective experience, and its performance on different physiological signal channels has time delay and intensity inconsistency. How to understand the semantic association between different modal signals and effectively fuse these heterogeneous information is a key challenge in multi-modal learning. The invention realizes dynamic weight allocation and semantic alignment between different modal signals by designing a cross-modal fusion network based on attention mechanism, which can adaptively adjust the importance of different modalities according to clinical scenarios and individual differences, significantly improving the accuracy and generalization ability of evaluation.

[0012] Individualized pain assessment and development state adaptation. Infants are in a rapid development stage, and infants of different ages / gestational ages have significant differences in pain response patterns. At the same time, disease state and medication will also affect pain expression. How to design an evaluation system with development sensitivity and individual adaptability is a core problem in clinical application. The invention establishes a parameter adjustment framework based on developmental neuroscience, integrates prior medical knowledge and individual characteristics through Bayesian inference network, and realizes automatic calibration and individual adjustment of evaluation parameters.

[0013] The overall architecture of the system of the invention consists of six core functional modules, forming a complete closed-loop system from signal acquisition to pain assessment. The system mathematically constructs a nonlinear mapping function from the multi-dimensional physiological signal space to the pain intensity score space, solving the key technical problem of objective quantification of subjective experience. The overall architecture of the system is shown in Figure 1 .

[0014] The multi-modal signal acquisition module adopts distributed sensing technology and non-contact monitoring scheme to minimize the interference to the infants. The module includes four subsystems: (1) The facial image acquisition system uses a binocular vision system composed of a high-resolution RGB camera (1920x1080@30fps) and a depth camera (640x480@60fps), combined with active infrared light supplement technology, to achieve stable imaging under various lighting conditions; (2) The audio acquisition system uses a high-sensitivity MEMS microphone array (sensitivity -38dBV / Pa) to enhance the target sound through adaptive beamforming technology, with a sampling rate of 48kHz and a quantization accuracy of 24bit; (3) The physiological parameter monitoring system integrates with clinical monitoring devices through a wireless communication interface to obtain physiological indicators such as heart rate, blood pressure, respiration, and blood oxygen in real time, while using a dedicated wireless wearable sensor to supplement the collection of galvanic skin response (GSR) and body temperature changes; (4) The behavioral response capture system combines a pressure sensing mattress and a micro motion sensor to monitor the frequency and intensity of limb movements. All the above sensing devices are synchronized to the microsecond level through a unified time server, and the collected data is transmitted to the central processing unit through a low-latency network.

[0015] The signal preprocessing module designs dedicated signal purification algorithms for different types of physiological signals. For facial image sequences, the system first applies adaptive histogram equalization to enhance contrast, then uses Gaussian Mixture Model (GMM) for background separation, and finally smoothes the facial motion trajectory through a Kalman filter. For audio signals, the system applies adaptive noise cancellation algorithm and band-pass filter (300-8000Hz) to reduce environmental noise, then performs sound event detection through energy detection and zero-crossing rate analysis. For physiological parameter signals, the system uses wavelet transform to remove baseline drift and high-frequency interference, and marks and interpolates missing data points through an outlier detection algorithm. For behavioral activity data, the system applies Principal Component Analysis (PCA) for dimensionality reduction and motion artifact removal. The preprocessed multi-modal signals are time-aligned to establish a unified analysis time window and form a synchronized multi-channel data stream. The complete multi-modal signal processing flow is shown in Figure 2

[0016] ​The multimodal feature extraction module realizes the mapping from the preprocessed signals to the pain feature representation. In terms of facial feature extraction, the system first localizes 68 facial landmarks by an improved FaceNet model, and then calculates the activation intensity of 19 facial action units (AUs) based on FACS (Facial Action Coding System), focusing on analyzing the spatiotemporal variation patterns of the core AUs related to pain (such as AU4, AU6 / 7, AU9, AU10 / 11, and AU20 / 25 / 26 / 27). In terms of acoustic feature extraction, the system combines short-time Fourier transform (STFT) and continuous wavelet transform (CWT) to analyze the time-frequency features of the crying sound, extracts acoustic feature vectors including fundamental frequency trajectory, spectral energy distribution, harmonic structure, and formant features, and captures the timbre characteristics of the sound through mel-frequency cepstral coefficients (MFCC). In terms of physiological feature extraction, the system applies multi-scale entropy and wavelet transform to analyze the nonlinear dynamics of heart rate variability and respiratory signals, and calculates indicators such as blood pressure fluctuation characteristics and oxygen saturation drop amplitude. In terms of behavioral feature extraction, the system quantifies the limb movement pattern based on optical flow method and pose estimation technology, and calculates activity intensity index, body rigidity, and stress response characteristics.

[0017] These features together constitute a high-dimensional feature space which can be represented as:

[0018]

[0019] where Φ is the feature extraction function that maps the original signal s(t) to the feature space. To improve the discriminative performance of the feature representation, the system applies feature selection and nonlinear dimensionality reduction techniques, combining the maximum relevance minimum redundancy (mRMR) algorithm and t-SNE method to retain a subset of features with the highest pain recognition ability, and constructs a low-dimensional but high information density feature representation:

[0020]

[0021] where Ψ is the dimensionality reduction mapping function, is the reduced feature space.

[0022] The deep fusion analysis module is the core of the invention, which realizes the mapping from the multimodal features to the pain assessment results. This module uses a multi-stream attention network architecture, which includes four key components. The detailed network architecture is shown in Figure 3

[0023] The modality-specific feature extractor designs a dedicated deep neural network for the characteristics of each modality signal. The facial features use a 3D convolutional neural network (3D-CNN) to capture the temporal changes of expressions: ​

[0024]

[0025] where X face is the sequence of face images, θ f is the network parameter. The acoustic features model the temporal acoustic features using a long short-term memory network (LSTM):

[0026]

[0027] where X audio is the acoustic sequence, θ a is the network parameter, θ a contains the weight matrices W f ,W i ,W o ,W c and the bias vectors b f ,b i ,b o ,b c ; T is the time step, d audio is the acoustic feature dimension;

[0028] The input acoustic sequence, d input is the input feature dimension;

[0029] The physiological features extract the contextual features of the physiological signals using a bidirectional gated recurrent unit (BiGRU) network:

[0030]

[0031] The behavioral features capture multi-scale activity patterns using a temporal convolutional network (TCN):

[0032]

[0033] The cross-modal context resolver captures the temporal dependency between different modal signals through a self-attention mechanism. For each modal feature sequence F m , the internal temporal dependency is first computed through a self-attention mechanism:

[0034]

[0035] C m = A m V m

[0036] where d k is the dimension of the key vector used to scale the attention scores, Q m , K m and V mrespectively, are obtained by linear projection from F m Q m = F m W Q ,K m = F m W K ,V m = F m W V where W Q ,W K ,W V are learnable projection matrices.

[0037] Then the information flow between different modalities is calculated by cross-modal attention:

[0038] E i,j = W i,j [C i ; C j ]

[0039] G i,j = σ(E i,j )

[0040] where W i,j is the inter-modal mapping matrix, G i,j is the gating unit between modality i and modality j, and σ is the sigmoid activation function.

[0041] The attention-guided modality fusion layer dynamically adjusts the importance weight of each modality according to different clinical scenarios and signal quality. The system designs a meta-attention network to learn the attention weight assigned to each modality:

[0042]

[0043] where f a is the attention score function, q is the query vector representing the current evaluation context, and M is the total number of modalities. The weighted multi-modal feature fusion is represented as:

[0044]

[0045] The pain assessment reasoning engine generates a pain probability distribution and intensity score based on the fused features. The system uses a multi-task learning framework to simultaneously optimize two related tasks: pain classification and pain intensity regression:

[0046]

[0047] where is the pain classification probability (no pain / mild / moderate / severe), is the pain intensity score on a 0-10 scale.

[0048] Classification task related parameters: Weight matrix of pain classification task, where d fused is the dimension of fused features, n class is the number of classes of pain classification; Bias vector of pain classification task; Regression task related parameters: Weight matrix of pain intensity regression task, output a single pain intensity score; Bias scalar of pain intensity regression task;

[0049] wherein; Fused feature vector, Classification task weight matrix, Classification task bias vector, Regression task weight matrix, Regression task bias scalar, Pain classification probability distribution, 0-10 scale pain intensity score;

[0050] The joint loss function of multi-task learning is defined as:

[0051]

[0052] wherein is the cross-entropy loss, is the mean square error loss, is the regularization term, λ1, λ2 and λ3 are hyperparameters to balance different losses.

[0053] The individualized calibration module realizes the individualized adjustment of the evaluation results, and adjusts the pain evaluation parameters according to the development state and clinical characteristics of infants. This module uses the Bayesian network framework to integrate medical prior knowledge and individual characteristics into the evaluation process:

[0054]

[0055] wherein Pain is the pain state, Signals is the observed physiological signal, and Ind is the individual characteristic variable (such as age, gestational age, disease state, medication, etc.). The conditional probability distribution P(Signals|Pain, Ind) is estimated from the training data, and the prior probability P(Pain|Ind) is provided by the medical knowledge graph.

[0056] Specifically, the age / gestational age adaptation unit adjusts the assessment weight according to the development stage of the infant. The system divides the infant into six development stages: extremely preterm infants (<32 weeks gestational age), preterm infants (32-37 weeks gestational age), full-term newborns (0-1 month), early infancy (1-6 months), middle infancy (6-12 months), and toddlerhood (12-36 months), and sets specific modal weight and pain expression mode priors for each stage:

[0057]

[0058] where g a is the age adjustment function, and Age is the age / gestational age variable.

[0059] The disease impact compensation unit considers the modulating effect of specific diseases (such as respiratory distress syndrome, nervous system diseases, postoperative state, etc.) on pain expression, and adjusts the assessment sensitivity through a disease-symptom knowledge graph:

[0060] S adj =S orig ·f d (Disease)

[0061] where S adj is the adjusted pain score, S orig is the original score, and f d is the disease adjustment factor.

[0062] The drug effect evaluation unit analyzes the degree of inhibition of sedatives and analgesics on pain response. The system establishes a pharmacodynamic model of drugs, estimates the current drug concentration and its effect on pain expression according to the drug type, dose, and administration time:

[0063]

[0064] where I drug is the drug influence index, D is the number of drugs used simultaneously, C i (t) is the estimated concentration of drug i at time t, E i is the effect coefficient of drug i, and w i is the weight factor.

[0065] The baseline behavior feature library stores the baseline behavior features of the infant in a non-pain state, which is used to calculate the individualized feature deviation:

[0066] ΔF=F current -F baseline

[0067] where F current is the current feature vector, and F baselineis the stored baseline feature. The system uses this deviation rather than the absolute feature value for evaluation, effectively controlling the interference caused by individual differences.

[0068] The result output and early warning module provides an intuitive visual interface and intelligent early warning function. The hierarchical pain assessment display interface displays the pain intensity on a 0-10 scale and is divided into four levels: no pain (0-1), mild (2-3), moderate (4-6), and severe (7-10), with different color coding (green, yellow, orange, red) for intuitive display. The pain trend analysis chart shows the dynamic changes of the pain score within 24 hours, helping medical staff identify pain patterns and intervention effects. The intervention recommendation generator recommends individualized interventions based on evidence-based medical evidence and clinical guidelines, according to the pain assessment results, individual characteristics of infants and young children, and past treatment responses, including drug interventions (type, dosage, administration route) and non-drug interventions (body position adjustment, environmental control, soothing techniques, etc.). The pain threshold warning system triggers medical staff intervention through visual and audio reminders when the pain score exceeds the preset threshold (default is 6) or rises rapidly within a short period of time (increases by 3 points within 30 minutes).

[0069] In terms of system implementation, the present application adopts a layered heterogeneous computing architecture, which fully balances the real-time requirements and computational complexity. The front-end acquisition device uses a low-power microcontroller to implement signal acquisition and preliminary preprocessing, and transmits the data to the edge computing node through a low-latency wireless network. The edge node uses an ARM architecture processor combined with a neural network accelerator (NPU) to complete signal preprocessing and preliminary feature extraction. The cloud server uses a GPU cluster to run complex deep learning models to handle multi-modal fusion and inference tasks. The system uses a master-slave architecture to ensure that the front-end device can maintain basic functions even in the case of network fluctuations or disconnections.

[0070] In terms of model training and verification, the present application uses a multi-center clinical data set containing data from the NICU and PICU of three tertiary children's hospitals, a total of 500 infants and young children of different ages / gestational ages, covering 35 common procedural pain scenarios (such as blood sampling, lumbar puncture, catheter insertion, etc.) and various disease-related pains. Each case includes synchronized multi-modal data records and scores from experienced pain specialist nurses using standard scales (such as NIPS, CRIES, etc.), and some cases also include independent scores from multiple assessors for consistency analysis. The model uses five-fold cross-validation to evaluate performance, and on an independent test set, the system achieves a pain recognition accuracy of 92.7% and a pain intensity score Pearson correlation coefficient of 0.89, significantly better than single-modal methods and traditional machine learning algorithms. In particular, in the analysis of subgroups of infants and young children at different developmental stages, the system showed stable performance, proving the effectiveness of the individualized calibration module.

[0071] The present application realizes a clinical application-oriented system deployment scheme, including an integrated bedside monitoring system and a portable follow-up evaluation system. The integrated system is suitable for inpatient environments such as NICU / PICU, integrates all sensing devices around the baby bed or incubator, and obtains medical record information and medication data through a hospital information system (HIS) interface. The system performs automatic evaluation every 5 minutes and continuously monitors physiological parameter fluctuations to trigger additional evaluations. The system pushes the evaluation results to the nurse workstation and the doctor's mobile terminal through the hospital network, and integrates with the electronic medical record system to automatically record the pain evaluation results and interventions. The portable system is designed for home environment and outpatient follow-up, consisting of a simplified sensor kit and a smartphone application. Parents or caregivers can use the system to regularly or on-demand evaluate infants at home, and the results are transmitted to the cloud platform through the mobile network and can be selectively shared with medical professionals for remote consultation. The system deployment in different application scenarios is shown in Figure 4 .

[0072] The present application also realizes an incremental learning mechanism to ensure continuous optimization of system performance over time. The system regularly collects clinically verified evaluation data and compares it with the original evaluation results to update the model parameters through an online learning algorithm:

[0073]

[0074] where θ t is the current model parameter, is the newly collected data, and η is the learning rate. To avoid the problem of catastrophic forgetting, the system uses a flexible weight merging strategy:

[0075]

[0076] where F i is the diagonal element of the Fisher information matrix, reflecting the importance of the parameter, is the previously optimized parameter. This continuous learning mechanism enables the system to adapt to new clinical environments and patient populations, continuously improving evaluation accuracy.

[0077] To protect medical data security and patient privacy, the present application designs a multi-level data protection mechanism. All sensing data are encrypted at the collection end, and end-to-end encryption protocols are used during transmission. The system supports federated learning deployment mode, allowing medical institutions to collaborate in training evaluation models without sharing original patient data:

[0078]

[0079] where is the local loss function of the kth participant, n kis the number of samples, n is the total number of samples. Each participant only shares the model gradient instead of the original data. The facial image and sound data are immediately destroyed after feature extraction, and the system only saves the de-identified feature vector for analysis.

[0080] The present application has the beneficial effects of:

[0081] 1. Multimodal comprehensive evaluation capability: The present application innovatively integrates the information of facial expression, voice feature, physiological parameter and behavioral response in four dimensions, and constructs a comprehensive infant pain representation system. Compared with the existing method of analyzing only a single mode, the multimodal fusion significantly improves the robustness and accuracy of the system. In clinical verification, the multimodal system improves the accuracy by 15.3% and the correlation coefficient by 0.18 compared with the best single mode method, especially in the case of greater environmental interference or poor signal quality of a certain mode, the system can still maintain reliable evaluation performance.

[0082] 2. Individualized dynamic evaluation framework: The present application breaks through the limitations of the "one-size-fits-all" evaluation mode of the prior art, and establishes an individualized evaluation framework based on developmental neuroscience. The system can automatically adjust the evaluation parameters according to the infant's age / gestational age, disease state and medication, so that the evaluation results are more in line with the actual situation of the individual. Clinical verification shows that individualized calibration reduces the accuracy difference of infants at different developmental stages from 18.7% to 5.2%, significantly enhancing the universal applicability of the system.

[0083] 3. Continuous real-time monitoring capability: Unlike traditional scale scoring relying on manual observation, the present application realizes continuous real-time monitoring of pain, and can identify intermittent pain, chronic pain and dynamic changes in pain intensity. The system can issue an early warning within an average of 1.8 minutes after the occurrence of pain, about 15 minutes earlier than conventional nursing observation, creating a valuable time window for timely intervention. Long-term monitoring data also support pain pattern analysis and intervention effect evaluation, providing an objective basis for optimizing pain management strategies.

[0084] 4. Clinical decision support function: The present application not only provides pain evaluation results, but also integrates evidence-based intervention recommendations and warning mechanisms. The system can recommend individualized drug and non-drug intervention programs according to the individual characteristics of the infant and the pain evaluation results, and trigger an early warning in a timely manner when the pain intensifies. Clinical application shows that pain management assisted by the system reduces inappropriate analgesia / sedation by 37.8%, reduces excessive drug use by 28.3%, and improves the comfort score of the patient by 31.5%.

[0085] 5. Extensible learning architecture: The present application adopts an incremental learning mechanism, which can continuously learn and optimize from clinical practice. With the extension of use time and data accumulation, the system performance is continuously improved, and the evaluation accuracy is improved by 4.7% after 3 months of clinical application, and by 7.2% after 6 months. The system also supports model migration, which can quickly adapt and optimize according to the characteristics of different medical institutions and patient groups, reducing the deployment difficulty and adaptation cost.

[0086] In summary, the present application provides a pain intelligent assessment system and method based on multi-modal physiological signals, which realizes innovation breakthroughs in objective quantitative evaluation, individualized parameter adjustment, continuous real-time monitoring and clinical decision support. The system provides a new technical means for accurate assessment and management of pain by integrating multi-dimensional physiological signals and advanced algorithms, which is expected to significantly improve the quality of pain management, reduce pain-related complications, and promote the healthy development and quality of life of infants. The present application is not only suitable for NICU, PICU and pediatric wards of medical institutions at all levels, but also can be popularized to primary medical institutions and home care environment, and has broad application prospect and social benefit. BRIEF DESCRIPTION OF DRAWINGS

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.

[0088] Figure 1 is a system overall architecture diagram;

[0089] Figure 2 is a multi-modal signal processing flowchart;

[0090] Figure 3 is a deep fusion network architecture diagram;

[0091] Figure 4 is a system application scenario diagram. DETAILED DESCRIPTION

[0092] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0093] The application provides a pain intelligent evaluation system and method based on multi-modal physiological signals, which can realize objective quantitative evaluation and continuous monitoring of pain. The composition, working principle and application scenario of the application will be described in detail below in conjunction with specific embodiments. Please refer to Figures 1 to 4 Understand the system architecture and workflow.

[0094] Embodiment one: neonatal intensive care unit (NICU) application scenario

[0095] In the neonatal intensive care unit (NICU) application scenario, the system can be deployed to monitor the pain state of postoperative and critically ill newborns. The system is composed of a group of sensing devices, an edge computing unit and a central server at the physical level. The sensing devices include a high-definition camera assembly installed above the baby incubator, a directional microphone array placed at the bedside, a physiological parameter acquisition module wirelessly connected to the monitor, and a pressure sensor array laid under the mattress. Among them, the camera assembly is a binocular system composed of an RGB camera and a depth camera, the resolution of the RGB camera is 1920x1080, the frame rate is 30fps, the field of view is 83°, and the installation height is 30cm; the depth camera uses structure light technology, the resolution is 640x480, the frame rate is 60fps, the measurement range is 0.2-1.0m, and the accuracy is ±1mm. The microphone array is composed of 4 high-sensitivity MEMS microphones, with a sensitivity of-38dBV / Pa, a signal-to-noise ratio of 65dB, a frequency response of 20Hz-20kHz, a sampling rate of 48kHz, and a quantization accuracy of 24bit. The physiological parameter acquisition module acquires monitor data in real time through medical device communication protocols (such as HL7), including heart rate, blood pressure, respiratory rate, oxygen saturation and body temperature. The mattress pressure sensor array is composed of 64 thin film pressure sensors in an 8x8 grid, with a sensitivity of 0.1N / cm 2 , a sampling rate of 100Hz, and can detect infant body movement and posture changes.

[0096] These sensing devices are connected to the edge computing unit, which is composed of an industrial-grade ARM processor (8-core 2.2GHz) with a dedicated neural network accelerator, with 4GB RAM and 128GB storage built-in, running a customized Linux system. The edge unit is responsible for real-time data acquisition, preliminary preprocessing and feature extraction, and is connected to the central server through the hospital's internal gigabit Ethernet. The central server is configured with an Intel Xeon processor, 128GB RAM, an NVIDIA Tesla V100 GPU and a 4TB RAID storage array, responsible for complex algorithm processing, model training and result storage.

[0097] At the software level, the system workflow can be divided into five stages: data acquisition, signal preprocessing, feature extraction, fusion analysis, and result output. The system adopts a pipeline architecture, with modules running in parallel to minimize latency. Based on clinical requirements, the system is configured to perform a regular assessment every 5 minutes, while continuously monitoring heart rate, blood oxygen, and activity parameters, triggering additional assessments when these indicators change significantly (e.g., heart rate increases by >20%, activity intensity suddenly increases).

[0098] In the data acquisition stage, the system first confirms that all sensors are in normal state and synchronizes the global clock. For facial image acquisition, the system uses an adaptive exposure control algorithm to automatically adjust camera parameters based on ambient light intensity and uses infrared fill light to ensure imaging quality during NICU night mode. The system uses an advanced infant face detection algorithm that combines RGB and depth information to construct a robust face tracking mechanism, which can adapt to frequent head movements and partial occlusions (e.g., respiratory tubing, monitoring patches). The system captures 5 frames of facial images per second for analysis, while recording continuous 30-second audio. Physiological parameters are collected at a frequency of 1 Hz, including heart rate, blood pressure, and blood oxygen data, as well as mattress pressure distribution maps. All data streams are labeled with global timestamps to ensure accurate alignment of multi-modal data.

[0099] In the signal preprocessing stage, the system applies specialized processing algorithms for different types of signals. For facial images, the system first applies Gaussian filtering (σ = 1.5) to reduce noise, then enhances image contrast through adaptive histogram equalization:

[0100] p out (i)=p in (C·F cum (i))

[0101] where p in and p out are the pixel values of input and output images, F cum is the cumulative histogram function, and C is the contrast limiting factor. The system uses an improved Viola-Jones algorithm and a deformable part model (DPM) to accurately locate the facial region, then applies an affine transformation to normalize the face to a unified coordinate system. To handle partial occlusion issues, the system implements an occlusion detection and recovery algorithm based on depth information, which can identify occlusions caused by medical devices (e.g., nasal cannula, endotracheal tube, monitoring electrodes) and compensate for them in subsequent analysis.

[0102] For audio signals, the system first applies an adaptive noise cancellation algorithm to reduce NICU environmental noise (e.g., monitor alarm sounds, ventilator sounds). This algorithm is based on the least mean square error (LMS) criterion, and the weight update formula for the adaptive filter is:

[0103] wn+1 = w n + μe n x n

[0104] where w n is the filter coefficient vector at the n-th iteration, μ is the step size parameter, e n is the estimation error, x n is the input signal vector. The system then applies a band-pass filter between 300-8000 Hz, covering the main spectral range of infant cries. By short-time energy analysis and zero-crossing rate computation, the system automatically detects cry segments and separates them from ambient sounds:

[0105]

[0106] Segments with short-time energy exceeding an adaptive threshold and zero-crossing rate within a certain range are labeled as valid cries.

[0107] For physiological parameter signals, the system applies wavelet transform to remove baseline drift and high-frequency interference. Taking heart rate variability (HRV) analysis as an example, the system first extracts the R-R interval sequence from electrocardiogram or pulse waveform, then obtains an equidistant heart rate time series by interpolation and resampling, and finally applies Daubechies wavelet (db4) for multi-scale decomposition to extract HRV features in different frequency bands. The system implements a missing data detection and interpolation algorithm, which can handle data loss caused by temporary failure of monitoring devices or poor signal quality.

[0108] For pressure sensing data, the system smooths the original pressure distribution map by two-dimensional Gaussian filtering, and then applies a background subtraction algorithm to separate the pressure changes caused by infant activity and the baseline pressure distribution. The system computes the pressure change integral between consecutive frames to quantify the overall activity level of the infant:

[0109]

[0110] where P t (i,j) is the pressure value at position (i,j) at time t, M and N are the dimensions of the sensor array.

[0111] In the feature extraction stage, the system extracts pain-related feature vectors from the preprocessed multi-modal signals. Face feature extraction first locates 68 facial key points through an improved FaceNet model to form a facial geometric representation. Then the system calculates the activation intensity of facial action unit based on FACS, focusing on the core AUs related to pain: AU4 (brow furrow): I AU4 = f(d innerbrow , θ brow )

[0112] AU6 / 7 (Eyelid Tightening): Calculated based on periorbital muscle condition and eye opening / closing degree. AU6 / 7 =g(A eye C eyelid AU9 (nasal wrinkles): Measures the degree of nasal alar bulge and nasal bridge wrinkles. AU9 =h(H nasal W nostril )

[0113] AU10 / 11 (Deepening of the nasolabial folds): Analysis of texture changes and geometric deformations in the nasolabial region I AU10 / 11 =k(d nasolabial ,T fold AU20 / 25 / 26 / 27 (Mouth Movements Collection): Comprehensive measurement of mouth opening and closing, lip shape changes, and degree of extension. AU20+ =m(A mouth ,d lip ,θ corner )

[0114] The system not only calculates static activation strengths but also analyzes their dynamic characteristics, including activation duration, rate of change, and AU co-occurrence patterns. The system constructs a temporal AU feature vector F. AU =[I AU1 ,I AU2 ,...,I AU43 ] t=1:T The system captures facial expression evolution patterns across a series of frames. Furthermore, it employs a deep convolutional neural network to directly extract high-level feature representations from facial images, using a pre-trained VGG-Face2 model fine-tuned for infant facial features to extract 4096-dimensional facial depth features.

[0115] Acoustic feature extraction combines time-domain and frequency-domain analysis to comprehensively capture the acoustic characteristics of crying sounds. In the time domain, the system calculates the signal energy profile, zero-crossing rate, duration pattern, and energy abrupt change points. In the frequency domain, the system applies short-time Fourier transform (25ms frame length, 10ms frame shift) to calculate the power spectral density and extracts a series of spectral parameters.

[0116] Spectral centroid: The "brightness" of a sound;

[0117] Spectrum bandwidth: Describes the degree of dispersion in the spectrum;

[0118] Spectral roll-off: Indicates the concentration of spectral energy distribution, such as the frequency corresponding to 85% of the energy;

[0119] Noise ratio: Reflects the ratio of periodic components to noise components in sound;

[0120] The system focuses on the fundamental frequency (F0) characteristics of the cries, estimating the F0 trajectory using autocorrelation methods:

[0121]

[0122] The system extracts statistical features of the F0, including mean, standard deviation, range, contour shape, etc., which are highly correlated with the pain state. In addition, the system computes the Mel-frequency cepstral coefficients (MFCCs), capturing the vocal tract resonance characteristics:

[0123]

[0124] where S m is the Mel-spectral energy. The system extracts 13 base MFCC coefficients and their first and second order differences, totaling 39-dimensional MFCC feature vectors.

[0125] Physiological parameter feature extraction focuses on heart rate variability (HRV), respiratory variability, and blood oxygenation dynamics. HRV analysis includes time-domain indices (SDNN, RMSSD, etc.), frequency-domain indices (LF, HF, LF / HF ratio, etc.), and nonlinear indices (sample entropy, approximate entropy, Poincaré plot analysis, etc.). Taking sample entropy as an example, its calculation formula is:

[0126]

[0127] where A m (r) and B m (r) are the number of matched pattern pairs under pattern dimensions m+1 and m, respectively. An increase in SampEn value is generally associated with the pain state. The system also analyzes indices such as respiratory frequency variability, blood pressure fluctuation characteristics, and skin electrical activity (if available), to build a comprehensive set of physiological response features.

[0128] Behavioral feature extraction is based on pressure distribution data and video motion analysis, quantifying the infant's activity patterns. The system calculates activity intensity indices (statistical features of the aforementioned activity scores), activity spectral features (periodicity of activity analyzed through Fourier transform), posture change frequency, and body rigidity, etc. The system applies optical flow methods to analyze motion patterns in the video, calculating statistical features of the motion vector field to distinguish between normal activity and stress-induced activity patterns.

[0129] To improve the discriminative performance of feature representation, the system applies feature selection and dimensionality reduction techniques. The maximum relevance minimum redundancy (mRMR) algorithm is used to select a feature subset with maximum information and minimum redundancy:

[0130]

[0131] where I(i,C) is the mutual information of feature i and class C, and I(i,j) is the mutual information between feature i and feature j. The system then applies principal component analysis (PCA) and t-SNE for nonlinear dimension reduction, retaining the most discriminative feature combinations.

[0132] In the deep fusion analysis stage, the system employs a multi-stream attention network to integrate multi-modal features and output the pain assessment results. This architecture contains four parallel modal-specific feature extraction branches, each optimized for a specific modality signal: the facial feature branch uses 3D-CNN to capture the spatio-temporal changes of expressions: The acoustic feature branch uses LSTM to model the temporal features of sounds: The physiological feature branch uses a BiGRU network to process multi-parameter physiological signals: The behavioral feature branch uses a TCN network to analyze multi-scale activity features:

[0133] Each branch first models the temporal dependency through self-attention mechanisms, and then the features are integrated through a cross-modal attention network. This network dynamically assesses the reliability and importance of each modality and assigns corresponding weights:

[0134]

[0135] where a m is the attention weight assigned to modality m. In the NICU environment, the system can be configured to give facial expressions and sound features higher weights (about 30-40%) under normal circumstances, physiological parameters next (20-30%), and behavioral features the lowest weight (10-20%); but in special circumstances (such as when the face is heavily obscured or the baby is sedated), the system can automatically adjust the weight distribution, increasing the weight of reliable modalities.

[0136] The fused features pass through a fully connected layer and softmax / sigmoid activation functions, simultaneously outputting pain classification results (no pain / mild / moderate / severe) and 0-10 scale pain intensity scores. The system employs a multi-task learning framework to optimize these two related tasks:

[0137]

[0138] The system can further consider the individual characteristics of the infant. Newborns are divided into three groups according to gestational age: extremely preterm infants (<32 weeks), preterm infants (32-37 weeks), and full-term infants (>37 weeks), and each group uses adjusted evaluation parameters. The system also adjusts the evaluation threshold according to the disease state (such as respiratory distress syndrome, necrotizing enterocolitis, postoperative state, etc.) and medication (such as opioid drugs, non-opioid analgesics, and sedatives). In particular, for infants receiving sedation therapy, the system increases the weight of physiological parameters and reduces the weight of facial expressions and behavioral characteristics to compensate for the inhibitory effect of drugs on overt behavior.

[0139] The system can establish a personalized baseline behavior profile for each infant, recording its reference characteristics in a non-pain state, and subsequent evaluations are based on the degree of deviation from the baseline rather than absolute characteristic values. This individualized comparison can significantly improve evaluation accuracy, especially for infants with neurological abnormalities or chronic diseases.

[0140] In the result output and warning stage, the system provides an intuitive visualization interface. The main interface displays the current pain score (0-10 scale) and pain level (no pain / mild / moderate / severe) in the form of a dashboard and uses different color coding (green / yellow / orange / red). The interface displays a 24-hour pain trend chart, marking the time points of drug intervention and operation events, helping medical staff analyze pain patterns and intervention effects. The system also generates detailed analysis reports, including the results of each modality feature analysis, the degree of deviation from the baseline of key indicators, and the main factors leading to the current score. For chronic pain management, the system provides long-term trend analysis functions to identify periodic patterns and long-term evolution trends of pain through time series analysis.

[0141] When the evaluation results indicate the presence of moderate to severe pain (score ≥4) or a rapid increase in pain score over a short period of time (increase ≥3 points within 30 minutes), the system can automatically trigger a warning, notifying medical staff through visual and audio reminders. Based on the evaluation results and individual characteristics of the infant, the system generates intervention suggestions, including drug intervention (recommended drug type, dosage, and administration regimen) and non-drug intervention (such as posture adjustment, environmental control, and soothing techniques). These suggestions can follow evidence-based medicine principles and clinical practice guidelines, while considering the individual characteristics of the patient and their past treatment response.

[0142] Through this embodiment, the system can achieve objective and quantitative assessment and continuous monitoring of infant pain in the NICU environment, improving the accuracy and timeliness of pain recognition and providing precise decision support for clinical pain management.

[0143] Embodiment Two: Community Medical and Family Application Scenarios

[0144] In community medical institutions and home environments, the present application can realize a set of lightweight infant pain assessment system. The system is composed of a portable assessment terminal and a smartphone application, suitable for outpatient follow-up and home care scenarios. The portable terminal integrates a small high-definition camera (1080p), a directional microphone, and a simplified physiological parameter monitoring device (portable pulse oximeter), connected to the smartphone through Bluetooth. The smartphone application guides the user to correctly place the sensor and complete the 30-60 second data collection process, then performs local preliminary analysis through a lightweight neural network model, and the results are uploaded to the cloud for more detailed evaluation.

[0145] To adapt to the computing resource limitations of mobile terminals, the system uses model compression and knowledge distillation techniques to transfer the capabilities of the full model to the lightweight network:

[0146]

[0147] where P T and P S are the output distributions of the teacher model (full model) and student model (lightweight model), respectively. The lightweight model reduces the model size to 15% of the original while maintaining 87% of the original performance, reducing the computational complexity by about 90%, enabling real-time operation on mid-end smartphones.

[0148] The home version system provides caregivers with a simplified interface and detailed usage guidance. The assessment results are displayed in a simple way (no pain / mild discomfort / obvious pain), and intervention suggestions suitable for non-professionals are provided. The system also contains educational content to help parents understand infant pain expression and basic relief techniques. Parents can choose to share the assessment results through a secure connection to medical professionals, supporting remote consultation and treatment plan adjustment.

[0149] The portable system can be configured for outpatient follow-up and home care programs. After the system completes the assessment, it can identify pain events that require medical intervention and remind caregivers in a timely manner. Through this early detection mechanism, many diseases (such as acute otitis media, urinary tract infection, etc.) can be avoided. Delayed diagnosis and unnecessary complications.

[0150] Another important feature of the system is the incremental learning capability, which continuously optimizes the assessment model through clinical feedback. The system regularly compares the assessment results with clinical judgments, collects significant differences for analysis, and updates the model parameters. To protect patient privacy, the system uses a federated learning framework, allowing medical institutions to collaborate to improve the model without sharing raw patient data:

[0151]

[0152] where is the local loss function of the kth participant, n k is the number of its samples, and n is the total number of samples. Each participant only shares the model gradient instead of the raw data.

[0153] In addition, the system is equipped with perfect data management and security mechanisms. All patient data are strictly encrypted and stored, and multiple authentication is required for access. Facial images and voice recordings are encrypted or destroyed immediately after feature extraction, and the system only saves de-identified feature vectors for analysis. For clinical research needs, the system supports anonymous data export, helping medical institutions to carry out pain management quality improvement projects and clinical research.

[0154] Through this embodiment, the present application proves that the system is not only suitable for advanced medical institution environment, but also can play an important role in primary medical and family scenes, providing comprehensive technical support for accurate assessment and timely intervention of infant pain. The modular design and scalable architecture of the system make it have good adaptability and popularization value, and can meet the multi-level needs from neonatal intensive care to family daily care, making important contributions to improve the quality of infant pain management and life quality. The complete system application scenarios are shown in Figure 4 .

[0155] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device embodiment, the above description is only the preferred embodiment of the present application, and since it is basically similar to the method embodiment, it is described more simply. The relevant part can be referred to the part of the method embodiment. The above description is only the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any change or replacement that can be easily thought of by those skilled in the art within the technical range disclosed by the present application, and without departing from the principle of the present application, should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-modal physiological signal based pain assessment system, characterized by: The system comprises a multi-modal signal acquisition module for synchronously acquiring a face image sequence, a sound signal, a physiological parameter, and a behavior reaction data of a target; a signal preprocessing module for pre-processing the acquired original signals by noise reduction, filtering, segmentation, feature point marking, and data calibration; a multi-modal feature extraction module for extracting a feature vector related to pain from the pre-processed signals of various types; a deep fusion analysis module for combining multi-modal feature information and calculating a pain probability and a pain intensity score through a deep learning algorithm; an individualized calibration module for adjusting the evaluation parameters according to the state of a user; a result output and early warning module for displaying a pain evaluation result, recommending intervention measures, and triggering an early warning mechanism.

2. The multi-modal physiological signal based pain assessment system as claimed in claim 1, wherein: The multi-modal signal acquisition module comprises a non-contact face image acquisition unit adopting a binocular camera system combined with a high-resolution RGB camera and a depth camera to capture the facial expression changes of the target; an audio acquisition unit adopting a high-sensitivity directional microphone array to record the crying and sound features of the target; a physiological parameter monitoring unit for communicating with medical equipment through wireless connection to acquire heart rate, blood pressure, respiratory rate, blood oxygen saturation, and skin electric reaction physiological indicators; and a behavior reaction capturing unit for recording the limb activity, body posture changes, and stress reaction through a pressure sensor and a motion capture system.

3. The multi-modal physiological signal based pain assessment system as claimed in claim 1, wherein, The multi-modal feature extraction module comprises a face feature extraction unit extracting facial action unit features including eyebrow frown degree, eyelid tightening degree, nasolabial sulcus deepening degree, and mouth corner stretching degree; an acoustic feature extraction unit extracting acoustic features including fundamental frequency trajectory, spectral energy distribution, harmonic structure, sound duration, and prosody features; a physiological feature extraction unit extracting physiological features including heart rate variability, respiratory rate change rate, blood pressure fluctuation features, and blood oxygen saturation drop amplitude; and a behavior feature extraction unit extracting behavior features including limb activity index, body rigidity, struggle reaction intensity, and pacification difficulty.

4. The multi-modal physiological signal based pain assessment system as claimed in claim 1, wherein: The deep fusion analysis module adopts a multi-flow attention network architecture, comprising a modality-specific feature extractor, a cross-modality context parser, an attention-guided modality fusion layer, and a pain assessment reasoning engine.

5. The multi-modal physiological signal based pain assessment system as claimed in claim 1, wherein, The individualized calibration module comprises an age or gestational age adaptation unit, a disease influence compensation unit, a drug effect evaluation unit, and a baseline behavior feature library.

6. The multi-modal physiological signal based pain assessment system as claimed in claim 1, wherein, The result output and early warning module is configured with a graded pain assessment display interface, a pain intensity display interface in a 0-10 point system and divided into four levels of no pain, mild, moderate, and severe; a pain trend analysis chart showing the change trend of the pain intensity over time; and An intervention recommendation generator recommends drug and non-drug interventions according to the pain assessment results and the target individual characteristics; A pain threshold warning system triggers a warning when the pain intensity exceeds a preset threshold or rapidly rises in a short time.

7. A method for pain assessment based on multi-modal physiological signals, characterized in that, The method comprises the following steps: 1) synchronously collecting a facial image sequence, a sound signal, a physiological parameter, and a behavioral response data; 2) preprocessing the collected original signals, and performing time alignment on the preprocessed multi-modal signals to establish a unified analysis time window and form a synchronous multi-channel data stream; 3) extracting a feature vector related to pain from each type of preprocessed signal; These The features collectively form a high-dimensional feature space Can be expressed as: wherein Φ is a feature extraction function that maps the original signal s(t) to a feature space, and a feature selection and nonlinear dimension reduction technique is applied, combined with the maximum relevance minimum redundancy mRMR algorithm and the t-SNE method, to retain a feature subset with the highest pain recognition ability, and to construct a low-dimensional but high information density feature representation; where Ψ is a dimensionality reduction mapping function, is the dimensionality reduced feature space; 4) combining multi-modal feature information, performing deep fusion analysis, and calculating pain probability and pain intensity score through a deep learning algorithm; the deep fusion analysis adopts a multi-stream attention network architecture, comprising the following processes: 4.1) a modality-specific feature extractor designs a dedicated deep neural network for the characteristics of each modality signal, 4.11) a 3D convolutional neural network 3D-CNN is used to capture the temporal changes of facial features: where X face is a sequence of face images, θ f is a network parameter; 4.12) a long short-term memory network LSTM is used to model the temporal acoustic features: where X audio is an acoustic sequence, θ a is a network parameter, θ a contains the weight matrices W f , W i , W o , W c and bias vectors b f , b i , b o , b c ; T is the time step, d audio is the acoustic feature dimension; input acoustic sequence, d input is the input feature dimension; 4.13) a bidirectional gated recurrent unit BiGRU network is used to extract the context features of physiological signals: where X phys is a physiological sequence, θ p is a network parameter; 4.41) a time convolution network TCN is used to capture multi-scale activity patterns: where X behav is a sequence of actions, θ b is a network parameter; 4.2) Cross-modal context parser captures temporal dependencies between different modal signals by self-attention mechanism, for each modal's feature sequence F m Firstly, compute the intra-temporal dependencies by self-attention mechanism: C m = A m V m where d k is the dimension of the key vector used to scale the attention scores, Q m , K m and V m are the query, key and value matrices, obtained by linear projections from F m ; Q m = F m W Q , K m = F m W K , V m = F m W V : where W Q , W K , W V are learnable projection matrices; Then, the information flow between different modalities is calculated through cross-modal attention: E i,j = W i,j [C i ; C j ] G i,j = σ (E i,j ) where W i,j is the inter-modal mapping matrix, G i,j is the gating unit between modality i and modality j, and σ is the sigmoid activation function; 4.3) an attention-guided modality fusion layer dynamically adjusts the importance weight of each modality according to different clinical scenarios and signal quality, designs a meta-attention network, and learns the attention weight assigned to each modality: where f a is the attention score function, q is the query vector, represents the current evaluation context, and M is the total number of modalities; the weighted multi-modal feature fusion is represented as: 4.4) a pain assessment reasoning engine generates a pain probability distribution and an intensity score based on the fused features, and adopts a multi-task learning framework to simultaneously optimize two related tasks of pain classification and pain intensity regression: Classification task related parameters: Weight matrix for the pain classification task, where d fused is the dimension of the fused features, n class is the number of classes for the pain classification; Bias vector for the pain classification task; Regression task related parameters: Weight matrix for the pain intensity regression task, outputting a single pain intensity score; Bias scalar for the pain intensity regression task; wherein; the multi-modal fused feature vector, the classification task weight matrix, the classification task bias vector, the regression task weight matrix, the regression task bias scalar, the pain classification probability distribution, 0-10 numeric rating scale pain intensity score; The joint loss function of multi-task learning is defined as: wherein is a cross-entropy loss, is a mean squared error loss, is a regularization term, and λ1, λ2, and λ3 are hyperparameters that trade off the different losses. 5) individual calibration is performed; 6) display the pain assessment results, recommend intervention measures, and trigger the warning mechanism if necessary.

8. The multi-modal physiological signal based pain assessment method as claimed in claim 1, wherein: The preprocessing process of step 2) includes noise reduction, filtering, segmentation, feature point marking, and data calibration, specifically including: 2.1) Preprocessing of facial image sequence: first, apply adaptive histogram equalization to enhance contrast, then use Gaussian mixture model GMM for background separation, and finally smooth the facial motion trajectory through Kalman filter; 2.2) Audio signal preprocessing: apply adaptive noise cancellation algorithm and 300-8000Hz band-pass filter to reduce environmental noise, then perform sound event detection through energy detection and zero-crossing rate analysis; 2.3) Physiological parameter signal preprocessing: use wavelet transform to remove baseline drift and high-frequency interference, and mark and interpolate missing data points through an outlier detection algorithm; 2.4) For the pre-processing of behavioral activity data, principal component analysis (PCA) is applied to reduce dimensionality and remove motion artifacts.

9. The multi-modal physiological signal based pain assessment method as claimed in claim 1, wherein: Step 3) The process of extracting feature vectors related to pain specifically includes: 3.1) Facial feature extraction, first locate multiple facial landmarks through an improved FaceNet model, then calculate FACS-based facial action unit (AU) activation intensity, focusing on analyzing the spatiotemporal variation patterns of core AUs related to pain; 3.2) Acoustic feature extraction, the system combines short-time Fourier transform (STFT) and continuous wavelet transform (CWT) to analyze the time-frequency features of crying, extract acoustic feature vectors including fundamental frequency trajectory, spectral energy distribution, harmonic structure, and formant features, and capture the timbre characteristics of sound through mel-frequency cepstral coefficients (MFCC); 3.3) Physiological feature extraction, the system applies multi-scale entropy and wavelet transform to analyze the nonlinear dynamics of heart rate variability and respiratory signals, calculates blood pressure fluctuation characteristics and oxygen saturation drop amplitude, etc; 3.4) Behavioral feature extraction, based on optical flow method and pose estimation technology, quantify limb movement patterns, calculate activity intensity index, body rigidity, and stress response characteristics.

10. The multi-modal physiological signal based pain assessment method as claimed in claim 1, wherein: Step 5) The individualized calibration process specifically adjusts evaluation parameters according to individual characteristic variables; this module uses a Bayesian network framework to integrate medical prior knowledge and individual characteristics into the evaluation process: where Pain is the pain state, Signals is the observed physiological signal, Ind is the individual characteristic variable, the conditional probability distribution P(Signals|Pain, Ind) is estimated from training data, and the prior probability P(Pain|Ind) is provided by the medical knowledge graph; the baseline behavior feature library stores the baseline behavior features of infants in a non-pain state, which is used to calculate individualized feature deviations: ΔF = F current - F baseline where F current is the current feature vector, F baseline is the stored baseline feature; The individual characteristic variables include age, gestational age, disease status, and medication, and the specific evaluation method is as follows: 5.1) The age or gestational age adaptation unit adjusts the evaluation weight according to the infant's developmental stage, the system divides infants into six developmental stages: infants less than 32 weeks of gestational age, infants 32-37 weeks of gestational age, full-term newborns 0-1 months old, early infants 1-6 months old, mid-infants 6-12 months old, and toddler stage 12-36 months old, and sets specific modal weight and pain expression mode prior for each stage: where g a is an age adjustment function, Age is a gestational age variable; 5.2) The disease influence compensation unit considers the regulatory effect of specific diseases on pain expression, adjusts the evaluation sensitivity through the disease-symptom knowledge graph: S adj = S orig · f d (Disease) where S adj is the adjusted pain score, S orig is the raw score, f d is the disease adjustment factor; 5.3) The drug effect evaluation unit analyzes the degree of inhibition of sedatives and analgesics on pain response, the system establishes a pharmacokinetic model, estimates the current drug concentration and its effect on pain expression according to the drug type, dose, and administration time: where I drug is the drug impact index, D is the number of drugs used simultaneously, C i (t) is the estimated concentration of drug i at time t, E i is the effect coefficient of drug i, w i is the weight factor.

Citation Information

Cited By

  • Multi-mode body feeling information processing method and system for physiological state evaluation

    CN121400782A

  • A multi-modal somatosensory information processing method and system for physiological state assessment

    CN121400782B

  • Newborn health monitoring system based on multi-modal data fusion

    CN121641461A

  • Multidimensional pain assessment method fusing eye tracking and physiological signals

    CN122681426A