Inplanatable multi-modal information system for postoperative chronic pain assessment

By designing a multimodal information system that combines facial expressions and physiological signals, the timing, comprehensiveness and multimodal data integration of postoperative chronic pain assessment in the prior art is solved, and personalized and accurate pain assessment is achieved, improving the accuracy and interpretability of the assessment.

CN120072289APending Publication Date: 2025-05-30JIANGSU APON MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510083479.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing postoperative chronic pain assessment methods lack the ability to integrate timing, comprehensiveness and multimodal data, resulting in errors and deviations in the evaluation and difficulty in dynamically capturing pain changes.

Method used

An interpretable multimodal information system was designed, combining facial expression data and physiological signals, and data processing and fusion were carried out through the facial expression algorithm framework and physiological signal algorithm module to achieve personalized pain assessment.

Benefits of technology

The system can accurately and personalize the evaluation of chronic pain after surgery, reduce individual differences in interference, improve the accuracy and interpretability of the evaluation, and is suitable for different patient groups, enhancing the credibility and practicality of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072289A_ABST
    Figure CN120072289A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable multi-modal information system for postoperative chronic pain assessment, which is characterized by comprising a facial expression algorithm framework for acquiring dynamic change of a facial expression facial action unit (AU) and a physiological signal algorithm module for receiving physiological signals, and the multi-modal fusion framework is used for fusing the facial expression algorithm framework and the physiological signal algorithm module. The method has the effects of accuracy, individuation and interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics extraction, and particularly to an interpretable multimodal information system for postoperative chronic pain assessment. Background Art

[0002] In the fields of facial expression feature extraction, biosignal processing, and chronic pain assessment; as well as in the fields of computer vision, facial expression recognition, and multimodal data fusion technologies. Currently, the process of pain quantification and assessment includes two core aspects: measurement and evaluation. In terms of measurement, existing pain quantification methods mainly rely on patients' self-reports, which include the Numerical Rating Scale (NRS), Visual Analogue Scale (VAS), etc. These methods can effectively evaluate indicators such as the intensity, nature, and duration of pain, and help doctors understand the causes and nature of pain, thereby enabling differential diagnosis and treatment. Evaluation is through the external manifestation of physiological data such as patients' self-narration and pain expressions. However, most of these traditional methods rely on patients' subjective reports, so there are often errors and biases, especially when evaluating the pain levels at different time points or among different patients. Traditional methods for postoperative chronic pain assessment include patients' self-reports, medical staff observation, questionnaire assessment, and physiological index monitoring, etc. Self-report allows patients to subjectively score through a pain scale, which is simple and direct, but the results are affected by emotions, culture, and expression ability, lacking consistency; medical staff observation judges through external features such as expressions and postures, which is suitable for patients with limited expression ability, but highly relies on experience and is vulnerable to subjective biases; the questionnaire method obtains patients' feelings through detailed pain descriptions, providing comprehensive information, but it is time-consuming and difficult to quantify; physiological index monitoring uses data such as heart rate and blood pressure to provide objective evidence, but a single index is vulnerable to interference from other factors and cannot comprehensively reflect the pain state. These methods have their own advantages and disadvantages, but generally lack the ability of timeliness, comprehensiveness, and multimodal data integration, which limits their application in complex pain scenarios. In the scenario of postoperative chronic pain assessment, patients' pain feelings are often affected by multiple factors, including activity status, psychological state, etc., resulting in more complex pain manifestations. Postoperative pain is usually divided into immediate pain and chronic pain, among which chronic pain has a particularly serious impact on patients, and its intensity changes over time and with the increase in activity. This makes traditional pain measurement methods have certain limitations when dealing with the dynamic changes and individual differences of patients' postoperative pain. The postoperative pain recovery status is crucial for patients' subsequent diagnosis and analgesia plan, and doctors need to formulate the next stage of treatment plan based on this information. However, currently, no multi-modal postoperative chronic pain detection system with high interpretability and acceptable to clinical practice has been proposed. Although existing pain assessment tools such as NRS and VAS provide important clinical data, they still strongly rely on patients' self-reports. The pain perception abilities of different patients vary greatly, and patients' subjective scores are also easily interfered by factors such as psychology and emotions. Therefore, how to capture the changes in patients' postoperative chronic pain in an objective and dynamic way has become an important challenge in the field of pain assessment. In addition, existing methods usually do not have the ability to be highly interpretable to clinical medical staff and continuously monitor the postoperative chronic pain level of patients, and enable them to timely know the pain recovery trend of patients.Meanwhile, the application of multimodal data, such as the combination of physiological signals (electrodermal activity EDA, heart rate variability HRV) and facial expression data, has not been fully considered. These physiological signals and facial expressions are closely related during the pain response process. However, due to different data sources, they usually face the problem of time alignment, which limits their combined application in pain assessment. Summary of the Invention

[0003] In view of the above problems, the purpose of the present invention is to provide an interpretable multimodal information system for postoperative chronic pain assessment that is accurate, personalized, and interpretable.

[0004] To achieve the above purpose, the present invention provides an interpretable multimodal information system for postoperative chronic pain assessment. The system includes a facial expression algorithm framework for obtaining the dynamic changes of facial action units AU of human faces, a physiological signal algorithm module for receiving physiological signals, and a multimodal fusion framework for fusing the facial expression algorithm framework and the physiological signal algorithm module.

[0005] In some embodiments, the specific steps of the facial expression algorithm framework are as follows:

[0006] (1) A video recording camera is placed on a shelf above the front face of the patient, facing the front face of the lying patient to record the facial expressions of the patient, and receiving the facial expression video data of the patient.

[0007] (2) Then, a sliding window is equipped and only continuous frames are obtained, and at the same time, it must be ensured that the order between these frames is continuous; the length of the sliding window can be freely set by medical staff according to the needs of the clinical scenario; each frame needs to be detected for the human face by the existing mature Dlib and InsightFace face detection toolkits.

[0008] (3) Next, the frames filtered by the sliding window will be corrected and cropped for focusing on the human face part by the functions provided by InsightFace, removing the noisy background part in the video frames.

[0009] (4) Finally, the system obtains a series of processed continuous facial expression frames.

[0010] In some embodiments, based on obtaining the valid frames under each of the windows, the well-known OpenFace facial action unit extractor in the industry is used to process the facial features of each frame, and the intensity and frequency information of the action units AU listed in the AU information table extracted by Openface can be obtained. According to the facial features of the current patient, medical staff can freely combine and select the AUs related to pain to measure the pain of the patient's face according to the patient's individual characteristics.

[0011] The OpenFace extractor, based on deep learning and facial key point detection algorithms, can accurately identify the dynamic changes of facial action units, generating data including AU intensity values (such as AU04 eyebrow tightening, AU07 orbicularis oculi muscle contraction, etc.) and the corresponding occurrence frequencies. By traversing all valid frames within the window, the current invention system only focuses on the intensity values of each AU. Assuming the number of AUs selected by medical staff is, then for each sliding window of length, dimensional intensity values can be obtained, and the dynamic changes of facial expressions are reflected by calculating the AU differences between adjacent frames. Specifically, the system subtracts the AU intensity of the previous frame from the AU intensity of the next frame for each of the AUs, obtaining a matrix difference matrix of shape 1×, representing the AU intensity difference between the current frame and the previous frame, denoted as:

[0012] Diff = [AUx 1 , AUx 2 ,..., AUx n .

[0013] In some embodiments, to further enhance the accuracy of evaluation, considering the individual differences in the pain levels of patients, the system also needs to model the intensity baseline values of AUs in the static state of the patient. The specific method is to collect the AU intensity values of multiple static frames and calculate their average value to obtain a reference matrix base of shape 1×, which represents the facial expression characteristics of the patient in a pain-free state.

[0014] In some embodiments, the total AU difference can be represented by calculating the sum of the element-wise absolute differences between the current frame and the reference;

[0015] This can be represented by the L1 norm:

[0016]

[0017] Where:

[0018] S t is the total AU difference of the t-th frame;

[0019] and base (i) are the values of the i-th AU in the current frame and the reference respectively;

[0020] The purpose of this step is to accurately quantify the dynamic changes of pain-induced Facial Action Units (AUs) based on frame-by-frame dynamic changes and in combination with the patient's personalized static baseline. This process can comprehensively consider the temporal dimension changes of facial expressions and the patient's specific baseline characteristics, further improving the accuracy of pain assessment. In particular, by subtracting the AU difference value of the individual baseline from each frame, the explosive change points of facial expressions during pain onset can be effectively captured through curves. This difference-based calculation method enables direct comparison of AU differences at different time points, thereby avoiding the interference of individual static levels and ensuring accurate identification of pain dynamic changes.

[0021] In some embodiments, the physiological signal algorithm module includes a data stream, a prediction module, a personalization module, and an output module.

[0022] The data stream section is responsible for receiving physiological signals from signal acquisition devices, specifically including two types of signals: Electrodermal Activity (EDA) and Blood Volume Pulse (BVP). The Electrodermal Activity (EDA) can capture changes in skin sweat gland activity and is closely related to the excitatory level of the autonomic nervous system. Therefore, it plays an important role in evaluating states such as emotions, stress, and pain. The Blood Volume Pulse (BVP) monitors changes in blood flow through optical means and provides key information such as heart rate and blood oxygen saturation. It plays an indispensable role in health monitoring and physiological state assessment.

[0023] The prediction module includes data preprocessing, signal decomposition, feature extraction, and discriminants, and these components not only progress sequentially but also interact with each other and work together.

[0024] In the data preprocessing stage, the system will automatically detect and eliminate abnormal data caused by signal acquisition device failures or improper wearing to ensure the accuracy and reliability of the data. First, the electrodermal signal data is downsampled to 32 Hz. This frequency range of 4 - 32 Hz is considered a safe downsampling interval according to research literature. Selecting 32 Hz is to consider the data volume requirements of edge computing capabilities while retaining signal integrity. Subsequently, a 1-second median filter is applied to smooth the data to eliminate outliers. Then, z-score normalization is performed on the EDA signal to eliminate individual differences. During this process, the mean and standard deviation of the EDA signal of the current patient in a calm state need to be obtained from the personalization module to perform subject-independent normalization on the signal in the pain state. If the calm state signal of the current object cannot be obtained, z-score normalization is directly used for normalization operations. This series of processing steps ensures the accuracy of the data and the generalization ability of the model.

[0025] The signal decomposition part uses CVX signal decomposition. The CVX algorithm is an EDA signal decomposition method based on convex optimization, aiming to extract more meaningful features from the electrodermal activity (EDA) signal, such as skin conductance level (SCL) and skin conductance response (SCR). By constructing a mathematical model, this algorithm decomposes the signal into low-frequency components and rapidly changing transient components, thus accurately capturing the important changes in the signal. The CVX algorithm not only improves the robustness and accuracy of decomposition but also can effectively handle noise and non-stationary signals. After signal analysis, the long-term change and transient change curves of the signal can be obtained. At this time, the low-frequency signal in the data is subjected to Butterworth filtering to eliminate the artificial artifacts of the EDA transient signal.

[0026] In the feature extraction step, the algorithm constructs three feature groups to comprehensively capture the characteristics of the EDA signal.

[0027] First, the response intensity type features include Phasic Peaks Amplitude Maximum, Phasic Maximum, Tonic Maximum, Phasic Area Under Curve, and Tonic Area Under Curve. These features reveal the maximum response intensity and total response amount of the EDA signal during monitoring, helping to quantify the overall activity intensity of the autonomic nervous system and the possible intensity of emotional responses.

[0028] Second, the response frequency / amplitude type features include Phasic Driver Number of Peaks, Rise TimeAverage, Tonic Standard Deviation, and Phasic Standard Deviation. These features mainly describe the change amplitude and response frequency of the EDA signal, helping to understand the frequency and response speed of the signal, revealing the fluctuations and response frequencies of the autonomic nervous system at a specific time, and the standard deviation reflects the volatility of the signal, thus reflecting the stability and fluctuation of the physiological state. Finally, the overall trend type features include Tonic Average and Phasic Average. These features describe the overall trend and baseline level of the EDA signal. The calculation of the average value smooths all values into an overall trend, used to reflect the basic physiological state of an individual during the entire monitoring period. Such features help to understand the overall physiological state, are not affected by individual extreme values, and are suitable for long-term observation and comparison of the baseline activity level of the individual's autonomic nervous system. Through these features, the present invention can deeply understand the activities of the autonomic nervous system from different dimensions and provide rich information for monitoring and analysis.

[0029] The discriminant analysis step involves data and label pairs extracted from three different data sources; the present invention uses the least squares method to construct a linear discriminant model. The advantage of using the least squares method is that it provides a simple and intuitive way to estimate the parameters in a linear model. This estimation method has good interpretability in statistics and can generate a robust model, which is convenient for understanding and application;

[0030] HRV heart rate variability analysis was performed on the blood volume pulse BVP signal, and three classic time series features pNN50, Mean RR interval, and Standard Deviation of NN Intervals were selected; these features are widely used in the medical field for heart rate analysis and are familiar to doctors; this selection not only improves the medical interpretability of the features but also provides the possibility for the system to be accepted and understood by doctors in clinical applications, further enhancing the credibility and practicality of the system.

[0031] In some embodiments, in the multimodal fusion framework, the feature layer outputs of two pre-trained models are respectively extracted, and the parameters of their feature extraction parts are frozen to ensure feature stability. Then, an MLP multi-layer perceptron is designed as the fusion module, and the features of the two modalities are concatenated and input, and the parameters of the MLP are trained to capture the non-linear correlation and potential interaction relationships between the features;

[0032] To construct the training data, physiological signal and facial expression data pairs from different subjects are selected to ensure that they correspond to the same pain label and reflect the same pain state; during the training process, the MLP uses the pain label as the supervision signal and learns the weight relationship and mapping method of the fusion features by optimizing the loss function; the output module of the postoperative chronic pain assessment system provides a long-term pain assessment report for the patient by aggregating the pain prediction results at multiple time points and combining time series analysis and statistical methods; this module uses a sliding window to smooth the data, calculates the average pain intensity, fluctuation range, and trend change, and generates a time series graph and a comprehensive pain index.

[0033] The beneficial effects of the present invention are accurate, personalized, and interpretable. Since traditional pain assessment methods usually rely on fixed standard scales, they ignore the differences in facial expressions and physiological characteristics of different patients in a static state. The present invention realizes label-free expression-based personalized pain assessment by combining individualized baseline facial expression data (i.e., the baseline matrix). During the dynamic assessment process, the system eliminates the interference of individual differences on pain assessment by subtracting the baseline from the specific baseline value of each patient, making the pain assessment of each patient more accurate and conforming to their physiological state. This method is particularly suitable for postoperative patients lying flat, newborn infants, aphasic patients, and other patients who cannot express pain verbally. It improves the sensitivity and accuracy of pain assessment for different individuals through a personalized model, effectively avoiding the limitations of traditional general scales. In addition, the personalized part of the physiological signal algorithm in the system normalizes the EDA signal by combining the specific calm state signal of the patient, eliminating the influence of individual differences on pain assessment and ensuring more accurate individualized pain monitoring. At the same time, through multi-dimensional feature extraction (response intensity, frequency, amplitude, and overall trend) and multi-signal fusion (EDA and BVP signals), the robustness and accuracy of the model are effectively improved. The linear discriminant model constructed by the least squares method increases the clinical interpretability of the assessment results, helping doctors better understand and apply the system. Overall, the personalized processing method significantly improves the applicability, accuracy, and clinical practicality of the system in different patient groups. In addition, with high clinical interpretability, different from many traditional deep learning models, the present invention provides interpretable pain assessment results for doctors by combining the analysis of action unit differences in facial expressions and the comprehensive assessment of physiological signals. In traditional "black box" artificial intelligence methods, the assessment results often lack sufficient interpretability, and it is difficult for doctors to understand how the system arrives at specific assessment results, which poses a certain obstacle to clinical decision-making support. However, the present invention clearly demonstrates the impact of each facial action unit and physiological characteristic on the pain assessment result through an intuitive difference matrix and baseline subtraction method, enabling doctors to clearly understand the specific contributions of different physiological and facial expression characteristics in pain assessment. Through this transparent process, doctors can more easily make clinical decisions based on the assessment results, enhancing the credibility and operability of the system in practical applications. The technical key points of this system:

[0034] (1) Personalized baseline modeling and baseline subtraction method

[0035] The present invention proposes a method for constructing a personalized baseline. By calculating the average AU intensity of a patient in a static state, a personalized baseline matrix is established, and the baseline subtraction method is adopted during the dynamic evaluation process to eliminate the interference of individual static differences. This innovative feature enables pain assessment to be more in line with the individual characteristics of patients, improving the accuracy and reliability of the assessment results, and is particularly suitable for special groups with limited language expression (such as newborn infants or aphasia patients). By calculating the mean and standard deviation of physiological signals of pain patients in a calm state, the signal to be predicted is normalized to eliminate the individual differences in the signals.

[0036] (2) Dynamic analysis of the difference matrix:

[0037] The present invention accurately models the dynamic changes of facial expressions, quantifies the changes in facial expressions using the AU difference matrix (the difference in AU intensity between the current frame and the previous frame), and further improves the sensitivity to pain mutation points (such as the onset of severe pain). This method can capture the instantaneous changes in pain on a frame-by-frame basis, thereby more accurately identifying the dynamic changes in pain than the prior art.

[0038] (3) Interpretable pain assessment results

[0039] Different from traditional deep learning "black box" models, the present invention enables the system evaluation results to have high interpretability through specific analysis of each facial action unit and physiological data. Doctors and clinical staff can understand and trace the specific process and results of pain assessment, and thus use the assessment results with more confidence in clinical decision-making, enhancing the clinical applicability and trust of the system.

[0040] (4) Multimodal fusion technology

[0041] By fusing physiological signals and facial expression data, the problem of multimodal data consistency is solved, and multi-dimensional information related to pain can be comprehensively captured, thereby improving the robustness and accuracy of pain assessment. Secondly, using the trained model to extract features and fusing them through MLP can effectively model the non-linear correlation and potential interaction relationship between the two modal features, optimizing the multimodal feature space. Finally, through time series analysis to continuously monitor chronic pain after surgery, the system can not only track the change trend of pain intensity, but also timely detect the risk of pain deterioration, provide a comprehensive and intuitive pain assessment report, and provide reliable postoperative pain management support for patients and the medical team.

[0042] Therefore, the system learns the non-linear correlation between non-single object multimodal features through a multi-layer perceptron (MLP). The multimodal fusion method of the present invention effectively solves the problem of data alignment between facial expressions and physiological signals, and realizes an accurate, personalized and interpretable postoperative chronic pain assessment scheme. Brief Description of the Drawings

[0043] Figure 1 It is a schematic structural diagram of the recording of human facial expressions in the present invention;

[0044] Figure 2 It is a flowchart for obtaining valid human face pictures within a window in the present invention;

[0045] Figure 3 It is a chart of AU information extracted by Openface in the present invention;

[0046] Figure 4 It is a framework diagram of the multi-modal fusion algorithm in the present invention. Detailed Description of the Invention

[0047] The following further describes the invention in detail with reference to the drawings.

[0048] The application object of the present invention is the patient group in the lying and resting state after surgery (with a small head swing amplitude). Such people cannot express the degree of pain smoothly and accurately by themselves in the postoperative scenario. The system device of this invention is: a video recording camera is mounted on a shelf above the front face of the patient, making it face the front face of the lying patient for recording human facial expressions (as shown below Figure 1 shown)

[0049] The system is equipped with a sliding window for the received human face expression video data of the patient as Figure 2 shown (the length can be freely set by medical staff according to clinical scenario requirements) and only obtains continuous frames (each frame needs to be detected by the existing mature Dlib[1] and InsightFace[2] human face detection toolkits), and at the same time, it must be ensured that the order between these frames is continuous). Then, the frames filtered by the sliding window will be corrected and cropped by the functions provided by InsightFace to focus on the human face part, removing the noisy background part in the video frame. Finally, the system obtains continuous human face expression frames after processing the frames.

[0050] On the basis of obtaining valid frames under each window, the well-known OpenFace3 human face action unit extractor in the industry is used to process the facial features of each frame, and the intensity and frequency information of the action units (AUs) as shown in the Figure 3 AU information table extracted by Openface can be extracted. According to the current facial features of the patient, medical staff can freely combine and select the AUs related to pain to measure the pain of the patient's face according to the patient's personality characteristics.

[0051] The OpenFace extractor, based on deep learning and facial key point detection algorithms, can accurately identify the dynamic changes of facial action units, generating data including AU intensity values (such as AU04 brow tightening, AU07 orbicularis oculi contraction, etc.) and corresponding occurrence frequencies. By traversing all valid frames within the window, the current invention system only focuses on the intensity values of each AU. Assuming the number of AUs selected by medical staff is, for each sliding window of length, dimensional intensity values can be obtained, and the dynamic changes of facial expressions are reflected by calculating the AU differences between adjacent frames. Specifically, the system subtracts the AU intensity of the previous frame from the AU intensity of the next frame for each of the AUs, obtaining a matrix of shape 1× (difference matrix) representing the AU intensity difference between the current frame and the previous frame, denoted as:

[0052] Diff = [AUx 1 , AUx 2 ,..., AUx n .

[0053] To further enhance the accuracy of the assessment and considering the individual differences in the pain levels of patients, the system also needs to model the intensity baseline values of AUs in the static state of the patients. The specific method is to collect the AU intensity values of multiple static frames and calculate their average value to obtain a 1× reference matrix (base), and this reference value represents the facial expression characteristics of the patient in a pain-free state.

[0054] In practical applications, the total AU difference can be represented by calculating the sum of the element-wise absolute differences between the current frame and the reference.

[0055] This can be represented by the L1 norm:

[0056]

[0057] Where:

[0058] S t is the total AU difference of the t-th frame;

[0059] and base (i) are the values of the i-th AU in the current frame and the reference respectively;

[0060] The purpose of this step is to precisely quantify the dynamic changes in pain-induced Facial Action Units (AUs) based on frame-by-frame dynamic variations, combined with the patient's personalized static baseline. This process can comprehensively consider the temporal dimension changes of facial expressions and the patient's specific baseline characteristics, further improving the accuracy of pain assessment. In particular, by subtracting the AU difference value of the individual baseline from each frame, the explosive change points of facial expressions during pain onset can be effectively captured (through curves). This difference-based calculation method enables direct comparison of AU differences at different time points, thus avoiding the interference of individual static levels and ensuring the accurate identification of pain dynamic changes.

[0061] 2. Physiological Signals

[0062] In this section, the physiological signal algorithm system is designed with four core components: data stream, prediction module, personalization module, and output module. The data stream section is responsible for receiving physiological signals from signal acquisition devices, specifically including two types of signals: electrodermal activity (EDA) and blood volume pulse (BVP). Electrodermal activity (EDA) can capture changes in skin sweat gland activity and is closely related to the excitatory level of the autonomic nervous system. Therefore, it plays an important role in evaluating states such as emotions, stress, and pain. Blood volume pulse (BVP) monitors changes in blood flow through optical means and provides key information such as heart rate and blood oxygen saturation. It plays an indispensable role in health monitoring and physiological state assessment.

[0063] As Figure 3 shown, the prediction module is the core of the system and consists of four key components: data preprocessing, signal decomposition, feature extraction, and discriminant. These components not only progress sequentially but also interact with each other and work together.

[0064] In the data preprocessing stage, the system will automatically detect and eliminate abnormal data caused by signal acquisition device failures or improper wearing to ensure the accuracy and reliability of the data. First, the electrodermal signal data is downsampled to 32Hz. This frequency range (4 - 32Hz) is considered a safe downsampling interval according to research literature. Selecting 32Hz is to consider the data volume requirements of edge computing capabilities while retaining signal integrity. Subsequently, a 1-second median filter is applied to smooth the data to eliminate outliers. Then, z-score normalization is performed on the EDA signal to eliminate individual differences. In this process, the mean and standard deviation of the EDA signal in the calm state of the current patient need to be obtained from the personalization module to perform subject-independent normalization on the signals in the pain state. If the calm state signal of the current object cannot be obtained, z-score normalization is directly used for the normalization operation. This series of processing steps ensures the accuracy of the data and the generalization ability of the model.

[0065] In the signal decomposition part, the system adopts CVX signal decomposition. The CVX algorithm is an EDA signal decomposition method based on convex optimization, aiming to extract more meaningful features from the electrodermal activity (EDA) signal, such as skin conductance level (SCL) and skin conductance response (SCR). By constructing a mathematical model, this algorithm decomposes the signal into low-frequency components and rapidly changing transient components, thus accurately capturing the important changes in the signal. The CVX algorithm not only improves the robustness and accuracy of the decomposition but also can effectively handle noise and non-stationary signals. After signal analysis, the long-term change and instantaneous change curves of the signal can be obtained. At this time, the low-frequency signal in the data is subjected to Butterworth filtering to eliminate the artificial artifacts of the EDA instantaneous signal.

[0066] In the feature extraction section, the system algorithm constructs three feature groups to comprehensively capture the characteristics of the EDA signal. First, the response intensity type features include Phasic Peaks Amplitude Maximum, Phasic Maximum, Tonic Maximum, Phasic Area Under Curve, and Tonic Area Under Curve. These features reveal the maximum response intensity and total response amount of the EDA signal during the monitoring period, helping to quantify the overall activity intensity of the autonomic nervous system and the possible emotional response intensity. Second, the response frequency / amplitude type features include Phasic Driver Number of Peaks, Rise Time Average, Tonic Standard Deviation, and Phasic Standard Deviation. These features mainly describe the change amplitude and response frequency of the EDA signal, helping to understand the frequency and response speed of the signal, revealing the fluctuations and response frequency of the autonomic nervous system at a specific time, and the standard deviation reflects the volatility of the signal, thus reflecting the stability and fluctuation of the physiological state. Finally, the overall trend type features include Tonic Average and Phasic Average. These features describe the overall trend and baseline level of the EDA signal. The calculation of the average value smooths all values into an overall trend, which is used to reflect the basic physiological state of an individual during the entire monitoring period. This type of feature helps to understand the overall physiological state, is not affected by individual extreme values, and is suitable for long-term observation and comparison of the baseline activity level of the individual's autonomic nervous system. Through these features, the present invention can deeply understand the activities of the autonomic nervous system from different dimensions, providing rich information for monitoring and analysis.

[0067] The discriminant analysis step involves data and label pairs extracted from three different data sources. The present invention uses the least squares method to construct a linear discriminant model. The advantage of using the least squares method is that it provides a simple and intuitive way to estimate the parameters in a linear model. This estimation method has good interpretability in statistics and can generate a robust model, which is easy to understand and apply.

[0068] For BVP (Blood Volume Pulse) signals, the system performs HRV (Heart Rate Variability) analysis and selects three classic time series features (pNN50, Mean RR interval, Standard Deviation of NN Intervals). These features are widely used in the medical field for heart rate analysis and are familiar to doctors. This selection not only improves the medical interpretability of the features but also provides the possibility for the system to be accepted and understood by doctors in clinical applications, further enhancing the credibility and practicality of the system.

[0069] 3. Multimodal Fusion (Physiological Signals and Facial Expressions)

[0070] Based on the idea of multimodal data fusion, a framework that combines a physiological signal algorithm and a facial expression algorithm is designed (as Figure 4 shown). In this framework, the feature layer outputs of two pre-trained models are respectively extracted, and the parameters of their feature extraction parts are frozen to ensure feature stability. Then, an MLP (Multi-Layer Perceptron) is designed as a fusion module. The features of the two modalities are concatenated and input, and the parameters of the MLP are trained to capture the non-linear correlation and potential interaction relationships between the features.

[0071] To construct the training data, pairs of physiological signal and facial expression data from different subjects are selected to ensure that they correspond to the same pain label and reflect the same pain state. During the training process, the MLP uses the pain label as the supervision signal and learns the weight relationship and mapping method of the fusion features by optimizing the loss function (such as mean square error or cross-entropy).

[0072] The advantages of this method are as follows: on the one hand, it uses data pairs composed of multimodal data of different patients for training, solving the problem of data consistency in multimodality; on the other hand, it models the non-linear correlation between the features of the two modalities through the MLP, optimizing the feature space of multimodal fusion. In addition, this method can effectively combine the advantages of the two modalities, improve the robustness and accuracy of pain assessment, and enhance the adaptability of the model to complex pain patterns, providing a feasible solution for the popularization of multimodal fusion models in practical applications.

[0073] For example, for the pain assessment from the 1st to the 7th day after surgery, the system shows that the average pain level is 5.2 (moderate), and the maximum pain level is 8.0 (severe), which occurred on the 2nd day and then showed a gradually decreasing trend; meanwhile, the fluctuation range is from 3.5 to 8.0. The system further recommends continuous monitoring after the 5th day. If the pain does not improve or worsens, the medical team should be contacted. The output form of this module is intuitive and comprehensive, providing reliable postoperative pain management support for patients and the medical team.

[0074] The above are only some embodiments of the present invention. For those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the invention.

Claims

1. An interpretable multimodal information system for postoperative chronic pain assessment, characterized in that: The system includes a facial expression algorithm framework for acquiring dynamic changes of facial action units AU, a physiological signal algorithm module for receiving physiological signals, and a multimodal fusion framework for fusing the facial expression algorithm framework and the physiological signal algorithm module.

2. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 1, characterized in that: The specific steps of the facial expression algorithm framework are as follows: (1) A video recording camera is mounted on a shelf above the front of the patient, so that it is facing the front of the patient lying flat to record the facial expression of the patient and receive the facial expression video data of the patient; (2) After that, a sliding window is provided and only continuous frames are obtained, and the order between these frames must be ensured to be continuous; the length of the sliding window can be freely set by medical staff according to the needs of clinical scenarios; each frame must be detected as a face by the existing mature Dlib and InsightFace face detection toolkits; (3) Then, the frames filtered by the sliding window will be corrected and cropped by the function provided by InsightFace after focusing on the face part, and the noise background part in the video frame will be removed; (4) Ultimately, the system obtains a series of processed facial expression frames.

3. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 2, characterized in that: On the basis of obtaining the valid frames under each of the windows, the facial features of each frame are processed using the OpenFace facial action unit extractor known in the industry, and the intensity and frequency information of the action unit AU listed in the AU information table extracted by Openface can be extracted; according to the facial features of the current patient, the medical staff can freely combine and select the AU related to pain to measure the pain of the patient's face according to the patient's personality characteristics; The OpenFace extractor is based on deep learning and facial key point detection algorithms, and can accurately identify the dynamic changes of facial action units, generate AU intensity values ​​(such as AU04 eyebrow tightening, AU07 orbicularis oculi contraction, etc.) and corresponding frequency of occurrence data; by traversing all valid frames in the window, the system of this invention currently only focuses on the intensity value of each AU; assuming that the number of AUs selected by the medical staff is, then for each sliding window of length, the individual intensity values ​​can be obtained, and the dynamic changes of facial expressions are reflected by calculating the AU differences between adjacent frames; specifically, the system uses the individual AU intensities of the next frame minus the AU intensities of the previous frame to obtain a matrix difference matrix with a shape of 1×, which represents the AU intensity difference between the current frame and the previous frame, recorded as: Diff=[AUx1,AUx2,...,AUx n ] 4. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 3, characterized in that: To further enhance the accuracy of the assessment, taking into account the individual differences in the degree of pain among patients, the system also needs to model the AU intensity baseline value when the patient is in a static state; the specific method is to collect the AU intensity values ​​of multiple static frames and calculate their average value to obtain a 1× benchmark matrix base, which represents the facial expression characteristics of the patient in a pain-free state.

5. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 3, characterized in that: The total AU difference may be represented by calculating the sum of element-wise absolute differences between the current frame and the reference; This can be expressed using the L1 norm: in: S t is the sum of AU differences in the tth frame; and base (i) are the values ​​of the i-th AU in the current frame and the baseline, respectively; The purpose of this step is to accurately quantify the dynamic changes of facial action units (AU) induced by pain on the basis of frame-by-frame dynamic changes combined with the patient's personalized static baseline; this process can comprehensively consider the temporal dimension changes of facial expressions and the patient's specific baseline characteristics, further improving the accuracy of pain assessment; in particular, by subtracting the AU difference value of the individual baseline from each frame, the explosive change point of facial expression during the onset of pain can be effectively captured through the curve; this difference-based calculation method enables the direct comparison of AU differences at different time points, thereby avoiding the interference of individual static levels and ensuring the accurate identification of dynamic changes in pain.

6. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 1, characterized in that: The physiological signal algorithm module includes a data stream, a prediction module, a personalization module and an output module; The data flow link is responsible for receiving physiological signals from the signal acquisition device, including two signals: galvanic skin response (EDA) and blood volume pulse (BVP); the galvanic skin response (EDA) can capture changes in the activity of the skin sweat glands and is closely related to the excitement of the autonomic nervous system, so it plays an important role in assessing states such as emotions, stress and pain; the blood volume pulse (BVP) monitors changes in blood flow through optical means, providing key information such as heart rate and blood oxygen saturation, and plays an indispensable role in health monitoring and physiological status assessment; The prediction module includes data preprocessing, signal decomposition, feature extraction and discriminant, and these components are not only sequentially progressive, but also intertwined and work together; In the data preprocessing stage, the system automatically detects and removes abnormal data caused by signal acquisition equipment failure or improper wearing to ensure the accuracy and reliability of the data; first, the skin electrical signal data is downsampled to 32Hz. This frequency range of 4-32Hz is considered to be a safe downsampling interval according to research literature; 32Hz is selected to preserve signal integrity while taking into account the demand for data volume for edge computing capabilities; then, a 1-second median filter is applied to smooth the data to eliminate outliers; then, the EDA signal is z-score normalized to eliminate individual differences; in this process, the mean and standard deviation of the EDA signal of the current patient in a calm state need to be obtained from the personalization module in order to subject-independently normalize the signal of the pain state; if the calm state signal of the current subject cannot be obtained, z-score normalization is directly used for normalization; this series of processing steps ensures the accuracy of the data and the generalization ability of the model; The signal decomposition part adopts CVX signal decomposition. The CVX algorithm is an EDA signal decomposition method based on convex optimization, which aims to extract more meaningful features from the skin electrical activity EDA signal, such as skin conductance level SCL and skin conductance response SCR; the algorithm decomposes the signal into low-frequency components and rapidly changing transient components by constructing a mathematical model, so as to accurately capture important changes in the signal; the CVX algorithm not only improves the robustness and accuracy of the decomposition, but also can effectively process noise and non-stationary signals; after signal analysis, the long-term change and the time-varying curve of the signal can be obtained, and at this time, the low-frequency signal in the data is subjected to butterworth filtering to eliminate the artificial artifacts of the EDA transient signal; The feature extraction link described is that the algorithm constructs three feature groups to fully capture the characteristics of the EDA signal; First, the response intensity features include Phasic Peaks Amplitude Maximum, Phasic Maximum, Tonic Maximum, Phasic Area Under Curve, and Tonic Area Under Curve. These features reveal the maximum response intensity and total response amount of the EDA signal during the monitoring period, which helps to quantify the overall activity intensity of the autonomic nervous system and the possible emotional response intensity. Secondly, the response frequency / amplitude features include Phasic Driver Number of Peaks, Rise TimeAverage, Tonic Standard Deviation, and Phasic Standard Deviation. These features mainly describe the change amplitude and response frequency of the EDA signal, help understand the frequency and response speed of the signal, and reveal the fluctuation and response frequency of the autonomic nervous system within a specific time. The standard deviation reflects the volatility of the signal, thereby reflecting the stability and fluctuation of the physiological state. Finally, the overall trend features include Tonic Average and Phasic Average. These features describe the overall trend and baseline level of the EDA signal. The calculation of the average value smoothes all values ​​into an overall trend to reflect the basic physiological state of the individual during the entire monitoring period. This type of feature helps to understand the overall physiological state and is not affected by individual extreme values. It is suitable for long-term observation and comparison of the baseline activity level of the individual autonomic nervous system. Through these features, the present invention can deeply understand the activities of the autonomic nervous system from different dimensions, providing rich information for monitoring and analysis. The discriminant analysis step involves data and label pairs extracted from three different data sources; the present invention uses the least squares method to construct a linear discriminant model. The advantage of using the least squares method is that it provides a simple and intuitive method to estimate the parameters in the linear model. This estimation method has good interpretability in statistics and can produce a robust model that is easy to understand and apply; The blood volume pulse BVP signal was subjected to HRV heart rate variability analysis, and three classic time series features pNN50, Mean RR interval, and Standard Deviation of NN Intervals were selected; these features are widely used in heart rate analysis in the medical field and are familiar to doctors; this selection not only improves the medical interpretability of the features, but also provides the possibility for the system to be accepted and understood by doctors in clinical applications, further enhancing the credibility and practicality of the system.

7. The interpretable multimodal information system for postoperative chronic pain assessment according to claim 1, characterized in that: The multimodal fusion framework extracts the feature layer outputs of two trained models respectively and freezes the parameters of the feature extraction part to ensure feature stability. Then, an MLP multi-layer perceptron was designed as a fusion module to concatenate the features of the two modalities and train the parameters of the MLP to capture the nonlinear correlation and potential interaction between the features. In order to construct training data, pairs of physiological signals and facial expression data from different subjects were selected to ensure that they corresponded to consistent pain labels and reflected the same pain state. During the training process, MLP uses pain labels as supervision signals and optimizes the loss function to learn the weight relationship and mapping method of fusion features. The output module of the postoperative chronic pain assessment system aggregates the pain prediction results of multiple time points and combines time series analysis with statistical methods to provide patients with a long-term pain assessment report. This module uses a sliding window to smooth data, calculate the average pain intensity, fluctuation range and trend changes, and generate a time series graph and a comprehensive pain index.

Citation Information

Cited By

  • Closed-loop olfactory stimulation awareness assessment method, equipment and medium

    CN120788582A

  • Closed-loop olfactory stimulation awareness assessment method, device, and medium

    CN120788582B

  • Acute pain assessment method and system based on EDA signal

    CN121867707A

  • An acute pain assessment method and system based on EDA signals

    CN121867707B