Multi-mode pain grade automatic evaluation method

Through the multimodal data fusion method, combining image and speech input, deep learning and attention mechanisms are used to extract facial expressions and speech features to generate pain level evaluation results, solving the problem of insufficient accuracy and robustness of the traditional single modal evaluation method, and achieving higher pain evaluation accuracy and stability.

CN120030410APending Publication Date: 2025-05-23TIANJIN HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510109633.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional pain rating assessment methods rely on single modal data, making it difficult to fully reflect the patient's true pain level, especially in complex or emotional situations, where accuracy and robustness are limited.

Method used

The multimodal data fusion method is adopted, combining image and speech input, and facial expressions and speech characteristics are extracted through deep convolutional neural networks, bidirectional long and short-term memory networks and attention mechanisms, and pain level evaluation results are generated through the weighted fusion model.

Benefits of technology

It improves the accuracy and robustness of pain assessment, can more comprehensively reflect the patient's pain level, and is especially suitable for situations where patients cannot effectively express pain, reducing the subjective judgment error of medical staff.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030410A_ABST
    Figure CN120030410A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal pain grade automatic evaluation method, which automatically evaluates the pain grade of a patient by acquiring a facial expression image, a voice signal and physiological data of the patient and utilizing a multi-modal data fusion technology and a deep learning model. The method comprises the following steps: firstly, preprocessing multi-modal data to extract key features; inputting the extracted features into a multi-modal fusion network, and performing feature weighting and fusion based on an attention mechanism; and finally, outputting the pain level of the patient through the classifier. Compared with the prior art, the method has the advantages that multi-modal data can be comprehensively analyzed, the accuracy and objectivity of pain grade evaluation are improved, subjective errors of manual evaluation are reduced, and the method has a good clinical application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and multimodal data processing, and in particular to a multimodal pain level automatic assessment method. Background Art

[0002] With the development of artificial intelligence technology, smart healthcare has gradually become a research hotspot. Pain level assessment is of great significance in clinical diagnosis and treatment, and can help medical staff judge the patient's condition and provide appropriate treatment. However, traditional pain assessment usually relies on subjective feedback from patients or observations by medical staff, which is difficult to standardize and has large subjective errors. This assessment method is even more limited when the patient cannot express his or her pain normally (such as patients in coma or with speech disorders).

[0003] Current pain level assessment methods mainly focus on technologies based on a single modality, such as analyzing facial features in images through expression recognition technology to determine pain levels, or evaluating patient voice features based on voice emotion analysis technology. However, single-modality data may not be sufficient to fully reflect the patient's true pain level, especially in complex or emotional situations, where the accuracy and robustness of single-modality assessments are limited. Therefore, automatic pain assessment methods that rely on a single modality have obvious bottlenecks in practical applications.

[0004] Based on this, developing a multimodal automatic pain level assessment method has important application value. By combining image and voice input, this method can more comprehensively obtain the patient's facial expressions and voice features, thereby more accurately assessing the patient's pain level. The multimodal assessment method can make up for the shortcomings of a single modality and achieve higher assessment accuracy. It is particularly suitable for scenarios where patients cannot clearly express their pain, and has broad application prospects. Summary of the invention

[0005] The purpose of the present invention is to provide a method for automatic assessment of pain levels based on multimodal data, which automatically analyzes and evaluates the patient's pain level by combining image and voice input, thereby improving the accuracy and robustness of the assessment, and is particularly suitable for situations where patients cannot accurately express their pain.

[0006] A multimodal pain level automatic assessment method comprises the following steps:

[0007] S1, collect multimodal data of patients, including facial expression images, voice signals and physiological data;

[0008] S2, preprocessing the collected multimodal data to extract facial expression features, voice features and physiological features respectively;

[0009] S3, input the extracted multimodal features into the fusion model, and weight and fuse the features based on the attention mechanism;

[0010] S4. Analyze the fused features through the classifier and output the patient's pain level assessment results.

[0011] Preferably, step S1 is implemented in the following steps:

[0012] S1.1, collect facial expression images of patients through a high-resolution camera to generate facial video sequences;

[0013] S1.2, collecting the patient's voice signal through a microphone and storing it as a voice waveform file;

[0014] S1.3. Collect the patient's heart rate, blood oxygen saturation and skin conductance data through physiological monitoring equipment.

[0015] Preferably, the high-resolution camera used in step S1.1 is a high-definition motion capture camera with infrared function.

[0016] Preferably, the extraction of facial expression features in step S2 is based on a deep convolutional neural network model; the extraction of speech features is based on Mel spectrum coefficients; and the extraction of physiological features is based on time domain and frequency domain analysis methods.

[0017] Preferably, the deep convolutional neural network model is a ResNet50 network optimized by transfer learning.

[0018] Preferably, the fusion model in step S3 adopts a multimodal attention fusion network based on a self-attention mechanism to perform weighted fusion of different modal features.

[0019] Preferably, the attention weights set in the fusion model are adaptively adjusted by multimodal data.

[0020] Preferably, the classifier in step S4 is a two-layer fully connected neural network, and its output includes three levels of mild pain, moderate pain and severe pain.

[0021] The beneficial effects of the present invention are:

[0022] 1. The present invention utilizes multimodal fusion analysis of images and voices, which can more comprehensively reflect the patient's pain level compared to single-modality pain assessment methods, improve the accuracy and reliability of assessment, and is particularly suitable for situations where patients cannot effectively express pain;

[0023] 2. By combining transfer learning with pre-trained models (such as ResNet50 and VGGish models), the data and computing costs required for model training are significantly reduced, while the efficiency and accuracy of feature extraction are improved, ensuring the robustness of the system in different environments;

[0024] 3. The bidirectional long short-term memory network (BiLSTM) and attention mechanism are introduced to focus on the key parts of the speech features, so that the pain-related features in the speech signal can be more accurately identified, further improving the sensitivity and accuracy of the assessment;

[0025] 4. The multi-layer perceptron is used to fuse and classify multimodal information, realizing the comprehensive analysis of image and voice features, making the output pain level results more stable and reliable, providing medical staff with a more valuable reference basis for judgment;

[0026] 5. The present invention automatically encrypts the sequence images after system recognition to protect the privacy of the research object, and selects three-dimensional chaotic Logistic mapping for image encryption. The structure of three-dimensional chaotic Logistic mapping is more complex than that of low-dimensional chaotic sequence system, and the generated chaotic sequence is more random, which improves the security of image encryption;

[0027] 6. The automatic assessment method of the present invention reduces the subjective judgment errors of medical staff, effectively makes up for the shortcomings of existing pain assessment methods, and has strong practicality and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0029] Figure 1 It is a flow chart of a multi-modal pain level automatic assessment method based on multi-modal input of the present invention;

[0030] Figure 2 It is a VGGish network structure diagram of a multimodal pain level automatic assessment method of the present invention;

[0031] Figure 3 This is the original model training result graph;

[0032] Figure 4 This is the optimized model training result graph;

[0033] Figure 5 It is the image encryption effect. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0035] A multimodal pain level automatic assessment method, such as Figure 1 As shown, the specific implementation steps are as follows:

[0036] Step 1: Preprocessing and feature extraction of image data;

[0037] Step 1.1 collect the patient's facial image as input data, and ensure that the image is clear and contains complete facial expression information;

[0038] Step 1.2: Use the pre-trained ResNet50 model to extract image features through transfer learning. The transfer learning process uses the ImageNet dataset for pre-training and fine-tunes on the target dataset to obtain key features related to pain.

[0039] Step 1.3 further optimizes the image features through the residual network layer and inputs them into the Softmax classifier to generate the pain classification result 1 based on the image features.

[0040] Step 2: Preprocessing and feature extraction of speech data;

[0041] Step 2.1 collects the patient's speech data and converts the speech signal into a Mel-spectrogram to extract the time-frequency characteristics of the speech;

[0042] Step 2.2: Use the VGGish model to extract features from the Mel-spectrogram, and optimize the feature expression by adjusting the frequency weighting method and octave bandwidth setting;

[0043] Step 2.3 uses a bidirectional long short-term memory network (BiLSTM) to capture the contextual association of speech features and combines the attention mechanism to give higher weights to pain-related features;

[0044] Step 2.4 inputs the processed speech features into the Softmax classifier to generate a pain classification result 2 based on the speech features.

[0045] Step 3: Processing and evaluation of objective standards for sound intensity;

[0046] Step 3.1 collect the patient's sound intensity data and calculate the equivalent continuous sound pressure level, peak sound pressure level and maximum sound pressure level respectively;

[0047] Step 3.2 conducts a comprehensive analysis of the three sound pressure levels according to the different weights of the sound features to generate an independent classification result 3, which serves as the third objective criterion for pain level assessment.

[0048] Step 4: Decision layer weighted fusion and pain level output;

[0049] Step 4.1: input the image classification result (result 1), the speech classification result (result 2), and the classification result of the sound intensity standard (result 3) into the decision layer;

[0050] Step 4.2: In the decision layer, weights are assigned according to the contribution of different modalities to pain assessment (such as image weight, speech weight, and sound intensity weight);

[0051] Step 4.3: weighted fusion of the three results to calculate the weighted final comprehensive pain level assessment result;

[0052] Step 4.4 outputs the final comprehensive pain level result to provide accurate and objective pain level reference information for medical staff;

[0053] Step 4.5 After the video sequence pain results are output, the images are encrypted to protect the patient's privacy.

[0054] Through the above implementation steps, the present invention uses three independent objective criteria of image, voice and sound intensity for weighted fusion to generate a more comprehensive pain level assessment result.

[0055] In a method of a multimodal pain level automatic assessment method of the present invention:

[0056] The role of step 1 is to preprocess and extract features from the patient's facial image, obtain key image features related to the pain level, and provide a basis for subsequent multimodal feature fusion and pain level assessment. The pre-trained ResNet50 model was used, and the feature extraction of facial images was combined with the transfer learning method. At the same time, the extracted image features were further optimized through the residual network layer, and finally input into the Softmax classifier to generate preliminary classification results based on the image. ResNet50 uses residual learning technology to solve the gradient vanishing problem in deep networks and can effectively extract high-level features in facial expressions. Transfer learning not only significantly reduces the training data requirements and computational costs by fine-tuning the model parameters trained on large-scale datasets such as ImageNet, but also improves the model adaptability and performance of specific tasks. The advantages of this method are high feature extraction efficiency, low data requirements, strong adaptability, and good robustness under different image qualities or environments. By providing high-quality image features, step 1 lays a solid technical foundation for multimodal evaluation methods.

[0057] The role of step 2 is to preprocess and extract features of the patient's speech signal, extract speech features related to the pain level, and provide high-quality input for subsequent multimodal feature fusion and classification. This step extracts the time-frequency features of speech by converting the speech signal into a mel-spectrogram, and uses the pre-trained VGGish model to extract speech features. At the same time, the bidirectional long short-term memory network (BiLSTM) is combined to capture the contextual association of speech features, and the attention mechanism is introduced to give higher weights to key features related to pain. The VGGish model makes speech feature extraction efficient and accurate through transfer learning and specific optimization adjustments. BiLSTM combines the previous and next time series information to enhance the expressiveness of speech data, while the attention mechanism significantly improves the ability to focus on pain features. The advantages of this method are high feature extraction accuracy and strong robustness to noise. It is especially suitable for scenarios where the patient's speech signal may be affected by the environment or emotions, providing a reliable speech feature foundation for multimodal pain assessment.

[0058] The purpose of step 3 is to add an objective measurement standard to the pain level assessment by analyzing the patient's sound intensity data. This step collects the patient's sound signal and calculates the characteristic values ​​such as equivalent continuous sound pressure level, peak sound pressure level and maximum sound pressure level respectively, and performs a comprehensive analysis based on the weights of different sound pressure levels to generate independent classification results. Through a comprehensive and objective evaluation of sound intensity, the volume characteristics of the patient's pain expression can be effectively captured, avoiding errors caused by subjective judgment in traditional methods. This step uses a comprehensive analysis of multiple sound pressure features to further improve the accuracy and robustness of the assessment, and provides important supplementary information for subsequent multimodal feature fusion. Its advantage is that the objectivity of the evaluation results is improved through physical quantitative indicators, which is particularly suitable for situations where speech features may be interfered with by emotions or unclear expressions.

[0059] The role of step 4 is to generate a comprehensive pain level assessment result by weighted fusion of multimodal classification results, and finally provide accurate and reliable reference information for medical staff. In this step, the image classification results, speech classification results, and sound intensity assessment results are input into the decision layer, weights are assigned according to the contribution of each modality to pain assessment, and the results are weighted fused to calculate the final pain level assessment result. After the result is output, the image is encrypted. The advantage of this step is that it can make full use of the complementary characteristics of multimodal data to make up for the shortcomings of the single modality method, and at the same time, it can adapt to the needs of different scenarios through weight adjustment to achieve higher assessment accuracy and stability. The pain level assessment result finally generated is more comprehensive and objective, providing solid support for actual clinical applications. The image is encrypted through three-dimensional chaotic Logistic mapping to protect the privacy of patients.

[0060] from Figure 3 It can be seen that in the process of training the multimodal pain recognition and classification model, this study first used the ResNet50+VGGish model to train on the self-built pain dataset. The results showed serious overfitting. After 500 iterations, the accuracy of the training set was close to 100%, but the accuracy of the validation set was only 80%.

[0061] from Figure 4 It can be seen that this project uses the ResNet50 network + BiLSTM and the VGGish model optimized by the channel attention mechanism to train on the same dataset. The results show that the model does not overfit or underfit. After 500 iterations, the accuracy of the model on the training set and the validation set is maintained at 80%, and the model has good robustness.

[0062] from Figure 5It can be seen from the figure that this project uses three-dimensional chaotic Logistic mapping to encrypt video sequence images. The structure of three-dimensional chaotic Logistic mapping is more complex than that of low-dimensional chaotic sequence system, and the generated chaotic sequence is more random.

[0063] Table 1 Comparison of the impact of optimization algorithms on model training results

[0064]

[0065] As can be seen from Table 1: This project uses SGDM, Adam, RMSProp, and AdaGrad optimization algorithms to analyze and compare their effects in model training. The results show that the Adam optimization algorithm has the fastest convergence speed (15 cycles), and both the training loss and the validation loss are better than the results of other optimization algorithms in model training.

[0066] Table 2 Comparison of the impact of learning rate strategy on model training effect

[0067]

[0068] As can be seen from Table 2: This project tested the effect of the optimizer Adam learning rate strategy on the model training effect. This project compared the effect of the fixed learning rate and the segmented learning rate in Adam optimization on the model training effect. The results showed that the final training accuracy of the segmented learning rate was 88.7%, and the verification accuracy was 86.5%, both of which were better than the fixed learning rate. The segmented learning rate can fully reflect the advantages of the Adam learning strategy. Therefore, the Adam optimizer segmented learning rate strategy was selected for the training of the model in the target database of this project.

[0069] Table 3. Comparison of optimized network model performance

[0070]

[0071] As can be seen from Table 3: This study compares the performance of different optimization networks in model training. This paper selects LSTM, BiLSTM and GRU for model training in the speech pain expression database. The results show that LSTM has the fastest training completion time (1.8h), but LST has the lowest accuracy (75.2%) in the model training results. The model training results show that BiLSTM has the highest accuracy (80.8%), and BiLSTM also has the smallest loss value (0.190) and the best FI-Score value (78.6%). After comprehensive analysis of the above three optimization networks, this study finally selects the BiLST network as the optimization network of the VGGish model to improve the performance of the improved model.

[0072] Table 4 Comparison of attention mechanisms and their impact on model performance

[0073]

[0074] It can be seen from Table 4 that this project uses the attention mechanism to optimize the model network and improve the robustness and generalization ability of the model network recognition system. This study selects the main attention mechanisms currently used in image recognition and classification systems, namely: first-order attention mechanism (channel attention mechanism) and sparse attention mechanism, as well as the model training results without attention mechanism optimization. The results show that the model training time is the shortest (2.5h) without attention mechanism optimization, but the accuracy of model training is relatively low (76.5%). Comparing the first-order attention mechanism and the sparse attention mechanism model optimization training, the results show that the first-order attention mechanism loss value is higher (0.370), but the accuracy (81.7%) and F1-Score value (81.5%) of the first-order attention mechanism optimization model training are better than the sparse attention model optimization results.

[0075] Table 5 Precision, Recall and F1-score test results of different categories

[0076] category Data volume Precision Recall F1-score Pain2 125 0.82 0.80 0.81 Pain3 128 0.83 0.82 0.82 Pain4 130 0.85 0.83 0.84

[0077] As can be seen from Table 5, the final test of the improved model was conducted using some data from the BioVid thermal pain database. The BioVid thermal pain database is a multimodal pain database consisting of facial expression videos and physiological data (skin conductivity, electrocardiogram, electromyographic signals, and electroencephalogram signals) of 90 volunteers. The project team downloaded the BioVid thermal pain database after obtaining the consent of the developer. The pain intensity of the database uses a four-classification method, namely Pain1: the subject's pain threshold (the lowest pain value that the subject can recognize), P4: the subject's pain tolerance (the subject's highest acceptable pain level), and PA2 and PA3 are defined as intermediate intensities through linear interpolation between PA1 and PA4. Since the pain intensity recognition and classification model designed in this study uses a three-classification method, for elderly perioperative patients with hip fractures who have clinical pain symptoms, according to the characteristics of the research subjects, this study selected the Pain2, Pain3, and Pain4 data sets from the BioVid database during the model testing process, shuffled the data of the three categories as a whole, and then trained the model. The confusion matrix test results show that the accuracy of the model in the Pain4 category in the BioVid database reaches 85%, the prediction effect is optimal, and the model has good generalization ability and robustness.

[0078] The overall structure algorithm code of the present invention is as follows:

[0079] #Step 1: Image input and processing

[0080] image_features = ResNet50(image_input)

[0081] image_features = TransferLearning(image_features, target_dataset)

[0082] classification_result_1 = SoftmaxClassifier(image_features)

[0083] # Step 2: Audio input and processing

[0084] log_mel_spec = GenerateLogMelSpectrogram(audio_input)

[0085] balanced_audio_data = DataBalancing(log_mel_spec)

[0086] audio_features = VGgish(balanced_audio_data)

[0087] time_series_features = BiLSTM(audio_features, AttentionMechanism = True)

[0088] classification_result_2 = SoftmaxClassifier(time_series_features)

[0089] # Step 3: Sound pressure level calculation

[0090] pressure_level_results = CalculateSoundPressureLevels(audio_input)

[0091] classification_result_3 = SoftmaxClassifier(pressure_level_results)

[0092] # Step 4: Weighted summation

[0093] weights = [Weight1, Weight2, Weight3]

[0094] final_result=WeightedSum([classification_result_1,classification_result_2,classification_result_3],weights)

[0095] #Step 5: Output the results

[0096] print(f"Pain Level:{final_result}")

[0097] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multimodal pain level automatic assessment method, characterized in that: The following steps are involved: S1. Collect multimodal data of patients, including facial expression images, voice signals and physiological data; S2, preprocessing the collected multimodal data to extract facial expression features, voice features and physiological features respectively; S3, input the extracted multimodal features into the fusion model, and weight and fuse the features based on the attention mechanism; S4. Analyze the fused features through the classifier and output the patient's pain level assessment results.

2. The multimodal pain level automatic assessment method according to claim 1, characterized in that: The step S1 is specifically implemented according to the following steps: S1.1, collect facial expression images of patients through a high-resolution camera to generate facial video sequences; S1.2, collecting the patient's voice signal through a microphone and storing it as a voice waveform file; S1.

3. Collect the patient's heart rate, blood oxygen saturation and skin conductance data through physiological monitoring equipment.

3. A multimodal pain level automatic assessment method according to claim 2, characterized in that: The high-resolution camera used in step S1.1 is a high-definition motion capture camera with infrared function.

4. The multimodal pain level automatic assessment method according to claim 1, characterized in that: The extraction of facial expression features in step S2 is based on a deep convolutional neural network model; the extraction of speech features is based on Mel spectrum coefficients; The extraction of physiological features is based on time domain and frequency domain analysis methods.

5. A multimodal pain level automatic assessment method according to claim 4, characterized in that: The deep convolutional neural network model is a ResNet50 network optimized by transfer learning.

6. A multimodal pain level automatic assessment method according to claim 1, characterized in that: The fusion model in step S3 adopts a multimodal attention fusion network based on a self-attention mechanism to perform weighted fusion of different modal features.

7. A multimodal pain level automatic assessment method according to claim 6, characterized in that: The attention weights set in the fusion model are adaptively adjusted by multimodal data.

8. A multimodal pain level automatic assessment method according to claim 1, characterized in that: The classifier in step S4 is a two-layer fully connected neural network, and its output includes three levels of mild pain, moderate pain and severe pain.

Citation Information

Cited By

  • Early warning method and device for recognizing Alzheimer based on voiceprint features

    CN122266399A