Pain recognition system based on voice emotion detection

Through a pain recognition system based on speech emotion detection, using voice signals for pain assessment, the problems of strong subjectivity and high cost of traditional methods are solved, and efficient and accurate pain assessment is achieved.

CN119993214APending Publication Date: 2025-05-13JIANGSU APON MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510159353.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional pain assessment methods are highly subjective and cannot quantify the pain degree in real time and objectively. The biological signal detection method is costly and complex to use.

Method used

A pain recognition system based on speech emotion detection is adopted, and the patient's voice signal is obtained through the speech acquisition module, preprocessing and emotional feature extraction is performed, long-term and short-term memory networks and attention mechanisms are used for emotion classification, and pain intensity is predicted in combination with the XGBoost regression model.

Benefits of technology

Non-contact, low-cost and real-time pain assessment is achieved, improving the efficiency and accuracy of pain assessment, especially for patients who cannot express pain clearly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993214A_ABST
    Figure CN119993214A_ABST
Patent Text Reader

Abstract

The invention discloses a pain recognition system based on voice emotion detection. The system comprises a voice acquisition module for acquiring an original voice signal from a patient, a voice preprocessing module which is connected with the voice acquisition module and performs preprocessing operation on the voice signal, and a voice emotion classification module which is connected with the voice preprocessing module and performs accurate recognition on the emotion type and intensity in the voice signal, and the pain recognition module is connected with the voice emotion classification module and performs pain scoring after emotion classification judgment. The method has the advantages of non-contact performance, low cost and real-time performance, and has the effect of high efficiency and accuracy of pain assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information recognition, and in particular to a pain recognition system based on voice emotion detection. Background Art

[0002] Pain is the physiological and psychological response of the human body to adverse stimuli. Pain assessment plays an important role in medical diagnosis and treatment, especially for some patients who cannot express pain directly (such as infants, the elderly, critically ill patients, etc.). Traditional assessment methods mainly rely on patients' subjective descriptions (such as visual analog scales, NRS scores, etc.) and medical staff's observations, but this method is highly subjective and relies on patients' cognitive and language expression abilities, and cannot quantify the degree of pain in real time and objectively. In recent years, a pain detection method that uses biological signals (such as electrocardiogram, electroencephalogram, electromyography, etc.) has gradually emerged, but this type of method requires additional wearable devices to identify pain information, which is not only costly to use, but also complicated to operate. Summary of the invention

[0003] In view of the above problems, the object of the present invention is to provide a pain recognition system based on voice emotion detection which has the advantages of non-contact, low cost and real-time performance, and high efficiency and accuracy in pain assessment.

[0004] To achieve the above-mentioned purpose, the present invention provides a pain recognition system based on voice emotion detection, wherein the system includes a voice acquisition module for acquiring original voice signals from patients, a voice preprocessing module connected to the voice acquisition module and performing preprocessing operations on the voice signals, a voice emotion classification module connected to the voice preprocessing module and accurately identifying the emotion type and intensity in the voice signals, and a pain recognition module connected to the voice emotion classification module and performing pain scoring after emotion classification.

[0005] In some implementations, the speech preprocessing module includes the following processing method:

[0006] (1) Speech denoising: After speech signal collection, the speech signal will be affected by environmental noise. The preprocessing module uses spectrum subtraction denoising technology to reduce background noise.

[0007] (2) Framing and windowing: In order to better analyze the local features of the speech signal, the signal is divided into multiple short-time frames, and then the Hamming window function is applied to each frame to reduce spectral leakage and enhance the accuracy of feature extraction;

[0008] (3) Emotional feature extraction: Speech carries rich emotional information, including pitch, volume, speaking speed, and intonation.

[0009] In some implementations, the spectral subtraction denoising technique is implemented as follows:

[0010] First, the system acquires the noise spectrum in the signal;

[0011] Then, the noise part is removed by the following formula:

[0012] S clean (f) = S noisy (f)-α·N(f)

[0013] Among them, S clean (f) is the denoised signal, S noisy (f) is the original noisy speech signal, N(f) is the spectrum of the noise, and α is the coefficient that controls the strength of noise suppression.

[0014] In some embodiments, the formula for the short-time Fourier transform is as follows:

[0015]

[0016] Where X(t,f) is the short-time Fourier transform result at time t and frequency f, x[n] is the sample point of the speech signal, and w[nt] is the windowing function.

[0017] In some embodiments, the emotion feature extraction step:

[0018] (1) Mel-frequency cepstral coefficients (MFCC): This process involves converting the speech signal to the Mel scale, calculating its power spectrum, and obtaining cepstral coefficients through discrete cosine transform. Mel-frequency cepstral coefficients (MFCC) capture the nonlinear response of the human auditory system and are suitable for expressing changes in emotional states.

[0019]

[0020] Among them, X k is the spectral component of the signal, K is the number of frequency bands, and N is the dimension of MFCC;

[0021] (2) Fundamental frequency F0: The fundamental frequency reflects the pitch of speech and can significantly affect the expression of emotions, especially pain-related emotional fluctuations;

[0022]

[0023] Where T is the size of the time window, fundamental_frequency(t) is the fundamental frequency value at time t;

[0024] (3) Energy feature: The energy intensity of speech reflects the intensity of the speaker’s emotions. Emotional changes caused by pain are usually accompanied by changes in speech energy. Therefore, by calculating the energy value of each short-time frame, the speaker’s emotional state can be further revealed. The calculation of energy features can be completed by summing the squares of all sample points in each frame, that is:

[0025]

[0026] Among them, x(t) is the time domain waveform of the speech signal, and E is the energy.

[0027] In some embodiments, the speech emotion classification module selects the long short-term memory network LSTM as the main time series model in order to accurately identify the type and intensity of emotions, especially the continuity of emotions and the persistence of pain, and combines the attention mechanism to solve the problem of the timing of emotion changes. The recent emotional fluctuations have a greater impact on the current emotion discrimination results;

[0028] The LSTM network is based on the fact that emotions are a dynamic process, and the fluctuation of emotions is often closely related to the emotional state of the previous moment. The LSTM network can maintain and update the long-term and short-term memory in the time series through a gating mechanism, so that the emotional state at each moment is not only affected by the current input, but also by the feedback of the historical emotional state.

[0029] The attention mechanism ensures that the module can dynamically assign different importance weights to different time steps when processing time series data; the role of the attention mechanism is to calculate the emotional state h of each time step according to the emotional state h of each time step. t and a query vector q to calculate the weighting coefficient α t , the weighted coefficient indicates which emotional change at the current moment contributes most to the final judgment result; the query vector q can be a context vector summarized from the past emotional state, indicating the evolution trend of the current emotion;

[0030] (1) Calculate attention weights: For each time step t, use the multi-layer perceptron scoring function and nonlinear mapping capabilities to calculate the current emotional state h t and the similarity of the query vector q; specifically, firstly transform the hidden state h at time step t t and query vector q to form a vector x t =[h t ; q], and then the concatenated vector is transformed nonlinearly through a multi-layer perceptron to calculate the score (h t ,q), its expression is:

[0031] score(h t,q)=W1·σ(W0·x t +b0)+b1

[0032] Among them, W1 and W0 are the weight matrices in the multilayer perceptron, σ(·) is the activation function, b0 and b1 are bias terms, and x t =[h t ; q] is the merged input vector;

[0033] Subsequently, the score values ​​are converted into weights that can be used for the attention mechanism. The score values ​​of all time steps need to be softmax normalized to obtain the attention weight α for each time step. t ;

[0034] The process is as follows:

[0035]

[0036] Through softmax normalization, the weights α of all time steps t will be adjusted to the range of [0,1], and the weight sum is 1; thus, the attention weight α at each time step t It means that in the process of emotion discrimination, the current emotional state h t The relative importance to the final judgment result;

[0037] (2) Weighted summation: Next, based on the calculated attention weight α t , perform weighted summation on the LSTM outputs of all time steps to obtain the weighted emotion feature vector h final , which contains information weighted according to the importance of the emotional evolution:

[0038]

[0039] This weighted emotion feature vector will be used as the final input for emotion discrimination.

[0040] The weighted emotion feature vector output by the LSTM network and the attention mechanism is passed into a fully connected layer to classify the emotion; since there are only two categories of emotion classification, positive and negative, the output of the classifier is a probability value p negative , indicating the probability that the current emotion is negative.

[0041] In some implementations, the LSTM is configured as a bidirectional LSTM, which means that the LSTM can not only understand the past emotional evolution, but also capture the potential information of future emotions; the LSTM outputs h at each time step. tRepresents the emotional state of the current time step, which will be passed to the subsequent attention mechanism to calculate the weighted representation of the emotion.

[0042] In some embodiments, if the pain recognition module classifies the emotion as positive, it directly outputs 0; after the emotion is determined to be negative in the emotion classification stage, the system will enter the pain score prediction stage; in this stage, the XGBoost regression model is used to predict the intensity of pain by combining the emotion features and other dynamic features in the speech signal: the score range is 1 to 10; the core task of pain score prediction is to combine the fluctuation of emotion, the physiological features and dynamic changes in the speech, and calculate the pain score through the regression model;

[0043] (1) Input features:

[0044] The regression model for pain score is trained and predicted based on the following features: probability p of negative emotion negative , emotional feature vector h final , MFCC features, F0, energy fluctuation;

[0045] Therefore, the final input feature vector X becomes:

[0046] X=[p negative ,h final ,MFCC,F0,Energy]

[0047] (2) XGBoost regression model

[0048] The XGBoost regression model is used to predict the intensity P of pain based on these input features, and the output is the predicted value, ranging from 1 to 10. XGBoost fits the regression problem by training a series of decision trees, and the model will improve the prediction performance in an integrated manner.

[0049] The basic formula of XGBoost is:

[0050]

[0051] Where f(X) is the predicted pain score of the model, X is the input feature vector, and θ k is each tree T k The weight of (X), T k (X) is the output of the kth decision tree;

[0052] In XGBoost, each tree T k (X) are all learned by optimizing the loss function, where the loss function L is the mean square error. In order to improve the stability of training and prevent overfitting, XGBoost uses a regularization term Ω to control the complexity of the model.

[0053] The loss function is defined as:

[0054]

[0055] Among them, y i is the actual pain score, f(X i ) is the score predicted by the XGBoost model, and λ is the regularization parameter to avoid overfitting;

[0056] (3) XGBoost training

[0057] The model is trained by gradient boosting. In each iteration, the model updates the tree structure by minimizing the loss function.

[0058] The training process can be described by the following steps:

[0059] A. Initialization: Start with the initial model and initialize it to zero;

[0060] B. Calculate gradient: Calculate the gradient deviation of each sample based on the prediction error of the current model;

[0061] These gradient values ​​reflect the degree of error of each sample under the current model:

[0062]

[0063] C. Build a new tree: Build a new decision tree by calculating the gradient and second-order derivative information to correct the error of the current model.

[0064] D. Update model: Each iteration adjusts the model parameters θ k , by gradually reducing the error, we finally get a powerful model that can predict pain scores.

[0065] (4) Output pain score

[0066] After XGBoost training, the final regression model will be able to predict the pain intensity score P based on the input emotional features, negative emotion probability and audio features; the score P is a numerical value, the larger the score, the more intense the pain.

[0067] In some embodiments, the voice collection module is a microphone for collecting ambient audio signals, and the microphone is connected to a smart phone, a smart speaker, an ear-worn device, or a medical device.

[0068] The beneficial effects of the present invention are that it has the advantages of non-contact, low cost and real-time, and the effect of high efficiency and accuracy of pain assessment. Since the voice signal is a natural, non-contact signal source, it contains rich emotional and physiological information. Pain can significantly affect human voice characteristics, such as speech speed, pitch, stress and tremor. By using artificial intelligence and emotion detection technology, the pain signal in the voice is automatically identified, which has the advantages of non-contact, low cost and real-time, and can significantly improve the efficiency and accuracy of pain assessment. The emotional features in the voice signal are used to assist in judging the patient's pain level, which is particularly suitable for patients who cannot clearly express pain. That is, the following methods are used: (1) Pain score prediction method based on voice emotion analysis: using voice emotion analysis and audio features (MFCC, F0, energy fluctuation) input, and predicting the pain score through the XGBoost regression model. (2) Multi-feature fusion pain assessment model: Fusion of emotional features and audio features (MFCC, F0, etc.) for pain score regression prediction to improve prediction accuracy. (3) Non-invasive real-time pain assessment: Using voice signals for emotion analysis and pain score prediction to achieve real-time, non-invasive pain assessment. Therefore, the use of LSTM and attention mechanism in dynamic emotion-pain relationship modeling can dynamically capture the temporal changes of emotions, especially the impact of recent emotional fluctuations on pain, thereby improving the responsiveness and accuracy of pain scoring. It achieves the advantages of non-contact, low cost and real-time, as well as the efficiency and accuracy of pain assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a block flow diagram of the present invention. DETAILED DESCRIPTION

[0070] The invention will be further described in detail below in conjunction with the accompanying drawings.

[0071] like Figure 1 As shown, a pain recognition system based on speech emotion detection mainly includes a speech acquisition module, a preprocessing module, a speech emotion classification module and a pain recognition module, among which the emotion classification module and the pain recognition module are key components.

[0072] 1. Voice collection module

[0073] The voice acquisition module is the front-end component of the system, responsible for acquiring the original voice signal from the patient. The voice acquisition module is compatible with various types of voice input devices, including but not limited to smartphones, smart speakers, ear-worn devices, and professional medical equipment in hospitals. When the voice acquisition module receives voice input, it first identifies the type of device and performs adaptive configuration according to the technical parameters of the device (such as sampling rate, bit depth, etc.) to ensure that a stable quality voice signal is obtained.

[0074] 2. Speech signal preprocessing module

[0075] After the speech signal is collected, the system first performs a series of preprocessing operations on the signal to remove background noise and extract important features that help with emotion and pain recognition.

[0076] Speech denoising: After speech signal collection is completed, the speech signal will be affected by environmental noise. The preprocessing module uses spectrum subtraction denoising technology to reduce background noise. The implementation method of spectrum subtraction denoising technology is as follows:

[0077] First, the system acquires the noise spectrum in the signal;

[0078] Then, the noise part is removed by the following formula:

[0079] S clean (f) = S noisy (f)-α·N(f)

[0080] Among them, S clean (f) is the denoised signal, S noisy (f) is the original noisy speech signal, N(f) is the spectrum of the noise, and α is the coefficient that controls the strength of noise suppression.

[0081] Framing and windowing: In order to better analyze the local features of the speech signal, the signal is divided into multiple short-time frames, and then a window function (Hamming window) is applied to each frame to reduce spectrum leakage and enhance the accuracy of feature extraction. The formula for short-time Fourier transform is as follows:

[0082]

[0083] Where X(t,f) is the short-time Fourier transform result at time t and frequency f, x[n] is the sample point of the speech signal, and w[nt] is the windowing function.

[0084] Emotional feature extraction: Speech carries a wealth of emotional information, including pitch, volume, speaking speed, intonation, etc. In the present invention, the extraction of emotional features is achieved through the following steps:

[0085] Emotion feature extraction steps:

[0086] (1) Mel-frequency cepstral coefficients (MFCC): This process involves converting the speech signal to the Mel scale, calculating its power spectrum, and obtaining cepstral coefficients through discrete cosine transform. Mel-frequency cepstral coefficients (MFCC) capture the nonlinear response of the human auditory system and are suitable for expressing changes in emotional states.

[0087]

[0088] Among them, Xk is the spectral component of the signal, K is the number of frequency bands, and N is the dimension of MFCC;

[0089] (2) Fundamental frequency F0: The fundamental frequency reflects the pitch of speech and can significantly affect the expression of emotions, especially pain-related emotional fluctuations;

[0090]

[0091] Where T is the size of the time window, fundamental_frequency(t) is the fundamental frequency value at time t;

[0092] (3) Energy feature: The energy intensity of speech reflects the intensity of the speaker’s emotions. Emotional changes caused by pain are usually accompanied by changes in speech energy. Therefore, by calculating the energy value of each short-time frame, the speaker’s emotional state can be further revealed. The calculation of energy features can be completed by summing the squares of all sample points in each frame, that is:

[0093]

[0094] Among them, x(t) is the time domain waveform of the speech signal, and E is the energy.

[0095] 3. Emotion Classification Module

[0096] In order to accurately identify the type and intensity of emotions, especially the continuity of emotions and the persistence of pain, the voice emotion classification module selects the long short-term memory network LSTM as the main time series model, combined with the attention mechanism to solve the problem of the timing of emotion changes. The recent emotion fluctuations have a greater impact on the current emotion discrimination results;

[0097] The LSTM network is based on the fact that emotions are a dynamic process, and the fluctuation of emotions is often closely related to the emotional state of the previous moment. The LSTM network can maintain and update the long-term and short-term memory in the time series through a gating mechanism, so that the emotional state at each moment is not only affected by the current input, but also by the feedback of the historical emotional state.

[0098] The attention mechanism ensures that the module can dynamically assign different importance weights to different time steps when processing time series data; the role of the attention mechanism is to calculate the emotional state h of each time step according to the emotional state h of each time step. t and a query vector q to calculate the weighting coefficient α t , the weighted coefficient indicates which emotional change at the current moment contributes most to the final judgment result; the query vector q can be a context vector summarized from the past emotional state, indicating the evolution trend of the current emotion;

[0099] (1) Calculate attention weights: For each time step t, use the multi-layer perceptron scoring function and nonlinear mapping capabilities to calculate the current emotional state h t and the similarity of the query vector q; specifically, firstly transform the hidden state h at time step t t and query vector q to form a vector x t =[h t ; q], and then the concatenated vector is transformed nonlinearly through a multi-layer perceptron to calculate the score (h t ,q), its expression is:

[0100] score(h t ,q)=W1·σ(W0·x t +b0)+b1

[0101] Among them, W1 and W0 are the weight matrices in the multilayer perceptron, σ(·) is the activation function, b0 and b1 are bias terms, and x t =[h t ; q] is the merged input vector;

[0102] Subsequently, the score values ​​are converted into weights that can be used for the attention mechanism. The score values ​​of all time steps need to be softmax normalized to obtain the attention weight α for each time step. t ;

[0103] The process is as follows:

[0104]

[0105] Through softmax normalization, the weights α of all time steps t will be adjusted to the range of [0,1], and the weight sum is 1; thus, the attention weight α at each time step t It means that in the process of emotion discrimination, the current emotional state h t The relative importance to the final judgment result;

[0106] (2) Weighted summation: Next, based on the calculated attention weight α t , perform weighted summation on the LSTM outputs of all time steps to obtain the weighted emotion feature vector h final , which contains information weighted according to the importance of the emotional evolution:

[0107]

[0108] This weighted emotion feature vector will be used as the final input for emotion discrimination.

[0109] The weighted emotion feature vector output by the LSTM network and the attention mechanism is passed into a fully connected layer to classify the emotion; since there are only two categories of emotion classification, positive and negative, the output of the classifier is a probability value p negative , indicating the probability that the current emotion is negative.

[0110] 4. Pain Recognition Module

[0111] If the pain recognition module classifies the emotion as positive, it will directly output 0. If the emotion is determined to be negative in the emotion classification stage, the system will enter the pain score prediction stage. In this stage, the XGBoost regression model is used to predict the intensity of pain by combining the emotion features and other dynamic features in the speech signal: the score range is 1 to 10. The core task of pain score prediction is to integrate the fluctuation of emotion, the physiological features in the speech, and the dynamic changes, and calculate the pain score through the regression model. (1) Input features:

[0112] The regression model for pain score is trained and predicted based on the following features: probability p of negative emotion negative , emotional feature vector h final , MFCC features, F0, energy fluctuation;

[0113] Therefore, the final input feature vector X becomes:

[0114] X=[p negative ,h final ,MFCC,F0,Energy]

[0115] (2) XGBoost regression model

[0116] The XGBoost regression model is used to predict the intensity P of pain based on these input features, and the output is the predicted value, ranging from 1 to 10. XGBoost fits the regression problem by training a series of decision trees, and the model will improve the prediction performance in an integrated manner.

[0117] The basic formula of XGBoost is:

[0118]

[0119] Where f(X) is the predicted pain score of the model, X is the input feature vector, and θ k is each tree T k The weight of (X), T k (X) is the output of the kth decision tree;

[0120] In XGBoost, each tree T k(X) are all learned by optimizing the loss function, where the loss function L is the mean square error. In order to improve the stability of training and prevent overfitting, XGBoost uses a regularization term Ω to control the complexity of the model.

[0121] The loss function is defined as:

[0122]

[0123] Among them, y i is the actual pain score, f(X i ) is the score predicted by the XGBoost model, and λ is the regularization parameter to avoid overfitting;

[0124] (3) XGBoost training

[0125] The model is trained by gradient boosting. In each iteration, the model updates the tree structure by minimizing the loss function.

[0126] The training process can be described by the following steps:

[0127] A. Initialization: Start with the initial model and initialize it to zero;

[0128] B. Calculate gradient: Calculate the gradient deviation of each sample based on the prediction error of the current model;

[0129] These gradient values ​​reflect the degree of error of each sample under the current model:

[0130]

[0131] C. Build a new tree: Build a new decision tree by calculating the gradient and second-order derivative information to correct the error of the current model.

[0132] D. Update model: Each iteration adjusts the model parameters θ k , by gradually reducing the error, we finally get a powerful model that can predict pain scores.

[0133] (4) Output pain score

[0134] After XGBoost training, the final regression model will be able to predict the pain intensity score P based on the input emotional features, negative emotion probability and audio features; the score P is a numerical value, the larger the score, the more intense the pain.

[0135] The above are only some embodiments of the present invention. For those skilled in the art, several modifications and improvements can be made without departing from the inventive concept of the present invention, which all belong to the protection scope of the invention.

Claims

1. A pain recognition system based on speech emotion detection, characterized in that: The system includes a speech acquisition module for acquiring original speech signals from patients, a speech preprocessing module connected to the speech acquisition module and performing preprocessing operations on the speech signals, a speech emotion classification module connected to the speech preprocessing module and accurately identifying the emotion type and intensity in the speech signals, and a pain recognition module connected to the speech emotion classification module and performing pain scoring after determining the emotion classification.

2. A pain recognition system based on speech emotion detection according to claim 1, characterized in that: The processing method included in the speech preprocessing module is as follows: (1) Speech denoising: After speech signal collection, the speech signal will be affected by environmental noise. The preprocessing module uses spectrum subtraction denoising technology to reduce background noise. (2) Framing and windowing: In order to better analyze the local features of the speech signal, the signal is divided into multiple short-time frames, and then the Hamming window function is applied to each frame to reduce spectral leakage and enhance the accuracy of feature extraction; (3) Emotional feature extraction: Speech carries rich emotional information, including pitch, volume, speaking speed, and intonation.

3. A pain recognition system based on speech emotion detection according to claim 2, characterized in that: The spectral subtraction denoising technology is implemented as follows: First, the system acquires the noise spectrum in the signal; Then, the noise part is removed by the following formula: S clean (f)=S noisy (f)-α·N(f) Among them, S clean (f) is the denoised signal, S noisy (f) is the original noisy speech signal, N(f) is the spectrum of the noise, and α is the coefficient that controls the strength of noise suppression.

4. The pain recognition system based on speech emotion detection according to claim 2, characterized in that: The formula for the short-time Fourier transform is as follows: Where X(t,f) is the short-time Fourier transform result at time t and frequency f, x[n] is the sample point of the speech signal, and w[nt] is the windowing function.

5. The pain recognition system based on speech emotion detection according to claim 2, characterized in that: The emotion feature extraction step: (1) Mel-frequency cepstral coefficients (MFCC): This process involves converting the speech signal to the Mel scale, calculating its power spectrum, and obtaining cepstral coefficients through discrete cosine transform. Mel-frequency cepstral coefficients (MFCC) capture the nonlinear response of the human auditory system and are suitable for expressing changes in emotional states. Among them, X k is the spectral component of the signal, K is the number of frequency bands, and N is the dimension of MFCC; (2) Fundamental frequency F0: The fundamental frequency reflects the pitch of speech and can significantly affect the expression of emotions, especially pain-related emotional fluctuations; Where T is the size of the time window, fundamental_frequency(t) is the fundamental frequency value at time t; (3) Energy feature: The energy intensity of speech reflects the intensity of the speaker’s emotions. Emotional changes caused by pain are usually accompanied by changes in speech energy. Therefore, by calculating the energy value of each short-time frame, the speaker’s emotional state can be further revealed. The calculation of energy features can be completed by summing the squares of all sample points in each frame, that is: Among them, x(t) is the time domain waveform of the speech signal, and E is the energy.

6. The pain recognition system based on speech emotion detection according to claim 2, characterized in that: In order to accurately identify the type and intensity of emotions, especially the continuity of emotions and the persistence of pain, the speech emotion classification module selects the long short-term memory network LSTM as the main time series model, combined with the attention mechanism to solve the problem of the timing of emotion changes. The recent emotion fluctuations have a greater impact on the current emotion discrimination results; The LSTM network is because emotions are a dynamic process, and the fluctuation of emotions is often closely related to the emotional state of the previous moment; the LSTM network can maintain and update the long-term and short-term memory in the time series through a gating mechanism, so that the emotional state at each moment is not only affected by the current input, but also by the feedback of the historical emotional state; The attention mechanism is to ensure that the module can dynamically assign different importance weights to different time steps when processing time series data; the role of the attention mechanism is to calculate the emotional state h of each time step according to the emotional state h of each time step. t and a query vector q to calculate the weighting coefficient α t , the weighted coefficient indicates which emotional change at the current moment contributes the most to the final judgment result; This query vector q can be a context vector summarized from past emotional states, indicating the evolution trend of current emotions; (1) Calculate attention weights: For each time step t, use the multi-layer perceptron scoring function and nonlinear mapping capabilities to calculate the current emotional state h t and the similarity of the query vector q; specifically, firstly transform the hidden state h at time step t t and query vector q to form a vector x t =[h t ; q], and then the concatenated vector is transformed nonlinearly through a multi-layer perceptron to calculate the score (h t ,q), its expression is: score(h t ,q)=W1·σ(W0·x t +b0)+b1 Among them, W1 and W0 are the weight matrices in the multilayer perceptron, σ(·) is the activation function, b0 and b1 are bias terms, and x t =[h t ; q] is the merged input vector; Subsequently, the score values ​​are converted into weights that can be used for the attention mechanism. The score values ​​of all time steps need to be softmax normalized to obtain the attention weight α for each time step. t ; The process is as follows: Through softmax normalization, the weight α of all time steps t will be adjusted to the range of [0,1], and the weight sum is 1; thus, the attention weight α at each time step t It means that in the process of emotion discrimination, the current emotional state h t The relative importance to the final judgment result; (2) Weighted summation: Next, based on the calculated attention weight α t , perform weighted summation on the LSTM outputs of all time steps to obtain the weighted emotion feature vector h final , which contains information weighted according to the importance of the emotional evolution: This weighted emotion feature vector will be used as the final input for emotion discrimination; The weighted emotion feature vector output by the LSTM network and the attention mechanism is passed into a fully connected layer to classify the emotion; since there are only two categories of emotion classification, positive and negative, the output of the classifier is a probability value p negative , indicating the probability that the current emotion is negative.

7. The pain recognition system based on speech emotion detection according to claim 6, characterized in that: The LSTM network is set as a bidirectional LSTM network, which means that the LSTM network can not only understand the past emotional evolution, but also capture the potential information of future emotions; the LSTM network outputs h at each time step. t Represents the emotional state of the current time step, which will be passed to the subsequent attention mechanism to calculate the weighted representation of the emotion.

8. The pain recognition system based on speech emotion detection according to claim 1, characterized in that: If the pain recognition module classifies the emotion as positive, it will directly output 0; after the emotion is determined to be negative in the emotion classification stage, the system will enter the pain score prediction stage; in this stage, the XGBoost regression model is used to predict the intensity of pain by combining the emotion features and other dynamic features in the speech signal: the score range is 1 to 10; the core task of pain score prediction is to combine the fluctuation of emotion, the physiological features in the speech, and the dynamic changes, and calculate the pain score through the regression model; (1) Input features: The regression model for pain score is trained and predicted based on the following features: probability p of negative emotion negative , emotional feature vector h final , MFCC features, F0, energy fluctuation; Therefore, the final input feature vector X becomes: X=[p negative ,h final ,MFCC,F0,Energy] (2) XGBoost regression model The XGBoost regression model is used to predict the intensity P of pain based on these input features, and the output is the predicted value, ranging from 1 to 10; XGBoost fits the regression problem by training a series of decision trees, and the model will improve the prediction performance through integration; The basic formula of XGBoost is: Where f(X) is the predicted pain score of the model, X is the input feature vector, and θ k is each tree T k The weight of (X), T k (X) is the output of the kth decision tree; In XGBoost, each tree T k (X) is learned by optimizing the loss function, where L is the mean square error. In order to improve the stability of training and prevent overfitting, XGBoost uses a regularization term Ω to control the complexity of the model. The loss function is defined as: Among them, y i is the actual pain score, f(X i ) is the score predicted by the XGBoost model, and λ is the regularization parameter to avoid overfitting; (3) XGBoost training The model is trained by gradient boosting. In each iteration, the model updates the tree structure by minimizing the loss function. The training process can be described by the following steps: A. Initialization: Start with the initial model and initialize it to zero; B. Calculate gradient: Calculate the gradient deviation of each sample based on the prediction error of the current model; These gradient values ​​reflect the degree of error of each sample under the current model: C. Build a new tree: Build a new decision tree by calculating the gradient and second-order derivative information to correct the error of the current model; D. Update model: Each iteration adjusts the model parameters θ k , by gradually reducing the error, we finally get a powerful model that can predict pain scores; (4) Output pain score After XGBoost training, the final regression model will be able to predict the pain intensity score P based on the input emotional features, negative emotion probability and audio features; the score P is a numerical value, the larger the score, the more intense the pain.

9. The pain recognition system based on speech emotion detection according to claim 1, characterized in that: The voice collection module is a microphone for collecting ambient audio signals, and the microphone is connected to a smart phone, a smart speaker, an ear-worn device or a medical device.