A method for classifying emotions in the elderly, computer equipment and media

By inputting the fusion features of facial expression and voice data into the emotion classification model, the problems of accuracy and applicability of emotion classification for the elderly are solved, and the accurate classification and interpretable analysis of the emotions of the elderly are achieved.

CN116933129BActive Publication Date: 2025-10-28PEKING UNION MEDICAL COLLEGE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310869023.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-14
Publication Date
2025-10-28
Estimated Expiration
2043-07-14

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately distinguish and classify emotions such as depression, anxiety, and apathy in older adults, leading to inappropriate interventions. Furthermore, existing emotion classification models are not applicable to older populations and lack interpretability.

Method used

By acquiring video and audio data of facial expressions from elderly individuals, facial expression features, audio features, and text features are extracted, fused, and then input into a trained emotion classification model. The model is then visualized using the SHAP interpretation tool to achieve emotion classification.

Benefits of technology

It achieves accurate classification of emotions in the elderly, including normal, depression only, anxiety only, apathy only, depression combined with anxiety, depression combined with apathy, and depression, anxiety and apathy combined, thus improving the interpretability and applicability of the classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116933129B_ABST
    Figure CN116933129B_ABST
Patent Text Reader

Abstract

This invention discloses a method, computer device, and medium for classifying emotions in the elderly, relating to the fields of speech processing and image processing. The method includes: acquiring facial expression video data and speech data of an elderly person whose emotion needs to be identified; extracting features from the facial expression video data to obtain facial expression features; the facial expression features include the presence and intensity representation of each specific facial movement; extracting features from the speech data to obtain speech features and text features; fusing the facial expression features, speech features, and text features to obtain fused features; and inputting the fused features into a trained emotion classification prediction model to obtain emotion classification results; the emotion classification results include normal, depression only, anxiety only, apathy only, depression combined with anxiety, depression combined with apathy, and depression, anxiety, and apathy combined. Through the above method, this invention achieves the identification and classification of emotions in the elderly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of speech processing and image processing, and in particular to a method for classifying the emotions of the elderly, a computer device, and a medium. Background Technology

[0002] As the population ages, mental health has become a crucial aspect of promoting healthy aging. Along with the decline in various physical functions, older adults often experience mental and psychological problems, such as depression, anxiety, and apathy. These emotional issues can lead to a decrease in the quality of life and cognitive function in the elderly, and an increased burden on caregivers. Therefore, addressing the emotional problems of the elderly is a pressing practical need.

[0003] Depression, anxiety, and apathy are often overlooked in the elderly, with caregivers only noticing and reporting them to doctors when symptoms become severe. These three emotional problems often overlap and intersect, making them difficult to distinguish clinically. Intervention strategies differ for different emotional problems; misusing antidepressants for apathy can worsen symptoms. Therefore, accurate classification of emotional problems is crucial for providing targeted psychological interventions. Summary of the Invention

[0004] The purpose of this invention is to provide a method, computer device, and medium for classifying the emotions of the elderly.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A method for classifying the emotions of the elderly, the method comprising:

[0007] Acquire video and voice data of facial expressions of elderly individuals whose emotions are to be identified;

[0008] Feature extraction is performed on the facial expression video data to obtain facial expression features; the facial expression features include the presence representation and intensity representation of each specific facial action, the presence representation is used to indicate whether the specific facial action exists, and the intensity representation is used to indicate the intensity of the specific facial action;

[0009] Feature extraction is performed on the speech data to obtain speech features and text features; the speech features include frequency features, energy features, spectral features, and time-related speech features; the text features are obtained by performing sentiment analysis on the speech data;

[0010] The facial expression features, speech features, and text features are fused to obtain fused features;

[0011] The fused features are input into the trained emotion classification prediction model to obtain the emotion classification results; the emotion classification results include normal, depression only, anxiety only, apathy only, depression combined with anxiety, depression combined with apathy, and depression, anxiety, and apathy combined; the trained emotion classification model is a model trained with the sample fused features as input and the sample emotion classification results corresponding to the sample fused features as labels.

[0012] Optionally, the facial expression video data is the facial video data of the elderly person whose emotion is to be identified, captured by a camera in a natural state, with the elderly person maintaining a straight gaze forward.

[0013] Optionally, the voice data is the voice data of the elderly person whose emotion is to be identified, collected when the elderly person performs the picture description task.

[0014] Optionally, the time-related speech features include time features related to speech rate.

[0015] Optionally, the time-related speech features may also include speech duration, pause duration, and speech ratio.

[0016] Optionally, before inputting the fused features into the trained emotion classification model, the method further includes:

[0017] The facial expression features, speech features, and text features in the fused features are normalized to eliminate the differences between different features.

[0018] Optionally, the method further includes:

[0019] The model constructs a first emotion classification model based on logistic regression, a second emotion classification model based on random forest, a third emotion classification model based on support vector machine, a fourth emotion classification model based on K nearest neighbor, a fifth emotion classification model based on Naive Bayes, a sixth emotion classification model based on gradient boosting machine, and a seventh emotion classification model based on extreme gradient boosting machine.

[0020] The first emotion classification model, the second emotion classification model, the third emotion classification model, the fourth emotion classification model, the fifth emotion classification model, the sixth emotion classification model, and the seventh emotion classification model are trained based on the training sample set, and the emotion classification model with the highest accuracy is selected as the emotion classification prediction model; the training sample set includes several sample fusion features and the emotion classification label corresponding to each sample fusion feature.

[0021] Optionally, after obtaining the emotion classification result, the method further includes:

[0022] The SHAP interpretation tool was used to visualize and analyze the emotion classification results.

[0023] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the above-described method for classifying the emotions of the elderly.

[0024] The present invention also provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor using the above-described method for classifying the emotions of the elderly.

[0025] According to specific embodiments provided by the present invention, the following technical effects are disclosed: The present invention provides a method, computer device, and medium for classifying emotions in the elderly. The method includes: acquiring facial expression video data and speech data of an elderly person whose emotions are to be identified; extracting features from the facial expression video data to obtain facial expression features; the facial expression features include an existence representation and an intensity representation for each specific facial movement, wherein the existence representation indicates whether a specific facial movement exists, and the intensity representation indicates the intensity of a specific facial movement; extracting features from the speech data to obtain speech features and text features; the speech features include frequency features, energy features, spectral features, and time-related speech features; the text features are obtained through emotion analysis of the speech data; fusing the facial expression features, speech features, and text features to obtain fused features; and inputting the fused features into a trained emotion classification prediction model to obtain the emotion classification result. Through the above method, the present invention achieves the identification and classification of emotions in the elderly. Attached Figure Description

[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0027] Figure 1 This is a schematic diagram of the process for classifying the emotions of the elderly provided in an embodiment of the present invention;

[0028] Figure 2 A schematic diagram of a specific facial expression unit provided in an embodiment of the present invention;

[0029] Figure 3 A schematic diagram illustrating the determination of the emotion classification prediction model and the visualization interpretation process provided in this embodiment of the invention;

[0030] Figure 4 This is a schematic diagram of the structure of a computer device provided by the present invention.

[0031] Symbol explanation:

[0032] 1000 - Computer equipment; 1001 - Processor; 1002 - Communication bus; 1003 - User interface; 1004 - Network interface; 1005 - Memory. Detailed Implementation

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0034] The purpose of this invention is to provide a method, computer device, and medium for classifying emotions in the elderly. By extracting facial expression features, speech features, and text features from video and speech data of elderly individuals whose emotions need to be identified, and fusing these features to obtain fused features, a trained emotion classification prediction model is used to predict the emotion classification result, thus achieving the identification and classification of emotions in the elderly. Furthermore, the emotion classification results in this invention include normal, depression only, anxiety only, apathy only, depression combined with anxiety, depression combined with apathy, and a combination of depression, anxiety, and apathy, taking into account the complexity of emotional issues in real-world situations. Existing emotion classification models are mostly designed for middle-aged and young adults, but the speech and facial expressions of the elderly differ from those of other age groups, making existing models unsuitable for classifying emotions in the elderly. Compared to existing machine learning emotion classification models, which are often unpredictable and difficult to interpret, this invention utilizes the SHAP interpretation tool to visualize and analyze the emotion classification results.

[0035] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0036] like Figure 1 As shown, the present invention provides a method for classifying the emotions of the elderly, the method comprising:

[0037] S1: Obtain video and audio data of facial expressions of elderly people whose emotions are to be identified.

[0038] S2: Extract features from the facial expression video data to obtain facial expression features; the facial expression features include the presence representation and intensity representation of each specific facial action, the presence representation is used to indicate whether the specific facial action exists, and the intensity representation is used to indicate the intensity of the specific facial action.

[0039] S3: Feature extraction is performed on the speech data to obtain speech features and text features; the speech features include frequency features, energy features, spectral features, and time-related speech features; the text features are obtained through sentiment analysis of the speech data. The time-related speech features include time features related to speech rate. The time-related speech features also include speech duration, pause duration, and speech ratio.

[0040] S4: The facial expression features, the voice features, and the text features are fused to obtain fused features.

[0041] S5: Input the fusion features into the trained emotion classification prediction model to obtain the emotion classification results; the emotion classification results include normal, depression only, anxiety only, apathy only, depression combined with anxiety, depression combined with apathy, and depression, anxiety, and apathy combined; the trained emotion classification model is a model trained with the sample fusion features as input and the sample emotion classification results corresponding to the sample fusion features as labels.

[0042] In S1, the facial expression video data is the facial video data captured by the camera of the elderly person whose emotion is to be identified, while they are looking straight ahead in a natural state. Specifically, in this embodiment, each elderly person is required to look straight ahead for 60 seconds, keeping their head as still as possible. The USB network camera is placed 70 centimeters directly in front of the elderly person whose emotion is to be identified, and the position is kept fixed. The captured facial video data is recorded at 25 frames per second, with a resolution of 640×480 pixels, and saved in WAV format.

[0043] The voice data refers to the voice data of the elderly person whose emotion is to be identified, collected when performing the picture description task. The microphone is set 20 cm in front of the elderly person whose emotion is to be identified, and the voice data is recorded at 48 kHz, 16-bit, and saved in WAV format in the recording device (iFlytek, SR702).

[0044] For facial expression video data, the publicly available facial behavior analysis toolkit OpenFace 3.0 was used to extract facial expression features. OpenFace 3.0 is developed based on the internationally recognized Facial Action Coding System (FACS), which classifies facial movements by facial appearance and can deconstruct the code of any anatomically oriented facial expression into the specific facial action unit (AU) that produces that expression. In this embodiment, 19 facial action units were extracted, and the activation (presence) and intensity of each specific facial action unit were calculated. The information of the 19 facial actions is shown in Table 1, and the schematic diagrams of each specific facial action unit are shown in the following order: Figure 2 (a)-(s). The presence and intensity of AUs (Activities of Interest) in each frame of facial expression video data of elderly individuals to be identified are provided using OpenFace 3.0. The presence of an AU is encoded as: 0 indicates absence, and 1 indicates presence. The intensity of an AU is rated on a 5-point scale, ranging from 1 to 5, where 1 is the minimum intensity and 5 is the maximum intensity. After obtaining the presence and intensity of each AU in each frame of the facial expression video data, the average presence and intensity of each AU during the task are calculated. The average presence value of each AU is used as the presence representation of each AU, and the average intensity value of each AU is used as the intensity representation of each AU.

[0045] Table 1 Facial Action Coding and Emotional Valence

[0046]

[0047]

[0048] In Table 1, a represents negative; b represents positive; c represents positive or negative; and d represents neutral.

[0049] For the aforementioned speech data, based on the OpenSmile toolkit, the Extended Version Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) was used to extract speech features. eGeMAPS is highly accurate in speech emotion recognition tasks and has shown stable performance across different databases. This embodiment extracted four types of standard acoustic features, including frequency features (such as pitch, jitter, and formants), energy features (such as loudness and noise ratio), spectral features (such as alpha ratio, spectral slopes of 0-500Hz and 500-1500Hz, and relative energies of formants 1, 2, and 3), and time features related to speech rate. Additionally, this embodiment used PyDub to extract six time-related speech features (temporal domain features), including speech duration, pause duration, and speech ratio. These time-related speech features, along with the speech rate-related time features, constitute the time-related speech features. This embodiment provides an overview, definition, and functional explanation of the extracted speech features, as shown in Table 2.

[0050] This embodiment also extracts text features from the speech data and uses Sentiment Knowledge Enhanced Pre-Training for Sentiment Analysis (SKEP) on a publicly available database to perform sentiment judgment on each image description. SKEP has been validated with good accuracy on 14 publicly available Chinese and English databases and can perform sentiment analysis on different tasks such as words, sentences, and paragraphs. Based on the image description task, each task corresponds to a specific text feature. In this embodiment, the text features are represented by 1-6, meaning that the corresponding text features are determined based on the task described by the elderly person whose emotion is to be identified, as detailed in Table 2.

[0051] Table 2 Overview and Explanation of Speech Features

[0052]

[0053]

[0054]

[0055]

[0056] Before inputting the fused features into the trained emotion classification model, this embodiment further includes: normalizing the facial expression features, speech features, and text features in the fused features to eliminate the differences between different features.

[0057] Specifically: After extracting facial expression features, speech features, and text features, the facial, speech, and text features are merged into a single feature set through a simple concatenation, resulting in fused features. Then, z-score normalization is used to normalize the facial expression features, speech features, and text features in the fused features to eliminate differences between different features.

[0058] like Figure 2 As shown, the method for classifying the emotions of the elderly also includes:

[0059] The model constructs a first emotion classification model based on logistic regression, a second emotion classification model based on random forest, a third emotion classification model based on support vector machine, a fourth emotion classification model based on K nearest neighbor, a fifth emotion classification model based on Naive Bayes, a sixth emotion classification model based on gradient boosting machine, and a seventh emotion classification model based on extreme gradient boosting machine.

[0060] The first emotion classification model, the second emotion classification model, the third emotion classification model, the fourth emotion classification model, the fifth emotion classification model, the sixth emotion classification model, and the seventh emotion classification model are trained based on the training sample set, and the emotion classification model with the highest accuracy is selected as the emotion classification prediction model; the training sample set includes several sample fusion features and the emotion classification label corresponding to each sample fusion feature.

[0061] Specifically: First, acquire the speech data and facial expression video data required for training:

[0062] For each elderly individual, the activation and intensity of 19 facial motion units (AUs) were extracted, as shown in Table 1. Facial expression videos were recorded at 25 frames per second, and OpenFace provided the presence and intensity of AUs for each elderly individual per frame. To obtain a more comprehensive measurement, this embodiment calculated the average intensity and presence representation of each AU across all frames in the facial expression video data. The average intensity and presence representation of all AUs constitute the facial expression features.

[0063] The audio data employed a picture description task, consisting of six images selected from the Chinese Affective Picture System (CFAPS), compiled by the State Key Laboratory of Cognitive Neuroscience and Learning at Beijing Normal University. The images comprised three emotional valences (neutral, positive, and negative) and two image types (facial expressions and scene images). Valence, arousal, and dominance were measured in 30 healthy elderly participants, with results largely consistent with the original population's assessment of the picture materials, indicating the effectiveness of these materials in the elderly population. To ensure recording quality, the microphone was positioned 20 cm directly in front of the participants. All audio files were recorded at 48 kHz, 16-bit, and saved in WAV format on the recording device (iFlytek SR702).

[0064] The above audio data is preprocessed to obtain a speech signal, and speech features (frequency features, energy features, spectral features, and temporal features) and text features are obtained.

[0065] Voice features, text features, and facial expression features are fused to obtain fused features, thus forming a training sample set. These fused features are then input into seven machine learning models for training, and the models are tuned using hyperparameters. Based on model evaluation metrics, the best model is selected from the seven machine learning models. This process includes three parts: feature selection, emotion grouping and imbalanced data processing, and the construction of an emotion classification prediction model. These three parts are described in detail below:

[0066] (1) Feature selection

[0067] After extracting facial expression features, speech features, and text features, these features are merged into a single feature set through a simple connection; this feature set is called the fused feature set. Then, z-score normalization is used to normalize the feature set. Normalization can eliminate differences between different features and also reduce the computational cost and training time of the model.

[0068] For the fused features, feature selection is performed. Feature selection is a necessary step in constructing an emotion classification model. A modality has several features, and it is unknown which of these features are useful for identifying emotion problems, thus requiring the extraction of a large number of features. However, to obtain statistically correct and reliable results, the amount of data required to support these results usually increases exponentially with the dimensionality, while the number of samples available in practice is often limited; this phenomenon is called the curse of dimensionality. To avoid the curse of dimensionality, the original feature set needs to be dimensionality reduced. Feature selection involves choosing a subset of features with good classification performance and relatively representative features from the original data dimensions. In this embodiment, feature selection is used to remove highly correlated features, and the candidate library of speech, text, and facial expression features for elderly emotion problems is reorganized, simplified, and optimized to a certain extent, laying the foundation for the next step of constructing an emotion classification model for the elderly.

[0069] First, this embodiment employs classic statistical methods—correlation analysis (Pearson, Spearman)—to identify the features associated with depression, anxiety, and apathy for each modality. For features with statistically insignificant correlations or controversial features, due to the large number of features, this embodiment uses machine learning algorithms to further mine the objective data. To overcome the method selection bias introduced by different feature selection methods in machine learning, this embodiment uses three feature selection methods: Lasso (linear algorithm), Boruta (correlation), and RFS (nonlinear algorithm). After obtaining the features selected by the three methods, the union of the selected features is taken. Based on this, considering the features selected by correlation analysis and machine learning, an emotion classification model for depression, anxiety, apathy, and their combined states is constructed, thereby eliminating invalid features.

[0070] (2) Emotion grouping and imbalanced data processing

[0071] Because human emotions fluctuate, in addition to collecting objective data (voice data and facial expression video data) on the day of the experiment, it was also necessary to collect subjective emotion scales, that is, to collect emotion scales for each elderly person. Specifically:

[0072] Depression, anxiety, and apathy were measured using the Public Health Questionnaire-9 (PHQ-9), the Generalized Anxiety Disorder Scale (GAD-7), and the Apathy Evaluation Scale (AES), respectively. Emotional classification labels for each sample were obtained using these scales.

[0073] 1) Depression measurement tools

[0074] The Patient Health Questionnaire Depression Scale was used to assess depressive mood. Developed by American scholars Spitzer et al. and translated into Chinese by Bian Cuidong, the scale is an important tool for screening and assessing depression. It contains nine items and evaluates participants' feelings over the past two weeks. The scale uses a 4-point Likert scale, with scores from 0 to 3 representing "nearly every day," and a total score of 0-27. Scores of 5, 10, 15, and 20 represent "mild," "moderate," "severe," and "profound" depression, respectively. The scale's Cronbach's α and test-retest reliability in the Chinese population were both 0.86.

[0075] 2) Anxiety measurement tools

[0076] Anxiety levels were measured using the Generalized Anxiety Scale (GAS). Developed by Spitzer et al. and translated into Chinese by He Xiaoyan et al., the scale contains seven items to assess patients' emotional distress over the past two weeks. A Likert scale of 4 points was used, with scores from 0 to 3 representing "never" to "almost daily." The total score ranged from 0 to 21, with scores of 5, 10, and 15 representing "mild," "moderate," and "severe" anxiety, respectively. The scale's sensitivity and specificity for identifying anxiety were 89.0% and 82.0%, respectively.

[0077] 3) Indifference measurement tools

[0078] Apathy was measured using the Apathy Evaluation Scale (AES-C), developed by Marin and translated into Chinese by Dong Shuhui et al. This scale includes three versions: Clinical Version (AES-C), Patient Version (AES-S), and Informant Version (AES-I). Studies showed that the Clinical Version was most sensitive to patients with MCI (Minor Inflammatory Disease), therefore, the Clinical Version (AES-C) was used in this embodiment. The scale consists of 18 items, assessing three dimensions: interests, activities, and daily life over the past four weeks. A Likert scale of 4 was used, with "1" representing "not at all" to "very much." The scale score ranged from 0 to 72 points; a score ≥35 was considered apathy; the higher the score, the more severe the apathy. The AES-C is a well-validated scale that effectively distinguishes between apathy and depression in older adults. The scale has a sensitivity of 75.3% and a specificity of 75.8%. Cronbach's α is 0.85, the split-half coefficient is 0.84, and the content validity is 0.94.

[0079] Based on the cutoff scores of the depression, anxiety, and apathy assessment scales, the emotional cutoff scores of the elderly were divided into seven emotion groups: normal, depression only, anxiety only, apathy only, depression with anxiety, depression with apathy, and depression, anxiety, and apathy combined. Based on these emotion groupings, emotion classification labels were assigned to the fusion features in the training sample set to form a complete training sample set, which was then used to train the constructed emotion classification prediction model.

[0080] It should be noted that due to the imbalance in the number of participants in each emotion group, this embodiment also employs undersampling and oversampling techniques to process the sample data in the training sample set. Undersampling randomly selects some samples from the majority sample and removes them, but it has the drawback that the excluded samples may contain important information, resulting in poor model performance. Oversampling is a method of randomly sampling from the minority class to add new samples, but it reduces the variability of the samples. T-link (Tomek Link) undersampling pairs the majority class with the nearest neighbor of another class, forming a T-link, and then removes this pair, creating a clear dividing line. SMOTE (Synthetic Minority Oversampling Technique) oversampling selects a random minority class sample A and its nearest neighbor B, and then randomly selects a point C from the line connecting A and B as a new minority class sample. These two techniques are currently widely used in imbalanced medical data. This embodiment first performs SMOTE oversampling, and then performs T-Link undersampling. The sampling process uses the SMOTE and T-link functions from the Imbalance package, written in Python, to perform oversampling and undersampling respectively to handle the imbalanced data.

[0081] (3) Constructing an emotion classification model

[0082] Multiclass classification is a machine learning method used to handle classification problems with two or more categories. Since depression, anxiety, and apathy have multiple categories, and there are instances of merging emotion problems, this embodiment uses a multiclass classification method to construct an emotion classification model.

[0083] In constructing the emotion classification prediction model, this embodiment employs seven machine learning methods: Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbor (KNN), Naive Bayes (NB), Gradient Boosting Machine (GBM), and Extreme Gradient Boosting (XGBoost).

[0084] Based on the above seven machine learning methods, we construct the first emotion classification model, the second emotion classification model, the third emotion classification model, the fourth emotion classification model, the fifth emotion classification model, the sixth emotion classification model, and the seventh emotion classification model, respectively.

[0085] After constructing the above 7 emotion classification models, the first emotion classification model, the second emotion classification model, the third emotion classification model, the fourth emotion classification model, the fifth emotion classification model, the sixth emotion classification model, and the seventh emotion classification model are trained based on the training sample set, and the emotion classification model with the highest accuracy is selected as the emotion classification prediction model; the training sample set includes several sample fusion features and the emotion classification label corresponding to each sample fusion feature.

[0086] To achieve optimal model performance, hyperparameter optimization is employed for each emotion classification model. This involves selecting the most relevant hyperparameters and iteratively adjusting them to achieve the best model performance, while keeping the remaining values ​​at their default values. This embodiment utilizes the mature hyperparameter optimization package Hyperopt in Python, which can achieve results superior to manual tuning in a relatively short time. To ensure the reliability of the results, 5000 bootstrap samplings are performed to train the classifier, and the hyperparameters of the classification model are determined based on the optimal Area Under the Receiver Operating Characteristic Curve (AUC).

[0087] In machine learning, overfitting degrades a model's predictive performance, typically occurring when the model is overly complex. This example uses 10-fold cross-validation to validate a sentiment classification model. 10-fold cross-validation involves randomly dividing the original data into 10 parts, using 9 parts as the training set and the remaining part as the validation set each time, repeating this process. Cross-validation allows for the adjustment of hyperparameters, and the model evaluation metric is the average of the 10 calculations. It effectively utilizes the data and prevents overfitting.

[0088] For binary classification machine learning models, this embodiment uses accuracy, precision, recall, F1 score, and area under the ROC curve to evaluate the performance of the sentiment classification model. Specifically:

[0089] Accuracy: The percentage of correctly classified samples out of the total number of samples classified. The formula is as follows:

[0090]

[0091] Among them, True Positive (TP): the actual example is positive, but the model classifies it as positive; False Negative (FN): the actual example is positive, but the model classifies it as negative; True Negative (TN): the actual example is negative, but the model classifies it as negative; False Positive (FP): the actual example is negative, but the model classifies it as positive.

[0092] Because accuracy can be affected by sample type imbalance, it is usually evaluated together with other metrics such as precision, recall, and F1 score.

[0093] Precision is defined as the ratio of correctly classified positive samples (true positives) to the total number of samples classified as positive by the model. The formula is as follows:

[0094]

[0095] Recall: Recall, commonly used in the medical field, is the ratio of the number of correctly classified positive samples to the total number of positive samples. Recall measures a model's ability to detect positive samples. The higher the recall, the more positive samples are detected. Its calculation formula is as follows:

[0096]

[0097] Specificity, also known as the true negative rate, is the percentage of individuals who are actually healthy but are correctly diagnosed as healthy by a diagnostic test. The formula is as follows:

[0098]

[0099] F1 score: The F1 score is the harmonic mean of precision and recall, thus providing a comprehensive evaluation of a model's performance in both areas. It represents a balance between achieving high levels of both precision and recall. The calculation formula is as follows:

[0100]

[0101] For multi-class classification models, it is necessary to calculate the average values ​​of multiple classes to measure the overall classification performance of the model. Three methods are used: macro-average, micro-average, and weighted-average. Macro-average refers to the arithmetic mean of each statistical indicator value across all classes; the corresponding Macro F1 score is the arithmetic mean of multiple F1 scores. Classes refer to emotion groups, with one emotion group corresponding to one class. In this example, there are seven classes. Micro-average involves averaging the corresponding elements of each confusion matrix to obtain TP, FP, TN, and FN, calculating the score for each class, and then calculating a weighted average, where each class has the same weight. Weighted-average assigns different weights to different classes based on the number of samples in each class when the samples are imbalanced. In this example, the emotion groups in the emotion problem are imbalanced, therefore, weighted-average is used to evaluate the performance of the multi-class classification model.

[0102] This embodiment uses a weighted F1 score to evaluate the overall model, which is the weighted average of the F1 scores for all categories. For the classification performance of specific emotion groups, the evaluation metrics for binary classification are used: accuracy, precision, recall, and F1 score to evaluate the model, thereby selecting the best emotion classification model as the emotion classification prediction model. The specific implementation process uses Python data analysis packages Scikit-learn, Scipy, NumPy, Pandas, SMOTE, TomeLinks, and Hyperopt.

[0103] Based on the selected emotion classification prediction model, emotion classification is performed to obtain the emotion classification results for elderly people whose emotions are to be identified. This embodiment further includes, after obtaining the emotion classification results:

[0104] The emotion classification results were visualized and analyzed using the SHAP (Shapley Additive Explanation) tool.

[0105] This paper employs the interpretable AI explanation method—SHAP—to explain the sentiment classification prediction model. Inspired by cooperative game theory, an additive explanation model is constructed, where all features are considered "contributors." SHAP provides rich visualization options for calculating and presenting the magnitude and direction of the influence of model features. Compared with other model interpretability methods, it can provide both global and local explanations and has a relatively complete theoretical foundation.

[0106] Based on the previously developed emotion classification prediction model, this embodiment uses SHAP values ​​to interpret the feature importance in the classification model. According to the SHAP values, SHAP bar charts are determined based on their absolute values, and these bar charts sequentially display the feature importance. Furthermore, the SHAP summary plot provides an overview of feature importance and feature effect. Taking the training sample set as an example, for instance... Figure 3 As shown, each point on the summary diagram represents a feature and the SHAP value of an instance; the vertical axis sorts features according to the sum of the SHAP values ​​of all samples, and the horizontal axis represents the SHAP values ​​(representing the distribution of the influence of features on the model output); each point represents a sample, the sample size is stacked vertically, and the color represents the feature value.

[0107] This invention uses a multi-classification machine learning method to address the complex emotional problems of the elderly, achieving accurate classification of their emotions. It employs SHAP to improve the interpretability of the emotion classification model, making it more aligned with clinical practice.

[0108] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the above-described method for classifying the emotions of the elderly.

[0109] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided in this application. For example... Figure 4As shown, computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. Furthermore, computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 4 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0110] exist Figure 4 In the computer device 1000 shown, the network interface 1004 provides network communication functions; the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application stored in the memory 1005 to implement the elderly emotion classification method described in the above embodiments, which will not be described in detail here.

[0111] The present invention also provides a computer-readable storage medium storing a computer program that is adapted to be loaded by a processor and executed by the elderly emotion classification method described in the above embodiments, which will not be described in detail here.

[0112] The above program can be deployed and executed on a single computer device, or deployed and executed on multiple computer devices located in one location, or executed on multiple computer devices distributed across multiple locations and interconnected through a communication network. Multiple computer devices distributed across multiple locations and interconnected through a communication network can form a blockchain network.

[0113] The aforementioned computer-readable storage medium can be an internal storage unit of the computer device, such as a hard drive or memory. It can also be an external storage device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flashcard. Furthermore, the computer-readable storage medium can include both internal and external storage units of the computer device. This computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. It can also be used to temporarily store data that has been output or will be output.

[0114] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0115] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for classifying the emotions of the elderly, characterized in that, The method includes: Facial expression video data and voice data of elderly individuals whose emotions were to be identified were acquired. The facial expression video data consisted of video footage of the elderly individuals' faces while they were looking straight ahead in a natural state. Each elderly individual was required to maintain this gaze for 60 seconds, keeping their head as still as possible. A USB network camera was placed 70 cm directly in front of the elderly individual and kept in a fixed position. The facial video data was recorded at 25 frames per second with a resolution of 640×480 pixels and saved in WAV format. Voice data was collected using a picture description task, which included six pictures selected from the Chinese Emotional Picture System. The pictures consisted of three emotional valences (neutral, positive, and negative) and two image types: facial expression and scene images. Valence, arousal, and dominance were measured in 30 healthy elderly individuals. The results were largely consistent with the original population scores for the picture materials, indicating that the materials were effective among the elderly. To ensure recording quality, the microphone was positioned 20 cm directly in front of the participants. All audio files were recorded at 48kHz, 16-bit and saved in WAV format on the recording device. Feature extraction is performed on the facial expression video data to obtain facial expression features; the facial expression features include the presence representation and intensity representation of each specific facial action, the presence representation is used to indicate whether the specific facial action exists, and the intensity representation is used to indicate the intensity of the specific facial action; The speech data is subjected to feature extraction to obtain speech features and text features; the speech features include frequency features, energy features, spectral features, and time-related speech features; the text features are obtained by performing sentiment analysis on the speech data; the time-related speech features include speech rate-related time features, speech duration, pause duration, and speech ratio. The facial expression features, the speech features, and the text features are fused to obtain the fused features; The model constructs a first emotion classification model based on logistic regression, a second emotion classification model based on random forest, a third emotion classification model based on support vector machine, a fourth emotion classification model based on K nearest neighbor, a fifth emotion classification model based on Naive Bayes, a sixth emotion classification model based on gradient boosting machine, and a seventh emotion classification model based on extreme gradient boosting machine. The first emotion classification model, the second emotion classification model, the third emotion classification model, the fourth emotion classification model, the fifth emotion classification model, the sixth emotion classification model, and the seventh emotion classification model are trained based on the training sample set, and the emotion classification model with the highest accuracy is selected as the emotion classification prediction model; the training sample set includes several sample fusion features and the emotion classification label corresponding to each sample fusion feature; The fused features are input into the trained emotion classification prediction model to obtain the emotion classification results; the emotion classification results include normal, depression only, anxiety only, apathy only, depression combined with anxiety, depression combined with apathy, and depression, anxiety, and apathy combined; the trained emotion classification model is a model trained with the sample fused features as input and the sample emotion classification results corresponding to the sample fused features as labels. Before feature extraction from the facial expression video data, facial motion encoding is set as follows: raising the inner corner of the eyebrow, raising the outer corner of the eyebrow, frowning, raising the upper eyelid, lifting the cheek, tightening the eyelid, wrinkling the nose, raising the upper lip, deepening the nasolabial fold, pulling the corners of the mouth, tightening the corners of the mouth, turning the corners of the mouth down, raising the lower lip, stretching the corners of the mouth, tightening the lips, separating the lips, lowering the chin, pursing the lips, and blinking; using OpenFace Version 3.0 provides the presence and intensity of Active Animations (AUs) in each frame of facial expression video data of elderly people whose emotions are to be identified. The presence of AUs is encoded as follows: 0 indicates absence, and 1 indicates presence. The intensity of AUs is rated on a 5-point scale, ranging from 1 to 5, with 1 being the minimum intensity and 5 being the maximum intensity. The intensity representations corresponding to the facial action codes are as follows: AU1, AU2, AU4, AU5, AU6, AU7, AU9, AU10, AU11, AU12, AU14, AU15, AU17, AU20, AU23, AU25, AU26, AU28, and AU45. The text feature extraction process includes: using publicly available database sentiment knowledge to enhance the pre-trained SKEP and performing sentiment judgment on the description content of each image; For the aforementioned speech data, speech features are extracted using the extended Geneva Minimalist Acoustic Parameter Set eGeMAPS based on the OpenSmile toolkit. The speech features include: frequency-related features, energy-related features, time-domain features, and spectral features; the frequency-related features include: pitch, jitter, formants, and formant 1-3 bandwidth; the energy-related features include: amplitude perturbation, loudness, and signal-to-noise ratio; the time-domain features include: peak loudness per second, audible portion per second, average audible length per second, standard deviation of audible portion per second, average silent portion length, standard deviation of silent portion length, isotonic lines, duration, speech duration, pause duration, speech ratio, speech rate, and speech speed; the spectral features include: alpha ratio, Hammarberg index, relative energy of formants 1-3, spectral slope of 0-500Hz and 500Hz, energy difference of H1-H2 and H1-A3 harmonics, Mel-frequency cepstral coefficients, and spectral flux; Because human emotions fluctuate to some extent, in addition to collecting voice data and facial expression video data on the day of the experiment, subjective emotion scales were also collected, that is, emotion scales were collected for each elderly person; depression, anxiety and apathy were measured by the Patient Health Questionnaire Depression Scale, Generalized Anxiety Scale and Apathy Rating Scale respectively; and emotion classification labels were obtained for each sample.

2. The method for classifying emotions in the elderly according to claim 1, characterized in that, The facial expression video data refers to the facial video data of the elderly person whose emotion is to be identified, captured by the camera while the elderly person is looking straight ahead in a natural state.

3. The method for classifying emotions in the elderly according to claim 1, characterized in that, The voice data refers to the voice data of the elderly person whose emotion is to be identified, collected when the elderly person is performing a picture description task.

4. The method for classifying emotions in the elderly according to claim 1, characterized in that, Before inputting the fused features into the trained emotion classification model, the following steps are also included: The facial expression features, speech features, and text features in the fused features are normalized to eliminate the differences between different features.

5. The method for classifying emotions in the elderly according to claim 1, characterized in that, After obtaining the emotion classification results, the following is also included: The SHAP interpretation tool was used to visualize and analyze the emotion classification results.

6. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method according to any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-mode senior citizen emotion recognition fusion model based on video image facial expressions and voices and establishment method thereof

    CN114582000A