Depression and suicide propensity recognition method fusing limb language, micro-expression and language
Patent Information
- Application Number
- CN202010764410.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-02
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2040-08-02
AI Technical Summary
从技术上讲,音频和视频(F. Xu, J.Zhang and J. Z. Wang, “Microexpression Identification and CategorizationUsing a Facial Dynamics Map,” IEEE Transactions on Affective Computing, vol.8, issue 2, pp. 1-1, 2017.)很容易获得,但容易受到噪声的影响
(1)本发明将多模态数据与文本层对齐。文本中间表示和所提出的融合方法形成了一个融合语音、肢体动作和面部表情的框架。该方法降低了语音、肢体动作和面部表情的维数,将三类信息统一为一个分量。
Smart Images

Figure CN112101097B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of emotion recognition, and specifically relates to a method for recognizing depression and suicidal tendencies by integrating body language, micro-expressions, and language. Background Technology
[0002] Human emotions can be identified in multiple ways, such as electrocardiogram (ECG) and electroencephalogram (EEG) (K. Takahashi, "Remarks on emotion recognition from multi-modal bio-potentialsignals"). Proc. IEEE Int. Conf. Ind. Technol. (ICIT) (See , vol. 3, pp. 1138-1143, Jun. 2004.), speech, facial expressions, etc. Among various emotional signals, physiological signals are widely used in emotion recognition. In recent years, human movement has also become a new feature.
[0003] Traditionally, there are two methods: one is to measure the physiological indicators of an object through contact (J. Kim, and E. André, “Emotion recognition based on physiological changes in music listening,” IEEE Transactions on Pattern Analysis & Machine Intelligence, vol. 30, no. 12, pp. 2067-2083, 2008.), and the other is to observe the physiological characteristics of an object using non-contact methods. In fact, while non-invasive methods are generally preferred, objects can mask their emotions. Technically, audio and video (F. Xu, J. Zhang and JZ Wang, “Microexpression Identification and Categorization Using a Facial Dynamics Map,” IEEE Transactions on Affective Computing, vol. 8, issue 2, pp. 1-1, 2017.) are readily available, but are susceptible to noise. In principle, detecting emotions solely through static posture or dynamic action (H. Wallbott, “Bodily Expressions of Emotion,” European J. Social Psychology, vol. 28, pp. 879-896, 1998; M. Coulson, “Attributing Emotion to Static Body Postures: Recognition Accuracy, Confusions, and Viewpoint Dependence,” J. Nonverbal Behavior, vol. 28, no. 2, pp. 117-139, 2004; J. Burgoon, L. Guerrero, and K. Floyd, Nonverbal Communication. Allyn and Bacon, 2010.) has lower computational complexity but also leads to lower recognition accuracy. Therefore, fusing these features is essential. By fusing multimodal features, the emotion category of the person being tested can be better identified. Summary of the Invention
[0004] To address the above problems, this invention proposes a method for identifying depression and suicidal tendencies by integrating body language, micro-expressions, and language. This method effectively integrates features such as body movements, facial expressions, and language, and categorizes these features for emotion classification. This allows for more efficient and accurate detection of whether a person has depressive symptoms or suicidal intent. First, a Kinect device with an infrared camera is used to collect information such as speech, body movements, and facial expressions. Prosody and spectral features extracted from the speech are used to convert the speech into a text description, including information such as intonation, pitch, and speech rate. Convolutional Neural Networks (CNN) and Bidirectional Long Short-Term Memory Conditional Random Fields (Bi-LSTM-CRF) are used to analyze static and dynamic human movements, respectively, and feature extraction and dimensionality reduction are performed on the facial images. Finally, the speech, body movements, and facial expressions are fused into the text description, and self-organizing maps (SOM) and compensation layers are used to understand the behavior and identify emotions.
[0005] The objective of this invention is achieved by at least one of the following technical solutions.
[0006] A method for identifying depression and suicidal tendencies that integrates body language, micro-expressions, and language includes the following steps: S1. Use a Kinect with an infrared camera to collect video and audio, and convert the video and audio information into feature text descriptions respectively; S2. The feature text descriptions generated in step S1 are fused, and the processing results are classified for sentiment using self-organizing map (SOM) and compensation layer. S3. Mark the individuals who may have depressive symptoms or suicidal tendencies obtained in step S2 and observe them.
[0007] Further, in step S1, the video information includes body movements and facial expression information extracted from the video, and the body movements include static movements and dynamic movements; the audio information includes spectrum, prosody and sound wave information extracted from the speech audio, and the spectrum information and prosody information are used to obtain speech markers, and the sound wave information is used to obtain speech content.
[0008] Furthermore, in step S1, the extraction of feature text descriptions specifically includes the following steps: S1.1. Use a convolutional neural network (CNN) to identify static motion and generate a text description of the static motion features; S1.2 Utilize Kinect to detect human skeletal data in real time, calculate human behavioral characteristics, complete dynamic motion recognition, and generate dynamic motion feature text descriptions. S1.3. Use local recognition methods to identify facial expressions and generate text descriptions of facial activity features; S1.4 Complete the recognition of speech markers and speech content, and generate a text description of language features.
[0009] Further, in step S1.1, single frames are selected from the collected video and input into a convolutional neural network (CNN) for training and testing; all individual frames from the video are input into the trained CNN to obtain static motion with emotional features, and the static motion with emotional features is input into a softmax classifier for classification to complete the recognition of static motion; the softmax function is calculated as follows: (1) in, For the first i The weight matrix of the feature text. b Represents bias.
[0010] Furthermore, Convolutional Neural Networks (CNNs) utilize partial filters to compute convolutions, that is, they perform inner product operations using local submatrices of the input terms and local filters, and the output is a convolution matrix; the hidden layers in a Convolutional Neural Network (CNN) include two convolutional layers and two pooling layers; The formula for calculating convolutional layers is as follows: (2) in, l Indicates the first l One convolutional layer, i The convolution output matrix represents the first... i The value of each component. j This indicates the number of corresponding output matrices; j The value of varies between 0 and N, where N represents the number of convolution output matrices; f It is non-linear. sigmoid Type function; The pooling layer uses average pooling. The input to the average pooling layer comes from the previous convolutional layer, and the output is used as the input to the next convolutional layer. The calculation formula is as follows: (3) in, This represents the local output after the pooling process is complete, derived from the mean of the local small matrix of size n×n in the previous layer.
[0011] Further, in step S1.2, firstly, human body localization and tracking are completed using Kinect to obtain the joints of the skeleton; the 15 skeleton joints are numbered from top to bottom and from left to right; since the position signal of the skeleton is time-varying, their definition is unclear when encountering occlusion, so frame sequences are extracted from the video and input into interval Kalman filtering to improve the accuracy of the skeleton position; then, a bidirectional long short-term memory network (Bi-LSTM-CRF) with a conditional random field layer is used to analyze the motion sequence of the 15 skeleton points respectively to obtain dynamic motion with emotional features; For a bidirectional long short-term memory neural network, given an input sequence Where t represents the t-th coordinate, and T represents a total of T coordinates, the output calculation formula of the hidden layer of the Long Short-Term Memory Neural Network is as follows: (4) in, This represents the output of the hidden layer at time t. This is the weight matrix from the input layer to the hidden layer. This is the weight matrix from hidden layer to hidden layer. For the bias of the hidden layer, This represents the activation function; although LSTM can capture information from long-term sequences, it only considers one direction. Bi-LSTM is used to strengthen this two-way relationship, with the first layer being a forward LSTM and the second layer a backward LSTM. Finally, the dynamic motion with emotional characteristics is input into the Softmax classifier in step S1.1 for classification.
[0012] Further, in step S1.3, based on the information of the face image frames captured by Kinect, the various segmented regions of the face are obtained; the original images of each segmented part of the face are processed into normalized standard images, and two-dimensional Gabor wavelets are used for feature extraction. Linear Discriminant Analysis (LDA) algorithm is used for dimensionality reduction to extract the most discriminative low-dimensional features from the high-dimensional feature space. Based on the extracted low-dimensional features, all samples of the same category are collected to separate other samples as much as possible, that is, the features with the largest ratio of dispersion between sample classes to dispersion within sample classes are selected; finally, the face image frames after feature extraction by Gabor wavelets and dimensionality reduction by LDA are classified using an open-source OpenFace neural network to obtain the facial expression recognition results.
[0013] Further, in step S1.4, firstly, the Kinect directly collects speech, and the collected speech is used to reduce noise using a Wiener-based noise filter. Then, the noise-reduced speech is input into a backpropagation neural network (BPNN) (BPNN is a type of feedforward neural network. BPNN adds a backpropagation algorithm to the structure of a feedforward network) for training to obtain speech with prosodic and spectral features. Finally, the speech with prosodic and spectral features is input into a Softmax classifier for classification to obtain the speech recognition result.
[0014] Furthermore, step S2 specifically includes the following steps: S2.1. Use an LSTM neural network to embed the feature text descriptions collected in step S1 into a fixed-size feature vector arranged in chronological order; the LSTM neural network is the Bi-LSTM feedforward LSTM network in step 1.2. S2.2. The self-organization mapping (SOM) algorithm is used to normalize the feature vectors in step S2.1. S2.3. Since the Self-Organizing Map (SOM) layer is an ambiguous layer that loses information, a compensation layer is used to compensate for the information loss. That is, for different classification results in the SOM, the compensation layer must be combined with a specific layer. The size of each layer is the same as the competition layer of the SOM network, and all nodes have their own weights. ,s represents the s-th class of the corresponding compensation layer, and t represents The t-th node within the layer; compensation layers are not shared, each layer corresponds to the SOM output of a specific type of classification; the formula for the multiplication result is as follows: (5) In the formula, These are the input weights of the s-th layer node. Let be the weight of the node at layer s. This is used to limit the compensation ratio between -1 and 1; Since similar characteristics may belong to the same class, a global optimization should be performed on the target result. The global optimization objective is: (6) The first item for: (7) The first item This is used to minimize the error between the label and the prediction result. The second term is... This is used to minimize the error between the labels and the SOM network results. (Third term) This is used to minimize the difference between the input signal and the output signal of the SOM network.
[0015] Furthermore, in step S3, based on the output of step S2, the likelihood of depressive mood and suicidal tendency is determined, high-risk individuals are marked, observed, and given certain psychological counseling.
[0016] Compared with the prior art, the present invention has the following advantages: (1) This invention aligns multimodal data with the text layer. The intermediate text representation and the proposed fusion method form a framework that integrates speech, body language, and facial expressions. This method reduces the dimensionality of speech, body language, and facial expressions, unifying the three types of information into a single component.
[0017] (2) Depth information enhances the robustness and accuracy of motion detection.
[0018] (3) The present invention takes into account both static and dynamic body movements, thus achieving higher efficiency.
[0019] (4) This invention utilizes Kinect for data acquisition, which is non-invasive, high-performance, and easy to operate. Attached Figure Description
[0020] Figure 1 This is a flowchart of a method for identifying depression and suicidal tendencies that integrates body language, micro-expressions, and language, as described in an embodiment of the present invention. Detailed Implementation
[0021] The specific implementation of the present invention will be further described below with reference to the embodiments and accompanying drawings, but the implementation of the present invention is not limited thereto.
[0022] Example: Methods for identifying depression and suicidal tendencies that integrate body language, micro-expressions, and language, such as Figure 1 As shown, it includes the following steps: S1. Use a Kinect with an infrared camera to collect video and audio, and convert the video information and audio information into feature text descriptions respectively; the video information includes body movements and facial expression information extracted from the video, and the body movements include static movements and dynamic movements; the audio information includes spectrum, prosody and sound wave information extracted from the speech audio, the spectrum information and prosody information are used to obtain speech markers, and the sound wave information is used to obtain speech content.
[0023] The extraction of feature text descriptions specifically includes the following steps: S1.1. Use a convolutional neural network (CNN) to identify static motion and generate a text description of the static motion features; Single frames are selected from the collected videos and input into a convolutional neural network (CNN) for training and testing. All individual frames from the videos are then input into the trained CNN to obtain static motion with emotional features. This static motion with emotional features is then input into a softmax classifier for classification, completing the static motion recognition. The softmax function is calculated using the following formula: (1) in, For the first i The weight matrix of the feature text. b Represents bias.
[0024] The formula for calculating convolutional layers is as follows: (2) in, l Indicates the first l One convolutional layer, i The convolution output matrix represents the first... i The value of each component; j This indicates the number of corresponding output matrices; j The value of varies between 0 and N, where N represents the number of convolution output matrices; f It is a non-linear... sigmoid Type function; The pooling layer uses average pooling. The input to the average pooling layer comes from the previous convolutional layer, and the output is used as the input to the next convolutional layer. The calculation formula is as follows: (3) in, This represents the local output after the pooling process is complete, derived from the mean of the local small matrix of size n×n in the previous layer.
[0025] S1.2 Utilize Kinect to detect human skeletal data in real time, calculate human behavioral characteristics, complete dynamic motion recognition, and generate dynamic motion feature text descriptions. First, human body localization and tracking are performed using Kinect to obtain the joints of the skeleton. The 15 skeleton joints are numbered from top to bottom and from left to right. Since the position signal of the skeleton is time-varying, its definition is unclear when encountering occlusion. Therefore, frame sequences are extracted from the video and input into interval Kalman filtering to improve the accuracy of skeleton position. Then, a bidirectional long short-term memory network (Bi-LSTM-CRF) with a conditional random field layer is used to analyze the motion sequence of the 15 skeleton points to obtain dynamic motion with emotional features. For a bidirectional long short-term memory neural network, given an input sequence Where t represents the t-th coordinate, and T represents a total of T coordinates, the output calculation formula of the hidden layer of the Long Short-Term Memory Neural Network is as follows: (4) in, This represents the output of the hidden layer at time t. This is the weight matrix from the input layer to the hidden layer. This is the weight matrix from hidden layer to hidden layer. For the bias of the hidden layer, This represents the activation function; although LSTM can capture information from long-term sequences, it only considers one direction. Bi-LSTM is used to strengthen this two-way relationship, with the first layer being a forward LSTM and the second layer a backward LSTM. Finally, the dynamic motion with emotional characteristics is input into the Softmax classifier in step S1.1 for classification.
[0026] S1.3. Use local recognition methods to identify facial expressions and generate text descriptions of facial activity features; Based on the information from the face image frames captured by Kinect, the various segmented regions of the face are obtained. The original images of each segmented part of the face are processed into normalized standard images, and two-dimensional Gabor wavelets are used for feature extraction. Linear Discriminant Analysis (LDA) algorithm is used for dimensionality reduction to extract the most discriminative low-dimensional features from the high-dimensional feature space. Based on the extracted low-dimensional features, all samples of the same category are collected to separate other samples as much as possible, i.e., the features with the largest ratio of dispersion between sample classes to dispersion within sample classes are selected. Finally, the face image frames after feature extraction by Gabor wavelets and dimensionality reduction by LDA are classified using an open-source OpenFace neural network to obtain the facial expression recognition results.
[0027] S1.4 Complete the recognition of speech markers and speech content, and generate a text description of language features; First, Kinect directly collects speech. The collected speech is then filtered to reduce noise using a Wiener-based noise filter. Next, the noise-reduced speech is input into a Backpropagation Neural Network (BPNN), which is a type of feedforward neural network. BPNN adds a backpropagation algorithm to the structure of a feedforward network to train the speech and obtain speech with prosodic and spectral features. Finally, the speech with prosodic and spectral features is input into a Softmax classifier for classification to obtain the speech recognition result.
[0028] S2. The feature text descriptions generated in step S1 are fused, and the processing results are classified for sentiment using self-organizing maps (SOM) and compensation layers; specifically, the following steps are included: S2.1. Use an LSTM neural network to embed the feature text descriptions collected in step S1 into a fixed-size feature vector arranged in chronological order; the LSTM neural network is the Bi-LSTM feedforward LSTM network in step 1.2. S2.2. The self-organization mapping (SOM) algorithm is used to normalize the feature vectors in step S2.1. S2.3. Since the Self-Organizing Map (SOM) layer is an ambiguous layer that loses information, a compensation layer is used to compensate for the information loss. That is, for different classification results in the SOM, the compensation layer must be combined with a specific layer. The size of each layer is the same as the competition layer of the SOM network, and all nodes have their own weights. ,s represents the s-th class of the corresponding compensation layer, and t represents The t-th node within the layer; compensation layers are not shared, each layer corresponds to the SOM output of a specific type of classification; the formula for the multiplication result is as follows: (5) In the formula, These are the input weights of the s-th layer node. Let be the weight of the node in layer s. This is used to limit the compensation ratio between -1 and 1; S2.4 Since similar characteristics may belong to the same class, a global optimization of the target result is required. The global optimization objective is: (6) The first item for: (7) The first item This is used to minimize the error between the label and the prediction result. The second term is... This is used to minimize the error between the labels and the SOM network results. (Third term) This is used to minimize the difference between the input signal and the output signal of the SOM network.
[0029] S3. Based on the output of step S2, determine the likelihood of depressive mood and suicidal tendency, and mark and observe high-risk individuals.
Claims
1. A method for identifying depression and suicidal tendencies by integrating body language, micro-expressions, and language, characterized in that... Includes the following steps: S1. Use a Kinect with an infrared camera to collect video and audio, and convert the video information and audio information into feature text descriptions respectively; the video information includes body movements and facial expression information extracted from the video, and body movements include static movements and dynamic movements; the audio information includes spectrum, prosody, and sound wave information extracted from the speech audio, the spectrum information and prosody information are used to obtain speech markers, and the sound wave information is used to obtain speech content; the extraction of feature text descriptions specifically includes the following steps: S1.1 A convolutional neural network (CNN) is used to identify static motion and generate textual descriptions of static motion features. Specifically, single frames are selected from the collected videos and input into the CNN for training and testing. All single frames from the videos are then input into the trained CNN to obtain static motion with emotional features. These emotionally-featured static motions are then input into a Softmax classifier for classification, completing the static motion identification. The softmax function is calculated using the following formula: (1) in, For the first i The weight matrix of the class feature text description, b Represents bias; S1.2 Utilize Kinect to detect human skeletal data in real time, calculate human behavioral characteristics, complete dynamic motion recognition, and generate dynamic motion feature text descriptions. S1.
3. Use local recognition methods to identify facial expressions and generate text descriptions of facial activity features; S1.4 Complete the recognition of speech markers and speech content, and generate a text description of language features; S2. The feature text descriptions generated in step S1 are fused, and the processing results are classified for sentiment using self-organizing maps (SOM) and compensation layers; specifically, the following steps are included: S2.
1. Use an LSTM neural network to embed the feature text descriptions collected in step S1 into a fixed-size feature vector arranged in chronological order; the LSTM neural network is the Bi-LSTM feedforward LSTM network in step 1.
2. S2.
2. The self-organizing map algorithm is used to normalize the feature vectors in step S2.
1. S2.
3. Since the self-organizing map layer is a fuzzy layer that loses information, a compensation layer is used to compensate for the information loss. That is, for different classification results in the SOM, the compensation layer must be combined with a specific layer. The size of each layer is the same as the competition layer of the SOM network, and all nodes have their own weights. ,s represents the s-th class of the corresponding compensation layer, and t represents The t-th node within the layer; compensation layers are not shared, each layer corresponds to the SOM output of a specific type of classification; the formula for the multiplication result is as follows: (5) In the formula, These are the input weights of the s-th layer node. Let be the weight of the node at layer s. This is used to limit the compensation ratio between -1 and 1; S2.4 Perform global optimization on the target result. The global optimization objective is: (6) The first item for: (7) The first item The first term is used to minimize the error between the label and the prediction result; the second term is... This is used to minimize the error between the label and the SOM network result; the third term This is used to minimize the difference between the input signal and the output signal of the SOM network; S3. Mark the individuals who may have depressive symptoms or suicidal tendencies obtained in step S2 and observe them.
2. The method for identifying depression and suicidal tendencies by integrating body language, micro-expressions, and language according to claim 1, characterized in that, In step S1.3, based on the information of the face image frames captured by Kinect, the various segmented regions of the face are obtained. The original images of each segmented part of the face are processed into normalized standard images, and two-dimensional Gabor wavelets are used for feature extraction. Dimensionality reduction is performed using linear discriminant analysis (LDA) algorithm to extract the most discriminative low-dimensional features from the high-dimensional feature space. All samples of the same category are collected based on the extracted low-dimensional features to separate other samples as much as possible, i.e., the feature with the largest ratio of dispersion between sample classes to dispersion within sample classes is selected. Finally, the face image frames after feature extraction using Gabor wavelets and dimensionality reduction using LDA are classified using an open-source OpenFace neural network to obtain the facial expression recognition results.
3. The method for identifying depression and suicidal tendencies by integrating body language, micro-expressions, and language according to claim 1, characterized in that, In step S1.4, firstly, the Kinect directly collects speech, and then uses a Wiener-based noise filter to reduce the noise in the speech. Then, the noise-reduced speech is input into the backpropagation neural network for training to obtain speech with prosodic and spectral features. Finally, the speech with prosodic and spectral features is input into the Softmax classifier for classification to obtain the speech recognition result.
4. The method for identifying depression and suicidal tendencies by integrating body language, micro-expressions, and language according to claim 1, characterized in that, In step S3, based on the output of step S2, the likelihood of depressive mood and suicidal tendency is determined, and high-risk individuals are marked and observed.
Citation Information
Patent Citations
Emotion recognition method, intelligent device and computer readable storage medium
CN111164601A
Bimodal emotion recognition method fusing multiple deep learning models
CN111292765A