Depression emotion grade judgment method based on Transform algorithm
By combining the Transformer algorithm and decision tree model, using the multi-head self-attention mechanism and proxy token clustering, the emotional characteristics of depression are extracted and judged, and the problems of large calculation overhead and opaque results in the existing technology are solved, and fast and accurate judgment of emotions level and model interpretability are achieved.
Patent Information
- Application Number
- CN202510092322.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-03
AI Technical Summary
In the judgment of emotional ratings of depression, the problem of high calculation overhead, lack of credibility in the results, and opaque decision paths, it is difficult to quickly and accurately judge emotional ratings.
Combining deep learning and machine learning, using the feature extraction ability of the Transformer algorithm and the interpretability of the decision tree, emotional features are extracted through the multi-head self-attention mechanism and proxy token clustering, and using the decision tree model to judge emotions level.
It improves the accuracy and robustness of emotional rank judgments, enhances the interpretability of the model, reduces calculation overhead, and achieves fast and transparent emotional rank judgments.
Smart Images

Figure CN120089333A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning and data mining, and particularly to a method for judging the depression mood level based on the Transformer algorithm. Background Art
[0002] Depression is a common psychological disorder with diverse symptoms, seriously affecting the daily life of patients. Traditional assessment methods usually rely on clinical interviews and standard questionnaires, which are time-consuming and highly subjective. Therefore, it is particularly important to develop a data-driven method for mood level assessment.
[0003] Currently, most emotion recognition systems are based on visual image emotion recognition or speech system emotion recognition. Then, a large emotion data model is established through Transformer, and then feature extraction is performed on the captured target information to achieve emotion recognition. In traditional methods, on the one hand, due to the black-box nature of Transformer, the results of emotion recognition lack sufficient credibility, and the decision-making path is also opaque. On the other hand, the computational cost of the Transformer model is relatively large, which may lead to slow system response, especially in applications that require rapid emotion recognition judgment. Therefore, how to quickly judge the mood level of depression patients and enhance the interpretability of the judgment model is an important challenge. Summary of the Invention
[0004] To solve the above technical problems, the present invention combines deep learning with machine learning, and utilizes the feature extraction ability of Transformer and the advantage of low time complexity of decision trees, which can quickly and effectively complete the extraction and determination of different mood levels, providing a new idea for the establishment of a mood judgment model. The depression mood level judgment system and method based on the Transformer algorithm enable the system to combine the feature extraction ability of deep learning with the interpretability of traditional decision trees, enhancing the discrimination ability and transparency of the model in depression mood judgment. At the same time, proxy token clustering is used to improve the discrimination ability of mood features.
[0005] The present invention provides a method for judging the depression mood level based on the Transformer algorithm, which specifically includes the following steps: S1: Preprocess the collected user-related data to obtain a data coding sequence; S2: Use the Transformer encoder to perform relevant emotion feature extraction on the data coding sequence based on the multi-head self-attention mechanism and proxy tokens, and output a feature vector; S3: Input the feature vector into a decision tree to construct an EJ decision tree model for mood judgment; S4: Input the user data to be judged into the constructed EJ decision tree model to judge the user's emotion level.
[0006] Optionally, step S1 is specifically as follows: Collect the user's online language comments, voices, and image actions as user-related data, clean the collected user-related data to remove noise and unnecessary information, then perform standardization and data augmentation, and encode the processed user-related data to obtain a data encoding sequence.
[0007] Optionally, step S2 specifically includes: Use the data encoding sequence of step S1 as the input sequence, convert each token in the input sequence into an embedding vector, the token obtains a context representation through an embedding layer, extract token features based on the embedding representation, use hierarchical clustering to group the tokens into multiple clusters, and each cluster contains multiple tokens; select one or more representative tokens as proxy tokens in each cluster according to the features of the tokens, and generate queries, keys, and values for each proxy token; In the multi-head self-attention mechanism, different linear transformations are used for each head to generate queries, keys, and values, and the attention weight output of each head is obtained and then merged. Combine the outputs of all heads to obtain the final representation and input it into the fully connected layer; The feature vector is subjected to feature extraction in the feed-forward neural network FNN; using self-supervised learning, first pre-train the emotion features of the data through the Transformer encoder, then verify whether the distribution of the emotion features meets the expectations through a dimensionality reduction method, add an emotion feature importance analysis module to verify the independence of each emotion feature and its correlation with the user's emotion level; the output of the FNN is multi-dimensional, and each dimension corresponds to a specific emotion feature, including emotional vocabulary, speech rate, pitch, loudness, sentence length, sentence complexity, facial expression, body posture, and gesture. Therefore, after being processed by the feed-forward neural network FNN, the final feature vector is obtained.
[0008] Optionally, the construction method in step S3 is specifically as follows: Use the C4.5 algorithm to construct the EJ decision tree.
[0009] Optionally, the user's emotion levels in step S4 include normal, mild depression, moderate depression, and severe depression.
[0010] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention proposes a method for judging the depression emotion level based on the Transformer algorithm, which can improve the feature representation ability. By introducing proxy tokens into the self-attention layer of the Transformer, the system can more effectively capture important features and context information in the input data. This improvement can significantly enhance the accuracy and robustness of user emotion level determination. This method enables the model to make more detailed distinctions for different emotion expressions when dealing with complex emotional states.
[0011] 2. The present invention proposes a method for judging the depression emotion level based on the Transformer algorithm, which enhances the discriminative ability of the model. By combining the Transformer encoder and the EJ decision tree model, the system can combine the feature extraction ability of deep learning with the interpretability of traditional decision trees. The decision tree can clearly reveal the relationship between features and the user's emotion level, enhancing the discriminative ability and transparency of the model in depression emotion judgment, and facilitating doctors and researchers to understand and trust the output of the model.
[0012] 3. The present invention proposes a method for judging the depression emotion level based on the Transformer algorithm, which optimizes the computational efficiency and enhances the interpretability of the model. The self-attention mechanism aggregates local information through proxy token clustering, avoiding over-reliance on complex convolutional or recurrent neural network structures, thereby reducing the computational amount and training time. At the same time, the decision tree model has good interpretability and can clearly show the basis for judging the user's emotion level, helping clinicians understand the logic of the model's judgment and providing support for further emotion intervention.
[0013] These beneficial effects jointly promote the development of the depression emotion judgment system in terms of accuracy, interpretability, and real-time performance, providing a more advanced and practical technical solution for clinical mental health assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic flow chart of a method for judging the depression emotion level based on the Transformer algorithm proposed by the present invention; Figure 2 It is a schematic network framework diagram of a method for judging the depression emotion level based on the Transformer algorithm proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0015] The technical solution of the present invention will be further described below in conjunction with the drawings and through specific embodiments.
[0016] As Figure 1As shown in the figure, a method for judging the depression emotion level based on the Transformer algorithm provided by an embodiment of the present invention specifically includes the following steps: S1: Preprocess the collected user-related data to obtain a data coding sequence; S2: Use a Transformer encoder to extract relevant emotion features from the data coding sequence based on the multi-head self-attention mechanism and proxy tokens, and output a feature vector; S3: Input the feature vector into a decision tree to construct an EJ decision tree model for emotion judgment; S4: Input the user data to be judged into the constructed EJ decision tree model to judge the user's emotion level.
[0017] As Figure 2 shown, step S1 is specifically: collect the user's online language comments, voices, and image actions as a data set, preprocess the collected data, clean the collected data to remove noise and unnecessary information, then standardize and augment the data, and encode the preprocessed data to obtain a data coding sequence.
[0018] The Transformer encoder in step S2 includes: an embedding layer and position encoding, a multi-head self-attention mechanism, residual connections and layer normalization, and a feed-forward neural network. Refer to Figure 2 the network framework schematic diagram to further explain the Transformer encoder.
[0019] Step S2 specifically includes S2.1 to S2.3.
[0020] S2.1 Tokenization and proxy tokens: Convert each token in the input sequence into an embedding vector. The token obtains a context representation through an embedding layer. Based on the embedding representation, extract the token features and group the tokens into multiple clusters using hierarchical clustering. Each cluster contains multiple tokens. Select one or more representative tokens as proxy tokens in each cluster according to the token features, and generate a query ( ), key ( ), and value ( ) for each proxy token. Use the relationship between the emotion features and the user's emotion level to judge the patient's emotion level, including normal, mild depression, moderate depression, and severe depression.
[0021] The input sequence and the proxy token sequence are as follows: , , Calculate the attention scores for the proxy token queries and keys, and apply normalization to obtain the attention weights: , , Among them, X represents the input sequence, n represents the number of tokens, B represents the sequence of proxy tokens, m represents the number of proxy tokens, , , are the corresponding linear transformation weight matrices, represents the query matrix, key matrix, and value matrix of the proxy tokens, represents the dimension of the key, represents performing a normalization operation on each row of the matrix, represents calculating the attention for each head.
[0022] S2.2 The multi-head attention layer calculation of the proxy tokens is similar to the standard multi-head attention, but it uses the proxy tokens B instead of the original input sequence. In the multi-head self-attention mechanism, different linear transformations are used for each head to generate the query, key, and value: For the i-th head: , The output of each head is: , The final multi-head output is combined as: , Combining the outputs of all heads to obtain the final representation and inputting it into the fully connected layer: , Among them, are the linear transformation parameters of the output, , , represent the different weight matrices used for each head, h is the number of heads, is used to concatenate or splice multiple elements, strings, arrays, or lists, etc., to generate a new result, represents concatenating the outputs of each head and generating the final output through a linear layer, represents the input to the fully connected layer.
[0023] S2.3 In step S2 in the feed-forward neural network (FNN) specifically undergoes the following processing for emotion feature extraction: The feed-forward neural network of the Transformer encoder consists of two fully connected layers (FC). The core calculation formula is as follows: , , Among them, , are weight matrices, , are bias terms, is the activation function represents the final output.
[0024] The feedforward neural network provides non - linear transformation through two fully - connected layers plus .
[0025] Using self - supervised learning, first pre - train the emotional features of the data through Transformer, and then verify whether the distribution of the emotional features meets the expectations through the dimensionality reduction method. Add a feature importance analysis module to verify the independence of each emotional feature and its correlation with the user's emotional level. Extract emotional features in the feedforward neural network. The output of the FNN is multi - dimensional, and each dimension corresponds to a specific emotional feature, such as emotional vocabulary, speech rate, pitch, loudness, sentence length, sentence complexity, facial expression, body posture, and gestures. Therefore, after being processed by the feedforward neural network, the final output is as follows: , where represents the output value of the i - th emotional feature.
[0026] Step S3 is specifically as follows: Input the feature vector output by the Transformer encoder into the EJ decision tree model. Since the output data is multi - dimensional, the C4.5 algorithm is used to construct the EJ decision tree to directly handle multi - classification problems. The information gain ratio is used to select the optimal feature split, and the EJ decision tree model is trained. The EJ decision tree judges the patient's emotional level (including normal, mild depression, moderate depression, severe depression) by learning the relationship between emotional features and the user's emotional level. Please refer to Figure 2 below for a detailed description of the EJ decision tree construction process.
[0027] S3.1 Combine the emotional level of the initial sample and the feature vector calculated by the Transformer encoder for the sample for feature selection.
[0028] S3.2 Calculate the information gain ratio for each feature and select the feature with the highest information gain ratio as the split feature of the current node.
[0029] For an emotional feature A, the formula for information gain is: , where D is the current data set, Values(A) are all possible values of the emotional feature A, is the subset when the value of the emotional feature A is v, is the entropy of the dataset D, defined as: , where is the probability of the i-th class.
[0030] To prevent the information gain from favoring features with more values, the gain ratio is used to adjust the impact of the information gain by the "uniformity" of the split: , where is the entropy of the emotion feature A: , S3.3 Split the dataset into multiple subsets according to the selected emotion features.
[0031] S3.4 For each subset, repeat steps 3.2 and 3.3 until the stopping condition is met (no more features are available for selection).
[0032] S3.5 When the recursion stops, the generated leaf nodes will correspond to the final user emotion levels, including normal, mild depression, moderate depression, and severe depression.
[0033] For step S4, based on the trained EJ decision tree model, evaluate the user data to be judged and output the user emotion level.
Claims
1. A method for judging depression emotion level based on Transformer algorithm, characterized in that: The specific steps include: S1: pre-process the collected user-related data to obtain a data coding sequence; S2: Use the Transformer encoder to extract relevant emotional features from the data encoding sequence based on the multi-head self-attention mechanism and proxy tokens, and output the feature vector; S3: input the feature vector into a decision tree to construct an EJ decision tree model for emotion judgment; S4: Input the user data to be judged into the constructed EJ decision tree model to judge the user's emotion level.
2. The method according to claim 1, characterized in that The step S1 specifically includes: collecting user online language comments, voice and image actions as user-related data, cleaning the collected user-related data to remove noise and unnecessary information, and then standardizing and data enhancing the data, encoding the processed user-related data to obtain a data encoding sequence.
3. The method according to claim 1, characterized in that The step S2 specifically includes: The data encoding sequence of step S1 is used as an input sequence, each token in the input sequence is converted into an embedding vector, the token obtains context representation through an embedding layer, and the token features are extracted based on the embedding representation. The tokens are grouped into multiple clusters using hierarchical clustering, and each cluster contains multiple tokens; one or more representative tokens are selected in each cluster as proxy tags according to the features of the tokens, and a query, key and value are generated for each proxy tag; In the multi-head self-attention mechanism, different linear transformations are used for each head to generate queries, keys, and values. The attention weight output of each head is merged and then the outputs of all heads are combined to obtain the final representation and input into the fully connected layer. The feature vector is extracted in the feedforward neural network FNN; self-supervised learning is used to pre-train the data for emotional features through the Transformer encoder, and then the dimension reduction method is used to verify whether the distribution of emotional features meets expectations, and an emotional feature importance analysis module is added to verify the independence of each emotional feature and its correlation with the user's emotional level; the output of FNN is multi-dimensional, and each dimension corresponds to a specific emotional feature, including emotional vocabulary, speaking speed, pitch, loudness, sentence length, sentence complexity, facial expressions, body postures and gestures, so after being processed by the feedforward neural network FNN, the final feature vector is obtained.
4. The method according to claim 3, characterized in that The construction method in step S3 is specifically: using the C4.5 algorithm to construct the EJ decision tree.
5. The method according to claim 1, characterized in that: The user emotion levels in step S4 include normal, mild depression, moderate depression, and severe depression.
Citation Information
Cited By
Method for accurately predicting clinical prognosis of colorectal cancer patient
CN121329977A
A method for accurately predicting the clinical prognosis of colorectal cancer patients
CN121329977B