A method for constructing a domain dictionary for depressive behavior characteristics
By constructing emotional, behavioral characteristics and emoji word collections, using TF-IDF, PMI and WoBERT models, refine behavioral characteristics and generate Chinese dictionary for depression, the problem of correlation between behavioral characteristics and condition in depression text is solved, and detection accuracy and diagnostic support are improved.
Patent Information
- Application Number
- CN202311015057.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-12
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-08-12
AI Technical Summary
The existing dictionary of depression field cannot effectively solve the correlation between behavioral characteristics and patient condition in Chinese depression texts, and rules and statistics-based methods cannot fully cover complex semantic phenomena and ambiguity.
By constructing emotional word collections, behavioral feature word collections and emoji word collections, using TF-IDF, PMI, WoBERT models and label propagation algorithms, we will refine behavioral characteristics, generate Chinese depression field dictionary, integrate emotional features and behavioral characteristics, build a semantic graph based on the similarity between words, and set tags to obtain accurate behavioral feature word collections.
It improves the accuracy of depression tendency detection and enhances the diagnostic support and therapeutic diagnosis of patients with depression.
Smart Images

Figure CN117033660B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and specifically to a method for constructing a domain dictionary for depression behavior characteristics. Background Art
[0002] Depression is a common mental disorder, and early detection of depression can effectively prevent the occurrence of severe depression. More and more depression patients express their inner emotions and feelings through social media platforms, discuss their conditions, and some users even seek psychological diagnosis through this. Therefore, studying depression texts helps to deeply explore the symptoms and courses of depression. Depression behavior characteristics are a series of characteristics shown by depression patients at the behavioral, emotional, and psychological levels, which have a certain correlation with the patient's condition. Studying depression behavior characteristics and applying them to depression tendency detection helps to provide support for users with severe depression tendencies and also helps medical staff further understand the patient's condition and treatment diagnosis.
[0003] Currently, the research on depression domain dictionaries mainly includes rule-based methods and statistic-based methods. The rule-based method uses rules to find matching words in the text and cannot cover complex semantic phenomena; the statistic-based method uses machine learning and statistical word frequency techniques to mine depression-related vocabulary from a large amount of text data but cannot solve the ambiguity problem caused by the context of domain vocabulary; at the same time, most of the research on Chinese depression domain dictionaries ignores the correlation between behavior characteristics and the patient's condition in depression texts.
[0004] Aiming at the above deficiencies, the present invention proposes a method for constructing a domain dictionary for depression behavior characteristics. The difference of the present invention lies in that, aiming at the sparse information of depression topic posts, exploring the corresponding relationship between the patient's condition and behavior characteristics to refine the behavior characteristics, considering the context information to deeply mine domain vocabulary, and integrating the emotional characteristics and behavior characteristics of depression users into the dictionary to improve the accuracy of depression tendency detection. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for constructing a domain dictionary for depression behavior characteristics, calculating similarity based on an existing dictionary to obtain an emotional word set, then refining the behavior characteristics by using the correlation between the patient's condition and behavior characteristics, generating word vectors corresponding to posts and behavior type seed words through a pre-trained model, calculating similarity to obtain a behavior type candidate word set, then constructing a semantic graph based on word similarity, using the label propagation algorithm to automatically set labels for candidate words to obtain a behavior characteristic word set, and finally collecting negative emotion emojis to construct an emoji word set, and combining the obtained word sets to obtain a Chinese depression domain dictionary.
[0006] The related definitions involved in the present invention are as follows:
[0007] Definition 1: Behavioral characteristics of cognitive awareness behavior type: Cognitive awareness behavior refers to the behaviors such as unwillingness to chat, hiding, self - blame, and the manifestation of cognitive dysfunction of patients due to mild depressive tendencies, resulting in social avoidance, escape psychology, and mental breakdown.
[0008] Definition 2: Behavioral characteristics of somatic representation type: Somatic representation refers to the somatic symptoms and somatic anxiety easily caused by moderate depressive tendencies in patients.
[0009] Definition 3: Behavioral characteristics of harm awareness behavior type: Harm awareness behavior refers to the behaviors of self - harm or harming others by patients with severe depressive tendencies, who have the idea of punishing themselves to vent their emotions and have severe self - abasement and self - guilt.
[0010] The present invention adopts the following technical solutions to achieve the invention purpose:
[0011] A method for constructing a domain dictionary for depressive behavior characteristics, comprising the following steps:
[0012] (1) Pre - processing of depressive topic posts to construct an emotional seed word set and a behavioral seed word set;
[0013] Obtain some Weibo depressive topic posts, perform pre - processing operations on them, use the TF - IDF algorithm to extract high - frequency behavioral and emotional words from depressive texts, and obtain an emotional seed word set and a behavioral seed word set after screening.
[0014] (2) Obtain an emotional word set by calculating similarity;
[0015] Calculate the similarity between negative emotional words in the general Chinese emotion dictionary and emotional seed words through the point - mutual information algorithm to obtain an emotional word set.
[0016] (3) Refine behavioral characteristics based on the correspondence between behavioral characteristics and disease conditions to obtain a behavioral candidate word set;
[0017] Construct the correspondence between the disease conditions of depressive patients and behavioral characteristics, define labels to refine behavioral characteristics, use the WoBERT model to obtain the text word vectors corresponding to the behavioral seed words and calculate the cosine similarity between the word vectors to obtain a behavioral candidate word set.
[0018] (4) Construct a semantic graph and set labels for candidate words to obtain a behavioral characteristic word set;
[0019] Construct a semantic graph based on the word - to - word similarity between seed words and candidate words, and set labels for candidate words through the label propagation algorithm to obtain a behavioral characteristic word set.
[0020] (5) Construct an emoji word set, and merge the word sets to obtain a Chinese dictionary for the field of depression;
[0021] Collect the commonly used negative emotion emojis of Weibo users to obtain an emoji word set, and merge the emotion word set, behavior feature word set and emoji word set to obtain a Chinese dictionary for the field of depression.
[0022] Among them, in the step (1), the specific operations for preprocessing depression topic posts and constructing an emotion seed word set and a behavior seed word set are as follows:
[0023] (1.1) Randomly extract some posts from the "Depression Super Topic" on Sina Weibo and perform unified annotation.
[0024] (1.2) Delete irrelevant information such as stop words and links, replace special symbols with the '^' symbol, delete meaningless or depression-irrelevant texts, and retain the Chinese meaning annotations corresponding to behavior words, emotion words and emojis.
[0025] (1.3) Use the Jieba word segmentation tool to perform word segmentation and part-of-speech tagging on the depression topic text corpus.
[0026] (1.4) Use the TF-IDF algorithm to extract high-frequency words from the text, and obtain the behavior seed word set and the emotion seed word set respectively through screening.
[0027] Among them, in the step (2), the specific steps for obtaining the emotion word set by calculating similarity are as follows:
[0028] (2.1) Extract negative emotion category words from the Dalian University of Technology Emotion Lexical Ontology as specific words.
[0029] (2.2) Calculate the similarity between the specific words in the word library and the emotion seed words through the point mutual information algorithm, add the words with similarity greater than the threshold 1 to the emotion word set, and combine the emotion seed word set to finally obtain the emotion word set.
[0030] The calculation process of the point mutual information algorithm is as follows:
[0031]
[0032] Parameter description: P(w E1 ) represents the probability that the word in the emotion seed word set appears alone, and P(w2) refers to the probability that the specific word introduced in this article appears alone in the word library. P(w E1 , w2) is the probability that w E1 and w2 appear in the corpus at the same time.
[0033] In order to select words that have semantic relevance to sentiment-related seed words, in this patent, the threshold for calculating point mutual information similarity is set to 1, and specific vocabulary in the thesaurus with a threshold exceeding 1 is selected and added to the sentiment word set.
[0034] Among them, in the step (3), the specific steps for refining the behavior features based on the correspondence between behavior features and medical conditions to obtain the behavior candidate word set are as follows:
[0035] (3.1) Construct a binary correspondence between behavior features and the patient's medical conditions, refine and define labels for the behavior features according to the severity of the medical conditions, assign different weights to the labels, and apply them to the behavior-related seed words.
[0036] (3.2) Use the WoBERT model to generate word vectors corresponding to the behavior-related seed words and the depression post texts respectively.
[0037] (3.3) Set the threshold of cosine similarity, select the text vocabulary above the threshold, and filter out the text vocabulary below the threshold.
[0038] To ensure the similarity between words and select more accurate domain vocabulary, in this patent, the optimal threshold of 0.5 is selected to select similar words as candidate words, that is, when the cosine similarity between the behavior-related seed word vector and the text word vector is greater than the threshold of 0.5, the corresponding text vocabulary is added to the behavior candidate word set.
[0039] (3.4) Calculate the similarity through cosine similarity, and select the text vocabulary with a similarity greater than the threshold to the behavior-related seed words to form the behavior candidate word set.
[0040] The calculation formula of cosine similarity is as follows:
[0041]
[0042] Parameter description: S represents the behavior-related seed word vector; T represents the text dynamic word vector; n represents the number of dimensions; S i and T i represent the value of the word vector in the i-th dimension.
[0043] Among them, in the step (4), the specific steps for constructing a semantic graph and setting labels for candidate words to obtain the behavior feature word set are as follows:
[0044] (4.1) Establish a semantic graph based on the word similarity between the behavior-related seed words and the behavior candidate words. This semantic graph consists of nodes and edges. Among them, the nodes include the behavior-related seed words with known labels and the behavior candidate words with unknown labels, and the edges between the nodes represent the similarity relationship between two words.
[0045] (4.2) Set labels for candidate behavior words through the label propagation algorithm. The label propagation algorithm calculates the labels of unknown label nodes based on the relationships between nodes and updates them to labels similar to adjacent nodes. Eventually, each behavior feature obtains a unique label, and a set of behavior feature words is obtained.
[0046] Let the seed words of the behavior class be i and the candidate words of the behavior class be j. Then, the calculation formula for the label weight of the candidate word is as follows:
[0047]
[0048] Parameter description: n is the number of nodes in the semantic graph, W[j] represents the possible label weight of candidate word j, T[i][j] represents the transition probability matrix from seed word i to candidate word j, and V[i] represents the initial label of node i before iteration.
[0049] Among them, in the step (5), the specific steps for constructing an emoji word set and merging the word sets to obtain a Chinese depression domain dictionary are as follows:
[0050] (5.1) Collect the Chinese meaning annotations corresponding to the negative emotion emojis on Weibo to obtain an emoji word set.
[0051] (5.2) Merge the emotion word set obtained in step (2.2), the behavior feature word set obtained in step (4.2), and the emoji word set obtained in step (5.1) to finally obtain a Chinese depression domain dictionary. Description of the Drawings
[0052] Figure 1 It is a flowchart of a method for constructing a domain dictionary for depression behavior characteristics;
[0053] Figure 2 It is a schematic diagram of the construction process of an emotion word set;
[0054] Figure 3 It is a schematic diagram of the construction process of a candidate behavior word set;
[0055] Figure 4 It is an example diagram of constructing a behavior feature word set based on a behavior class seed word set and a behavior class candidate word set; Detailed Implementation Modes
[0056] The following further explains the present invention through specific embodiments.
[0057] Embodiment 1: The present invention provides a method for constructing a domain dictionary for depression behavior characteristics, as Figure 1 shown. The specific steps are as follows:
[0058] S1. Preprocess the depression topic posts, and construct an emotion class seed word set and a behavior class seed word set;
[0059] S1.1. First, obtain the depressive disorder topic post data through web crawlers. Randomly select some user posts in the "Depressive Disorder Super Topic" on Sina Weibo and perform unified annotation.
[0060] S1.2. Delete meaningless or depressive disorder - unrelated texts and phrases, delete stop words, pictures, links, and videos, replace special symbols with the '^' symbol, and retain the action words, emotion words, and the Chinese meaning annotations corresponding to emoticons.
[0061] S1.3. Use the Jieba word - segmentation tool to perform word - segmentation and part - of - speech tagging operations on the pre - processed posts. Use the paddle mode in the Jieba word - segmentation tool and utilize the PaddlePaddle deep - learning framework to train a network model to achieve word - segmentation and part - of - speech tagging.
[0062] S1.4. Extract the high - frequency words of the posts through the TF - IDF algorithm. After screening, obtain the action - type seed word set and the emotion - type seed word set respectively. The calculation formula of the TF - IDF algorithm is as follows:
[0063] TF - IDF = TF×IDF
[0064]
[0065] Among them, TF refers to the probability that a given word appears in a document, and IDF refers to the inverse document frequency, which is a measure of the general importance of a word; assume that there are k words in document j; n i,j represents the frequency of word i in document j, |j| represents the total number of words in document j, n k,j is the word t k appearing in file N i The number of times it appears; N represents the total number of documents, that is, the total number of depressive disorder topic texts, and N i represents the total number of documents containing word i.
[0066] S2. Based on the calculation of word - to - word similarity, obtain the emotion - type word set. The following explanations are combined with Figure 2 as follows:
[0067] S2.1. Select the emotion words of the five negative emotion categories of "sorrow, disgust, surprise, fear, anger" in the emotion word ontology library of Dalian University of Technology as specific words.
[0068] S2.2. Calculate the similarity between specific words and sentiment-class seed words through the point mutual information algorithm. To select words with semantic relevance, set the threshold to 1, select words with a similarity greater than the threshold for word expansion, and set the sentiment word weight to 1. The calculation formula of the point mutual information algorithm is as follows:
[0069]
[0070] where P(w E1 ) represents the probability of a word in the sentiment-class seed word set appearing alone, and P(w2) refers to the probability of a specific word in the introduced word library in this article appearing alone; P(w E1 , w2) is the probability that w E1 and w2 appear in the corpus at the same time; if the two are independent, the value of PMI is 0, and if the two have semantic relevance, the value of PMI is greater than 1.
[0071] S3. Refine the behavior features based on the correspondence between behavior features and medical conditions to obtain a set of candidate behavior words. Combined with Figure 3 the following is an explanation:
[0072] S3.1. Construct the correspondence between <behavior features, medical conditions>, and the correspondences are <cognitive awareness behavior type behavior features, mild depressive tendency>; <somatization representation type behavior features, moderate depressive tendency>; <harm awareness behavior type behavior features, severe depressive tendency>. Based on this, define the behavior features and their weights in the dictionary as: cognitive awareness behavior type, weight is 1; somatization representation type, weight is 3; harm awareness behavior type, weight is 5.
[0073] S3.2. Input the preprocessed depression posts and behavior-class seed words into the WoBERT model respectively to obtain the text word vectors and the corresponding word vectors of the seed words. The WoBERT model modifies the tokenizer and adds a pre-tokenizer operation to tokenize Chinese words.
[0074] S3.3. Set the cosine similarity threshold to 0.5. When the cosine similarity value between the word vector corresponding to the behavior-class seed word and the text word vector is greater than the threshold 0.5, add the corresponding text word to the set of candidate behavior words.
[0075] S3.4. Calculate the similarity through cosine similarity, and select the text words with a similarity greater than the threshold to the behavior-class seed words to form a set of candidate behavior words. The cosine similarity calculation formula is as follows:
[0076]
[0077] where S represents the behavior-class seed word vector; T represents the text dynamic word vector; n represents the number of dimensions; Si With T i represents the value of the word vector in the i-th dimension.
[0078] S4. Construct a semantic graph, set labels for candidate words to obtain a set of behavior feature words. Combine Figure 4 The following is an explanation:
[0079] S4.1. According to the above calculation of the similarity between word vectors, construct a semantic graph based on the similarity between words. Each node represents a behavior feature word, including the seed words of behavior classes with known labels and the candidate words of behavior classes with unknown labels. The similarity relationship between words forms the edges. The semantic graph structure is represented as follows:
[0080] G = (X, E), where X represents the set of nodes in the graph, the number of seed words is n, and E is the adjacency matrix of graph G.
[0081] S4.2. Construct a transition probability matrix from seed words to candidate words. Through the transition probability matrix, guide the process of label propagation, set labels for candidate words of behavior classes, and finally each behavior feature obtains a unique label to obtain a set of behavior feature words. The calculation formula of the transition probability matrix is as follows:
[0082]
[0083] Among them, n is the number of nodes in the semantic graph, T[i][j] represents the transition probability matrix from seed word i to candidate word j, and SIM(w i , w j ) represents the similarity between w i and w j calculated by cosine similarity.
[0084] The calculation formula for the label weight of candidate word j is as follows:
[0085]
[0086] Among them, n is the number of nodes in the semantic graph, T[i][j] represents the transition probability matrix from seed word i to candidate word j, and V[i] represents the initial label of node i before iteration.
[0087] S5. Construct a set of emoji words, merge the sets of words to obtain a Chinese depression domain dictionary, and combine Figure 1 The following is an explanation:
[0088] S5.1. Collect the Chinese meaning annotations corresponding to the commonly used negative emotion emoji of Weibo users on Sina Weibo and set their weights to 1 to obtain a set of emoji words.
[0089] S5.2. Merge the sets of emotion words, behavior feature words and emoji words to obtain a Chinese depression domain dictionary.
[0090] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0091] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for constructing a domain dictionary for depressive behavior characteristics, characterized in that It includes the following steps: Step 1: Preprocess the depression topic posts, and construct the emotional seed word set and the behavioral seed word set; preprocess the crawled depression topic posts, perform word segmentation and part-of-speech tagging operations using the Jieba word segmentation tool, calculate the high-frequency words through the TF-IDF algorithm, and obtain the behavioral seed word set and the emotional seed word set respectively after screening; Step 2: Obtain the emotional word set based on similarity calculation; calculate the similarity between the negative emotional words in the general Chinese emotion dictionary and the emotional seed words through the pointwise mutual information algorithm, and add the words with similarity greater than the threshold to the emotional word set; Step 3: Refine the behavioral features based on the correspondence between behavioral features and conditions, and obtain the behavioral candidate word set; generate the word vectors corresponding to the depression posts and the behavioral seed words through the WoBERT model, and calculate the cosine similarity between the two to obtain the behavioral candidate word set; Step 4: Construct a semantic graph, and set labels for the candidate words to obtain the behavioral feature word set; construct a semantic graph based on the word similarity between the behavioral seed words and the behavioral candidate words, and then automatically set labels for the behavioral candidate words through the label propagation algorithm; Step 5: Construct the emoji word set, and merge the word sets to obtain the Chinese depression domain dictionary; Collect the commonly used negative emotional emojis on Weibo to obtain the emoji word set, merge the emotional word set, the behavioral feature word set, and the emoji word set, and finally obtain the Chinese depression domain dictionary.
2. The method for constructing a domain dictionary for depressive behavior characteristics according to claim 1, characterized in that Step 1 includes: Step 1.1 Preprocess the depression topic posts: After obtaining some depression topic posts, delete the meaningless and depression-unrelated posts, delete the stop words, pictures, links, and videos, replace the special symbols with '^', and retain the Chinese meaning annotations corresponding to the behavioral words, emotional words, and emojis. Perform word segmentation and part-of-speech tagging using the Jieba word segmentation tool, aiming to ensure the integrity of the domain vocabulary to the greatest extent; Step 1.2 Construct the emotional seed word set and the behavioral seed word set: Use the TF-IDF algorithm to extract the high-frequency words from the preprocessed posts, and obtain the emotional seed word set and the behavioral seed word set respectively after screening. The TF-IDF calculation formula is as follows: TF-IDF = TF × IDF Among them, TF refers to the probability that a given word appears in a document, and IDF refers to the inverse document frequency, which is a measure of the general importance of a word. Assume that there are k words in document j, n i,j represents the frequency of word i in document j, |j| represents the total number of words in document j, n k,j is the word t k appears in file N i The number of occurrences; N represents the total number of documents, that is, the total number of texts on the topic of depression, N i represents the total number of documents containing word i in the document.
3. The method for constructing a domain dictionary for depressive behavior characteristics according to claim 1, characterized in that Step 2 includes: Step 2.1 Extract specific words from the general Chinese emotion dictionary: Select the emotion words of the five negative emotion categories of "sorrow, disgust, surprise, fear, anger" in the emotion word ontology library of Dalian University of Technology; Step 2.2 Construct the emotional word set: Calculate the similarity between these words and the emotional seed word set through the pointwise mutual information algorithm, and select the words with semantic relevance to the emotional seeds, that is, the words with PMI value greater than 1 for expansion to realize the construction of the emotional word set. The calculation formula of the mutual point information algorithm is as follows: Among them, P(w E1 ) represents the probability of a word in the sentiment seed word set appearing alone; P(w2) refers to the probability of a specific word in the thesaurus introduced in this paper appearing alone; P(w E1 , w2) is the probability that w E1 and w2 appear in the corpus simultaneously.
4. The method for constructing a domain dictionary for depressive behavior characteristics according to claim 1, characterized in that Step 3 includes: Step 3.1 Set behavior feature labels: Construct a binary tuple <behavior feature, condition> to represent the correlation between behavior features and the patient's condition. Based on this, define labels to refine the behavior features and apply them to the behavior type seed word set. The labels include cognitive awareness behavior type, somatic representation type, and harm awareness behavior type. According to the severity of the condition, set the weights of the three types of behavior features in the dictionary to 1, 3, and 5 respectively; Step 3.2 Construct the behavior type candidate word set: Obtain the word vectors corresponding to the depression posts and behavior type seed words through the WoBERT model for the preprocessed posts and behavior type seed words respectively; Step 3.3 Set the threshold of cosine similarity: To ensure the similarity degree between words to select more accurate domain vocabulary, set the threshold to 0.
5. When the cosine similarity between the behavior type seed word vector and the text word vector is greater than the threshold 0.5, add the corresponding word in the text to the behavior type candidate word set; Step 3.4 Calculate the similarity between word vectors through cosine similarity: Select the post corresponding words with similarity greater than the threshold to form the behavior type candidate word set. The calculation formula of cosine similarity is as follows: Among them, S represents the behavior - type seed word vector; T represents the text dynamic word vector; n represents the number of dimensions; S i and T i represents the value of the word vector in the i - th dimension.
5. The method for constructing a domain dictionary for depressive behavior characteristics according to claim 1, characterized in that Step 4 includes: Step 4.1 Construct a semantic graph based on word similarity: Construct a semantic graph according to the similarity between the behavior type seed words and the word vectors corresponding to the behavior type candidate words in Step 3. The graph structure is represented as follows: G=(X, E), where X represents the node set of the graph structure, the nodes include the behavior type seed words with known labels and the behavior type candidate words with unknown labels, and E is the adjacency matrix of graph G. Each edge represents the semantic similarity relationship between nodes; Step 4.2 Set labels through the label propagation algorithm: Construct a transition probability matrix, and through the iteration of the label propagation algorithm, update the nodes with unknown labels to the labels of the words with the highest similarity among their adjacent nodes. Let the seed word be i and the candidate word be j, then the transition probability matrix T[i][j] from the seed word i to the candidate word j is as follows: where n is the number of semantic graph nodes, SIM(w i , w j ) represents the similarity between w i and w j calculated by cosine similarity; the label weight calculation formula for candidate word j is as follows: where n is the number of semantic graph nodes, W[j] represents the possible label weight of the candidate word j, T[i][j] represents the transition probability matrix from the seed word i to the candidate word j, and V[i] represents the initial label of node i before iteration.
6. The domain dictionary construction method for depressive behavior characteristics according to claim 1, wherein Step 5 includes: Step 5.1 Construct an emoji word set: Collect the Chinese meaning annotations corresponding to the commonly used negative emotion emojis on Weibo to form an emoji word set; Step 5.2 Obtain the Chinese depression domain dictionary: Merge the emotion word set obtained in Step 2, the behavior feature word set obtained in Step 4, and the emoji word set obtained in Step 5 to complete the construction of the Chinese depression domain dictionary.