Social media user depression detection method and system based on symptom sequence
By combining symptoms and emotional characteristics in depression detection, using multiple deep learning models for feature extraction and fusion, the problem that the existing technology cannot fully capture user emotional changes is solved, and higher accuracy and comprehensiveness of depression detection are achieved.
Patent Information
- Application Number
- CN202510627068.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing depression detection technologies are unable to fully capture subtle mood changes in user language and complex relationships between multiple variables, resulting in insufficient detection accuracy and comprehensiveness.
A social media user depression detection method based on symptom sequence is proposed. By combining symptom characteristics and emotional characteristics, using sentence embedding models, pre-trained sentiment analysis models and bidirectional gated recurrent unit networks, data cleaning, feature extraction, feature fusion and multivariate time series classification are carried out to improve the accuracy of detection.
By combining symptoms and emotional characteristics, the user's depression state can be more comprehensively portrayed, the prediction accuracy and comprehensiveness of the model can be improved, the dynamic changes and contextual relationships of user emotions can be captured, and the reliability of depression detection can be enhanced.
Smart Images

Figure CN120148874A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and particularly to a method and system for detecting depression of social media users based on symptom sequences. Background Art
[0002] The popularity of social media provides researchers with a dynamic observation window, enabling people to share their life and emotional states anytime and anywhere. This real-time nature offers unprecedented opportunities for the early detection and intervention of depression. However, traditional clinical diagnosis methods have many limitations. On the one hand, questionnaire surveys and clinical interviews rely on patients' self-descriptions of subjective feelings, which are easily affected by patients' subjective will and social expectations, making the diagnosis results inaccurate. On the other hand, the limited professional psychiatric medical resources cannot meet the needs of large-scale screening, resulting in many depression patients not being diagnosed and treated in a timely manner. In addition, some existing automatic depression detection technologies, such as methods based on behavioral feature analysis, often can only capture some depressive symptoms, and the comprehensiveness of data and the accuracy of detection still need to be improved.
[0003] Depression patients often show persistent low mood, and this emotional state is manifested in language as the frequent use of negative emotion words, including direct negative words and implicit emotional expressions. Traditional methods often rely on patients' self-reports in interviews or questionnaires, but these methods cannot capture the subtle emotional changes in language and cannot quantify the dynamic development of emotional states.
[0004] A multivariate time series refers to a sequence of observed values that change over time on multiple variables. Different from univariate time series, multivariate time series contain the dynamic change information of multiple related variables. For example, in depression prediction, users' emotional characteristics, symptom characteristics, etc. can all be regarded as different variables, which change over time and jointly constitute a multivariate time series. Traditional methods usually can only process univariate data and cannot capture the complex relationships between multiple variables. Summary of the Invention
[0005] In view of the above situation, the main purpose of the present invention is to propose a method and system for detecting depression of social media users based on symptom sequences to solve the above technical problems.
[0006] The present invention proposes a method for detecting depression of social media users based on symptom sequences, and the method includes the following steps: Step 1: Input the original published articles of users and given tweet sequences, clean the tweet sequences to obtain the cleaned tweet sequences; Step 2: Input the cleaned tweet sequence and the text sequence of depressive symptom descriptions in the DSM-5 manual into the sentence embedding model for encoding to obtain the similarity score sequence between the posts and the symptom descriptions; Step 3: Calculate the dynamic symptom indicators for the similarity score sequence between the posts and the symptom descriptions to obtain the set of dynamic symptom indicators; Step 4: Input the cleaned tweet sequence into the pre-trained model for sentiment analysis to generate an emotion score sequence, and input the emotion score sequence into a bidirectional gated recurrent unit network (Bi-GRU) for emotion capture processing to obtain the context-related emotion sequence; Step 5: Perform feature fusion on the similarity score sequence between the posts and the symptom descriptions, the set of dynamic symptom indicators, and the context-related emotion sequence to obtain the overall feature representation of the user tweet sequence; Step 6: Input the overall feature representation of the user tweet sequence into a multivariate time series classifier for classification to obtain the classification result.
[0007] The present invention also proposes a depressive detection system for social media users based on symptom sequences, and the system includes: A data cleaning module for: Inputting the original published articles of users and given tweet sequences, cleaning the tweet sequences to obtain the cleaned tweet sequences; A symptom feature extraction module for: Inputting the cleaned tweet sequence and the text sequence of depressive symptom descriptions in the DSM-5 manual into the sentence embedding model for encoding to obtain the similarity score sequence between the posts and the symptom descriptions; A symptom dynamic indicator extraction module for: Calculating the dynamic symptom indicators for the similarity score sequence between the posts and the symptom descriptions to obtain the set of dynamic symptom indicators; An emotion feature extraction module for: Inputting the cleaned tweet sequence into the pre-trained model for sentiment analysis to generate an emotion score sequence, and inputting the emotion score sequence into a bidirectional gated recurrent unit network for emotion capture processing to obtain the context-related emotion sequence; A feature fusion module for: Performing feature fusion on the similarity score sequence between the posts and the symptom descriptions, the set of dynamic symptom indicators, and the context-related emotion sequence to obtain the overall feature representation of the user tweet sequence; A prediction module for: Inputting the overall feature representation of the user tweet sequence into a multivariate time series classifier for classification to obtain the classification result.
[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By combining symptom features and emotional features, the present invention can more comprehensively depict the depressive state of users and improve the prediction accuracy of the model; patients with depression will show specific features both in terms of emotions and symptoms. Relying solely on features in one aspect may miss some important information. The present invention combines the two, comprehensively considering the emotional changes expressed by users on social media and the semantic information related to depressive symptoms, so as to more comprehensively evaluate the degree of the depressive state of users; using the SentenceBert model to encode the cleaned tweet sequence and the text sequence of depressive symptom descriptions in the DSM-5 manual, and calculating the similarity sequence of the high-dimensional semantic vectors of the two as the symptom score sequence, can more accurately capture the semantic information related to depressive symptoms in the user's tweets and improve the accuracy and effectiveness of symptom feature extraction; the SentenceBert model has strong semantic understanding ability and can map the text into a high-dimensional semantic space, so as to better measure the similarity between texts; by comparing with the depressive symptom descriptions in the DSM-5, the content in the user's tweets that may reflect depressive symptoms can be accurately found; 2. The present invention obtains an emotional score sequence through the SentimentBert model and inputs this sequence into a Bi-GRU to obtain a context-related emotional score sequence, which can better capture the dynamic changes and context relationships of users' emotions and improve the accuracy and comprehensiveness of emotional feature extraction; the SentimentBert model can perform sentiment analysis on tweets and give emotional scores, while the Bi-GRU can consider context information, making the emotional features more accurately reflect the emotional state of users at different time points and their mutual influences; fusing the symptom sequence, dynamic indicators, and context-related emotional score sequence to obtain the overall representation sequence of the user's tweets can comprehensively consider the influence of various features on depression detection and improve the prediction performance of the model; this fusion strategy organically combines different types of features, making the overall representation sequence contain rich information and being able to more comprehensively describe the depressive state of users, thus providing more powerful support for subsequent classification prediction; 3. The present invention uses a multivariate time series classifier to classify and predict the overall representation sequence, which can make full use of the dynamic change information in the multivariate time series data and improve the accuracy and reliability of the depression classification result; the multivariate time series classifier can simultaneously process the time series data of multiple variables, capture the correlation and dynamic change trend between variables, and thus more accurately classify the depressive state of users.
[0009] The additional aspects and advantages of the present invention will be partially given in the following description, partially will become obvious from the following description, or can be understood through the embodiments of the present invention. Brief Description of the Drawings
[0010] Figure 1 This is the flowchart of the steps of the method for detecting depression of social media users based on symptom sequences proposed by the present invention.
[0011] Figure 2 This is the architecture diagram of the method for detecting depression of social media users based on symptom sequences proposed by the present invention.
[0012] Figure 3 This is the schematic structural diagram of the system for detecting depression of social media users based on symptom sequences proposed by the present invention. Detailed implementation manners
[0013] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals are the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0014] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed as some ways to implement the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0015] Please refer to Figure 1 , an embodiment of the present invention proposes a method for detecting depression of social media users based on symptom sequences. The method includes the following steps: Step 1: Input the original posted articles of the user and give a tweet sequence, and perform data cleaning on the tweet sequence to obtain the cleaned tweet sequence.
[0016] Step 2: Input the cleaned tweet sequence and the text sequence of the depression symptom descriptions in the DSM-5 manual into the sentence embedding model for encoding respectively to obtain the similarity score sequence between the post and the symptom description.
[0017] Please refer to Figure 2 , in step 2, inputting the cleaned tweet sequence and the text sequence of the depression symptom descriptions in the DSM-5 manual into the sentence embedding model for encoding respectively to obtain the similarity score sequence between the post and the symptom description specifically includes the following steps: Input the cleaned tweet sequence and the text sequence of the depression symptom descriptions in the DSM-5 manual into the sentence embedding model for encoding respectively to obtain the high-dimensional semantic vector representations of the user's tweets and the high-dimensional semantic vector representations of the symptom description texts; Calculate the cosine similarity between the high-dimensional semantic vector representation of the user's tweet and the high-dimensional semantic vector representation of the symptom description text to obtain the cosine similarity between the user's post and the symptom description text; Summarize the cosine similarities between the user's posts and the symptom description texts to obtain a similarity score sequence of the posts and the symptom descriptions.
[0018] Input the cleaned tweet sequence and the sequence of depressive symptom description texts in the DSM-5 manual into the sentence embedding model for encoding respectively, to obtain the high-dimensional semantic vector representations of the user's tweets and the symptom description texts respectively. The corresponding relationships in the process are as follows: ; where, represents the high-dimensional semantic vector representation of the user's tweet obtained through the sentence embedding model , represents the th tweet of the user, represents the high-dimensional semantic vector representation of the symptom description text obtained through the sentence embedding model , represents the th symptom description text; It should be noted that is more efficient when processing sentence pairs because it only needs to encode each sentence independently once, rather than jointly encoding each pair of sentences.
[0019] In the step of calculating the cosine similarity between the high-dimensional semantic vector representation of the user's tweet and the high-dimensional semantic vector representation of the symptom description text to obtain the cosine similarity between the user's post and the symptom description text, the corresponding relationships in the process are as follows: ; where, represents the cosine similarity calculated between the th tweet of the user and the th symptom description text, represents the cosine similarity calculation, represents the high-dimensional semantic vector representation of the th tweet of the user, represents the high-dimensional semantic vector representation of the th depressive description text, represents taking the modulus of the vector; In the step of summarizing the cosine similarities between the user's posts and the symptom description texts to obtain a similarity score sequence of the posts and the symptom descriptions, the corresponding relationships in the process are as follows: ; Among them, represents the sum of the cosine similarities between the th tweet of the user and 11 symptom description texts, represents the cosine similarity between the th tweet of the user and 1 symptom description text, represents the similarity score sequence between the post and the symptom descriptions, represents the sum of the cosine similarities between the 1st tweet of the user and 11 symptom description texts, represents the th tweet of the user and the sum of the cosine similarities between the user and 11 symptom description texts, represents the total number of user tweets.
[0020] Furthermore, the DSM-5 was published by the American Psychiatric Association. It provides standardized classification and diagnostic criteria for the diagnosis of mental disorders and is an important tool for mental health professionals such as psychiatrists and psychologists to diagnose mental disorders. It details the diagnostic criteria for various mental disorders, including the specific manifestations of symptoms, duration, severity, etc., to improve the reliability and consistency of diagnosis.
[0021] Step 3: Calculate the dynamic symptom indicators for the similarity score sequence between the post and the symptom descriptions to obtain a set of dynamic symptom indicators.
[0022] In the said Step 3, calculating the dynamic symptom indicators for the similarity score sequence between the post and the symptom descriptions to obtain a set of dynamic symptom indicators specifically includes the following steps: Calculate the dynamic symptom indicators for the similarity score sequence between the post and the symptom descriptions to separately obtain the average symptom score, symptom variability, average rise rate of all rising processes in the sequence, and average recovery rate of all recovery processes in the sequence; Summarize and aggregate the average symptom score, symptom variability, average rise rate of all rising processes in the sequence, and average recovery rate of all recovery processes in the sequence to obtain a set of dynamic symptom indicators.
[0023] When calculating the dynamic symptom indicators for the similarity score sequence between the post and the symptom descriptions, separately obtaining the average symptom score, symptom variability, average rise rate of all rising processes in the sequence, and average recovery rate of all recovery processes in the sequence, the corresponding relationships are as follows: ; Among them, represents the average symptom score, represents the standard deviation of the entire sequence, represents the symptom variability, represents the baseline interval of the symptom score, represents the The rising rate of an ascending process Denote the Score value of the nearest baseline point of the Denote the index of the ascending process Denote the index of the score peak Denote the th score peak in the user tweet sequence Denote the Time point corresponding to the nearest baseline point of the Denote the Recovery rate of the Denote the index of the recovery process Denote the Score value of the nearest baseline point of the Denote the th score valley in the user tweet sequence Denote the Time point of the nearest baseline point of the Denote the Time point corresponding to the th score valley Denote the number of valid ascending processes in the sequence Denote the average recovery rate of all recovery processes in the sequence Denote the number of valid recovery processes in the sequence; In the step of aggregating and collecting the average symptom score, symptom variability, average rising rate of all ascending processes in the sequence, and average recovery rate of all recovery processes in the sequence to obtain the dynamic index set of symptoms, the following relational expressions exist for the corresponding processes: ; where Denote the dynamic index set of symptoms
[0024] Step 4: Input the cleaned tweet sequence into the pre-trained model for sentiment analysis to generate an emotion score sequence, and input the emotion score sequence into a bidirectional gated recurrent unit network for emotion capture processing to obtain a contextually related emotion sequence
[0025] In step 4, input the cleaned tweet sequence into the pre-trained model for sentiment analysis to generate an emotion score sequence, and input the emotion score sequence into a bidirectional gated recurrent unit network for emotion capture processing to obtain a contextually related emotion sequence. The following relational expressions exist for the corresponding processes: ; Among them, represents the user's tweet sequence, represents the user's first tweet, represents the th tweet of the user, represents the obtained emotional score sequence, represents being processed by the Emotional BERT model, represents the context-related emotional score sequence, represents being processed by the bidirectional gated recurrent unit network, represents the th tweet of the user to obtain the context-related emotional score.
[0026] Step 5: Perform feature fusion on the similarity score sequence between the post and the symptom description, the set of dynamic indicators of the symptoms, and the context-related emotion sequence to obtain the overall feature representation of the user's tweet sequence.
[0027] In the said Step 5, perform feature fusion on the similarity score sequence between the post and the symptom description, the set of dynamic indicators of the symptoms, and the context-related emotion sequence to obtain the overall feature representation of the user's tweet sequence. The corresponding relationship in the process is as follows: ; Among them, represents the overall feature representation of the user's tweet sequence.
[0028] Step 6: Input the overall feature representation of the user's tweet sequence into a multivariate time series classifier for classification to obtain a classification result.
[0029] In the said Step 6, input the overall feature representation of the user's tweet sequence into a multivariate time series classifier for classification to obtain a classification result. The corresponding relationship in the process is as follows: ; Among them, represents the classification result output by the classifier, represents being processed by the multivariate time series classifier.
[0030] Please refer to Figure 3 , the present invention also proposes a social media user depression detection system based on a symptom sequence. The system includes: A data cleaning module, used for: Input the user's original published article and given tweet sequence, perform data cleaning on the tweet sequence to obtain the cleaned tweet sequence; A symptom feature extraction module, used for: The sequence of cleaned tweets and the sequence of text descriptions of depressive symptoms in the DSM-5 manual are respectively input into a sentence embedding model for encoding to obtain a sequence of similarity scores between the posts and the symptom descriptions; A symptom dynamic index extraction module, configured to: Calculate symptom dynamic indices for the sequence of similarity scores between the posts and the symptom descriptions to obtain a set of dynamic indices of the symptoms; An emotion feature extraction module, configured to: Input the sequence of cleaned tweets into a pre-trained model for sentiment analysis to generate a sequence of emotion scores, and input the sequence of emotion scores into a bidirectional gated recurrent unit network for emotion capture processing to obtain a contextually related emotion sequence; A feature fusion module, configured to: Fuse the sequence of similarity scores between the posts and the symptom descriptions, the set of dynamic indices of the symptoms, and the contextually related emotion sequence to obtain an overall feature representation of the user tweet sequence; A prediction module, configured to: Input the overall feature representation of the user tweet sequence into a multivariate time series classifier for classification to obtain a classification result.
[0031] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0032] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0033] The above-described embodiments merely represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. A method for detecting depression in social media users based on symptom sequences, characterized in that: The method comprises the following steps: Step 1: Input the original article published by the user and give a tweet sequence, perform data cleaning on the tweet sequence, and obtain the cleaned tweet sequence; Step 2: Input the cleaned tweet sequence and the text sequence of depressive symptom descriptions in the DSM-5 manual into the sentence embedding model for encoding to obtain a similarity score sequence between the post and the symptom description; Step 3: Calculate the symptom dynamic index for the similarity score sequence between the post and the symptom description to obtain a set of dynamic indexes for the symptom; Step 4: Input the cleaned tweet sequence into the pre-trained model for sentiment analysis to generate a sentiment score sequence, and input the sentiment score sequence into the bidirectional gated recurrent unit network for sentiment capture processing to obtain a context-associated sentiment sequence; Step 5: Feature fusion is performed on the similarity score sequence between the post and the symptom description, the dynamic indicator set of the symptom, and the context-related emotion sequence to obtain the overall feature representation of the user's tweet sequence; Step 6: Input the overall feature representation of the user's tweet sequence into the multivariate time series classifier for classification to obtain the classification result.
2. The method for detecting depression in social media users based on symptom sequences according to claim 1, characterized in that: In step 2, the cleaned tweet sequence and the text sequence of depression symptom descriptions in the DSM-5 manual are respectively input into the sentence embedding model for encoding to obtain a similarity score sequence between the post and the symptom description, which specifically includes the following steps: The cleaned tweet sequence and the text sequence describing depression symptoms in the DSM-5 manual are respectively input into the sentence embedding model for encoding, and the high-dimensional semantic vector representation of the user's tweets and the high-dimensional semantic vector representation of the symptom description text are obtained respectively; The cosine similarity between the high-dimensional semantic vector representation of the user's tweet and the high-dimensional semantic vector representation of the symptom description text is calculated to obtain the cosine similarity between the user's post and the symptom description text; The cosine similarities between user posts and symptom description texts are summarized to obtain a similarity score sequence between posts and symptom descriptions.
3. The method for detecting depression in social media users based on symptom sequences according to claim 2, characterized in that: The cleaned tweet sequence and the text sequence describing depression symptoms in the DSM-5 manual are respectively input into the sentence embedding model for encoding, and the high-dimensional semantic vector representation of the user's tweets and the high-dimensional semantic vector representation of the symptom description text are obtained respectively. The relationship between the corresponding processes is as follows: ; in, Represents the sentence embedding model The obtained high-dimensional semantic vector representation of the user's tweets, Indicates user Tweets, Represents the sentence embedding model The obtained high-dimensional semantic vector representation of the symptom description text is Indicates Symptom description text; In the step of calculating the cosine similarity between the high-dimensional semantic vector representation of the user's tweet and the high-dimensional semantic vector representation of the symptom description text, and obtaining the cosine similarity between the user's post and the symptom description text, the corresponding process has the following relationship: ; in, Indicates user Tweets and The cosine similarity calculated from the symptom description texts, represents the cosine similarity calculation, Indicates user The high-dimensional semantic vector representation of a tweet, Indicates A high-dimensional semantic vector representation of a depression description text. The modulus represents the orientation quantity; In the step of summarizing the cosine similarities between user posts and symptom description texts to obtain a similarity score sequence between posts and symptom descriptions, the corresponding process has the following relationship: ; in, Indicates user The sum of the cosine similarities between this tweet and 11 symptom description texts, Indicates user The cosine similarity between a tweet and a symptom description text, represents the similarity score sequence between the post and the symptom description, represents the sum of the cosine similarities between the user’s first tweet and the 11 symptom description texts, Indicates user The sum of the cosine similarities between this tweet and 11 symptom description texts, Represents the total number of tweets by the user.
4. The method for detecting depression in social media users based on symptom sequences according to claim 3, characterized in that: In step 3, the symptom dynamic index is calculated for the similarity score sequence between the post and the symptom description to obtain a set of dynamic indexes of the symptom, which specifically includes the following steps: The symptom dynamic index was calculated for the similarity score sequence between the post and the symptom description, and the average symptom score, symptom variability, the average rise rate of all rising processes in the sequence, and the average recovery rate of all recovery processes in the sequence were obtained respectively; The average symptom score, symptom variability, the average value of the rising rate of all rising processes in the sequence, and the average value of the recovery rate of all recovery processes in the sequence are summarized and collected to obtain a dynamic indicator set of symptoms.
5. The method for detecting depression in social media users based on symptom sequences according to claim 4, characterized in that: The symptom dynamic index is calculated for the similarity score sequence between the post and the symptom description, and the average symptom score, symptom variability, the average value of the rising rate of all rising processes in the sequence, and the average value of the recovery rate of all recovery processes in the sequence are obtained. The relationship between the corresponding processes is as follows: ; in, represents the average symptom score, represents the standard deviation of the entire series, Indicates variability of symptoms, represents the baseline interval of symptom scores, Indicates The rate of rise of the rising process, Indicates The score value of the nearest baseline point of the rising process, represents the index of the ascending process, represents the index of the peak score, Indicates the first Peak score The corresponding time point, Indicates The time point corresponding to the nearest baseline point of the rising process, Indicates The recovery rate of the recovery process, Indicates the index of the recovery process, Indicates The score value of the nearest baseline point of the recovery process, Indicates the first The score valley, represents the index of the score valley, Indicates The time point of the most recent baseline point of the recovery process, Indicates The time point corresponding to the score valley value is represents the average value of the rising rate of all rising processes in the sequence, represents the number of effective rising processes in the sequence, represents the average recovery rate of all recovery processes in the sequence, Indicates the number of valid recovery processes in the sequence; In the step of summarizing the average symptom score, symptom variability, the average value of the rising rate of all rising processes in the sequence, and the average value of the recovery rate of all recovery processes in the sequence to obtain the dynamic indicator set of the symptom, the relationship between the corresponding processes is as follows: ; in, A dynamic collection of metrics that represent symptoms.
6. The method for detecting depression in social media users based on symptom sequences according to claim 5, characterized in that: In step 4, the cleaned tweet sequence is input into the pre-trained model for sentiment analysis to generate a sentiment score sequence, and the sentiment score sequence is input into the bidirectional gated recurrent unit network for sentiment capture processing to obtain a context-related sentiment sequence. The corresponding process has the following relationship: ; in, represents a sequence of tweets from a user, Indicates the user's first tweet. Indicates the user's Tweets, represents the obtained emotion score sequence, Indicates that it has been processed by the emotional BERT model. represents a sequence of contextual sentiment scores, It means that it is processed by a bidirectional gated recurrent unit network. Indicates the user's The sentiment score of the context associated with a tweet.
7. The method for detecting depression in social media users based on symptom sequences according to claim 6, characterized in that: In step 5, the similarity score sequence between the post and the symptom description, the dynamic indicator set of the symptom and the emotion sequence associated with the context are feature fused to obtain the overall feature representation of the user's tweet sequence. The corresponding process has the following relationship: ; in, Represents the overall feature representation of a user's tweet sequence.
8. The method for detecting depression in social media users based on symptom sequences according to claim 7, characterized in that: In step 6, the overall feature representation of the user's tweet sequence is input into the multivariate time series classifier for classification to obtain the classification result. The relationship between the corresponding process is as follows: ; in, Represents the classification result of the output of the classifier, Indicates that it has been processed by a multivariate time series classifier.
9. A depression detection system for social media users based on symptom sequences, characterized in that: The system applies any one of the symptom sequence-based depression detection methods for social media users according to claims 1 to 8, and the system comprises: Data cleaning module, used to: Input the original article published by the user and give a tweet sequence, perform data cleaning on the tweet sequence, and obtain the cleaned tweet sequence; Symptom feature extraction module, used to: The cleaned tweet sequence and the text sequence of depressive symptom descriptions in the DSM-5 manual are respectively input into the sentence embedding model for encoding to obtain a sequence of similarity scores between the posts and the symptom descriptions; Symptom dynamic indicator extraction module, used for: Calculate the symptom dynamic index for the similarity score sequence between the post and the symptom description to obtain a set of dynamic indexes for the symptom; Emotion feature extraction module, used for: The cleaned tweet sequence is input into the pre-trained model for sentiment analysis to generate a sentiment score sequence, which is then input into a bidirectional gated recurrent unit network for sentiment capture processing to obtain a context-associated sentiment sequence. Feature fusion module, used for: The similarity score sequence between posts and symptom descriptions, the dynamic indicator set of symptoms, and the sentiment sequence associated with the context are fused to obtain the overall feature representation of the user's tweet sequence. Prediction module for: The overall feature representation of the user's tweet sequence is input into the multivariate time series classifier for classification to obtain the classification result.