Depression Detection Method Based on Text Analysis

By combining eye movement data, facial video and autonomic nerve signals, the automatic detection of depression is solved, and the problem of low recognition rate in the prior art is improved, and the detection ability of depression is improved, especially in the younger age group.

CN114300120BActive Publication Date: 2025-08-08WOMIN HIGH-TECH (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111522158.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2025-08-08
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

The detection and recognition rate of depression in the prior art is low, especially in hospitals above prefecture-level cities and cities, less than 20%, and less than 10% of patients receive drug treatment, making it difficult to effectively carry out the medical prevention and treatment of depression.

Method used

By obtaining eye movement data, facial videos and autonomic nerve signals when answering the questionnaire, using eye movement data to determine the target text segment, determining the weight of depression based on the target text segment, and combining autonomic nerve signals to detect depression on facial videos to achieve automatic detection.

Benefits of technology

Automatic detection of depression is achieved, detection recognition rate is improved, depression can be detected earlier, especially among young people, and the effectiveness of medical prevention and treatment is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114300120B_ABST
    Figure CN114300120B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for detecting depression based on text analysis. The method comprises: obtaining eye movement data, facial videos, and autonomic nerve signals from a user as they answer a questionnaire; determining a target text segment based on the eye movement data; determining a depression susceptibility weight based on the target text segment; and performing depression detection on the facial video based on the depression susceptibility weight and the autonomic nerve signals to obtain a depression detection result. The method provided by the present invention detects whether a user suffers from depression using the user's eye movement data, facial videos, and autonomic nerve signals as they answer the questionnaire, thereby achieving automatic detection of depression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of psychological assessment, and in particular to a method for detecting depression based on text analysis. Background Art

[0002] Currently, depression is the second most common illness in humans, second only to cardiovascular disease. Approximately 800,000 people commit suicide each year due to depression. Simultaneously, depression is beginning to strike younger people (including university students and even elementary and middle school students). However, medical treatment and prevention for depression still suffers from a low recognition rate. Hospitals at the prefecture level and above have a recognition rate of less than 20%, and less than 10% of patients receive relevant medication. Therefore, depression detection is crucial for its medical prevention. Summary of the Invention

[0003] (1) Technical issues to be solved

[0004] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a method for depression detection based on text analysis.

[0005] (2) Technical solution

[0006] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:

[0007] A method for detecting depression based on text analysis, the method comprising:

[0008] S101, obtaining eye movement data, facial video, and autonomic nerve signals of the user when answering the questionnaire;

[0009] S102, determining a target text segment based on the eye movement data;

[0010] S103, determining a depression morbidity weight based on the target text segment;

[0011] S104 , performing depression detection on the facial video based on the depression morbidity weight and the autonomic nerve signal to obtain a depression detection result.

[0012] Optionally, the questionnaire includes multiple questions, wherein each question includes a stem and options;

[0013] The S102 includes:

[0014] determining a target text segment corresponding to each question based on the eye movement data;

[0015] The S103 includes:

[0016] The depression morbidity weight is determined based on each target text segment.

[0017] Optionally, determining a target text segment corresponding to each question based on the eye movement data includes:

[0018] For any question,

[0019] Determining the eye movement trajectory of the user when answering any of the questions according to the eye movement data;

[0020] A target word is selected from any of the questions according to the eye movement trajectory, and the target word is used to form a target text segment corresponding to the any of the questions.

[0021] Optionally, selecting a target word from any one of the questions based on the eye movement trajectory and forming a target text segment corresponding to any one of the questions with the target word includes:

[0022] Determining each word where the eyes linger and how long they linger based on the eye movement trajectory;

[0023] Determine the segmentation corresponding to each word on which the eyes linger according to the pre-segmentation result of any of the questions, and determine the time value of the segmentation corresponding to each word on which the eyes linger according to the length of time the eyes linger on each word;

[0024] According to the order in which the eyes rest, the segmented words corresponding to the words where the eyes rest are used to form a target text segment corresponding to any of the questions, wherein each segmented word in the target text segment has a time value.

[0025] Optionally, determining the segmentation corresponding to each word on which the eyes linger based on the pre-segmentation result of any of the questions, and determining the time value of the segmentation corresponding to each word on which the eyes linger based on the duration of the lingering eyes on each word, includes:

[0026] For any word that your eyes rest on in any of the questions above,

[0027] If the any character is the character on which the gaze first stops, the segmentation corresponding to the any character is confirmed in the pre-segmentation result of the any question, and is used as the segmentation corresponding to the any character; the duration of the gaze stopping on the any character is determined as the time value of the segmentation corresponding to the any character, and the stop order of the segmentation corresponding to the any character is determined to be 1;

[0028] If any of the characters is not the first character where the gaze stops, the segmented word corresponding to the character, the time value of the segmented word corresponding to the character, and the stop order of the segmented word corresponding to the character are determined according to the previous stop position of the gaze.

[0029] Optionally, determining the segmented word corresponding to any character, the time value of the segmented word corresponding to any character, and the order of the segmented words corresponding to any character according to the previous stay position of the gaze includes:

[0030] If the previous stop position is not the position where the eyes stopped at the previous word, the word corresponding to the any word is confirmed in the pre-segmentation result of any question, and is used as the word corresponding to the any word. At the same time, a place-filling participle is added before the word corresponding to the any word; the duration of the eyes stopping at the any word is determined as the time value of the word corresponding to the any word, and the difference between the first moment the eyes stop at the any word and the last moment the eyes stop at the previous word is determined as the time value of the place-filling participle; the stop order of the place-filling participle is determined as the stop order of the word corresponding to the previous word plus 1, and the stop order of the word corresponding to the any word is determined as the stop order of the place-filling participle plus 1;

[0031] If the previous stay position is the position of the previous character where the eyes stayed, then when the previous character where the eyes stayed is the same as any character, the participle corresponding to the previous character where the eyes stayed is used as the participle corresponding to the any character, and the stay time of the participle corresponding to the previous character where the eyes stayed is increased by the stay time of the eyes staying on the any character; when the previous character where the eyes stayed is different from the any character, the participle corresponding to the any character is confirmed in the pre-segmentation result of any question, and is used as the participle corresponding to the any character, the stay time of the eyes staying on the any character is determined as the time value of the participle corresponding to the any character, and the stay order of the participle corresponding to the any character is determined as the stay order of the participle corresponding to the previous character plus 1.

[0032] Optionally, determining the depression morbidity weight based on each target text segment includes:

[0033] Determine the influence of each target text segment;

[0034] Determine the weight of depression = ∑ u α u *β u *ε u ;

[0035] Among them, u is the target text segment identifier, α u is the weight of the target text segment u, β u is the influence of the target text segment u, ε u is the answer score of the question corresponding to the target text segment u.

[0036] Optionally, determining the influence of each target text segment includes:

[0037] For any target text segment, the influence of the target text segment is determined by the following formula:

[0038]

[0039] Where v is the word segmentation identifier in any target text segment, γ v is the weight of the word v in any target text segment, t v is the time value of the word v in any target text segment, t max is the maximum time value of all word segments in any target text segment, W v is the influence value of the predetermined participle v;

[0040] If the participle v is not a complement participle, then If the participle v is a complement participle, then

[0041] n1 is the number of non-place-filling participles, n2 is the number of place-filling participles, t1 is the sum of the time values of the non-place-filling participles, t2 is the sum of the time values of the place-filling participles, and N is the total number of participles in any target text segment.

[0042] Optionally, the S104 includes:

[0043] S104-1, forming a signal set from the autonomic nerve signals; wherein each element in the signal set corresponds to an autonomic nerve signal value collected at a time, and the elements in the signal set are arranged from far to near according to the collection time;

[0044] S104-2, determining the difference between each two adjacent elements in the signal set to form a signal difference set, where the difference is the value of the latter element minus the value of the previous element;

[0045] S104-3, determining the standard deviation σ of all elements in the signal difference set Δ ;

[0046] S104-4, determining the element with the largest value in the signal difference set The element with the smallest sum

[0047] S104-5, determining the element a with the largest value in the signal set max The element a with the smallest sum min , and the time t corresponding to the element with the largest value max The time t corresponding to the element with the smallest value min ;

[0048] S104-6, determine the emotion coefficient

[0049] S104-7, performing depression detection on the facial video based on the emotion coefficient and the depression morbidity weight to obtain a depression detection result.

[0050] Optionally, the S104-7 includes:

[0051] identifying micro-expressions in each frame of the facial video;

[0052] Determine the degree of variation between micro-expressions in each frame;

[0053] Determine the maximum number of consecutive frames whose degree of change is not greater than a change threshold;

[0054] Determine the detection value = maximum number * the depression morbidity weight * I1;

[0055] If the detection value is greater than the depression threshold, it is determined that depression is detected.

[0056] (3) Beneficial effects

[0057] The method obtains eye movement data, facial video, and autonomic nerve signals from the user as they answer a questionnaire; determines a target text segment based on the eye movement data; determines a depression weight based on the target text segment; and performs a depression test on the facial video based on the depression weight and autonomic nerve signals to obtain a depression test result. The method provided by the present invention uses eye movement data, facial video, and autonomic nerve signals from the user as they answer the questionnaire to determine whether the user has depression, thereby achieving automatic depression detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 A flowchart of a method for detecting depression based on text analysis provided by one embodiment of the present invention;

[0059] Figure 2 A schematic diagram of an eye movement trajectory provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0060] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.

[0061] Currently, depression is the second most common illness in humans, second only to cardiovascular disease. Approximately 800,000 people commit suicide each year due to depression. Simultaneously, depression is beginning to strike younger people (including university students and even elementary and middle school students). However, medical treatment and prevention for depression still suffers from a low recognition rate. Hospitals at the prefecture level and above have a recognition rate of less than 20%, and less than 10% of patients receive relevant medication. Therefore, depression detection is crucial for its medical prevention.

[0062] Based on this, the present invention provides a method for detecting depression based on text analysis, which includes: obtaining eye movement data, facial videos and autonomic nerve signals of users when answering questionnaires; determining a target text segment based on the eye movement data; determining a depression susceptibility weight based on the target text segment; performing depression detection on facial videos based on the depression susceptibility weight and autonomic nerve signals to obtain a depression detection result, thereby realizing automatic detection of depression.

[0063] In the specific implementation, a questionnaire will be provided to the user. Figure 1 The method shown detects whether the user has a depressive tendency, and then determines whether the user suffers from depression.

[0064] See also Figure 1 The implementation process of the method for detecting depression based on text analysis provided in this embodiment is as follows:

[0065] S101, obtaining the user's eye movement data, facial video, and autonomic nerve signals when answering the questionnaire.

[0066] This step uses existing solutions to obtain eye movement data, facial videos, and autonomic nerve signals from users as they answer the questionnaire. For example, eye movement data can be obtained using an eye tracker. This will not be elaborated on here.

[0067] In addition, the questionnaire includes multiple questions, wherein each question includes a stem and options. In other words, the questionnaire is multiple-choice questions, and users need to make a choice.

[0068] In addition, when designing the questionnaire, each question is assigned a value that represents its contribution to the diagnosis of depression. The more a question reflects the degree of depression, the larger its value, which serves as the weight for the target text segment corresponding to that question.

[0069] After designing the questionnaire, we preprocessed it, performing word segmentation to identify the corresponding words for each question. For example, the question "Do you have difficulty falling asleep?" was segmented to identify "you," "whether," "sleep," and "difficulty." The word segmentation scheme is an existing one and will not be detailed here.

[0070] S102: Determine a target text segment based on the eye movement data.

[0071] Since the questionnaire includes multiple questions, this step will determine the target text segment corresponding to each question based on the eye movement data.

[0072] Each question includes a stem and options.

[0073] Specifically, for any question,

[0074] 1. Determine the user's eye movement trajectory when answering any question based on eye movement data.

[0075] This step will be implemented based on the existing method.

[0076] For example, for the question "Do you have difficulty falling asleep? A. Yes B. No", the eye movement trajectory obtained is Figure 2 As shown. They are 1-2-3-4-5-6-7-8 respectively.

[0077] Figure 2 This is just a schematic diagram; the actual eye movement trajectory may be more complicated.

[0078] 2. Select target words from any question based on the eye movement trajectory, and form the target words into a target text segment corresponding to any question.

[0079] Specifically,

[0080] 1) Based on the eye movement trajectory, determine the words where the eyes linger and how long they linger.

[0081] For example, the word for the eye gaze that is labeled 1 (for the convenience of description, it is recorded as stay position 1) is "yes", and the stay duration T1 at stay position 1 is recorded. The eye gaze that is labeled 2 (for the convenience of description, it is recorded as stay position 2) does not correspond to a word, so it is not processed. The word for the eye gaze that is labeled 3 (for the convenience of description, it is recorded as stay position 3) is "no", and the stay duration T3 at stay position 3 is recorded. The word for the eye gaze that is labeled 4 (for the convenience of description, it is recorded as stay position 4) is "enter", and the stay duration T4 at stay position 4 is recorded. The word for the eye gaze that is labeled 5 (for the convenience of description, it is recorded as stay position 5) is also "enter", and the stay duration T5 at stay position 5 is recorded. The word for the eye gaze that is labeled 6 (for the convenience of description, it is recorded as stay position 6) is "sleep", and the stay duration T6 at stay position 6 is recorded. The eye gaze stop numbered 7 (for the sake of convenience, recorded as stop position 7) does not correspond to a word, so it is not processed. The eye gaze stop numbered 8 (for the sake of convenience, recorded as stop position 8) has a word "yes", and the stop duration T8 of stop position 8 is recorded.

[0082] Finally, the corresponding table shown in Table 1 is obtained.

[0083] Table 1

[0084]

[0085] It should be noted that if the stop position corresponds to a punctuation mark, it will not be processed. In other words, as long as it does not correspond to text, it will not be processed.

[0086] What is not processed can be the blank shown in Table 1, or a preset character, such as "0", or "NULL", etc. This embodiment does not limit this.

[0087] 2) According to the pre-segmented result of any question, determine the segments corresponding to each character where the eye gaze stays, and determine the time value of the segments corresponding to each character where the eye gaze stays according to the stay duration of each character where the eye gaze stays.

[0088] Specifically, for any character where the eye gaze stays in any question

[0089] · If any character is the first character where the eye gaze stays, then

[0090] 1.1 Confirm the segment corresponding to any character in the pre-segmented result of any question, and use it as the segment corresponding to any character.

[0091] Since after designing the questionnaire, the questionnaire will be preprocessed, such as segmenting words, to obtain the words corresponding to each question in the questionnaire. This step is to find the word corresponding to the character in the obtained words, so as to obtain the word where the character at the eye gaze position is located, and then perform analysis based on the word.

[0092] 1.2 Determine the stay duration of any character where the eye gaze stays as the time value of the segment corresponding to any character, and determine the stay order of the segment corresponding to any character as 1.

[0093] For example, the "yes" corresponding to the stay position 1, which is the first character where the eye gaze stays (it should be noted that in actual application, the first character where the eye gaze stays may not be the first stay position. If the first stay is at a blank, or a symbol position, then it is not the first character where the eye gaze stays. Here, it is only based on whether the first stay is on a character).

[0094] Then, confirm that the segment corresponding to "yes" in the pre-segmented result is "whether", and use "whether" as the segment corresponding to the character "yes" (the character corresponding to the stay position 1). Determine T1 as the time value of the segment "whether" (the segment corresponding to the character corresponding to the stay position 1), and determine the stay order of the segment "whether" (the segment corresponding to the character corresponding to the stay position 1) as 1. As shown in Table 2.

[0095] Table 2

[0096]

[0097] · If any character is not the first character where the eye gaze stays, then​​​​​For example, “no,” “enter,” “enter,” “sleep,” and the last option, “yes,” are all words that are not the first ones that the eyes linger on.

[0100] The specific implementation plan is:

[0101] If the previous stop position is not the position where the eyes stopped before the word, then

[0102] 2.1.1 Confirm the participle corresponding to any character in the pre-segmentation results of any question and use it as the participle corresponding to any character. At the same time, add a placeholder participle before the participle corresponding to any character.

[0103] 2.1.2 The duration of gaze resting on any character is determined as the time value of the participle corresponding to any character, and the difference between the first moment of gaze resting on any character and the last moment of gaze resting on the previous character is determined as the time value of the complement participle.

[0104] 2.1.3 The order of the complement participle is determined as the order of the participle corresponding to the previous character when the eyes stay on it plus 1, and the order of the participle corresponding to any character is determined as the order of the complement participle plus 1.

[0105] For example, for "no" in stop position 3, its previous stop position is stop position 2, and stop position 2 is not the position of the word before the eyes stop (i.e., stop position 1). Then, confirm that the participle corresponding to "no" is "whether" in the pre-segmentation result, and use "whether" as the participle corresponding to the word "no" (the word corresponding to stop position 3). At the same time, add a place-filling participle before the participle corresponding to "no" (the word corresponding to stop position 3) (such as adding place-filling participle 1, the added place-filling participle can be filled with a preset symbol, or NULL, or not filled with content, which is not limited in this embodiment). Determine T3 as the time value of "whether" (the participle corresponding to the word corresponding to stop position 3), and the difference between the first moment when the eyes stay on "no" (the word corresponding to stop position 3) and the last moment when the eyes stay on "yes" (the word corresponding to stop position 1) is determined as the time value of place-filling participle 1 (such as being recorded as T 补1 The order of the eye's gaze resting on the participle corresponding to "yes" (i.e., the participle corresponding to the character corresponding to rest position 1) is determined to be the order of the eye's gaze resting on the participle corresponding to "yes" (i.e., 1) plus 1 = 2. The order of the eye's gaze resting on "yes" (i.e., the participle corresponding to the character corresponding to rest position 3) is determined to be the order of the eye's gaze resting on the participle corresponding to "yes" (i.e., 2) plus 1 = 3. As shown in Tables 3 and 4.

[0106] Table 3

[0107]

[0108] Table 4

[0109] Complementary participle Complementary participle 1 Time value <![CDATA[T 补1 ]]> Stay order 2

[0110] If the previous stay position is the position of the character before the eye fixation, then

[0111] 2.2.1 When the character before the eye fixation is different from any character, confirm the segmentation corresponding to any character in the pre-segmented result of any question, and use it as the segmentation corresponding to any character. Determine the stay duration of the eye fixation on any character as the time value of the segmentation corresponding to any character, and determine the stay order of the segmentation corresponding to any character as the stay order of the eye fixation on the segmentation corresponding to the previous character plus 1.

[0112] For example, for the character "入" at the stay position 4, its previous stay position is the stay position 3, and the stay position 3 is the position of the character before the eye fixation (i.e., the stay position 3). Moreover, the character before the eye fixation (i.e., the character "否" corresponding to the stay position 3) is different from "入" (i.e., the character corresponding to the stay position 4). At this time, confirm that the segmentation corresponding to "入" in the pre-segmented result is "入睡", and use "入睡" as the segmentation corresponding to the character "入". Determine T4 as the time value of "入睡" (the segmentation corresponding to the character at the stay position 4), and determine the stay order of "入睡" (the segmentation corresponding to the character at the stay position 1) as the stay order of the eye fixation on the segmentation corresponding to the previous character (i.e., the segmentation corresponding to the character at the stay position 3, which is "是否") plus 1 = 4. As shown in Table 5.

[0113] Table 5

[0114]

[0115] 2.2.2 When the character before the eye fixation is the same as any character, use the segmentation corresponding to the character before the eye fixation as the segmentation corresponding to any character, and increase the stay duration of the segmentation corresponding to the character before the eye fixation by the stay duration of the eye fixation on any character.

[0116] For example, for the character "入" at the stay position 5, its previous stay position is the stay position 4, and the stay position 4 is the position of the character before the eye fixation (i.e., the stay position 4). However, the character before the eye fixation (i.e., the segmentation corresponding to the character at the stay position 4, which is "入") is the same as "入" (the segmentation corresponding to the character at the stay position 5). At this time, use the segmentation corresponding to the character before the eye fixation (i.e., the segmentation corresponding to the stay position 4) as the segmentation corresponding to the character at the stay position 5 (that is, the "入" corresponding to the stay position 4 and the "入" corresponding to the stay position 5 have the same corresponding segmentation, both being "入睡"), and increase the stay duration of the segmentation corresponding to the character before the eye fixation by the stay duration of the eye fixation on any character (i.e., T4 + T5). As shown in Table 6.

[0117] Table 6

[0118]

[0119] For another example, for the word "sleep" at stay position 6, its previous stay position is stay position 5. Stay position 5 is the position of the word before the eye fixation (i.e., stay position 5). Moreover, the word before the eye fixation (i.e., the word "enter" corresponding to stay position 5) is different from "sleep" (i.e., the word corresponding to stay position 6). At this time, confirm that the segmentation of "sleep" in the pre-segmented result is "fall asleep", and take "fall asleep" as the segmentation corresponding to the word "sleep". Determine the time value of T6 as the time value of "fall asleep" (the segmentation corresponding to the word at stay position 6), and determine the stay order of "fall asleep" (the segmentation corresponding to the word at stay position 6) as the stay order of the segmentation corresponding to the word before the eye fixation (i.e., the segmentation "fall asleep" corresponding to the word at stay position 5) plus 1 = 5. As shown in Table 7.

[0120] Table 7

[0121]

[0122] It can be seen from this that only when the words corresponding to the two consecutive stay positions are the same, they can share one segmentation. Otherwise, they are processed as two segmentations. Even if two different words correspond to the same segmentation, they will still be processed as two segmentations.

[0123] For the word "is" at stay position 8, its previous stay position is stay position 7. Stay position 7 is not the position of the word before the eye fixation (i.e., stay position 6). Then, confirm that the segmentation of "is" in the pre-segmented result is "A. is", and take "A. is" as the segmentation corresponding to the word "is" (the word at stay position 8). At the same time, add a padding segmentation before the segmentation corresponding to the word "is" (e.g., add padding segmentation 2). Determine the time value of T8 as the time value of "A. is" (the segmentation corresponding to the word at stay position 8), and determine the time value of the padding segmentation 2 as the difference between the first moment when the eyes fixate on "is" (the word at stay position 8) and the last moment when the eyes fixate on "sleep" (the word at stay position 6) (e.g., denoted as T 补2 ). Determine the stay order of the padding segmentation 2 as the stay order of the segmentation corresponding to "sleep" (i.e., the segmentation "fall asleep" corresponding to the word at stay position 6) plus 1 = 6, and determine the stay order of "A. is" (the segmentation corresponding to the word at stay position 8) as the stay order of the padding segmentation 2 (i.e., 6) plus 1 = 7. As shown in Table 8 and Table 9.

[0124] Table 8

[0125]

[0126] Table 9

[0127] Complementary participle Complementary participle 1 Complementary participle 2 Time value <![CDATA[T 补1 ]]> <![CDATA[T 补2 ]]> Stay order 2 6

[0128] It should be noted that in this embodiment, when determining word segmentation, it is not judged only based on characters, but based on characters and their positions. For example, for the same character '是' (meaning 'is'), the '是' at停留位置1 (it's not clear what '停留位置1' exactly means, maybe 'position 1 of eye fixation'?) is different from the '是' at停留位置8, and their corresponding word segmentations are also different.

[0129] 3) According to the order of eye fixation, form the target text segment corresponding to each question by combining the word segmentations corresponding to each character where the eye fixation stays.

[0130] Among them, each word segmentation in the target text segment has a time value.

[0131] The finally obtained target text segment is shown in Table 10.

[0132] Table 10

[0133] Stay order 1 2 3 4 5 6 7 Corresponding participles whether Complementary participle 1 whether Falling asleep Falling asleep Complementary participle 2 A. Yes Time value <![CDATA[T1]]> <![CDATA[T 补1 ]]> <![CDATA[T3]]> <![CDATA[T4+T5]]> <![CDATA[T6]]> <![CDATA[T 补2 ]]> <![CDATA[T8]]>

[0134] S103. Determine the weight of depression based on the target text segment.

[0135] Since in step S102, the target text segment corresponding to each question will be determined according to the eye movement data, in this step, the weight of depression will be determined based on each target text segment.

[0136] Specifically,

[0137] 1. Determine the influence degree of each target text segment.

[0138] For any target text segment, determine the influence degree of any target text segment through the following formula:

[0139]

[0140] Among them, v is the word segmentation identifier in any target text segment, γ v is the weight of the word segmentation v in any target text segment, t v is the time value of the word segmentation v in any target text segment, t max is the maximum time value of all word segmentations in any target text segment, W v is the pre-determined influence value of the word segmentation v.

[0141] If the word segmentation v is not a padding word segmentation, then If the word segmentation v is a padding word segmentation, then

[0142] n1 is the number of non-padding word segmentations, n2 is the number of padding word segmentations, t1 is the sum of the time values of non-padding word segmentations, t2 is the sum of the time values of padding word segmentations, and N is the total number of word segmentations in any target text segment.

[0143] Take the target text segment shown in Table 10 as an example, which includes 7 participles (i.e., N=7). Among them, the number of non-place-complementing participles is 5 (i.e., n1=5), and the number of place-complementing participles is 2 (i.e., n2=2). The sum of the time values of the non-place-complementing participles is T1+T3+T4+T5+T6+T8 (i.e., t1=T1+T3+T4+T5+T6+T8), and the sum of the time values of the place-complementing participles is T 补1 +T 补2 (i.e. t2 = T 补1 +T 补2 ).

[0144] Then for non-complementary participles,

[0145] For the complement participle,

[0146] Among them, v is the word segmentation mark in any target text segment, γ v is the weight of the word v in any target text segment, t v is the time value of word v in any target text segment, t max is the maximum time value of all word segments in any target text segment, W v is the influence value of the predetermined word v.

[0147] The frequency of each word in patients with depression can be counted in advance, and the relationship between words and word frequencies can be recorded. v When the word frequency corresponding to the participle v is used as W v .

[0148] When calculating the weights in this step, we take into account the inattention of patients with depression, which is reflected in the large number of place-filling participles and long time values (the place-filling participles do not correspond to words in the corresponding rest positions, reflecting the lack of concentration). Therefore, the weight of each participle is determined by the number and duration, reflecting the weight's representation of the tendency towards depression.

[0149] 2. Determine the weight of depression = ∑ u α u *β u *ε u .

[0150] Among them, u is the target text segment identifier, α u is the weight of the target text segment u, β u is the influence of the target text segment u, ε u is the answer score of the question corresponding to the target text segment u.

[0151] When designing the questionnaire, a value is determined for each question, which represents the contribution of each question to the discrimination of depression. The higher the degree of depression a question can reflect, the larger its value is, and this value is used as the weight α of the target text segment corresponding to the question. u .

[0152] Similarly, when designing the questionnaire, the degree to which each answer to each question represents the tendency towards depression will be determined. For example, each answer corresponds to a score, and the higher the score, the more severe the tendency towards depression. This value is ε u .

[0153] S104, performing depression detection on the facial video based on the depression disease weight and the autonomic nerve signal to obtain a depression detection result.

[0154] S104-1, forming the autonomic nerve signal into a signal set.

[0155] Among them, each element in the signal set corresponds to the autonomic nerve signal value collected at a moment, and the elements in the signal set are arranged from far to near according to the collection time.

[0156] For example, the signal set S includes 5 elements, and the set S={S0, S1, S2, S3, S4}.

[0157] S104-2, determining the difference between each two adjacent elements in the signal set to form a signal difference set.

[0158] The difference is the value of the next element minus the value of the previous element.

[0159] For example, the signal difference set ΔS includes 4 elements, and the set ΔS={S1-S0, S2-S1, S3-S2, S4-S3}.

[0160] If S1-S0 is denoted as a0, S2-S1 is denoted as a1, S3-S2 is denoted as a2, and S4-S3 is denoted as a3, then ΔS={a0, a1, a2, a3}.

[0161] That is, any element in the set ΔS (such as a j ), whose value is S j+1 -S j , that is, a j =S j+1 -S j .

[0162] S104-3, determine the standard deviation σ of all elements in the signal difference set Δ .

[0163] That is, determine the standard deviation σ of all elements in the signal difference set ΔS Δ .

[0164] For example, if ΔS = {a0, a1, a2, a3}, then

[0165]

[0166] S104-4, determine the element with the largest value in the signal difference set The element with the smallest sum

[0167] Confirm

[0168] Among them, max{} is the maximum value function, and min{} is the minimum value function.

[0169] S104-5, determine the element a with the largest value in the signal set max The element a with the smallest sum min , and the time t corresponding to the element with the largest value max The time t corresponding to the element with the smallest value min .

[0170] That is to determine a max =max{S0, S1, S2, S3, S4}, a min =min{S0, S1, S2, S3, S4}.

[0171] a max The corresponding collection time is t max , a min The corresponding collection time is t min .

[0172] S104-6, determine the emotion coefficient

[0173] The clinical manifestations of depression are bad mood and unhappiness in real life, long-term low mood, from the initial melancholy to the final grief, inferiority, pain, pessimism, and world-weariness. It feels like every day of living is a desperate torture to oneself, negative, escapist, and finally even has suicidal attempts and behaviors.

[0174] Patients with depression do not actively interact with the outside world, that is, they are slow to respond to external stimuli. The autonomic nerve signal value can reflect the user's response to the current emotional stimulus. Therefore, the longer the time between the maximum response and the minimum response (such as t max -t min The greater the value of the maximum response, the greater the possibility of depression. max -a minThe smaller the value), the less sensitive they are to external stimuli, and the greater the possibility of depression.

[0175] In addition, patients with depression may also overreact. Δ Characterizes the discrete degree of change in the autonomic nerve signal value between two times. If σ Δ The larger the value, the more obvious the mood swings are and the greater the possibility of depression. It represents the gap between the maximum and minimum changes. The larger the gap, the more intense the reaction and the greater the possibility of depression.

[0176] S104-7, performing depression detection on the facial video based on the emotion coefficient and the depression weight, and obtaining a depression detection result.

[0177] Specifically,

[0178] 1. Identify micro-expressions in each frame of a facial video.

[0179] This step adopts the existing micro-expression recognition solution, which will not be described here.

[0180] 2. Determine the degree of variation between micro-expressions in each frame.

[0181] Existing micro-expression analysis methods are also used here to determine the expression changes between the previous and next frames.

[0182] The degree of change can be represented by multiple factors, for example, the degree of change is the number of micro-expression feature points that have changed, or the degree of change is the number of micro-expression feature points that have changed / the total number of micro-expression feature points, or the degree of change is the average distance between the micro-expression feature points that have changed, where the average distance is ∑ the position difference between each micro-expression feature point in the previous and next frames.

[0183] 3. Determine the maximum number of consecutive frames whose degree of change is not greater than the change threshold.

[0184] The change threshold is an empirical value that can be set in advance or obtained through training with sample data.

[0185] In addition, the continuous frames include all frames involved in the degree of change not greater than the change threshold. In other words, if the degree of change is obtained by taking the difference between two frames, then both frames are involved.

[0186] In this step, the degree of change of two adjacent frames is calculated in sequence, and the relationship between each degree of change and the change threshold is determined respectively.

[0187] For example, as shown in Table 11:

[0188] Table 11

[0189]

[0190] In the data shown in Table 1, there are three consecutive frames whose degree of change is not greater than the change threshold. The first segment is frames F0 and F1 corresponding to D0, the second segment is frames F2, F3 and F4 corresponding to D2 and D3, and the third segment is frames F5, F6, F7 and F8 corresponding to D5, D6 and D7.

[0191] Then the maximum number is 4 (ie F5, F6, F7 and F8).

[0192] The maximum number of consecutive frames represents the longest number of frames during which the user remains unresponsive. Frames also exist in a temporal sequence, meaning the number of frames per minute in a video is fixed, and the number of frames can reflect the length of time, or the longest period of time during which the user remains unresponsive to emotional stimulation. The longer this time, the greater the likelihood of depression.

[0193] 4. Determine the detection value = maximum number * depression weight * I1.

[0194] Among them, I1 is the emotion coefficient.

[0195] 5. If the detection value is greater than the depression threshold, depression is determined to be detected.

[0196] The depression threshold is an empirical value that can be set in advance or obtained through training with sample data.

[0197] The method of this embodiment obtains eye movement data, facial video, and autonomic nerve signals of the user when answering a questionnaire; determines a target text segment based on the eye movement data; determines a depression susceptibility weight based on the target text segment; and performs depression detection on the facial video based on the depression susceptibility weight and the autonomic nerve signals to obtain a depression detection result, thereby realizing automatic detection of depression.

[0198] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0199] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0200] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.

[0201] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0202] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0203] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention shall also include such modifications and variations.

Claims

1. A method for detecting depression based on text analysis, characterized in that: The method comprises: S101, obtaining eye movement data, facial video, and autonomic nerve signals of a user when answering a questionnaire; the questionnaire includes a plurality of questions, each of which includes a stem and options; S102, determining a target text segment corresponding to each question based on the eye movement data; Specifically, for any question, the eye movement trajectory of the user when answering the question is determined based on the eye movement data; based on the eye movement trajectory, each word on which the eye rests and the length of time the eye rests are determined; based on the pre-segmented word result of the question, the word segmentation corresponding to each word on which the eye rests is determined, and the time value of each word segmentation corresponding to each word is determined based on the length of time the eye rests on each word; according to the order in which the eye rests, the word segmentation corresponding to each word on which the eye rests is formed into a target text segment corresponding to the question, wherein each word segmentation in the target text segment has a time value; S103, determining a depression morbidity weight based on each target text segment; S104 , performing depression detection on the facial video based on the depression morbidity weight and the autonomic nerve signal to obtain a depression detection result.

2. The method according to claim 1, characterized in that The step of determining the segmentation corresponding to each word on which the eyes linger based on the pre-segmentation result of any of the questions, and determining the time value of the segmentation corresponding to each word on which the eyes linger based on the duration of the lingering time of each word, includes: For any word that your eyes rest on in any of the questions above, If the any character is the character on which the gaze first stops, the segmentation corresponding to the any character is confirmed in the pre-segmentation result of the any question, and is used as the segmentation corresponding to the any character; the duration of the gaze stopping on the any character is determined as the time value of the segmentation corresponding to the any character, and the stop order of the segmentation corresponding to the any character is determined to be 1; If any of the characters is not the first character where the gaze stops, the segmented word corresponding to the character, the time value of the segmented word corresponding to the character, and the stop order of the segmented word corresponding to the character are determined according to the previous stop position of the gaze.

3. The method according to claim 2, characterized in that The step of determining the segmented word corresponding to any character, the time value of the segmented word corresponding to any character, and the sequence of the segmented words corresponding to any character according to the previous stop position of the gaze comprises: If the previous stop position is not the position where the eyes stopped at the previous word, the word corresponding to the any word is confirmed in the pre-segmentation result of any question, and is used as the word corresponding to the any word. At the same time, a place-filling participle is added before the word corresponding to the any word; the duration of the eyes stopping at the any word is determined as the time value of the word corresponding to the any word, and the difference between the first moment the eyes stop at the any word and the last moment the eyes stop at the previous word is determined as the time value of the place-filling participle; the stop order of the place-filling participle is determined as the stop order of the word corresponding to the previous word plus 1, and the stop order of the word corresponding to the any word is determined as the stop order of the place-filling participle plus 1; If the previous stay position is the position of the previous character where the eyes stayed, then when the previous character where the eyes stayed is the same as any character, the participle corresponding to the previous character where the eyes stayed is used as the participle corresponding to the any character, and the stay time of the participle corresponding to the previous character where the eyes stayed is increased by the stay time of the eyes staying on the any character; when the previous character where the eyes stayed is different from the any character, the participle corresponding to the any character is confirmed in the pre-segmentation result of any question, and is used as the participle corresponding to the any character, the stay time of the eyes staying on the any character is determined as the time value of the participle corresponding to the any character, and the stay order of the participle corresponding to the any character is determined as the stay order of the participle corresponding to the previous character plus 1.

4. The method according to claim 3, characterized in that The step of determining the depression morbidity weight based on each target text segment includes: Determine the influence of each target text segment; Determine the weight of depression = ∑ u α u *β u *ε u ; Among them, u is the target text segment identifier, α u is the weight of the target text segment u, β u is the influence of the target text segment u, ε u is the answer score of the question corresponding to the target text segment u.

5. The method according to claim 4, characterized in that Determining the influence of each target text segment includes: For any target text segment, the influence of the target text segment is determined by the following formula: Where v is the word segmentation identifier in any target text segment, γ v is the weight of the word v in any target text segment, t v is the time value of the word v in any target text segment, t max is the maximum time value of all word segments in any target text segment, W v is the influence value of the predetermined participle v; If the participle v is not a complement participle, then If the participle v is a complement participle, then n1 is the number of non-place-filling participles, n2 is the number of place-filling participles, t1 is the sum of the time values of the non-place-filling participles, t2 is the sum of the time values of the place-filling participles, and N is the total number of participles in any target text segment.

6. The method according to claim 1, wherein The S104 includes: S104-1, forming a signal set from the autonomic nerve signals; wherein each element in the signal set corresponds to an autonomic nerve signal value collected at a time, and the elements in the signal set are arranged from far to near according to the collection time; S104-2, determining the difference between each two adjacent elements in the signal set to form a signal difference set, where the difference is the value of the latter element minus the value of the previous element; S104-3, determining the standard deviation σ of all elements in the signal difference set Δ ; S104-4, determining the element with the largest value in the signal difference set The element with the smallest sum S104-5, determining the element a with the largest value in the signal set max The element a with the smallest sum min , and the time t corresponding to the element with the largest value max The time t corresponding to the element with the smallest value min ; S104-6, determine the emotion coefficient S104-7, performing depression detection on the facial video based on the emotion coefficient and the depression morbidity weight to obtain a depression detection result.

7. The method according to claim 6, characterized in that Said S104-7 includes: identifying micro-expressions in each frame of the facial video; Determine the degree of variation between micro-expressions in each frame; Determine the maximum number of consecutive frames whose degree of change is not greater than a change threshold; Determine the detection value = maximum number * the depression morbidity weight * I1; If the detection value is greater than the depression threshold, it is determined that depression is detected.

Citation Information

Patent Citations

  • Method and device for assessing information flow creativity

    CN110432915A

  • Device, system and method for determining a stress level of a user

    CN112673434A

  • Psychological evaluation method based on dynamic video data

    CN113707294A