A method for generating a mental health screening and conversation dataset

CN117912667BActive Publication Date: 2026-09-29GUANGDONG ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY LAB (GUANGZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311313541.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-11
Publication Date
2026-09-29
Estimated Expiration
2043-10-11

AI Technical Summary

Technical Problem

[0005]1、通过门诊或网上留言的方式进行收集速度较慢,很难收集到大量、不同患者画像群体的对话数据

Benefits of technology

[0056]在本发明实施例中提出的方法能够大量生成不同抑郁症人群,不同背景画像下的医生患者对话数据实例,有助于相关心理异常类疾病的标准化量表辅助诊断方法的研究和实施。因而可对抑郁症等心理疾病在人群的大规模筛查、门诊诊断效率提升产生十分重要的意义。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117912667B_ABST
    Figure CN117912667B_ABST
Patent Text Reader

Abstract

The application discloses a mental health screening and dialogue data set generation method, comprising the following steps: S1, constructing a doctor professional large language model based on an existing open source large language model; S2, using the doctor professional large language model to simulate a doctor role to initiate a dialogue with a patient simulating a real crowd portrait, recording the dialogue content, and obtaining a mental abnormality disease scale auxiliary diagnosis dialogue data set. By using the doctor professional large language model to simulate the doctor role to initiate the dialogue with the patient simulating the real crowd portrait, the dialogue content is recorded, so that a large number of doctor-patient dialogue data instances of different depression crowds and different background portraits can be generated, which is helpful for the research and implementation of a standardized scale auxiliary diagnosis method of related mental abnormality diseases. Thus, the mental diseases such as depression can be screened in a large scale in a crowd, and the efficiency of outpatient diagnosis is improved, which is of great significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to smart healthcare technology, specifically to a method for mental health screening and generating dialogue datasets. Background Technology

[0002] In modern society, depression, as a common mental illness, places a significant burden on patients and society. Accurate diagnosis of depression is crucial for treatment and recovery. However, traditional diagnostic methods for depression primarily rely on the judgment of clinicians, some biochemical tests, and screening results from auxiliary scales. In areas with limited medical resources, this approach hinders rapid screening and intervention for mental illnesses, potentially leading to more serious consequences.

[0003] To address this issue, utilizing large language models and digital human technology for automated scale-assisted diagnosis of mental illnesses such as depression is of great significance. By analyzing the diagnostic dialogue data of patients during clinical visits, key semantic features and emotional indicators can be extracted, providing an objective, accurate, and efficient method for the rapid and assisted diagnosis and screening of depression. However, existing depression diagnostic dialogue datasets are relatively limited, with small data samples and a lack of population diversity. This restricts the training and optimization of depression diagnostic algorithms.

[0004] Currently, there are relatively few publicly available datasets of dialogues used for depression diagnosis. The main problems with collecting dialogue data and using existing methods for auxiliary diagnosis of depression are as follows:

[0005] 1. Collecting data through outpatient clinics or online messages is slow and makes it difficult to gather a large amount of dialogue data from different patient profiles.

[0006] 2. Depression diagnosis methods based on machine learning and other computer big data technologies have a certain degree of uninterpretability due to the black box nature of the related algorithms, which is not conducive to their widespread use in the medical field, where explanations and accountability to patients are required.

[0007] 3. Traditional computer technologies based on state machines and knowledge graphs mainly rely on predefined rules and structured knowledge. Data generation requires a large amount of semantic information from the patient profile knowledge graph to be input in advance to expand the patient profile, which is still relatively inefficient.

[0008] 4. Traditional chatbot responses are limited by predefined rules and knowledge graphs, making it impossible to provide creative or personalized answers. Compared to large language models, they lack the ability to understand and generate natural language, thus failing to produce diverse and human-like responses. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for mental health screening and dialogue dataset generation, thereby enriching the dialogue dataset for depression diagnosis and providing valuable resources for related research and clinical practice.

[0010] To achieve the above objectives, the technical solution of the present invention is as follows:

[0011] A method for mental health screening and dialogue dataset generation includes:

[0012] S1. Construct a professional large language model for doctors based on existing open-source large language models;

[0013] S2. Use the aforementioned professional language model of doctors to simulate the role of doctors to initiate dialogue with patients who are simulated real-life population profiles, record the dialogue content, and obtain a dialogue dataset for auxiliary diagnosis of psychological abnormality disease scale.

[0014] Further, step S1 includes:

[0015] The basic architecture of the open-source large language model is a transformer, including a multi-head self-attention mechanism, relative position encoding, and a feedforward neural network. The multi-head self-attention mechanism focuses on parts of the input text while generating each word in the output. The multi-head self-attention mechanism captures contextual information by calculating the interactions and correlations between different positions in the input sequence. The relative position encoding captures the relative distances between different positions in the input sequence. The feedforward neural network performs non-linear transformations and feature mappings on the hidden states at each position.

[0016] The open-source large language model was tuned using a publicly available depression clinic dialogue dataset, and self-supervised training of the model was performed using patient-doctor dialogue-response data pairs from the dataset.

[0017] During the adjustment process, a loss function is set, and the gradient of the loss function with respect to the model parameters is calculated through the backpropagation algorithm. An optimizer is then used to update the model parameters so that the model can be gradually optimized.

[0018] LoRA was used to adjust the optimized open-source large language model to obtain the final simulated doctor professional large language model.

[0019] Furthermore, the multi-head self-attention mechanism is represented by the calculation process of the interaction and correlation between different positions in the input sequence as follows:

[0020] Attention(Q, K, V) = softmax((QK^T + DK^T) / sqrt(d_k))V

[0021] Where Q, K, and V represent the query, key, and value of the input sequence, respectively, and d_k represents the dimension of each attention head.

[0022] Furthermore, the calculation of the relative position encoding used to capture the relative distance between different positions in the input sequence is expressed by the following formula:

[0023] PE_{(pos, 2i)} = sin(pos / 10000^(2i / d_model))

[0024] PE_{(pos, 2i+1)} = cos(pos / 10000^(2i / d_model))

[0025] Where PE_{(pos, 2i)} and PE_{(pos, 2i+1)} represent the relative position encoding of position pos and dimensions 2i and 2i+1, respectively; d_model represents the dimension of the model.

[0026] Furthermore, the computational process of the feedforward neural network for performing nonlinear transformation and feature mapping on the hidden state at each position is represented as follows:

[0027] FN(x) = max(0, xW_1 + b_1)W_2 + b_2

[0028] Where x represents the hidden state of the input, and W_1, W_2, b_1, and b_2 are the parameters of the model.

[0029] Furthermore, the basic principle of LoRA is to freeze the pre-trained model weight parameters. While freezing the original model parameters, additional network layers are added to the model, and only the parameters of these newly added network layers are trained. The loss function at this point is expressed as:

[0030] Loss = Loss_task_specific_head + λ * Loss_pretrained_model

[0031] Where Loss_task_specific_head is the loss function for the task-specific head network, and Loss_pretrained_model is the loss function for the pre-trained model, which uses a predefined loss function:

[0032] L(θ) = -∑[y * log(p) + (1-y) * log(1-p)]

[0033] Where θ is the model parameter, y is the label, p is the probability predicted by the model; λ is the weight coefficient of the two loss functions, used to balance the importance of the two; through backpropagation and parameter updates, the model will continuously optimize its weights and biases to better adapt to the specific task.

[0034] Further, step S2 includes:

[0035] S2.1. By setting patient profiles using the prompt words of the doctor's professional language model, diverse dialogues can be generated to simulate people with depression from different background profiles.

[0036] S2.2 At the start of the simulated dialogue in the depression clinic, ask the user a fixed question to collect basic information;

[0037] S2.3. Extract several questions from the preset scale question bank according to the established strategy to construct the main question flow of the depression clinic dialogue; automatically insert clinical diagnostic scale questions into several rounds of simulated dialogue between doctors and patients, and output the questions in the main question flow in sequence with the generated content of the doctor's professional language model to the user. When there are no questions left in the main question flow and the corresponding subsequent steps are completed, the clinic dialogue ends.

[0038] S2.4 After receiving the user's response, record the dialogue data.

[0039] Furthermore, the method also includes:

[0040] S3. Quantitatively evaluate the diagnostic dialogue dataset for the psychological abnormality disease scale, including:

[0041] S3.1. Call the doctor's professional language model to score, extract the semantics of key information in the patient's answer as the basis for scale scoring, and record the corresponding scoring results as the basis for subsequent scale-assisted diagnosis.

[0042] S3.2 If the patient's answer is vague and the validity of the answer cannot be accurately assessed, the doctor's professional language model is used to re-initiate the dialogue and ask questions, and the patient is reminded that the previous answer was not understood by them.

[0043] S3.3 If there are still questions remaining in the main question flow in step S2.3, use the doctor's professional large language model to generate subsequent questions based on the current question and the user's response to the question, and repeat the above steps S3.1-S3.2 until all scale questions are asked, and output the depression scale auxiliary diagnosis results with quantitative assessment.

[0044] Furthermore, the method also includes:

[0045] S4. Filter the diagnostic dialogue dataset for the psychological abnormality disease scale, including...

[0046] S4.1. For questions where some patients simulate answering questions that repeat what the doctor said, use n-gram consistency filtering;

[0047] S4.2 Use sentiment analysis to filter dialogue data from simulated conversations of patients with depression that do not show obvious negative emotions;

[0048] S4.3 uses the Dirichlet assignment topic modeling technique to filter depression topics in the dialogue dataset.

[0049] Furthermore, the method also includes:

[0050] S5. Screening for mental health disorders through digital human deployment applications, including:

[0051] S5.1. Using a professional medical language model, doctors can act as doctors and interact with patients in the real world via voice. The text generated by the professional medical language model is processed in a computer that processes data using text-to-speech technology.

[0052] S5.2. Call the Sadtalker digital human technology framework to generate a digital human image, and combine the generated digital human image with voice to interact with the real-world patient through display and sound devices.

[0053] S5.3. Record the patient's voice through a sound recording device, and use speech-to-text technology to communicate with the doctor's professional large language model running in the computer background. Repeat steps S5.1-5.3 until all depression-related scale questions have been asked and accurate answers from the patient have been received.

[0054] S5.4. Call the auxiliary model to evaluate the patient's answers to the scale questions, and finally output a report on the auxiliary diagnosis of depression scale, and give the patient a reminder of the risk of mental abnormality based on the results of quantitative scoring.

[0055] Compared with the prior art, the advantages of this invention are as follows:

[0056] The method proposed in this invention can generate a large number of doctor-patient dialogue data instances under different background profiles and different depression groups, which is helpful for the research and implementation of standardized scale-assisted diagnostic methods for related mental disorders. Therefore, it can have a very important impact on large-scale screening of mental illnesses such as depression and improving the efficiency of outpatient diagnosis. Attached Figure Description

[0057] Figure 1A flowchart illustrating the main steps of the mental health screening and dialogue dataset generation method provided in this embodiment of the invention;

[0058] Figure 2 This is a schematic diagram illustrating the computational process of a multi-head self-attention mechanism.

[0059] Figure 3 The flowchart illustrates the steps of a method for generating a mental health screening and dialogue dataset according to a preferred embodiment of the present invention. Detailed Implementation

[0060] Example:

[0061] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0062] See Figure 1 As shown, the method for generating mental health screening and dialogue datasets provided in this embodiment mainly includes the following steps:

[0063] S1. Construct a professional large language model for doctors based on existing open-source large language models;

[0064] S2. Use the aforementioned professional language model of doctors to simulate the role of doctors to initiate dialogue with patients who are simulated real-life population profiles, record the dialogue content, and obtain a dialogue dataset for auxiliary diagnosis of psychological abnormality disease scale.

[0065] Thus, by using a professional medical language model to simulate a doctor's role in initiating dialogues with patients in simulated real-world population profiles and recording the dialogue content, a large number of doctor-patient dialogue data instances with different individuals suffering from depression and with varying background profiles can be generated. This is helpful for the research and implementation of standardized scale-based diagnostic methods for related mental disorders. Therefore, it can have significant implications for large-scale screening of depression and other mental illnesses and for improving the efficiency of outpatient diagnosis.

[0066] In one specific embodiment, step S1 includes:

[0067] This open-source large language model is based on the Transformer architecture, which is the core of the model. It relies on a multi-layered stack of self-attention mechanisms, enabling the model to weigh the importance of different words in a sentence when processing each word. This self-attention mechanism helps the model effectively capture long-range dependencies in the text. It includes multi-head self-attention mechanisms, relative position encoding, and feedforward neural networks.

[0068] Among them, the multi-head self-attention mechanism focuses on parts of the input text while generating each word in the output; such as Figure 2As shown, the multi-head self-attention mechanism captures contextual information by calculating the interactions and correlations between different positions in the input sequence. The calculation process can be represented as follows:

[0069] Attention(Q, K, V) = softmax((QK^T + DK^T) / sqrt(d_k))V

[0070] Where Q, K, and V represent the query, key, and value of the input sequence, respectively, and d_k represents the dimension of each attention head.

[0071] The Transformer uses relative positional encoding to capture the relative distances between different positions in the input sequence. The calculation of relative positional encoding can be expressed by the following formula:

[0072] PE_{(pos, 2i)} = sin(pos / 10000^(2i / d_model))

[0073] PE_{(pos, 2i+1)} = cos(pos / 10000^(2i / d_model))

[0074] Where PE_{(pos, 2i)} and PE_{(pos, 2i+1)} represent the relative position encoding of position pos and dimensions 2i and 2i+1, respectively. d_model represents the dimension of the model.

[0075] A feedforward neural network is used to perform nonlinear transformations and feature mappings on the hidden state at each location. The computation process can be represented as follows:

[0076] FFN(x) = max(0, xW_1 + b_1)W_2 + b_2

[0077] Where x represents the hidden state of the input, and W_1, W_2, b_1, and b_2 are the parameters of the model.

[0078] These are the core components of the Transformer architecture. Through relative position encoding, multi-head self-attention mechanism, and feedforward neural network, Transformer can capture long-range dependencies in sequence data and generate accurate output in natural language processing tasks.

[0079] The open-source large language model was fine-tuned using publicly available outpatient dialogue datasets for depression, such as the D4 or CMDC Chinese dataset. The model was then trained in a self-supervised manner using patient-doctor dialogue-response data pairs from the dataset.

[0080] During fine-tuning, the first step is to define a loss function suitable for the specific task. The loss function measures the model's performance on a particular task; the target is to minimize the loss function. In general network fine-tuning, the commonly used loss function is cross-entropy loss.

[0081] L(θ) = -∑[y * log(p)]

[0082] Where θ is the network parameter, y is the label (usually a one-hot vector), and p is the probability predicted by the model.

[0083] This loss function measures the model's performance on classification problems by calculating the cross-entropy between the labels and the model's predictions. By minimizing the cross-entropy loss, the parameters can be gradually adjusted to make the model's predictions closer to the actual labels.

[0084] The core of fine-tuning is to calculate the gradient of the loss function with respect to the model parameters using the backpropagation algorithm, and then use an optimizer to update the model parameters, gradually converging them to better performance. Specifically, for each parameter θ, the parameter value is updated based on the gradient ∇L(θ) of the loss function:

[0085] θ_new = θ_old - learning_rate * ∇L(θ_old)

[0086] Here, θ_old is the current parameter value, and learning_rate is the learning rate, used to control the step size or speed of parameter updates. By calculating the gradient ∇L(θ_old) of the loss function, we can obtain the derivative information of the parameter θ_old. Then, we multiply it by the learning rate learning_rate, and subtract the result from the current parameter value θ_old to obtain the new parameter value θ_new.

[0087] This update process typically occurs in each training iteration and is repeated multiple times until a predetermined number of iterations is reached or a convergence condition is met. By continuously adjusting the parameters in the direction of gradient descent, the model gradually optimizes, causing the loss function to decrease gradually, thereby improving the model's performance and accuracy.

[0088] LoRA (Low-Rank Adaptation of Large Language Models) is used to fine-tune an optimized open-source large language model. LoRA is a parameter-efficient fine-tuning method. The basic principle of LoRA is to freeze the pre-trained model weights and parameters. While freezing the original model parameters, additional network layers are added to the model, and only the parameters of these newly added layers are trained. Because the number of these new parameters is relatively small, this significantly reduces the cost of fine-tuning while achieving similar results to full model fine-tuning. Therefore, the loss function can be expressed as:

[0089] Loss = Loss_task_specific_head(predicted value, label) + λ * Loss_pretrained_model(predicted value, output of pretrained model)

[0090] Where Loss_task_specific_head is the loss function for the task-specific head network, and Loss_pretrained_model is the loss function for the pre-trained model, which can be a predefined loss function.

[0091] L(θ) = -∑[y * log(p) + (1-y) * log(1-p)]

[0092] Here, θ represents the model parameters, y represents the label (ground truth), and p represents the probability predicted by the model. λ is the weight coefficient of the two loss functions, used to balance their importance. Through backpropagation and parameter updates, the model continuously optimizes its weights and biases to better adapt to the specific task.

[0093] In this way, after appropriate fine-tuning, the final professional language model for doctors is obtained, which enables them to complete the required tasks more accurately.

[0094] Specifically, step S2 above includes:

[0095] S2.1. By using the prompt words of the doctor's professional large language model to set up patient profiles, such as basic demographic characteristics such as age, gender, education level, occupation, income, marital status and childbearing, as well as additional supplementary attributes such as mood, mental state, and life changes, it can effectively simulate the generation of diverse dialogues for people with depression with different background profiles.

[0096] S2.2 Basic Information Collection: At the start of the simulated dialogue in the depression clinic, ask the user a fixed question to collect basic information, including age, occupation, etc. Extract and record this information for use in subsequent steps.

[0097] S2.3. Main Question Flow Construction: Following a predetermined strategy, several questions are extracted from a pre-set question bank, such as the PHQ-9 questions for depression, to construct the main question flow for a depression clinic dialogue. Clinical diagnostic scale questions are automatically inserted into several rounds of simulated dialogue between the doctor and patient. The questions in the main question flow, combined with the content generated by the large language model, are sequentially output to the user. The clinic dialogue ends when there are no remaining questions in the main question flow and all subsequent steps have been completed.

[0098] S2.4 After receiving the user's response, record the dialogue data.

[0099] As a preferred embodiment of this invention, such as Figure 3 As shown, the above-mentioned method for mental health screening and dialogue dataset generation further includes: S3, quantitatively evaluating the dialogue dataset for auxiliary diagnosis of mental abnormality disease scale, including:

[0100] S3.1. Call the professional doctor's large language model to score, extract the semantics of key information in the patient's answer as the basis for scale scoring, and record the corresponding scoring results as the basis for subsequent scale-assisted diagnosis.

[0101] S3.2 If the patient's answer is vague or its effectiveness cannot be accurately assessed, the large language model simulating a professional doctor initiates a new dialogue, prompting the patient that their previous answer was not comprehensible. This method effectively reduces invalid and low-quality answers, further filtering the data. Similarly, this method can be used to assess the effectiveness and rationality of the simulated dialogue, recording the model's comprehensibility of each other's generated content, whether the scale questions were answered correctly, and whether the doctor's answers were insightful. See the table below for details:

[0102] Reasonableness of the patient's response % 94.0% 95.7% 95.2% Like a doctor % 95.6% 97.5% 96.4% Dialogue containing caring words 82.6% 86.9% 89.1% Total number of dialogues (rounds) 314 653 545

[0103] S3.3, Subsequent Question Generation: If there are still questions remaining in the main question flow from step S2.3, the doctor's professional large language model is used to generate subsequent questions based on the current system questions and user responses. This process is repeated until all scale questions have been asked. Finally, a depression scale-based auxiliary diagnostic result with quantitative assessment is output, as shown in the table below:

[0104] Number of people 24 48 40 Percentage % 21.4 42.8 35.7

[0105] As a preferred embodiment of this invention, such as Figure 3As shown, the above-mentioned method for mental health screening and dialogue dataset generation further includes: S4, filtering the dialogue dataset for auxiliary diagnosis of mental abnormality disease scale, including:

[0106] S4.1, n-gram consistency filtering

[0107] For the issue of some patients repeating what the doctor said in their simulated answers, n-gram consistency filtering was used on this low-quality dialogue data. n-gram is a commonly used text representation method to capture continuous relationships between words or characters. n-gram can be used to compare the similarity between two text segments. First, each text segment is divided into n consecutive word or character sequences, and then their frequencies are calculated. By comparing the number and frequency of common n-grams between two text segments, their similarity can be measured. The results of filtering out low-quality data using this method are as follows:

[0108] After N-gram consistency filtering (round) 1485 Percentage % 98.2%

[0109] S4.2 Sentiment Analysis

[0110] This technology was used to filter dialogue data from simulated conversations involving patients with depression that did not show obvious negative emotions.

[0111] Sentiment analysis filtering (round) 1498 Percentage % 99.1%

[0112] S4.3, Latent Dirichlet Allocation (LDA) Topic Modeling Technique

[0113] This technique can be used to further filter depression topics within dialogue datasets, mapping text data (such as dialogues) to a latent topic space. For depression-related dialogues, LDA topic modeling can help determine what topics are discussed in these dialogues.

[0114] The process of using LDA for topic modeling is as follows: First, each dialogue is treated as a document, consisting of multiple words, which can be vocabulary or phrases from the dialogue. Next, the LDA model uses unsupervised learning methods to discover latent topics. LDA assumes that each word in the document is generated from a certain topic, and that a dialogue may contain multiple topics simultaneously.

[0115] During training, the LDA model automatically learns the distribution of topics and the word distribution for each topic in each dialogue. A trained LDA model can be used to identify topics in dialogues to evaluate their effectiveness and quality related to depression. The effectiveness and quality of the dialogue can be assessed by observing the topic distribution and the prevalence of depression-related topics.

[0116] LDA topic modeling and filtering (round) 1468 Percentage % 97.1%

[0117] The final dataset result after filtering is as follows:

[0118] Total number of dialogue rounds 1435 1402 1468 Average number of rounds 26.3 25.8 26.8 Total number of people in the conversation 224 112 112 On average, the number of topics generated per conversation related to each scale question 2.8 - - Average conversation length 112 89 135

[0119] MotiVAte mental health 4000 13.46 - ESConv Emotional support 1053 - 29.8 DAIC-WOZ Pain Analysis 189 - - CMDC Depression counseling 852 12 - D4 Depression counseling 1339 21.6 30.9 This method Depression diagnosis 1435 26.3 26.8

[0120] As a preferred embodiment of this invention, such as Figure 3 As shown, the above-mentioned method for mental health screening and dialogue dataset generation also includes: S5, screening for mental disorders through digital human deployment applications, including:

[0121] S5.1. Using a large language model of doctors to play the role of doctors and interact with patients in the real world through voice, the text generated by the large language model of doctors is first processed in a computer that processes data through text-to-speech technology.

[0122] S5.2. Call the Sadtalker digital human technology framework to generate a digital human image, and combine the generated digital human image with voice to interact with the real-world patient through display and sound devices.

[0123] S5.3. Record the patient's voice through a sound recording device, and use speech-to-text technology to communicate with a large doctor model running in the computer background. Repeat steps S5.1-5.3 until all depression-related scale questions have been asked and accurate answers from the patient have been received.

[0124] S5.4. Call the auxiliary model to evaluate the patient's answers to the scale questions, and finally output a report on the auxiliary diagnosis of depression scale. Based on the quantitative score results, give patients using this device a reminder of the risk of depression, and assist doctors in screening.

[0125] In summary, the implementation of this method can overcome the limitations of traditional methods for diagnosing depression, providing doctors and patients with a reliable and intelligent diagnostic screening tool for depression, and helping more patients obtain accurate diagnoses and effective treatments as early as possible.

[0126] The above embodiments are merely illustrative of the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made based on the essence of the content of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for mental health screening and dialogue dataset generation, characterized in that, include: S1. Construct a professional large language model for doctors based on existing open-source large language models; S2. Use the doctor's professional language model to simulate the doctor's role and initiate a dialogue with a patient who is a simulated real-person profile. Record the dialogue content to obtain a dialogue dataset for auxiliary diagnosis of psychological abnormality disease scale. Step S1 includes: The basic architecture of the open-source large language model is a transformer, including a multi-head self-attention mechanism, relative position encoding, and a feedforward neural network. The multi-head self-attention mechanism focuses on parts of the input text while generating each word in the output. The multi-head self-attention mechanism captures contextual information by calculating the interactions and correlations between different positions in the input sequence. The relative position encoding captures the relative distances between different positions in the input sequence. The feedforward neural network performs non-linear transformations and feature mappings on the hidden states at each position. The open-source large language model was tuned using a publicly available depression clinic dialogue dataset, and self-supervised training of the model was performed using patient-doctor dialogue-response data pairs from the dataset. During the adjustment process, a loss function is set, and the gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm. An optimizer is then used to update the model parameters so that the model can be gradually optimized. LoRA was used to adjust the optimized open-source large language model to obtain the final simulated doctor professional large language model; Step S2 includes: S2.

1. Set up patient profiles using the prompt words of the doctor's professional language model to simulate diverse dialogues among people with depression from different background profiles. S2.2 At the start of the simulated dialogue in the depression clinic, ask the user a fixed question to collect basic information; S2.

3. Extract several questions from the preset question bank of the scale according to the established strategy to construct the main question flow of the depression clinic dialogue; automatically insert clinical diagnostic scale questions into several rounds of simulated dialogue between doctors and patients, and output the questions in the main question flow in sequence with the generated content of the doctor's professional language model to the user. When there are no questions left in the main question flow and the corresponding subsequent steps are completed, the clinic dialogue ends. S2.4 After receiving the user's response, record the dialogue data.

2. The method for mental health screening and dialogue dataset generation as described in claim 1, characterized in that, The relative position encoding used to capture the relative distance between different positions in the input sequence is calculated using the following formula: PE_{(pos, 2i)} = sin(pos / 10000^(2i / d_model)) PE_{(pos, 2i+1)} = cos(pos / 10000^(2i / d_model)) Where PE_{(pos, 2i)} and PE_{(pos, 2i+1)} represent the relative position encoding of position pos and dimensions 2i and 2i+1, respectively; d_model represents the dimension of the model.

3. The method for mental health screening and dialogue dataset generation as described in claim 1, characterized in that, The calculation process of the feedforward neural network for nonlinear transformation and feature mapping of the hidden state at each position is represented as follows: FN(x) = max(0, xW_1 + b_1)W_2 + b_2 Where x represents the hidden state of the input, and W_1, W_2, b_1, and b_2 are the parameters of the model.

4. The method for mental health screening and dialogue dataset generation as described in claim 1, characterized in that, The basic principle of LoRA is to freeze the pre-trained model weights. With the original model parameters frozen, additional network layers are added to the model, and only the parameters of these newly added network layers are trained. The loss function in this case is expressed as: Loss = Loss_task_specific_head + λ * Loss_pretrained_model Where Loss_task_specific_head is the loss function for the task-specific head network, and Loss_pretrained_model is the loss function for the pre-trained model, which uses a predefined loss function: L(θ) = -∑[y * log(p) + (1-y) * log(1-p)] Where θ is the model parameter, y is the label, p is the probability predicted by the model, and λ is the weight coefficient of the two loss functions, used to balance their importance.

5. The method for mental health screening and dialogue dataset generation as described in claim 1, characterized in that, The method further includes: S3. Quantitatively evaluate the diagnostic dialogue dataset for the psychological abnormality disease scale, including: S3.

1. Call the doctor's professional language model to score, extract the semantics of key information in the patient's answer as the basis for scale scoring, and record the corresponding scoring results as the basis for subsequent scale-assisted diagnosis. S3.2 If the patient's answer is vague and the validity of the answer cannot be accurately assessed, the doctor's professional language model is used to re-initiate the dialogue and ask questions, and the patient is reminded that the previous answer was not understood by them. S3.3 If there are still questions remaining in the main question flow in step S2.3, use the doctor's professional large language model to generate subsequent questions based on the current question and the user's response to the question, and repeat the above steps S3.1-S3.2 until all scale questions are asked, and output the depression scale auxiliary diagnosis results with quantitative assessment.

6. The method for mental health screening and dialogue dataset generation as described in claim 5, characterized in that, The method further includes: S4. Filter the diagnostic dialogue dataset for the psychological abnormality disease scale, including... S4.

1. For questions where some patients simulate answering questions that repeat what the doctor said, use n-gram consistency filtering; S4.2 Use sentiment analysis to filter dialogue data from simulated conversations of patients with depression that do not show obvious negative emotions; S4.

3. Use the Dirichlet assignment topic modeling technique to filter depression topics in the dialogue dataset.

Citation Information

Patent Citations

  • Depression diagnosis dialogue data set generation method, electronic equipment and storage medium

    CN114334163A

  • Voice communication sentiment analysis intelligent session system, method, device and medium

    CN116778921A