Model training method and bipolar disorder personalized treatment dialogue generation method and device

By collecting and analyzing the speech data of patients with bipolar disorder, and using multimodal language models and deep learning technology, personalized treatment dialogues are generated, which solves the problems of accuracy and matching in diagnosis and treatment in existing technologies and improves the treatment effect of bipolar disorder.

CN122369501APending Publication Date: 2026-07-10THE SECOND HOSPITAL AFFILIATED TO WENZHOU MEDICAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE SECOND HOSPITAL AFFILIATED TO WENZHOU MEDICAL COLLEGE
Filing Date
2026-03-12
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In the current technology, the diagnosis and treatment of bipolar disorder are not accurate enough, mainly relying on the doctor's subjective judgment and the treatment dialogue is poorly matched, resulting in poor treatment effect.

Method used

By collecting sample speech data labeled with bipolar disorder classification tags, semantic, mood, and tone features are extracted using a multimodal language model. Combined with a Transformer encoder and a gated recurrent unit network module, a target classification model is trained to generate personalized therapeutic dialogues.

Benefits of technology

It improves the accuracy of bipolar disorder classification, reduces reliance on doctors' subjective diagnoses, and enhances the matching of treatment dialogues and treatment outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369501A_ABST
    Figure CN122369501A_ABST
Patent Text Reader

Abstract

The application provides a model training method and a bipolar disorder personalized treatment dialogue generation method and device, belonging to the technical field of deep learning, aiming at the problems that the current diagnosis accuracy of bipolar mood disorder is not high and the matching of treatment dialogue with patients is poor, the model training can be performed on a multi-modal language model according to a voice signal with significant characteristics of bipolar mood disorder, and a target classification model for bipolar mood disorder classification is obtained. Since the voice signal contains characteristics such as semantics, tone, pronunciation rhythm and other characteristics directly related to bipolar mood disorder expressed by the patient, the target classification model can be used for accurate classification of bipolar mood disorder, and when the target classification model is used for classification of bipolar mood disorder, the excessive dependence on the subjective diagnosis of doctors can be avoided, and the accuracy of the classification prediction of bipolar mood disorder is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to a model training method, a method and apparatus for generating personalized treatment dialogues for bipolar disorder. Background Technology

[0002] Bipolar disorder (BD), also known as bipolar disorder or formerly known as manic-depressive illness, is a common mental health condition belonging to the category of mood disorders. Its core characteristic is drastic fluctuations in mood between two extremes: from extremely elevated, hyperactive manic / hypomanic phases to extremely depressed, hopeless phases, sometimes with mixed episodes (simultaneous manic and depressive symptoms). Globally, approximately 1 in 200 people are affected (about 37 million). In China, the lifetime prevalence among adults is about 0.6-0.7%, with a near-equal male-to-female ratio. First onset typically occurs between the ages of 15 and 30. It is a chronic, relapsing disease; without proper treatment, the relapse rate is very high, and the suicide risk is significantly higher than in the general population (up to 10-20 times). Bipolar disorder has become a serious health problem. Currently, the diagnosis and assessment of bipolar disorder mainly rely on the subjective judgment of psychologists or psychiatrists, which often leads to errors in the diagnosis of bipolar disorder and a low accuracy rate. At the same time, the manifestations of bipolar disorder often vary greatly among different patients, but a uniform dialogue template is often used for patients with bipolar disorder. The treatment dialogue is poorly matched with the patient, which leads to poor treatment results for bipolar disorder. Summary of the Invention

[0003] In view of the above problems, embodiments of this application provide a model training method, a method for generating personalized treatment dialogues for bipolar disorder, an apparatus, an electronic device, and a readable storage medium to overcome or at least partially solve the above problems.

[0004] In a first aspect, embodiments of this application provide a model training method, the method comprising: A sample speech dataset of patients was collected, which includes multiple sample speech samples labeled with bipolar disorder classification tags; Input the sample speech into the multimodal language model to obtain the sample classification prediction results of the multimodal language model output sample speech; Based on the sample classification prediction results and bipolar disorder classification labels, the multi-class cross-entropy loss value of the multimodal language model is determined. Based on the multi-class cross-entropy loss value, the parameters of the multimodal language model are iteratively trained and optimized to obtain the target classification model; the target classification model is used to determine the bidirectional affective disorder category of the target patient based on the target patient's target speech data.

[0005] Optionally, the sample speech is input into a multimodal language model to obtain the sample classification prediction results of the multimodal language model outputting the sample speech, including: Semantic features, mood features, and tone features are extracted from the sample speech. Semantic features represent the contextual information of the semantics of the sample speech, mood features represent the emotional expression of the sample speech, and tone features represent the prosodic parameters of the sample speech. Semantic features, mood features, and tone features are input into a multimodal language model to obtain the sample classification prediction results of the multimodal language model's output sample speech.

[0006] Optionally, the multimodal language model includes a Transformer encoder and a gated recurrent unit network module. Semantic features, mood features, and tone features are input into the multimodal language model to obtain the sample classification prediction results of the output sample speech of the multimodal language model, including: Using the Transformer encoder, semantic encoding vectors of semantic features, tone encoding vectors of mood features, and tone encoding vectors of tone features are extracted respectively. Using a gated recurrent unit network module, based on semantic encoding vector, mood encoding vector and tone encoding vector, the sample classification prediction result of sample speech is determined.

[0007] Optionally, semantic features, mood features, and tone features of the sample speech are extracted, including: Data augmentation is performed on the sample speech to obtain augmented sample speech; data augmentation includes adding background noise or speed adjustment. Extract samples to enhance the semantic features, mood features, and tone features of speech.

[0008] Optionally, a patient speech dataset is collected, including: Voice samples from real patients were collected from clinical records or mobile devices and categorized by experts. The categorization labels included at least one of manic, depressive, and normal states.

[0009] Secondly, embodiments of this application provide a method for generating personalized treatment dialogues for bipolar disorder, the method comprising: Collect target speech data of the target patient at the time of onset of illness; The target speech data is input into the target classification model to obtain the target classification prediction result output by the target classification model; wherein, the target classification model is obtained based on the model training method described in any of the above items; Based on the target classification prediction results, personalized treatment dialogue information is generated that matches the target patient.

[0010] Optionally, the target speech data is input into the target classification model to obtain the target classification prediction result output by the target classification model, including: Extract the target semantic features, target mood features, and target tone features from the target speech data; Using a target classification model, the target classification prediction results are determined based on the target semantic features, target tone features, and target pitch features.

[0011] Thirdly, embodiments of this application provide a model training apparatus, the apparatus comprising: The first acquisition module is used to acquire sample speech datasets from sample patients. The sample speech datasets include multiple sample speech samples labeled with bipolar disorder classification tags. The first input / output module is used to input sample speech into the multimodal language model and obtain the sample classification prediction result of the multimodal language model output sample speech; The first determination module is used to determine the multi-class cross-entropy loss value of the multimodal language model based on the sample classification prediction results and the bipolar disorder classification label; The training module is used to iteratively train and optimize the parameters of the multimodal language model based on the multi-class cross-entropy loss value to obtain the target classification model; the target classification model is used to determine the bidirectional affective disorder category of the target patient based on the target patient's target speech data.

[0012] Fourthly, embodiments of this application provide a personalized treatment dialogue generation device for bipolar disorder, the device comprising: The second acquisition module is used to acquire the target patient's speech data at the time of the onset of the disease. The second input / output module is used to input the target speech data into the target classification model and obtain the target classification prediction result output by the target classification model. The generation module is used to generate personalized treatment dialogue information that matches the target patient based on the target classification prediction results.

[0013] Fifthly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the model training method as described above, or the personalized treatment dialogue generation method for bipolar disorder.

[0014] Sixthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the model training method described above, or the personalized treatment dialogue generation method for bipolar disorder.

[0015] The specific beneficial effects are as follows: This application embodiment collects sample speech datasets from sample patients, including multiple sample speech samples labeled with bipolar disorder classification tags. The sample speech is input into a multimodal language model to obtain the sample classification prediction results of the multimodal language model's output sample speech. Based on the sample classification prediction results and the bipolar disorder classification tags, the multi-class cross-entropy loss value of the multimodal language model is determined. Based on the multi-class cross-entropy loss value, the parameters of the multimodal language model are iteratively trained and optimized to obtain a target classification model. The target classification model is used to determine the bipolar disorder category of the target patient based on the target patient's target speech data. The multimodal language model can be trained based on speech signals with significant characteristics of bipolar disorder to obtain a target classification model that can be used for bipolar disorder classification. Since the speech signal contains features directly related to bipolar disorder, such as the patient's semantics, tone, and rhythm, the target classification model can be used for accurate classification of bipolar disorder. When using the target classification model to classify bipolar disorder, it can avoid over-reliance on the doctor's subjective diagnosis and effectively improve the accuracy of bipolar disorder classification prediction. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic flowchart of a model training method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the model loss value decrease curve provided in the embodiments of this application; Figure 3 A flowchart illustrating a personalized treatment dialogue generation method for bipolar disorder provided in this application embodiment; Figure 4 This is a logic block diagram of a model training device provided in an embodiment of this application; Figure 5 This is a logic block diagram of a personalized treatment dialogue generation device for bipolar disorder provided in an embodiment of this application; Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0019] like Figure 1 The model training method provided in this application embodiment includes: Step S101: Collect sample speech datasets from sample patients. The sample speech datasets include multiple sample speech samples labeled with bipolar disorder classification tags. Step S102: Input the sample speech into the multimodal language model to obtain the sample classification prediction result of the multimodal language model output sample speech; Step S103: Based on the sample classification prediction results and the bipolar disorder classification label, determine the multi-class cross-entropy loss value of the multimodal language model; Step S104: Based on the multi-class cross-entropy loss value, iteratively train and optimize the parameters of the multimodal language model to obtain the target classification model; the target classification model is used to determine the bidirectional affective disorder category of the target patient based on the target patient's target speech data.

[0020] In embodiments of this application, speech data can be collected from multiple "sample patients" (i.e., a representative patient group used for training) to form a dataset. This dataset is not random but structured, including multiple speech samples, each labeled to indicate the corresponding bipolar disorder category (e.g., normal, manic, depressive, etc.). In practice, this may involve recording patients' daily conversations or standard interviews (e.g., Hamilton Depression Rating Scale interviews) using recording devices (e.g., mobile apps), followed by manual labeling by psychiatrists or experts. The dataset typically requires hundreds to thousands of samples to ensure the model's generalization ability. Labeled sample speech with bipolar disorder classification tags means that each speech sample is assigned a "label," which is a category identifier. For example, the label might be "BD-I - Manic Episode," "BD-II - Depressive Episode," or "Normal." The labeling process is usually performed by experts to ensure accuracy.

[0021] In embodiments of this application, collected speech samples can be input into a pre-existing multimodal language model. After processing, the model outputs a predictive classification result. For example, inputting a patient's speech might result in the model outputting "BD-I manic phase, probability 0.85". Multimodal models can handle various features of speech, such as audio waveforms, semantic content (through speech-to-text conversion), and even combine other modalities (such as text or images, if expanded). Multimodal language models can handle multiple data types, such as speech, text, and images. Typical examples include Transformer-based models (such as CLIP or variants of AudioBERT), which not only understand text but also extract speech features (e.g., MFCC acoustic features, speech rate, pitch). In BD classification, it can be used to detect rapid speech during manic phases or monotonous intonation during depressive phases.

[0022] In the embodiments of this application, the multi-class cross-entropy loss value can be used to evaluate the difference between the prediction results of the comparison model (sample classification prediction results) and the true labels (bipolar disorder classification labels). If the prediction is accurate, the loss is small; otherwise, it is large, which can be used to guide model improvement. The multi-class cross-entropy loss value is a loss function used for multi-class classification problems, expressed as follows: (Equation 1) In Equation 1 above, C is the number of categories. It is the one-hot encoding of the real label (e.g., [0,1,0] represents the second class). This refers to the model's predicted probability. Cross-entropy measures the difference between two probability distributions, and the multi-class version is suitable for scenarios with more than two categories (such as different subtypes of BD). In the context of BD, it encourages the model to output a probability distribution consistent with the true diagnosis.

[0023] In the embodiments of this application, the model parameters can be adjusted using the multi-class cross-entropy loss value through backpropagation, and the model can be updated by minimizing the loss function through multiple epochs. An optimizer (such as Adam) uses the gradient descent algorithm to adjust the parameters until the model performance stabilizes, resulting in the final target classification model. Learnable variables within the model include weight matrices and biases.

[0024] In the embodiments of this application, a sample speech dataset of sample patients is collected. This dataset includes multiple sample speech samples labeled with bipolar disorder classification tags. The sample speech is input into a multimodal language model to obtain the sample classification prediction result of the multimodal language model outputting the sample speech. Based on the sample classification prediction result and the bipolar disorder classification tags, the multi-class cross-entropy loss value of the multimodal language model is determined. Based on the multi-class cross-entropy loss value, the parameters of the multimodal language model are iteratively trained and optimized to obtain a target classification model. The target classification model is used to determine the bipolar disorder category of the target patient based on the target patient's target speech data. The multimodal language model can be trained based on speech signals with significant characteristics of bipolar disorder to obtain a target classification model that can be used for bipolar disorder classification. Since the speech signal contains features directly related to bipolar disorder, such as the semantics, tone, and rhythm of the patient's speech, the target classification model can be used for accurate classification of bipolar disorder. When using the target classification model to classify bipolar disorder, it can avoid over-reliance on the doctor's subjective diagnosis and effectively improve the accuracy of bipolar disorder classification prediction.

[0025] In some embodiments, step S102 includes: Step S121: Extract semantic features, mood features and tone features of the sample speech, where semantic features represent the contextual information of the semantics corresponding to the sample speech, mood features represent the emotional expression of the sample speech, and tone features represent the prosodic parameters of the sample speech. Step S122: Input semantic features, mood features, and tone features into the multimodal language model to obtain the sample classification prediction results of the multimodal language model output sample speech.

[0026] In embodiments of this application, three features can be separated and quantified from each sample audio file. These features are not the original waveform, but rather processed numerical representations (e.g., vectors or tensors) used to capture speech biomarkers for BD. The purpose of extraction is to reduce the dimensionality of high-dimensional audio data into interpretable features, facilitating subsequent model processing. Why three? Because BD symptoms are multimodal: semantics reflects cognitive content (e.g., depressive thoughts), tone reflects emotional state (e.g., irritability), and pitch reflects physiological rhythm (e.g., energy level). After extraction, these features can be analyzed independently or fused. Semantic features characterize the contextual information of the semantics corresponding to the sample audio. Semantics refers to the meaning and logical structure of language, including lexical meaning, sentence coherence, topic distribution, etc. Contextual information refers to the preceding and following relationships and implied meanings of utterances. In BD, this captures the patient's thought content (e.g., grandiose delusional words during manic episodes). Tone features characterize the emotional expression of the sample audio. Tone refers to the emotional coloring and attitude during speech, such as warmth, anger, or sadness. Emotional expression includes intensity, polarity, and type (e.g., positive / negative). In BD, this reflects emotional state (e.g., the desperate tone of a depressive episode). Tone features characterize the prosodic parameters of sample speech. Tone, broadly speaking, refers to the "musical" element of speech, while prosodic parameters are acoustic physical quantities such as pitch variation and rhythm. In BD, this quantifies physiological manifestations (e.g., the rapid prosody of a manic episode). In code, these three feature vectors can be stored as NumPy arrays or PyTorch tensors.

[0027] In the embodiments of this application, the three extracted features can be used as input and processed using a multimodal language model. These features are fused internally within the model, outputting a classification prediction for the sample. This replaces directly inputting the raw audio, improving efficiency and interpretability because the features have been preprocessed. In BD training, this step occurs during the forward propagation of each batch and is used to calculate the loss. Feature fusion can take one of the following approaches: early fusion, intermediate fusion, or late fusion. Early fusion refers to directly concatenating the features into a single vector, e.g., `fused = torch.cat([semantic, tone, prosody], dim = 1)`; intermediate fusion refers to using a Cross-Attention layer to allow semantic queries to capture tone key-value pairs, capturing interactions (e.g., negative semantic reinforcement of a depressed tone); late fusion refers to processing the features separately through sub-modules (e.g., semantics into BERT, tone into CNN), and then weighted averaging or using an MLP for fusion.

[0028] In the embodiments of this application, semantic features, mood features, and tone features are extracted from sample speech. Semantic features represent the contextual information of the semantics corresponding to the sample speech, mood features represent the emotional expression of the sample speech, and tone features represent the prosodic parameters of the sample speech. The semantic features, mood features, and tone features are input into a multimodal language model to obtain the sample classification prediction results of the multimodal language model output sample speech. This can clarify the information source in the multimodal language model, improve the accuracy of information processing of the multimodal language model to a certain extent, and thus accelerate the model training process and improve the model training efficiency.

[0029] In some embodiments, the multimodal language model includes a Transformer encoder and a gated recurrent unit network module. Step S122 includes: Step S1221: Using the Transformer encoder, extract the semantic encoding vector of semantic features, the tone encoding vector of tone features, and the tone encoding vector of tone features, respectively. Step S1222: Using the gated recurrent unit network module, based on the semantic encoding vector, mood encoding vector and tone encoding vector, determine the sample classification prediction result of the sample speech.

[0030] In embodiments of this application, the multimodal language model includes a Transformer encoder and a gated recurrent unit network module. The Transformer encoder can independently process the three previously extracted raw features (semantics, mood, and tone) to generate corresponding encoded vectors. These vectors are compressed and enhanced representations of the raw features, capturing deeper patterns (e.g., capturing dependencies within features through self-attention mechanisms). In a BD context, this helps amplify symptom signals: for example, semantic vectors reinforce the negative context of depressive episodes, while tone vectors highlight prosodic variations of manic episodes. The core component of the Transformer model, consisting of multiple layers of self-attention and a feed-forward network, is used to process sequential data and extract high-level representations. It does not include a decoder part and is only used to encode the input. The Transformer's advantage lies in parallel processing and capturing long-range dependencies, and it is often used in BERT-like models in speech / text tasks. The extraction of the three features can be performed in parallel.

[0031] In the embodiments of this application, a Gated Recurrent Unit (GRU) network module can be used to fuse three encoded vectors, capture their temporal interactions, and output the final classification prediction result. GRU excels at processing sequential data and can model the dynamic changes in speech (e.g., the transition from depression to mania). In BD classification, this step integrates multimodal signals to generate more accurate predictions (e.g., episodic category). The Gated Recurrent Unit Network Module (GRU) is a variant of a Recurrent Neural Network (RNN) that uses update and reset gates to control information flow and avoid gradient vanishing. It is simpler and more efficient than LSTM and suitable for temporal tasks such as speech classification. A module refers to a GRU layer or stacked GRUs. The sample classification prediction result refers to the model's final prediction output, typically a probability distribution (e.g., softmax output), representing the probability of the BD category (e.g., [normal, BD-I mania, BD-II depression]). In practice, the combination of Transformer and GRU can improve accuracy by 10-15% in multimodal speech tasks (compared to single MLP), especially for dynamic symptoms such as mixed episodes.

[0032] In the embodiments of this application, by utilizing a Transformer encoder, semantic encoding vectors of semantic features, tone encoding vectors of mood features, and tone encoding vectors of pitch features are extracted respectively. Using a gated recurrent unit network module, based on the semantic encoding vectors, tone encoding vectors, and pitch encoding vectors, the sample classification prediction result of the sample speech is determined. A multimodal language model can be composed of a Transformer encoder and a gated recurrent unit network module. Furthermore, the semantic encoding vectors of semantic features, tone encoding vectors of mood features, and tone encoding vectors of pitch features can be processed sequentially through the Transformer encoder and the gated recurrent unit network module to output the sample classification prediction result. This can, to a certain extent, improve the information processing capability and prediction performance of the multimodal language model.

[0033] In some embodiments, step S121 may include: Step S1211: Perform data augmentation on the sample speech to obtain sample augmented speech; data augmentation includes adding background noise or speed adjustment. Step S1212: Extract semantic features, mood features, and tone features of the sample enhanced speech.

[0034] In the embodiments of this application, one or more enhancement techniques can be applied to generate variant versions of speech (sample-enhanced speech) starting from the original sample speech. The purpose of enhancement is to expand the dataset size, introduce diversity, and prevent the model from overfitting to clean data, thereby improving performance in the real world (such as noisy or variable-speed speech). In BD, this simulates the speech variations of patients under different environments or emotions (e.g., faster speech rate under stress). Background noise can be added using the librosa, torchaudio, or audiomentations libraries by calling noise functions; regarding variable-speed enhancement, human speaking speed varies with emotion, age, and culture (e.g., in BD, manic speech is fast > 5 syllables / second, while depressive speech is slow < 3 syllables / second). Variable-speed enhancement introduces these variations, making the model less sensitive to speed and improving generalization (accuracy can be improved by 3-8%).

[0035] In embodiments of this application, enhanced speech can be used as input to extract three features, as described previously. This step ensures the model learns from diverse data, improving robustness to noise or speed variations. In BD classification, post-enhancement extraction better captures true symptoms (e.g., detecting the monotonous tone of depression even in noise).

[0036] In the embodiments of this application, sample-enhanced speech is obtained by performing data augmentation on sample speech; data augmentation includes adding background noise or speed adjustment, and extracting semantic features, mood features and tone features of sample-enhanced speech, which can improve the richness of the data used for model training and improve the model's generalization ability and prediction performance.

[0037] In some embodiments, the acquisition of a patient voice dataset includes: Voice samples from real patients were collected from clinical records or mobile devices and categorized by experts. The categorization labels included at least one of manic, depressive, and normal states.

[0038] In the embodiments of this application, a smartphone application can be used to collect patient sample voice recordings. The dataset consists of natural conversation recordings (total duration exceeding 2000 hours, similar in structure to the PRIori dataset or AVEC2018BDCorpus) collected through clinical collaboration from approximately 150 BD patients and 50 healthy controls. Each recording is labeled with a clinician assessment status tag under the DSM-5 standard.

[0039] In the embodiments of this application, by collecting real patient sample voices from clinical records or mobile devices and labeling them with expert annotations, the data reliability and diversity of real patient sample voices can be significantly improved.

[0040] like Figure 2 As shown, Figure 2 The graph shows the loss value decreasing during model training. "Epoch" represents the training epoch, the blue line represents the cross-entropy loss during training, and the orange line represents the validation loss (Val Loss). Each epoch contains 1000 samples. The graph shows that after two training epochs, the loss value has decreased to a considerably low level. Simultaneously, the validation loss did not increase, indicating that the model has not overfitted.

[0041] like Figure 3 As shown, Figure 3 A method for generating personalized treatment dialogues for bipolar disorder, provided in this application embodiment, includes: Step S301: Collect target speech data of the target patient at the time of onset; Step S302: Input the target speech data into the target classification model to obtain the target classification prediction result output by the target classification model; wherein, the target classification model is obtained based on the model training method described in any of the above items; Step S303: Based on the target classification prediction results, generate personalized treatment dialogue information that matches the target patient.

[0042] In the embodiments of this application, a target classification model obtained through the above-described model training method can be used to predict the classification of bipolar disorder in target patients. Specifically, for abusive speech, firstly, target speech data at the time of the onset of the illness is collected. Then, the target speech data can be input into the target classification model to obtain the target classification prediction result output by the model. Subsequently, to provide corresponding dialogue therapy to the target patient, personalized treatment dialogue information corresponding to the target classification prediction result can be generated to match the personalized treatment dialogue information with the target patient. The core difference between treatment dialogues for different bipolar disorder classifications lies in their different focuses. For example, if the classification prediction result is mania, the dialogue features used by the therapist may include a gentle, calm, and non-confrontational tone, avoiding stimulation and directly challenging grandiose thoughts (which may cause defensiveness or exacerbate mania), with an emphasis on emotional stability, reality testing, and establishing compliance. If it is a depressive phase, the dialogue features used by the therapist may include empathy, encouragement, and structuring, helping the patient challenge negative automatic thoughts, activate behavior, and rebuild hope. An example dialogue is as follows: "Scenario: The patient is in a hypomanic phase, with abundant energy, jumping thoughts, reduced sleep, and planning multiple 'big projects,' but is beginning to experience impulsive spending and interpersonal conflicts. The therapist uses psychoeducation and IPSRT elements."

[0043] Therapist: You seem particularly energetic today, and you're speaking very quickly. I noticed you mentioned that you've only slept 3-4 hours a night for the past few days, but you still feel very energetic. Could you tell me what has been most exciting for you lately? Patient: This is amazing! I've finally figured it out! I'm going to start a company, write a book, and even take a round-the-world trip within three months! Last night I wrote a 30-page business plan in one go, I feel invincible! Therapist (remaining calm, smiling, not interrupting, but gently guiding): It sounds like you have a lot of great ideas in your head, which is really encouraging. We've talked before that bipolar disorder can sometimes cause the brain's "excitement switch" to be turned up too high, resulting in a particularly fast, numerous, and compelling stream of thoughts. Do you feel that these thoughts are different from your usual ones? Patient: It's different? It's faster and clearer! I used to always procrastinate, but now I can finish in one night what others do in a month! Therapist (not directly denying, but providing information): Yes, many people have described this "super productivity," like the brain being accelerated. But we also know that if this continues, with reduced sleep and excessive energy depletion, it might suddenly drop, leading to exhaustion or even depression. Would you like to explore together how to make these good ideas "slower" in their implementation, instead of burning through your energy all at once? For example, we could try writing down our ideas first, and then deciding which ones are most worthwhile to do tomorrow? Patient (slightly hesitant): Hmm...maybe. But I really can't stop now! Therapist: I understand that feeling of not being able to stop. We can start with small things, like trying to put your phone on airplane mode tonight, going to bed before 10 pm, even if it's just lying in bed. Many people find that forcing themselves to have a regular schedule makes their minds much clearer, not more chaotic. Would you like to try this "little experiment"? We'll see the results next time. Its key strategies are to avoid arguing or saying "you are pathological," but to frame the symptoms using the "accelerated bipolar pattern"; to emphasize the importance of sleep / rhythm; to use behavioral experiments to reduce resistance; and to establish a treatment alliance.

[0044] In the embodiments of this application, target speech data of the target patient at the time of onset is collected; the target speech data is input into the target classification model to obtain the target classification prediction result output by the target classification model; based on the target classification prediction result, personalized treatment dialogue information matching the target patient is generated. The target classification model can be used to classify and predict the bipolar disorder of the target patient, and corresponding treatment dialogue can be generated according to the classification prediction result, which can improve the matching degree between the treatment dialogue and the target patient to a certain extent.

[0045] In some embodiments, step S302 may include: Step S321: Extract the target semantic features, target mood features, and target tone features from the target speech data; Step S322: Using the target classification model, based on the target semantic features, target tone features, and target pitch features, determine the target classification prediction result.

[0046] In the embodiments of this application, feature extraction can be performed before the model processes the data, extracting the target semantic features, target tone features, and target pitch features of the target speech data. The target semantic features, target tone features, and target pitch features are then input into the target classification model for processing, so as to avoid the model recognizing the corresponding features on its own.

[0047] In the embodiments of this application, by extracting the target semantic features, target tone features, and target pitch features from the target speech data, and using the target classification model, the target classification prediction result is determined based on the target semantic features, target tone features, and target pitch features. This avoids problems such as feature ambiguity caused by the model automatically identifying and extracting the corresponding features, thereby improving the classification prediction accuracy of the target classification model.

[0048] refer to Figure 4 , Figure 4 This application provides a logic block diagram of a model training apparatus 400, which may include: The first acquisition module 401 is used to acquire sample speech datasets of sample patients. The sample speech datasets include multiple sample speech samples labeled with bipolar disorder classification tags. The first input / output module 402 is used to input sample speech into the multimodal language model and obtain the sample classification prediction result of the multimodal language model output sample speech; The first determining module 403 is used to determine the multi-class cross-entropy loss value of the multimodal language model based on the sample classification prediction results and the bipolar disorder classification label. Training module 404 is used to iteratively train and optimize the parameters of the multimodal language model based on the multi-class cross-entropy loss value to obtain the target classification model; the target classification model is used to determine the bidirectional affective disorder category of the target patient based on the target patient's target speech data.

[0049] Optionally, the first input / output module 402 includes: The first extraction submodule is used to extract semantic features, mood features and tone features of sample speech. Semantic features represent the contextual information of the semantics corresponding to the sample speech, mood features represent the emotional expression of the sample speech, and tone features represent the prosodic parameters of the sample speech. The first input-output submodule is used to input semantic features, mood features, and tone features into the multimodal language model to obtain the sample classification prediction results of the multimodal language model output sample speech.

[0050] Optionally, the multimodal language model includes a Transformer encoder and a gated recurrent unit network module, with the first input and output in the module, including: The first extraction unit is used to extract the semantic encoding vector of semantic features, the tone encoding vector of mood features, and the tone encoding vector of tone features using the Transformer encoder. The determination unit is used to determine the sample classification prediction result of sample speech based on semantic encoding vector, mood encoding vector and tone encoding vector using the gated recurrent unit network module.

[0051] Optionally, the first extraction submodule includes: The enhancement unit is used to perform data enhancement on the sample speech to obtain enhanced sample speech; data enhancement includes adding background noise or speed adjustment. The second extraction unit is used to extract semantic features, mood features, and tone features of the sample enhanced speech.

[0052] Optionally, the first acquisition module includes: The acquisition submodule is used to acquire real patient sample voices from clinical records or mobile devices, and to annotate them with expert labels. The labels include at least one of manic, depressive and normal states.

[0053] The model training device in this application embodiment can be integrated into an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices besides a terminal. For example, the electronic device can be a GPUbox, mobile phone, tablet computer, robot, ultra-mobile personal computer (UMPC), or personal digital assistant (PDA), etc., and can also be a server, personal computer (PC), etc. This application embodiment does not specifically limit the specific implementation.

[0054] The model training device in this application embodiment can be a device with an operating system. This operating system can be Android, Linux, Windows, or other possible operating systems; this application embodiment does not specifically limit it.

[0055] The model training device provided in this application embodiment can achieve... Figure 1 and Figure 2 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0056] Reference Figure 5 , Figure 5 This application provides a logic block diagram of a personalized treatment dialogue generation device for bipolar disorder, wherein the device 500 may include: The second acquisition module 501 is used to acquire the target patient's speech data at the time of the onset of the disease. The second input / output module 502 is used to input the target speech data into the target classification model and obtain the target classification prediction result output by the target classification model. The generation module 503 is used to generate personalized treatment dialogue information that matches the target patient based on the target classification prediction results.

[0057] Optionally, the second input / output module 502 includes: The second extraction submodule is used to extract the target semantic features, target mood features, and target tone features of the target speech data. The second determination submodule is used to determine the target classification prediction result based on the target semantic features, target tone features, and target pitch features using the target classification model.

[0058] The personalized treatment dialogue generation device for bipolar disorder in this application embodiment can be integrated into an electronic device, or a component of an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices besides a terminal. For example, the electronic device can be a GPUbox, mobile phone, tablet computer, laptop computer, handheld computer, robot, ultra-mobile personal computer (UMPC), or personal digital assistant (PDA), etc., and can also be a server, personal computer (PC), etc. This application embodiment does not specifically limit the specific implementation.

[0059] The personalized treatment dialogue generation device for bipolar disorder in this application embodiment can be a device with an operating system. This operating system can be Android, Linux, Windows, or other possible operating systems; this application embodiment does not specifically limit it.

[0060] The bipolar disorder personalized treatment dialogue generation device provided in this application embodiment can achieve… Figure 3 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0061] This application provides an electronic device, see [link to relevant documentation] Figure 6 The electronic device 60 includes: a processor 601, a memory 602, and a computer program 6021 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the program, it implements the model training method of the foregoing embodiments, or the personalized treatment dialogue generation method for bipolar disorder.

[0062] This application also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps in the model training method or the bipolar disorder personalized treatment dialogue generation method disclosed in this application.

[0063] This application also provides a computer program product that, when run on an electronic device, enables a processor to implement the steps of the model training method or the bipolar disorder personalized treatment dialogue generation method disclosed in this application.

[0064] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0065] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0066] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0068] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0069] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0070] The above provides a detailed description of the model training method, the personalized treatment dialogue generation method and apparatus for bipolar disorder provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and its core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A model training method, characterized in that, The method includes: Collect sample speech datasets from sample patients, the sample speech datasets including multiple sample speech samples labeled with bipolar disorder classification tags; The sample speech is input into a multimodal language model to obtain the sample classification prediction result of the sample speech output by the multimodal language model; Based on the sample classification prediction results and the bipolar disorder classification labels, the multi-class cross-entropy loss value of the multimodal language model is determined. Based on the multi-class cross-entropy loss value, the parameters of the multimodal language model are iteratively trained and optimized to obtain the target classification model; the target classification model is used to determine the bidirectional affective disorder category of the target patient based on the target patient's target speech data.

2. The method according to claim 1, characterized in that, The step of inputting the sample speech into a multimodal language model to obtain the sample classification prediction result of the sample speech output by the multimodal language model includes: Extract semantic features, mood features, and tone features from the sample speech, wherein the semantic features represent the contextual information of the semantics corresponding to the sample speech, the mood features represent the emotional expression of the sample speech, and the tone features represent the prosodic parameters of the sample speech. The semantic features, the mood features, and the tone features are input into a multimodal language model to obtain the sample classification prediction results of the sample speech output by the multimodal language model.

3. The method according to claim 2, characterized in that, The multimodal language model includes a Transformer encoder and a gated recurrent unit network module. The step of inputting the semantic features, the mood features, and the tone features into the multimodal language model to obtain the sample classification prediction result of the sample speech output by the multimodal language model includes: Using a Transformer encoder, the semantic encoding vector of the semantic feature, the tone encoding vector of the mood feature, and the tone encoding vector of the tone feature are extracted respectively. Using the gated recurrent unit network module, the sample classification prediction result of the sample speech is determined based on the semantic encoding vector, the mood encoding vector, and the tone encoding vector.

4. The method according to claim 1, characterized in that, The extraction of semantic features, mood features, and tone features from the sample speech includes: The sample speech is augmented with data to obtain augmented speech; the data augmentation includes adding background noise or speed adjustment. Extract the semantic features, mood features, and tone features of the sample speech to enhance speech.

5. The method according to claim 1, characterized in that, The collected patient voice dataset includes: Real patient voice samples were collected from clinical records or mobile devices and categorized by experts. The categorization labels included at least one of manic, depressive, and normal states.

6. A method for generating personalized treatment dialogues for bipolar disorder, characterized in that, The method includes: Collect target speech data of the target patient at the time of onset of illness; The target speech data is input into a target classification model to obtain the target classification prediction result output by the target classification model; wherein, the target classification model is obtained based on the model training method according to any one of claims 1 to 5; Based on the target classification prediction results, personalized treatment dialogue information matching the target patient is generated.

7. The method according to claim 6, characterized in that, The step of inputting the target speech data into the target classification model and obtaining the target classification prediction result output by the target classification model includes: Extract the target semantic features, target mood features, and target tone features from the target speech data; Using a target classification model, the target classification prediction result is determined based on the target semantic features, the target tone features, and the target pitch features.

8. A model training device, characterized in that, The device includes: The first acquisition module is used to acquire sample speech datasets from sample patients, the sample speech datasets including multiple sample speech samples labeled with bipolar disorder classification tags; The first input / output module is used to input the sample speech into the multimodal language model and obtain the sample classification prediction result of the sample speech output by the multimodal language model; The determination module is used to determine the multi-class cross-entropy loss value of the multimodal language model based on the sample classification prediction results and the bipolar disorder classification label; An iterative module is used to iteratively train and optimize the parameters of the multimodal language model based on the multi-class cross-entropy loss value to obtain a target classification model; the target classification model is used to determine the bidirectional affective disorder category of the target patient based on the target patient's target speech data.

9. A personalized treatment dialogue generation device for bipolar disorder, characterized in that, The device includes: The second acquisition module is used to acquire the target patient's speech data at the time of the onset of the disease. The second input / output module is used to input the target speech data into the target classification model and obtain the target classification prediction result output by the target classification model. The generation module is used to generate personalized treatment dialogue information that matches the target patient based on the target classification prediction results.