Multi-mode psychological problem detection system based on psychological interviews
By designing a multimodal psychological problem detection system, integrating voice, video, text and physiological signal data, and using artificial intelligence and deep learning technology for data fusion and analysis, the problems of insufficient data fusion, insufficient personalization, limited real-time and scalability in the existing technology are solved, and more accurate and efficient psychological problem detection is achieved.
Patent Information
- Application Number
- CN202510302855.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems such as insufficient data fusion, insufficient personalization, and limited real-time and scalability in psychological problem detection.
A multimodal psychological problem detection system based on psychological interviews was designed, including digital doctor interaction module, multimodal emotion analysis module, structured probability model, semantic analysis module and additional medical information provision module. The system integrates voice, video, text and physiological signal data, uses artificial intelligence and deep learning technologies to fusion and analysis, dynamically adjusts the diagnostic path, and realizes personalized evaluation.
It improves the accuracy and efficiency of psychological problem detection, achieves more accurate personalized evaluation, and can conduct user self-diagnosis without the guidance of a doctor, with a wide range of application prospects.
Smart Images

Figure CN120148866A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of medical health, psychology, computer science, human-computer interaction, etc., and particularly relates to a multi-modal psychological problem detection system based on psychological interviews. Background Art
[0002] With the increase in social pressure and the acceleration of the pace of life, mental health problems have increasingly become the focus of global attention. Traditional methods for detecting psychological problems mainly rely on face-to-face interviews and questionnaires by psychiatrists. Although these methods are effective, they have the following limitations:
[0003] Time and space limitations: Face-to-face interviews require appointments and physical presence, which limit the access to psychiatrists for users in remote locations or with tight schedules.
[0004] Strong subjectivity: Traditional psychological assessments rely on the subjective descriptions of patients and the empirical judgments of psychiatrists, which may lead to biases.
[0005] Single data source: Assessments are based only on language and text information, ignoring the impact of non-verbal information (such as facial expressions, speech intonation, body language, etc.) on mental states.
[0006] Poor real-time performance: Traditional methods are difficult to monitor changes in mental states in real time and cannot provide timely intervention measures.
[0007] In recent years, with the development of artificial intelligence and multi-modal data analysis technologies, psychological problem detection systems based on multi-modal data have gradually become a research hotspot. These systems can more comprehensively evaluate an individual's mental state by integrating multiple data sources (such as voice, video, text, physiological signals, etc.). However, the existing technologies still have the following problems:
[0008] Insufficient data fusion: Existing systems still have deficiencies in the fusion and processing of multi-modal data and are difficult to fully utilize the complementary information between different modalities.
[0009] Lack of personalization: Existing systems usually adopt general models and are difficult to perform personalized analysis and diagnosis according to the specific situation of an individual.
[0010] Limited real-time performance and scalability: Existing systems have performance bottlenecks in processing large-scale data and real-time analysis. Summary of the Invention
[0011] The purpose of the present invention is to provide a multi-modal psychological problem detection system based on psychological interviews to solve the technical problems of insufficient data fusion, lack of personalization, and limited real-time performance and scalability in the detection of psychological problems in the existing technology.
[0012] To solve the above technical problems, the specific technical solution of the present invention is as follows:
[0013] A multi-modal psychological problem detection system based on psychological interviews, the system includes a digital doctor interaction module, a multi-modal emotion analysis module, a structured probability model, a semantic analysis module, and an additional medical information providing module;
[0014] The digital doctor interaction module receives the text from the semantic analysis module, synthesizes it into a digital human picture, voice and plays it, and interacts with the patient; the digital doctor interaction module collects audio information through a microphone, collects video information through a camera, and sends the collected video information and audio information to the multi-modal emotion analysis module;
[0015] The multi-modal emotion analysis module receives the video signal and audio signal collected by the digital doctor interaction module, uses speech recognition technology to transcribe the text signal, and calculates the final prediction result by synthesizing the video signal, audio signal and text signal as the user emotion intensity score corresponding to the current problem and sends it to the semantic analysis module;
[0016] The additional medical information providing module collects the additional medical information of the patient, and injects the additional medical information into the large language model of the semantic analysis module through prompt engineering;
[0017] The semantic analysis module is responsible for generating user-friendly answers and follow-up questions, obtaining the disease probability according to the dialogue semantic content, additional medical information, and the user emotion intensity score corresponding to the current problem sent by the multi-modal emotion analysis module, and inputting the disease probability into the structured probability model;
[0018] The structured probability model is modeled through a Markov decision process, regards the possible answers to each question as state nodes, combines the disease probability input by the semantic analysis module, and determines the selection of the next question through the state transition probability.
[0019] Furthermore, the multi-modal emotion analysis module uses Librosa as an audio feature extractor to extract the features of audio information to obtain audio features; uses OpenFace to perform face key point detection, head pose estimation, gaze estimation, and facial action unit recognition on video information to extract video features; uses the Whisper speech recognition tool to transcribe the audio information into text information; uses BERT to extract the features of the text information to obtain text features; the multi-modal emotion analysis module obtains the final prediction result as the user emotion intensity score corresponding to the current problem by performing feature fusion on the feature data of the three modalities of video features, audio features, and text features.
[0020] Furthermore, the semantic analysis module realizes the following functions:
[0021] Mental State Perception Dialogue Engine: Adopting a dynamic context awareness mechanism, based on the multi-round dialogue records between the user and the model on topics related to the MINI scale, it extracts emotional keywords and semantic implicit features in real time;
[0022] Incremental User Portrait Modeling: After each round of dialogue, the dialogue records are fused with the user's historical portrait data through a large model;
[0023] Intelligent Scale Question Jumping System: Based on the current answer content, it calculates the relevance scores of each question in the MINI scale and realizes question jumping in cooperation with a structured probability model.
[0024] Furthermore, the relevance scores of each question in the MINI scale are calculated in the following way:
[0025] Score i = cos(Question i , text) · Score
[0026] Score i = 0.7 · Score i + 0.3 · Score i-1
[0027] Among them, Score i is the score of the i-th question in the scale, Question i represents the i-th question, text represents the dialogue record, cos is the calculation function of cosine similarity, which is used to calculate the relevance between the i-th question and the dialogue record, Score represents the preset score benchmark evaluation parameter generated by the large language model through prompt engineering. This parameter comprehensively integrates three-dimensional information: the semantic content of the previous i rounds of dialogue, user portrait modeling, and the user's emotional intensity score corresponding to the current question sent by the multi-modal emotion analysis module. Finally, it is the benchmark evaluation parameter generated by the large language model through prompt engineering.
[0028] Furthermore, the state of the structured probability model is expressed as:
[0029] S t = (Q t , H t , P t )
[0030] Among them, Q t is the current major question number (1 - 12), H t is the historical answer sequence, P t ∈[0, 1] is the disease probability vector of 12 dimensions, and the disease probability of the i-th dimension is expressed as P t(i), The probability of illness is calculated by the large language model and sent to the structured probability model. The probability of illness will be dynamically updated in its entirety as the conversation progresses. Define action A t , representing the judgment of three states: (1) continue to ask the current question; (2) end the current question; (3) go back to the previous question; The transition probability P(S t+1 |S t ,A t ) represents the probability of transferring to the next state S t when action A t is executed in state S t+1 . This probability is based on historical data and the initially set probability, and is dynamically updated during the Q&A process.
[0031] Compared with the prior art, the present invention has the following beneficial technical effects:
[0032] 1) The innovation of the present invention lies in abstracting the open-ended doctor-patient Q&A process into a model, and quantitatively describing the psychological consultation process through this model. Different from traditional methods, the system can simulate the decision-making reasoning of doctors during the diagnosis process, dynamically adjust the diagnostic path, so as to achieve a more accurate personalized psychological problem assessment. In addition, combined with multi-modal data analysis and artificial intelligence technology, the system can improve the accuracy and efficiency of diagnosis by deeply analyzing data such as patients' expressions and voices. This innovation not only provides more advanced diagnostic tools for medical institutions, but also provides a convenient self-detection method for individual users, and has broad application prospects.
[0033] 2) The present invention can synthesize texts that conform to human communication habits; can synthesize realistic and interactive digital human videos; patients only need to talk to the digital human to achieve automatic psychological problem diagnosis, and can automatically adjust the scale results according to the collected patient information; finally, the system will output a diagnostic report; the present invention can enable users to self-diagnose without the guidance of a doctor. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 is a schematic diagram of the multi-modal psychological problem detection system based on psychological interviews of the present invention.
[0036] Figure 2 is a schematic diagram of feature fusion of the multi-modal emotion analysis module of the present invention.
[0037] Figure 3 This is a schematic diagram of the interaction system of the semantic analysis module of the present invention.
[0038] Figure 4 This is a schematic diagram of the interaction of the structured probability model of the present invention. Specific implementation manners
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] The multi-modal psychological problem detection system based on psychological interviews proposed by the present invention includes a digital doctor interaction module, a multi-modal emotion analysis module, a structured probability model, a semantic analysis module, and an additional medical information providing module, as Figure 1 shown.
[0041] Among them, the structured probability model is the core of the system. Based on the Mini-International Neuropsychiatric Interview (MINI), it systematically asks patients through a pre-designed interrogation "script" to diagnose their potential mental diseases. Its feature is that each question has a clear standardized process, and the patient's answer will determine the selection or jump path of the next question. The structured probability model is to perform mathematical modeling and encapsulation on the basis of the Mini-International Neuropsychiatric Interview, so that the system not only depends on the immediate answer, but also combines the patient's outpatient case and the previous answer trajectory to achieve personalized logical judgment. During the implementation process, the system can automatically backtrack. If an incorrect jump or abnormal answer is detected, it can roll back to the previous question to re-evaluate the patient's state.
[0042] The semantic analysis module is responsible for generating user-friendly answers and follow-up questions. For example, when the structured probability model has difficulty in choosing a branch at a certain node, the semantic analysis module is responsible for repeatedly generating follow-up questions until the system has enough confidence in the route selection.
[0043] The multi-modal emotion analysis module is used to select the node jump route in the structured probability model. This module will receive two signals (video, audio) collected by the digital doctor interaction module, and use speech recognition technology to transcribe the text signal. The final prediction result is calculated by integrating the video signal, audio signal and text signal and sent to the semantic analysis module as the user emotion intensity score corresponding to the current question.
[0044] Throughout the process, the digital doctor interaction module receives the text from the semantic analysis module, synthesizes it into the digital human's picture, voice and plays them, and interacts with the patient. During the whole interaction process, the patient will be recorded in real time.
[0045] The digital doctor interaction module is a device equipped with a screen, a camera, and a microphone. A virtual digital human will appear on the screen of the device. This digital human is generated by highly realistic artificial intelligence technology. Not only are its appearance and movements very natural, but its language expression and emotional reactions are also close to those of real people. Users will communicate with this virtual digital human just as they would with a real doctor.
[0046] The virtual digital human of the digital doctor interaction module is synthesized by face synthesis of image generation technology. It uses computer vision and machine learning algorithms to create or modify face images to make them look real and natural. The present invention uses the TalkAsYou model to complete the face driving work. TalkAsYou mainly includes a feature extraction module, a personalized speaking style learning module, a motion generation module, a lip motion optimization module, and a neural rendering module.
[0047] The feature extraction module is used to extract the semantic and acoustic features of the audio and the 3DMM coefficients of the face in the video frame.
[0048] The personalized speaking style learning module dynamically fuses the multi-frame information of the reference video based on audio correlation, so as to learn the personalized speaking style of the person in the reference video.
[0049] The motion generation module generates a motion sequence of the face, that is, the 3DMM coefficients of the face, based on multi-modal features, according to the given Gaussian noise, the motion prior containing the personalized speaking style, and the driving audio features.
[0050] In order to alleviate the problem that the model focuses on the naturalness of the overall face motion during motion generation, resulting in a poor audio-lip synchronization effect, the lip motion optimization module further optimizes the lip motion generated by the motion generation module by using a lip motion optimization module.
[0051] The neural rendering module maps the face motion sequence (i.e., the 3DMM coefficients) generated by the model into a realistic two-dimensional face video.
[0052] In the communication with the virtual digital human, the digital doctor interaction module realizes the questioning of the user's psychological problems and the collection of the user's question answering information. The patient's expressions and voices will be collected and analyzed in real time. The system uses deep learning technology to comprehensively analyze the patient's facial expressions, speech rate, tone, and answering content to determine whether the patient's answers to each question in the scale are valid. This information includes not only the analysis of language content but also the interpretation of non-verbal information, such as the emotional color of the voice and the subtle changes in facial expressions, which are important bases for evaluating the patient's psychological state.
[0053] The digital doctor interaction module collects audio information through a microphone and video information through a camera, and sends the collected video information and audio information to the multi-modal emotion analysis module.
[0054] The multi-modal emotion analysis module receives the audio information and video information collected by the digital doctor interaction module. It uses Librosa as an audio feature extractor to extract the features of the audio information and obtain audio features; it uses OpenFace to perform face key point detection, head pose estimation, gaze estimation, and facial action unit recognition on the video information to extract video features. The multi-modal emotion analysis module uses the Whisper speech recognition tool to transcribe the audio information into text information. Subsequently, BERT is used to extract the features of the text information to obtain text features.
[0055] The multi-modal emotion analysis module performs feature fusion on the feature data of the three modalities of video features, audio features, and text features.
[0056] The feature fusion of the multi-modal emotion analysis module adopts a modality decoupling fusion method, as Figure 2 shown. The multi-modal emotion analysis module maps the features of each modality to two different feature spaces. The first is the Modality-Invariant space, where cross-modal representation learning learns their commonalities and reduces the modality gap. The second is the Modality-Specific space, where each modality representation learns its modality-specific characteristics. The following are the specific implementation algorithms and formulas:
[0057] For each input sample, the input contains the feature sequences of three modalities: text features (L), video features (V), and audio features (A). The text feature sequence is represented as where, represents the set of real numbers, T l represents the length of the text feature sequence, and d l represents the size of the text features. The video feature sequence is represented as where, T v represents the length of the video feature sequence, and d vIndicates the size of the video feature. The audio feature sequence is represented as where T a represents the length of the audio feature sequence, and d a represents the size of the audio feature.
[0058] For the feature sequence U of each modality l , U v , U a , it is mapped to a fixed-length feature vector through a temporal feature extractor Extractor (such as a bidirectional long short-term memory (LSTM) network), and the mapping method is as follows:
[0059]
[0060] where represents the text feature vector, d h represents the length of the feature vector, represents the parameters of the text LSTM layer, and Extractor represents the temporal feature extractor; represents the video feature vector, represents the parameters of the video LSTM layer; represents the audio feature vector, represents the parameters of the audio LSTM layer.
[0061] Then project each feature vector into two subspaces to obtain hidden vectors, and the projection method is as follows:
[0062]
[0063] where represents the hidden vector of the text modality invariant subspace feature, E c represents the modality invariant subspace projection layer, and θ c represents the parameters of the modality invariant subspace projection layer, represents the hidden vector of the text modality unique subspace feature, E p represents the modality unique subspace projection layer, and θ p represents the parameters of the modality unique subspace projection layer; represents the hidden vector of the video modality invariant subspace feature, represents the hidden vector of the video modality unique subspace feature; represents the hidden vector of the audio modality invariant subspace feature, represents the hidden vector of the audio modality unique subspace feature.
[0064] After projecting the feature sequences of the three modalities into two subspaces, use Transformer for feature fusion, and use the scaled dot-product attention mechanism:
[0065]
[0066] Among them, Attention(Q, K, V) represents the attention function, where Q, K, and V are the query, key, and value matrices respectively; softmax represents the softmax function, and T represents the transpose.
[0067] Transformer can compute attention in parallel, and each attention module is called a "head". The expression of the i-th head is as follows:
[0068]
[0069] Among them, represents the query projection matrix, represents the key projection matrix, represents the value projection matrix.
[0070] Stack the hidden vectors of 6 modalities into a matrix Then, make each modality hidden vector execute the cross-modal multi-head attention mechanism. This cross-modal attention mechanism can enable each representation to learn the representations of other modalities and has a synergistic effect. Set Then the new matrix generated by Transformer is
[0071]
[0072] Among them, represents the hidden vector of the invariant subspace feature of the text modality after attention processing, represents the hidden vector of the invariant subspace feature of the video modality after attention processing, represents the hidden vector of the invariant subspace feature of the audio modality after attention processing, represents the hidden vector of the unique subspace feature of the text modality after attention processing, represents the hidden vector of the unique subspace feature of the video modality after attention processing, represents the hidden vector of the unique subspace feature of the audio after attention processing; represents vector concatenation, MultiHead represents the multi-head attention function, and θ att represents the parameter of the multi-head attention function, and θ att ={W q , W k , W v , W o}, n represents the number of multi-head attention heads, and W o represents the projection matrix.
[0073] Finally, the outputs of the Transformer are concatenated into an output vector
[0074] Then, the function generates the final prediction result. G represents the task prediction function, and θ out represents the parameters of the task prediction function. The final prediction result is sent to the semantic analysis module as the user's emotional intensity score corresponding to the current question.
[0075] The additional medical information providing module collects the patient's additional medical information (such as case history, self-rating scale, self-evaluation), and then injects the additional medical information into the large language model of the semantic analysis module through prompt engineering.
[0076] The large language model of the semantic analysis module adopts a dialogue pre-training model and is used in combination with the Mini-International Neuropsychiatric Interview (MINI). The large language model is implemented based on the DeepSeek-R1-Distill-Llama-8B architecture.
[0077] As Figure 3 shown, the large language model mainly realizes the following functions:
[0078] Firstly, the psychological state perception dialogue engine: adopting a dynamic context awareness mechanism, based on the multi-round dialogue records of the user and the model on topics related to the MINI scale (preliminary mental health screening scale), real-time extraction of emotion keywords (such as anxiety / depression index) and semantic implicit features.
[0079] The present invention realizes an intelligent interaction mechanism similar to that of a psychologist by constructing an empathy-based prompt architecture and a multi-dimensional authenticity evaluation system. In order to make the large model more in line with a psychologist, reasonable prompts are designed for the large language model so that the large language model can guide the user to express their psychological problems in a euphemistic way like a doctor, rather than asking straightforwardly in a mechanical manner. At the same time, in order to judge whether the user is hiding the truth or expressing unclearly, the large language model also needs to deeply understand the reasons for the user's psychological problems. For the user's answer, real-time extraction of emotion keywords (such as anxiety / depression index), and then through the way of prompts, it acts on the large language model in a reverse way, dynamically adjusting the questions of the large language model, so that the system can adjust the inquiry depth and emotional orientation according to the real-time changes of the user's psychological state, and realize the dynamic personalized adaptation of the dialogue process.
[0080] Secondly, incremental user portrait modeling: After each round of dialogue, the dialogue records are fused with the user's historical portrait data through the large model.
[0081] The present invention realizes accurate psychological assessment through dynamic user portrait construction and iterative optimization mechanism. After each round of dialogue, the large language model integrates the semantic features (such as keyword frequency, topic distribution), emotional indicators (anxiety / depression index) and behavioral patterns (response delay, word ambiguity) of the current dialogue with historical data in multiple dimensions, and gradually constructs a three-dimensional user portrait matrix including a demographic attribute layer, such as: gender, age, etc.; a psychological characteristic layer, such as: mood, emotional fluctuations, etc.; an interactive behavior layer, such as: word ambiguity (frequency of use of fuzzy qualifiers, such as "probably", "possibly", and density of negative words), and corrective behavior (number of self-corrections and content reversal ratio). This portrait not only guides the adjustment of dialogue strategies in real time, but also after each round of dialogue, the gradually improved user portrait will assist the model in generating all the problem scores of the MINI scale. Finally, using these scores, we will obtain an evaluation report with both clinical validity and personalized interpretation.
[0082] Third, the intelligent jump system for scale questions: Based on the current answer content, the relevance score of each question in the MINI scale is calculated, and the question jump is achieved in conjunction with the structured probability model.
[0083] The present invention realizes accurate psychological assessment through an intelligent scale question jump system. According to the MINI scale, the present invention designs 12 major questions to judge 12 psychological problem dimensions, such as anxiety, etc. Under each major question, there are a variable number of small questions. Each round of conversation of the user will affect each question of the scale, and the specific impact depends on the cosine similarity between the conversation content and the scale. The scores generated by each round of conversation will use weighted summation to affect the final score. The specific implementation process is as follows:
[0084] Score i =cos(Question i ,text)·Score
[0085] Score i =0.7 Score i +0.3 Score i-1
[0086] Among them, Score i is the score of the i-th question in the scale, Question iDenote the i-th question, text represents the conversation record, cos is the calculation function of cosine similarity, which is used to calculate the relevance between the i-th question and the conversation record. Score represents the preset score base generated by the large language model through prompt engineering. This base comprehensively integrates the semantic content of the previous i rounds of conversations, additional medical information, and the user's emotional intensity score corresponding to the current question sent by the multi-modal sentiment analysis module in three dimensions of information. Finally, the large language model generates the benchmark evaluation parameters for 12 questions through prompt engineering. Finally, multiply by the cosine similarity respectively to obtain the scores of 12 questions. In order to update the scores of each question in real time, we use the method of weighted summation to update the scores, Score i-1 represents the score of the (i - 1)-th question. Finally, we use these 12 scores as the disease probability to input into the structured probability model to achieve personalized jumping of the questionnaire and improve the psychological evaluation ability of the model.
[0087] During the evaluation process, the large language model maintains an absolutely objective and neutral attitude, not affected by personal biases and emotions. Respect the privacy and rights of patients and ensure the confidentiality of the evaluation process.
[0088] The structured probability model is modeled through mathematical models such as the Markov decision process (MDP). Regarding the possible answers to each question as state nodes, the state transition probability is used to determine the selection of the next question. In the traditional questionnaire system, the order of questions is usually linear or predefined. Regardless of how the user answers, the jumping path of the questions will not be dynamically adjusted. However, for complex application scenarios, especially in the fields of health diagnosis, psychological evaluation, etc., relying solely on the answers to the current question is often insufficient to capture the true state and needs of the patient. The structured probability model introduces a high-order Markov decision process and combines a backtracking mechanism. By considering the previous multiple answers of the patient and combining multi-modal inputs, it realizes intelligent and personalized questionnaire jumping, as Figure 4 shown.
[0089] In the multi-round conversation record, for each small question among the 12 major questions, the answer content of each small question is scored with a disease probability score (0 - 1) in 12 dimensions. In order to maximize the diagnostic accuracy within a limited number of questions, a high-order Markov decision process (MDP) is introduced. Combining the current state and historical information, it dynamically determines the questioning strategy, avoids the limitations of the fixed path, and combines the reward function to ensure that there will be no unlimited questioning or premature ending. Define the state:
[0090] S t =(Q t ,H t ,P t )
[0091] where, Q tIs the current major problem number (1 - 12), H t Is the historical answer sequence (including answers to major questions that have been answered), P t ∈[0,1] is the disease probability vector in 12 dimensions, and the disease probability of the i-th dimension is denoted as P t (i). The disease probability is calculated by the large language model and sent to the structured probability model. The disease probability will be dynamically updated in its entirety as the conversation progresses. Define the action A t , representing the judgment of three states: (1) Continue to ask the current question; (2) End the current question; (3) Backtrack to the previous question. The transition probability P(S t+1 |S t ,A t ) represents the probability of transitioning to the next state S t when the action A t is executed in the state S t+1 . This probability is based on historical data and the initially set probability, and is dynamically updated during the Q&A process.
[0092] As an auxiliary, introduce a backtracking error correction mode to avoid inaccurate estimation of disease probability due to insufficient questioning, and further ensure the comprehensiveness of the evaluation. When the disease probability P t [i] < 0.6 (a preset threshold, which will be adjusted according to the individual characteristics of the interviewee), and the score difference ΔP t [i] < 0.1 after two consecutive rounds of major question answers, trigger the system to roll back to the first small question of the major question in this dimension and ask again. And continue to ask until P t [i] is stable (ΔP t [i] < 0.05 for three consecutive rounds) or reaches the maximum number of questions (15 times) to avoid infinite loops.
[0093] As a complete process: First, the initial probability P 0 = [0.5], representing the neutral probability in 12 dimensions, providing the starting point for the record chart. The historical record H 0 = [] records the subsequent answers. Subsequently, start generating a small question under the current major question based on patient characteristics (such as age, gender), obtain the patient's answer A t , input it into the Bayesian network, and calculate the posterior probability:
[0094]
[0095] Among them, D i is the disease state, P(A t |D i ) is the conditional probability defined by experts, P(D i |A 0 ) = P(Di ), P(D i ) is the prior probability based on statistical data, that is, it provides the starting point for the posterior probability (P(D i |A 1 , …, A t )) and finally updates P t (i) = P(D i = 1|A 1 , …, A t ), that is, the probability of being diagnosed with the disease in a specific dimension (the disease state D i = 1, indicating being diagnosed with the disease), and the historical answer H t = H t-1 + A t , H t-1 represents the historical answer to the previous round of questions. When P(D i |A 1 , …, A t ) = P(D i ) indicates that H is empty, and the Q value is updated using Q-learning:
[0096]
[0097] Among them, Q(S t , A t ) represents the Q value when the state is S t and the answer is A t . The Q value is updated according to each round of answers. The parameter α is the learning rate (usually α = 0.1), and γ is the discount factor (usually γ = 0.9). During training, the ε-greedy strategy (exploration rate 0.1) is used. R t represents the immediate reward, which is assigned according to the effectiveness of the question (+1, -0.5, +10). Then, represents the question path with the highest questioning efficiency. The optimal path selects A t by choosing the largest Q. Then, the next question is proposed, and the reward function is updated:
[0098]
[0099] Among them, Value represents the value that balances the two dimensions. w 1 represents the weight of the disease accuracy (0.5), represents the accuracy of the disease under this question. Since each dimension is equally important for the final disease probability, the final accuracy is obtained by adding all dimensions. w 2 represents the weight of the questioning efficiency (0.5), and N represents the current number of questions. The highest Value value is selected to control the number of questions. When P t (i) changes less than 0.05 for 3 consecutive rounds (i.e., ΔPt [i] < 0.05) or reaches the maximum number of questions (15 times), the answer to this round of questions ends. When P t [i] < 0.6 (a preset threshold that will be adjusted according to the different personalizations of the interviewees), and after two consecutive rounds of question answering, ΔP t [i] < 0.1 (small change in score), the system is triggered to roll back to the first question of the questions in this dimension and re-ask. After all questions are asked, the score vector P j (j is the big question number, 1 - 12) is recorded after each big question is completed. The average value of the disease probabilities for each big question is calculated to obtain the final probability. The calculation formula is as follows:
[0100]
[0101] Finally, the system is based on P final to provide diagnostic suggestions. For example, if P final > 0.7, it is recommended to further give a warning to the doctor.
[0102] The structured probability model module provides the mathematical model and theoretical basis. The results of the multi-modal sentiment analysis module and the semantic analysis module jointly affect the selection of the dialogue route in the structured psychological interview. The patient characteristics provided by the semantic analysis module, as background information, guide the direction and focus of the interview; while the multi-modal sentiment analysis module monitors the patient's reactions in real time, enabling the interview to be flexibly adjusted to ensure the pertinence and adaptability of the questions. Such an interactive method makes the interview more user-friendly, provides a brand-new diagnostic experience for the patient, and also improves the accuracy of the assessment.
[0103] In the communication with the virtual digital human, when the system model believes that it has obtained an effective answer to a certain scale question, it will automatically jump to the next question. If it is analyzed that the answers to multiple questions are abnormal, it will backtrack until all scale questions are answered. During the whole process, the Q&A content of the virtual digital human is generated by a fine-tuned language large model, ensuring the natural fluency and professionalism of the communication.
[0104] After all scale questions are properly processed, the system will give a global diagnostic result based on the model according to the patient's expressions, voices, and text answers. The answer to each question will contribute a local diagnostic result, and these results will be comprehensively considered by the system together with the diagnostic results given by the scale.
[0105] Finally, the system will generate a detailed analysis report, which not only includes the diagnostic results of the scale, but also comprehensively analyzes the local and global diagnoses, providing a comprehensive and in-depth reference for psychological disease diagnosis for doctors or patients.
[0106] In addition, the design of this system fully considers the privacy protection of patients. All data collection and processing strictly comply with relevant privacy protection regulations to ensure the security of patient information.
[0107] It can be understood that the present invention is described through some embodiments. As is known to those skilled in the art, without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A multimodal psychological problem detection system based on psychological interview, characterized in that: The system includes a digital doctor interaction module, a multimodal sentiment analysis module, a structured probability model, a semantic analysis module, and an additional medical information provision module; The digital doctor interaction module receives the text from the semantic analysis module, synthesizes it into digital human images and sounds, and plays them to interact with patients; The digital doctor interaction module collects audio information through a microphone and video information through a camera, and sends the collected video information and audio information to the multimodal sentiment analysis module; The multimodal sentiment analysis module receives the video and audio signals collected by the digital doctor interaction module, and transcribes them into text signals using speech recognition technology. It then calculates the final prediction result based on the video, audio and text signals as the user sentiment intensity score corresponding to the current question and sends it to the semantic analysis module. The additional medical information providing module collects additional medical information of the patient and injects the additional medical information into the large language model of the semantic analysis module through the prompt word engineering; The semantic analysis module is responsible for generating humanized answers and follow-up questions. It calculates the probability of illness based on the semantic content of the conversation, additional medical information, and the user's emotional intensity score corresponding to the current question sent by the multimodal sentiment analysis module, and inputs the probability of illness into the structured probability model. The structured probability model is modeled through the Markov decision process, which regards the possible answers to each question as state nodes, combines the probability of illness input by the semantic analysis module, and determines the choice of the next question through the state transition probability.
2. The multimodal psychological problem detection system based on psychological interview according to claim 1 is characterized in that: The multimodal sentiment analysis module uses Librosa as an audio feature extractor to extract features from audio information and obtain audio features. It uses OpenFace to perform facial key point detection, head posture estimation, line of sight estimation, and facial action unit recognition on video information to extract video features. Use Whisper speech recognition tool to transcribe audio information into text information; use BERT to extract features from text information to obtain text features; The multimodal sentiment analysis module obtains the final prediction result as the user sentiment intensity score corresponding to the current question by fusing the feature data of the three modalities: video features, audio features, and text features.
3. The multimodal psychological problem detection system based on psychological interview according to claim 1 is characterized in that: The semantic analysis module implements the following functions: Psychological state-aware dialogue engine: It uses a dynamic context-aware mechanism to extract emotional keywords and semantic implicit features in real time based on multiple rounds of dialogue records between users and models on topics related to the MINI scale. Incremental user profile modeling: After each round of conversation, the conversation record is integrated with the user's historical profile data through a large model; Intelligent jump system for scale questions: Based on the current answer content, the relevance score of each question in the MINI scale is calculated, and the question jump is achieved by combining with the structured probability model.
4. The multimodal psychological problem detection system based on psychological interview according to claim 3 is characterized in that: The relevance score for each question in the MINI scale is calculated as follows: Score i =cos(Question i ,text)·Score Score i =0.7·Score i +0.3·Score i-1 Among them, Score i is the score of the i-th question in the scale, Question i represents the i-th question, text represents the conversation record, cos is the calculation function of cosine similarity, which is used to calculate the correlation between the i-th question and the conversation record, and Score represents the preset score benchmark evaluation parameter generated by the large language model through the prompt word engineering. This base integrates the semantic content of the previous i rounds of conversations, user portrait modeling, and the user emotion intensity score corresponding to the current question sent by the multimodal sentiment analysis module. The benchmark evaluation parameter is finally generated by the large language model through the prompt word engineering.
5. The multimodal psychological problem detection system based on psychological interview according to claim 1 is characterized in that: The state of the structured probabilistic model is represented as: S t =(Q t ,H t ,P t ) Among them, Q t is the number of the current major problem (1 to 12), H t is the historical answer sequence, P t ∈[0,1] is a 12-dimensional disease probability vector, and the disease probability of the i-th dimension is expressed as P t (i) The disease probability is calculated by the large language model and sent to the structured probability model. The disease probability will be dynamically updated as the conversation progresses. Define action A t , representing three states of judgment: (1) continue to ask the current question; (2) end the current question; (3) go back to the previous question; the transition probability P(S t+1 |S t ,A t ) indicates that in state S t Next, execute action A t Then transfer to the next state S t+1 This probability is based on historical data and the most initialized set probability, and is dynamically updated during the question-answering process.
Citation Information
Cited By
Intelligent analysis accompanying system for pain symptoms of middle-aged and elderly people
CN120853832A
Psychological interview language interaction structured analysis method and system
CN121528217A