Multi-mode interview mental health screening method, device, equipment, medium and product
Through multimodal interview technology, combined with large language models, human-computer interaction and deep learning algorithms, the user's voice and facial emotional characteristics are identified, which solves the accuracy and efficiency problems of existing mental health screening, improves user experience and screening efficiency, and promotes the popularization of mental health.
Patent Information
- Application Number
- CN202510720845.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
AI Technical Summary
Existing mental health screening methods have problems such as limited accuracy of screening results, long time consumption and poor user experience. This is mainly because they rely on standardized scales and manual interviews, lack interactivity and personalization, and find it difficult to capture users' true emotions and physiological reactions.
A multimodal interview method is adopted, and inquiry statements for natural dialogue scenarios are generated through a large language model. User audio and video data are collected in combination with human-computer interaction technology. Deep learning algorithms are used to identify voice and facial emotion features, and machine learning models are imported for evaluation to generate a mental health screening report.
It improves the accuracy and efficiency of screening results, enhances user experience, reduces subjective bias, and promotes the popularization of mental health and the accessibility of services.
Smart Images

Figure CN120636795A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of human-computer interaction and big data analysis technology, and specifically relates to a multimodal interview mental health screening method, device, equipment, medium and product. Background Art
[0002] Health is a state of complete physical, mental, and social well-being, not simply the absence of illness or infirmity. Therefore, mental health screening is the process of assessing a person's mental health and well-being using scientific psychometric methods.
[0003] Currently, existing mental health screening methods are primarily based on mental health screening scales, which collect user responses through manual interviews or self-administered questionnaires. These scales are then combined with scientific psychometric methods to generate mental health screening results and reports. While this existing approach offers advantages such as standardization (i.e., scales are typically rigorously standardized and validated, resulting in high reliability and validity) and systematicity (i.e., scales contain multiple items or questions, enabling a comprehensive assessment of an individual's mental state), it also presents the following challenges:
[0004] (1) The accuracy of screening results is limited: the filling and scoring of the scale often rely on the individual's subjective feelings and memories, which may have certain subjective biases; and because they mainly rely on standardized scale questionnaires, although these questionnaires are clearly structured and easy to quantify, they often lack interactivity and personalization, which may cause users to feel resistance when filling in the questionnaires, thus affecting the accuracy of the screening results; in addition, scale questionnaires can usually only obtain users' subjective statements, and it is difficult to capture the users' real emotions and physiological reactions when answering;
[0005] (2) The screening process is time-consuming: that is, it may take a long time to complete a complete psychological scale, which may not be suitable for time-sensitive screening scenarios.
[0006] In summary, how to provide a new solution for mental health screening that can combine psychological assessment with artificial intelligence, human-computer interaction, and big data analysis technologies to effectively improve the accuracy and efficiency of screening results, enhance user experience, and promote the popularization of mental health is a topic that technical personnel in this field urgently need to study. Summary of the Invention
[0007] The purpose of the present invention is to provide a multimodal interview mental health screening method, device, computer equipment, computer-readable storage medium and computer program product to solve the problems of limited accuracy of screening results, time-consuming screening process and low screening efficiency in existing mental health screening methods.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] First, a multimodal interview mental health screening method is provided, including:
[0010] Obtaining a mental health screening scale, wherein the mental health screening scale comprises a plurality of multiple-choice questions, each of which comprises a stem and at least two alternative answers;
[0011] For each multiple-choice question in the mental health screening scale, based on the corresponding question content, the first language model is used to randomly generate a corresponding inquiry sentence suitable for a natural conversation scenario;
[0012] For each of the multiple-choice questions, human-computer interaction technology is applied to present a corresponding inquiry statement to the user, and corresponding user response data is collected, wherein the user response data contains synchronized user answer audio data and user facial video data;
[0013] For each of the multiple-choice questions, convert the corresponding user answer audio data into answer text, and apply the second largest language model to identify the corresponding user-selected answer from the corresponding at least two alternative answers based on the answer text;
[0014] For each of the multiple-choice questions, based on the corresponding user answer audio data, a speech emotion recognition technology based on a first deep learning algorithm is applied to identify the corresponding user speech emotion feature;
[0015] For each of the multiple-choice questions, applying facial emotion recognition technology based on a second deep learning algorithm based on the corresponding user facial video data to identify the corresponding user facial emotion features;
[0016] Importing the user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features into a mental health screening assessment model pre-trained based on a machine learning algorithm, and outputting a mental health screening assessment result;
[0017] Based on the mental health screening assessment results, a mental health screening report for the user is generated and output.
[0018] Based on the above invention content, a new solution for mental health screening is provided which combines psychological assessment with artificial intelligence, human-computer interaction and big data analysis technology. That is, first, for each multiple-choice question in the mental health screening scale, according to the corresponding question content, a large language model is applied to randomly generate corresponding inquiry statements suitable for natural dialogue scenarios, and human-computer interaction technology is applied to present the corresponding inquiry statements to the user and collect the corresponding user response data. Then, based on the user response data, multimodal information including the user's selected answers, the user's voice emotion characteristics and the user's facial emotion characteristics is extracted. Finally, the multimodal information is imported into a mental health screening assessment model pre-trained based on a machine learning algorithm, and the mental health screening assessment results are output and a mental health screening report for the user is generated accordingly. This can avoid human subjectivity, shorten the screening process, effectively improve the accuracy of the screening results and the screening efficiency, enhance the user experience and promote the popularization of mental health, and facilitate practical application and promotion.
[0019] In one possible design, for each of the multiple-choice questions, human-computer interaction technology is applied to present corresponding query statements to the user, including:
[0020] Build user-oriented virtual digital humans;
[0021] For each of the multiple-choice questions, the TTS voice broadcast algorithm is applied to convert the corresponding inquiry sentence into an inquiry voice signal suitable for a natural and friendly dialogue scenario, and the inquiry voice signal is presented in a spoken form and / or an action-assisted form by the virtual digital human.
[0022] In one possible design, the first deep learning algorithm adopts an LSTM-based deep learning algorithm, and the user voice emotion features include user speaking speed features, user volume features and / or user voice intonation change features.
[0023] In one possible design, the second deep learning algorithm adopts a deep learning algorithm based on the Transformer architecture, and the user facial emotion features include user expression features, user eye features and / or user smile features.
[0024] In one possible design, the user's selected answers to the multiple-choice questions, the user's voice emotion features, and the user's facial emotion features are imported into a mental health screening assessment model pre-trained based on a machine learning algorithm, and the mental health screening assessment results are output, including:
[0025] For each mental health dimension among the multiple mental health dimensions used for participating in the mental health screening, the user's selected answers to the multiple-choice questions, the user's voice emotion features, and the user's facial emotion features are imported into a corresponding mental health scoring model pre-trained based on the first machine learning algorithm, and a corresponding mental health score value is output;
[0026] And / or, the user's selected answers to the multiple-choice questions, the user's voice emotion features and the user's facial emotion features are imported into a mental health status recognition model pre-trained based on a second machine learning algorithm, and the mental health status recognition results presented in the form of confidence are output.
[0027] In one possible design, after outputting the mental health screening assessment results, the method further includes:
[0028] Based on the mental health screening assessment results, corresponding mental health advice and / or psychology experts are matched, and the matching results are pushed to the user.
[0029] In a second aspect, a multimodal interview mental health screening device is provided, which includes a screening scale acquisition unit, an inquiry statement generation unit, a human-computer interactive dialogue unit, an answer selection determination unit, a voice emotion recognition unit, a facial emotion recognition unit, a psychological screening assessment unit, and a screening report generation unit;
[0030] The screening scale acquisition unit is used to acquire a mental health screening scale, wherein the mental health screening scale includes a plurality of multiple-choice questions, each of which includes a stem and at least two alternative answers;
[0031] The inquiry statement generating unit is communicatively connected to the screening scale acquiring unit, and is configured to apply the first language model to randomly generate corresponding inquiry statements suitable for natural conversation scenarios for each multiple-choice question in the mental health screening scale according to the corresponding question content;
[0032] The human-computer interaction dialogue unit is communicatively connected to the query statement generation unit, and is used to apply human-computer interaction technology to each of the multiple-choice questions, present a corresponding query statement to the user, and collect corresponding user response data, wherein the user response data includes synchronized user answer audio data and user facial video data;
[0033] The selected answer determination unit is communicatively connected to the human-computer interaction dialogue unit and is configured to convert the corresponding user answer audio data into an answer text for each of the multiple-choice questions, and apply the second largest language model to identify the corresponding user selected answer from the corresponding at least two alternative answers based on the answer text;
[0034] The speech emotion recognition unit is communicatively connected to the human-computer interaction dialogue unit and is configured to apply a speech emotion recognition technology based on a first deep learning algorithm to each of the multiple-choice questions based on the corresponding user answer audio data to identify the corresponding user speech emotion feature;
[0035] The facial emotion recognition unit is communicatively connected to the human-computer interaction dialogue unit and is configured to apply a facial emotion recognition technology based on a second deep learning algorithm to each of the multiple-choice questions based on the corresponding user facial video data to identify corresponding user facial emotion features;
[0036] The psychological screening and evaluation unit is communicatively connected to the selected answer determination unit, the voice emotion recognition unit, and the facial emotion recognition unit, respectively, and is used to import the user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features into a psychological health screening and evaluation model pre-trained based on a machine learning algorithm, and output a psychological health screening and evaluation result;
[0037] The screening report generating unit is communicatively connected to the psychological screening evaluation unit, and is used to generate and output the user's psychological health screening report based on the psychological health screening evaluation result.
[0038] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a transceiver communicatively connected in sequence, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the multimodal interview mental health screening method as described in the first aspect or any possible design of the first aspect.
[0039] In a fourth aspect, the present invention provides a computer-readable storage medium having instructions stored thereon. When the instructions are run on a computer, the multimodal interview mental health screening method as described in the first aspect or any possible design of the first aspect is executed.
[0040] In a fifth aspect, the present invention provides a computer program product comprising a computer program or instructions, which, when executed by a computer, implement the multimodal interview mental health screening method as described in the first aspect or any possible design of the first aspect.
[0041] Beneficial effects of the above scheme:
[0042] (1) The present invention creatively provides a new scheme for mental health screening by combining psychological assessment with artificial intelligence, human-computer interaction and big data analysis technology. That is, first, for each multiple-choice question in the mental health screening scale, according to the corresponding question content, a large language model is applied to randomly generate corresponding inquiry statements suitable for natural dialogue scenarios, and human-computer interaction technology is applied to present the corresponding inquiry statements to the user and collect the corresponding user response data. Then, based on the user response data, multimodal information including the user's selected answers, the user's voice emotion characteristics and the user's facial emotion characteristics is extracted. Finally, the multimodal information is imported into a mental health screening assessment model pre-trained based on a machine learning algorithm, and the mental health screening assessment results are output and a mental health screening report for the user is generated based on the results. This can avoid human subjectivity, shorten the screening process, effectively improve the accuracy of the screening results and the screening efficiency, enhance user experience and promote the popularization of mental health.
[0043] (2) It can improve screening accuracy: that is, by combining the scale content with the dialogue form and using multimodal information analysis technology, this solution can more comprehensively capture the user's psychological state, reduce the screening error caused by the user's subjective concealment or misunderstanding of the scale questions, and thus improve the accuracy of screening;
[0044] (3) Enhanced user experience: The natural dialogue form of the virtual digital human makes the screening process more friendly and seamless, reduces the user's psychological pressure, and improves the acceptance and participation of the screening. At the same time, personalized feedback and suggestions also enhance the user's sense of participation and satisfaction;
[0045] (4) Improved screening efficiency: The automated and intelligent screening process reduces manual intervention and improves screening efficiency. In addition, by continuously optimizing the conversation content and evaluation model, the system can continuously adapt to the needs of different user groups, further improving the efficiency and effectiveness of screening.
[0046] (5) It can promote the popularization of mental health: that is, this program provides a convenient and efficient mental health screening tool, which helps to promote the popularization of mental health knowledge and the accessibility of mental health services, contributes to building a harmonious society and improving the mental health level of the whole people, and is convenient for practical application and promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1A flowchart of the multimodal interview mental health screening method provided in an embodiment of the present application.
[0049] Figure 2 A schematic diagram of the structure of a multimodal interview mental health screening device provided in an embodiment of the present application.
[0050] Figure 3 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be briefly introduced below in conjunction with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structures of the drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative work. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0052] It should be understood that although the terms first, second, etc. may be used herein to describe various objects, these objects should not be limited by these terms. These terms are merely used to distinguish one object from another. For example, a first object can be referred to as a second object, and similarly, a second object can be referred to as a first object without departing from the scope of the exemplary embodiments of the present invention.
[0053] It should be understood that the term "and / or" that may appear in this document is merely a description of the association relationship between associated objects, indicating that there may be three relationships. For example, A and / or B can indicate three situations: A exists alone, B exists alone, or A and B exist at the same time. For another example, A, B and / or C can indicate the existence of any one of A, B and C or any combination of them. The term " / and" that may appear in this document describes another type of association object relationship, indicating that there may be two relationships. For example, A / and B can indicate two situations: A exists alone or A and B exist at the same time. In addition, the character " / " that may appear in this document generally indicates that the previous and next associated objects are in an "or" relationship.
[0054] Example
[0055] like Figure 1As shown, the multimodal interview mental health screening method provided in the first aspect of this embodiment can be, but is not limited to, executed by a computer device with certain computing resources, such as a cloud server, a personal computer (PC, a multi-purpose computer with a size, price and performance suitable for personal use; desktops, laptops, small laptops, tablets and ultrabooks are all personal computers), a smart phone, a personal digital assistant (PDA) or a wearable device. Figure 1 As shown, the multimodal interview mental health screening method may include, but is not limited to, the following steps S1 to S8.
[0056] S1. Obtain a mental health screening scale, wherein the mental health screening scale includes but is not limited to multiple choice questions, and the multiple choice questions include but are not limited to question content and at least two alternative answers.
[0057] In step S1, the mental health screening scale is an existing scale questionnaire used in traditional mental health screening methods, including, but not limited to, the following multiple-choice question: "Have you had insomnia recently?" (i.e., as the question stem), "A. Never; B. Rarely; C. Sometimes; D. Often" (i.e., as four alternative answers). In addition, the mental health screening scale can be read conventionally.
[0058] S2. For each multiple-choice question in the mental health screening scale, according to the corresponding question content, the first language model is applied to randomly generate a corresponding inquiry statement suitable for natural conversation scenarios.
[0059] In step S2, the large language model is a deep learning model with billions to hundreds of billions of parameters. It captures language statistical laws from massive texts through unsupervised learning and simulates the human language cognitive process. Its core capabilities include text generation, semantic understanding, and contextual reasoning. Therefore, the first large language model can be applied to randomly generate inquiry statements suitable for natural dialogue scenarios for each question content; for example, for the question content "Have you had insomnia recently?", the following inquiry statements can be randomly generated: "How has your sleep been recently? How often do you have insomnia on average?" or "Have you fallen asleep quickly this month?", etc. (that is, how to ask questions naturally is random). In addition, for example, the first large language model can be, but is not limited to, fine-tuned using a BERT (Bidirectional Encoder Representations from Transformers) series model or a T5 series model (that is, fine-tuned to be suitable for natural dialogue scenarios).
[0060] S3. For each of the multiple-choice questions, human-computer interaction technology is applied to present the corresponding inquiry statement to the user, and the corresponding user response data is collected, wherein the user response data includes but is not limited to synchronized user answer audio data and user facial video data, etc.
[0061] In step S3, the human-computer interaction techniques (HCI) refer to technologies that enable human-computer dialogue in an effective manner through computer input and output devices. Specifically, they may include, but are not limited to, presenting inquiry statements to the user through voice playback, and specifically, but are not limited to, conventionally capturing synchronized audio data of the user's answers and video data of the user's face through audio and video capture devices (e.g., microphones and cameras). To achieve a natural and friendly conversation with the user and avoid psychological pressure caused by directly using the original scale text, HCI techniques are preferably applied to each multiple-choice question to present the corresponding inquiry statement to the user, including but not limited to the following steps S31-S32.
[0062] S31. Build user-oriented virtual digital humans.
[0063] In step S31, the specific process of constructing the virtual digital human is based on existing technologies, preferably aiming for a highly realistic virtual digital human. Furthermore, a matching virtual digital human can be constructed based on the user's personal preferences. For example, if the user prefers a certain type of anime character, that type of anime character can be constructed as a matching virtual digital human.
[0064] S32. For each of the multiple-choice questions, the TTS voice broadcast algorithm is applied to convert the corresponding inquiry sentence into an inquiry voice signal suitable for a natural and friendly dialogue scenario, and the inquiry voice signal is presented in a spoken form and / or an action-assisted form by the virtual digital human.
[0065] In step S32, the TTS (Text-To-Speech) voice broadcast algorithm is an existing algorithm that converts text into sound. Examples of inquiries suitable for natural and friendly conversations include, but are not limited to, a female voice with a certain degree of affinity and magnetic voice. Furthermore, the gesture assistance may include, but is not limited to, preset candid gestures, proactive waving, and / or smiling.
[0066] S4. For each of the multiple-choice questions, the corresponding user answer audio data is converted into an answer text, and based on the answer text, the second largest language model is applied to identify the corresponding user-selected answer from the corresponding at least two alternative answers.
[0067] In step S4, the specific process of text conversion can be conventionally implemented through existing speech-to-text recognition technology. Since the large language model also has semantic understanding and contextual reasoning capabilities, the second large language model can be applied to identify the corresponding user-selected answer from the corresponding at least two alternative answers based on the answer text; for example, for the query sentence: "How is your sleep recently? How often do you suffer from insomnia on average?", if the answer text is "Once last week", then "C. Sometimes" can be used as the user-selected answer; for another example, for the query sentence: "Have you fallen asleep quickly this month?", if the answer text is "Basically no problem", then "B. Rarely" can be used as the user-selected answer. In addition, the second large language model and the first large language model may be the same model or not.
[0068] S5. For each of the multiple-choice questions, according to the corresponding user answer audio data, apply the speech emotion recognition technology based on the first deep learning algorithm to identify the corresponding user speech emotion characteristics.
[0069] In step S5, deep learning specifically refers to machine learning based on deep neural network models and methods; specifically, the first deep learning algorithm can be but is not limited to a deep learning algorithm based on LSTM (Long Short-Term Memory), and the user voice emotion features include but are not limited to user speaking speed features, user volume features and / or user voice tone change features, etc. These features are used to reflect the user's true emotions and psychological state from the voice dimension, and can be obtained based on the conventional capture and analysis of the LSTM model.
[0070] S6. For each of the multiple-choice questions, according to the corresponding user facial video data, apply the facial emotion recognition technology based on the second deep learning algorithm to identify the corresponding user facial emotion features.
[0071] In step S6, specifically, the second deep learning algorithm can be but is not limited to a deep learning algorithm based on the Transformer architecture, and the user's facial emotional features include but are not limited to user expression features, user eye features and / or user smile features, etc. These features are used to reflect the user's true emotions and psychological state from the dimension of facial muscle changes, and can be obtained based on the conventional capture and analysis of the Transformer model.
[0072] S7. Import the user's selected answers to the multiple-choice questions, the user's voice emotion features, and the user's facial emotion features into a mental health screening assessment model pre-trained based on a machine learning algorithm, and output the mental health screening assessment results.
[0073] In step S7, the user's selected answers to each multiple-choice question, the user's voice emotion characteristics, and the user's facial emotion characteristics are used as multimodal information to be imported into the mental health screening assessment model as model input items. The machine learning algorithm is a core artificial intelligence algorithm that specifically studies how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. It is the fundamental way to make computers intelligent. Specifically, the machine learning algorithm preferably adopts, but is not limited to, machine learning algorithms based on graph neural networks, support vector machines, K-nearest neighbor methods, stochastic gradient descent methods, multivariate linear regression, multilayer perceptrons, decision trees, backpropagation neural networks, or radial basis function networks to quickly and accurately identify patterns in data. Therefore, based on a certain amount of sample data (i.e., the model input is multimodal information, and the model output is the mental health screening assessment label corresponding to the multimodal information), the mental health screening assessment model that has passed the verification can be trained through conventional calibration and verification modeling methods (the specific process includes the calibration process and verification process of the model, that is, first comparing the model simulation results with the measured data, and then adjusting the model parameters according to the comparison results so that the simulation results are consistent with the actual process).
[0074] In step S7, in order to achieve the purpose of enriching the evaluation indicators to accurately evaluate the user's mental state, preferably, the user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features are imported into a mental health screening evaluation model pre-trained based on a machine learning algorithm, and a mental health screening evaluation result is output, including but not limited to: for each mental health dimension in the multiple mental health dimensions used for participating in the mental health screening, the user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features are imported into a corresponding mental health scoring model pre-trained based on a first machine learning algorithm, and the corresponding mental health score value is output; and / or, the user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features are imported into a mental health state recognition model pre-trained based on a second machine learning algorithm, and the mental health state recognition result is output in the form of confidence. The aforementioned multiple mental health dimensions include but are not limited to mental health dimensions such as anxiety, depression, and / or stress; the mental health scoring model and the mental health state recognition model are specific forms of the mental health screening evaluation model, and can also be obtained through conventional training based on a certain amount of corresponding sample data through conventional calibration and validation modeling. In addition, the first machine learning algorithm or the second machine learning algorithm preferably adopts, but is not limited to, a machine learning algorithm based on graph neural network, support vector machine, K nearest neighbor method, stochastic gradient descent method, multivariate linear regression, multilayer perceptron, decision tree, back propagation neural network or radial basis function network.
[0075] S8. Generate and output the user's mental health screening report based on the mental health screening assessment results.
[0076] In step S8, the specific generation process of the mental health screening report can be conventionally implemented using existing technologies, for example, the mental health score values in mental health dimensions such as anxiety, depression and / or stress and / or the probability of determining that there are mental health problems / risks (i.e., the confidence level of the mental health status identification results) are automatically filled in the mental health screening report, so as to obtain a personalized feedback report for the user. At the same time, in order to achieve the ability to provide corresponding mental health advice or guide users to seek professional help, after outputting the mental health screening assessment results, the method also includes but is not limited to: matching corresponding mental health advice and / or psychology experts based on the mental health screening assessment results (i.e., the system has some professional psychology experts with expertise in different fields, and can match the most suitable experts), and pushing the matching results (specifically the matched mental health advice and / or psychology expert recommendation information, etc.) to the user. The specific process of the aforementioned matching may be, but is not limited to: taking the mental health score values and / or the probability of judging the presence of mental health problems / risks in mental health dimensions such as anxiety, depression and / or stress as a vector, and then calculating the distances to different mental health advice and / or psychology experts based on the vector, and finally taking the mental health advice and / or psychology expert corresponding to the shortest distance as the matching result. In addition, the conversation content and evaluation model can be continuously optimized based on user feedback (for example, during the conversation, the user's satisfaction can be collected. If the user is dissatisfied, the case is recorded and used as a training sample to optimize the first language model) to improve the accuracy of screening and user experience.
[0077] Therefore, based on the multimodal interview mental health screening method described in the aforementioned steps S1 to S8, a new scheme for mental health screening is provided that combines psychological assessment with artificial intelligence, human-computer interaction and big data analysis technology. That is, first, for each multiple-choice question in the mental health screening scale, according to the corresponding question content, a large language model is applied to randomly generate corresponding inquiry statements suitable for natural dialogue scenarios, and human-computer interaction technology is applied to present the corresponding inquiry statements to the user and collect the corresponding user response data. Then, based on the user response data, multimodal information including the user's selected answers, the user's voice emotion characteristics and the user's facial emotion characteristics is extracted. Finally, the multimodal information is imported into a mental health screening assessment model pre-trained based on a machine learning algorithm, and the mental health screening assessment results are output and a mental health screening report for the user is generated based on this. This can avoid human subjectivity, shorten the screening process, effectively improve the accuracy of the screening results and the screening efficiency, enhance the user experience and promote the popularization of mental health, and facilitate practical application and promotion.
[0078] like Figure 2As shown, the second aspect of this embodiment provides a virtual device for implementing the multimodal interview mental health screening method described in the first aspect, comprising a screening scale acquisition unit, an inquiry statement generation unit, a human-computer interaction dialogue unit, an answer selection determination unit, a voice emotion recognition unit, a facial emotion recognition unit, a psychological screening assessment unit, and a screening report generation unit;
[0079] The screening scale acquisition unit is used to acquire a mental health screening scale, wherein the mental health screening scale includes a plurality of multiple-choice questions, each of which includes a stem and at least two alternative answers;
[0080] The inquiry statement generating unit is communicatively connected to the screening scale acquiring unit, and is configured to apply the first language model to randomly generate corresponding inquiry statements suitable for natural conversation scenarios for each multiple-choice question in the mental health screening scale according to the corresponding question content;
[0081] The human-computer interaction dialogue unit is communicatively connected to the query statement generation unit, and is used to apply human-computer interaction technology to each of the multiple-choice questions, present a corresponding query statement to the user, and collect corresponding user response data, wherein the user response data includes synchronized user answer audio data and user facial video data;
[0082] The selected answer determination unit is communicatively connected to the human-computer interaction dialogue unit and is configured to convert the corresponding user answer audio data into an answer text for each of the multiple-choice questions, and apply the second largest language model to identify the corresponding user selected answer from the corresponding at least two alternative answers based on the answer text;
[0083] The speech emotion recognition unit is communicatively connected to the human-computer interaction dialogue unit and is configured to apply a speech emotion recognition technology based on a first deep learning algorithm to each of the multiple-choice questions based on the corresponding user answer audio data to identify the corresponding user speech emotion feature;
[0084] The facial emotion recognition unit is communicatively connected to the human-computer interaction dialogue unit and is configured to apply a facial emotion recognition technology based on a second deep learning algorithm to each of the multiple-choice questions based on the corresponding user facial video data to identify corresponding user facial emotion features;
[0085] The psychological screening and evaluation unit is communicatively connected to the selected answer determination unit, the voice emotion recognition unit, and the facial emotion recognition unit, respectively, and is used to import the user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features into a psychological health screening and evaluation model pre-trained based on a machine learning algorithm, and output a psychological health screening and evaluation result;
[0086] The screening report generating unit is communicatively connected to the psychological screening evaluation unit, and is used to generate and output the user's psychological health screening report based on the psychological health screening evaluation result.
[0087] The working process, working details and technical effects of the aforementioned device provided in the second aspect of this embodiment can be found in the multimodal interview mental health screening method described in the first aspect, and will not be described in detail here.
[0088] like Figure 3 As shown, the third aspect of this embodiment provides a computer device for executing the multimodal interview mental health screening method as described in the first aspect, including a memory, a processor and a transceiver that are sequentially connected in communication, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the multimodal interview mental health screening method as described in the first aspect. For example, the memory may include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, a first-in-first-out memory (FIFO) and / or a first-in-last-out memory (FILO), etc.; the processor may include, but is not limited to, a microprocessor of the STM32F105 series. In addition, the computer device may also include, but is not limited to, a power module, a display screen and other necessary components.
[0089] The working process, working details and technical effects of the aforementioned computer device provided in the third aspect of this embodiment can be found in the multimodal interview mental health screening method described in the first aspect, and will not be described in detail here.
[0090] A fourth aspect of this embodiment provides a computer-readable storage medium storing instructions containing the multimodal interview mental health screening method described in the first aspect. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, execute the multimodal interview mental health screening method described in the first aspect. The computer-readable storage medium refers to a data storage medium, which may include, but is not limited to, a floppy disk, a CD, a hard disk, a flash memory, a USB flash drive, and / or a memory stick. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device.
[0091] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment can be found in the multimodal interview mental health screening method described in the first aspect, and will not be repeated here.
[0092] A fifth aspect of this embodiment provides a computer program product, including a computer program or instructions, which, when executed by a computer, implements the multimodal interview mental health screening method described in the first aspect. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0093] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A multimodal interview mental health screening method, characterized in that: include: Obtaining a mental health screening scale, wherein the mental health screening scale comprises a plurality of multiple-choice questions, each of which comprises a stem and at least two alternative answers; For each multiple-choice question in the mental health screening scale, based on the corresponding question content, the first language model is used to randomly generate a corresponding inquiry sentence suitable for a natural conversation scenario; For each of the multiple-choice questions, human-computer interaction technology is applied to present a corresponding inquiry statement to the user, and corresponding user response data is collected, wherein the user response data contains synchronized user answer audio data and user facial video data; For each of the multiple-choice questions, convert the corresponding user answer audio data into answer text, and apply the second largest language model to identify the corresponding user-selected answer from the corresponding at least two alternative answers based on the answer text; For each of the multiple-choice questions, based on the corresponding user answer audio data, a speech emotion recognition technology based on a first deep learning algorithm is applied to identify the corresponding user speech emotion feature; For each of the multiple-choice questions, applying facial emotion recognition technology based on a second deep learning algorithm based on the corresponding user facial video data to identify the corresponding user facial emotion features; Importing the user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features into a mental health screening assessment model pre-trained based on a machine learning algorithm, and outputting a mental health screening assessment result; Based on the mental health screening assessment results, a mental health screening report for the user is generated and output.
2. The multimodal interview mental health screening method according to claim 1, characterized in that: For each of the multiple-choice questions, human-computer interaction technology is applied to present corresponding query statements to the user, including: Build user-oriented virtual digital humans; For each of the multiple-choice questions, the TTS voice broadcast algorithm is applied to convert the corresponding inquiry sentence into an inquiry voice signal suitable for a natural and friendly dialogue scenario, and the inquiry voice signal is presented in a spoken form and / or an action-assisted form by the virtual digital human.
3. The multimodal interview mental health screening method according to claim 1, characterized in that: The first deep learning algorithm adopts an LSTM-based deep learning algorithm, and the user voice emotion features include user speaking speed features, user volume features and / or user voice intonation change features.
4. The multimodal interview mental health screening method according to claim 1, characterized in that: The second deep learning algorithm adopts a deep learning algorithm based on the Transformer architecture, and the user facial emotion features include user expression features, user eye features and / or user smile features.
5. The multimodal interview mental health screening method according to claim 1, characterized in that: The user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features are imported into a mental health screening assessment model pre-trained based on a machine learning algorithm, and the mental health screening assessment results are output, including: For each mental health dimension among the multiple mental health dimensions used for participating in the mental health screening, the user's selected answers to the multiple-choice questions, the user's voice emotion features, and the user's facial emotion features are imported into a corresponding mental health scoring model pre-trained based on the first machine learning algorithm, and a corresponding mental health score value is output; And / or, the user's selected answers to the multiple-choice questions, the user's voice emotion features and the user's facial emotion features are imported into a mental health status recognition model pre-trained based on a second machine learning algorithm, and the mental health status recognition results presented in the form of confidence are output.
6. The multimodal interview mental health screening method according to claim 1, characterized in that: After outputting the mental health screening assessment results, the method further includes: Based on the mental health screening assessment results, corresponding mental health advice and / or psychology experts are matched, and the matching results are pushed to the user.
7. A multimodal interview mental health screening device, characterized in that: It includes a screening scale acquisition unit, an inquiry statement generation unit, a human-computer interaction dialogue unit, an answer selection determination unit, a voice emotion recognition unit, a facial emotion recognition unit, a psychological screening assessment unit, and a screening report generation unit; The screening scale acquisition unit is used to acquire a mental health screening scale, wherein the mental health screening scale includes a plurality of multiple-choice questions, each of which includes a stem and at least two alternative answers; The inquiry statement generating unit is communicatively connected to the screening scale acquiring unit, and is configured to apply the first language model to randomly generate corresponding inquiry statements suitable for natural conversation scenarios for each multiple-choice question in the mental health screening scale according to the corresponding question content; The human-computer interaction dialogue unit is communicatively connected to the query statement generation unit, and is used to apply human-computer interaction technology to each of the multiple-choice questions, present a corresponding query statement to the user, and collect corresponding user response data, wherein the user response data includes synchronized user answer audio data and user facial video data; The selected answer determination unit is communicatively connected to the human-computer interaction dialogue unit and is configured to convert the corresponding user answer audio data into an answer text for each of the multiple-choice questions, and apply the second largest language model to identify the corresponding user selected answer from the corresponding at least two alternative answers based on the answer text; The speech emotion recognition unit is communicatively connected to the human-computer interaction dialogue unit and is configured to apply a speech emotion recognition technology based on a first deep learning algorithm to each of the multiple-choice questions based on the corresponding user answer audio data to identify the corresponding user speech emotion feature; The facial emotion recognition unit is communicatively connected to the human-computer interaction dialogue unit and is configured to apply a facial emotion recognition technology based on a second deep learning algorithm to each of the multiple-choice questions based on the corresponding user facial video data to identify corresponding user facial emotion features; The psychological screening and evaluation unit is communicatively connected to the selected answer determination unit, the voice emotion recognition unit, and the facial emotion recognition unit, respectively, and is used to import the user's selected answers to each multiple-choice question, the user's voice emotion features, and the user's facial emotion features into a psychological health screening and evaluation model pre-trained based on a machine learning algorithm, and output a psychological health screening and evaluation result; The screening report generating unit is communicatively connected to the psychological screening evaluation unit, and is used to generate and output the user's psychological health screening report based on the psychological health screening evaluation result.
8. A computer device, characterized in that: It includes a memory, a processor and a transceiver that are communicatively connected in sequence, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the multimodal interview mental health screening method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are run on a computer, the multimodal interview mental health screening method as described in any one of claims 1 to 6 is executed.
10. A computer program product comprising a computer program or instructions, characterized in that When executed by a computer, the computer program or the instructions implement the multimodal interview mental health screening method as described in any one of claims 1 to 6.