Explainable machine learning human interaction methods, systems, media, and devices
By optimizing the response quality of generative AI through human-computer interaction methods based on interpretable machine learning, the problems of inconsistent responses and insufficient prompting engineering in personalized and immediate support of mental health chatbots are solved, and efficient mental health services that match or even surpass those of human counselors are achieved.
Patent Information
- Application Number
- CN202311213360.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-19
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-09-19
AI Technical Summary
Existing mental health chatbots suffer from inconsistent response quality, lack of effective feedback mechanisms, and insufficient prompting engineering in providing personalized and immediate psychological support. In particular, they struggle to understand systems, cultures, and languages when dealing with professional fields, resulting in insufficient adaptability and stability in real-world scenarios.
This study employs a human-computer interaction approach based on interpretable machine learning. By acquiring and analyzing language cues in mental health counseling data, and utilizing feature sensitivity analysis and cue engineering, the response quality of generative AI is optimized. This includes language pattern feature extraction and prediction models, and the design of targeted cue strategies to improve AI performance in mental health question-and-answer scenarios.
It significantly improves the response quality of generative AI in mental health Q&A, making it match or even surpass the performance of human counselors in some aspects. It solves the problems of adaptability and stability of generative AI in real-world scenarios, provides efficient and accurate mental health services, reduces costs and expands service coverage.
Smart Images

Figure CN117252265B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of human-computer interaction, and particularly relates to a human-computer interaction method, system, medium and equipment based on interpretable machine learning. BACKGROUND
[0002] For a long time, there has been a significant shortage of utilization of mental health services, especially in low- and middle-income countries. Although online counseling is gradually adopted by the public, the services provided by professional counselors still cannot meet the increasing demand for psychological treatment, resulting in a lack of timely and effective support for many people in need of psychological assistance.
[0003] In recent years, to solve this problem, providing personalized and on-demand mental health counseling agents is considered to potentially solve this dilemma. However, especially considering the significant progress of generative AI technology like ChatGPT, whether "AI-based counseling agent technology" can effectively support in such a dilemma still needs further technical exploration and verification.
[0004] Although artificial intelligence (AI) has been widely integrated in various fields of automation and daily life, it is still cautious to directly apply it to real mental health problems. These reservations mainly stem from two different concerns: the effectiveness of AI in providing services and its ability to meet the qualifications required to provide such services. Among these issues, the inherent ability of AI, such as communication ability, is a major prerequisite. For the application of AI in mental health, we should abandon the dichotomy of either-or. Because some mental health problems do not require specialized treatment, but only some emotional support and direct advice. In this context, the usefulness of mental health chat robots is considered an indispensable asset.
[0005] The advent of generative artificial intelligence (particularly ChatGPT) technology marks a significant breakthrough in the field of AI-based chatbots. Research has shown that ChatGPT has the ability to pass medical exams. With the growing demand for mental health services and the increasing capabilities of artificial intelligence in recent years, the development of digital mental health interventions (DMHIs) has been driven. DMHIs are technology-driven tools and services, usually driven by AI, designed to help diagnose, treat, and manage mental health conditions. One of the interventions is chatbots, which are AI-based computer programs that can engage in conversation with people. Chatbots have been integrated into DMHIs to support diagnosis and screening, symptom management and behavior modification, and service delivery. In addition, some conversational agents or chatbots are designed as emotionally intelligent machines and establish therapeutic alliances with users. However, concerns about the accuracy of information, the immaturity of technology, the violation of ethical frameworks, and the questioning of human authenticity have limited the application of chatbots in the field of mental health. Currently, some of these issues have been addressed by related technological advances, and the potential of chatbots in the field of mental health has received increasing attention.
[0006] The application of chatbots in the field of mental health services does have some advantages. First, chatbots can provide anytime, anywhere, personalized, low-threshold, and self-managed mental support and treatment. Second, chatbots can replace traditional face-to-face therapy in some cases. For example, those in need of treatment may not be able to receive treatment quickly due to the shortage of mental health professionals, especially in rural and low-income areas. Third, chatbots can provide a relatively anonymous and private form of therapy, making it easier for patients to accept (because of the fear of stigmatization). In fact, some studies have shown that some people may actually prefer to interact with chatbots rather than mental health professionals, which may encourage those who usually do not seek treatment to receive treatment. In addition, chatbots are considered to be less biased than humans, which can promote self-disclosure and allow greater conversational flexibility.
[0007] ChatGPT is a generative language model tool launched by OpenAI that enables the public to have a conversation with a machine on a wide range of topics. Compared to past chatbots, the most impressive enhancement of ChatGPT is its contextual understanding and the ability to allow for subsequent corrections. Currently, ChatGPT can accurately understand what people input (questions need to be logically coherent in language) and output what people want to know (except for unethical or illegal content), which corresponds to the most critical feature of a mental health domain chat agent, that is, the interactive tool is generated in natural language without constraints, rather than pre-defined content. Logically, the capabilities of ChatGPT can meet these requirements. Some researchers highlight the reasonable reasoning ability and the provision of effective clinical insights exhibited by ChatGPT, thereby enhancing trust and interpretability. However, they remain cautious about the impact ChatGPT can have on healthcare in terms of preserving medical expertise.
[0008] In fact, the quality of ChatGPT's responses depends largely on the input prompts provided by the user, which is often referred to as prompt engineering. Prompt engineering is an increasingly important skill for effective conversations with large language models like ChatGPT. Prompts are instructions given to large language models to perform rules, automate processes, and ensure a specific quality (and quantity) of generated output.
[0009] While psychotherapy and social support have been proven to be effective as valuable therapeutic modalities, many vulnerable individuals encounter significant barriers that limit their access to therapy and counseling services. One possible solution is to shift from the traditional one-on-one counseling model to a one-to-many model to maximize counselor resources without sacrificing the quality of services. Text-based counseling, especially online counseling, provides a flexible and convenient platform for mental health professionals and patients to communicate. This approach can effectively meet the privacy and convenience needs of individuals while reducing physical space and time constraints. In the past few years, text-based counseling has become increasingly popular, especially among the younger generation. One of the reasons for this is that this form can eliminate or reduce the stigma associated with psychotherapy. In addition, online counseling provides a valuable opportunity for those living in areas far from mental health resources that may be geographically inaccessible. While text-based counseling provides convenience for many people, it also presents some new challenges. For example, text-based communication can result in the loss of emotional information as it lacks non-verbal cues. In addition, due to the anonymity of the online environment, ensuring the safety and confidentiality of counseling becomes more complex.
[0010] However, for some patients, text-based counseling can be the best option available to them. Especially considering its ability to provide timely feedback and wide accessibility, this makes mental health chatbots, such as ChatGPT, have a potential field of application. One significant advantage of text-based interventions is their accessibility and immediacy. Users can prefer immediate support when seeking support, even if this support is provided by trained volunteers or humans. Whereas the previous online mental health community interaction model was asynchronous. Therefore, in this case, clarifying the impact of chatbots, such as ChatGPT, on the effectiveness of mental health interventions can have a significant impact on reducing costs and increasing the acceptance of mental health services.
[0011] Natural language processing (NLP) is a subset of artificial intelligence that enables machines to learn the structure of data implicitly from written and spoken language. Due to the clinical relevance and ease of recording of human language in mental health counseling, NLP has become a key tool with the potential to provide therapeutic feedback for both clinicians and patients in psychotherapy. Linguistic Inquiry and Word Count (LIWC) can be classified as a tool that applies NLP. LIWC works by reporting the percentage of words in a text file that belong to its dictionary-defined grammatical, psychological, and content categories. In addition to classifying and counting words, LIWC provides a statistical overview of word usage for predefined and psychologically meaningful categories. The functionality of LIWC and other similar programs or algorithms has enabled researchers to begin studying language use patterns between individuals with depression and other mental health conditions.
[0012] One relatively common application of NLP in text-based counseling is the automatic detection of therapeutic alliance. Therapeutic alliance is an important predictor of the outcome of psychotherapy, including online text-based counseling. However, the current gold standard for alliance assessment is self-report and qualitative coding of observer-based therapeutic interactions, which are labor-intensive and time-consuming. The development of LIWC technology and psychotherapy language cues provides a foundation for automatically identifying therapeutic alliance. Some existing research has explored language cues for therapeutic alliance. For example, the frequency of the therapist's first-person pronoun use and the patient's syntactic transformation are potential indicators of therapeutic alliance in speech. The level of language style matching (LSM), defined as the degree of similarity in the functional word rate in a dialogue interaction, is considered to reflect the extent to which conversation partners automatically coordinate language style to achieve common goals between the patient and the mentor, and is an effective predictor of therapeutic alliance. In addition, affective words, discrepancy words, and language synchrony have also been shown to be language cues for identifying therapeutic alliance. There are also experiments that found that weaker alliances are characterized by negation and reflection words.
[0013] While language models like ChatGPT show great potential for conversational interactions, their effectiveness on real mental health issues has not been fully validated and there is a risk of generating low-quality or inconsistent responses with human experts. The limitations of the model in understanding the actual problem in depth may result in inconsistent quality of responses, just fluent and grammatically correct does not mean truly beneficial or relevant. The existing background technology has a series of shortcomings in mental health applications: there is often a lack of effective feedback loop in current systems, which hinders the continuous improvement of output quality. Prompt engineering also faces major challenges: how to accurately provide prompts that help the model generate high-quality answers? There is no clear and efficient method. Especially in the face of professional fields with less training data for large language models, the model may encounter difficulties in understanding institutions, culture and language. Even if the model may perform well in experimental environments, its adaptability and stability in real-world scenarios remain an important consideration. SUMMARY
[0014] The present invention addresses the shortcomings of existing mental health question and answer technology, and proposes a new human-computer interaction method, system, medium and device that integrates natural language processing, interpretable machine learning and prompt engineering. This method aims to optimize the performance of generative AI in mental health question and answer scenarios by in-depth analysis of language clues, to better meet the needs of users seeking psychological counseling.
[0015] The present invention is implemented as an interpretable machine learning-based human-computer interaction method for evaluating, understanding and optimizing the quality and user experience of a dialogue system, which can provide convenient and effective data support for the human-computer interaction process of the dialogue system. The method includes: obtaining mental health counseling data of generative artificial intelligence; identifying psychological linguistics clues, topic consistency, language style matching and emotional similarity language patterns from the question and answer of the seeker in the counseling data and the interaction process text between them. Obtain effective language pattern features for regression prediction. Use feature sensitivity analysis to refine the prediction model and identify the interpretive analysis results of the feature impact on the effectiveness of psychological counseling; extract operable language model features to develop targeted prompts. Give the targeted prompts and seeker questions to the generative artificial intelligence model to obtain optimized responses.
[0016] Further, comprising:
[0017] Using prediction and comparative analysis methods, obtain and predict the quality score of mental health question and answer of generative AI technology, then compare it with the answers of human experts, and analyze the average difference between the two.
[0018] Using language clue analysis methods, identify key language clues that can affect perceived helpfulness in question and answer.
[0019] Using the prompt engineering strategy generation method, according to the analysis results, a targeted prompt strategy is generated for the generative AI, thereby improving its performance in the mental health Q&A scenario.
[0020] Further, the main challenge and difficulty of the data set based on the interpretable machine learning human-computer interaction method for predicting the perceptual utility of artificial intelligence in mental health Q&A is to obtain mental health Q&A data with labels; the mental health Q&A text in the data set is preprocessed, and the word segmentation, stop word removal, and keyword extraction tasks are completed; the toolkit is used to segment the mental health Q&A text, a mental field dictionary is constructed using Gensim, and the embedded representation of the mental health theme is derived.
[0021] The Chinese complex emotion analysis tool is used to identify 7 emotion statistics from the Q&A text, namely, positive, joy, sadness, anger, fear, disgust, and surprise; the cLIWC is used to extract statistical features of psychological language clues from the Q&A text; the language survey and Word count are used to calculate the theme consistency, language style similarity, and emotional similarity in the Q&A text, and describe the therapeutic alliance in the mental health Q&A text; 20 groups of mental health-related Q&A texts are manually checked, and all extracted feature values are verified.
[0022] Further, the interpretable machine learning human-computer interaction method based on the interpretable machine learning method for large-scale Q&A discourse analysis uses computational linguistics, and uses a feature set of the questioner's question, the counselor's answer, and the synchronous interaction between the two to construct a pre-explanation model; predict and explain the perceptual helpfulness of mental health Q&A; the feature set includes LIWC, theme consistency, language style, and emotional similarity features;
[0023] LIWC captures the psychological dimensions of user language use, measures and captures the language behavior of visitors and counselors, including language clues such as emotional expression, mental state, social relationship, cognitive process, and vocabulary features; theme consistency is a feature based on text content, reflecting the consistency of the theme in the text;
[0024] Language style matching is a feature based on language style, which measures the similarity of language style in the text; emotional similarity reflects whether the emotions of the visitor and the counselor in the conversation are consistent, and the feature is extracted using an emotion analysis method.
[0025] Further, the human-computer interaction method based on interpretable machine learning utilizes the perceived helpfulness as the output target of the prediction model, and uses the language exploration and word count of the help-seeker and the counselor, as well as the therapeutic alliance features, as inputs, to propose an interpretable prediction model using interpretable machine learning techniques.
[0026] Further, five commonly used regression algorithms are selected, including ridge regression, least absolute shrinkage and selection operator, support vector regression, random forest and eXtreme gradient boosting; ridge regression and lasso are regression techniques for handling collinearity and overfitting; ridge regression restricts model coefficients through a penalty term, achieving automatic variable selection and enhancing model interpretability; support vector regression is an application of support vector machines, handling non-linear and high-dimensional data problems, and has been used for perceived helpfulness prediction; the ensemble method improves prediction accuracy and prevents overfitting by combining multiple decision trees and applying regularization terms.
[0027] Further, the evaluation indicators of the human-computer interaction method based on interpretable machine learning include root mean square error RMSE and mean absolute percentage error MAPE, defined by formulas (1)-(2); root mean square error RMSE is a measure of the difference between predicted and actual values; mean absolute percentage error MAPE is a measure method that quantifies prediction accuracy in percentage;
[0028]
[0029]
[0030] where n represents the total number of observations in the mental health question and answer dataset, y i represents the actual perceived help of the i-th observation in the dataset i is the predicted perceived helpfulness of the i-th observation output by the regression model.
[0031] Another object of the present application is to provide a computer device comprising a memory and a processor, the memory storing a computer program, the computer program being executed by the processor to cause the processor to execute the human-computer interaction method based on interpretable machine learning.
[0032] Another object of the present application is to provide a computer-readable storage medium storing a computer program, the computer program being executed by a processor to cause the processor to execute the human-computer interaction method based on interpretable machine learning.
[0033] Another object of the present application is to provide an information data processing terminal for self-help mental health services, which is used to implement the human-computer interaction method based on interpretable machine learning.
[0034] It is another object of the present application to provide an explainable machine learning-based human-computer interaction system, which comprises:
[0035] An information analysis module is responsible for predicting mental health questions and answers, and comparing them with human consultants, thereby analyzing the differences between the generative AI and human consultants.
[0036] A clue determination module aims to determine the key language clues that affect the perceived helpfulness of mental health questions and answers, and analyze how these clues affect the perceived helpfulness.
[0037] A strategy prompt module is responsible for designing specific prompt strategies for generative AI, through which the performance of AI in the mental health question and answer scenario is significantly improved.
[0038] In combination with the above technical solutions and the technical problems solved, the technical solutions protected by the present application have the following advantages and positive effects:
[0039] First, the present application tests the response quality of generative AI for real-world mental health problems. In order to optimize its response, we use an explainable machine learning (XML) based target prompt engineering method to effectively improve the response quality of generative AI. The present application constructs a dataset including a large number of mental health related questions and answers, and then uses machine learning methods to establish a model to evaluate the perceived helpfulness of generative AI. Through these steps, we find that the optimized generative AI can be on par with human consultants in answering certain mental health questions, and even exceed human consultants in some aspects.
[0040] Second, from the perspective of the product as a whole, the technical solution proposed by the invention has the following technical effects and advantages: The invention proposes and verifies a new method of target prompting engineering based on interpretable machine learning, which significantly improves the performance of generative AI. We use advanced XML technology to compare and analyze the responses of generative AI and real humans in mental health question and answer. This unique method allows us to accurately identify the differences between the spontaneous responses of generative AI and real human responses. Although large prophetic models have less training data in specific professional fields, using interpretable machine learning to identify these differences not only provides us with in-depth insights into AI responses, but also provides clear guidelines for further optimizing generative AI. Therefore, the technical solution of the invention successfully addresses the difficulties of generative AI in understanding institutions, culture and language, enabling the model to produce high-quality responses when facing these challenges. Therefore, this technical solution not only improves the response quality of generative AI in mental health question and answer, but also provides valuable reference for future artificial intelligence development and application, making it closer to real human responses and perceptions, thereby achieving more efficient, accurate and beneficial user experience.
[0041] Third, in order to continuously optimize the output quality of generative AI, the invention builds a dataset including a large number of mental health-related questions and answers, and provides an effective feedback loop mechanism. This mechanism can be continuously optimized according to actual application scenarios and user feedback, effectively solving the problems in the background technology.
[0042] More importantly, we have targetedly solved a series of problems in the background technology. First, for adaptability and stability in real scenarios, the invention not only performs well in experimental environments, but also significantly improves its adaptability and stability in real scenarios. Second, the invention provides strong support for the response ability of generative AI in human subjective emotional expression. When facing a wide range of mental health problems, its answer quality can match the average level of human consultants, and even surpass them in some topics, proving that the invention not only performs well in experimental environments, but also significantly improves its adaptability and stability in real scenarios. These results enhance the confidence in the contribution potential of generative AI in the field of human mental health.
[0043] Fourth, as the invention's creative auxiliary evidence, it is also reflected in the following important aspects:
[0044] (1) The expected benefits and commercial value of the technical solution of the present application after transformation are: the demand for psychological health consultation continues to grow, but due to the shortage of professional consultants, the supply is facing great challenges. The technical solution of the present application provides users with high-quality responses similar to humans by using generative AI, especially ChatGPT. Since ChatGPT performs even better than human experts on low-risk and no-risk issues, it means that it has great commercial potential, both greatly reducing costs and covering a wide range of demand groups. The expected benefits are to provide more extensive, more economical and more immediate psychological health services, and the commercial value lies in attracting more users, providing continuous services, and expanding to other application scenarios.
[0045] (2) The technical solution of the present application fills the technical gap in the industry at home and abroad: although the application of generative AI is welcomed in many fields, its in-depth application in the field of psychological health is still cautious. The existing technology only stays at producing fluent and grammatically correct responses, but does not delve into how to ensure its practicality, reliability and consistency with human experts. The present application fills this gap by in-depth research and verification, providing a new, efficient and effective solution for generative AI in psychological health.
[0046] (3) Does the technical solution of the present application solve the technical problems that people have long been eager to solve but have always failed to succeed: the present application solves two long-standing problems: how to provide real-time and effective psychological health consultation services for a wide range of people, and how to ensure that the response based on generative AI is not only fluent and grammatically correct, but also real, accurate and of practical value. Through the establishment of a prediction model, the prompt engineering based on interpretable machine learning and extensive verification, the present application successfully provides answers to these problems.
[0047] (4) Does the technical solution of the present application overcome technical bias: the technical solution of the present application not only focuses on the quality of the answers generated by generative AI, but also overcomes the technical bias of only producing fluent and grammatically correct responses. By introducing target prompt engineering based on interpretable machine learning (XML), the present application ensures that the response of AI is more real, accurate and helpful, thus meeting the real-world psychological health needs. In addition, through in-depth research and verification, the authenticity of the output and the consistency with human experts are ensured, further overcoming the limitations of existing technology.
[0048] Fifthly, the specific technical progress of each step of the present application is:
[0049] S101: Predictive psychological help questions and answers are based on ChatGPT, and comparative analysis of average psychological health questions and answers of human consultants and ChatGPT is carried out.
[0050] In this step, ChatGPT is used as a predictive model for psychological help and question answering, and is compared with human consultants for psychological health question answering. This allows for a comparison of the performance of the model and human consultants in psychological health question answering, and provides a benchmark for subsequent technical improvements.
[0051] S102: The key language cues that affect the perceived helpfulness of psychological health question answering are determined, as well as the impact of the key language cues on perceived helpfulness.
[0052] In this step, the key language cues that affect the perceived helpfulness of psychological health question answering are determined through analysis of the content of the question answering. These language cues may include emotional expression, the way suggestions are made, understanding of the problem, etc. At the same time, the degree of influence of these key language cues on perceived helpfulness is evaluated to determine which factors are crucial for improving the performance of the model.
[0053] S103: Design a prompt strategy for ChatGPT to improve its performance in psychological health question answering through targeted prompts.
[0054] In this step, a prompt strategy is designed to improve the performance of ChatGPT in psychological health question answering based on its shortcomings. These prompts can be reminders for key language cues, such as emphasizing emotional understanding when answering questions or paying attention to tone and expression when making suggestions. Through targeted prompts, ChatGPT can better understand user needs and provide more effective psychological help.
[0055] Through these steps of technical progress, the embodiment of the present invention provides a method of human-computer interaction based on interpretable machine learning, which can improve the performance of ChatGPT in psychological health question answering and provide more perceived helpfulness in answering. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 is a flowchart of the method of human-computer interaction based on interpretable machine learning provided by the embodiment of the present invention;
[0057] Figure 2 is a schematic diagram of the method of human-computer interaction based on interpretable machine learning provided by the embodiment of the present invention;
[0058] Figure 3 is a schematic diagram of the distribution of perceived helpfulness scores of generative AI and human consultants in different risk levels of psychological health topics provided by the embodiment of the present invention;
[0059] Figure 4is an average comparison diagram of the perceived helpfulness scores of the generative AI and human counselors under different mental health topics provided by the embodiments of the present application;
[0060] Figure 5 is a diagram of the top 40 important features and performance of the prediction model composed of different numbers of top-level features provided by the embodiments of the present application;
[0061] Figure 6 is a shape diagram of the perceived helpfulness of the global explainability of the mental health question and answer provided by the embodiments of the present application;
[0062] Figure 7 is a distribution diagram of the perceived helpfulness scores of the responses of the target prompt generative AI, the original generative AI and the human counselor from different risk levels and mental health topics provided by the embodiments of the present application;
[0063] Figure 8 is a multiple comparison diagram of the perceived helpfulness scores of the target prompt generative AI, the original generative AI and the human counselor across different mental health topics provided by the embodiments of the present application. DETAILED DESCRIPTION
[0064] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0065] As shown in Figure 1 , the human-computer interaction method based on explainable machine learning provided by the embodiments of the present application comprises the following steps:
[0066] S101: Predicted mental health question and answer based on ChatGPT, and average comparison analysis of mental health question and answer human counselor and ChatGPT;
[0067] S102: Determine the key language clues that affect the perceived helpfulness of the mental health question and answer, and the influence of the key language clues on the perceived helpfulness;
[0068] S103: Design a prompt strategy for ChatGPT, and improve the performance of ChatGPT in the mental health question and answer through targeted prompting.
[0069] Embodiment 1:
[0070] 1. Method and experiment
[0071] 1.1 In this invention, the invention adopts an interpretable machine learning method to predict the perceived utility of ChatGPT in the context of mental health Q&A, and further compares it with the efficacy of human counselors. Next, the invention will introduce the dataset, input features, regression algorithm, interpretation method and evaluation metrics used to build, interpret and experiment the prediction model. In addition, the invention will also summarize the methods used and the research conducted.
[0072] 1.2 Dataset
[0073] The main challenge and difficulty of predicting the perceived utility of artificial intelligence in mental health Q&A through machine learning is centered around obtaining mental health Q&A data with labels. The data used in this invention comes from the Ye Xinli community, which is one of the largest online counseling platforms in China (https: / / www.xinli001.com / ), and nearly 40 million people from 137 countries around the world seek mental support services. The mental health data in this invention comes from the Q&A section of this platform. Individuals seeking psychological help can post their questions and problems anonymously in this invention to seek mental health services and support from platform counselors. The collected dataset includes 10,903 different psychological help-seeking questions and 19,682 counselor responses. The data spans from November 3, 2022 to March 30, 2023.
[0074] The survey report published by the well-known online mental health counseling platform, Xinli001.com / Public / 2020 / ), shows that among users of online mental health counseling, female visitors are three times more than male visitors, and early adult users (21-35 years old) account for 77.57%. Using the gpt-3.5-turbo model, the invention obtained ChatGPT's answers to the above help-seeking questions. Table 1 is a descriptive statistical analysis of the mental health question data used in this invention. Table S1 provides an overview of the data volume of ChatGPT and human counselor Q&As, which covers a wide range of mental health problems (questions), topics (topics) and crisis risk levels (Corbitt-Hall et al., 2016).
[0075] Overall, these datasets adequately cover help-seekers across various mental health problems, topics, and crisis risk levels. In addition, these counselors demonstrate a high level of expertise, thereby providing strong support for whether ChatGPT can match the level of human counselors in online mental health Q&A.
[0076] Table 1. Descriptive statistics of the mental Q&A dataset
[0077]
[0078] The present invention pre-processes the mental health question-answer text in the above dataset. The Chinese Natural Language Toolkit, jieba, provides an easy-to-use interface for Chinese corpus and lexical resources, which can complete the tasks of word segmentation, stop word removal, keyword extraction, etc. The present invention uses this toolkit to segment the above mental health question-answer text. A mental health domain dictionary is constructed using Gensim, and the embedded representation of the mental health topic is derived.
[0079] The Chinese Complex Sentiment Analysis Tool (Cnsenti) is used to identify 7 emotion statistics (positive, joy, sadness, anger, fear, disgust, surprise) from the question-answer text. The cLIWC (Chinese Language Inquiry Word Count) is used to extract statistical features of psychological language cues from the question-answer text. Based on previous research of the present invention, the present invention uses language inquiry and word count (LIWC) to calculate topic consistency, language style similarity, and emotional similarity in question-answer text to describe the therapeutic alliance in mental health question-answer text. The present invention confirms the accuracy of the feature extraction results of the present invention by manually checking 20 sets of mental health-related question-answer texts. All extracted feature values are verified.
[0080] 1.3 Feature Engineering
[0081] The present invention adjusts the concept of perceived helpfulness from the perspective of the quality of counseling services to the public to evaluate the quality of mental health questions and answers. Using computational linguistics and interpretable machine learning methods suitable for large-scale question-answer discourse analysis, a pre-interpretation model is constructed. The present invention uses the questions of the questioner, the answers of the counselor, and the synchronous interaction between the two, which are all related to mental health questions and answers, to predict and explain the perceived helpfulness of mental health questions and answers. The present invention includes LIWC, topic consistency, language style, emotional similarity features, etc.
[0082] LIWC is a text-based analysis tool that can capture the psychological dimensions of user language use. It provides a reliable method for the present invention to measure and capture the language behavior of visitors and counselors, including language cues such as emotional expression, mental state, social relationship, cognitive process, and lexical features such as vocabulary use and sentence construction. Topic consistency is a text content-based feature that reflects the consistency of the topic discussed in the text, helping the present invention to understand whether the visitor counselor question-answer is around the same topic.
[0083] Language style match is a language style-based feature that measures the similarity of language styles in text. It helps to understand whether the language style of the visitor and the counselor is consistent and whether this match affects the public's evaluation of the perceived helpfulness of mental health Q&A. Emotional similarity reflects whether the emotions of the visitor and the counselor are consistent in the conversation. This feature is extracted using an emotion analysis method, which can help the invention understand the influence of emotional interaction between the visitor and the counselor on perceived helpfulness.
[0084] In summary, the invention selects the above-mentioned language cues based on their availability in large-scale mental health Q&A text analysis and their potential role in predicting perceived helpfulness. In the machine learning models built below, these language cues form a key feature set that can predict and explain the perceived helpfulness of mental health Q&A.
[0085] 1.4 Explainable machine learning methods
[0086] Using perceived helpfulness as the output target of the prediction model of the invention, and using language inquiry and word count (LIWC) of the seeker and counselor, as well as the therapeutic alliance feature, as input, the invention uses explainable machine learning techniques to develop explainable prediction models. These models aim to assess the perceived helpfulness in mental health Q&A.
[0087] The invention selects five commonly used regression algorithms, including ridge regression, least absolute shrinkage and selection operator (LASSO), support vector regression (SVR), random forest (RF) and eXtreme gradient boosting (XGBoost). Ridge regression and lasso are regression techniques used to deal with collinearity and overfitting. Ridge regression restricts model coefficients by a penalty term, reducing the noise sensitivity of the model, while lasso allows coefficients to become zero, thus achieving automatic variable selection and enhancing the interpretability of the model. Support vector regression is an application of support vector machines that can effectively handle non-linear and high-dimensional data problems and has been used for perceived helpfulness prediction. Ensemble methods such as random forest and XGBoost improve prediction accuracy and prevent overfitting by combining multiple decision trees and applying regularization terms, respectively.
[0088] The present application uses cross-validation recursive feature elimination (RFECV) and XGBoost model fitting to select effective features from the original feature set, and finally filters out 338 features that significantly affect the prediction of perceived helpfulness. In addition, the present application uses GridSearchCV to select hyperparameters with the best prediction performance from the parameter set. In addition, in order to further improve the prediction accuracy, interpretability and generalization performance of the model, the present application adopts the method of sensitivity analysis and feature pruning. The present application selects an important feature subset from the effective feature set and re-trains the model using these features.
[0089] Artificial intelligence better understands different types of mental aid requests through additional prompts, and takes targeted response strategies, which is crucial for building a trustworthy mental health Q&A robot. Shapley value is widely used in cooperative game theory, which can represent the degree of influence of a feature on the change of model output. SHAP value greatly improves the transparency of machine learning and has been applied to many research and industrial scenarios. Therefore, the present application uses SHAP value to explain the prediction model established by the present application. They can help the present application understand the size and direction of the contribution of each feature to the prediction result.
[0090] Therefore, the model of the present application considers the influence of relevant language prompt characteristics on mental health Q&A, and through the impact value of these characteristics, the present application can understand the contribution of language behavior to the prediction of the willingness of consultants to answer and the direction of its influence. With this information, the present application can design a prompt strategy for ChatGPT.
[0091] 1.4 Evaluation index
[0092] The prediction performance of the model constructed in the present application is evaluated using four widely used regression model indicators, including root mean square error (RMSE) and mean absolute percentage error (MAPE), as defined in equations (1)-(2). Root mean square error (RMSE) is a measure of the difference between predicted and actual values. On the other hand, mean absolute percentage error (MAPE) is a popular measurement method that quantifies prediction accuracy in percentage, providing a relative measure of prediction error:
[0093]
[0094]
[0095] where n represents the total number of observations in the mental health Q&A dataset, y i represents the actual perceived help of the i-th observation in the dataset, is the predicted perceived helpfulness of the i-th observation output by the regression model.
[0096] The embodiment of the present application provides a human-computer interaction system based on an interpretable machine learning, which comprises:
[0097] An information analysis module is used for predicting psychological help questions and answers based on ChatGPT, and performing average comparison analysis on psychological health questions and answers of a human consultant and ChatGPT;
[0098] A clue determination module is used for determining key language clues affecting the perceived helpfulness of psychological health questions and answers, and the influence of the key language clues on the perceived helpfulness;
[0099] A strategy prompt module is used for designing a prompt strategy for ChatGPT, and improving the performance of ChatGPT in psychological health questions and answers through targeted prompting.
[0100] Through the comparative analysis of ChatGPT and human consultants in the psychological health question and answer environment, the present application discloses a series of significant findings, indicating that the embodiment has great advantages. The following are these findings described based on experimental data and graphs:
[0101] Improving perceived helpfulness: Analysis of variance (ANOVA) and Tukey HSD test show that through the targeted prompting process, the perceived helpfulness of ChatGPT at all risk levels is significantly improved. In particular, at a lower risk level, the prompted ChatGPT even slightly outperforms the human consultant in terms of perceived helpfulness.
[0102] Approachability to human consultants: At a high psychological crisis risk level, the performance of the prompted ChatGPT and the human consultant does not differ significantly, indicating that the machine's response has reached a similar effect to that of a professional.
[0103] Stability of score distribution: Compared with the original ChatGPT model, the prompted ChatGPT shows more concentrated characteristics in the distribution of perceived helpfulness scores, and overall shows greater consistency, with its response being more stable compared to the human consultant's score.
[0104] Improvement in thematic areas: In certain specific psychological health thematic areas, such as popular topics, career choice strategies, communication and interpersonal relationship boundaries, the prompted ChatGPT improves the perceived helpfulness by 11%-15%, which is significantly higher than the unprompted version.
[0105] Exceeding human consultants in certain areas: Although in some specific areas, the performance of the human consultant is still slightly better, in other topics such as career management and work stress, the prompted ChatGPT improves the performance by more than 10% compared to the human consultant.
[0106] In summary, the present application significantly improves the performance of ChatGPT in the psychological health question and answer process by targeted prompting, making it even reach or exceed the level of human consultants in some areas. This brings a potential tool to the field of psychological health consultation, which is expected to be further optimized and applied in the future.
[0107] The present application has achieved some positive effects during research and development or use, and indeed has great advantages compared with the prior art. The following content is described in combination with the data, graphs and the like of the test process.
[0108] 2. Results
[0109] 2.1 Performance analysis and optimization of perceived helpfulness prediction model
[0110] In the present application, the present application establishes an explanatory prediction model of perceived helpfulness of psychological health question and answer through three steps: (1) the present application trains and tests prediction models using complete effective prediction features and five different classification algorithms, and then selects the best model in overall performance; (2) the best model is evaluated for the influence of each feature on the prediction performance by using sensitivity analysis technology, and then the features are pruned to only retain the most influential features and the optimal prediction model; (3) the present application selects a small number of important features from the pruned feature set, and re-trains and tests the best classifier to obtain an explanatory prediction model. The present application will report the detailed results of each stage below. The present application develops an optimal prediction model and applies it to determine the perceived helpfulness rating of ChatGPT response. Then, the present application compares these scores with those of human consultants to evaluate whether ChatGPT can match their level in online psychological health question and answer.
[0111] 2.1.1 Training and testing regression variables
[0112] The present application uses different regression algorithms to establish different prediction models on the effective feature set, and evaluates the performance of each prediction model by RMSE and MAPE values, as shown in Table 2. From the table, it can be seen that the RMSE value and MAPE value of the prediction model based on XGBoost are the lowest, which are 0.4863 and 17.13% respectively. The prediction performance of XGBoost is better than all other prediction models, and the prediction performance of random forest (RF), support vector regression (SVR), ridge regression and Lasso regression algorithms decreases in turn.
[0113] Table 2 Performance of different algorithms
[0114]
[0115] 2.1.2 Comparative analysis of ChatGPT and human counselors' perceived helpfulness in mental health Q&A
[0116] The ChatGPT score results automatically predicted by the optimal prediction model show that ChatGPT without targeted prompts performs worse than the average level of human counselors in perceived helpfulness in mental health Q&A (Table 3, Figure 3 ). Specifically, in all types of mental health Q&A, the answers provided by human counselors are significantly better than ChatGPT in perceived helpfulness (p < 0.01; two-tailed Student's t-test; Table 3). In addition, the present invention found that human counselors showed strong dichotomous division in perceived helpfulness in mental health Q&A, while ChatGPT showed stronger consistency; ChatGPT's lowest performance was higher than humans, but its highest performance was lower than humans (see Figure 3 ).
[0117] Table 3 Multiple comparisons of ChatGPT and human counselors' answer perceived helpfulness scores
[0118]
[0119] The present invention further analyzed the percentage difference in perceived helpfulness between ChatGPT and human mental health counselors in four different levels of psychological crisis, as Figure 4 shown. The present invention found that the gap between ChatGPT and human counselors was greater in low-risk mental health topics, with human counselors' reactions being more than 15% more helpful than ChatGPT's reactions (e.g., family relationships, growing up experiences, parent communication, personality-related, and marriage). The experiment found that human counselors were more effective than ChatGPT in dealing with medium-risk mental health issues, with their reaction efficiency exceeding 10%. These topics covered a wide range, from social software to disease diagnosis, stress, family trauma, etc. In high-risk mental health topics, human counselors' reactions were more than 6.9% more helpful than ChatGPT's reactions (e.g., self-harm and suicide). Although ChatGPT performed worse than human counselors in most topics, it performed better or similarly to human counselors in career Q&A topics (e.g., midlife crisis, career burnout, work stress, and workplace communication). Overall, ChatGPT has not yet proven to be as capable as human counselors in mental health Q&A without prompts.
[0120] 2.1.3 Feature pruning and model refinement based on optimization analysis
[0121] Considering that prompting engineering is a key optimization technique for large language model-driven artificial intelligence, it is necessary to explore how to purposefully design prompt engineering, and then discuss the performance of the optimized ChatGPT in the mental health Q&A scenario. Therefore, the present application combines the SHAP value method for sensitivity analysis and feature pruning of the prediction model to improve the explainability of the prediction model. First, based on the constructed prediction model, the SHAP value of different features is calculated using the SHAP method, and the feature importance is sorted; second, the present application adds features in the prediction model according to the importance order, and uses the MAPE value to evaluate the relative importance of each feature and the performance of the model of different feature combinations, as shown in Figure 5 The present application finds that when the number of features is 60, the prediction model reaches the optimal performance, with MAPE of 0.1685 and root mean square error (RMSE) of 0.4815; when the number of features is 17, the MAPE and RMSE of the model are 0.1691 and 0.4911, respectively, reaching 99.662% (MAPE) and 97.985% (RMSE) of the optimal performance.
[0122] Therefore, the present application establishes an explainable prediction model based on the first 17 features, and evaluates the performance of the model using RMSE and MAPE, as shown in Table 4. As can be seen from the table, the RMSE of the pruned prediction model is slightly higher than that of the original prediction model (0.4847), which is 0.4913, and the MAPE is 16.9312%, which is better than that of the original model (17.0877%), indicating that the pruned prediction model performs well.
[0123] Table 4 Comparison of model performance before and after feature pruning
[0124]
[0125] 2.2 Explainable analysis of perceived helpfulness in mental health Q&A and generation of ChatGPT prompting strategy
[0126] 2.2.1 Influence of different language cues on perceived helpfulness in mental health Q&A sessions
[0127] In order to design the prompting strategy of ChatGPT, the present application uses SHAP value to perform global explainability analysis of the model, and discusses the influence of different features on perceived helpfulness. Specifically, the present application calculates the cumulative SHAP value of the first 17 features, and sorts them according to their influence, as shown in Table 5. In addition, in order to more intuitively analyze the degree and direction of the influence of each feature on perceived helpfulness, the present application collects all individual samples in the data set into a global explainability SHAP plot, as shown in Figure 6 As can be seen from the figure, the influence of each feature on perceived helpfulness can be mainly divided into two types:
[0128] Table 5 Features that affect perceived helpfulness and their ranking
[0129]
[0130]
[0131] (1) Positively affect perceived helpfulness. Among these features, the ones that have the most impact on perceived helpfulness are the reward amount (Reward), the total number of words (WordCount), and the number of views (Views). The higher the reward, the more words, and the more perspectives, the more they can improve the helpfulness of the answer. In addition, the more use of language features such as psychological words (Psych), prep-end words (PrepEnd), multi-functional words (MultiFun), perceived course words (Percept), period (Pefiod), quotation marks (Quot), etc., the more they can improve the perceived helpfulness.
[0132] (2) Negatively affect perceived helpfulness. This feature is mainly focused on language cues such as adverbs (Adverb) and tentative words (Tentat); at the same time, parentheses (Paren), word diversity (WordSentence), other punctuation (OtherP), present words (tNow), and professional topics (Profession) also belong to this type of influence.
[0133] Note: In this graph, the graph is composed of thousands of individual points from the training dataset, with higher values being more red and lower values being more blue. This is described by the "Feature Value" bar on the right side of each graph. Therefore, if the points on one side of the center line become more and more red or blue, it means that increasing or decreasing the value will move the predicted perceived helpfulness in this direction.
[0134] 2.2.2 Generation of ChatGPT prompt strategies based on language cues
[0135] To further analyze and validate the differences in language use between ChatGPT and human counselors and their impact on perceived helpfulness, the present invention conducted a mean comparison analysis of the language feature values of ChatGPT and human counselors. Table 6 shows the results of the two-tailed t-test of the language features used in the answers of human counselors and ChatGPT. The present invention found that ChatGPT was lower than humans in text length and function word use (p<0.001), used too many temporal words (p<0.01), which reduced perceived helpfulness. However, ChatGPT outperformed humans in period use and sentence conciseness (p<0.001), showing its unique advantage in language conciseness. These findings helped the present invention establish additional targeted prompts that can effectively improve perceived helpfulness. Therefore, the present invention further identified and validated three types of key operational features: the present invention named the first type as engagement features, such as answer length and use of punctuation, which are positively correlated with perceived helpfulness. The present invention named the second type as cognitive process features, such as function words and post words, which can enhance sentence logic and improve perceived helpfulness, while temporal words increase uncertainty and reduce perceived helpfulness. The present invention named the third type as psychological theme features, such as the use of psychological vocabulary, which reflects the psychological care of the respondents and enhances the relevance of the answers to the topic and improves perceived helpfulness.
[0136] For these three core features, the present invention added three specific prompts to the first overall prompt as follows:
[0137] I hope you can be a mental health counselor. I will provide you with a user seeking help for mental health problems, and I will ask you to provide advice and suggestions for their emotions, stress, anxiety, and other mental health problems. You should use your knowledge of cognitive behavioral therapy, meditation techniques, mindfulness exercises, and other psychological treatment methods to create personalized strategies to improve their overall mental health. Please answer according to the following prompts:
[0138] Try to increase the number of words in your answers as much as possible to show your attention and engagement with the current problem (including but not limited to more specific examples and details to better explain certain concepts and points of view; quote relevant research and materials in your answers to support your points and suggestions) to help users better understand the problem and solutions.
[0139] Show coherence and certainty in your answers, use more conjunctions and function words, and provide thorough analysis and reliable suggestions for the current mental health problem, and help those seeking help to take better action.
[0140] In your answers and the visitors' attention to the theme of mental health, use terms to explain the problem and propose solutions to improve the professionalism of the answers.
[0141] Table 6 presents a comparative analysis of the average values of perceived helpfulness transmitted between ChatGPT and human counselors in the context of mental health questions and answers
[0142]
[0143] 2.3 Comparative analysis of post-prompted ChatGPT and human counselors in the context of mental health questions and answers
[0144] This invention presents a comparative analysis of the perceived helpfulness of mental health questions and answers (Q&A) generated by human counselors and answers by ChatGPT, both in its original version and in a version improved through a targeted prompting process.
[0145] 2.3.1 Comparative analysis of perceived helpfulness of ChatGPT and human counselors before and after targeted prompting in the context of mental health questions and answers
[0146] This invention uses Analysis of Variance (ANOVA) and Tukey's Honestly Significant Difference (HSD) test to compare the average perceived helpfulness scores of the reactions generated by the three participants. The objective of this invention is to assess to what extent the targeted prompting process improved ChatGPT's performance and to determine whether ChatGPT after prompting is significantly different from human counselors in terms of perceived helpfulness of the reactions. As shown in Table 7, the first key finding is that ChatGPT's reactions after prompting were significantly better than before optimization (p<0.001) at all four risk levels. This indicates that the prompting process significantly improved the perceived helpfulness of the reactions generated by ChatGPT. Interestingly, after the prompting process, there was no significant difference between ChatGPT and human counselors in terms of reactions at higher levels of psychological crisis risk. In addition, at lower risk levels, ChatGPT's reactions were even slightly better than those of human counselors (p<0.001).
[0147] Table 7 presents a multiple comparison of the average perceived helpfulness of different types of mental health questions and answers.
[0148]
[0149] To further understand the overall performance of human counselors, ChatGPT, and ChatGPT after prompting in the context of mental health questions and answers, this invention tests the distribution of perceived helpfulness scores at different levels of psychological crisis risk. Figure 7The distributions show that the prompted ChatGPT scores are concentrated around the median, while the original ChatGPT model scores are skewed towards the 25th percentile and the median. Regardless of whether it has undergone targeted prompts, ChatGPT's responses exhibit greater consistency than human counselors, albeit with a lower bound.
[0150] As Figure 8 shown, the percentage difference in perceived helpfulness scores for different mental health topics, following post-hoc ANOVA and Tukey HSD multiple comparison analysis, shows that scores overall improved after prompting. The average score for ChatGPT after prompting was at least 7% higher than the un-prompted version. In certain topic areas, particularly those with no risk of mental crisis, such as popular topics and career choice strategies, the perceived helpfulness scores increased by more than 11%. On lower-risk topics, such as communication and interpersonal relationship boundaries, scores improved by more than 15%. Therefore, the present invention can conclude that the machine learning-based prompting process significantly improved ChatGPT's performance in the mental health question-answering process.
[0151] In addition, the present invention found that ChatGPT's performance after prompting in the mental health question-answering context was comparable, if not superior, to human counselors. Despite these advances, human counselors still outperformed ChatGPT by more than 4.2% on some low-risk topics (such as family, family relationships, parent relationships, parent communication, growing up, personality traits, personality refinement) and medium-risk topics (such as family trauma, lost love). Notably, ChatGPT outperformed human counselors by more than 10% on low-risk topics such as career management and work stress. On no-risk topics such as career choice strategies, ChatGPT's performance was 9% higher than that of human counselors.
[0152] By refining the prompting process, it is possible to further improve ChatGPT's performance and make it a valuable tool in the field of mental health counseling. Future research can focus on improving the model's performance in high-risk scenarios, where human counselors still have a slight advantage.
[0153] 2.3.2 Comparative analysis of ChatGPT and human counselor key language cues usage after prompting engineering
[0154] To further determine the mental health question-answering behind the aforementioned differences, the present invention analyzed the differences in average language cue eigenvalues between human counselors, ChatGPT, and the target prompted ChatGPT on different risk topics using ANOVA and Tukey HSD multiple comparison methods. The results are shown in Table 8.
[0155] Firstly, the application compares the differences in language cues before and after applying the target prompt in ChatGPT's answers. Specifically, in terms of participation, as can be seen from the table, the total number of words contained in ChatGPT's responses is 40.496% higher than before the prompt, and the frequency of quotation usage increases by 82.845%. These differences in language cues all indicate that ChatGPT's participation is enhanced after the prompt. In terms of psychological themes, the former uses more psychological words (3.986%, p<0.001), which to some extent improves the readability and professionalism of the reply text. However, the former shows mixed results in cognitive processes, reducing the use of present tense words (16.976%, p<0.001) to help enhance perception; but its use of multi-functional words (8.122%, p<0.001), postpositions (2.935, p<0.001) and temporary words (21.999%, p<0.001) is not conducive to enhancing the helpfulness of perception.
[0156] Secondly, the application compares and analyzes the differences in language cues between ChatGPT and human counselors after the target prompt. Specifically, in terms of participation, the former's total number of answers is slightly lower than the latter (2.229%, p<0.01), uses fewer adverbs (39.67%, p<0.001), but uses more periods (41.57%, p<0.001). This shows that the difference in investment is not significant, and ChatGPT's answer language is more concise; in terms of psychological themes, the former uses more psychological words (24.53%, p<0.001). In cognitive processes, the former also shows different results, reducing the use of present tense words (-79.178%, p<0.001) to help enhance perception, but using multi-functional words (-22.016%, p<0.001) and temporary words (47.368%, p<0.001) is not conducive to enhancing the helpfulness of perception.
[0157] Overall, ChatGPT can significantly optimize its language use behavior in psychological health question and answer based on targeted prompt strategies, achieving the performance of human counselors in this field. However, when considering the optimal use of language cues related to participation and psychological themes, the optimization of language cues related to cognitive processes is not considered.
[0158] Table 8 Comparison of the average values of features affecting the helpfulness of perception between human counselors and ChatGPT before and after the target prompt.
[0159]
[0160] 3. Conclusion
[0161] In this invention, the invention tested the ability of ChatGPT to handle real online mental health questions from the Chinese mental health community (YiXinLi, xinli001.com). These findings are interesting and exciting, ChatGPT is like a beginning student, and with special help can perform better. Initially, ChatGPT's answers to mental health questions were considered less helpful than those provided by humans. However, after implementing customized prompts containing basic elements that help to be helpful, ChatGPT outperformed human performance on a total scale. Specifically, ChatGPT outperformed human counselors in handling low-risk mental health problems, but did not significantly differ from human levels in moderate-risk and high-risk problems. In the following sections, the invention will provide a detailed explanation of the research results and an overview of the prospects of the study.
[0162] 3.1 ChatGPT's ability to solve mental health problems
[0163] As a large language model, ChatGPT has gained praise and attention for its strong ability in language understanding and organizing output. Its vast knowledge base and strong natural language processing capabilities enable it to handle a variety of complex language tasks, even when communicating content contains emotional components that require more cognitive resources. The invention obtained ChatGPT's answers to more than 10,000 real mental health questions and compared them with human answers to confirm its ability to understand and answer these questions. These findings further indicate that, when given appropriate prompts, ChatGPT can handle these mental health problems, possibly reaching or exceeding human levels.
[0164] It should be noted that ChatGPT does not handle all mental health problems as well as humans. The invention found that, after receiving prompts, ChatGPT performed better than humans in handling no-risk and low-risk problems, but did not significantly differ from human levels in handling moderate-to-high-risk problems. One possible explanation is that ChatGPT's training data and model structure may have some limitations. Although ChatGPT's training scale is large, it is still limited by the quality and diversity of the training data. If there is a lack of moderate-to-high-risk examples in the training data, ChatGPT may not be able to accurately understand and answer such problems.
[0165] Human consultants who deal with severe risk issues also have rich domain knowledge and professional experience. They improve their crisis intervention skills through long-term learning and practice. In contrast, ChatGPT is a language processing system based on model training. Although it can learn from a large amount of text data, its domain knowledge and experience are still limited and cannot be compared with human consultants. As a content generation tool, ChatGPT cannot replace professional human consultants. Although it can provide some basic information and support to users when dealing with moderate and high-risk mental health issues, it still needs the involvement of human consultants in decision-making and treatment planning. As ChatGPT-generated content: a tool is just a tool, not a substitute for professional mental health consultants. Although it can provide some basic information and support to users when dealing with moderate and high-risk mental health issues, decision-making and treatment planning still require the involvement of human consultants.
[0166] 3.2 ChatGPT's language understanding and expression abilities
[0167] The breadth of ChatGPT's understanding is commendable. The psychological topics involved in this invention are very diverse, including personal, family, school, workplace, social level, etc. (see Table S1). ChatGPT can handle these topics through unified prompts, while a human consultant may not necessarily have such broad abilities and insights. At the same time, ChatGPT's understanding of details still has a lot of room for improvement. In the prompt engineering part of this invention, it was found that ChatGPT could not effectively target each feedback prompt (although the overall improvement after the prompt was significant). For example, in the prompt, the invention asked to reduce the use of exploratory words, but the result was that the use of these words actually increased in the output content after the prompt. One possible explanation is that ChatGPT, as a language generation model, is subject to default language rules when outputting text. Therefore, when faced with prompt requests about syntactic and semantic details, ChatGPT must adhere to the default rules. Another more direct explanation is that GPT-3.5 still has relatively limited algorithmic capabilities when reacting to each specific prompt. For higher versions (e.g., GPT-4.0), this concern will no longer be necessary.
[0168] Furthermore, ChatGPT's performance varies significantly across different environments. For example, ChatGPT performs outstandingly in handling career-related issues compared to more specialized mental health issues such as depression and anxiety. This difference can be attributed to the need for different coping strategies in mental health support, which depends on the issue. For instance, addressing career mental health issues might focus more on providing specific action plans and suggestions, while addressing depression and anxiety issues requires empathy and emotional support. In fact, large-scale predictive models have shown the necessary capabilities for mental health support, such as exploration, empathy, and providing suggestions.
[0169] However, their adaptability in different question-and-answer environments leaves room for improvement. Although ChatGPT has specific mental health counseling capabilities, it still lacks the ability to apply these skills in context. To improve this situation, the invention proposes two directions for improvement: on the one hand, the invention needs to find more training data specifically targeting depression and anxiety issues to improve the model's performance on these topics; on the other hand, the invention needs to optimize the model's adaptability to enable it to choose the appropriate counseling strategy depending on the question-and-answer situation.
[0170] 3.3 ChatGPT's answerability is stable and malleable
[0171] Based on the collected real mental health issues, the invention compared the perceived helpfulness of responses from humans, ChatGPT, and ChatGPT to the same question. The results showed that the perceived helpfulness of human counselor responses was polarized, concentrated below the median level, while the upper limit was relatively high. In contrast, ChatGPT-generated responses showed more consistency, with a concentration around the median, especially for prompted ChatGPT (see Figure 6 ). One possible explanation is that the help-seekers' questions were expressed spontaneously without a fixed question framework and covered a wide range of topics. Coupled with the different experiences and abilities of human counselors, the perceived helpfulness of human responses was polarized. Although ChatGPT did not have predefined response templates, its training and generation process followed certain frameworks and methods, which might lead to similar responses to specific queries. ChatGPT's overall performance was more stable than that of human counselors. Moreover, ChatGPT even outperformed humans in the face of no-risk and low-risk mental health issues. These results have important implications: ChatGPT has stable mental health handling capabilities, and its output content is malleable.
[0172] 3.4 Timely engineering settings
[0173] Prompt engineering is the foundation of effectively and fully utilizing ChatGPT's capabilities. A key task of rapid engineering is to find key enabling elements that are consistent with the target. In this invention, a predictive model of perceived helpfulness for mental health problems was created using NLP and ML. The invention uses sensitivity analysis to prune features to enhance the interpretability of the model. Ultimately, the invention obtained a predictive model with 17 key features. These characteristics and their effects provide a clear basis for setting targeted prompt content. Subsequently, the invention translates the interpretable machine learning results directly into ChatGPT prompts. Comparative analysis shows that ChatGPT has made corresponding improvements in the content generated after prompting (see Table 7), confirming the effectiveness of this prompt engineering.
[0174] When using interpretable machine learning to analyze prompt content, as seen in this invention, there may be a large amount of data for machine learning at times. In this case, a theory-based approach can be a viable option. For example, in the context of mental health Q&A topics, if there is an established theoretical framework outlining factors that affect the effectiveness of perceived answers, researchers can use these factors to create prompts for ChatGPT to produce better answers. This also provides a way to objectively verify certain theoretical frameworks. However, further discussion of this issue is beyond the scope of this invention.
[0175] 3.5 Prompt engineering effectiveness verification
[0176] After getting customized prompts, how can users confirm that ChatGPT has understood and made improvements targeted? This invention uses NLP techniques to carefully compare ChatGPT output before and after prompting. The experiment found that prompt items significantly affected the overall perceived helpfulness of the response (Hariri, 2023). Currently, since user interaction with ChatGPT mainly involves text-based content, it is a challenge for users to directly evaluate the effectiveness of prompt engineering, especially in complex prompt needs. Therefore, natural language processing can be used as a scientific method to evaluate the effectiveness of rapid engineering. It can objectively assess and review the impact of rapid strategies on the quality of ChatGPT responses. The results of this invention not only effectively verify ChatGPT's ability to understand modified prompts, but also provide a specific and feasible solution for the development and verification of prompt engineering.
[0177] Overall, this study ultimately demonstrates ChatGPT's ability to understand and respond to text-based mental health issues. Moreover, when combined with precise prompting strategies, ChatGPT shows potential for targeted and substantial enhancement. These results provide preliminary scientific support for the practice of large language models or chatbots in the mental health field and important inspiration for realizing a truly artificial intelligence counselor. It is worth noting that these conclusions should be interpreted and promoted cautiously, and this study does not advocate using them to promote "artificial intelligence threat" claims. In the short term, hardware-based computer computing power is limited, but prompting engineering is "unlimited." ChatGPT has extensive knowledge beyond any individual, and all ordinary users need to do is explore how to prompt it to get accurate answers. Without a doubt, all these attempts must be carried out within an ethical, moral, and legal framework.
[0178] First, the mental health dataset in this study contains a variety of questions and answers, such as work-related questions and family-related questions. Different areas of mental health Q&A may have different language styles and helpfulness patterns of perceived response. However, when building the prediction model of this study, this study did not perform category division (prediction error is 17.13%), which may limit the external validity of the results of this study (this study still compared the performance of ChatGPT when facing different types of questions). Therefore, in future research, it is advocated to conduct specialized prediction modeling for segmented mental health areas. Second, the prediction model developed in this study includes some marginal clues, such as the number of questioner rewards and the number of post views. These factors may affect the perceived usefulness of an answer, but cannot be translated into the content of ChatGPT clues, thereby weakening the effectiveness of immediate engineering. Even so, ChatGPT has made significant improvements under limited prompts. Third, there are many differences between mental health Q&A and psychological counseling. Although some counseling terms used in counseling, such as mental health counseling and counselor, are used in the study, this study does not recommend directly extending the research results to the field of mental health counseling. Fourth, this study uses ChatGPT (GPT-3.5 version) developed by OpenAI, and whether the above results are applicable to other large language models or higher versions still needs to be verified.
[0179] The present invention uses a machine learning method to establish a prediction model based on a real-world mental health question and answer dataset. Using this model, the invention compares the responses of ChatGPT and human counselors within the same framework. The results show that ChatGPT's answers to mental health problems are lower than those of humans. Subsequently, the invention develops a prompt engineering based on interpretable machine learning. Prompting ChatGPT achieves the average level of human counselors on high-risk and medium-risk issues, and even exceeds the average level on low-risk and no-risk issues. The present invention provides preliminary evidence for the ability of ChatGPT to solve real mental health problems, and proves the way to build instant engineering using interpretable machine learning. This paves the way for the practical application of ChatGPT in the field of mental health.
[0180] Example 2: Image classification system based on machine learning
[0181] Data collection and preprocessing
[0182] Collect a dataset containing images of different categories, ensuring that the dataset has sufficient samples and representativeness.
[0183] Preprocess the images, such as resizing, cropping, and standardizing, to facilitate input into the machine learning model.
[0184] Feature extraction and selection
[0185] Use a convolutional neural network (CNN) as a feature extractor to obtain high-level feature representations of images by performing feature extraction on a pre-trained CNN model.
[0186] Select appropriate feature selection methods, such as principal component analysis (PCA) or mutual information (Mutual Information), as needed to reduce feature dimensionality and improve classification performance.
[0187] Model training and evaluation
[0188] Use the extracted image features as input and use supervised learning algorithms (such as support vector machines, random forests, or deep neural networks) for model training.
[0189] Divide the dataset into training and test sets, use the training set for model training, and then use the test set to evaluate the classification performance of the model.
[0190] According to the evaluation results, perform model optimization, such as adjusting hyperparameters, increasing data samples, or using ensemble learning methods, etc.
[0191] Example 3: Autonomous driving system based on machine learning
[0192] Data acquisition and preprocessing
[0193] Data collection
[0194] Data preprocessing
[0195] Feature extraction and selection
[0196] Feature extraction
[0197] Feature selection
[0198] Model training and optimization
[0199] Model training
[0200] Algorithm selection
[0201] Model optimization
[0202] Conclusion
[0203] It should be noted that the embodiments of the present application can be realized by hardware, software or a combination of software and hardware. The hardware part can be realized by special logic; the software part can be stored in a memory and executed by a suitable instruction execution system, such as a microprocessor or a specially designed hardware. Those skilled in the art can understand that the above-mentioned devices and methods can be realized by computer executable instructions and / or included in processor control code, such as carrier media, such as magnetic disk, CD or DVD-ROM, programmable memory, such as read-only memory (firmware), or data carrier, such as optical or electronic signal carrier. The device and its modules of the present application can be realized by hardware circuit, such as ultra-large scale integrated circuit or gate array, semiconductor, such as logic chip, transistor, etc., or programmable hardware device, such as field programmable gate array, programmable logic device, etc., can also be realized by software executed by various types of processors, and can also be realized by the combination of the above hardware circuit and software, such as firmware.
[0204] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited thereto, any modification, equivalent replacement and improvement within the technical range disclosed by the present application and within the spirit and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A human-machine interaction method based on interpretable machine learning, characterized in that, The method comprises the following steps: Obtaining mental health counseling data of generative artificial intelligence, the main challenge and difficulty of predicting the perceived usefulness of artificial intelligence in mental health question and answer through machine learning is to obtain mental health question and answer data with labels; Preprocessing the dataset, including word segmentation, stop word removal, keyword extraction; using tools such as Gensim to process mental health question and answer text, build a mental health domain dictionary, and export an embedded representation of mental health topics; using emotion analysis tools to identify 7 emotion statistics from question and answer text, including positive, joy, sadness, anger, fear, disgust, surprise; using cLIWC to extract statistical features of psycholinguistic cues from question and answer text; using language surveys and word counts to calculate topic coherence, language style similarity, and emotional similarity in question and answer text to describe the therapeutic alliance in mental health question and answer text; verifying all extracted feature values by manually checking the question and answer text related to mental health; identifying language patterns from counseling data, including the mental language cues of the questioner's question, the counselor's answer, and the interaction process between the two, topic coherence, language style matching, and emotional similarity; obtaining effective language pattern features for regression prediction; Optimizing the prediction model using feature sensitivity analysis and identifying the impact of features on the effectiveness of psychological counseling; using computational linguistics and interpretable machine learning methods suitable for large-scale question and answer discourse analysis to build a prediction model; using the questioner's question, the counselor's answer, and the synchronous interaction between the two to predict and explain the perceived helpfulness of mental health question and answer; this includes LIWC, topic coherence, language style, and emotional similarity features; using perceived helpfulness as the output target of the prediction model, and using the language exploration and word counts of the seeker and the counselor, as well as the therapeutic alliance features, as inputs, using interpretable machine learning techniques to develop an interpretable prediction model; selecting five commonly used regression algorithms, including ridge regression, least absolute shrinkage and selection operator, support vector regression, random forest, and eXtreme gradient boosting; Ridge regression and LASSO are used to handle problems of collinearity and overfitting; support vector regression is an application of support vector machines for handling non-linear and high-dimensional data problems and has been used for perceived helpfulness prediction; ensemble methods improve prediction performance by combining multiple models, such as random forest and eXtreme gradient boosting; Based on the results of feature sensitivity analysis, extract actionable language model features for targeted prompts; these prompts can include how to improve the expression of questions and how to improve the quality of answers; These prompts are transmitted to the generative artificial intelligence model along with the seeker's question to generate optimized answers; The generated answers not only meet the needs of the seeker, but also take into account the extracted language model features to create high-quality mental health question and answer; extract actionable language model features to develop targeted prompts; and these prompts are transmitted to the generative artificial intelligence model along with the seeker's question to obtain optimized answers.
2. The method of claim 1, wherein, The prediction results are subjected to feature sensitivity analysis to determine the key features affecting the prediction results; through the feature sensitivity analysis, it can be determined which mental health question and answer features are most important to predict the perceived helpfulness.
3. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the method of any one of claims 1-2.
4. An information data processing terminal for self-help mental health services, characterized by The information data processing terminal of the self-service mental health service is used to implement the method of any one of claims 1-2.
5. The human-machine interaction system based on interpretable machine learning according to any one of claims 1-2, wherein, The self-service mental health service information data processing terminal comprises: An information analysis module is configured to predict mental health question and answer based on ChatGPT, and to compare and analyze the average mental health question and answer of human consultants and ChatGPT; A clue determination module is configured to determine key language clues affecting the perceived helpfulness of mental health question and answer, and the influence of key language clues on perceived helpfulness; A strategy prompting module is configured to design a prompting strategy for ChatGPT, and to improve the performance of ChatGPT in mental health question and answer through targeted prompting.