Online psychological health self-help guide dialogue method

By constructing personalized temporal psychological models and multimodal causal inference maps, the problem of depth and personalized understanding in online psychological dialogue systems is solved, enabling dynamic and accurate assessment and personalized guidance of users' psychological states.

CN121416005APending Publication Date: 2026-01-27周洁
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511593617.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing online psychological dialogue systems lack diachronic comprehensive analysis of users' multimodal information, resulting in a lack of depth and personalization in understanding users, and generating generic guidance responses that lack dynamic empathy.

Method used

By integrating users' explicit language expressions and implicit interactive behaviors, a personalized temporal mental model is constructed. A sequence model is used to capture the evolution of mental states. Combined with a multimodal causal inference graph, a dynamic empathic context summary is generated, enabling a deep and diachronic understanding of the dialogue system.

Benefits of technology

The system improves the real-time accuracy of psychological state assessment, proactively explores potential triggers leading to the current state, generates highly personalized and targeted guidance responses, and enhances the intervention effect of self-help tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121416005A_ABST
    Figure CN121416005A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and psychological health, and discloses an online psychological health self-help guide dialogue method, which comprises the following steps: acquiring multi-modal data of a user, including content data and behavior metadata; constructing and dynamically updating a personalized time sequence psychological model based on the data, wherein the model contains a sequence model for capturing state time sequence evolution and a multi-modal causal inference map for modeling a causal relationship; then, generating a dynamic common situation context abstract representing the current psychological state of the user based on the model; and finally, combining the abstract with the current input of the user, and generating a guided dialogue through a large language model. According to the method, through deep diachronic analysis fusing languages and behaviors, guiding replies which are highly personalized and have dynamic estrus sharing ability can be generated, and the effectiveness of self-service psychological services is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and mental health technology, specifically to an online self-guided dialogue method for mental health. Background Technology

[0002] With the accelerating pace of society and increasing public awareness of mental health issues, digital and intelligent mental health services are developing rapidly. Among them, artificial intelligence-based online dialogue systems (i.e., chatbots) are emerging support tools that have shown great application potential due to their ability to provide 24 / 7, convenient, low-cost, and highly anonymous support channels, attracting widespread attention and research investment from the industry.

[0003] However, existing online psychological dialogue systems still face significant challenges in providing in-depth and effective personalized guidance. Most current mainstream chatbots rely on immediate analysis of the user's current single-turn input. This means the system's understanding and response are limited to isolated dialogue fragments, lacking long-term memory and diachronic understanding of the user's state. When users repeatedly mention related concerns, the system struggles to connect these scattered pieces of information into a coherent personal history, resulting in responses that remain at a generalized and fragmented level, failing to reflect a deep understanding of the user's personal experiences and long-term patterns. This limitation in understanding is further exacerbated by their singular reliance on information modalities. Most systems focus on analyzing the explicit textual content of the user's input—what the user "said" (content data)—while generally ignoring the rich signals implicit in the interaction process—how the user "said" (behavioral metadata). Hesitation, pauses, and frequency of modifications when inputting text, or choosing to interact frequently late at night, are valuable signals reflecting the user's true emotional state, focus, and cognitive load. The neglect of this implicit information by existing technologies is tantamount to abandoning another important dimension of understanding users' psychological state, making the overall profile of users one-sided and incomplete.

[0004] Therefore, although existing AI models can generate fluent and even seemingly empathetic language, this empathy is often superficial and immediate, based on a passive response to specific emotional words in the current discourse, rather than on a deep, dynamic model of the individual's unique psychological world. Due to a lack of understanding of the potential causal relationships between users' emotional fluctuations, key life events, and behavioral patterns, the guiding suggestions provided by the system are difficult to truly address the core issues, thus limiting its role in promoting long-term, meaningful self-growth and change. In conclusion, there is an urgent need in this field for a new technical solution that can overcome the bottlenecks in the depth of understanding and personalization capabilities of current dialogue systems, achieving truly dynamic, accurate, and effective self-help guidance for mental health through diachronic comprehensive analysis of users' multimodal information. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an online self-guided dialogue method for mental health. It aims to solve the technical problem that existing online psychological dialogue systems rely solely on the analysis of users' single-turn, single-modality (text) content, resulting in a lack of diachronic depth and multi-dimensional perspective in understanding users, and thus generating generic guidance responses that lack dynamic empathy.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an online mental health self-help guidance dialogue method. This method integrates users' explicit language expressions and implicit interactive behaviors through an innovative technical framework to construct a dynamically evolving personalized temporal psychological model, thereby achieving a deep and diachronic understanding of the user's state and generating a dialogue that truly has context awareness and adaptive guidance capabilities.

[0007] The technical solution provided by this invention includes the following steps: First, acquire the user's content data and behavioral metadata during the human-computer interaction process. In one embodiment, the content data may include unstructured text or voice input by the user, as well as structured psychological records filled out by the user (e.g., emotion self-rating, sleep logs, etc.). Meanwhile, the behavioral metadata may include at least one of the following data reflecting the user's interaction state: Input rhythm characteristics: such as the user's typing speed, the frequency of input pauses, and the number of times the text is modified.

[0008] Interaction timing characteristics: such as the time when the session occurs and its duration.

[0009] Environmental indicators: such as the category of background acoustic environment perceived through the device's microphone.

[0010] Next, based on the content data and the behavioral metadata, a personalized temporal psychological model corresponding to the user is constructed and dynamically updated. The construction and updating process of this model is one of the core aspects of this invention, and specifically may include: The acquired heterogeneous data is represented in a unified vector format. Specifically, content data is processed into content feature vectors, and behavioral metadata is processed into behavioral metadata feature vectors. These two vectors are then fused to generate a fused psychological state vector that comprehensively represents the user's overall psychological state at a specific moment.

[0011] A sequence model is used to process time series data composed of fused mental state vectors from multiple moments. This sequence model can capture the evolution of mental states over time and update a hidden state. The hidden state is a cumulative vector representation containing the user's entire interaction history from the initial moment to the present, which can be regarded as a mathematical representation of the user's long-term mental trajectory. This process can be summarized by the following formula: ; in, This indicates the hidden state at the current moment. This represents the fused mental state vector at the current moment. This indicates the hidden state at the previous moment. This represents the update function for the sequence model.

[0012] In a preferred embodiment, the personalized temporal psychological model further includes a multimodal causal inference graph. The graph's nodes include event nodes (such as "work deadlines") and emotion nodes (such as "anxiety") extracted from the content data, as well as behavioral pattern nodes (such as "high-frequency late-night interactions") abstracted from the behavioral metadata. Nodes are connected by weighted directed edges, representing the potential influence relationships between different psychological elements. The construction of this graph enables the system not only to know about changes in user state but also to infer the underlying causes of those changes.

[0013] Then, based on the personalized temporal mental model, a dynamic empathic context summary representing the user's current mental state is generated. This step transforms the model's complex internal states into usable background knowledge. For example, the system can analyze the deviation of the current hidden state from its historical baseline and, combined with a multimodal causal inference graph, trace back high-weight paths associated with the current emotional state to locate event nodes or behavioral pattern nodes that may trigger the state, ultimately forming a text summary describing the user's current state and its potential causes.

[0014] Finally, a guided dialogue is generated by combining the user's current input with the dynamic empathic context summary. In one embodiment, this step may specifically involve: combining the user's current input query with the generated dynamic empathic context summary to construct a contextualized prompt rich in deep background information; then, inputting the contextualized prompt into a large language model, which generates the final, highly personalized and guided dialogue content.

[0015] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the method described above.

[0016] The present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.

[0017] This invention provides an online self-guided dialogue method for mental health. It has the following beneficial effects: 1. This invention constructs a comprehensive user understanding model that far surpasses the capabilities of analyzing text alone by simultaneously collecting content data and behavioral metadata. This dual-input mechanism enables the system to capture potential inconsistencies between speech and behavior, thereby gaining a more acute insight into the user's unexpressed true emotional state or cognitive load, and improving the immediate accuracy of psychological state assessment.

[0018] 2. By leveraging sequence models (such as LSTM) in personalized temporal psychological models, this invention overcomes the core shortcomings of traditional dialogue systems: "immediacy" and "forgetfulness." This model treats each user interaction as a node on a continuous time series, achieving long-term accumulation and dynamic extraction of the user's psychological evolution trajectory.

[0019] 3. The multimodal causal inference graph intuitively models the potential influence relationships between specific life events, behavioral patterns, and emotional responses through dynamically updated edge weights. This capability allows the system to move beyond passively responding to user emotions and proactively explore potential triggers leading to the current state.

[0020] 4. This invention significantly improves the relevance and empathic quality of responses generated by Large Language Models (LLMs) through the intermediate step of "dynamic empathic context summarization." LLMs are no longer generated "blindly" based on isolated user input, but are instead "anchored" to a deep context rich in analysis of individual history, current state, and potential triggers. Therefore, the responses generated by the system are no longer generic comforting phrases, but truly personalized guidance that reflects a deep understanding of the user's unique experiences.

[0021] 5. Based on the in-depth analysis results of PCPM, the system can proactively and adjust the direction of the conversation through LLM. This "guidance" is not a rigid instruction, but is achieved through context-rich questions, reminders, or empathetic feedback, empowering users to more effectively explore themselves and solve problems, thereby substantially improving the intervention effect of self-help tools. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the system architecture and information flow of the present invention; Figure 2 This is a flowchart of the multimodal data feature vectorization process of the present invention; Figure 3 This is a schematic diagram of the internal structure of the personalized temporal psychological model of the present invention; Figure 4 This is a schematic diagram of the guided dialogue generation process of the present invention. Detailed Implementation

[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Please refer to the attached figures*-*. This embodiment of the invention provides an online self-help guidance dialogue method for mental health, which aims to generate highly personalized and dynamically empathetic guided dialogues by conducting in-depth and diachronic comprehensive analysis of users' language content and interactive behavior.

[0025] The client is the terminal through which the user directly interacts, such as a smartphone, tablet, or personal computer running a browser with a specific application installed. The server is responsible for executing the core computational tasks of the method of this invention. Its hardware environment may include one or more central processing units (CPUs), graphics processing units (GPUs) for accelerating deep learning model operations, high-speed memory, and storage units for persistently storing user data and models. The server-side software environment may be deployed on an operating system such as Linux and includes relational or non-relational databases, as well as deep learning frameworks such as TensorFlow or PyTorch. The client and server communicate via a secure transmission protocol (such as HTTPS) to ensure the secure transmission of sensitive user data.

[0026] In this embodiment, the overall information processing flow of the online mental health self-help guidance dialogue method may include the following steps: Step S101: Obtain the user's content data and behavioral metadata during the human-computer interaction process.

[0027] Step S102: Based on the content data and the behavioral metadata, construct and dynamically update a personalized temporal psychological model corresponding to the user.

[0028] Step S103: Based on the personalized temporal psychological model, generate a dynamic empathic context summary representing the user's current psychological state.

[0029] Step S104: Combine the user's current input with the dynamic empathy context summary to generate a guided dialogue.

[0030] The following will be combined with the appendix Figure 1 The system architecture shown above provides a detailed explanation of the above process.

[0031] During system operation, step S101 is first executed by the data acquisition module deployed on the client side. This module has a dual function: firstly, it receives explicit information actively input by the user through the interactive interface, i.e., content data, such as text messages typed by the user, recorded voice clips, or structured questionnaire answers submitted in guided exercises. Secondly, this module runs silently in the background to capture implicit signals unconsciously generated by the user during interaction, i.e., behavioral metadata, such as recording the user's typing speed, pause patterns, message editing frequency, and the time and duration of the interaction.

[0032] Once the data acquisition module obtains the content data and behavioral metadata, this data is securely transmitted to the server. The server-side feature processing module receives this raw data and performs a series of standardization and vectorization preprocessing operations. This step aims to transform heterogeneous data from different sources and in different formats into a unified mathematical representation that can be processed by subsequent models, namely, content feature vectors and behavioral metadata feature vectors.

[0033] After the feature vector is generated, it is passed to the core of the system, the personalized temporal psychological model construction and update module, which executes step S102. This module is key to achieving deep, diachronic user understanding in this invention. It receives the feature vector at the current moment and combines it with the user's historical model state stored in the database to dynamically update the model. Internally, this module uses a sequence model to capture the evolution trajectory of the user's psychological state over time and utilizes a graph structure to explicitly model the complex relationships between the user's emotions, life events, and behavioral patterns.

[0034] In each interaction round requiring a response, the Dynamic Empathy Context Summary Generation module initiates and executes step S103. This module acts as a translator, decoding the complex internal states of the personalized temporal mental model (e.g., high-dimensional hidden state vectors and graph structures) into concise, easily understood natural language text. This text, the Dynamic Empathy Context Summary, condenses deep insights into the user's current state, such as a comparison of current emotion with its long-term baseline, and potential triggers that might lead to this state.

[0035] Finally, the guided dialogue generation module executes step S1O4. It integrates the user's latest input query with the dynamic empathic context summary generated in the previous step to construct an information-rich, contextualized prompt. This prompt is then fed into a large language model. Based on this prompt containing in-depth background information, the large language model generates the final guided dialogue response. This response is not only linguistically fluent and natural but also reflects a profound understanding and care for the user's unique experiences and current situation. The generated response is ultimately transmitted back to the client via the network and presented to the user, thus completing a full interactive loop with dynamic empathy capabilities.

[0036] In this embodiment, step S101, the synchronous acquisition and preprocessing of multimodal user data, is achieved through a data acquisition module deployed on the client. For content data, when a user inputs text in the dialogue interface, the module obtains the text string by listening to the submission event of the input box; when a user uses the voice input function, the module calls the voice recognition service interface provided by the device operating system to convert the voice stream into text data in real time; when the system presents a structured psychological record form (such as a CBT thinking record form containing several questions) to the user, the module collects the answers to each field submitted by the user.

[0037] For behavioral metadata, the collection process is sequential. The data acquisition module has a built-in timer and event listener to calculate the time interval between each character input by the user, thus obtaining the typing speed and pause frequency. By comparing the final sent text with the draft during the input process, the number of deletions and modifications can be calculated. Each interaction with the server is timestamped for analyzing the interaction timing characteristics. In addition, after the user uses the device for the first time and agrees to authorization, the module acquires a background audio sample through the device's microphone interface for a very short sampling time. Then, using a small audio classification model running on the device (e.g., based on the MobileNet architecture), the audio sample is classified into one of several preset environmental categories (such as "quiet," "noisy traffic," or "human voice environment"), thus obtaining environmental indicative features. All raw data undergoes preliminary cleaning and formatting locally before transmission.

[0038] In this embodiment, one of the core steps in step S102 is the feature vectorization of the dual-stream data. The feature processing module on the server side is responsible for this task. For content data, the system uses a pre-trained Transformer model, such as BERT (Bidirectional Encoder Representations from Transformers), to encode the received text string into a dense vector of fixed dimensions, i.e., the text embedding vector. Simultaneously, the numerical values ​​in the structured psychological records (e.g., emotion ratings from 1 to 5) are subjected to min-max normalization, mapping their value ranges to intervals. These normalized values ​​are then concatenated into a structured data vector. The final content feature vector It is composed of these two parts:

[0039] in, This indicates a vector concatenation operation.

[0040] For behavioral metadata, the feature processing module processes the received metrics to generate behavioral metadata feature vectors. For example, continuous numerical features such as typing speed and pause frequency are Z-score standardized to ensure they follow a standard normal distribution with a mean of 0 and a standard deviation of 1. For categorical features such as environmental indicators, one-hot encoding is used to convert them into multi-dimensional binary vectors. All processed behavioral meta-features are concatenated to form a behavioral meta-feature vector. .

[0041] In this embodiment, the second core step of step S102 is the construction and dynamic updating of the PCPM. This step is performed by the Personalized Temporal Mental Model Construction and Update Module. First, this module converts the content feature vector generated in the previous step into a PCPM. With behavioral meta-feature vector By splicing the elements along the dimensional lines, a higher-dimensional fusion mental state vector is formed. This vector serves as the current time step. A panoramic snapshot of the user's state is input into the subsequent sequence model.

[0042]

[0043] Next, this module employs a Long Short-Term Memory (LSTM) network as a sequence model to capture the dynamic evolution of the fused mental state vector over time. At each time step, the LSTM network... take over As input, and based on its internal gating mechanism, it selectively forgets historical information, updates the current state, and outputs a hidden state. This hidden state It is a non-linear, highly condensed, cumulative representation of all psychological state information of a user from the beginning of use to the present. The core computational process of LSTM includes: Forgotten Gate: The decision is based on the cell state at the previous moment. Which information is discarded?

[0044] Input Gate: It determines which new information to store in the current cell state.

[0045] New information generated: A new candidate information vector is created based on the current input and the historical hidden state.

[0046] Cell status update: This combines historical information with new information to form the cell state at the current moment.

[0047] Output gate: It determines which information to output from the current cell state.

[0048] Hidden status update: This generates the final hidden state output.

[0049] Here, It is the Sigmoid function. It is the hyperbolic tangent function. It is the Hadamard product (element-by-element product). and These represent the input weight matrix and the cyclic weight matrix for different gated units, respectively. It is the corresponding bias vector.

[0050] At the same time, the module also dynamically maintains a multimodal causal inference graph. The nodes of this graph are generated as follows: event nodes (such as "arguing with family") are extracted by running a Named Entity Recognition (NER) model on the user's input text; emotion nodes (such as "sadness") are identified by a text sentiment classification model; and behavioral pattern nodes (such as "input hesitation pattern") are generated by using the user's historical behavioral meta-feature vector sequence. Clustering algorithms such as K-Means are applied to group similar behavioral patterns into cluster centers. The edge weights of the graph are updated according to a co-occurrence statistical formula that considers time decay to reflect the potential influence strength between different nodes.

[0051] In this embodiment, step S103: the generation of dynamic empathy context summary is completed by the dynamic empathy context summary generation module. This module first processes the latest hidden state... Perform decoding. For example, [the following is a list of steps / methods]: The module compares the value of a specific dimension (which is associated with emotional valence during model training) with the moving average of that dimension over a past period. If the current value is significantly lower than the average, it can be determined that the user's emotion is fluctuating negatively. Subsequently, this module... The system executes a graph search algorithm. For example, using the currently identified negative emotion node as the target, it applies a reverse version of Dijkstra's algorithm to find the source node pointing to it with the highest sum of path weights. This source node is considered a potential trigger causing the current emotional state. Finally, the system populates these analysis results (such as "emotional state: depressed, significantly below baseline" and "potential trigger: event 'project failure'") into a preset natural language template to form a dynamic empathic context summary.

[0052] In this embodiment, step S104: the generation and interaction of guided dialogue is handled by the guided dialogue generation module. The first step performed by this module is to construct a structured, contextualized prompt. This prompt text typically contains three parts: System Role Instructions: A fixed text that defines the AI's role as an empathetic and supportive psychological partner.

[0053] Dynamic Empathic Context Summary: A summary of in-depth analysis about the user's current state generated in the previous step.

[0054] Current input from the user: The latest text query sent by the user.

[0055] These three parts of text are concatenated into a complete prompt, which is then sent to a large language model (LLM). For example, a GPT (Generative Pre-trained Transformer) model that has been fine-tuned according to instructions. Upon receiving this context-rich prompt, the LLM can generate a response that is not only linguistically coherent and emotionally empathetic, but also strategically targeted. For instance, if the contextual summary points to a trigger, the LLM's response might guide the user to explore that trigger; or if the summary mentions a coping technique the user has successfully used in the past, the LLM might suggest that the user try it again. The final generated response is then transmitted back to the client via the server and displayed to the user, thus completing an interaction.

[0056] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An online self-help dialogue method for mental health, characterized in that, include: Acquire user content data and behavioral metadata during human-computer interaction; Based on the content data and the behavioral metadata, a personalized temporal psychological model corresponding to the user is constructed and dynamically updated; Based on the personalized temporal psychological model, a dynamic empathic context summary representing the user's current psychological state is generated; Guided dialogue is generated by combining the user's current input with the dynamic empathic context summary.

2. The method according to claim 1, characterized in that, The content data includes: unstructured text or voice input by the user, and structured psychological records filled in by the user.

3. The method according to claim 1, characterized in that, The behavioral metadata includes at least one of the following: input rhythm features, interaction timing features, or environmental indicator features.

4. The method according to claim 1, characterized in that, The steps of constructing and dynamically updating the personalized temporal psychological model include: vectorizing the content data and the behavioral metadata into feature vectors to obtain content feature vectors and behavioral metadata feature vectors; and fusing the content feature vectors and the behavioral metadata feature vectors to generate a fused psychological state vector representing the user's comprehensive psychological state.

5. The method according to claim 4, characterized in that, The step of constructing and dynamically updating the personalized temporal psychological model further includes: processing the time series of the fused psychological state vector using a sequence model to update a hidden state that encodes the user's historical psychological trajectory.

6. The method according to claim 1, characterized in that, The personalized temporal psychological model also includes a multimodal causal inference graph; the nodes of the graph include event nodes and emotion nodes extracted from the content data, as well as behavioral pattern nodes abstracted from the behavioral metadata.

7. The method according to claim 1, characterized in that, The step of generating a dynamic empathic context summary includes: analyzing the personalized temporal mental model to identify deviations between the user's current mental state and its long-term baseline; and inferring potential triggers that cause the deviation based on the personalized temporal mental model.

8. The method according to claim 7, characterized in that, The personalized temporal psychological model includes a multimodal causal inference graph, and the step of inferring potential triggers for the deviation includes: in the multimodal causal inference graph, tracing back the high-weighted path associated with the current psychological state to locate the event node or behavioral pattern node that triggered the state.

9. The method according to claim 1, characterized in that, The steps for generating guided dialogue include: combining the user's current input with the dynamic empathic context summary to construct contextualized prompts.

10. The method according to claim 9, characterized in that, The step of generating guided dialogue further includes: inputting the contextualized prompts into a large language model to generate the guided dialogue.