Psychotherapy-healing-oriented large model dialogue agent
By integrating user input and historical dialogue, using structured coding and feature fusion technology, matching psychological empathy strategies, and building situational empathy Prompt, the emotional understanding and safety problems of large language models in psychological healing are solved, and the depth and reliability of psychological support are improved.
Patent Information
- Application Number
- CN202510812666.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing large language models lack in-depth understanding and dynamic tracking of users' individual emotional states in psychological healing applications, making it difficult to provide high-quality and contextual psychological support, and there are security and ethical problems.
By integrating user input, historical dialogue summary and emotional memory, using structured coding and feature fusion technology, matching predefined empathy strategies based on emotional signals and semantic analysis, situational empathy Prompt is constructed, and combining multiple candidate generation and empathy quality evaluation mechanisms to ensure that the response strategy conforms to the logic of psychological intervention.
It realizes in-depth analysis of user emotions, cognitive status and potential needs, improves the depth and reliability of psychological support, ensures the professionalism and safety of responses, and overcomes the shortcomings of traditional models in emotion tracking, professional integration and safety control.
Smart Images

Figure CN120353899A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of healing conversations, and more specifically, to a large model dialogue agent oriented towards psychological healing. Background Art
[0002] With the accelerating social pace and increasing life pressure, the public's demand for mental health services has been continuously rising. Traditional psychological healing relies on professional counselors, but there are problems such as uneven distribution of resources, high costs, and "stigma", which limit the accessibility of services. In recent years, the development of artificial intelligence, especially large language models, has provided a new path for the field of mental health. Large language models have powerful language understanding and generation capabilities, can simulate human conversations, and provide a technical basis for constructing psychological support agents. The research and development of dialogue agents oriented towards psychological healing aims to make up for the shortage of professional resources, provide users with convenient, low-cost, and private psychological support, help relieve negative emotions, enhance psychological resilience, and guide the establishment of a positive cognitive model to improve the overall mental health level.
[0003] However, directly applying general large language models to psychological healing faces multiple challenges. Although current dialogue systems can generate fluent and seemingly empathetic responses, their "empathy" is mostly a patterned imitation, lacking in-depth understanding and dynamic tracking of the user's individual emotional state, making it difficult to capture emotional changes and potential needs in continuous conversations, resulting in superficial and un-targeted responses. In addition, general models do not deeply integrate psychological theories and professional empathy strategies, and the interaction logic is not optimized for the characteristics of psychological healing. It is often difficult to maintain a stable healer role in long-term conversations, and also lacks the ability to adjust response strategies in real time according to user feedback, and cannot provide high-quality and context-specific support in aspects such as empathy expression, emotion validation, and cognitive guidance. At the same time, how to ensure the safety and ethics of the generated content and avoid causing secondary harm to users due to inappropriate responses is also an issue that must be taken seriously during the application process.
[0004] Therefore, there is an urgent need for a large model dialogue agent solution oriented towards psychological healing that can deeply understand the user's emotional state, dynamically select and execute appropriate empathy strategies, and continuously optimize the response quality to improve the authenticity, consistency, and healing depth of the conversation. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed.
[0006] According to one aspect of the present application, there is provided a large model dialogue agent oriented towards psychological healing, which includes: a user input acquisition module for obtaining the text information currently input by the user; a dialogue and emotional memory extraction module for obtaining the historical dialogue summary and previous emotional state representation of the user; an emotional state update module for calculating and updating the user's emotional state representation based on the currently input text information, the historical dialogue summary, and the previous emotional state representation; a dynamic empathy response selection module for inputting the updated user emotional state representation and the currently input text information into an empathy strategy selection module to obtain a selected empathy strategy; a contextualized empathy Prompt construction module for constructing a contextualized empathy Prompt based on the selected empathy strategy, the updated user emotional state representation, and the currently input text information; and an empathy response generation module for inputting the contextualized empathy Prompt into a basic large language model to obtain an empathy response text.
[0007] Compared with the prior art, the large model dialogue agent oriented towards psychological healing provided by the present application integrates real-time input, historical dialogue summary, and emotional memory, and uses structured coding and feature fusion technologies to achieve in-depth analysis of the user's core emotions, cognitive states, and potential needs, overcoming the problem of superficial emotional understanding of general models. Based on the matching of emotional signals and semantic analysis results with predefined empathy strategy tags, it ensures that the response strategy conforms to the logic of psychological intervention, such as the precise invocation of professional strategies such as verification, support, and care, enhancing the professionalism of cognitive guidance. Through key element identification and instruction embedding, psychological theories are transformed into executable dialogue frameworks, combined with multi-candidate generation and empathy quality assessment mechanisms based on semantic embedding, which not only improves the fit between the response and the user's emotional state, but also guarantees content security through a double verification mechanism. This solution realizes the transformation from surface-level dialogue to deep psychological state intervention, and effectively solves the defects of traditional models in emotion tracking, professional integration, and security control by continuously optimizing emotional representation and strategy matching, significantly enhancing the depth and reliability of psychological support. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0009] Figure 1 It is a block diagram of a large model dialogue agent oriented towards psychological healing according to an embodiment of the present application.
[0010] Figure 2Data flow diagram of the large model dialogue agent oriented to psychological healing according to an embodiment of the present application.
[0011] Figure 3 Block diagram of the emotional state update module in the large model dialogue agent oriented to psychological healing according to an embodiment of the present application.
[0012] Figure 4 Data flow diagram of the scenario-based empathy Prompt construction module in the large model dialogue agent oriented to psychological healing according to an embodiment of the present application. Detailed implementation manners
[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0014] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0015] In view of the problems in the above background art, the present application is proposed. Figure 1 Block diagram of the large model dialogue agent oriented to psychological healing according to an embodiment of the present application. Figure 2 Data flow diagram of the large model dialogue agent oriented to psychological healing according to an embodiment of the present application. Specifically, as Figure 1 and Figure 2As shown in the figure, the large model dialogue intelligent agent 100 oriented to psychological healing according to the embodiments of the present application includes: a user input acquisition module 110 for acquiring the text information currently input by the user; a dialogue and emotional memory extraction module 120 for acquiring the historical dialogue summary and previous emotional state representation of the user; an emotional state update module 130 for calculating and updating the user's emotional state representation based on the currently input text information, the historical dialogue summary, and the previous emotional state representation; a dynamic empathy response selection module 140 for inputting the updated user's emotional state representation and the currently input text information into an empathy strategy selection module to obtain a selected empathy strategy; a contextualized empathy Prompt construction module 150 for constructing a contextualized empathy Propmt based on the selected empathy strategy, the updated user's emotional state representation, and the currently input text information; and an empathy response generation module 160 for inputting the contextualized empathy Propmt into a basic large language model to obtain an empathy response text.
[0016] Specifically, the user input acquisition module 110 is used to acquire the text information currently input by the user. It should be understood that the core of psychological healing lies in understanding and responding to the user's inner state, and the text information currently input by the user, whether it is the direct expression of emotions, the description of specific troubles, or the continuation of the previous dialogue, carries the most immediate and direct psychological activities and demand signals of the user at this moment. Accurately acquiring this information is the premise for the subsequent series of complex processes to be carried out accurately. As mentioned in the background art, existing general models lack in-depth understanding and dynamic tracking of the user's individual emotional state and are difficult to capture the emotional changes and potential needs in continuous conversations. Therefore, only by ensuring the complete and unbiased acquisition of the user's input in each round can a solid foundation be laid for subsequent in-depth understanding and personalized, contextualized responses, thereby solving the deficiencies of general models in psychological healing applications pointed out in the background art and ultimately achieving the goal of effective psychological healing.
[0017] Specifically, the user inputs the content they want to express through the user interface of the psychological healing dialogue agent, such as the chat window in a smartphone application or the text input box on a web page. For example, the user may input a text: "I have been feeling very anxious recently. Things at work are always overwhelming me, and I can't sleep well at night." When the user finishes inputting and selects to send (for example, clicks the "Send" button or presses the Enter key), the user input collection module is activated. This module first receives the original text string input by the user from the front-end user interface through a preset event listening mechanism or interface call. The preset or determined values here include but are not limited to: the encoding format for the system to receive text, which is set to the common UTF-8 encoding to ensure correct processing of various characters including Chinese; and the possible maximum input character length limit, for example, set to 2000 characters, which is determined by the system designer considering the processing ability of the large language model and the appropriate length of user expression to prevent abnormal input from affecting the system stability. After the module receives text information such as "I have been feeling very anxious recently. Things at work are always overwhelming me, and I can't sleep well at night", it will perform basic verification, such as checking whether the input is empty. Subsequently, the module directly uses this original text information as its processing output, that is, the text information currently input by the user.
[0018] Specifically, the dialogue and emotional memory extraction module 120 is used to obtain the historical dialogue summary and previous emotional state representation of the user. Correspondingly, psychological healing is not an isolated single communication, but a continuously evolving process. The historical dialogue summary of the user records the past communication topics, important events, and the responses of the agent, providing the necessary context for the current dialogue, avoiding the agent from repeating questions or giving responses inconsistent with the previous discussion, and ensuring the natural fluency of the interaction. More crucially, the previous emotional state representation, as a quantitative description of the user's emotional state after the previous interaction, is the cornerstone for understanding the user's emotional evolution trajectory, judging the progress of healing, and adjusting subsequent strategies. Only based on the history and previous state can the agent more profoundly understand the true meaning of the user's current input and provide more targeted and in-depth empathy responses.
[0019] Specifically, first, the module will access the conversation history database in the system background based on the unique identifier of the current user, such as the user ID. This database stores all the historical interaction records between the user and the agent. To improve efficiency and relevance, the system may not retrieve all the history, but extract the conversation of the most recent several rounds according to preset rules, or screen out the historical conversation fragments most relevant to the current input topic based on a certain algorithm (such as TF-IDF weighting or semantic similarity). Subsequently, these original conversation records will be fed into a pre-trained text summarization model, such as a summarization model based on the Transformer architecture, and the specific model selection and parameters are preset by the system developer. This model compresses the multi-round conversation into a refined text, that is, the historical conversation summary. For example, if the past three rounds of conversation discussed the user's insomnia and work stress, the summary might be "The user has been suffering from insomnia due to high work stress for a week." The length of this summary can be preset, for example, not exceeding 200 Chinese characters, to ensure information density and processing efficiency. In parallel, the module will obtain the previous emotional state representation. This previous emotional state representation does not arise out of thin air. It is calculated and stored by the emotional state update module in this solution at the end of the user's previous interaction with the agent. In other words, it is the result of the previous emotional state update. When specifically obtaining it, the module also retrieves the most recently saved emotional state representation from the user state database according to the user ID. This representation is in the form of a vector. For example, if the user's emotional state update information includes core emotions, cognitive states, potential needs, and state attributes, then this previous emotional state representation vector may be a numerical vector of a fixed dimension, and each dimension corresponds to these different aspects. For example, in a preset 10-dimensional vector, the first 5 dimensions may represent the probability distributions of five core emotions (such as happiness, sadness, anger, anxiety, calmness) using Softmax-normalized values, and the subsequent dimensions may represent the degree of cognitive distortion, the intensity of supportive needs, the level of self-efficacy, etc. These specific dimension divisions and meanings are predefined during system design. For example, the system may retrieve such a vector: [0.1, 0.6, 0.1, 0.15, 0.05, 0.7, 0.8, 0.3], where the first five values may represent the probability of happiness 0.1, the probability of sadness 0.6, etc., and the subsequent values represent other cognitive states, potential needs, and state attributes. This vector is directly output as the previous emotional state representation for use by subsequent modules.
[0020] Specifically, the emotional state update module 130 is used to calculate and update the user's emotional state representation based on the currently input text information, the historical conversation summary, and the previous emotional state representation. It should be understood that psychological healing emphasizes the continuous attention to and adaptation to the emotional changes of individuals. By integrating the user's latest expressions, the historical interaction background, and the representation information at the previous moment, this step can capture the evolution trajectory of the user's emotions and the current core needs, avoiding static or one-sided interpretations of the user's state. This comprehensive analysis generates a more accurate and personalized user emotional portrait, which not only reflects the user's current emotions, cognitions, and potential needs, but also provides a key basis for the subsequent selection of dynamic empathy responses and the generation of contextualized responses, making the interaction of the intelligent agent more targeted and healing, truly realizing the transformation from patterned imitation to deep empathy, and enhancing the authenticity, consistency, and depth of psychological healing conversations.
[0021] Specifically, in a specific example of the present application, Figure 3 is a block diagram of the emotional state update module in the large model dialogue intelligent agent oriented to psychological healing according to the embodiment of the present application. As Figure 3 shown, the emotional state update module 130 includes: a user input structured encoding unit 131, which is used to perform structured encoding on the currently input text information to obtain structured user input information; a historical conversation structured encoding unit 132, which is used to perform structured encoding on the historical conversation summary to obtain structured historical conversation summary; a user state feature integration unit 133, which is used to input the structured historical conversation summary, the structured user input information, and the previous emotional state representation into the feature fusion module of the trained user emotional state representation model to obtain user emotional state update information; an emotional state output unit 134, which is used to input the user emotional state update information into the output layer of the trained user emotional state representation model to obtain the updated user emotional state representation, wherein the output layer uses the Softmax function to vectorize the user emotional state update information.
[0022] Correspondingly, since the currently input text information is unstructured natural language, directly using it for feature fusion and calculation in complex machine learning models is not ideal in terms of both efficiency and effect. Structured encoding aims to transform this free text into a format that is more easily understood and processed by machines, extract the key semantics, emotions, intentions, etc. contained therein, and quantify or organize them into regular data forms, such as feature vectors or predefined attribute sets. In this way, the model can more deeply understand the nuances of the user input, rather than just staying at the surface vocabulary matching.
[0023] Specifically, the implementation of the user input structured encoding unit 131 is as follows: For the currently received input text information, "I have been feeling very anxious recently. Things at work always make me feel overwhelmed and I can't sleep well at night." First, this unit utilizes a pre-trained language model that is pre-loaded and configured, such as a pre-trained language model based on the Transformer architecture, like a specific version of the Chinese sentence vector model (such as ERNIE, etc. The specific model file path and parameters, such as the output vector dimension, are preset to 768 dimensions and are determined during system initialization). After receiving the input text string, inside the model, the text is first processed by a tokenizer into a sequence of tokens. These token sequences are then transformed into initial word vectors through the embedding layer of the model, and through the complex calculations of the self-attention mechanism of multiple layers of Transformer encoders and the feed-forward neural network, the context dependencies between words are captured. Finally, the model will adopt a specific strategy, such as average pooling the output vectors of all tokens in the last layer, to generate a dense vector with a fixed dimension representing the semantics of the entire input sentence. To obtain a richer structured representation, this unit parallelly invokes a pre-trained multi-label text classification model specifically for the psychological field. The label system of this model is predefined, for example, it includes multiple sub-dimensions such as anxiety, depression, stress, insomnia distress, self-negation, etc. For the input text, this classification model will output a vector consistent with the dimensions of the label system, where the value of each element ranges from zero to one, indicating the confidence level of the text belonging to the corresponding label. For example, the value corresponding to the anxiety dimension in the output vector may be 0.85, the insomnia distress dimension may be 0.7, and the depression dimension may be 0.2. In addition, a keyword extraction module is integrated. This module uses a preset algorithm (such as TF-IDF or TextRank, and the parameter such as the upper limit of the number of extracted keywords is preset to five) to identify and extract core words or phrases from the text, such as [anxiety, work pressure, overwhelmed, can't sleep well]. Finally, this semantic vector, multi-label sentiment confidence vector, and keyword list together constitute the structured user input information for subsequent modules to perform deeper analysis and processing.
[0024] It should be understood that although the historical conversation summary is already condensed information, its essence is still natural language text. In order to make more effective use of this historical information, it needs to be transformed into a structured or quantitative form that is more easily understood and processed by machines. Structured encoding can extract deep features such as the core semantics, key entities, and potential sentiment tendencies in the historical summary and represent them as numerical vectors or structured data with specific attributes. This not only improves the richness and accuracy of feature expression but also enables the historical information to be effectively aligned and fused with the current user input information and the previous sentiment state representation in the same semantic space or feature dimension.
[0025] Specifically, the implementation of the historical dialogue structured encoding unit 132 is as follows: Upon receiving the historical dialogue summary, "The user has suffered from insomnia for a week due to high work pressure recently." This unit will analyze and encode it using a preset natural language processing technique. A common implementation method is to utilize a pre-trained language model, such as the ERNIE model. The specific version (such as ERNIE-Gram) and parameters of this model are pre-determined and configured during the system development phase. When the historical dialogue summary is input into this ERNIE model, the model first segments the text into a sequence of tokens through its built-in tokenizer. These tokens are transformed into initial vector representations through the embedding layer and then processed through a multi-layer Transformer encoder structure. In each layer of the Transformer, the self-attention mechanism captures the long-range dependencies and context information among the tokens. Finally, the model generates a dense vector of a fixed dimension according to a preset pooling strategy, such as average pooling of all token output vectors. This vector is the deep semantic embedding representation of the historical dialogue summary. In addition, this unit also combines named entity recognition technology to extract the key entities and their attributes in the summary, such as identifying work pressure, insomnia, and duration of one week. Therefore, the final output structured historical dialogue summary is a data structure that includes the above semantic embedding vector and a list or dictionary containing these key entities and their attributes. This structured representation enables the historical information to be more precisely utilized by subsequent modules.
[0026] Correspondingly, it is difficult to comprehensively grasp the user's true state based on a single source of information, such as just the current user input. The structured historical dialogue summary provides the necessary interaction background and the user's long-term topics; the structured user input information reflects the user's latest direct expression; while the previous emotional state representation records the emotional situation at the previous moment. Fusing the features of these three through a specially trained model can generate a more accurate, rich, and personalized user emotional state update information than any single information source. Specifically, in an example of this application, the user emotional state update information includes core emotions, cognitive states, potential needs, and state attributes.
[0027] Specifically, the implementation of the user state feature integration unit 133 is as follows: First, a trained user emotional state representation model is obtained. The training of this model depends on a specially constructed dataset containing a large number of real or simulated psychotherapy conversations. In this dataset, each sample includes: a conversation history, the user's current round input, and the real user emotional state corresponding to this current input and history, annotated by a psychology expert or a specially trained annotator. During the training process, the model is a multi-input multi-output network architecture based on Transformer. Specific hyperparameters such as the specific number of layers and the number of attention heads are preset and tuned by the model designer in advance to use the structured user input information, the structured historical conversation summary, and the previous emotional state representation as inputs. In particular, at the initial stage of training, the previous emotional state representation is a zero vector or randomly initialized, and in subsequent rounds, the predicted output of the model at the previous moment is used. The optimization goal of the model is to make the output emotional state representation as close as possible to the real emotional state annotated by the expert. A combination of loss functions such as the cross-entropy loss function and the mean squared error loss function is used as the loss function, and multi-round iterative training is performed through optimization algorithms such as backpropagation and gradient descent until the model reaches the expected performance metrics on the validation set.
[0028] In the inference stage, the structured historical conversation summary, the structured user input information, and the previous emotional state representation are fed into the feature fusion module of this trained user emotional state representation model. A variety of fusion strategies are adopted inside this module. For example, feature vectors from different sources are concatenated and then non-linearly transformed through several fully connected layers; or a more complex attention mechanism, such as cross-attention or self-attention layers, is used to enable the model to learn the mutual influence and respective importance weights between different information sources, thereby dynamically integrating these features. For example, the attention mechanism can calculate the association strength between the anxiety keywords in the current input and the insomnia symptoms in the historical summary, and accordingly adjust the judgment of the user's core emotion. After being processed by the feature fusion module, the model will finally output the user emotional state update information, which is a structured data that clearly includes the evaluation and representation of the user's current core emotions, such as judged as 70% anxiety, 20% fatigue, cognitive state, such as the existence of catastrophic thinking, potential needs, such as seeking comfort and support, hoping to get solutions, and state attributes, such as high emotional intensity and low self-efficacy.
[0029] It should be understood that, in order to convert the unnormalized user emotional state update information generated inside the model into a standardized vector representation that is structured, interpretable, and convenient for downstream tasks, the present application will use the Softmax function to vectorize the user emotional state update information. That is, the deficiencies of existing models in empathy performance and dynamic understanding mentioned in the background art, and a clear and quantifiable emotional state representation is the basis for achieving accurate empathy. The application of the Softmax function can convert the raw scores output by the model into a probability distribution. This vectorized and probabilized representation not only makes the update of the user emotional state representation have an intuitive statistical meaning, but also provides a standardized input for the subsequent empathy strategy selection module, ensuring the accuracy and consistency of decisions, thereby enhancing the professionalism and effectiveness of the psychological healing agent.
[0030] Specifically, the implementation of the emotional state output unit 134 is as follows: The user emotional state update information at this stage is the raw prediction values or unactivated scores of the model inside for each emotional state dimension of the user, such as core emotion, cognitive state, potential needs, and state attributes. For example, if the pre-set core emotion categories include happy, sad, angry, anxious, and calm, then the part of the user emotional state update information regarding the core emotion is a vector containing five raw values. When this set of values is sent to the output layer as part of the user emotional state update information, the output layer will apply the Softmax function to it. The calculation method of the Softmax function is: First, perform an exponential operation on each raw value in the vector, that is, e to the power of this value; then, divide the result of each exponential operation by the sum of the results of all raw value exponential operations. Through such calculations, the above-mentioned raw values, after being processed by the Softmax function, will be converted into a probability distribution vector, for example, [0.1, 0.2, 0.05, 0.5, 0.15]. Each element value in this newly generated five-dimensional vector is between 0 and 1, and the sum of all element values is equal to 1, respectively representing the probabilities that the user is currently in the five core emotions of happy, sad, angry, anxious, and calm. Similarly, for the predictions of other classification-defined attributes in the user emotional state update information, their corresponding raw scores will also be converted into probability vectors through the Softmax function. Finally, the vectors of each part after being processed by Softmax will be combined to form the final updated user emotional state representation. This representation is a vector with a fixed dimension and normalized values, comprehensively reflecting the user's current multi-faceted psychological state, and the number of its dimensions is determined by the pre-set emotional state classification system. For example, 5-dimensional core emotion + 3-dimensional cognitive state + 2-dimensional potential needs + 1-dimensional state attribute.
[0031] Specifically, the dynamic empathy response selection module 140 is configured to input the updated user emotional state representation and the currently input text information into the empathy strategy selection module to obtain a selected empathy strategy. Correspondingly, the background art points out that the general model's empathy is superficial and it is difficult to achieve continuous and effective healing conversations. Only relying on the updated user emotional state representation can grasp deep information such as the user's core emotions and cognitive states, but it may ignore the immediate focus and specific context of the current conversation; only relying on the currently input text information may only achieve surface responses and cannot touch the user's potential needs and complex emotions. Therefore, inputting the two in combination enables the empathy strategy selection module to comprehensively consider the user's overall emotional portrait, thereby selecting the most appropriate, user-touching, and healing-value empathy method to achieve personalized interaction with both depth and pertinence and overcome the limitations of the prior art.
[0032] Specifically, in a specific example of the present application, the dynamic empathy response selection module includes: an emotional state signal extraction unit for extracting key signals from the updated user emotional state representation; a user input semantic parsing unit for parsing the currently input text information to obtain keywords, themes, emotional colors, and potential intentions; and an empathy strategy matching and decision-making unit for performing matching analysis on the key signals, keywords, themes, emotional colors, and potential intentions with a predefined empathy strategy label list of the empathy strategy selection module to obtain the selected empathy strategy.
[0033] It should be understood that the background art points out that it is difficult for general large models to accurately grasp the dynamic changes of the user's deep emotional state. Although the generated updated user emotional state representation contains multi-dimensional information, directly inputting it as a whole into subsequent modules may not be efficient enough or lack focus. For this reason, in the present application, by extracting key signals from the updated user emotional state representation, such as the core emotion with the highest score, the most prominent potential need, or specific state attributes (such as a high crisis risk level), the system can quickly focus on the psychological aspects that the user most urgently needs to solve or should be concerned about currently. This provides a more clear and prioritized guiding direction for subsequent empathy strategy selection and empathy response generation.
[0034] Specifically, the implementation of the emotional state signal extraction unit is as follows: The signal extraction process relies on the parsing and comparison of each section of this vector. For example, for the core emotion part (such as the first five elements of the vector, corresponding to the probabilities of happiness, sadness, anger, anxiety, and calmness respectively), the unit will traverse these five values, find the maximum value and its corresponding index. If the vector is [0.1, 0.2, 0.05, 0.5, 0.15], then the maximum value is 0.5, and the corresponding index points to anxiety. Therefore, one of the key signals extracted is the core emotion: anxiety, intensity: 0.5. Similarly, for the potential demand part, the one with a higher value is found, and the most prominent potential demand, such as potential demand: seeking comfort, intensity: 0.6 will be extracted. The same applies to the state attributes. All these identified most significant dimensions and their corresponding values or labels together constitute the set of key signals output, for example: {highest core emotion: (anxiety, 0.5), most prominent demand: (seeking comfort, 0.6), crisis risk: high}.
[0035] Accordingly, in order to achieve accurate and in-depth empathy responses, the intelligent agent not only needs to understand the user's comprehensive internal emotional state, that is, the user's emotional state representation, but also needs to carefully grasp the content and manner specifically expressed in the user's current conversation turn. It is mentioned in the background technology that existing models are difficult to dynamically adjust empathy strategies, and this parsing step is to make up for this deficiency. By extracting keywords, summarizing the topics discussed, emotional colors, and potential intentions in the user's discourse, it can provide richer and more direct context clues for the empathy strategy selection module.
[0036] Specifically, the implementation of the user input semantic parsing unit is as follows: First, for keyword extraction, the system uses a preset algorithm such as TextRank or a sequence annotation model fine-tuned based on a pre-trained language model (such as BERT). If TextRank is used, the input text will be constructed into a word graph, and the word importance score will be iteratively calculated to select a predetermined number (for example, set to extract 3 to 5) of the highest-scoring words as keywords, such as [anxiety, work, breathlessness, poor sleep]. Secondly, for topic identification, the system will call a pre-trained topic model, such as the Latent Dirichlet Allocation model, the training topic categories of which are pre-defined, including common psychological healing topics such as work pressure, academic distress, interpersonal relationships, and emotional management. After the input text is processed by the model, the probability distribution of the text on each preset topic will be output, or the most relevant topic will be directly determined, such as work pressure and sleep distress. Thirdly, for the analysis of emotional color, the system will use a pre-trained sentiment classifier, which uses a deep learning architecture (such as CNN or LSTM, or a classification model based on BERT) and is trained on a large-scale sentiment annotation corpus. Its emotional label system is predetermined, such as positive, negative, neutral, or more detailed such as strongly negative, slightly negative, etc. For the example input, the model may output strongly negative. Finally, for the identification of potential intent, the system uses an intent recognition model. The model can be rule-based or a classifier trained by machine learning methods on a data set labeled with user intent (such as pre-set categories such as confiding, seeking advice, venting emotions, and information consultation). The input text "I feel very anxious recently, and things at work always make me breathless and I can't sleep well at night" is analyzed by this model. The potential intent that may be identified is to confide emotions and seek understanding. Finally, these extracted keyword lists, topic attribution, emotional color judgment, and potential intent recognition results will be integrated to form a structured data containing this information for use by subsequent modules.
[0037] It should be understood that the background technology emphasizes the deficiencies of existing large models in empathy strategy selection and dynamic adjustment. By matching information such as key signals, keywords, themes, emotional colors, and potential intentions with a pre-defined list of empathy strategy tags containing psychological healing expertise, the system can, based on multi-dimensional and in-depth user insights, select the most appropriate and effective empathy approach in the current situation. This matching analysis ensures that the agent's response is not a general statement, but rather can provide personalized and healing-value interactions targeting the user's core pain points, immediate expressions, and potential needs, thereby enhancing the professionalism of the conversation and the user experience. Specifically, in an example of the present application, the pre-defined list of empathy strategy tags includes validation, feedback and clarification, support and care, perspective transformation and empowerment, information and guidance, and security.
[0038] Specifically, the implementation of the empathy strategy matching decision unit is as follows: First, it receives the integrated input information. These inputs include (key signals: {highest core emotion: (anxiety, 0.5), most prominent need: (seeking comfort, 0.6), crisis risk: high}; keywords: [anxiety, work, overwhelmed, poor sleep]; theme: work stress and sleep problems; emotional color: strongly negative, potential intention: pouring out emotions and seeking understanding). At the same time, the module internally stores or can access a pre-defined list of empathy strategy tags, which was pre-set by psychology experts during system design and contains various subdivided empathy strategies, such as validation, like "Deep_Emotional_Validation" (deep emotional validation), feedback and clarification, support and care, like "Expressing_Care" (expressing care), perspective transformation and empowerment, information and guidance, and security, like "Providing_Crisis_Resources".
[0039] The matching analysis process can be carried out using a rule-based decision system or a small machine learning model (such as a decision tree, support vector machine, or small neural network). If a rule system is adopted, a series of "if-then" rules will be preset. For example, a high-priority rule might be: If the crisis risk in the key signal is high, then directly select "Escalation_Protocol_Trigger" or "Providing_Crisis_Resources" in the safety category. Another rule might be: If the highest core emotion in the key signal is anxiety, the most prominent need is "seeking comfort", the potential intention includes expressing emotions, and the emotional color is strongly negative, then increase the matching scores for the strategies of "Deep_Emotional_Validation" in the verification category and "Expressing_Care" in the support and care category. Keywords and themes, such as anxiety, work stress, and poor sleep, will further refine the application directions of these strategies. For example, it enables "Deep_Emotional_Validation" to specifically verify that anxiety and insomnia related to work stress are normal. The system will comprehensively evaluate all activated rules and the support intensity of each input for different strategies, and finally select one or a group of empathy strategy labels with the highest weighted scores as the selected empathy strategies. For example, under the current input conditions, considering the high crisis risk, the system may first preferentially select the safety category strategy, Providing_Crisis_Resources. If multiple strategies are allowed to be selected or there are secondary strategy selections, based on signals such as anxiety, seeking comfort, and expressing emotions, "Deep_Emotional_Validation" in the verification category may be selected as the main healing empathy response strategy. Or, if the model is designed to output a single main strategy and there is an independent channel for safety processing, then "Deep_Emotional_Validation" may be selected because it can highly match multiple signals such as anxiety, seeking comfort, and strong negativity. Finally, these two strategy labels will be used as the selected empathy strategies and passed to the subsequent modules.
[0040] Specifically, the Situational Empathy Prompt Construction Module 150 is configured to construct a situational empathy Prompt based on the selected empathy strategy, the updated user emotional state representation, and the currently input text information. It should be understood that the selected empathy strategy indicates the macro direction of the response, while the new user emotional state representation reveals the user's deep and quantified psychological state. The currently input text information provides the specific context and content focus of the instant conversation. These three elements are jointly constructed into a situational empathy Prompt, which ensures that when the large language model generates a response, its expression can accurately align with the preselected strategy, profoundly resonate with the user's overall emotional portrait, and closely revolve around the details of the user's current utterance, thereby generating an empathetic response with both strategicity and rich personalized understanding and specific context care, which has a healing effect and overcomes the defects of the general model's responses being superficial and lacking pertinence.
[0041] Specifically, in a specific example of the present application, Figure 4 is a data flow diagram of the situational empathy Prompt construction module in the large model dialogue intelligent agent oriented to psychological healing according to the embodiments of the present application. As Figure 4 shown, the Situational Empathy Prompt Construction Module 150 includes: a Prompt template selection unit 151 for selecting an initial Prompt template that matches the selected empathy strategy; a key state element recognition unit 152 for recognizing key state elements from the updated user emotional state representation and the currently input text information; an instruction generation and embedding unit 153 for translating the key state elements and the selected empathy strategy into instructions and embedding them into the initial Prompt template; a persona and constraint injection unit 154 for injecting personas and constraint conditions to obtain the situational empathy Prompt.
[0042] Correspondingly, considering that directly using the original strategy tags, emotional representations, and user inputs to drive the large language model may result in responses that lack the expression emphasis and communication postures required by specific empathy strategies, or it may be difficult to ensure the professionalism and consistency of the responses. By predefining and selecting an initial Prompt template that matches the selected empathy strategy, the system can ensure that the generated responses are consistent with the selected strategy in terms of the core guiding logic and emotional tone, and at the same time provide a standardized entry for accurately filling in the details of the user's specific information and emotional state subsequently, thereby effectively improving the quality and pertinence of the empathetic responses.
[0043] Specifically, the implementation of the Prompt template selection unit 151 is as follows: The initial Prompt template selection unit accesses a pre-set Prompt template library within the system. This library consists of a series of carefully designed text templates, each of which is explicitly associated with one or a group of policy tags in a predefined list of empathy policy tags. These templates are jointly created by psychologists and natural language processing engineers, aiming to encapsulate the core communication patterns and professional considerations of different empathy policies, and placeholders for subsequent dynamically filling in specific user information are included within the templates. When this unit receives multiple selected empathy policies as shown in this example, namely "Providing_Crisis_Resources" and "Deep_Emotional_Validation", it will search in the template library according to the pre-set logic. Based on the corpus prompts provided by the user, the system will attempt to find a fusion template that can integrate the intentions of these two policies. This means that there may be dedicated templates pre-stored in the template library for specific, common policy combinations. For example, there is a template designed specifically for scenarios that require both providing crisis resources and conducting deep emotional validation. The retrieval key of this template may be a specific combination identifier of these two policy tags. If such a fusion template is successfully matched, the system will select it. For example, for the combination of "Providing_Crisis_Resources" and "Deep_Emotional_Validation", a selected fusion initial Prompt template may be as follows: "[System instruction: The user's current state simultaneously triggers a high-priority crisis response and deep emotional understanding. Please construct a response that first uses deep emotional validation to confirm the user's feelings, and then smoothly transitions to providing necessary crisis support information.] I have listened very attentively to your account of [placeholder for the key content summary of the user's current input text information], and I can understand how strongly this has made you feel [core emotion extracted from the updated user emotional state representation, such as anxiety] as well as [placeholder for other related feelings such as helplessness or exhaustion], especially when you talked about [placeholder for the specific details of the user's mentioned distress], that kind of [further specific description or impact of the emotion placeholder] feeling. Please know that it is completely understandable to have these emotions in the face of such a situation. At the same time, I am very concerned about your situation. If you feel that the pain at this moment is too much for you to bear alone, or if you need immediate help, here are some ways to provide you with support: [placeholder for specific crisis resources]. Please always remember that seeking help is a brave and important choice when you need it.] This selected fusion text template containing guiding instructions and placeholders is the initial Prompt template, which forms the basis for subsequent personalized Prompt construction.
[0044] It should be understood that the selected initial Prompt template contains placeholders that need to be dynamically filled, and these placeholders represent the user's unique and current psychological experiences and specific troubles. In order to make the subsequent generated situational empathy Prompts truly personalized and accurately empathetic, the core content that can specifically depict the user's state needs to be extracted from the most reliable and direct information sources. Updating the user's emotional state representation provides a structured and quantitative understanding of the user's internal emotions, cognitions, needs, and crisis levels, while the currently input text information directly reflects the user's explicit expressions and focus of attention at this moment. By identifying these key state elements, it is possible to ensure that the information filled into the Prompt template not only deeply reflects the user's internal state but also closely conforms to their current specific expressions, thus guiding the large language model to generate highly situational, targeted, and understanding empathetic responses.
[0045] Specifically, the implementation of the key state element recognition unit 152 is as follows: First, key information is extracted from the updated user emotional state representation. For example, based on a preset emotional state classification system, the unit will identify the item with the highest probability in the core emotion dimension, such as anxiety, with a probability value of 0.5. Similarly, the most prominent potential need is identified, such as seeking comfort, with a probability value of 0.6. At the same time, key state attributes are extracted, such as crisis risk: high. These preset dimensions and thresholds, such as what constitutes a "high" risk and what constitutes a "prominent" need, are determined based on psychological theories and practical experiences during system construction.
[0046] Second, from the currently input text information, specific expressions directly related to the user's current troubles are identified through natural language processing techniques (such as keyword extraction, semantic analysis, or dependency syntactic analysis). For example, from the sentence "I've been feeling very anxious lately. Things at work are always overwhelming me, and I can't sleep well at night", words describing the core emotional experience such as very anxious can be extracted, the specific event causing the trouble such as things at work, and the resulting somatic sensations and functional impacts such as being overwhelmed and not being able to sleep well at night. Finally, the information identified from these two sources is integrated into a set of key state elements. For this example, the output key state elements may be a structured object with the content: {core emotion: anxiety, emotion intensity: 0.5, main trouble source: things at work, specific physical feeling: being overwhelmed, related symptom: not being able to sleep well at night, main need: seeking comfort, crisis level: high}.
[0047] Accordingly, the abstract empathy strategy and personalized key state elements are transformed into explicit instructions for the large language model and seamlessly integrated into the preset framework of the initial Prompt template, so as to generate highly contextualized empathy responses that follow the macro strategy guidance and can accurately respond to the user's unique situation and deep needs, thereby enhancing the pertinence and healing effect of communication and avoiding the superficial responses that may be caused by general templates.
[0048] Specifically, the implementation of the instruction generation and embedding unit 153 is as follows: For the [placeholder for the key content summary of the user's current input text information] in the initial Prompt template, the system will combine the main source of trouble: things at work and the original input, and refine and fill it as: You are deeply troubled by things at work. For the placeholder of [the core emotion extracted from the updated user emotion state representation, such as anxiety], the anxiety in the key state elements will be directly used for replacement. For the placeholder of [other related feelings such as helplessness or exhaustion], the system will fill it as: and the resulting oppression and exhaustion according to the specific physical feelings: being out of breath and related symptoms: not sleeping well at night, in combination with the emotion association rule base (for example, the preset rule indicates that persistent physical discomfort and sleep disorders are usually accompanied by a sense of exhaustion or powerlessness). Then, the [placeholder for the specific details of the user's trouble mentioned] will be filled with a more specific expression, such as: Work matters are making you out of breath. And the [placeholder for the further specific description or impact of the emotion] will be filled by integrating the core emotion and symptoms as: This sense of anxiety is making it difficult for you to fall asleep at night and leaving you physically and mentally exhausted. Finally, the [placeholder for specific crisis resources] will select and fill in specific information from the preset and audited crisis resource database according to the crisis level: high and the strategy "Providing_Crisis_Resources", for example: XX Psychological Aid Hotline: 400 - XXX - XXXX, or visit the official XX mental health service platform at www.XXX. In this way, the original general template is transformed into a customized contextualized empathy Prompt that contains the user's specific situation information and explicit response instructions.
[0049] It should be understood that although the previous steps have constructed the Prompt content containing the specific user situation and strategy instructions, without clear personas and constraints, the responses of the large model may seem rigid, lack a specific sense of role, and may even deviate from the professional requirements of psychological healing. For example, over-providing personal opinions or exceeding the support scope. By injecting preset personas, such as a warm and professional healing assistant, and constraints, such as response length, prohibited words, and specific communication postures, the large model can be effectively guided to generate high-quality, safe and reliable responses that meet the specific requirements of the psychological healing scenario, thus completing the final contextualized empathy Prompt that can directly drive the large language model.
[0050] Specifically, the implementation of the persona and constraint injection unit 154 is as follows: Based on the previous step, the system retrieves the persona and constraints from a preset configuration library. The persona is predefined and aims to endow the dialogue agent with a stable and trustworthy virtual identity. For example, a preset persona could be: You are a senior psychological healing partner named 'Xin Wei'. Your communication style should be: warm, inclusive, patient, compassionate, and never judgmental. Your main task is to listen to, understand, and support the user's emotions, help them sort out their feelings, and provide general coping suggestions or resource information when appropriate. The constraints are also preset to regulate the behavioral boundaries of the large model. For example, the constraints may include: First, the response must strictly follow the requirements for content generation in the above system instructions. Second, the tone of the response must be consistent with the persona of 'Xin Wei'. Third, the length of the response should be controlled between three and five sentences, aiming for conciseness and refinement. Fourth, it is prohibited to provide any form of medical diagnosis or specific drug treatment advice. Fifth, it is prohibited to use any stimulating words that may cause negative emotions or panic in the user. Sixth, the conversation should focus on the emotions and troubles expressed by the user, avoiding unfounded topic switching. These preset persona descriptions and constraint condition texts will be combined with the initial Prompt template after embedding in a specific format as clear instructions for the large language model. Finally, the combined text is the contextualized empathy Prompt, which is the output of this step. For example, the final contextualized empathy Prompt may be as follows: "[System persona setting starts] You are a senior psychological healing partner named 'Xin Wei'. Your communication style should be: warm, inclusive, patient, compassionate, and never judgmental. Your main task is to listen to, understand, and support the user's emotions, help them sort out their feelings, and provide general coping suggestions or resource information when appropriate. [System persona setting ends][System constraint conditions start] First, the response must strictly follow the requirements for content generation in the following system instructions. Second, the tone of the response must be consistent with the persona of 'Xin Wei' (warm, inclusive, patient, compassionate, non-judgmental). Third, the length of the response should be controlled between three and five sentences, aiming for conciseness and refinement. Fourth, it is prohibited to provide any form of medical diagnosis or specific drug treatment advice. Fifth, it is prohibited to use any stimulating words that may cause negative emotions or panic in the user. Sixth, the conversation should focus on the emotions and troubles expressed by the user, avoiding unfounded topic switching. [System constraint conditions ends][Core conversation content instruction starts][System instruction: The user's current state simultaneously triggers a high-priority crisis response and in-depth emotional understanding. Please construct a response that first uses in-depth emotional verification to confirm the user's feelings and then smoothly transitions to providing necessary crisis support information.I listened very attentively to your outpouring about being deeply troubled by work matters. I can understand how intense the anxiety you feel is, and the oppression and exhaustion that come with it. Especially when you said that work affairs made you feel out of breath, and that this sense of anxiety made it difficult for you to fall asleep at night and left you physically and mentally exhausted. Please know that in the face of such a situation, it is completely understandable to have these emotions. At the same time, I am very concerned about your situation. If you feel that the pain at this moment is too much for you to bear alone, or if you need immediate help, here are some ways to provide you with support: XX Psychological Assistance Hotline: 400 - XXX - XXXX, or visit the official XX Mental Health Service Platform at www.XXX. Please always remember that seeking help is a brave and important choice when you need it. [End of the core conversation content instruction]” This complete Prompt will guide the large language model to generate an empathic response that meets all expectations.
[0051] Specifically, the empathic response generation module 160 is used to input the scenario-based empathic Prompt into the basic large language model to obtain an empathic response text. Correspondingly, the careful design of all the previous steps, including user understanding, strategy selection, key element extraction, template filling, and the injection of persona and constraints, jointly constructs a highly structured, information-rich, and goal-oriented scenario-based empathic Prompt. This Prompt guides the basic large language model on how to respond appropriately. By inputting this customized scenario-based empathic Prompt, it can ensure that the model strictly follows all the instructions contained therein, ensuring that the dialogue agent can truly implement the psychotherapeutic-oriented function designed by the patent.
[0052] Specifically, in a specific example of this application, the empathic response generation module includes: a candidate response generation unit, which is used to input the scenario-based empathic Prompt into the basic large language model to obtain a set of candidate responses; an empathy quality assessment unit, which is used to conduct empathy checks on each candidate response in the set of candidate responses to obtain a set of empathy check score values; and an optimal response selection unit, which is used to select the candidate response corresponding to the largest value in the set of empathy check score values as the empathic response text.
[0053] It should be understood that on the premise of following the previously constructed detailed instructions of the scenario-based empathic Prompt, a series of empathic response texts that may have slight semantic and expressive differences but all meet the core requirements are produced. This is to provide a selection space for possible subsequent screening or ranking mechanisms, thereby further improving the quality, adaptability, or diversity of the final response. In this application, the scenario-based empathic Prompt is input into the basic large language model to obtain a set of candidate responses.
[0054] Specifically, the implementation of the candidate response generation unit is as follows: The base large language model is based on the deep learning Transformer architecture, and its internal structure includes: an embedding layer that tokenizes the input Prompt text and converts it into word embedding vectors; one or more encoder (for encoder-decoder models) or only decoder (such as GPT) modules, each module internally stacked by a multi-head self-attention mechanism and a feed-forward neural network layer, supplemented by residual connections and layer normalization. The self-attention mechanism enables the model to capture the long-range dependencies and complex semantic associations between the elements within the Prompt. When the contextualized empathy Prompt is fed into the model, its processing process first converts the text sequence into an integer sequence, tokens or markers in the model's vocabulary through a tokenizer. These tokens are then mapped into high-dimensional vectors through the embedding layer. In a pure decoder model, these vector sequences directly enter the multi-layer Transformer decoder module. Each layer of the decoder calculates the association strength between each token in the sequence and other tokens through the self-attention mechanism, and combines the feed-forward network for non-linear transformation, so as to gradually understand the overall intention and detailed requirements of the Prompt. In the text generation stage, the model adopts an autoregressive approach, that is, predicting the next most likely token one by one. To obtain a set of candidate responses, the model does not simply select the token with the highest probability at each step during the decoding process, but introduces a random sampling strategy. For example, the nucleus sampling method can be used, where the model randomly samples from the smallest set of tokens whose cumulative probability reaches a certain preset threshold (such as ninety percent) at each step. By running this generation process multiple times, or using beam search and retaining multiple high-score sequences, the model can output a set of candidate responses containing multiple different expressions for the same contextualized empathy Prompt. For example, the set of candidate responses output may contain the following three texts: Candidate One: "I listened carefully to your troubles about work-related anxiety and exhaustion, and sensed your current difficulties. These emotions are completely understandable. If you feel it's hard to bear alone, remember there are resources here to support you: XX Psychological Assistance Hotline: 400-XXX-XXXX, and the XX official mental health service platform www.XXX. Seeking help is a brave step." Candidate Two: "Learning that work-related matters are making you so anxious and affecting your sleep, I can feel your hardship and exhaustion. Please accept these feelings at this moment, which is normal. If you need immediate support, these channels are always open for you: XX Psychological Assistance Hotline: 400-XXX-XXXX; XX official mental health service platform www.XXX. Remember, it's very important to seek help when needed." Candidate Three: "Hearing that you are deeply anxious due to work, to the point of being out of breath and unable to sleep at night, I deeply feel your heaviness. Please believe that it's reasonable for you to have these feelings.If the pain you're feeling right now seems overwhelming, the following resources can help: XX Psychological Support Hotline: 400-XXX-XXXX, and the official XX Mental Health Service Platform at www.XXX. Don't forget, reaching out for support is a positive step.
[0055] Accordingly, even if the base large language model generates responses based on detailed contextual empathy Prompts, there may still be differences in the depth, accuracy, and appropriateness of empathy conveyed by each candidate text. It is necessary to quantitatively score the empathy quality demonstrated by each candidate response. This ensures that the responses presented to users not only follow the instructions in terms of content but also meet professional standards for psychological healing at the level of emotional transmission, such as effectively validating the user's feelings, showing care, and aligning with the intent of the selected strategy.
[0056] Specifically, in a particular example of this application, the empathy quality assessment unit includes: a candidate response semantic embedding encoding sub-unit for performing semantic embedding encoding on each candidate response in the set of candidate responses to obtain a set of candidate response semantic embedding encoding vectors; an empathy strategy semantic embedding encoding sub-unit for performing semantic embedding encoding on the selected empathy strategy to obtain a selected empathy strategy semantic embedding encoding vector; a candidate response emotion recognition sub-unit for performing emotion recognition on each candidate response semantic embedding encoding vector in the set of candidate response semantic embedding encoding vectors to obtain a set of candidate response emotion recognition representation vectors; and a cosine similarity calculation sub-unit for performing similarity analysis on each candidate response emotion recognition representation vector in the set of candidate response emotion recognition representation vectors and the selected empathy strategy semantic embedding encoding vector to obtain the empathy check score value.
[0057] It should be understood that candidate responses and selected empathy strategies are unstructured data. To convert them into numerically computable vectors that can represent their deep semantics, and to objectively and quantitatively evaluate to what extent each candidate response conforms to and embodies the selected empathy strategy, thereby providing a solid mathematical basis for empathy checking, this application maps the high-dimensional sparse text information into a low-dimensional dense vector space by performing semantic embedding encoding on these two types of data.
[0058] Specifically, the implementation of the candidate response semantic embedding encoding subunit and the candidate response emotion recognition subunit is as follows: For each candidate response text in the set, the system will call a pre-trained sentence-level or paragraph-level semantic embedding model (for example, a pre-trained language model based on the transformer architecture, such as a model fine-tuned for specific tasks like BERT, RoBERTa, or Sentence-BERT, whose specific selection and parameters are preset and determined based on their performance in tasks such as semantic similarity calculation). The specific processing process of this model generally includes: First, the input text string is tokenized to convert it into a sequence of tokens that the model can recognize; then, these token sequences are converted into initial word vectors through the embedding layer of the model; next, these word vectors will pass through multiple layers of transformer modules, each layer containing a multi-head self-attention mechanism and a feed-forward neural network, and through these complex calculations, the internal context dependencies and deep semantic features of the text are captured; finally, the model will take the specific token of the last layer, the hidden state corresponding to the [CLS] token, to generate a dense vector of a fixed dimension, and this vector is the semantic embedding encoding vector of the candidate response. This process is executed for all candidate responses in the set, and finally, a set of candidate response semantic embedding encoding vectors containing multiple such vectors is obtained. In parallel, in order to also convert the policy labels into semantic vectors, the system first maps these policy labels to their pre-defined, more descriptive texts. For example, "Providing_Crisis_Resources" may correspond to the descriptive text: This policy aims to identify the crisis state of the user and promptly provide specific and actionable help resource information, such as hotline numbers, contact information of professional institutions, etc., to ensure the safety of the user. "Deep_Emotional_Validation" may correspond to: The core of this policy lies in deeply understanding and validating the emotions expressed by the user, making the user feel that their emotions are fully accepted and recognized, and emphasizing the reasonableness and universality of their feelings. These texts are concatenated into a longer text string, such as: "This policy aims to identify the crisis state of the user and promptly provide specific and actionable help resource information, such as hotline numbers, contact information of professional institutions, etc., to ensure the safety of the user. At the same time, the core of this policy lies in deeply understanding and validating the emotions expressed by the user, making the user feel that their emotions are fully accepted and recognized, and emphasizing the reasonableness and universality of their feelings" Then, these descriptive texts will be encoded through the exact same semantic embedding model used when processing candidate responses to obtain a selected empathy policy semantic embedding encoding vector of a fixed dimension.
[0059] Accordingly, since the candidate responses need to not only match the user's distress and the selected strategy in terms of semantic content, but also the emotional color they convey is crucial. Especially in the context of psychological healing, the emotional expression of the response is necessary. Therefore, in this step, the emotional tendency implicitly or explicitly expressed in each candidate response text is quantitatively characterized.
[0060] Specifically, the implementation of the candidate response emotion recognition sub-unit is as follows: For each candidate response semantic embedding encoding vector in this set, the system will input it into a pre-trained emotion recognition model. This emotion recognition model is a deep neural network. Its input layer is designed to receive vectors with the same dimension as the semantic embedding vector. It internally contains several fully connected layers, the activation function ReLU, and a final output layer. The output layer uses the Softmax function to normalize the output into a probability distribution. This emotion recognition model is trained on a large-scale text data labeled with emotion labels and can infer the corresponding emotion category from the input semantic representation. The preset emotion categories are defined according to the needs of psychological healing conversations. For example, they may include: empathic understanding, support and encouragement, neutral objectivity, confusion, accusation, etc. Assume there are five preset emotion categories. For each input candidate response semantic embedding encoding vector, the emotion recognition model will output a candidate response emotion recognition representation vector. The dimension of this vector is equal to the number of preset emotion categories (five dimensions in this example). Each element in the vector represents the probability that the candidate response belongs to the corresponding emotion category. For example, for the semantic embedding vector of candidate one, the emotion recognition model may output the emotion recognition representation vector [0.7, 0.15, 0.1, 0.03, 0.02], indicating that candidate one has a 0.7 probability of expressing empathic understanding, a 0.15 probability of expressing support and encouragement, and so on. This process will be executed one by one for all candidate response semantic embedding encoding vectors in the set, and finally form a set of candidate response emotion recognition representation vectors. It should be noted that this newly generated set of candidate response emotion recognition representation vectors and the selected empathic strategy semantic embedding encoding vectors generated in the previous step are two different types of vectors. The semantic vector of the strategy defines what should be said, while the emotion representation vector of the response reflects how the emotion of the response sounds. The combination of the two can comprehensively evaluate the quality of empathy.
[0061] Specifically, here, when calculating the cosine similarity between the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding coding vector, since the candidate response emotion recognition representation vector involves dialogue - type semantics, while the selected empathy strategy semantic embedding coding vector involves strategy - type semantics, their semantic spaces are misaligned, and when calculating the cosine similarity, it is necessary to ensure that the vector lengths of the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding coding vector are the same, which will also lead to inconsistent spatial density distributions. Therefore, it is expected to optimize this before calculating the cosine similarity.
[0062] More specifically, in a specific example of the present application, the cosine similarity calculation sub - unit is used to: calculate the associated gradient vectors between the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding coding vector to obtain a candidate response - selected empathy associated gradient vector and a selected empathy - candidate response associated gradient vector, that is: ; where is the candidate response emotion recognition representation vector, is the selected empathy strategy semantic embedding coding vector, and both are column vectors, is vector multiplication, is the transpose operation, is point - by - point subtraction by position, is the candidate response - selected empathy associated gradient vector, is the selected empathy - candidate response associated gradient vector, that is, by sharing the manifold through a common associated mapping to calculate the gradient operator on the manifold, so as to construct the local curvature for representing the spatial density distribution on the manifold through the manifold gradient constraint.
[0063] Based on the candidate response - selected empathy associated gradient vector and the selected empathy - candidate response associated gradient vector, perform a manifold dynamic focus mapping on the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding coding vector to obtain a manifold dynamic focus mapping coefficient, that is: ; where is point - by - point multiplication by position, is the vector two - norm, is the manifold dynamic focus mapping coefficient, and the manifold dynamic focus mapping coefficient performs spatial implicit association analysis under different spatial density distributions in a cosine - like manner through dynamic focus mapping based on the manifold gradient.
[0064] Based on the manifold dynamic focus mapping coefficient and in combination with the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding coding vector, perform an asymmetric spatial association on the cosine similarity between the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding coding vector to obtain an optimized cosine similarity as the empathy check score value, that is: ; where is the cosine similarity between and is a weight parameter, which is a hyperparameter selected within a preset range according to experience and determined through experimental tuning, used to balance the contributions of the basic cosine similarity and the dynamic adjustment term in the final empathy check score, is the vector length, is the optimized cosine similarity, that is, the empathy check score value.
[0065] That is to say, by analyzing the spatial implicit association under different spatial density distributions to enhance the mapping weight of the spatial alignment weight based on the dynamic focus, and using the analytic continuation of the key coefficient of the spatial association under asymmetry, to align the semantic space while keeping the spatial density distributions of the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding coding vector consistent, thereby improving the calculation accuracy of the cosine similarity.
[0066] It should be understood that through the previous complex empathy check process, the system has assigned a quantified empathy check score value to each candidate response. This score value comprehensively evaluates the fit of the candidate response with the selected empathy strategy at the semantic level, the suitability of the emotional color it conveys itself, and other potential quality dimensions. Therefore, selecting the one with the highest score is to select the response that is considered to be the most effective in conveying empathy, the most in line with the goal of psychotherapy, and the best in overall quality under the current evaluation system.
[0067] Specifically, the implementation of the optimal response selection unit is as follows: For the set of empathy check score values obtained after the empathy check on the above-mentioned set of candidate responses, it is [0.85, 0.79, 0.92], where 0.85 corresponds to candidate one, 0.79 corresponds to candidate two, and 0.92 corresponds to candidate three. The system will traverse and compare this set of score values to find the maximum value among them. In this example, the maximum value is 0.92. Once the maximum score value is determined, the system will, based on the position of this score value in the set or through a pre-established correspondence, find and select the candidate response associated with it. Since 0.92 corresponds to candidate three: Hearing that you are deeply anxious due to work, so much so that you are out of breath and can't sleep at night, I deeply feel your heaviness. Please believe that it is reasonable for you to have these feelings. If the pain at this moment makes it difficult for you to face, the following resources can provide help: XX Psychological Assistance Hotline: 400 - XXX - XXXX, and the official XX Mental Health Service Platform at www.XXX. Don't forget, taking the initiative to seek support is worthy of affirmation. Therefore, the system will output this candidate three text as the final empathy response text.
[0068] In summary, the large - model dialogue intelligent agent 100 oriented towards psychological healing according to the embodiments of the present application is elucidated. By integrating real - time input, historical dialogue summaries, and emotional memories, and adopting structured encoding and feature fusion technologies, it realizes in - depth analysis of the user's core emotions, cognitive states, and potential needs, overcoming the problem of superficial emotional understanding of general models. Matching predefined empathy strategy tags based on emotional signals and semantic analysis results ensures that the response strategy conforms to the logic of psychological intervention, such as the precise invocation of professional strategies such as verification - type, support - and - care - type, etc., enhancing the professionalism of cognitive guidance. Through key element identification and instruction embedding, psychological theories are transformed into executable dialogue frameworks, combined with a multi - candidate generation and empathy quality assessment mechanism based on semantic embedding, which not only improves the fit between the response and the user's emotional state, but also ensures content security through a double - verification mechanism. This solution realizes the transformation from surface - level dialogue to deep - level psychological state intervention. By continuously optimizing emotional representation and strategy matching, it effectively solves the deficiencies of traditional models in emotion tracking, professional integration, and security control, significantly enhancing the depth and reliability of psychological support.
[0069] As described above, the large model dialogue agent 100 oriented to psychological healing according to the embodiments of the present application can be implemented in various wireless terminals, such as a server with a large model dialogue algorithm oriented to psychological healing. In a possible implementation manner, the large model dialogue agent 100 oriented to psychological healing according to the embodiments of the present application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the large model dialogue agent 100 oriented to psychological healing can be a software module in the operating system of the wireless terminal, or can be an application program developed for the wireless terminal; of course, the large model dialogue agent 100 oriented to psychological healing can also be one of the many hardware modules of the wireless terminal.
[0070] Alternatively, in another example, the large model dialogue agent 100 oriented to psychological healing and the wireless terminal can also be separate devices, and the large model dialogue agent 100 oriented to psychological healing can be connected to the wireless terminal through a wired and / or wireless network, and transmit interaction information according to a predefined data format.
[0071] The above has described the various implementations of the present disclosure. The above description is exemplary and not exhaustive. And it is not limited to the disclosed implementations. Many modifications and changes are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described implementations.
Claims
1. A large model dialogue intelligent agent oriented towards psychological healing, characterized in that, Including: A user input acquisition module, which is used to obtain the text information currently input by the user; A dialogue and emotional memory extraction module, which is used to obtain the historical dialogue summary and the previous emotional state representation of the user; An emotional state update module, which is used to calculate and update the user's emotional state representation based on the currently input text information, the historical dialogue summary, and the previous emotional state representation; A dynamic empathy response selection module, which is used to input the updated user's emotional state representation and the currently input text information into an empathy strategy selection module to obtain a selected empathy strategy; A contextual empathy Prompt construction module, which is used to construct a contextual empathy Propmt based on the selected empathy strategy, the updated user's emotional state representation, and the currently input text information; An empathy response generation module, which is used to input the contextual empathy Propmt into a basic large language model to obtain an empathy response text.
2. The large model dialogue intelligent agent oriented to psychological healing according to claim 1, wherein The emotional state update module includes: a user input structured encoding unit, which is used to perform structured encoding on the currently input text information to obtain structured user input information; a historical dialogue structured encoding unit, which is used to perform structured encoding on the historical dialogue summary to obtain structured historical dialogue summary; a user state feature integration unit, which is used to input the structured historical dialogue summary, the structured user input information, and the previous emotional state representation into the feature fusion module of the trained user emotional state representation model to obtain user emotional state update information; an emotional state output unit, which is used to input the user emotional state update information into the output layer of the trained user emotional state representation model to obtain the updated user emotional state representation, where the output layer uses the Softmax function to vectorize the user emotional state update information.
3. The large model dialogue intelligent agent oriented to psychological healing according to claim 2, wherein The user emotional state update information includes core emotion, cognitive state, potential needs, and state attributes.
4. The large model dialogue intelligent agent oriented to psychological healing according to claim 1, wherein The dynamic empathy response selection module includes: an emotional state signal extraction unit, which is used to extract key signals in the updated user's emotional state representation; a user input semantic parsing unit, which is used to parse the currently input text information to obtain keywords, topics, emotional colors, and potential intentions; an empathy strategy matching decision unit, which is used to perform matching analysis on the key signals, keywords, topics, emotional colors, and potential intentions with the predefined empathy strategy label list of the empathy strategy selection module to obtain the selected empathy strategy.
5. The large model dialogue intelligent agent oriented to psychological healing according to claim 4, characterized in that, The predefined empathy strategy label list includes verification type, feedback and clarification type, support and care type, perspective conversion and empowerment type, information and guidance type, and security type.
6. The large model dialogue intelligent agent oriented to psychological healing according to claim 1, characterized in that The described situational empathy Prompt construction module includes: a Prompt template selection unit for selecting an initial Prompt template that matches the selected empathy strategy; a key state element identification unit for identifying key state elements from the updated user emotional state representation and the current input text information; an instruction generation and embedding unit for translating the key state elements and the selected empathy strategy into instructions and embedding them into the initial Prompt template; a persona and constraint injection unit for injecting a persona and constraint conditions to obtain the situational empathy Propmt.
7. The large model dialogue intelligent agent oriented to psychotherapy according to claim 1, wherein The described empathy response generation module includes: a candidate response generation unit for inputting the situational empathy Propmt into a basic large language model to obtain a set of candidate responses; an empathy quality assessment unit for performing empathy checks on each candidate response in the set of candidate responses to obtain a set of empathy check score values; an optimal response selection unit for selecting the candidate response corresponding to the largest value in the set of empathy check score values as the empathy response text.
8. The large model dialogue intelligent agent oriented to psychological healing according to claim 7, wherein The described empathy quality assessment unit includes: a candidate response semantic embedding encoding sub-unit for performing semantic embedding encoding on each candidate response in the set of candidate responses to obtain a set of candidate response semantic embedding encoding vectors; an empathy strategy semantic embedding encoding sub-unit for performing semantic embedding encoding on the selected empathy strategy to obtain a selected empathy strategy semantic embedding encoding vector; a candidate response emotion recognition sub-unit for performing emotion recognition on each candidate response semantic embedding encoding vector in the set of candidate response semantic embedding encoding vectors to obtain a set of candidate response emotion recognition representation vectors; a cosine similarity calculation sub-unit for performing similarity analysis on each candidate response emotion recognition representation vector in the set of candidate response emotion recognition representation vectors and the selected empathy strategy semantic embedding encoding vector to obtain the empathy check score value.
9. The large model dialogue intelligent agent oriented to psychological healing according to claim 8, characterized in that, The described cosine similarity calculation sub-unit is used to: calculate the association gradient vectors between the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding encoding vector to obtain a candidate response - selected empathy association gradient vector and a selected empathy - candidate response association gradient vector; based on the candidate response - selected empathy association gradient vector and the selected empathy - candidate response association gradient vector, perform a manifold dynamic focus mapping on the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding encoding vector to obtain a manifold dynamic focus mapping coefficient; Based on the manifold dynamic focus mapping coefficient and in combination with the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding encoding vector, perform an asymmetric space association on the cosine similarity between the candidate response emotion recognition representation vector and the selected empathy strategy semantic embedding encoding vector to obtain an optimized cosine similarity as the empathy check score value.
Citation Information
Patent Citations
Automatic prompt construction method based on man-machine conversation history and semantic retrieval
CN119441443A
Psychological accompanying dialogue method based on multi-agent collaboration and storage medium
CN119599029A
Memory optimization system for intelligent health care accompanying robot
CN119829294A
Intelligent query semantic understanding method based on information geometry and Riemannian manifold
CN119829740A
Large model dialogue system for psychotherapy healing
CN119830923A
Cited By
Multi-module cooperative intelligent role playing system
CN120542581A
A multi-module cooperative intelligent role-playing system
CN120542581B
Emotional dialogue generation method and system based on large model
CN120851041A
Adaptive emotion-sharing dialogue method based on quantitative emotion analysis
CN120910227A
Report interpretation method and system
CN121052260A