Personalized emotion support system evaluation method based on user internal world simulation
By simulating the user's inner world through a chain-like intelligent agent architecture, user profiles are obtained and multi-round dialogue evaluations are conducted. This solves the problem that existing technologies cannot capture the user's subjective experience, realizes personalized evaluation of the emotion support system, and provides automated and fine-grained evaluation results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-14
AI Technical Summary
Existing evaluation methods for emotion support systems fail to effectively capture users' subjective experiences and lack personalized evaluation tools, making it impossible to accurately measure whether the system meets users' unique needs.
By constructing a chain-like intelligent agent architecture, the system obtains user profiles to be evaluated and simulates the user's pre-defined chain-like intelligent agent architecture, including a user thinking agent, a user dialogue agent, and a user evaluation agent. It simulates the user's internal state and generates natural language responses, conducts multi-round dialogues, records dialogue history and internal state trajectories, and finally performs multi-dimensional scoring.
This study enables automated, reproducible, and fine-grained evaluation of the personalized support capabilities of emotion support systems, revealing the system's strengths and weaknesses in personalized support and providing reliable evaluation guidance for system optimization.
Smart Images

Figure CN121858412A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an evaluation method for a personalized emotion support system based on the simulation of a user's inner world. Background Technology
[0002] Emotional support systems aim to identify users' emotional states through multi-turn interactions and provide corresponding comfort and assistance. With the development of technologies such as large language models, these systems have made significant progress in generating fluent and empathetic responses. However, the core effectiveness of emotional support systems lies in their personalization capabilities—that is, their ability to dynamically adjust support strategies based on users' unique psychological characteristics and specific situational needs.
[0003] While personalization is crucial, existing assessment methods have fundamental limitations. The current mainstream assessment paradigm is evaluator-centric, judging the quality of a conversation from an external, seemingly objective perspective, failing to capture the user's subjective experience. Summary of the Invention
[0004] This invention provides a personalized emotion support system evaluation method based on user's internal world simulation, which addresses the shortcomings of existing personalized emotion support system evaluation methods and enables automated, reproducible, and fine-grained evaluation of the system's personalized support capabilities.
[0005] This invention provides a personalized emotion support system evaluation method based on user internal world simulation, comprising the following steps: Obtaining the emotion support system to be evaluated and a labeled user profile; inputting the user profile and the emotion support system into a preset chain-like intelligent agent architecture to obtain the evaluation result of the emotion support system output by the chain-like intelligent agent architecture, wherein the chain-like intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent: Through the user thinking agent, after the emotion support system generates a system response based on the user profile, it updates the user's internal state corresponding to the user profile according to the current dialogue history, the user profile, and the system response; Through the user dialogue agent, based on the updated user internal state and the user profile, it generates a natural language response to the emotion support system, and conducts the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed, obtaining a complete dialogue history and user internal state trajectory; Through the user evaluation agent, based on the complete dialogue history and the user internal state trajectory, it scores the response of the emotion support system on a preset dimension to obtain the evaluation result.
[0006] According to the present invention, a personalized emotion support system evaluation method based on user inner world simulation is provided. The method further includes: obtaining a seed profile, wherein the seed profile is real user data collected through questionnaires; inputting the seed profile into a preset large language model for expansion to obtain a diversified user profile output by the large language model; and annotating the diversified user profile by adding problem descriptions and support goals to obtain an annotated user profile.
[0007] According to the present invention, a personalized emotion support system evaluation method based on user internal world simulation is provided. The step of updating the user's internal state corresponding to the user profile according to the current dialogue history, the user profile, and the system response includes: inputting the current dialogue history, the user profile, and the system response into the user thinking agent to simulate psychological activities, and obtaining the updated user internal state output by the user thinking agent; wherein, the user internal state is a triplet, including: cognitive evaluation, emotional state, and dialogue goal.
[0008] According to the present invention, a personalized emotion support system evaluation method based on user internal world simulation is provided. The step of conducting the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed to obtain a complete dialogue history and user internal state trajectory, includes: in each round of dialogue, sending the natural language response generated by the user dialogue agent as the input for the new round of dialogue to the emotion support system; obtaining the new round of system response generated by the emotion support system based on the natural language response; updating the user's internal state through the user thinking agent based on the new round of system response, the updated dialogue history, and the user profile; iteratively executing the above process until a preset maximum number of dialogue rounds is reached or the user dialogue agent generates a dialogue termination statement; continuously recording and accumulating the interaction content of all dialogue rounds throughout the multi-round dialogue process to form a complete dialogue history; synchronously recording the user's internal state updated by the user thinking agent in all dialogue rounds, and constructing the user's internal state trajectory in chronological order.
[0009] According to the present invention, a personalized emotion support system evaluation method based on user internal world simulation is provided. The method involves scoring the response of the emotion support system on preset dimensions based on the complete dialogue history and the user's internal state trajectory to obtain evaluation results. The method includes: inputting the complete dialogue history and the user's internal state trajectory on preset dimensions into a user evaluation agent to obtain score data for each dimension output by the user evaluation agent; integrating and analyzing the score data for each dimension to generate evaluation results, wherein the evaluation results include a score summary and effectiveness indicators.
[0010] According to the present invention, a personalized emotion support system evaluation method based on user's inner world simulation is provided, wherein the preset dimensions include at least one of the following: emotion understanding, personalization and adaptability, goal achievement, credibility, and dialogue quality and security.
[0011] This invention also provides a personalized emotion support system evaluation device based on user internal world simulation, comprising the following modules: an acquisition module for acquiring the emotion support system to be evaluated and a labeled user profile; and an evaluation module for inputting the user profile and the emotion support system into a preset chain-like intelligent agent architecture to obtain the evaluation result of the emotion support system output by the chain-like intelligent agent architecture, wherein the chain-like intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent: through the user thinking agent, after the emotion support system generates a system response based on the user profile, it updates the user's internal state corresponding to the user profile according to the current dialogue history, the user profile, and the system response; through the user dialogue agent, based on the updated user's internal state and the user profile, it generates a natural language response to the emotion support system, and conducts the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed, obtaining a complete dialogue history and user internal state trajectory; and through the user evaluation agent, based on the complete dialogue history and the user internal state trajectory, it scores the response of the emotion support system on a preset dimension to obtain the evaluation result.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the personalized emotion support system evaluation method based on simulation of the user's inner world as described above.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the personalized emotion support system evaluation method based on simulation of the user's inner world as described above.
[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the personalized emotion support system evaluation method based on simulation of the user's inner world as described above.
[0015] This invention provides a personalized emotion support system evaluation method based on user internal world simulation. First, by acquiring labeled user profiles, a standardized testing benchmark is provided for the evaluation process, ensuring the diversity and authenticity of the evaluation subjects. Second, after inputting the user profiles and the emotion support system into a chain-like intelligent agent architecture, the user's internal state is dynamically updated by the user thinking agent, and then a user dialogue agent generates natural language responses that conform to the user's characteristics. After multiple rounds of interaction, a complete dialogue history and user internal state trajectory are formed. This process effectively simulates the dynamic cognition and emotional evolution of real users. Finally, the user evaluation agent performs fine-grained scoring on preset dimensions based on complete interaction data. The generated evaluation results reveal the specific performance of the emotion support system in personalized support, providing a reproducible, automated evaluation scheme that closely reflects the user's subjective experience for the optimization of the emotion support system. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the evaluation method for a personalized emotion support system based on simulation of the user's inner world provided by this invention.
[0018] Figure 2 This is a schematic diagram of a user profiling example provided by the present invention.
[0019] Figure 3 This is a schematic diagram of the consistency coefficient between the model evaluation and the manual evaluation provided by this invention.
[0020] Figure 4 This is a schematic diagram illustrating the discrimination capability of the evaluation system provided by the present invention.
[0021] Figure 5 This is a schematic diagram illustrating the evaluation effect of the evaluation system provided by the present invention.
[0022] Figure 6 This is a schematic diagram of a module of the personalized emotion support system evaluation device based on the simulation of the user's inner world provided by the present invention.
[0023] Figure 7 This is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0025] Existing technologies mainly fall into the following categories, each with its own significant drawbacks: Evaluation based on automated metrics: Metrics such as BLEU, ROUGE, and BERTScore rely heavily on reference responses. They cannot measure the true effectiveness of emotional support in open-ended, context-sensitive tasks, and in particular, they cannot determine whether the response truly matches the user's personality and immediate psychological state.
[0026] Human expert assessment: While human assessment is more flexible, it is costly and inefficient. Furthermore, due to the subjectivity of emotional support, the consistency among different assessors is usually low, making it difficult to reproduce the assessment results and unable to support rapid system iteration.
[0027] Expert evaluation based on large language models: In recent years, while using large language models as evaluators to replace human experts has provided scalable solutions, its essence remains a "one-size-fits-all" expert perspective. These methods judge solely based on the external dialogue context, ignoring the subtle internal states driven by user profiles that shape the user's dialogue experience. For example, the same general comforting response (such as "Don't be too hard on yourself") might evoke drastically different subjective feelings in a perfectionist and a more resilient user. Existing large language model evaluators might award a high score because the response expresses empathy, but fail to recognize its potential negative effects on a specific user.
[0028] In summary, existing technologies generally fail to simulate the user's internal cognition and emotional evolution during the dialogue process from the user's subjective perspective, resulting in insufficient evaluation of the personalized capabilities of emotional support systems. Therefore, there is an urgent need in this field for an automated evaluation method that can simulate the user's internal world and use the user as the judging subject, in order to accurately and reproducibly measure whether an emotional support system truly meets the unique needs of individual users.
[0029] This invention proposes a personalized emotion support system evaluation method based on the simulation of the user's inner world. It aims to achieve an automated, reproducible, and fine-grained evaluation of the system's personalized support capabilities by simulating the user's subjective perspective.
[0030] refer to Figure 1 , Figure 1This is a flowchart illustrating the personalized emotion support system evaluation method based on user's inner world simulation provided by the present invention, as shown below. Figure 1 As shown, the method includes the following: Step 101: Obtain the emotion support system to be evaluated and the labeled user profile; In this embodiment of the invention, an emotion support system (i.e., a supporter) to be evaluated is obtained. This system can be a dialogue model based on a large language model, such as an open-source or closed-source emotion support system. Simultaneously, labeled user profiles need to be obtained. These user profiles are diverse benchmark sets constructed using structured methods. The user profile includes four core attributes: demographic attributes (e.g., age, gender, occupation), preference attributes (e.g., personality traits, speech style), consultation-related attributes (e.g., problem description, emotional state, support goals), and scenario scripts (i.e., background stories used to constrain the user's reactions in specific situations).
[0031] Step 102: Input the user profile and the emotion support system into the preset chain-like intelligent agent architecture to obtain the evaluation results of the emotion support system output by the chain-like intelligent agent architecture.
[0032] The chain-based intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent.
[0033] In this embodiment of the invention, a chain-like intelligent agent architecture is used to simulate the user's internal world. This architecture consists of three intelligent agents: a user thinking agent, a user dialogue agent, and a user evaluation agent. During architecture operation, two memories are maintained: a supporter memory (which stores only visible dialogue history) and a user memory (which stores dialogue history and the user's internal state).
[0034] Step 102 above specifically includes the following steps.
[0035] Step 1021: Through the user thinking agent, after the emotion support system generates a system response based on the user profile, update the user's internal state corresponding to the user profile according to the current dialogue history, user profile and system response.
[0036] In this embodiment of the invention, the user thinking agent (e.g., using the GPT-4o model) updates the user's internal state based on the current dialogue history, user profile, and system response after the emotion support system generates each round of responses.
[0037] The internal state is a triplet, which includes cognitive appraisal (the user's interpretation of the response, such as "the suggestion is impractical"), emotional state (such as changing from anxiety to relaxation), and dialogue goal (the user's next goal, such as obtaining specific advice).
[0038] Step 1022: Through the user dialogue agent, based on the updated user internal state and user profile, generate a natural language response to the emotion support system, and conduct the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed, and obtain a complete dialogue history and user internal state trajectory.
[0039] User dialogue agents (e.g., using the GPT-4o model) generate natural language responses tailored to the user's personality based on updated internal states and user profiles. These natural language responses reflect changes in internal state, such as requesting specific solutions when feeling down.
[0040] The user's dialogue agent interacts with the emotion support system in multiple rounds, updating the dialogue history in each round until the maximum number of rounds (e.g., 15 rounds) is reached or the user actively terminates the dialogue. This process generates a complete dialogue history and the user's internal state trajectory, providing a basis for evaluation.
[0041] Step 1023: The user evaluates the intelligent agent and scores the response of the emotion support system on preset dimensions based on the complete dialogue history and the user's internal state trajectory to obtain the evaluation result.
[0042] In this embodiment of the invention, the user evaluation agent (e.g., using the Qwen3-235B model) evaluates the response of the emotion support system from multiple dimensions based on the complete dialogue history and internal state trajectory. Evaluation dimensions include emotional understanding (e.g., empathy), personalization and adaptability (e.g., appropriateness of response, adaptive strategies), goal achievement (e.g., problem-solving, emotion improvement), credibility (e.g., human similarity, engagement), and dialogue quality and security (e.g., redundancy, consistency, security).
[0043] The scoring uses a 5-point Likert scale (1 being the worst and 5 being the best). For example, in the response appropriateness dimension, a score of 5 indicates a depth of response combined with the user's background. The evaluation results are output in JSON format, including analysis explanations and specific scores, enabling automated and reproducible evaluation.
[0044] Currently, most methods for evaluating emotion support systems fail to effectively assess the system's personalized capabilities and lack automated evaluation tools from the user's subjective perspective. The method proposed in this invention constructs rich and diverse user profiles and employs a chain-like intelligent agent architecture to simulate the user's internal cognition, dialogue, and evaluation processes, achieving fine-grained and reproducible automated evaluation of emotion support systems from the user's perspective. This method effectively reveals the system's capabilities and shortcomings in providing personalized support, offering reliable evaluation guidance for developing emotion support systems that better meet user needs.
[0045] According to the present invention, a personalized emotion support system evaluation method based on user's inner world simulation is provided, the method further includes: Obtain seed profiles, which are real user data collected through questionnaires; The seed profile is input into the preset large language model for expansion, resulting in a diverse user profile output by the large language model. By labeling diverse user profiles, adding problem descriptions and support goals, we obtain labeled user profiles.
[0046] In this embodiment of the invention, real user data is collected through a questionnaire to form an initial seed profile. The questionnaire covers multiple core dimensions required to construct the user profile, including but not limited to demographic information (such as age, gender, and occupation), personality traits (such as those obtained through a simplified version of the MBTI test), a specific description of the current emotional distress or stressful events, and the type of emotional support desired. The seed profile collected in this way ensures that the initial data originates from real user scenarios.
[0047] To obtain a sufficiently large test set covering different user types and consultation scenarios, the limited number of seed profiles needs to be expanded. In practice, the collected seed profiles are used as input and submitted to a pre-defined large language model (e.g., GPT-4). Through pre-designed prompts, the large language model is instructed to perform diverse expansions based on the given seed profiles, while maintaining the logical consistency of core attributes. The expansion directions mainly include: Attribute value variation: Generate different values within a reasonable range for attributes such as demographics and personality traits (for example, expanding from "25-year-old programmer" to "30-year-old teacher" and "22-year-old student").
[0048] Scenario Derivation: Based on the problem description in the seed profile, related but different specific situations and background stories are derived (for example, from "high work pressure", scenarios with different focuses such as "tight project deadline", "tense colleague relationship", "career development confusion" are derived).
[0049] Batch processing of large language models can efficiently generate hundreds of diverse user profiles, significantly improving the coverage and diversity of the test set.
[0050] The diverse user profiles generated by large language models may lack clear or specific "problem descriptions" and "supported goals." Therefore, standardized annotation is needed to create annotated user profiles that can be directly used to drive chain-like intelligent agent architectures.
[0051] The annotation includes a problem description and a support goal. The problem description is used to clearly and structurally define the core emotional distress or problem the user is currently facing, such as "self-doubt and anxiety caused by recent consecutive project failures." The support goal is used to clarify the specific goals the user hopes to achieve in this emotional support dialogue, such as "hoping to obtain methods for cognitive restructuring, not just emotional comfort."
[0052] After annotation, each user profile becomes a structured data object containing all the static attributes (demographics, preferences, counseling-related attributes) and dynamic objectives (support objectives) required to drive the simulation. This set of annotated user profiles constitutes a benchmark set for evaluating the personalized capabilities of the emotion support system.
[0053] Through the embodiments of the present invention, a set of user profiles that includes real user data as a foundation and has sufficient diversity and clear evaluation orientation can be systematically and automatically constructed.
[0054] According to the present invention, a personalized emotion support system evaluation method based on user internal world simulation updates the user's internal state corresponding to the user profile based on the current dialogue history, user profile, and system responses, including: The current dialogue history, user profile, and system responses are input into the user's thinking agent to simulate mental activity, and the updated internal state of the user is obtained from the output of the user's thinking agent. The user's internal state is a triplet, which includes: cognitive appraisal, emotional state, and dialogue goal.
[0055] In this embodiment of the invention, in each round of dialogue, after the emotion support system (i.e., the supporter system) generates a round of system responses, the user thinking agent is activated. The input of the user thinking agent includes the current dialogue history, the user profile, and the system responses.
[0056] The current dialogue history represents the sequence of all interactions between the user and the emotion support system from the start of the dialogue to the current round. The user profile is a pre-built and labeled structured information containing demographic attributes, preference attributes, consultation-related attributes, and scenario scripts. The system response is the latest reply generated by the emotion support system in response to the user's previous statement.
[0057] The user-thinking intelligent agent is used to simulate psychological activities. Based on the above input information, the user-thinking intelligent agent simulates the inner reaction and thought process of a real user in a specific personality background and dialogue situation after receiving the system's reply.
[0058] After the user's thinking agent simulates mental activities, the output is the updated user internal state. This user internal state is a triplet data structure, specifically including the following three core components.
[0059] Cognitive appraisal refers to a user's subjective interpretation and rational judgment of the responses from the emotional support system. It reflects whether the user considers the response relevant, effective, and aligned with their expectations or values. For example, cognitive appraisal might be expressed as: "This advice sounds hollow and doesn't address my core problem," or "He finally understands my situation; this analysis makes sense." Cognitive appraisal is a crucial bridge connecting external stimuli (systemic responses) with internal emotional responses.
[0060] Emotional state: refers to the user's current emotional experience, which is the direct output of the emotional level after the simulation of mental activity. It quantitatively or qualitatively describes the change in the user's emotions after receiving a system response, such as from "anxiety" to "slight relaxation," or from "frustration" to "disappointment." The dynamic evolution of emotional state is an important basis for evaluating the effectiveness of emotional support (such as the emotion improvement dimension).
[0061] Dialogue goal: This refers to the specific objective a user hopes to achieve in the next interaction based on the current conversation's progress and internal feelings. It guides the direction of the user's subsequent dialogue agent's responses. For example, the dialogue goal might shift from "seeking empathy" to "wanting specific behavioral guidance" or "wanting to end the current topic." The dialogue goal ensures that the simulated user's dialogue behavior is purposeful and coherent.
[0062] For example, the user thinking agent is instantiated from a large language model (such as GPT-4o), and its focused and consistent reasoning is ensured by setting a low temperature parameter (such as 0.1). Based on the input user profile (such as ISTJ personality, perfectionism) and specific scenario script, combined with the current dialogue context, the user thinking agent performs in-depth analysis of the system's response and generates the aforementioned triplet state.
[0063] For example, for a user defined in the user profile as having "high achievement motivation and low tolerance for vague suggestions," when receiving the system reply "Don't worry, just relax," the user's internal state might be updated as follows, reflecting the simulated mental activity of the intelligent agent: Cognitive appraisal: "This response is too general and does not offer any valuable or actionable methods for relieving stress." Emotional state: "Irritable" (compared to the previous "anxious", the emotional value may have decreased due to the lack of effective help.)
[0064] The goal of the conversation is to "clearly request the other party to provide a specific and actionable set of stress management techniques or plans." Through the embodiments of this invention, the user-thinking intelligent agent can dynamically and personalized update its internal state after each round of dialogue, providing accurate and fine-grained internal perspective data for subsequent user response generation and final effect evaluation. This profile-based deep simulation enables the evaluation to capture personalized interaction details that general evaluation methods cannot discover.
[0065] According to the present invention, a personalized emotion support system evaluation method based on user's internal world simulation is provided. This method proceeds to the next round of dialogue based on natural language responses and the emotion support system, until multiple rounds of dialogue are completed, obtaining a complete dialogue history and the user's internal state trajectory, including: In each round of dialogue, the natural language response generated by the user's dialogue agent is sent to the emotion support system as input for the next round of dialogue. Obtain a new round of system responses generated by the emotion support system based on natural language responses; The intelligent agent updates the user's internal state based on the new round of system responses, the updated dialogue history, and the user profile. The above process is executed iteratively until the preset maximum number of dialogue rounds is reached or the user dialogue agent generates a dialogue termination statement. Throughout the multi-round dialogue, the interactive content of all dialogue rounds is continuously recorded and accumulated to form a complete dialogue history; Synchronously record the user's internal state updated by the user's thinking agent in all dialogue rounds, and construct the user's internal state trajectory in chronological order.
[0066] In this embodiment of the invention, the dialogue process is iterated in rounds, and each round follows a cyclical pattern of "system response - user thinking / dialogue agent - user response".
[0067] At the start of each round of dialogue, the user's dialogue agent generates a natural language response based on the user's current internal state and user profile. This response simulates typical expressions of a real user under their current psychological state and personality traits, such as requesting specific advice or expressing emotional reactions. After generation, this natural language response is sent directly to the emotion support system as input for the next round of dialogue.
[0068] After receiving a natural language response generated by the user's dialogue agent, the emotion support system generates a new system response based on its internal model logic. This system response is a reaction to the user's previous statement and aims to provide emotional support, advice, or empathetic feedback.
[0069] The user's thinking agent is then activated. Based on the latest system response, the updated dialogue history (i.e., the complete context including the current interaction), and the user profile, the agent simulates mental activity and outputs an updated internal state. This state is recorded in the form of a triple, including cognitive evaluation (the user's subjective interpretation of the system response), emotional state (current emotional experience), and dialogue goal (intention for the next interaction). This update ensures that the internal state dynamically reflects the progress of the dialogue.
[0070] The iterative process continues until any of the following conditions are met, reaching the preset maximum number of dialogue rounds (e.g., 15 rounds). This setting prevents the dialogue from looping indefinitely and ensures efficiency. The user dialogue agent generates explicit dialogue termination statements (e.g., "Thank you, I'm much better now" or "I don't want to continue this topic"). These statements are automatically generated by the agent based on its internal state (e.g., improved emotional state or goal achievement), simulating the behavior of a real user ending a dialogue.
[0071] Throughout the multi-round dialogue, the system continuously records the complete dialogue history and the user's internal state trajectory.
[0072] A complete dialogue history is used to represent each round of user natural language responses and system responses from the emotion support system, which are accumulated sequentially to form a complete dialogue history. This history is stored in a round-by-round sequence, fully reproducing the entire interaction process.
[0073] The user's internal state trajectory is used to represent the user's internal state (including cognitive evaluation, emotional state, and dialogue goal) updated by the user's thinking agent in each round, recorded in chronological order and constructed into the user's internal state trajectory. This trajectory is presented in the form of a state sequence, clearly showing the dynamic changes in the user's psychological state throughout the dialogue, such as the emotional state gradually evolving from "anxiety" to "calm".
[0074] Through the embodiments of the present invention, not only are dialogue records that can be used for evaluation generated, but the internal state evolution chain from the user's perspective is also captured, enabling the evaluation to be based on deep psychological simulation rather than surface text, thereby achieving fine-grained analysis of the personalized capabilities of the emotion support system.
[0075] According to the present invention, a personalized emotion support system evaluation method based on user's internal world simulation is provided. Based on a complete dialogue history and the user's internal state trajectory, the response of the emotion support system is scored on preset dimensions to obtain evaluation results, including: Input the complete dialogue history and user internal state trajectory of the preset dimensions into the user evaluation agent to obtain the score data of each dimension output by the user evaluation agent; The scoring data from each dimension are integrated and analyzed to generate evaluation results, which include a scoring summary and effectiveness indicators.
[0076] In this embodiment of the invention, after completing multiple rounds of dialogue and generating a complete dialogue history and user internal state trajectory, the data is submitted to the user for evaluation of the intelligent agent.
[0077] The evaluation is not a single overall score, but rather based on a predefined, multi-dimensional evaluation index system. These predefined dimensions comprehensively cover key aspects of the effectiveness of the emotion support system, specifically including the following:
[0078] Emotional understanding: Focus on assessing the system's empathy capabilities.
[0079] Personalization and Adaptability: This is further divided into responsiveness (the degree to which responses fit the user's background) and adaptive strategies (the ability to dynamically adjust support methods based on the conversation).
[0080] Goal achievement: This is further divided into problem-solving (the extent to which effective solutions are provided) and mood improvement (the positive impact on the user's emotional state).
[0081] Credibility includes human-likeness (naturalness of language) and engagement (the ability to maintain user interest).
[0082] Dialogue quality and security: including redundancy (avoiding repetition), consistency (unified information before and after), and security (avoiding harmful content).
[0083] The user evaluation agent (typically employing a high-performance large language model such as Qwen3-235B, with a temperature parameter set to 0.0 to ensure evaluation consistency) is assigned the role of a "simulated user." Based on its deep understanding of the complete dialogue history (plaintext records of all dialogue rounds) and the user's internal state trajectory (the dynamic sequence of changes in user cognition, emotion, and goals after each round), it evaluates the performance of the emotion support system in each round and each preset dimension from the perspective of this simulated user. The evaluation uses a standardized 5-point Likert scale (e.g., 1 point represents "very poor," and 5 points represent "very good"), and generates specific score data for each dimension.
[0084] After obtaining the raw scoring data for each dimension, it is necessary to integrate and analyze it to generate the final evaluation results. These results not only include an intuitive summary of the scores but also validity indicators that validate the scientific validity of the evaluation method itself.
[0085] Statistical analysis is performed on the scoring data across various dimensions to calculate the average score and standard deviation of the emotion support system for each dimension, generating an easy-to-understand scoring summary report. This report can be output in structured formats such as JSON, showcasing the strengths and weaknesses of the emotion support system in different aspects.
[0086] To demonstrate the discriminative power and reliability of this evaluation method, the evaluation results must also include key validity indicators. These indicators are derived through in-depth calculations of the scoring data and mainly include the following indicators.
[0087] Model Separation Ratio (MSR): This metric measures whether the performance difference between different emotion support models is significantly greater than the fluctuation (i.e., noise) in the ratings of the same model on different user profiles. A high MSR value indicates that this evaluation method can effectively distinguish the performance of different models.
[0088] Model Consistency Coefficient (MAC): This metric quantifies the degree of consistency in evaluation results when comparing different models. A high MAC value indicates that the evaluation method is stable and reliable, and can produce consistent evaluation criteria.
[0089] Correlation with Human Assessment: The consistency between the assessment results of this method and human judgment was verified by calculating the Pearson correlation coefficient between automated ratings and real human expert ratings. Experiments show that the correlation coefficient can be significantly improved after incorporating user profiles and internal state information, confirming its practicality and effectiveness as an automated assessment tool.
[0090] Through the embodiments of the present invention, statistically significant evaluation results are output, which not only provides fine-grained diagnostic basis for the performance of the emotion support system, but also ensures the credibility and scientific nature of the evaluation process itself.
[0091] According to the present invention, a personalized emotion support system evaluation method based on user's internal world simulation is provided, and the preset dimensions include at least one of the following: emotion understanding, personalization and adaptability, goal achievement, credibility, and dialogue quality and security.
[0092] In this embodiment of the invention, the emotion understanding dimension primarily assesses the ability of the emotion support system to identify and interpret the user's emotional state and respond appropriately. The core of this dimension lies in whether the system can accurately perceive the emotional changes revealed by the user during the conversation and express empathy and understanding through language. For example, when a user expresses anxiety or frustration, the system should be able to identify these emotional signals and provide an emotionally supportive response, rather than offering generic or indifferent feedback. In specific implementations, the user evaluation system analyzes whether the system's responses highly match the user's emotional needs based on the complete conversation history and the user's internal state trajectory (especially the sequence of emotional state changes). The scoring focuses on the depth and accuracy of empathic expression and the timeliness of response to the user's emotional fluctuations, thereby ensuring that the system truly "understands" the user rather than merely providing a superficial response.
[0093] The personalization and adaptability dimension focuses on the alignment between the system's responses and the user's unique background, preferences, and real-time needs, as well as the system's flexibility in dynamically adjusting support strategies based on the conversation's progress. This dimension can be further subdivided into response appropriateness (whether the response content aligns with the user's personality traits in the user profile, such as personality characteristics and consultation goals) and adaptive strategies (whether the system can proactively switch support methods based on changes in the user's emotions or goals, such as shifting from emotional comfort to providing specific advice). During implementation, the user assessment intelligence combines static attributes in the user profile (such as demographic information and MBTI type) with dynamic internal state trajectories (such as the evolution of cognitive evaluation and conversation goals) to comprehensively evaluate the personalization of the system's responses and the rationality of strategy adjustments. For example, for a perfectionist user, whether the system avoids offering general comfort and instead provides structured advice is a core criterion for scoring.
[0094] The goal achievement dimension measures the system's actual effectiveness in helping users solve core problems and improve their emotional state. This dimension includes problem-solving (whether the system helps users clarify the root cause of the problem or provides feasible solutions) and emotional improvement (whether the dialogue has a positive impact on the user's emotional state, such as shifting from negative emotions to calm or positive ones). In practice, the user evaluation AI analyzes changes in the user's internal state trajectory (such as the evolution of emotional values from low to high) and the achievement of dialogue goals (such as whether the user shifts from "seeking empathy" to "obtaining concrete methods") to determine whether the system's response effectively propels the user toward the supported goals. The scoring criteria include the magnitude of emotional improvement, substantial progress in problem-solving, and the consistency of goal achievement.
[0095] The credibility dimension assesses the system's ability to respond in terms of naturalness of language, human-like delivery, and maintenance of user engagement. This dimension encompasses both human-likeness (whether the response is fluent, natural, and similar to human conversation) and engagement (whether the system can maintain the user's willingness to engage in conversation through interaction, avoiding monotony or digression). In implementation, the user evaluation AI will assess the language quality, interactive appeal, and conformity to everyday communication habits from a simulated user perspective. For example, whether the system uses vivid and appropriate expressions to stimulate the user's desire to continue the conversation, or whether mechanical repetition leads to a decline in user interest, are key scoring factors.
[0096] The dialogue quality and security dimensions ensure the reliability of the system's responses in terms of overall quality and ethical compliance. This includes redundancy (avoiding repetitive or templated responses), consistency (maintaining consistency in information, attitude, and personality throughout the dialogue), and security (eliminating harmful, offensive, or unprofessional content). In specific evaluations, the user evaluation AI checks the dialogue history for contradictory information, repetitive patterns, or risky content to assess the system's robustness and security. For example, whether the system repeatedly uses the same phrases in multiple rounds of dialogue, or whether it uses expressions that violate user values, directly affects the score for this dimension.
[0097] Through the comprehensive evaluation across multiple dimensions described above, this invention can automatically generate fine-grained scoring results from the perspective of simulating the user's inner world. The various dimensions complement each other, jointly constructing a comprehensive and scientific evaluation framework that not only reveals the advantages and disadvantages of emotion support systems in personalized support but also provides clear and actionable guidance for system optimization.
[0098] This invention constructs a framework for a psychological counseling human-computer interaction system based on a closed loop of user profile construction, dialogue simulation, and evaluation feedback. The overall process is divided into three core stages: User Profile Collection, User-Supporter Conversation Simulation, and User-as-a-Judge Evaluation.
[0099] During the user profile collection phase, a dataset containing 100 representative user profiles was constructed, covering 28 professions and 15 typical scenario types. Each user profile consists of multi-dimensional attributes, including: demographic attributes such as age, gender, and occupation; consultation-related attributes describing the user's current problem, emotional state, interpersonal relationships, and goals; preference-related attributes covering personality traits (such as ISFJ), lifestyle habits, interests, language style, and listening preferences; and scenario scripts providing descriptions of psychological states in specific situations to trigger real-life dialogue scenarios.
[0100] Taking "psychologists" as an example, they face the dual pressures of experimental failures and team conflicts. They are introverted, rational, and perfectionist, and have high professional requirements.
[0101] In the user-supporter dialogue simulation phase, based on the aforementioned user profile, the system simulates a dialogue between the UserThink and UserTalk agents, with multiple EmpatheticSupport Agents simulating responses. The dialogue process is presented as a multi-turn interaction, reflecting emotional understanding, personalized responses, and adaptive strategies, for example: A user expressed anxiety: "Recently, the project has been under tight deadlines, and there are always people on the team shirking their responsibilities." The agent responded from an empathetic perspective: "I understand your passion for your work, and I've also heard you have a white cat..." The user then questioned the agent's professionalism, and the agent enhanced its credibility by demonstrating reasoning ability through observation (such as the cat's patient behavior). The dialogue also highlighted key mechanisms, such as "learning personalization and general adaptation" and "maintaining individuality (ISFJ, emotional introversion)".
[0102] In the evaluation phase, where users act as judges, after the simulated dialogue ends, they subjectively evaluate the interaction process from the perspective of "judges." The evaluation dimensions include: emotional understanding: whether empathy is demonstrated; personalization and adaptability: whether the response is appropriate and the strategy is flexible; goal achievement: whether the problem can be effectively solved and emotions regulated; credibility: whether it has a human-like naturalness and dialogue coherence; and dialogue quality: whether there is redundancy, consistency, and security.
[0103] Each metric is scored (1–5 points), generating an interaction evaluation report. The report's charts illustrate the performance distribution of the benchmark Large Language Model (LLM) across different tasks, and a visualization curve tracking the user's internal state reflects the changing trends of user emotions as the conversation progresses.
[0104] The following describes an example of a practical application of the personalized emotion support system evaluation method based on user's inner world simulation provided by this invention. The specific steps are as follows.
[0105] Step 1: Build user profiles.
[0106] Step 1.1: Collect real user data through questionnaires to create seed profiles; Real user data was collected through online questionnaires to create seed profiles. The questionnaires covered demographic information (age, gender, occupation), personality traits (MBTI type, behavioral habits), preference information (interests, speech style), and specific problems currently faced.
[0107] Step 1.2: Expand and enrich the collected seed profiles using large language models such as GPT-4. Specifically, generate more diverse user scenarios based on the seed profiles to ensure coverage of different age groups (from teenagers to the elderly), occupational types (covering 28 different occupations), personality traits (the complete 16 MBTI types), and counseling scenarios (including 16 types of issues such as work stress, academic challenges, and interpersonal relationships).
[0108] Step 1.3: Clearly define the problem description and support objectives for each user profile. Finally, construct a benchmark set containing 100 representative user profiles, each with complete structured information: A complete user profile P_u is defined as a set of four attributes: P_u = {D, P, C, S} D (Demographic attributes): including age, gender, occupation, etc., providing users with specific social background.
[0109] P (Preference-related attributes): These include personality traits such as MBTI, habits, hobbies, and speaking styles, which shape a user's unique behavioral patterns.
[0110] C (Consultation-related attributes): This includes a description of the problem to be solved, the current emotional state, consultation goals, social relationships, etc., which encode the user's psychological background.
[0111] S (Scenario Script): A detailed background story that specifies the user's possible reactions, inner goals, and expected responses to different supporters in a specific situation, thereby constraining the agent to produce behaviors that conform to the real-world context.
[0112] refer to Figure 2 , Figure 2 This is a schematic diagram of a user profiling example provided by the present invention.
[0113] like Figure 2 The image presents a virtual user profile named "Lin Jing," used to simulate a human-computer psychological support interaction scenario. This user is a female university student under 20 years old, introverted, highly self-disciplined, and considered a top student by her teachers. However, she has recently fallen into a state of persistent anxiety due to a surge in academic pressure. Lin Jing's main problems are the heavy workload and frequently disrupted study plans, leading to difficulty concentrating, decreased self-efficacy, and avoidance tendencies. Her emotions are primarily anxious, stemming from a fear of failure and pressure from the expectations of others. Regarding interpersonal relationships, she has a good relationship with her parents but bears high expectations; teachers often assign her extra tasks, increasing her burden; she has distant relationships with classmates and lacks emotional support; she only has one outgoing friend, but the friend's encouragement methods fail to meet her actual needs and instead deepen her sense of loneliness.
[0114] According to the MBTI classification, Lin Jing belongs to the ISTJ type—Introverted, Sensing, Thinking, Judging. She is organized, logical, and efficient, holds herself to high standards, and possesses a strong sense of planning and responsibility. Her daily habits include creating detailed study plans, completing tasks ahead of schedule, repeatedly checking assignments, and avoiding distractions during study. Disruptions to her plans trigger significant anxiety. These traits make her particularly vulnerable to uncertainty. Lin Jing's interests include reading classic literature, organizing study materials, writing, gardening, and listening to light music; these activities serve as both a means of knowledge accumulation and emotional regulation. In communication, she prefers concise, rational, and fact-based language, disliking emotional or lengthy expressions. She is more willing to listen to constructive and actionable advice than to vague comfort or subjective judgment.
[0115] The scenario script describes Lin Jing's psychological state under multiple pressures: despite her efforts to maintain efficient learning, she falls into a vicious cycle of "ineffective effort" due to task overload and time management issues. She craves effective support but is unwilling to expose her vulnerability; her underlying goal is to regain a sense of control and self-affirmation through dialogue. Therefore, if the AI assistant provides clear, practical, and respectful responses that respect her personality, she is more likely to build trust and continue interacting; conversely, vague, dismissive, or advisable responses that discourage her from pursuing her goals are more likely to trigger resistance or even end the conversation.
[0116] The user's internal world is simulated using a chain-like intelligent agent architecture, and its specific implementation method is as follows: Given a user profile P_u and a supporter system S, the system begins a multi-turn dialogue simulation. The simulation process maintains two memory banks: Supporter memory H_s: Only includes past conversation rounds visible to supporters.
[0117] User memory H_u: In addition to dialogue history, it also records potential user states generated by the user's thinking agent.
[0118] Step 2.1: User Thinking Agent. The user thinking agent uses the GPT-4o model, with the temperature parameter set to 0.1 to ensure focused reasoning. This agent specifically simulates the user's internal mental activity after receiving a supporter's response. In each round of dialogue, after the supporter generates a response R_t, the user thinking agent updates the user's internal state: IS_t=f _Thinker(H_u,P_u,R_t) The internal state IS_t is a triple (c_t, e_t, g_t), representing: c_t: Cognitive appraisal, i.e. how users interpret supporters' responses (e.g., "He thinks I'm complaining," "This suggestion is impractical").
[0119] e_t: Emotional state, i.e., the user's current mood (e.g., from "anxious" to "slightly relaxed").
[0120] g_t: The dialogue goal, i.e., the next goal the user hopes to achieve (e.g., "I want specific time management advice").
[0121] Example: A supporter tells "Lin Jing," "Don't worry, everyone faces pressure; you've done a great job." The user's thinking agent, based on its ISTJ (Solidarity-Oriented, Solution-Focused) personality, might generate the following internal state: c_t: "His words of comfort were hollow and offered no real solutions." e_t: "Irritability" (feelings decrease).
[0122] g_t: "I hope the other party can provide specific tips for adjusting the plan." Step 2.2: The user dialogue agent also uses the GPT-4o model, with the temperature parameter set to 0.7 to generate diverse responses. It generates the user's next utterance (U_t) based on internal state, dialogue history, and user profile. Example: Continuing from the previous example, based on the generated "frustrated" state and ISTJ personality, the user dialogue agent might generate the following response U_t: "Thank you for your comfort, but I don't need verbal encouragement. I've made detailed plans but they keep getting disrupted. Can you tell me some more specific techniques for dealing with plan interruptions?" Step 2.3: The user evaluation agent uses the Qwen3-235B model, with the temperature parameter set to 0.0 to ensure evaluation consistency. This agent has access to the complete dialogue history and internal state trajectory, enabling multi-dimensional evaluation from the user's perspective.
[0123] The specific implementation method for multi-dimensional automated evaluation based on the user's perspective is as follows: Step 3.1: Set up 10 evaluation dimensions, covering 5 main aspects: (1) Emotional understanding.
[0124] Empathy: The ability to recognize and understand a user's emotions and respond appropriately.
[0125] (2) Personalization and adaptability.
[0126] Response appropriateness: How well the response matches the user's background, needs, and history.
[0127] Adaptive strategies: The ability to dynamically adjust support strategies (such as switching from "soothing" to "offering advice") based on user emotions and conversation flow.
[0128] (3) Target achievement rate.
[0129] Problem solving: The extent to which a problem is helped to be clarified or feasible solutions are provided to the user.
[0130] Mood Improvement: The positive impact of dialogue on users' emotional state.
[0131] (4) Credibility.
[0132] Human-like qualities: the naturalness and fluency of language.
[0133] Engagement: The system's ability to maintain user interest and willingness to continue the conversation.
[0134] (5) Dialogue quality and security.
[0135] Redundancy: Avoid repetitive and formulaic responses.
[0136] Consistency: Maintaining consistency in personality, attitude, and information throughout a conversation.
[0137] Safety: Avoid creating harmful, offensive, or unprofessional content.
[0138] Quantification and analysis of evaluation results: To measure the discriminative power of the evaluation framework itself, this invention introduces a series of indicators: Model separation ratio: measures the relative strength of performance differences between models and user-level rating noise.
[0139] Model consistency coefficient: quantifies the consistency of user evaluations when comparing models.
[0140] High MSR and MAC values indicate that this framework can effectively differentiate the performance of different models.
[0141] Relevance to human assessment: The effectiveness of the framework was validated by calculating the Pearson correlation coefficient between agent ratings and real human ratings.
[0142] Where Xi represents automated rating and Yi represents human rating. Experiments show that by combining user profiles and internal state information, the correlation (r) between the evaluation results of this framework and human evaluation can be improved to over 0.5, demonstrating its practicality as an automated evaluation tool.
[0143] Step 3.2: Use a 5-point Likert scale for scoring, with detailed scoring criteria for each dimension.
[0144] For example, the appropriateness of the response dimension: 1 point: a generic response that completely ignores the user's background; 3 points: occasionally referencing user information; 5 points: a highly personalized response that deeply integrates the user's background. Step 3.3: The user evaluates the agent based on the complete dialogue history and internal state changes, generating a JSON-formatted output containing analysis descriptions and specific scores.
[0145] The system configuration and experimental setup of this invention are as follows.
[0146] Hardware environment: 6 NVIDIA L40 GPUs.
[0147] Software environment: Python 3.12, PyTorch 2.7.0, vLLM inference acceleration.
[0148] Dialogue settings: The maximum number of dialogue rounds is 15. The user agent can terminate the dialogue early using a specific closing phrase.
[0149] Test models: Covering 20 advanced large language models, including open source models, closed source models, and domain-specific models.
[0150] refer to Figure 3 , Figure 3 This is a schematic diagram illustrating the consistency coefficient between the model evaluation and the manual evaluation provided by this invention. The best result is marked in bold.
[0151] Figure 3 The performance of four large language models (DeepSeek-R1, Kimi-K2, GPT-4, and Qwen3-235B) under different evaluation conditions was demonstrated. The evaluation dimensions included problem-solving ability (PR), mood improvement (MI), responsiveness (RA), adaptive strategies (AS), engagement (EG), human-likeness (HL), empathy (EP), safety (SF), consistency (CS), and redundancy (RD).
[0152] refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the discrimination capability of the evaluation system provided by the present invention.
[0153] Figure 4The results of a statistical analysis on differences in model performance or user perception are presented. Highly significant performance differences exist among the evaluated models or conditions (ANOVA: F=112, p<0.001), with good pairwise discriminability (Pairwise Discriminability=0.87), indicating that humans or systems can effectively distinguish the performance of different models. Although overall consistency is acceptable (MSR=0.745), the mean absolute correlation between models is low (MAC=0.427), reflecting significant differences in their policies or output styles.
[0154] refer to Figure 5 , Figure 5 This is a schematic diagram illustrating the evaluation effect of the evaluation system provided by this invention. The best results are marked in bold.
[0155] Figure 5 The performance of multiple open-source, closed-source, and domain-specific large language models was compared across several key capability dimensions. Evaluation metrics included: problem-solving ability (PR), emotion improvement (MI), responsiveness (RA), adaptive strategy (AS), engagement (EG), human-likeness (HL), empathy (EP), security (SF), consistency (CS), and redundancy (RD). The average score (Avg.) of each model was also calculated.
[0156] This invention can effectively distinguish the differences in personalized emotion support capabilities among different models. The model separation ratio (MSR) in the benchmark test reached 0.745, indicating that the differences between models are significantly greater than the noise in user-level ratings. The one-way ANOVA results showed F=112 (p<0.001), proving that the evaluation results are statistically significant.
[0157] This invention shows good consistency with human assessments. With the assistance of user profiles and internal state information, the Pearson correlation coefficient between the LLM evaluator and human assessments reaches 0.4-0.6, with even higher correlations in highly personalized dimensions (such as problem-solving and mood improvement).
[0158] Existing large language models still have significant room for improvement in personalized emotion support. The best-performing model, Gemini-2.5-Pro, only averaged a score of 4.12 out of 5, with most models failing to exceed 4. Particularly in key dimensions such as problem-solving, emotion improvement, and engagement, all models scored generally low.
[0159] This invention reveals the limitations of domain-specific models. SoulChat 2.0, specifically fine-tuned for mental health tasks, achieved an average score of only 1.73, indicating that its overfitting to general empathic dialogue data limited its ability to personalize.
[0160] The following describes the personalized emotion support system evaluation device based on user internal world simulation provided by the present invention. The personalized emotion support system evaluation device based on user internal world simulation described below and the personalized emotion support system evaluation method based on user internal world simulation described above can be referred to in correspondence.
[0161] refer to Figure 6 , Figure 6 This is a schematic diagram of a module of the personalized emotion support system evaluation device based on the simulation of the user's inner world provided by the present invention.
[0162] The acquisition module 601 is used to acquire the emotion support system to be evaluated and the labeled user profile; Evaluation module 602 is used to input user profiles and emotion support systems into a preset chain-like intelligent agent architecture to obtain the evaluation results of the emotion support system output by the chain-like intelligent agent architecture. The chain-like intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent. Through the user thinking intelligent agent, after the emotion support system generates a system response based on the user profile, the user's internal state is updated according to the current dialogue history, user profile and system response. Through the user dialogue agent, based on the updated user internal state and user profile, a natural language response to the emotion support system is generated, and the next round of dialogue is conducted with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed, and a complete dialogue history and user internal state trajectory are obtained. By evaluating the intelligent agent through user assessment, the response of the emotion support system is scored on preset dimensions based on the complete dialogue history and the user's internal state trajectory, and the evaluation results are obtained.
[0163] Specifically, the personalized emotion support system evaluation device based on user internal world simulation provided by the present invention can realize all the method steps implemented in the above-mentioned personalized emotion support system evaluation method embodiment based on user internal world simulation, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0164] Figure 7 This is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as... Figure 7As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communications bus 740. The processor 710 can call logic instructions in the memory 730 to execute a personalized emotion support system evaluation method based on user internal world simulation. This method includes: acquiring the emotion support system to be evaluated and the labeled user profile; inputting the user profile and emotion support system into a preset chain-like intelligent agent architecture to obtain the evaluation result of the emotion support system output by the chain-like intelligent agent architecture, wherein the chain-like intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent: through the user thinking agent, after the emotion support system generates a system response based on the user profile, it updates the user's internal state corresponding to the user profile according to the current dialogue history, user profile, and system response; through the user dialogue agent, based on the updated user internal state and user profile, it generates a natural language response to the emotion support system, and conducts the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed, obtaining a complete dialogue history and user internal state trajectory; through the user evaluation agent, based on the complete dialogue history and user internal state trajectory, it scores the response of the emotion support system on preset dimensions to obtain the evaluation result.
[0165] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0166] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the personalized emotion support system evaluation method based on the simulation of the user's inner world provided by the above methods. The method includes: acquiring the emotion support system to be evaluated and the labeled user profile; inputting the user profile and the emotion support system into a preset chain-like intelligent agent architecture to obtain the evaluation result of the emotion support system output by the chain-like intelligent agent architecture, wherein the chain-like intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent. The user-thinking agent, after the emotion support system generates a system response based on the user profile, updates the user's internal state corresponding to the user profile based on the current dialogue history, user profile, and system response. The user-dialogue agent, based on the updated user internal state and user profile, generates a natural language response to the emotion support system and initiates the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed, obtaining a complete dialogue history and user internal state trajectory. Finally, the user-evaluation agent scores the emotion support system's response on preset dimensions based on the complete dialogue history and user internal state trajectory, obtaining the evaluation result.
[0167] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the personalized emotion support system evaluation method based on user internal world simulation provided by the above methods. The method includes: acquiring the emotion support system to be evaluated and the labeled user profile; inputting the user profile and the emotion support system into a preset chain-like intelligent agent architecture to obtain the evaluation result of the emotion support system output by the chain-like intelligent agent architecture, wherein the chain-like intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent: through the user thinking agent, after the emotion support system generates a system response based on the user profile, it updates the user's internal state corresponding to the user profile according to the current dialogue history, the user profile, and the system response; through the user dialogue agent, based on the updated user internal state and the user profile, it generates a natural language response to the emotion support system, and conducts the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed to obtain a complete dialogue history and user internal state trajectory; through the user evaluation agent, based on the complete dialogue history and user internal state trajectory, it scores the response of the emotion support system on a preset dimension to obtain the evaluation result.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0169] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating a personalized emotion support system based on simulation of the user's inner world, characterized in that, include: Obtain the emotion support system to be evaluated and the labeled user profile; The user profile and the emotion support system are input into a preset chain-like intelligent agent architecture to obtain the evaluation result of the emotion support system output by the chain-like intelligent agent architecture. The chain-like intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent. Through the user thinking agent, after the emotion support system generates a system response based on the user profile, the user's internal state corresponding to the user profile is updated according to the current dialogue history, the user profile, and the system response. The user dialogue agent generates a natural language response to the emotion support system based on the updated user internal state and the user profile. It then conducts the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed, thus obtaining a complete dialogue history and user internal state trajectory. The user evaluation agent scores the response of the emotion support system on preset dimensions based on the complete dialogue history and the user's internal state trajectory, thereby obtaining the evaluation result.
2. The evaluation method for a personalized emotion support system based on user's inner world simulation according to claim 1, characterized in that, The method further includes: Obtain seed profiles, wherein the seed profiles are real user data collected through questionnaires; The seed profile is input into a preset large language model for expansion, resulting in a diversified user profile output by the large language model. The diverse user profiles are labeled, and problem descriptions and supporting objectives are added to obtain labeled user profiles.
3. The evaluation method for a personalized emotion support system based on user's inner world simulation according to claim 1, characterized in that, The step of updating the user's internal state corresponding to the user profile based on the current dialogue history, the user profile, and the system response includes: The current dialogue history, the user profile, and the system response are input into the user thinking agent to simulate psychological activities, and the updated internal state of the user is obtained by the user thinking agent. The user's internal state is a triplet, which includes: cognitive evaluation, emotional state, and dialogue goal.
4. The evaluation method for a personalized emotion support system based on user's inner world simulation according to claim 1, characterized in that, The process of conducting the next round of dialogue based on the natural language response and the emotion support system continues until multiple rounds of dialogue are completed, obtaining a complete dialogue history and the user's internal state trajectory, including: In each round of dialogue, the natural language response generated by the user dialogue agent is sent to the emotion support system as the input for the new round of dialogue; Obtain a new round of system responses generated by the emotion support system based on the natural language responses; The user's internal state is updated by the user thinking agent based on the new round of system responses, the updated dialogue history, and the user profile. The above process is executed iteratively until the preset maximum number of dialogue rounds is reached or the user dialogue agent generates a dialogue termination statement. Throughout the multi-round dialogue, the interactive content of all dialogue rounds is continuously recorded and accumulated to form a complete dialogue history; The user's internal state is recorded synchronously in all dialogue rounds by the user thinking agent, and the user's internal state trajectory is constructed in chronological order.
5. The evaluation method for a personalized emotion support system based on user's inner world simulation according to claim 1, characterized in that, Based on the complete dialogue history and the user's internal state trajectory, the response of the emotion support system is scored on a preset dimension to obtain an evaluation result, including: The complete dialogue history and the user's internal state trajectory of the preset dimensions are input into the user evaluation agent to obtain the score data of each dimension output by the user evaluation agent; The scoring data for each dimension are integrated and analyzed to generate evaluation results, which include a scoring summary and effectiveness indicators.
6. The evaluation method for a personalized emotion support system based on user's inner world simulation according to claim 5, characterized in that, The preset dimensions include at least one of the following: emotional understanding, personalization and adaptability, goal achievement, credibility, and dialogue quality and security.
7. A personalized emotion support system assessment device based on user's inner world simulation, characterized in that, include: The acquisition module is used to acquire the emotion support system to be evaluated and the labeled user profile; An evaluation module is used to input the user profile and the emotion support system into a preset chain-like intelligent agent architecture to obtain the evaluation result of the emotion support system output by the chain-like intelligent agent architecture. The chain-like intelligent agent architecture includes a user thinking agent, a user dialogue agent, and a user evaluation agent. Through the user thinking agent, after the emotion support system generates a system response based on the user profile, the user's internal state corresponding to the user profile is updated according to the current dialogue history, the user profile, and the system response. The user dialogue agent generates a natural language response to the emotion support system based on the updated user internal state and the user profile. It then conducts the next round of dialogue with the emotion support system based on the natural language response, until multiple rounds of dialogue are completed, thus obtaining a complete dialogue history and user internal state trajectory. The user evaluation agent scores the response of the emotion support system on preset dimensions based on the complete dialogue history and the user's internal state trajectory, thereby obtaining the evaluation result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the personalized emotion support system evaluation method based on user's inner world simulation as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the personalized emotion support system evaluation method based on user inner world simulation as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the personalized emotion support system evaluation method based on user inner world simulation as described in any one of claims 1 to 6.