Conversation content scene suitability intelligent evaluation method based on text sentiment calculation

Through multimodal emotional computing and reinforcement learning optimization, the intelligent simulated patient system achieves accurate recognition and dynamic adaptation of emotional expressions, solving the problems of single emotional expression and insufficient adaptability of existing systems, and improving the interactive quality and adaptability of medical education.

CN120600355AActive Publication Date: 2025-09-05CHONGQING MEDICAL UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510719576.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-05
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Existing intelligent simulated patient systems lack emotional computing capabilities, have a single emotional expression, are unable to dynamically adapt to clinical scenarios, have a single data source, lack a systematic evaluation mechanism, have difficulty handling complex dialogue processes, and their application scenarios are limited to traditional medical practices.

Method used

A multimodal Transformer architecture is used to fuse text and speech features, construct a multimodal sentiment vector, combine KL divergence and cosine similarity to evaluate sentiment matching, establish an evaluation index system for sentiment consistency, semantic relevance, and interaction naturalness, and optimize the dialogue strategy through reinforcement learning and expert system feedback.

Benefits of technology

It achieves accurate emotion recognition and dynamic adaptation of intelligent simulated patient conversation content, improves the authenticity and adaptability of interaction, is applicable to a variety of medical scenarios, provides a scientific evaluation basis, and enhances the quality of medical training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600355A_ABST
    Figure CN120600355A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue content scene suitability intelligent evaluation method based on text sentiment calculation, and belongs to the technical field of dialogue sentiment detection. The method comprises the following steps: a corpus construction and situation labeling stage: labeling scene labels for dialogues of a corpus by using a multi-layer labeling system through an expert system; an emotion calculation model construction stage: evaluating the matching degree between the output emotion and the expected medical scene by adopting double indexes of KL divergence and cosine similarity, and optimizing a dialogue strategy; an emotion suitability evaluation mechanism establishment stage: generating a comprehensive situation suitability score for evaluating the situation suitability degree of the output content of the intelligent simulation patient ISP; and a real-time optimization and adjustment stage: performing real-time optimization on a dialogue strategy, language expression and emotion adaptation through reinforcement learning and expert system feedback, and dynamically adjusting a dialogue style according to case requirements. According to the method, the limitation of single text sentiment analysis is broken through, accurate recognition and dynamic adaptation of sentiment expression are realized, and the dialogue has higher sentiment naturalness and authenticity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dialogue emotion detection, and in particular to an intelligent evaluation method for dialogue content scenario adaptability based on text emotion calculation. Background Art

[0002] The concept of affective computing was proposed by Professor Rosalind Picard of MIT in 1997. This technology integrates knowledge from multiple disciplines, including computer science, psychology, cognitive science, and neuroscience. Its core goal is to empower computers with more natural, emotionally engaging interactions. Currently, affective computing technology demonstrates broad application potential in fields such as intelligent customer service, education, healthcare, public administration, and entertainment. Through emotion recognition technology, computers gain a deeper understanding of human emotions and can dynamically adjust their behavior and feedback based on user moods. This enables intelligent systems to factor user emotions into their decision-making processes, providing more personalized content and services. Furthermore, affective computing technology can enable intelligent systems to demonstrate a degree of empathy, enabling them to more effectively understand and support user needs, enhance their intelligence, and strengthen their ability to interact with human society, ultimately increasing their influence within society.

[0003] Text sentiment technology, also known as sentiment analysis, is a key branch of natural language processing (NLP). It focuses on collecting and analyzing public opinions, thoughts, and feelings about various topics, products, issues, and services. Its goal is to use computer technology to automatically identify and categorize the emotional tendencies inherent in text. The core of text sentiment technology lies in the processing and analysis of text data. Its methods can be broadly categorized into three main types: those based on sentiment lexicons, traditional machine learning, and deep learning. This technology has broad application prospects in numerous fields, including business analytics, healthcare, word-of-mouth management, public opinion monitoring, and customer feedback. Text sentiment technology enables automated processing and analysis of large amounts of text data, multimodal sentiment analysis, and the construction of personalized user sentiment models to more accurately understand users' emotional expressions and tendencies.

[0004] The Intelligent Simulation Patient (ISP) is a virtual patient simulator that integrates natural language processing (NLP) and intelligent tutoring system (ITS) technologies. Currently, IPS systems are developing towards a deeper integration of NLP and ITS technologies. By simulating real-world clinical scenarios, these systems enable students to practice clinical diagnostic reasoning in a virtual environment, gradually replacing some traditional clinical teaching activities in medical education. Faced with increasingly limited educational resources and increasingly complex clinical environments, IPS can effectively alleviate the challenge of clinical teaching models that struggle to meet student needs. Without incurring additional labor costs, IPS provides students with more realistic clinical scenarios and a more natural language interaction experience. Furthermore, the system can track student progress and performance and provide real-time feedback, thereby ensuring high-quality clinical training. Despite significant advances in NLP technology, IPS still may not accurately understand student input in certain situations and lacks a comprehensive evaluation mechanism for this output.

[0005] Contextual adaptability assessment is a method used to determine whether an object or system can operate or function effectively and rationally in a specific context. This assessment method primarily encompasses various types, including dimension-based, test type-based, subject-based, and data-driven. This method is applicable to assessment needs in fields such as scenario-based simulation teaching, medical device usability evaluation, and artificial intelligence. It comprehensively assesses contextual adaptability from multiple dimensions, avoiding the bias that can result from a single evaluation metric. Depending on the assessment object and objectives, appropriate assessment methods and metrics can be flexibly selected to improve the accuracy and effectiveness of the assessment. This assessment method is applicable to different operational modes (individual or team participation) and simulation scenarios, adapting flexibly to various needs. Furthermore, its results are instructive, providing clear direction and basis for improvement and optimization, helping to enhance system performance and user experience. In medical education, contextual adaptability assessment methods can reveal problems in the teaching process and provide clear feedback to instructors, helping them adjust teaching content, methods, and strategies to better meet students' learning needs.

[0006] The following problems exist in the existing technology: 1. Most existing intelligent simulated patient systems lack emotional computing capabilities. Some systems are inflexible in speech recognition and dialogue generation, have single emotional expression, unnatural communication, and cannot dynamically adapt to clinical situations.

[0007] 2. The data sources used by traditional intelligent simulated patient evaluation systems are relatively single. A fixed classification framework is often set when designing the model based on prior knowledge. The differences between different scenarios are not fully considered. There is a lack of a systematic situational adaptability evaluation mechanism. The evaluation dimension is single, making it difficult to comprehensively evaluate the quality of the output content.

[0008] 3. Existing systems often rely on simple decision trees or scripts for dialogue management, making them incapable of handling complex conversational flows. When responding to users, they typically follow fixed patterns and use mechanical expressions, failing to provide a flexible and natural interactive experience and lacking adaptive learning capabilities.

[0009] 4. Existing intelligent patient simulation systems are weak in processing real-time data and dynamic adjustment capabilities. Limited by text-only interfaces and a lack of multimodal data integration, they are difficult to adapt to new medical scenarios or changing medical needs. Their application scenarios are often limited to traditional medical practice teaching. Summary of the Invention

[0010] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for intelligently evaluating the contextual adaptability of conversation content based on text sentiment calculation.

[0011] The objective of the present invention is achieved through the following technical solution: a method for intelligently evaluating the contextual adaptability of conversation content based on text sentiment calculation, comprising the following steps: Corpus construction and context annotation phase: Clinical conversation texts and voices from various data sources and different medical scenarios are collected. An expert system uses a multi-layer annotation system to annotate the conversations in the corpus with context labels. During the emotional computing model construction phase, we simultaneously collected conversational text and speech data from the intelligent simulated patient ISP to construct a multimodal dataset. We used a multimodal Transformer architecture to fuse text and speech features to generate a multimodal emotion vector. KL divergence and cosine similarity were used to evaluate the match between the output emotion and the expected medical scenario, and to optimize the conversational strategy. Emotional adaptability evaluation mechanism establishment phase: Build an evaluation index system covering emotional consistency score, semantic relevance score, and interaction naturalness score. Apply the analytic hierarchy process and fuzzy comprehensive evaluation to weightedly calculate multiple indicators to generate a comprehensive contextual adaptability score, which is used to evaluate the contextual adaptability of the output content of the intelligent simulated patient ISP. Real-time optimization and adjustment stage: Through reinforcement learning and expert system feedback, dialogue strategies, language expression, and emotional adaptation are optimized in real time, and the dialogue style is dynamically adjusted according to case requirements.

[0012] Preferably, the corpus construction and context annotation stage further includes the following steps: We selected conversation data covering different disease types, patient age groups, medical institutions, and medical scenarios from medical training cases, patient interview records, and standardized patient simulation conversations, and increased the weight of conversation data on rare diseases and complex cases. A multi-layer labeling system including disease type, patient emotional state, and medical communication strategy is established; the expert system uses multi-person labeling and cross-validation methods to label each conversation data based on the multi-layer labeling system.

[0013] Preferably, the emotion computing model building stage further includes the following steps: Recording the conversation text of the intelligent simulated patient ISP, using natural speech processing methods to perform word segmentation, denoising, and standardization preprocessing on the conversation text, and collecting voice data corresponding to the content of the conversation text; the voice data includes pitch, speaking speed, and tone changes; Use the pre-trained BERT model to contextually encode conversation text, extract semantic vectors, and capture the sentiment and fine-grained emotion categories in the conversation text. Combined with TF-IDF and word embedding methods, the representation capabilities of text features are enhanced. The short-time spectrum features of speech signals are extracted through a deep convolutional neural network, and the temporal dependencies are captured by a BiLSTM network to generate speech emotion vectors. An end-to-end speech recognition method is used to extract the semantic information of speech data and enhance the feature fusion effect. Through the cross-modal attention layer, the association weights of text features and speech features are calculated, and dynamically fused through the attention mechanism to generate a fused multimodal sentiment vector; Calculate the difference in probability distribution between the multimodal emotion vector and the scene annotation emotion label to obtain the KL divergence; Combined with the vector normalization method, the cosine similarity is obtained by measuring the directional consistency between the multimodal emotion vector and the scene expected emotion vector in the semantic space; The KL divergence and cosine similarity are weighted to generate a sentiment consistency score, and a dynamic weight adjustment strategy is adopted to optimize the scoring model according to the needs of different scenarios.

[0014] Preferably, the emotional adaptability evaluation mechanism establishment stage further includes the following steps: The dual indicators of emotion classification accuracy and emotion intensity matching are used to measure the matching degree between the output emotion of the intelligently simulated patient and the expected patient emotion; The BERT model was used to calculate the semantic match between the intelligent simulated patient's ISP answer and the clinical context, and the semantic relevance score was obtained by combining semantic role labeling and dependency syntax analysis. A combination of language model scoring and manual evaluation is used to score the fluency, logic, and contextual coherence of the conversation to obtain an interaction naturalness score. The weight of each evaluation index was determined using AHP, and the weight distribution was optimized by combining expert scoring and statistical analysis; The scores of each evaluation indicator are fuzzy processed and combined with weights to generate a comprehensive situational adaptability score; Fuzzy membership function and fuzzy comprehensive evaluation matrix are used to improve the stability of evaluation results.

[0015] Preferably, the real-time optimization and adjustment stage further includes the following steps: Establish a reinforcement learning environment, use the conversational strategy of the intelligent simulated patient ISP as the learning object, and define the state space, action space, and reward function. Combined with feedback from the expert system, continuously optimize the voice expression and emotional adaptability of the intelligent simulated patient ISP. Use a combination of online and offline learning methods to improve the adaptability of the model. Analyze the characteristics and needs of the current medical situation, including disease type, patient emotions, and medical staff roles; use situational awareness methods to capture situational changes in real time; dynamically adjust the dialogue strategy of the intelligent simulated patient ISP based on the results of the situational analysis to ensure that its output content is highly consistent with real doctor-patient communication in terms of style, tone, and wording; use a multimodal generation model to generate dialogue content that meets situational needs; and continuously optimize the dialogue strategy of the intelligent simulated patient ISP by monitoring the dialogue effect in real time to improve its situational adaptability.

[0016] The beneficial effects of the present invention are: 1) This invention uses a multimodal text emotion computing model to accurately identify, analyze, and match the conversational emotions of intelligent simulated patients. By combining text and voice dual-modal data collection, it breaks through the limitations of single-text emotion analysis and achieves accurate recognition and dynamic adaptation of emotional expressions, making the conversation more emotionally natural and authentic.

[0017] 2) This paper constructs a multi-dimensional evaluation system, including multiple indicators such as emotional consistency, semantic relevance, and interaction naturalness, and uses hierarchical analysis method and fuzzy comprehensive evaluation algorithm to generate a comprehensive situational adaptability score, which can more comprehensively and accurately evaluate the adaptability of the output content of intelligent simulated patients.

[0018] 3) This invention introduces a reinforcement learning mechanism, combined with feedback from professionals, to dynamically adjust the conversational style, tone, and wording of the intelligent simulated patient, enabling it to adapt to different medical scenarios, optimize language expression and emotional adaptability, enhance the intelligence level of the system, and promote the authenticity of the interaction.

[0019] 4) This invention is not only suitable for clinical communication training of medical students, but can also be widely used in various medical scenarios such as general practitioner consultation simulation and psychological diagnosis and treatment scenarios. It has stronger versatility and scalability, and can meet the needs of different medical education and clinical simulations.

[0020] 5) Improving the contextual adaptability of intelligent patient simulators: This paper utilizes multiple data sources, including clinical conversation text and speech, to construct a patient simulator corpus and develop a contextual adaptability evaluation method. By optimizing the conversations using a reinforcement learning-based optimization mechanism, the contextual adaptability of intelligent patient simulators was effectively enhanced.

[0021] 6) Enhanced Emotional Naturalness in Interactive Training: This invention utilizes a multimodal text-based sentiment computing model, breaking through the limitations of traditional single-text sentiment analysis and accurately capturing the characteristics of users' emotional expressions. This multimodal emotion recognition and analysis capability enables the intelligent simulated patient to display more natural and authentic emotional responses when interacting with users, significantly improving the emotional naturalness of interactive training.

[0022] 7) Building an intelligent evaluation system to improve the quality of medical training: This evaluation system comprehensively considers multiple dimensions, including emotional consistency, semantic relevance, and natural interaction. Using the Analytic Hierarchy Process (AHP) and fuzzy comprehensive evaluation algorithm, it weights multiple evaluation indicators to generate a comprehensive contextual adaptability score, providing a scientific and accurate basis for the evaluation of medical training.

[0023] 8) Optimizing the adaptive learning capabilities of intelligent simulated patients: This invention is based on a reinforcement learning optimization mechanism and, combined with feedback from professionals, can dynamically adjust the dialogue strategy of intelligent simulated patients, enabling them to adapt to different medical scenarios.

[0024] 9) Applicable to various medical scenarios: The present invention can flexibly adapt to different medical needs. It is not only suitable for clinical communication training of medical students, but can also be widely used in various medical scenarios such as general practitioner consultation simulation and psychological diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flow chart of the intelligent evaluation method for dialogue content scenario adaptability based on text sentiment computing. DETAILED DESCRIPTION

[0026] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.

[0027] This paper evaluates the contextual adaptability of the output content of an intelligent patient simulator system. It proposes a method for evaluating the contextual adaptability of the output content of the intelligent patient simulator based on multimodal text sentiment computing and deep learning. This method addresses the key issues of existing intelligent patient simulator systems, such as limited emotional expression, insufficient dynamic adaptability, and a lack of a systematic evaluation mechanism in clinical interaction scenarios. This method utilizes text sentiment computing technology to identify, analyze, and match the emotions expressed by the intelligent patient simulator's conversations. This method constructs a cross-modal emotion consistency model, overcoming the limitations of single-text emotion recognition. It accurately captures the emotional expression characteristics of the intelligent patient simulator and establishes multi-dimensional indicators (emotional consistency, semantic relevance, and interaction naturalness) to quantitatively assess the adaptability of the output content to clinical scenarios. This method can be widely applied to clinical communication training for medical students, general practitioner consultation simulations, and psychological diagnosis and treatment scenarios, providing a novel intelligent solution for the medical education and technology field. This method effectively improves the interactive quality of the intelligent patient simulator, making its conversations more consistent with clinical context requirements, thereby enhancing the authenticity and effectiveness of doctor-patient communication training in medical education and training.

[0028] See Figure 1 The present invention provides a technical solution: an intelligent evaluation method for the contextual adaptability of dialogue content based on text sentiment calculation, comprising the following steps: Corpus construction and context annotation phase: Clinical conversation texts and voices from various data sources and different medical scenarios are collected. An expert system uses a multi-layer annotation system to annotate the conversations in the corpus with context labels. During the emotional computing model construction phase, we simultaneously collected conversational text and speech data from the intelligent simulated patient ISP to construct a multimodal dataset. We used a multimodal Transformer architecture to fuse text and speech features to generate a multimodal emotion vector. KL divergence and cosine similarity were used to evaluate the match between the output emotion and the expected medical scenario, and to optimize the conversational strategy. Emotional adaptability evaluation mechanism establishment phase: Build an evaluation index system covering emotional consistency score, semantic relevance score, and interaction naturalness score. Apply the analytic hierarchy process and fuzzy comprehensive evaluation to weightedly calculate multiple indicators to generate a comprehensive contextual adaptability score, which is used to evaluate the contextual adaptability of the output content of the intelligent simulated patient ISP. Real-time optimization and adjustment stage: Through reinforcement learning and expert system feedback, dialogue strategies, language expression, and emotional adaptation are optimized in real time, and the dialogue style is dynamically adjusted according to case requirements.

[0029] This embodiment utilizes affective computing technology, multimodal data fusion, and a scenario-adaptability evaluation method, combined with reinforcement learning optimization and other technical approaches, to effectively overcome the limitations of existing intelligent patient simulation systems in terms of expression, emotion recognition and analysis, scenario adaptability, and adaptive learning capabilities. Compared to existing technologies, this invention demonstrates significant advantages in terms of scenario adaptability, intelligent evaluation systems, and adaptive learning capabilities, providing a more realistic, natural, and intelligent solution for medical education and clinical simulation.

[0030] This paper leverages natural language processing, sentiment computing, and multimodal sentiment analysis to build a comprehensive system for evaluating the contextual adaptability of intelligent simulation content. This method effectively assesses the suitability of conversational content with intelligent simulated patients in diverse medical scenarios, thereby enhancing the realism, coherence, and adaptability of these interactions.

[0031] It mainly includes the following contents: 1. Corpus construction and context annotation: To ensure the contextual adaptability of the output content of the intelligent simulated patient, this method first needs to construct a standardized medical dialogue corpus and perform clinical context annotation. (1) Data collection: Collect clinical dialogue text and voice from various data sources such as medical training cases, patient interview records, standardized patient simulation dialogues, etc. Covering different medical scenarios including initial diagnosis, follow-up, emergency, preoperative, and postoperative. (2) Manual context labeling: Form a medical and nursing professional team and use a multi-layer annotation system to label the conversation content in the corpus, such as disease type (chronic disease, acute disease, psychological disease, etc.), patient emotional state (anxiety, anger, sadness, calmness, etc.), and medical communication strategies (comfort, explanation, guidance, empathy, etc.).

[0032] 2. Intelligent simulated patient emotion computing model: In order to evaluate whether the emotional expression of the intelligent simulated patient conversation matches the medical scenario, break through the limitations of single text emotion analysis, and achieve accurate recognition and dynamic adaptation of the emotional expression of the intelligent simulated patient, this method designs a multimodal text emotion computing model. The specific steps are as follows: (1) Synchronous acquisition of bimodal data: The text conversation content and speech output data (such as pitch, speaking speed, tone change, etc.) of the intelligent simulated patient are collected simultaneously to construct a time-aligned multimodal dataset. (2) Cross-modal feature fusion model: A multimodal Transformer architecture is used to dynamically fuse text features (BERT encoding) and speech features (acoustic features extracted based on CNN+BiLSTM) through an attention mechanism. The pre-trained BERT model is used to contextually encode the conversation text and extract the semantic vector; the BiLSTM+Attention network is used to capture the emotional tendency (positive, neutral, negative) and fine-grained emotion categories (such as anxiety, sadness, calmness) in the text. A deep convolutional neural network (CNN) is used to extract the short-time spectral features (MFCC, fundamental frequency, energy, etc.) of the speech signal, and combined with BiLSTM to capture temporal dependencies, a speech emotion vector (such as anger, calmness, and tension) is generated. A cross-modal attention layer is designed to calculate the association weights between text and speech features, and to generate a fused multimodal emotion vector. (3) Emotional consistency analysis: Quantitatively evaluate the matching degree between the output emotion of the intelligent simulated patient and the expected emotion of the medical scene, and optimize the dialogue strategy in real time. The KL divergence (Kullback-Leibler Divergence) and cosine similarity dual indicators are used for joint evaluation: KL divergence is used to calculate the difference in probability distribution between the emotion distribution output by the intelligent simulated patient (multimodal fusion result) and the emotion label of the scene. The smaller the difference, the higher the adaptability. Cosine similarity is used to measure the directional consistency between the output emotion vector and the expected emotion vector of the scene in the semantic space. The final emotion consistency is calculated by weighting the two indicators.

[0033] 3. Situational adaptability evaluation mechanism: In order to conduct a comprehensive situational adaptability evaluation of the output content of the intelligent simulated patient, this section designs a multi-dimensional evaluation mechanism, including: (1) Evaluation index system: The emotional consistency score is used to measure whether the emotions output by the intelligent simulated patient are consistent with the expected patient emotions; the semantic relevance score is used to calculate whether the intelligent simulated patient's answers are consistent with the clinical context (the BERT model performs semantic matching); the interaction naturalness score is used to score the fluency, logic, and contextual coherence of the dialogue. (2) Evaluation algorithm: The Analytic Hierarchy Process (AHP) and fuzzy comprehensive evaluation are used to perform weighted calculations on multiple evaluation indicators to generate a comprehensive situational adaptability score; the score range is set to 0-100, and the higher the score, the better the situational adaptability of the output content of the intelligent simulated patient.

[0034] 4. Optimization and real-time adjustment of intelligent simulated patient dialogue: Based on the above evaluation results, the present invention also proposes a reinforcement learning optimization mechanism to optimize the text generation strategy of the intelligent simulated patient, which specifically includes: (1) Feedback mechanism: Through the reinforcement learning environment, the dialogue strategy of the intelligent simulated patient is adjusted to enable it to adapt to different medical situations; combined with professional feedback, the language expression and emotional adaptability of the intelligent simulated patient are continuously optimized. (2) Adaptive situation adjustment: Based on the case requirements, the conversation style, tone, wording, etc. of the intelligent simulated patient are dynamically adjusted to make it more in line with the needs of real doctor-patient communication.

[0035] In some embodiments, the corpus construction and context annotation stage further includes the following steps: We selected conversation data covering different disease types, patient age groups, medical institutions, and medical scenarios from medical training cases, patient interview records, and standardized patient simulation conversations, and increased the weight of conversation data on rare diseases and complex cases. A multi-layer labeling system including disease type, patient emotional state, and medical communication strategy is established; the expert system uses multi-person labeling and cross-validation methods to label each conversation data based on the multi-layer labeling system.

[0036] In this example, clinical conversation text and voice data were extracted from a variety of data sources, including medical training cases, patient interview records, and standardized patient simulations. This data covers conversations across different disease types, patient age groups, and medical institutions. The collected data covers various medical scenarios, including initial visits, follow-up visits, emergency department visits, and pre- and post-operative settings, with a particular focus on conversations involving rare diseases and complex cases.

[0037] Manual context labeling: A multi-layered labeling system was designed, encompassing disease type (e.g., chronic, acute, psychological), patient emotional state (e.g., anxiety, anger, sadness, calmness), and doctor-nurse communication strategies (e.g., comfort, explanation, guidance, empathy). The labeling team meticulously annotated each conversation based on this labeling system. For example, a sentence like "I've been feeling anxious lately and having trouble sleeping at night" could be labeled "Disease type: psychological, patient emotional state: anxiety, doctor-nurse communication strategy: comfort, and conversation stage: description of the condition." Multiple annotations and cross-validation were used to ensure consistency.

[0038] In some embodiments, the emotion computing model building stage further includes the following steps: Recording the conversation text of the intelligent simulated patient ISP, using natural speech processing methods to perform word segmentation, denoising, and standardization preprocessing on the conversation text, and collecting voice data corresponding to the content of the conversation text; the voice data includes pitch, speaking speed, and tone changes; Use the pre-trained BERT model to contextually encode conversation text, extract semantic vectors, and capture the sentiment and fine-grained emotion categories in the conversation text. Combined with TF-IDF and word embedding methods, the representation capabilities of text features are enhanced. The short-time spectrum features of speech signals are extracted through a deep convolutional neural network, and the temporal dependencies are captured by a BiLSTM network to generate speech emotion vectors. An end-to-end speech recognition method is used to extract the semantic information of speech data and enhance the feature fusion effect. Through the cross-modal attention layer, the association weights of text features and speech features are calculated, and dynamically fused through the attention mechanism to generate a fused multimodal sentiment vector; Calculate the difference in probability distribution between the multimodal emotion vector and the scene annotation emotion label to obtain the KL divergence; Combined with the vector normalization method, the cosine similarity is obtained by measuring the directional consistency between the multimodal emotion vector and the scene expected emotion vector in the semantic space; The KL divergence and cosine similarity are weighted to generate a sentiment consistency score, and a dynamic weight adjustment strategy is adopted to optimize the scoring model according to the needs of different scenarios.

[0039] In this embodiment, dual-modal data is collected simultaneously using high-precision microphones and professional recording equipment to ensure the accuracy of the voice data. The conversation text generated by the intelligent simulated patient is recorded and pre-processed using natural language processing techniques, including word segmentation, noise removal, and normalization. The speech output data corresponding to the textual conversation content is collected, including acoustic features such as pitch, speaking rate, and inflection.

[0040] Cross-modal feature fusion model: A pre-trained BERT model is used to contextually encode conversational text, extracting semantic vectors that capture sentiment and fine-grained emotion categories. TF-IDF and word embedding techniques are also combined to further enhance the representational capabilities of text features. A deep convolutional neural network (CNN) extracts short-term spectral features of speech signals (such as MFCC, fundamental frequency, and energy). Combined with a BiLSTM network, these features capture temporal dependencies and generate speech emotion vectors (such as anger, calmness, and tension). End-to-end speech recognition technology is employed to extract semantic information from speech and enhance feature fusion. A cross-modal attention layer is designed to calculate the correlation weights between text and speech features. Dynamic fusion is performed through the attention mechanism to generate a fused multimodal emotion vector.

[0041] Emotional consistency analysis: KL divergence is calculated to measure the difference between the emotional distribution output by the intelligent simulated patient (multimodal fusion effect) and the probability distribution of the emotion labels annotated in the scene. The smaller the difference, the higher the compatibility. Smoothing techniques are used to avoid probability distributions reaching zero. Cosine similarity is calculated to measure the directional consistency between the output emotion vector and the expected emotion vector in the scene in semantic space. Vector normalization techniques are combined to improve the stability of similarity calculations. Emotional consistency scoring is a weighted calculation of KL divergence and cosine similarity to generate the final emotion consistency score. A dynamic weight adjustment strategy is used to optimize the scoring model based on the needs of different scenarios.

[0042] In some embodiments, the emotional adaptability evaluation mechanism establishment stage further includes the following steps: The dual indicators of emotion classification accuracy and emotion intensity matching are used to measure the matching degree between the output emotion of the intelligently simulated patient and the expected patient emotion; The BERT model was used to calculate the semantic match between the intelligent simulated patient's ISP answer and the clinical context, and the semantic relevance score was obtained by combining semantic role labeling and dependency syntax analysis. A combination of language model scoring and manual evaluation is used to score the fluency, logic, and contextual coherence of the conversation to obtain an interaction naturalness score. The weight of each evaluation index was determined using AHP, and the weight distribution was optimized by combining expert scoring and statistical analysis; The scores of each evaluation indicator are fuzzy processed and combined with weights to generate a comprehensive situational adaptability score; Fuzzy membership function and fuzzy comprehensive evaluation matrix are used to improve the stability of evaluation results.

[0043] In this embodiment, the emotional consistency score uses the dual metrics of emotion classification accuracy and emotional intensity matching to measure the degree to which the intelligent simulated patient's output matches the expected patient's emotion. The semantic relevance score uses the BERT model to calculate the semantic match between the intelligent simulated patient's responses and the clinical context, combining semantic role labeling and dependency syntax analysis to improve the accuracy of semantic matching. The interaction naturalness score uses a combination of language model scoring and manual evaluation to assess the fluency, logic, and contextual coherence of the conversation.

[0044] The evaluation algorithm uses AHP to determine the weights of each evaluation indicator, optimizing the weight distribution through a combination of expert scoring and statistical analysis. The scores of each evaluation indicator are fuzzified and combined with the weights to generate a comprehensive contextual adaptability score, ranging from 0 to 100, with higher scores indicating better contextual adaptability. Fuzzy membership functions and a fuzzy comprehensive evaluation matrix are used to improve the stability of the evaluation results.

[0045] In some embodiments, the real-time optimization and adjustment stage further includes the following steps: Establish a reinforcement learning environment, use the conversational strategy of the intelligent simulated patient ISP as the learning object, and define the state space, action space, and reward function. Combined with feedback from the expert system, continuously optimize the voice expression and emotional adaptability of the intelligent simulated patient ISP. Use a combination of online and offline learning methods to improve the adaptability of the model. Analyze the characteristics and needs of the current medical situation, including disease type, patient emotions, and medical staff roles; use situational awareness methods to capture situational changes in real time; dynamically adjust the dialogue strategy of the intelligent simulated patient ISP based on the results of the situational analysis to ensure that its output content is highly consistent with real doctor-patient communication in terms of style, tone, and wording; use a multimodal generation model to generate dialogue content that meets situational needs; and continuously optimize the dialogue strategy of the intelligent simulated patient ISP by monitoring the dialogue effect in real time to improve its situational adaptability.

[0046] In this example, a reinforcement learning environment was designed, using the conversational strategies of an intelligent simulated patient as the learning object. The state space, action space, and reward function were defined to ensure the effectiveness of reinforcement learning. Incorporating feedback from professionals, the language expression and emotional adaptability of the intelligent simulated patient were continuously optimized. A combination of online and offline learning was employed to enhance the model's adaptability.

[0047] Adaptive contextual adjustment analyzes the characteristics and needs of the current medical situation, including disease type, patient emotion, and the roles of doctors and nurses. Contextual awareness technology is used to capture contextual changes in real time. Based on the results of this contextual analysis, the intelligent patient simulator's conversational strategy is dynamically adjusted to ensure that its output, in terms of style, tone, and wording, closely aligns with real-life doctor-patient interactions. A multimodal generative model is used to generate conversational content tailored to the specific context. By monitoring conversational effectiveness in real time, the intelligent patient simulator's conversational strategy is continuously optimized to enhance its contextual adaptability.

[0048] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.

Claims

1. An intelligent evaluation method for the contextual adaptability of conversational content based on text sentiment computing, characterized by: The following steps are involved: Corpus construction and context annotation phase: Clinical conversation texts and voices from various data sources and different medical scenarios are collected. An expert system uses a multi-layer annotation system to annotate the conversations in the corpus with context labels. During the emotional computing model construction phase, we simultaneously collected conversational text and speech data from the intelligent simulated patient ISP to construct a multimodal dataset. We used a multimodal Transformer architecture to fuse text and speech features to generate a multimodal emotion vector. KL divergence and cosine similarity were used to evaluate the match between the output emotion and the expected medical scenario, and to optimize the conversational strategy. Emotional adaptability evaluation mechanism establishment phase: Build an evaluation index system covering emotional consistency score, semantic relevance score, and interaction naturalness score. Apply the analytic hierarchy process and fuzzy comprehensive evaluation to weightedly calculate multiple indicators to generate a comprehensive contextual adaptability score, which is used to evaluate the contextual adaptability of the output content of the intelligent simulated patient ISP. Real-time optimization and adjustment stage: Through reinforcement learning and expert system feedback, dialogue strategies, language expression, and emotional adaptation are optimized in real time, and the dialogue style is dynamically adjusted according to case requirements.

2. The method for intelligently evaluating the contextual adaptability of conversation content based on text sentiment computing according to claim 1 is characterized by: The corpus construction and context annotation stage also includes the following steps: We selected conversation data covering different disease types, patient age groups, medical institutions, and medical scenarios from medical training cases, patient interview records, and standardized patient simulation conversations, and increased the weight of conversation data on rare diseases and complex cases. A multi-layer labeling system including disease type, patient emotional state, and medical communication strategy is established; the expert system uses multi-person labeling and cross-validation methods to label each conversation data based on the multi-layer labeling system.

3. The method for intelligently evaluating the contextual adaptability of conversation content based on text sentiment computing according to claim 1 is characterized by: The emotional computing model construction phase also includes the following steps: Recording the conversation text of the intelligent simulated patient ISP, using natural speech processing methods to perform word segmentation, denoising, and standardization preprocessing on the conversation text, and collecting voice data corresponding to the content of the conversation text; the voice data includes pitch, speaking speed, and tone changes; Use the pre-trained BERT model to contextually encode conversation text, extract semantic vectors, and capture the sentiment and fine-grained emotion categories in the conversation text. Combined with TF-IDF and word embedding methods, the representation capabilities of text features are enhanced. The short-time spectrum features of speech signals are extracted through a deep convolutional neural network, and the temporal dependencies are captured by a BiLSTM network to generate speech emotion vectors. An end-to-end speech recognition method is used to extract the semantic information of speech data and enhance the feature fusion effect. Through the cross-modal attention layer, the association weights of text features and speech features are calculated, and dynamically fused through the attention mechanism to generate a fused multimodal sentiment vector; Calculate the probability distribution difference between the multimodal emotion vector and the scene annotation emotion label to obtain the KL divergence; Combined with the vector normalization method, the cosine similarity is obtained by measuring the directional consistency between the multimodal emotion vector and the scene expected emotion vector in the semantic space; The KL divergence and cosine similarity are weighted to generate a sentiment consistency score, and a dynamic weight adjustment strategy is adopted to optimize the scoring model according to the needs of different scenarios.

4. The method for intelligently evaluating the contextual adaptability of conversation content based on text sentiment computing according to claim 1 is characterized by: The emotional adaptability evaluation mechanism establishment stage further includes the following steps: The dual indicators of emotion classification accuracy and emotion intensity matching are used to measure the matching degree between the output emotion of the intelligently simulated patient and the expected patient emotion; The BERT model was used to calculate the semantic match between the intelligent simulated patient's ISP answer and the clinical context, and the semantic relevance score was obtained by combining semantic role labeling and dependency syntax analysis. A combination of language model scoring and manual evaluation is used to score the fluency, logic, and contextual coherence of the conversation to obtain an interaction naturalness score. The weight of each evaluation index was determined using AHP, and the weight distribution was optimized by combining expert scoring and statistical analysis; The scores of each evaluation indicator are fuzzy processed and combined with weights to generate a comprehensive situational adaptability score; Fuzzy membership function and fuzzy comprehensive evaluation matrix are used to improve the stability of evaluation results.

5. The method for intelligently evaluating the contextual adaptability of conversation content based on text sentiment computing according to claim 1 is characterized by: The real-time optimization and adjustment stage further includes the following steps: Establish a reinforcement learning environment, use the conversational strategy of the intelligent simulated patient ISP as the learning object, and define the state space, action space, and reward function. Combined with feedback from the expert system, continuously optimize the voice expression and emotional adaptability of the intelligent simulated patient ISP. Use a combination of online and offline learning methods to improve the adaptability of the model. Analyze the characteristics and needs of the current medical situation, including disease type, patient emotions, and medical staff roles; use situational awareness methods to capture situational changes in real time; dynamically adjust the dialogue strategy of the intelligent simulated patient ISP based on the results of the situational analysis to ensure that its output content is highly consistent with real doctor-patient communication in terms of style, tone, and wording; use a multimodal generation model to generate dialogue content that meets situational needs; and continuously optimize the dialogue strategy of the intelligent simulated patient ISP by monitoring the dialogue effect in real time to improve its situational adaptability.

Citation Information

Patent Citations

  • Dialogue digital person emotion style similarity evaluation method and system

    CN117150320A

  • Large medical emotion support model and generation method thereof

    CN119673384A

  • Microteaching actual effect evaluation system based on artificial intelligence

    CN119740916A

  • AI-based mental health counseling and customized follow-up management system

    KR102783562B1