Intelligent consultation interaction method and system integrated with man-machine cooperation
By monitoring the visual, verbal, and operational behaviors of medical professionals in real time and dynamically adjusting the presentation of risk warnings, the problem of risk warnings being ignored in existing systems is solved, thus improving the efficiency and security of intelligent medical consultations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN QIYUN TECHNOLOGY CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-17
AI Technical Summary
Existing intelligent medical consultation systems struggle to accurately distinguish between core and non-critical medical information when faced with patients who are emotionally stressed, have unclear expressions, or communicate indirectly. This leads to risk warnings being overlooked, impacting the accuracy and safety of diagnosis and treatment.
By acquiring visual behavior information, voice interaction information, and operational behavior information of medical professionals, their cognitive load status can be determined in real time, and the intensity and manner of risk warning presentation can be enhanced when the cognitive load is high, including adjustments to visual salience, spatial location, and dynamic effects.
It effectively avoids ignoring key risk information under high cognitive load, improves the accuracy and safety of diagnosis and treatment, ensures that risk warnings are noticed in a timely manner under high pressure, and reduces the risk of misdiagnosis or delayed treatment.
Smart Images

Figure CN121885246A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of intelligent consultation interaction integrating human-computer collaboration, and specifically to an intelligent consultation interaction method and system integrating human-computer collaboration. Background Technology
[0002] In the field of modern medical consultation, intelligent consultation and interaction systems that integrate artificial intelligence (AI) with human doctors aim to optimize the medical consultation process and improve diagnostic efficiency and accuracy by combining the analytical capabilities of AI with the professional judgment and empathy of human medical professionals. These systems typically utilize high-fidelity microphone arrays to capture doctor-patient conversations and record patients' nonverbal information via high-definition cameras. This data is transmitted in real-time to AI models for processing, assisting doctors in information gathering, generating preliminary diagnostic suggestions, and optimizing treatment plans. Key information and supplementary suggestions are then presented in a structured format on the doctor's interactive terminal display.
[0003] However, in real-world medical consultation settings, especially when dealing with patients who are emotionally anxious, unclear in their expression, or communicate indirectly, doctors often need to adopt flexible communication strategies to effectively soothe patients, build trust, and guide them to clearly articulate their condition. This might involve temporarily deviating from the system's preset questioning order or dialogue flow, engaging in non-medical, casual conversation, or repeatedly confirming the patient's feelings. While this flexible communication approach is crucial for improving patient experience and treatment adherence, it can pose challenges to the natural language understanding models of AI systems. Most current natural language understanding models are primarily trained on medical terminology and structured medical history descriptions. When processing conversational, emotional, and unstructured text, their ability to deeply understand context and filter key information is limited. They may identify doctors' reassuring language and patients' emotional expressions as "noise," making it difficult for the model to accurately distinguish between core medical information and non-critical information, thus affecting its accuracy. Summary of the Invention
[0004] The purpose of this invention is to address the aforementioned shortcomings by proposing an intelligent consultation interaction method and system that integrates human-computer collaboration.
[0005] The present invention adopts the following technical solution:
[0006] An intelligent consultation interaction method integrating human-computer collaboration, the method includes the following steps: acquiring visual behavior information, voice interaction information and operational behavior information of medical professionals during the consultation process;
[0007] Based on visual behavioral information, voice interaction information, and operational behavior information, the cognitive load status of medical professionals can be assessed in real time.
[0008] Based on cognitive load status, the presentation of risk warnings for auxiliary suggestions is adjusted, and the auxiliary suggestions are generated by the intelligent assistance system.
[0009] When cognitive load indicates that healthcare professionals are under high cognitive load, increase the intensity of risk warning presentation.
[0010] Through this technical solution, this application can sense the cognitive load status of medical professionals in real time and dynamically adjust the presentation of risk prompts accordingly. The prompt intensity is enhanced when the cognitive load is high, which effectively prevents doctors from ignoring key risk information under high pressure. This improves the practicality and safety of the intelligent assistance system and solves the problem in the prior art that the risk prompts are not prominent, which makes doctors easily ignore potential risks under high cognitive load.
[0011] This application also discloses an intelligent consultation interaction system integrating human-computer collaboration, applied to the aforementioned intelligent consultation interaction method integrating human-computer collaboration. The system includes:
[0012] The acquisition module acquires visual behavior information, voice interaction information, and operational behavior information of medical professionals during the consultation process;
[0013] The judgment module, based on visual behavior information, voice interaction information, and operational behavior information, judges the cognitive load status of medical professionals in real time.
[0014] The adjustment module adjusts the presentation of risk warnings for assistance suggestions based on cognitive load status. These assistance suggestions are generated by the intelligent assistance system.
[0015] The control module enhances the presentation intensity of risk warnings when the cognitive load status indicates that healthcare professionals are under high cognitive load.
[0016] Through this technical solution, this application provides a system capable of implementing the above-mentioned method. Through modular design, the system can efficiently acquire, judge, adjust and control the presentation of risk prompts, thereby effectively supporting medical professionals in their diagnosis and treatment work under high cognitive load and improving the efficiency and security of intelligent consultation interaction.
[0017] This application monitors doctors' cognitive load in real time and dynamically adjusts the visual salience, spatial location, and dynamic effects of risk prompts accordingly. This ensures that under high cognitive load, risk prompts are presented in a more prominent form (e.g., larger font, brighter colors, flashing effects, or placement in the screen area most frequently viewed by the doctor), effectively attracting the doctor's attention and successfully conveying the potential risks of AI-driven suggestions. This prevents doctors from unintentionally incorporating inaccurate information into their subsequent judgments without fully recognizing the potential risks of AI suggestions, thereby reducing the risk of misdiagnosis or delayed treatment and significantly improving the accuracy and safety of medical consultations.
[0018] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description
[0019] Figure 1 This is a flowchart of an intelligent consultation interaction method integrating human-computer collaboration according to the present invention.
[0020] Figure 2 This is a schematic diagram of the structure of an intelligent consultation and interaction system integrating human-computer collaboration according to the present invention. Detailed Implementation
[0021] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated in advance. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.
[0022] This embodiment provides an intelligent consultation interaction method and system integrating human-computer collaboration, combined with Figure 1 and Figure 2 As shown.
[0023] refer to Figure 1 An intelligent consultation interaction method integrating human-computer collaboration, the method includes the following steps:
[0024] Acquire visual behavior information, voice interaction information, and operational behavior information of medical professionals during the consultation process;
[0025] Based on visual behavioral information, voice interaction information, and operational behavior information, the cognitive load status of medical professionals can be assessed in real time.
[0026] Based on cognitive load status, the presentation of risk warnings for auxiliary suggestions is adjusted, and the auxiliary suggestions are generated by the intelligent assistance system.
[0027] When cognitive load indicates that healthcare professionals are under high cognitive load, increase the intensity of risk warning presentation.
[0028] Specifically, the "visual behavioral information" in this application refers to visual data captured by devices such as cameras, including eye movement trajectories, pupil diameter changes, and facial expressions, reflecting the attention focus, cognitive effort, and emotional state of medical professionals. "Voice interaction information" refers to auditory data recorded by devices such as microphone arrays, including speech rate, tone, volume, and content, revealing communication patterns, emotional fluctuations, and cognitive fluency. "Operational behavioral information" refers to behavioral data of medical professionals on electronic medical record systems, intelligent assistance systems, or other interactive interfaces, such as mouse clicks, keyboard input, and touchscreen operations, directly reflecting their task execution efficiency and interaction habits. All of the above information can be collected in real time using sensors, cameras, microphones deployed in the examination room, and software modules integrated with the doctor's workstation.
[0029] The implementation process first requires acquiring visual behavior, voice interaction, and operational behavior information of medical professionals during consultations. For example, eye-tracking devices installed in the consultation room can capture the real-time trajectory of the medical professional's eye movements and changes in pupil diameter; high-fidelity microphone arrays can record doctor-patient conversations and analyze the medical professional's voice; simultaneously, software modules integrated into the electronic medical record system can record various operational behaviors of the medical professional within the system. This acquisition of information is continuous and real-time, providing a data foundation for subsequent cognitive load assessment.
[0030] Secondly, based on the acquired visual behavior information, voice interaction information, and operational behavior information, the cognitive load status of medical professionals can be assessed in real time. For example, machine learning models can be used, taking the aforementioned multimodal data as input, and outputting the current cognitive load level of the medical professional through a pre-trained model. Specifically, when eye movement trajectories show frequent shifts in fixation point, continuously enlarging pupil diameter, increased speech rate, or increased hesitation and errors in operational behavior, the combined pattern of these indicators may indicate that the medical professional is in a state of high cognitive load.
[0031] Next, based on the determined cognitive load level, the presentation of risk warnings in the auxiliary suggestions is adjusted. These suggestions are generated by the intelligent assistance system. For example, when the cognitive load is low, the risk warnings can use a softer visual effect, such as displaying them in small font at the edge of the screen; when the cognitive load is moderate, the risk warnings can use a moderate intensity visual effect, such as displaying them in standard font and color in a fixed area of the screen.
[0032] Finally, when cognitive load indicates that healthcare professionals are under high cognitive load, the presentation intensity of risk warnings should be enhanced. For example, risk warnings can be presented as full-screen pop-ups, flashing borders, highlighted colors, larger fonts, or combined with voice prompts to ensure that under high cognitive load, risk warnings can effectively attract the attention of healthcare professionals and prevent them from being ignored.
[0033] The intelligent consultation interaction method integrating human-computer collaboration proposed in this application acquires visual behavior information, voice interaction information, and operational behavior information of medical professionals, and judges their cognitive load status in real time based on this information. Then, it adjusts the presentation of risk warnings in the auxiliary suggestions according to the cognitive load status, especially enhancing the presentation intensity of risk warnings under high cognitive load, thereby effectively solving the problem that risk warnings in existing intelligent consultation systems are easily ignored under high cognitive load.
[0034] Specifically, in traditional intelligent consultation systems, risk warnings are often presented in a fixed manner, failing to fully consider the high cognitive load that medical professionals may face during actual diagnosis and treatment. This fixed and inconspicuous presentation is easily overlooked by medical professionals in high cognitive load situations, leading to a failure to promptly obtain and process critical risk information, potentially resulting in misdiagnosis or delayed treatment.
[0035] This application significantly improves the effectiveness of risk warnings by introducing real-time assessments of the cognitive load state of healthcare professionals and dynamically adjusting the presentation of risk warnings accordingly. When healthcare professionals are under high cognitive load, the system can intelligently enhance the intensity of risk warning presentation, for example, by using more prominent visual effects (such as flashing, highlighting, and large font) or combining them with auditory cues to forcefully attract the healthcare professional's attention. This adaptive adjustment mechanism ensures that critical risk information can be effectively delivered and received even when healthcare professionals have limited attentional resources.
[0036] This application further proposes an intelligent consultation interaction method integrating human-computer collaboration. When an external interpersonal verbal communication event occurs in the consultation room, the method also includes the following steps:
[0037] Identify verbal communications between non-medical personnel and patients in the consultation room, and determine the urgency of the verbal communication based on the vocal characteristics of the communication and the immediate reactions of medical professionals.
[0038] When the urgency level indicated by verbal communication is high urgency, monitor the shift in the medical professional's gaze focus, changes in pupil diameter, facial expressions, and operational behaviors to confirm that the medical professional's attention has been forcibly diverted.
[0039] Once it is confirmed that the attention of medical professionals has been forcibly diverted, the cognitive isolation mode is activated, which minimizes, blurs or hides all non-critical information on the screen, displays critical risk information in the screen area, and plays isolation sound effects.
[0040] Monitor the recovery of attention among healthcare professionals and deactivate cognitive isolation mode once critical risk information has been confirmed to have been addressed.
[0041] Specifically, identifying verbal communication outside the doctor-patient relationship in the consultation room involves using speech recognition technology to monitor and analyze all speech within the consultation room in real time, distinguishing third-party speech such as conversations from nurses, family members, or other staff. The speech characteristics of verbal communication can include speech rate, volume, tone, and keywords. For example, faster speech, increased volume, a hurried tone, or the inclusion of keywords such as "urgent," "immediately," or "dangerous" can all serve as indicators of the level of urgency. The immediate reactions of medical professionals can be seen in their head turns, changes in eye contact, and adjustments in body posture; these reactions can help determine the extent to which communication affects their attention.
[0042] Furthermore, when the urgency of verbal communication is assessed as high urgency, the system initiates refined monitoring of the healthcare professional, including gaze focus shift, pupil diameter changes, facial expressions, and operational behaviors. Gait focus shift refers to the healthcare professional's gaze moving away from the current diagnostic interface or patient and turning towards an external communication source. Changes in pupil diameter can reflect a sharp change in cognitive load or emotional state. Facial expressions, such as furrowed brows and tense eyes, can also indicate a forced shift in attention. Operational behaviors, such as stopping current input or interrupting mouse and keyboard operations, are also manifestations of attention shift. By integrating this multimodal information, the system can accurately confirm whether a forced shift in the healthcare professional's attention has occurred.
[0043] Once the system confirms that the healthcare professional's attention has been forcibly diverted, it immediately activates a cognitive isolation mode. In this mode, all auxiliary information, background data, and secondary prompts on the screen that are not directly related to the current diagnostic task are minimized, blurred, or hidden to reduce visual distractions. Simultaneously, the system displays the most critical risk information in the current diagnostic process in a highly visually salient manner in a prominent area of the screen, such as the center or the area most frequently viewed by the healthcare professional. This information includes things like drug contraindications, allergy history, and abnormal results from important tests. Furthermore, a preset, non-invasive isolation sound effect is played, designed to gently guide the healthcare professional's auditory attention back to the critical information, rather than external distractions.
[0044] Finally, the system continuously monitors the recovery of the medical professional's attention. This includes, but is not limited to, monitoring whether their gaze refocuses on the key risk information on the screen, whether their pupil diameter stabilizes, whether their facial expression returns to calm, and whether they perform any actions related to the key risk information (such as clicking, confirming, or modifying medical records). Only when the system comprehensively determines that the medical professional's attention has been effectively restored and that the key risk information has been understood and processed will the cognitive isolation mode be lifted, allowing the screen to return to normal display, thereby ensuring the smooth progress of the diagnosis and treatment process.
[0045] This application's solution effectively overcomes the limitations of adjusting risk alerts solely based on the cognitive load of medical professionals by introducing a mechanism for identifying and processing external interpersonal verbal communication events. When high-urgency external communication occurs in the examination room, medical professionals' attention is easily forcibly diverted. Even if the system has enhanced the presentation of risk alerts based on their high cognitive load, key information may still be overlooked due to distraction. This application identifies the urgency of external communication and immediately activates a cognitive isolation mode upon confirming a forced shift in the medical professional's attention. This mode forcibly redirects the medical professional's attention back to the most critical diagnostic and treatment risks by minimizing non-critical information, highlighting key risk information, and supplementing it with isolation sound effects. Subsequently, by continuously monitoring the recovery of attention and confirming that key risk information has been processed, it ensures that key diagnostic and treatment risks can still be effectively perceived and handled under external interference, thereby avoiding potential medical errors.
[0046] In some preferred embodiments, assuming a doctor is diagnosing a patient, the intelligent assistance system displays risk warnings about the patient's history of drug allergies on the screen based on the doctor's cognitive load. At this moment, a nurse suddenly enters the examination room and reports to the doctor, in a rapid, loud voice, that another patient has experienced a sudden emergency.
[0047] First, the intelligent assistance system uses voice recognition technology to identify that the nurse's verbal communication is not a dialogue between the doctor and the patient. Based on the nurse's rapid speech, high volume, and urgent tone, combined with the doctor's immediate reaction of turning his head to the nurse, it determines that the verbal communication is a highly urgent event.
[0048] Subsequently, the system immediately monitors whether the doctor's gaze shifts from the current treatment interface to the nurse, whether the pupils dilate instantly, whether facial expressions show tension, and whether hand operations cease. When this multimodal information confirms that the doctor's attention has been forcibly diverted, the system immediately activates cognitive isolation mode.
[0049] In this mode, apart from the critical risk warning about the patient's drug allergy history, all other non-critical information (such as medical record details, historical test results, etc.) on the screen is minimized or blurred, while a gentle prompting sound plays. After processing the nurse's emergency report, the doctor's attention is drawn to the highlighted drug allergy risk warning on the screen. The doctor carefully reads and verbally confirms the risk, and adjusts the relevant medication prescription in the medical record system.
[0050] The system detected that the doctor's gaze was steadily focused on the risk warning area, that the doctor verbally expressed their understanding of the risk, and that they took appropriate action. When the system comprehensively determined that the doctor had a deep understanding and effective processing of the key risk information, the cognitive isolation mode was deactivated, and the screen returned to normal display. Through this process, even under external emergency interference, the doctor did not miss crucial drug allergy risk warnings, thus avoiding potential medical accidents.
[0051] This application further proposes the steps for monitoring the recovery of attention of healthcare professionals and, after confirming that key risk information has been processed, to remove the cognitive isolation mode, including:
[0052] When monitoring the recovery of attention of medical professionals, the following comprehensive judgments are made on the medical professionals: whether the medical professionals’ gaze is stably focused on the key risk information display area; whether the medical professionals’ verbal expression is analyzed through speech recognition to determine whether the medical professionals actively repeat or ask about key risk information, and whether they express understanding questions or confirmations about key risk information; at the same time, whether the medical professionals modify, supplement or confirm fields directly related to key risk information in the medical record system or related interactive interface after the cognitive isolation mode is activated.
[0053] When a comprehensive assessment of a healthcare professional's eye focus, verbal expression, and operational behavior indicates that the professional has a deep understanding and effective processing of key risk information, the cognitive isolation mode is lifted.
[0054] The aforementioned comprehensive assessment aims to evaluate the cognitive state of healthcare professionals and their processing of critical risk information from multiple dimensions. Specifically, determining whether a healthcare professional's gaze is stably focused on the area displaying critical risk information can be achieved through eye-tracking technology. For example, an infrared eye tracker can be used to capture the healthcare professional's eye movements in real time and analyze the duration, frequency, and stability of their gaze on the critical risk information area on the screen. Stable focus typically indicates that attention has been recovered from external distractions and is concentrated on the critical information on the screen.
[0055] Furthermore, by analyzing the verbal expressions of medical professionals through speech recognition, it's possible to determine whether they actively repeat or inquire about key risk information, and whether they express understanding questions or confirmations regarding that information. For example, when a medical professional verbally repeats key risk information (such as drug dosage, contraindications, etc.), or asks questions or makes confirmatory statements such as "The side effects of this drug are..., right?" or "The patient's allergy history needs to be reconfirmed," these can be identified as active cognitive processing of key risk information. This can be achieved through semantic analysis of the speech recognition results using natural language processing technology.
[0056] Furthermore, monitoring whether healthcare professionals modify, supplement, or confirm fields directly related to critical risk information in the medical record system or related interactive interfaces after activating cognitive isolation mode is a key indicator for assessing their behavioral processing. For example, if the critical risk information involves a patient's drug allergy history, and the healthcare professional enters, modifies, or confirms and saves the allergy history field in the medical record system after activating cognitive isolation mode, it indicates that they have translated the critical risk information into actual diagnostic and treatment actions. This can be achieved through real-time monitoring and analysis of system operation logs.
[0057] This application's solution addresses the limitations of single indicators in accurately assessing the recovery of attention and the depth of information processing by introducing a multimodal comprehensive judgment mechanism. When an external emergency forces a healthcare professional's attention to shift and activates a cognitive isolation mode, the system must ensure that their attention has truly recovered and that they have effectively processed critical risk information before safely lifting the isolation. Traditional attention monitoring may only focus on gaze, but focused gaze does not equate to deep understanding and processing. This solution constructs a more comprehensive and robust assessment model by combining information from three dimensions: focused gaze, verbal expression, and operational behavior. Focused gaze ensures the return of visual attention; verbal expression reflects cognitive understanding, memory, and active thinking; and operational behavior directly verifies whether information has been transformed into actual diagnostic and treatment decisions or records. It is precisely this multi-dimensional cross-validation that enables the system to more accurately determine whether healthcare professionals have achieved a deep understanding and effective processing of critical risk information, thereby avoiding the diagnostic and treatment risks that may arise from prematurely lifting the cognitive isolation mode due to misjudgment.
[0058] The steps for adjusting the presentation of risk warnings in auxiliary suggestions based on cognitive load include:
[0059] Record the response time, click frequency, and verbal confirmation of medical professionals to different risk warning presentation methods;
[0060] Based on response time, click frequency, and verbal confirmation, determine the individual preferences and historical interaction data of medical professionals;
[0061] Adjust the visual prominence, spatial location, and dynamic effects of risk warnings based on individual preferences and historical interaction data;
[0062] Based on eye movement trajectory data of medical professionals, identify the screen areas that medical professionals most frequently look at when under high cognitive load, and present risk warnings in the screen areas.
[0063] Based on the medical professionals' treatment habits and historical data, risk information relevant to the current treatment tasks is screened and presented.
[0064] Specifically, recording the response time, click frequency, and verbal confirmation of medical professionals to different risk warning presentation methods refers to the system continuously collecting data on the time required for medical professionals to respond (e.g., click confirmation, verbal repetition, eye contact) to risk warnings with different visual styles (such as color, size, and flashing frequency), locations, or dynamic effects, as well as the number of clicks and verbal confirmation content captured by speech recognition technology. This data is used to quantify medical professionals' acceptance, processing efficiency, and comprehension of specific presentation methods.
[0065] The system assesses individual preferences and historical interaction data of healthcare professionals based on response time, click frequency, and verbal confirmation. This analysis of recorded data allows the system to build a personalized interaction model for healthcare professionals. Individual preferences refer to the tendencies healthcare professionals exhibit towards specific risk warning presentation methods over long-term use, such as a preference for a particular color or screen area. Historical interaction data includes healthcare professionals' response patterns and effectiveness evaluations of various risk warnings under different cognitive load states. Machine learning algorithms can extract valuable patterns from this data to guide subsequent personalized adjustments.
[0066] In practical applications, the visual salience, spatial location, and dynamic effects of risk warnings are adjusted based on individual preferences and historical interaction data. Specifically, this means the system dynamically adjusts the visual attributes of risk warnings based on the identified individual preferences and historical interaction data of medical professionals. For example, visual salience can be enhanced or weakened by adjusting font size, color contrast, background brightness, or using highlighted borders. Spatial location refers to the display area of the risk warning on the screen, which can be placed in areas where medical professionals most frequently focus their attention or areas they habitually pay attention to. Dynamic effects can include flashing, fading, pop-up animations, etc., to attract attention when necessary and avoid distraction when unnecessary. The goal is to make the presentation of risk warnings more in line with the individual cognitive habits and current situational needs of medical professionals.
[0067] Furthermore, based on eye movement trajectory data of medical professionals, the screen area most frequently viewed by them under high cognitive load is identified, and risk warnings are presented in this area. Eye movement trajectory data can be acquired in real time using eye-tracking devices to analyze the focus, duration, and saccade path of medical professionals. Under high cognitive load, medical professionals have limited attentional resources; accurately presenting risk warnings in the screen area they most frequently view ensures that critical information is effectively received and prevents important warnings from being missed due to shifting gaze.
[0068] Furthermore, based on the medical professionals' treatment habits and historical data, the system filters and presents risk information relevant to the current treatment task. Treatment habits refer to the routine patterns of information acquisition and processing employed by medical professionals in specific diseases or treatment procedures. Historical data includes past diagnoses, treatment plans, and patient feedback. By combining this information with the current treatment task, the intelligent system can filter out risk information not directly related to the current treatment, presenting only the most critical and urgent risk alerts, avoiding information overload, and thus improving information processing efficiency.
[0069] This application's solution addresses the shortcomings of existing basic solutions in terms of the general applicability and low information transmission efficiency of risk warning presentation methods by comprehensively considering individual preferences of medical professionals, historical interaction data, real-time eye movement trajectories, and the clinical context. Specifically, by recording medical professionals' response data to different presentation methods, the system can learn and determine their personalized interaction habits, thus avoiding a "one-size-fits-all" presentation strategy. Furthermore, under high cognitive load conditions, by combining eye movement trajectory data, risk warnings are accurately presented in the screen area most frequently viewed by medical professionals, ensuring that key information is effectively captured even under conditions of limited attentional resources. In addition, filtering risk information based on clinical habits and historical data avoids unnecessary interference, making the presented risk warnings more targeted and timely, thereby significantly improving the effectiveness of risk warnings and the efficiency of information processing for medical professionals.
[0070] In some preferred embodiments, assuming a healthcare professional is processing a complex medical case and their cognitive load is assessed as high, the intelligent assistance system first retrieves the professional's historical interaction data. It finds that the professional responds quickly to flashing red risk warnings and clicks frequently. Eye-tracking data also shows that during periods of high cognitive load, their gaze is often focused on the lower right corner of the screen. Based on this data, the system presents risk warnings regarding drug interactions in the current medical case in a bright red flashing light in the lower right corner of the screen. Simultaneously, based on the professional's past treatment habits, the system only displays risk information directly related to the current patient's diagnosis and medication regimen, temporarily hiding or downplaying other non-urgent risk information. Once the professional's eyes are consistently focused on this area and they verbally confirm "I understand the drug interaction risks," the system records this interaction and further optimizes future presentation strategies based on their feedback.
[0071] This application further proposes an intelligent consultation interaction method integrating human-computer collaboration, which includes the following steps:
[0072] Real-time monitoring of patient emotional state, communication patterns, and communication strategies of medical professionals during doctor-patient communication to obtain information about the diagnosis and treatment context;
[0073] Continuously monitor the eye movement trajectory, pupil diameter changes, speech rate and operational behavior of medical professionals to obtain real-time information on their cognitive load levels;
[0074] When the information in the diagnosis and treatment context or the real-time cognitive load level changes, the visual salience, spatial location, and dynamic effect of the risk warning are dynamically adjusted according to the information in the diagnosis and treatment context and the real-time cognitive load level.
[0075] Specifically, contextual information in clinical practice refers to the environmental context information obtained during doctor-patient communication by analyzing the patient's emotional state, the communication patterns between the doctor and patient, and the communication strategies employed by medical professionals in real time. The patient's emotional state can be assessed using technologies such as facial expression recognition, voice tone analysis, and keyword recognition; communication patterns can include the fluency of the conversation, the number of interruptions, and the question-to-answer ratio; and the communication strategies of medical professionals can be evaluated based on their speech content, speaking speed, tone, and responses to patient questions. The acquisition of this information aims to comprehensively understand the background and atmosphere of the current clinical dialogue.
[0076] Furthermore, information on the immediate cognitive load level of healthcare professionals is obtained through continuous monitoring of their physiological and behavioral indicators. For example, eye movement trajectories can reflect their attention focus and scanning patterns; changes in pupil diameter are a physiological indicator of cognitive effort; speech rate can reflect their mental activity and stress level; and operational behaviors, including mouse clicks and keyboard input, reflect the efficiency and fluency of their interaction with the system. Through real-time analysis of this multimodal data, the cognitive load level of healthcare professionals at any given moment can be accurately assessed.
[0077] The dynamic adjustment of the visual salience, spatial location, and dynamic effects of risk warnings refers to the intelligent assistance system's immediate adaptive adjustment of the presentation of risk warnings based on the latest contextual and cognitive load data when the system detects significant changes in the diagnostic and treatment context or the immediate cognitive load level of medical professionals. For example, visual salience can refer to the color, font size, and flashing frequency of the warning box; spatial location can refer to the display area of the warning box on the screen, such as moving it from the edge of the screen to a more central position; and dynamic effects can include animation effects and gradient effects. This adjustment aims to ensure that risk warnings are perceived and processed by medical professionals at the most appropriate time and in the most effective way.
[0078] This application's solution effectively addresses the limitations of the aforementioned basic solutions in responding to real-time changes in the treatment process by introducing a real-time monitoring and dynamic adjustment mechanism for contextual information and the immediate cognitive load level of medical professionals. Specifically, by monitoring the patient's emotional state, doctor-patient communication patterns, and the communication strategies of medical professionals in real time, the system can construct a complete contextual understanding of the current treatment situation, thus avoiding the bias of relying solely on static or historical data. Simultaneously, continuous monitoring of the medical professional's eye movement trajectory, pupil diameter changes, speech rate, and operational behavior allows the system to accurately capture fluctuations in their immediate cognitive load during the treatment process, rather than relying solely on macroscopic cognitive load assessments. Therefore, when this crucial information changes, the system can immediately and dynamically adjust the visual salience, spatial location, and dynamic effects of risk warnings based on the latest treatment context and the medical professional's immediate cognitive load level, ensuring that the presentation of risk warnings highly matches the current situation, thereby maximizing the effectiveness of risk warnings without increasing the burden on medical professionals.
[0079] In some preferred embodiments, it is assumed that a healthcare professional is diagnosing a patient. The intelligent assistance system continuously monitors the patient and detects a sudden shift in facial expressions and tone of voice indicating significant anxiety while the patient describes symptoms, along with a slightly tense doctor-patient communication pattern. This is identified by the system as a change in the context of the diagnosis and treatment. Simultaneously, by monitoring the healthcare professional's eye movements and pupil diameter changes, the system detects a slight increase in their immediate cognitive load, suggesting they may be contemplating complex diagnostic questions. In this situation, if the system detects a medication contraindication risk warning highly relevant to the current symptoms, it will no longer present it solely based on the healthcare professional's historical preferences or general cognitive load. Instead, the system dynamically adjusts the presentation of the risk warning based on the patient's current anxiety context and the healthcare professional's slightly elevated immediate cognitive load. For example, the system may enhance the visual salience of the risk warning (e.g., by using a more striking color or a slight flicker), adjust its spatial location to a key screen area near the current focus of the medical professional's gaze, and may employ a brief, gentle dynamic effect to attract attention, ensuring that the critical risk information is perceived promptly and clearly at the moment when the medical professional needs to pay the most attention and can effectively handle it, thereby avoiding potential medical risks.
[0080] Obtaining contextual information for diagnosis and treatment also includes the following steps:
[0081] Collect patient characteristic information, including facial expressions, tone of voice, body movements, and key speech content;
[0082] The patient's characteristic information is matched with preset emotional and behavioral patterns to identify the patient's initial emotional state.
[0083] Analyze the patient's eye movement trajectory, pupil diameter changes, duration of facial micro-expressions, and consistency between verbal content and nonverbal behavior;
[0084] When there is inconsistency between verbal content and nonverbal behavior, or abnormal duration of facial micro-expressions, or changes in eye movement trajectory and pupil diameter indicating increased cognitive load, it is judged that the patient's emotional expression is performative.
[0085] When it is determined that a patient's emotional expression is performative, the patient's initial emotional state is modified, and the modified initial emotional state is used as part of the diagnostic and treatment context information. The emotional information corresponding to the patient's initial emotional state is marked as low confidence.
[0086] Specifically, when collecting patient characteristic information, facial expressions can be captured and analyzed in real time using facial recognition technology and cameras; voice tone can be captured using microphones and analyzed for acoustic features; body movements can be captured using posture recognition technology and depth sensors or cameras; and key speech content is extracted from the patient's verbal expression using speech recognition and natural language processing technologies. This characteristic information is used to comprehensively describe the patient's behavior during the consultation process.
[0087] Matching patient characteristics with pre-defined emotional and behavioral patterns involves using machine learning models, such as support vector machines, neural networks, or deep learning models, to process the collected patient characteristics and compare them with a pre-trained database containing various emotions (such as happiness, sadness, anxiety, anger, etc.) and their corresponding behavioral patterns, thereby initially identifying the patient's initial emotional state.
[0088] Furthermore, analyzing patients' eye movement trajectories, pupil diameter changes, duration of facial micro-expressions, and the consistency between verbal content and nonverbal behavior is to assess the authenticity of emotions at a deeper level. For example, eye movement trajectories and pupil diameter changes can reflect a patient's cognitive load and emotional arousal level. When a patient is in a state of emotional tension or cognitive effort, the pupil diameter usually increases, and eye movement patterns may also change. Facial micro-expressions refer to unconscious facial expressions that last for very short periods (usually less than 0.5 seconds). They often reveal true inner emotions. By measuring their duration and comparing them with a database of real emotional micro-expression features, abnormal or potentially performative expressions can be identified. The consistency between verbal content and nonverbal behavior is assessed by comparing the emotional tendencies expressed verbally (obtained through natural language processing) with the emotional tendencies expressed through facial expressions and body movements (obtained through visual analysis). If there is a significant difference between the two, it may indicate inconsistency in emotional expression.
[0089] Therefore, when there is inconsistency between verbal content and nonverbal behavior, abnormal duration of facial micro-expressions, or changes in eye movement trajectory and pupil diameter indicating increased cognitive load, these signs can all serve as evidence of the performative nature of the patient's emotional expression. For example, if a patient verbally claims to be in "extreme pain," but their facial expression does not show corresponding pain micro-expressions, or the micro-expressions are too long / too short, and the pupil diameter does not show a physiological response related to pain, then this can be judged as performative emotion.
[0090] As a preferred implementation, when a patient's emotional expression is determined to be performative, the initial emotional state is modified. This means the system does not fully accept the initially identified emotional state but adjusts it, for example, by reducing its intensity or changing its nature, to reflect its unnaturalness. Simultaneously, marking the emotional information corresponding to the patient's initial emotional state as having low confidence serves to remind the intelligent assistance system or medical professionals that this emotional information may not be entirely reliable and should be carefully considered in subsequent assistance suggestion generation and risk warning presentation.
[0091] This application further includes steps such as analyzing the patient's eye movement trajectory, pupil diameter changes, duration of facial micro-expressions, and consistency between verbal content and nonverbal behavior:
[0092] By analyzing the patient's eye movement trajectory in real time, identifying the patient's fixation point shift and fixation duration, and combining this with changes in the patient's pupil diameter, it is possible to determine whether there are physiological reactions related to emotional stress or cognitive effort. When such reactions are present, they serve as evidence of increased cognitive load in the patient.
[0093] Facial expression recognition technology is used to measure and record the duration of a patient's facial micro-expressions and compare them with a pre-set database of real emotional micro-expression features to identify abnormalities in the duration of facial micro-expressions.
[0094] Natural language processing technology is used to analyze the emotional tendency of patients' speech content, and the results of the emotional tendency analysis are compared with the emotional tendency expressed by facial expressions and body movements obtained through visual analysis, in order to assess the degree of consistency between the patient's speech content and nonverbal behavior in terms of emotional tendency.
[0095] Because genuine emotions are usually accompanied by specific physiological responses (such as pupil changes and eye movements) and unconscious micro-expressions, and verbal and nonverbal expressions tend to be consistent, while performative emotions may exhibit inconsistencies or abnormalities in these aspects, the system can accurately determine the performative nature of a patient's emotional expression when it detects inconsistencies between verbal and nonverbal communication, abnormal duration of micro-expressions, or physiological indicators showing increased cognitive load. Ultimately, by identifying performative emotions, the system can correct the patient's initial emotional state and mark it as low-confidence, thereby ensuring that the diagnostic and treatment context information relied upon by the intelligent assistance system is closer to the patient's actual situation.
[0096] In some preferred embodiments, suppose a patient presents with abdominal pain and, while communicating with a medical professional, verbally states that the pain is "severe and unbearable," accompanied by a furrowed brow and a curled-up posture. The intelligent assistance system first collects characteristic information such as the patient's facial expressions, tone of voice, body movements, and key verbal content. This initial identification may indicate that the patient is in an initial emotional state of "severe pain."
[0097] However, the proposed solution will be further analyzed. For example, monitoring using an eye-tracking device revealed that although the patient verbally expressed severe pain, the change in pupil diameter was not significant, and the eye movement trajectory did not show a persistent tension or avoidance response associated with severe pain. Simultaneously, the facial micro-expression recognition system detected that the patient's "painful" expression was abnormally short-lived, inconsistent with the typical sustained pattern of micro-expressions during genuine severe pain. Furthermore, the natural language processing system analyzed the patient's speech content and found that when describing pain, the patient's speech rate was steady, and the tone lacked obvious tremor or rapidity, which was inconsistent with the nonverbal behavior of the patient's claimed "unbearable" severe pain.
[0098] Based on the comprehensive analysis of the aforementioned multimodal data, the system determined that the patient's emotional expression was performative. Therefore, the system did not fully accept the initial emotional state of "severe pain," but rather modified it, for example, by reducing its intensity and marking it as low confidence. This modified patient emotional state, as part of the diagnostic context information, was transmitted to the intelligent assistance system. When adjusting the presentation of risk warnings in the assistance suggestions, the intelligent assistance system considered the low confidence level of the patient's emotional information and may place more emphasis on objective physiological indicators or examination results, rather than relying entirely on the patient's subjective description. This avoids misleading risk warnings due to the patient's performative emotions and ensures that medical professionals receive more accurate and reliable diagnostic context information.
[0099] Specifically, in the process of acquiring the above-mentioned diagnostic and treatment context information, the steps of analyzing the patient's eye movement trajectory, pupil diameter changes, duration of facial micro-expressions, and consistency between verbal content and nonverbal behavior may include the following.
[0100] By analyzing patients' eye movement trajectories in real time, the system identifies fixation point shifts and fixation durations, and combines this with changes in pupil diameter to determine the presence of physiological responses related to emotional stress or cognitive effort. When present, these responses serve as evidence of increased cognitive load. Facial expression recognition technology measures and records the duration of patients' facial micro-expressions, comparing them with a pre-defined database of real emotional micro-expression features to identify abnormalities in duration. Natural language processing technology is used to analyze the emotional tendency of patients' speech content, comparing the results with the emotional tendencies expressed by facial expressions and body movements obtained through visual analysis to assess the consistency between the patient's verbal content and nonverbal behavior in terms of emotional tendency.
[0101] This method involves real-time analysis of the patient's eye movement trajectory, identifying fixation point shifts and fixation duration, and combining this with changes in pupil diameter to assess the patient's cognitive load through objective physiological indicators. For example, eye-tracking devices can be used to capture the patient's fixation point, fixation duration, and pupil diameter data. When a patient's fixation point shifts frequently, fixation duration is too short or too long, and pupil diameter dilates significantly, these physiological responses are often seen as manifestations of emotional stress or cognitive effort, and thus can serve as strong evidence of increased cognitive load.
[0102] Furthermore, facial expression recognition technology is used to measure and record the duration of a patient's facial micro-expressions, comparing them with a pre-defined database of real emotional micro-expressions. The aim is to identify subtle abnormalities in the patient's emotional expression. Micro-expressions are involuntary facial expressions that are extremely short in duration (typically between 0.5 and 4 seconds) and often reveal true inner emotions. High-precision facial expression recognition algorithms can accurately capture and measure the duration of these micro-expressions. Comparing the measurement results with a feature database containing typical durations of real emotional micro-expressions can identify micro-expressions with abnormal durations, such as those that are too long or too short. This may suggest that the patient's emotional expression is not entirely genuine and contains performative elements.
[0103] Furthermore, natural language processing (NLP) techniques were used to analyze the emotional tendency of patients' verbal content, and the results were compared with the emotional tendencies expressed by facial expressions and body movements obtained through visual analysis. This aimed to assess the consistency between the patient's verbal and nonverbal emotional expressions. Specifically, an emotional analysis model was used to classify and assess the intensity of emotions in the patient's verbal descriptions. Simultaneously, computer vision techniques were used to analyze the patient's facial expressions and body movements to extract the emotional tendencies they expressed. The results of these two analyses were then cross-compared. If there is a significant inconsistency between the emotional content expressed in verbal content and the emotional content expressed in nonverbal behavior—for example, if a patient verbally claims to be relaxed and happy, but their facial expressions or body movements show tension or anxiety—it indicates that their emotional expression may be performative.
[0104] This application's solution, through multimodal data fusion analysis, enables a more comprehensive and in-depth understanding of the patient's true emotional state. Specifically, by monitoring the patient's eye movement trajectory and pupil diameter changes in real time, it can capture the patient's physiological responses during emotional stress or cognitive effort, providing objective evidence for assessing increased cognitive load. Simultaneously, using facial expression recognition technology to identify abnormalities in the duration of micro-expressions can effectively distinguish between genuine and performative emotions, as genuine micro-expressions are typically short-lived and difficult to fake. Furthermore, by using natural language processing technology to analyze the emotional tendency of verbal content and comparing it with the emotional tendency expressed by facial expressions and body movements obtained from visual analysis, it can reveal inconsistencies between the patient's verbal and nonverbal behaviors. Such inconsistencies are often important indicators of performative emotions. Through the above comprehensive analysis, this application can provide a detailed assessment of the patient's emotional expression from multiple dimensions, including physiological, micro-behavioral, and multimodal consistency, thereby improving the accuracy of judging the patient's true emotional state.
[0105] In this regard, this application further proposes steps for determining whether a patient's emotional expression is performative, including:
[0106] Obtain the patient's medical history, which includes diagnoses of diseases that affect facial expressions, speech, or body movements;
[0107] The patient's characteristic information is correlated with the patient's medical history information to determine whether the diseases contained in the patient's medical history information are related to the patient's emotional expression characteristics, thereby determining whether the patient's medical history information contains diseases that affect emotional expression;
[0108] When a patient's medical history indicates the presence of a disease that affects emotional expression, the patient's behavioral data is matched with disease symptom characteristics.
[0109] When a patient's behavioral data matches the characteristics of disease symptoms, it is determined that the behavior is caused by the disease.
[0110] When a patient's behavioral data does not match the characteristics of their disease symptoms, or when their medical history does not indicate the presence of a disease that affects their emotional expression, it is determined that the patient's emotional expression is performative.
[0111] Specifically, obtaining a patient's medical history involves extracting disease diagnosis information that may be related to the patient's facial expressions, verbal expressions, or body movements from the patient's electronic medical record system, historical medical records, or other reliable medical information sources. For example, these diagnoses may include facial nerve palsy, Parkinson's disease, stroke sequelae, dysarthria, aphasia, or certain mental illnesses, which may cause atypical or incoherent physiological manifestations when the patient expresses emotions. Furthermore, correlation analysis between the patient's characteristic information and their medical history aims to identify whether the patient's current behavior may be caused by their known diseases. For example, if the patient's medical history includes a diagnosis of facial nerve palsy, then the asymmetry or stiffness of their facial expressions may not stem from emotional performance but rather be a direct symptom of the disease.
[0112] When a patient's medical history indicates a disease affecting emotional expression, it's necessary to match the patient's behavioral data with disease symptom characteristics. This behavioral data refers to real-time facial expressions, voice tone, and body movements collected through sensors such as vision and speech sensors. Disease symptom characteristics refer to observable behavioral patterns or physiological indicators related to a specific disease; for example, facial nerve palsy may cause weakness or numbness in one side of the face, and Parkinson's disease may cause a mask-like face or limb tremors. By comparing real-time behavioral data with these pre-defined disease symptom characteristics, it can be determined whether the patient's abnormal behavior matches disease symptoms. Therefore, when the patient's behavioral data matches disease symptom characteristics, it can be determined that the behavior is caused by the disease, rather than an emotional performance. Conversely, when the patient's behavioral data does not match disease symptom characteristics, or when the patient's medical history does not indicate a disease affecting emotional expression, it can be more reliably determined that the patient's emotional expression is performative.
[0113] This application's solution addresses the potential misjudgment issues in traditional methods for assessing the performativity of patients' emotional expressions by incorporating consideration of patient medical history. Specifically, judging performativity solely based on inconsistencies between verbal and nonverbal behavior, abnormal duration of micro-expressions, or increased cognitive load fails to distinguish between atypical expressions caused by physiological illness and performances driven by subjective intent. This solution, by acquiring and correlating patient medical history, can identify disease diagnoses that may affect a patient's emotional expression. When such a disease is identified, the system further matches the patient's real-time behavioral data with the symptom characteristics of the corresponding disease. If the behavioral performance matches the disease symptom characteristics, the abnormal behavior is attributed to the disease, thus avoiding misjudgment as performativity. This mechanism ensures that the performativity of a patient's emotional expression can only be assessed after ruling out the influence of physiological illness, significantly improving the accuracy and reliability of the assessment.
[0114] In some preferred embodiments, assuming a patient's facial expressions appear stiff and asymmetrical during a consultation, and their speech is slightly delayed, the system might initially determine, based on the judgment logic in the aforementioned diagnostic and treatment context information acquisition method, that the patient's emotional expression is performative. However, with the solution of this application, the system first obtains the patient's medical history information. If the medical history information shows that the patient has been diagnosed with facial nerve palsy, the system will perform a correlation analysis between this diagnosis and the patient's current facial expressions and speech behavior. Further, the system will match the patient's facial stiffness and delayed speech, among other behavioral data, with the typical symptom characteristics of facial nerve palsy. If the matching results show that these behaviors highly match the symptom characteristics of facial nerve palsy, such as weakness or numbness of one side of the facial muscles, the system will determine that these behaviors are caused by facial nerve palsy, rather than a deliberate emotional performance by the patient. In this scenario, the system will not flag the patient's emotional expression as performative. Instead, it will adjust the diagnostic context information based on the patient's true emotional state (e.g., anxiety or frustration that may be caused by the illness) and adjust the presentation of risk warnings for auxiliary suggestions accordingly. This avoids misjudging the patient's emotions and ensures the accuracy and appropriateness of the auxiliary suggestions.
[0115] This application further proposes that the steps for determining whether a patient's emotional expression is performative include:
[0116] Obtain the patient's verbal description and emotional expression of the current medical situation;
[0117] Analyze whether there is concern about the treatment situation in the verbal description, and determine the degree of concern in the verbal description;
[0118] Assess the intensity and duration of emotional expressions and match them with the level of concern in verbal descriptions;
[0119] When a patient’s verbal description indicates concern about the treatment situation, and the intensity and duration of the emotional expression are consistent with the level of concern in the verbal description, the patient’s emotional expression is judged to be a genuine emotion.
[0120] When the verbal description does not indicate concern about the treatment situation, or the intensity and duration of the emotional expression are abnormal, or the intensity and duration of the emotional expression are inconsistent with the level of concern in the verbal description, the patient's emotional expression is judged to be performative.
[0121] Specifically, obtaining a patient's verbal description and emotional expression regarding the current medical situation involves converting the patient's speech during the consultation into text using speech recognition technology, and then using natural language processing technology to extract the patient's direct or indirect descriptions of the current disease, treatment plan, prognosis, and other medical situations, as well as the emotional vocabulary and expressions contained therein. Simultaneously, through technologies such as speech emotion recognition, facial expression recognition, and body language analysis, the patient's emotional expression when conveying this content is assessed, including its intensity (e.g., loudness of the voice, speech rate, and facial muscle activity) and duration.
[0122] Analyzing whether there is concern about the treatment situation in the verbal description and determining the degree of concern in the verbal description can be understood as using an emotion analysis model to conduct in-depth analysis of the patient's verbal text, identifying keywords, phrases, and sentence structures related to negative emotions such as "worry," "anxiety," "fear," and "uncertainty," and quantitatively or qualitatively assessing the degree of concern about the treatment situation in the patient's verbal expression based on the frequency, intensity, and context of these elements. For example, a worry index can be calculated using a pre-defined worry dictionary and grammatical rules.
[0123] In practical applications, assessing the intensity and duration of emotional expression and matching it with the level of worry in verbal descriptions involves comparing the intensity and duration of emotions identified through nonverbal behaviors (such as facial expressions, tone of voice, and body language) with the level of worry derived from verbal content analysis. For example, if a patient expresses high levels of worry verbally, but their facial expressions and tone of voice appear unusually calm or are too short-lived, a mismatch may exist. Conversely, if the level of worry in the verbal description is highly consistent with the intensity and duration of nonverbal emotional expression, it is more likely to be judged as genuine emotion.
[0124] This application's approach effectively overcomes the limitations of relying solely on physiological and behavioral characteristics for emotion judgment by introducing a comprehensive analysis of patients' verbal descriptions and emotional expressions. Because patients' verbal descriptions typically reflect their conscious cognition and understanding of the situation, while emotional expressions may contain deeper, sometimes unconscious, emotional outpourings, matching and comparing the two can create a mutually verifying mechanism. When the level of worry expressed in a patient's verbal content is highly consistent with the intensity and duration of their nonverbal emotional expression, it indicates that the patient's internal and external emotional expressions are coordinated, thus enhancing the confidence in judging the true emotion. Conversely, if there is a significant inconsistency between the two, such as a discrepancy between the level of worry expressed verbally and the intensity or duration of the actual emotional expression, it suggests that the patient may be concealing, exaggerating, or performing emotions, thus enabling a more accurate identification of performative emotions.
[0125] In some preferred embodiments, suppose a patient is informed of their diagnosis in an examination room. The patient verbally describes repeatedly that "I am very worried about this result and don't know what to do," and the analysis of their speech content indicates a "high" level of concern. However, facial expression recognition and voice emotion analysis reveal that while the patient's facial expression is slightly serious, it is short-lived, and the fluctuation in tone of voice is not strong. The intensity and duration of the emotional expression are assessed as "moderate to low," which is significantly inconsistent with the "high concern" stated in the verbal description. In this case, the solution of this application would determine that the patient's emotional expression is performative, suggesting that medical professionals may need to further investigate the patient's true concerns or the underlying reasons for their expression, rather than simply accepting their surface emotions.
[0126] If a patient verbally describes "I'm a little nervous, but I trust the doctor," with a level of "moderate" anxiety, and their facial expression is slightly tense, and their speech is a little fast, but the overall intensity and duration of their emotional expression match the "moderate nervousness" described verbally, then the proposed solution will determine that the patient's emotional expression is genuine, thereby providing medical professionals with more accurate feedback on the patient's emotions.
[0127] In some embodiments described above, this application proposes an intelligent consultation interaction method integrating human-computer collaboration. This method can dynamically adjust the presentation of risk warnings in auxiliary suggestions based on the cognitive load status of medical professionals. However, this method only describes its operational steps and does not explicitly provide the specific system architecture and functional modules for implementing these steps. In practical applications, the lack of a clear and integrated system framework may lead to inefficient or insufficiently coordinated implementation of data collection, cognitive load assessment, and risk warning adjustment, thereby affecting the stability and reliability of the overall solution.
[0128] refer to Figure 2This application proposes an intelligent consultation interaction system integrating human-computer collaboration, applied to the aforementioned intelligent consultation interaction method integrating human-computer collaboration. The system includes:
[0129] The acquisition module acquires visual behavior information, voice interaction information, and operational behavior information of medical professionals during the consultation process;
[0130] The judgment module, based on visual behavior information, voice interaction information, and operational behavior information, judges the cognitive load status of medical professionals in real time.
[0131] The adjustment module adjusts the presentation of risk warnings for assistance suggestions based on cognitive load status. These assistance suggestions are generated by the intelligent assistance system.
[0132] The control module enhances the presentation intensity of risk warnings when the cognitive load status indicates that healthcare professionals are under high cognitive load.
[0133] Specifically, the acquisition module is configured to comprehensively collect multimodal behavioral data from healthcare professionals during the consultation process. Visual behavioral information includes physiological indicators such as eye movement trajectories, fixation points, and pupil diameter changes, typically acquired in real-time using high-precision eye-tracking devices integrated into the monitor or clinic environment. Verbal interaction information refers to the content, rate, tone, and volume of verbal communication between healthcare professionals and patients, as well as with intelligent assistance systems; its acquisition usually relies on high-sensitivity microphone arrays and advanced speech recognition technology. Operational behavioral information covers all actions performed by healthcare professionals on electronic medical record systems, diagnostic software, or other interactive interfaces, such as mouse clicks, keyboard input, scrolling, and data modification; this information can be captured through system logs or screen behavior analysis tools. The acquisition of this multidimensional information aims to provide rich and real-time input for subsequent cognitive load assessment.
[0134] The judgment module is designed to comprehensively analyze visual behavior information, voice interaction information, and operational behavior information collected by the acquisition module to assess the cognitive load of healthcare professionals in real time. This module employs multimodal data fusion technology, combined with machine learning or deep learning models, such as recurrent neural networks, long short-term memory networks, or attention mechanism models, to perform feature extraction, weighted fusion, and pattern recognition on different types of data. Through training, the model can identify behavioral patterns associated with different cognitive load levels (e.g., low, medium, and high load) and output real-time cognitive load assessment results. This real-time judgment mechanism ensures that the system can instantly perceive changes in the psychological state of healthcare professionals.
[0135] In practical applications, the adjustment module is responsible for dynamically adjusting the presentation of risk warnings in the assistance suggestions generated by the intelligent assistance system based on the cognitive load status output by the judgment module. Assistance suggestions are typically based on patient medical records, treatment guidelines, drug interaction databases, and other information, providing decision support to healthcare professionals through the intelligent assistance system. The presentation of risk warnings can be adjusted in dimensions including, but not limited to, visual salience (e.g., by changing color saturation, font size, background contrast, flashing frequency, etc.), spatial location (e.g., placing the warning information in the center, edge, specific area of the screen, or as a pop-up), and dynamic effects (e.g., introducing gradients, animations, or vibration cues). The aim is to optimize the efficiency of risk information delivery without increasing additional cognitive load.
[0136] Furthermore, the control module is activated to enhance the presentation intensity of risk warnings when the judgment module indicates that the medical professional is in a state of high cognitive load. Enhancing presentation intensity means presenting the risk warnings in a more compelling and harder-to-ignore manner. For example, the font size of the risk warning can be enlarged to its maximum, the color set to a high-contrast warning color (such as red or orange), the flashing frequency increased, it placed in the most prominent visual area of the screen, and even accompanied by short, non-invasive auditory cues. The aim is to ensure that critical risk information can penetrate the cognitive barriers of medical professionals in high-pressure or information-overloaded diagnostic and treatment situations, and be effectively perceived and processed, thereby minimizing potential medical errors caused by excessive cognitive load.
[0137] The system solution proposed in this application achieves the effective implementation of the aforementioned integrated human-computer collaboration intelligent consultation interaction method through modular design.
[0138] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the scope of protection of the present invention. Therefore, all equivalent technical changes made based on the content of the present invention specification and drawings are included within the scope of protection of the present invention. Furthermore, the elements therein can be updated as technology develops.
Claims
1. A method for intelligent consultation interaction integrating human-computer collaboration, characterized in that, The method includes the following steps: Acquire visual behavior information, voice interaction information, and operational behavior information of medical professionals during the consultation process; Based on visual behavioral information, voice interaction information, and operational behavior information, the cognitive load status of medical professionals can be assessed in real time. Based on cognitive load status, the presentation of risk warnings for auxiliary suggestions is adjusted, and the auxiliary suggestions are generated by the intelligent assistance system. When cognitive load indicates that healthcare professionals are under high cognitive load, increase the intensity of risk warning presentation.
2. The intelligent consultation interaction method integrating human-computer collaboration as described in claim 1, characterized in that, When an external interpersonal verbal communication event occurs in the consultation room, the method also includes the following steps: Identify verbal communications between non-medical personnel and patients in the consultation room, and determine the urgency of the verbal communication based on the vocal characteristics of the communication and the immediate reactions of medical professionals. When the urgency level indicated by verbal communication is high urgency, monitor the medical professional's gaze shift, pupil diameter changes, facial expressions, and operational behavior to confirm that the medical professional's attention has been forcibly diverted; Once it is confirmed that the attention of medical professionals has been forcibly diverted, the cognitive isolation mode is activated, which minimizes, blurs or hides all non-critical information on the screen, displays critical risk information in the screen area, and plays isolation sound effects. Monitor the recovery of attention among healthcare professionals and deactivate cognitive isolation mode once critical risk information has been confirmed to have been addressed.
3. The intelligent consultation interaction method integrating human-computer collaboration as described in claim 2, characterized in that, Monitoring the recovery of attention among healthcare professionals and, after confirming that key risk information has been addressed, the steps to de-identify from cognitive isolation include: When monitoring the recovery of attention of medical professionals, the following comprehensive judgments are made on the medical professionals: whether the medical professionals’ gaze is stably focused on the key risk information display area; whether the medical professionals’ verbal expression is analyzed through speech recognition to determine whether the medical professionals actively repeat or ask about key risk information, and whether they express understanding questions or confirmations about key risk information; at the same time, whether the medical professionals modify, supplement or confirm fields directly related to key risk information in the medical record system or related interactive interface after the cognitive isolation mode is activated. When a comprehensive assessment of a healthcare professional's eye focus, verbal expression, and operational behavior indicates that the professional has a deep understanding and effective processing of key risk information, the cognitive isolation mode is lifted.
4. The intelligent consultation interaction method integrating human-computer collaboration as described in claim 1, characterized in that, The steps for adjusting the presentation of risk warnings in auxiliary suggestions based on cognitive load include: Record the response time, click frequency, and verbal confirmation of medical professionals to different risk warning presentation methods; Based on response time, click frequency, and verbal confirmation, determine the individual preferences and historical interaction data of medical professionals; Adjust the visual prominence, spatial location, and dynamic effects of risk warnings based on individual preferences and historical interaction data; Based on eye movement trajectory data of medical professionals, identify the screen areas that medical professionals most frequently look at when under high cognitive load, and present risk warnings in the screen areas. Based on the medical professionals' treatment habits and historical data, risk information relevant to the current treatment tasks is screened and presented.
5. The intelligent consultation interaction method integrating human-computer collaboration as described in claim 4, characterized in that, The method also includes the following steps: Real-time monitoring of patient emotional state, communication patterns, and communication strategies of medical professionals during doctor-patient communication to obtain information about the diagnosis and treatment context; Continuously monitor the eye movement trajectory, pupil diameter changes, speech rate, and operational behavior of medical professionals to obtain real-time information on their cognitive load levels; When the information in the diagnosis and treatment context or the real-time cognitive load level changes, the visual salience, spatial location, and dynamic effect of the risk warning are dynamically adjusted according to the information in the diagnosis and treatment context and the real-time cognitive load level.
6. The intelligent consultation interaction method integrating human-computer collaboration as described in claim 5, characterized in that, Obtaining contextual information for diagnosis and treatment also includes the following steps: Collect patient characteristic information, including facial expressions, tone of voice, body movements, and key speech content; The patient's characteristic information is matched with preset emotional and behavioral patterns to identify the patient's initial emotional state. Analyze the patient's eye movement trajectory, pupil diameter changes, duration of facial micro-expressions, and consistency between verbal content and nonverbal behavior; When there is inconsistency between verbal content and nonverbal behavior, or abnormal duration of facial micro-expressions, or changes in eye movement trajectory and pupil diameter indicating increased cognitive load, it is judged that the patient's emotional expression is performative. When it is determined that a patient's emotional expression is performative, the patient's initial emotional state is modified, and the modified initial emotional state is used as part of the diagnostic and treatment context information. The emotional information corresponding to the patient's initial emotional state is marked as low confidence.
7. The intelligent consultation interaction method integrating human-computer collaboration as described in claim 6, characterized in that, The steps for analyzing a patient's eye movement trajectory, pupil diameter changes, duration of facial micro-expressions, and consistency between verbal content and nonverbal behavior include: By analyzing the patient's eye movement trajectory in real time, identifying the patient's fixation point shift and fixation duration, and combining this with changes in the patient's pupil diameter, it is possible to determine whether there are physiological reactions related to emotional stress or cognitive effort. When such reactions are present, they serve as evidence of increased cognitive load in the patient. Facial expression recognition technology is used to measure and record the duration of a patient's facial micro-expressions and compare them with a pre-set database of real emotional micro-expression features to identify abnormalities in the duration of facial micro-expressions. Natural language processing technology is used to analyze the emotional tendency of patients' speech content, and the results of the emotional tendency analysis are compared with the emotional tendency expressed by facial expressions and body movements obtained through visual analysis, in order to assess the degree of consistency between the patient's speech content and nonverbal behavior in terms of emotional tendency.
8. The intelligent consultation interaction method integrating human-computer collaboration as described in claim 6, characterized in that, Determining whether a patient's emotional expression is performative also includes the following steps: Obtain the patient's medical history, which includes diagnoses of diseases that affect facial expressions, speech, or body movements; The patient's characteristic information is correlated with the patient's medical history information to determine whether the diseases contained in the patient's medical history information are related to the patient's emotional expression characteristics, thereby determining whether the patient's medical history information contains diseases that affect emotional expression; When a patient's medical history indicates the presence of a disease that affects emotional expression, the patient's behavioral data is matched with disease symptom characteristics. When a patient's behavioral data matches the characteristics of disease symptoms, it is determined that the behavior is caused by the disease. When a patient's behavioral data does not match the characteristics of their disease symptoms, or when their medical history does not indicate the presence of a disease that affects their emotional expression, it is determined that the patient's emotional expression is performative.
9. The intelligent consultation interaction method integrating human-computer collaboration as described in claim 6, characterized in that, Determining whether a patient's emotional expression is performative also includes the following steps: Obtain the patient's verbal description and emotional expression of the current medical situation; Analyze whether there is concern about the treatment situation in the verbal description, and determine the degree of concern in the verbal description; Assess the intensity and duration of emotional expressions and match them with the level of concern in verbal descriptions; When a patient’s verbal description indicates concern about the treatment situation, and the intensity and duration of the emotional expression are consistent with the level of concern in the verbal description, the patient’s emotional expression is judged to be a genuine emotion. When the verbal description does not indicate concern about the treatment situation, or the intensity and duration of the emotional expression are abnormal, or the intensity and duration of the emotional expression are inconsistent with the level of concern in the verbal description, the patient's emotional expression is judged to be performative.
10. An intelligent consultation interaction system integrating human-computer collaboration, applied to the intelligent consultation interaction method integrating human-computer collaboration as described in claim 1, characterized in that, The system includes: The acquisition module acquires visual behavior information, voice interaction information, and operational behavior information of medical professionals during the consultation process; The judgment module, based on visual behavior information, voice interaction information, and operational behavior information, judges the cognitive load status of medical professionals in real time. The adjustment module adjusts the presentation of risk warnings for assistance suggestions based on cognitive load status. These assistance suggestions are generated by the intelligent assistance system. The control module enhances the presentation intensity of risk warnings when the cognitive load status indicates that healthcare professionals are under high cognitive load.