A virtual diagnosis and treatment interaction method and system based on a patient agent
By constructing a structured, hierarchical patient intelligent agent, and combining cognitive boundary constraints and role characteristics, information disclosure is dynamically controlled, solving the problem of complete information output in virtual intelligent agent interaction. This achieves virtual diagnosis and treatment interaction that conforms to the natural progression law, improving realism and consistency.
Patent Information
- Application Number
- CN202610351391.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-21
- Publication Date
- 2026-06-12
AI Technical Summary
Existing virtual intelligent agent interaction systems lack a dynamic data disclosure control mechanism based on the depth of interaction and the cognitive boundaries of the intelligent agent role. This results in the complete output of information at once or information exceeding the deep-seated related information set by the role, which undermines the realism of multi-round human-computer interaction and the continuity of the conversation state.
A virtual diagnosis and treatment interaction method based on patient intelligent agents is adopted. By constructing a structured hierarchical organization of patient information cognitive boundary constraints and role characteristics, multi-source input data is continuously received to construct a session-level interaction context. The information range is determined within the cognitive boundary according to the interaction stage and information disclosure rules, and response results that conform to the natural progression law are generated and the session state is updated.
It achieves a highly realistic progressive consultation interaction loop, improves the consistency of roles and the realism of scenarios in multi-round virtual consultations, and overcomes the drawback of traditional models that mechanically throw out all medical records.
Smart Images

Figure CN122201571A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a virtual diagnosis and treatment interaction method and system based on a patient intelligent agent. Background Technology
[0002] With the continuous development of artificial intelligence and natural language processing technologies, vertical domain scenario simulation systems based on human-computer interaction are gradually becoming an important direction for professional skills training and intelligent auxiliary system testing. In these application scenarios, the system typically needs to construct virtual intelligent agents with specific role settings in order to highly reproduce the multi-turn dialogues, information collection, material retrieval, and logical reasoning processes in real business processes in a virtual environment, thereby providing a reliable digital interactive environment for professionals' skills training or human-computer collaborative testing.
[0003] Existing virtual intelligent agent interaction solutions typically employ static script configuration, rule engine matching, or direct generation of responses using a general-purpose large language model combined with full background data. These conventional implementations can receive user text or voice input and, within a given static knowledge base, provide corresponding answers through basic text retrieval matching or natural language generation techniques, thus meeting, to some extent, the technical requirements for basic question-and-answer interaction and static business scenario demonstration.
[0004] However, in real-world interaction scenarios, information providers are often limited by their own cognitive level, objective conditions, or expression habits. They often need to gradually release underlying, localized details based on the questioner's guidance and effective follow-up questions. However, existing virtual interaction systems lack a dynamic data disclosure control mechanism based on the interaction depth and the cognitive boundaries of the agent's role. This causes the agent to easily output all underlying structured data or deep-related information beyond its role settings when receiving an initial question. This makes the multi-round human-computer interaction process extremely mechanical and violates the natural progressive communication pattern, severely damaging the realism of the situation simulation system and the continuity of the conversation state. Summary of the Invention
[0005] This invention provides a virtual diagnosis and treatment interaction method and system based on a patient intelligent agent, which solves the defects of existing technologies that lack a dynamic data disclosure control mechanism based on the interaction depth and the cognitive boundary of the intelligent agent role, and are prone to outputting all underlying structured data or deep professional information beyond the role setting at once, so as to achieve a highly realistic and progressive consultation interaction closed loop that conforms to the natural progression law.
[0006] This invention provides a virtual diagnosis and treatment interaction method based on a patient intelligent agent, characterized in that it is applied in a virtual interaction environment pre-configured with patient intelligent agents, wherein the patient intelligent agent includes structured and hierarchically organized patient information cognitive boundary constraints and role characteristics, and the method includes: In the virtual interactive session, multi-source input data is continuously received, and a session-level interactive context is constructed based on historically disclosed information and the multi-source input data; Based on the interaction stage of the session-level interaction context and the preset information disclosure rules, the range of patient information that is currently allowed to be disclosed is determined within the scope of the cognitive boundary constraints. Based on the determined range of patient information currently allowed to be disclosed and the role characteristics of the patient agent, a response result matching the interaction phase is generated; Output the response result, and update the set of disclosed information and the set of information to be disclosed for the current session according to the currently allowed range of patient information to be disclosed.
[0007] According to the method provided by the present invention, the structured hierarchical patient information is pre-configured through the following steps: acquiring target case information and dividing the target case information into at least one of the following layers: basic identity layer, chief complaint and symptom layer, medical history layer, examination material layer, subjective perception layer, cognitive boundary layer and interaction strategy layer; The segmented information at each level is organized into a structured feature template, wherein the structured feature template includes at least one of the following: basic tags, disease topic tags, symptom expression mode tags, cognitive clarity tags, emotion and communication style tags, information disclosure priority tags, attachment material trigger tags, voice style tags, and conversation restrictions and boundaries tags.
[0008] According to the method provided by the present invention, the step of continuously receiving multi-source input data in a virtual interactive session and constructing a session-level interactive context based on historically disclosed information and the multi-source input data includes: The multi-source input data is received, which includes at least one of the following: doctor's voice input, doctor's text question input, material viewing request triggered by the doctor's terminal, and system status events during the session. The multi-source input data is processed in a unified manner to determine the current consultation round, the current doctor's question, and the current discussion topic; The current consultation round, the current doctor's question, the current discussion topic, and the historically disclosed information are jointly processed to construct the session-level interaction context.
[0009] According to the method provided by the present invention, the step of determining the range of patient information currently allowed to be disclosed within the cognitive boundary constraints based on the interaction stage of the session-level interaction context and preset information disclosure rules includes: Determine the semantic type of the current doctor's question and the consultation stage of the current discussion topic; Based on the judgment results and the cognitive boundary constraints, the patient information of the structured hierarchical organization is classified into multiple categories of information, which include at least one of the following: information that can be proactively disclosed, information that needs to be disclosed after further questioning, material-triggered information, information with incomplete patient cognition, and information that is prohibited from being proactively disclosed. According to the preset information disclosure rules, the information that can be proactively disclosed and the information that needs to be disclosed after further inquiry are classified, and the information that can be proactively disclosed and the information that needs to be disclosed after further inquiry are included in the scope of patient information that is currently allowed to be disclosed.
[0010] According to the method provided by the present invention, after classifying the patient information of the structured hierarchical organization into multiple categories based on the judgment result and the cognitive boundary constraints, the method further includes: When the current doctor's question includes information that the patient has incomplete knowledge or information that is prohibited from being disclosed proactively, the cognitive boundary constraint is triggered; Based on the triggered cognitive boundary constraint, the scope of currently allowed patient information disclosure is limited to include uncertain expressions, so as to constrain the generated response result to not exceed the scope of the patient's perception and expression.
[0011] According to the method provided by the present invention, the step of generating a response result matching the interaction phase based on the determined range of currently permitted patient information disclosure and the role characteristics of the patient agent includes: Extract the role features representing expressive style and emotional state from the patient's intelligent agent; Based on the questioning style of the current doctor's question and the role characteristics, and using the currently allowed range of patient information to be disclosed, a patient text response is generated, and the patient text response is used as the response result that matches the interaction stage.
[0012] According to the method provided by the present invention, after the step of taking the patient's text response as the response result matching the interaction phase, the method further includes: Extract the patient's gender, patient age group, and existing voice settings contained in the character characteristics; Based on the patient's gender, age group, existing voice settings, and current emotional state, the patient's text response is converted into corresponding speech output content; Maintain consistency in tone, speed, and expression style of the voice output content within the same virtual interactive session, and output the voice output content synchronously with the patient's text response.
[0013] According to the method provided by the present invention, after the steps of outputting the response result and updating the set of disclosed information and the set of information to be disclosed in the current session according to the currently allowed range of patient information to be disclosed, the method further includes: Monitor whether the preset conditions for triggering attachment feedback are met in the session-level interaction context. The preset conditions include the doctor asking about relevant materials, the doctor making a request to view them, or the interaction stage reaching the point where materials are disclosed. When the preset conditions are met, the examination report, test results or imaging data associated with the response result are extracted from the patient's intelligent agent as target feedback material. The target feedback material and the generated response results are linked and displayed.
[0014] According to the method provided by the present invention, after the steps of outputting the response result and updating the set of disclosed information and the set of information to be disclosed in the current session according to the currently allowed range of patient information to be disclosed, the method further includes: After the virtual interaction session ends, the continuously received multi-source input data is combined into a doctor's question sequence in chronological order, the continuously generated response results are combined into a patient agent response sequence in chronological order, the continuously determined changes in the range of currently allowed patient information disclosure are recorded and combined into an information disclosure trajectory, and the call records of the target feedback material displayed in conjunction with the output are combined into a material trigger trajectory. The combined doctor question sequence, patient agent response sequence, information disclosure trajectory, and material trigger trajectory are archived. Based on the results of the archiving process, the role characteristics in the patient's intelligent agent, the preset information disclosure rules, and the preset conditions for triggering attachment feedback are optimized and updated.
[0015] This invention also provides a virtual diagnosis and treatment interaction system based on patient intelligent agents, applied in a virtual interaction environment pre-configured with patient intelligent agents. The patient intelligent agent includes structured, hierarchically organized patient information cognitive boundary constraints and role characteristics. The system includes: A construction module is used to continuously receive multi-source input data in a virtual interaction session and construct a session-level interaction context based on historically disclosed information and the multi-source input data; The determination module is used to determine the range of patient information that is currently allowed to be disclosed within the cognitive boundary constraints, based on the interaction stage of the session-level interaction context and the preset information disclosure rules. A generation module is used to generate a response result that matches the interaction phase based on the determined range of currently permitted patient information disclosure and the role characteristics of the patient agent. The update module is used to output the response result and update the set of disclosed information and the set of information to be disclosed in the current session according to the range of patient information that is currently allowed to be disclosed.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the virtual diagnosis and treatment interaction method based on a patient intelligent agent as described above.
[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the virtual diagnosis and treatment interaction method based on a patient intelligent agent as described above.
[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the virtual diagnosis and treatment interaction method based on a patient intelligent agent as described above.
[0019] The virtual diagnosis and treatment interaction method and system based on patient intelligent agents provided by this invention provides a multi-dimensional underlying data foundation for accurately simulating real patients by enabling the patient intelligent agent to contain structured and hierarchical patient information, cognitive boundary constraints, and role characteristics. During the interaction process, by continuously receiving multi-source input data in the virtual interaction session and constructing a session-level interaction context based on historically disclosed information and multi-source input data, the system can comprehensively and coherently perceive the current consultation status. Furthermore, based on the interaction stage of the session-level interaction context and preset information disclosure rules, the system determines the range of patient information that can be disclosed within the cognitive boundary constraints, effectively preventing the premature disclosure of deep medical details and professional medical expressions beyond the cognitive ability of ordinary patients from the data flow level. Finally, based on the determined range of patient information that can be disclosed and the role characteristics of the patient intelligent agent, this invention generates a response result that matches the interaction stage, and after outputting the response result, updates the set of disclosed information and the set of information to be disclosed in the current session according to the range of patient information that can be disclosed. This rigorous contextual reasoning and set update mechanism overcomes the drawback of traditional models that mechanically discard all medical records at once, endows the patient agent with interactive characteristics that conform to real clinical logic, and significantly improves the role consistency and situational realism of multi-round virtual consultations. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1 This is one of the flowcharts of the virtual diagnosis and treatment interaction method based on patient intelligent agents provided by the present invention.
[0022] Figure 2 This is the second flowchart of the virtual diagnosis and treatment interaction method based on patient intelligent agents provided by the present invention.
[0023] Figure 3 This is a schematic diagram of the information disclosure and cognitive boundary control process provided by the present invention.
[0024] Figure 4 This is the third flowchart of the virtual diagnosis and treatment interaction method based on patient intelligent agents provided by the present invention.
[0025] Figure 5 This is a schematic diagram of the structure of the virtual diagnosis and treatment interaction system based on patient intelligent agents provided by the present invention.
[0026] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0028] In the field of virtual medical consultation interaction, existing technologies mainly rely on static case scripts, rule engine matching, or general large language models to generate patient responses. While these technologies can meet basic question-and-answer needs, they still have the following four technical shortcomings: First, the patient simulations lack realism.
[0029] Existing solutions mostly use fixed scripts or static question-and-answer databases to drive virtual patients. Their responses rely entirely on preset paths and cannot dynamically adjust their expressions based on the doctor's questioning style, order, or depth of follow-up questions. This mechanical interaction method fails to reflect the differences in expression habits, comprehension abilities, and behavioral characteristics of real patients during consultations due to individual differences (such as age, education level, and personality traits). As a result, the performance of virtual patients is highly templated and differs significantly from real clinical scenarios.
[0030] Second, there is a lack of information disclosure mechanisms that address the actual consultation process.
[0031] In real outpatient settings, patients typically don't provide all their medical information at once. Instead, they gradually supplement details such as symptoms, medical history, and examination results as the doctor guides and asks follow-up questions. However, current solutions generally adopt a "full information output" model, which involves presenting all relevant information from the case at once after the first round of questioning, or directly calling the complete answer during rule matching. This disclosure method violates the natural law of information unfolding layer by layer in clinical consultations, causing multiple rounds of interaction to lose their proper progressive logic.
[0032] Third, there is a lack of effective constraints on the boundaries of patients' cognition.
[0033] Real patients' expressive abilities are limited by their own level of medical knowledge, their perception of symptoms, and the clarity of their memory of past medical history. They can usually only express themselves based on subjective feelings and known information, and will not proactively provide medical diagnoses or professional inferences beyond their cognitive scope. However, existing solutions based on general large language models are prone to using their built-in medical knowledge base to output overly professional and comprehensive answers when generating responses. This causes a serious disconnect between the virtual patient's expression and the real cognitive ability of an ordinary patient, undermining the rationality of the role setting.
[0034] Fourth, it is difficult to support a complete virtual diagnosis and treatment loop.
[0035] A complete outpatient process involves not only verbal communication between doctors and patients, but also the retrieval of examination data, the viewing of imaging results, and the supplementation of past medical history. Most existing technological solutions focus only on the single modality of text-based question-and-answer, lacking coordinated design with aspects such as the return of examination data, voice interaction, and maintaining conversational continuity. For example, when a doctor requests to view lab results, existing solutions cannot automatically output the corresponding report; when voice interaction is required, existing solutions struggle to maintain consistency in the patient's tone, intonation, and expression style throughout the same conversation. This keeps the virtual consultation process at the level of "question-and-answer simulation," failing to replicate the complete closed loop of "consultation—data review—further assessment" in a real outpatient setting.
[0036] To address the aforementioned issues, this invention proposes a virtual medical consultation interaction method and system based on a patient intelligent agent. By constructing a patient intelligent agent containing structured hierarchical information, cognitive boundary constraints, and role characteristics, and by setting mechanisms such as information disclosure control, session state maintenance, and multimodal output linkage, a progressive interaction consistent with real-world consultation logic is achieved. This method can dynamically determine the scope of information that can be disclosed based on the doctor's questions during multiple rounds of consultation, maintain consistency in the patient role, and support voice and material-based feedback linkage, thereby significantly improving the realism and usability of virtual medical consultation.
[0037] Before describing the technical solutions of the embodiments of the present invention, the terms and concepts involved in the embodiments of the present invention will be explained illustratively.
[0038] Patient agent: A digital entity built in a virtual environment to simulate real patient interactions during consultations. This agent contains the patient's identity information, medical condition information, behavioral characteristics, cognitive boundaries, and expression style, and can dynamically generate responses that conform to the patient's role settings based on the doctor's questions in multi-turn dialogues.
[0039] Historically disclosed information: In the current virtual consultation session, this is the sum of all information the patient's agent has provided to the doctor up to the present moment, including symptom descriptions, medical history, and examination results. This information set is used to avoid duplicate or contradictory answers.
[0040] Structured hierarchical information: A data structure that organizes patient information in multiple levels according to attributes, hierarchy, and disclosure priority. It is usually divided into a basic identity layer, chief complaint and symptom layer, medical history layer, examination materials layer, subjective perception layer, cognitive boundary layer, and interaction strategy layer, so as to facilitate on-demand retrieval during interaction.
[0041] Cognitive boundary constraints: rules that limit the output of the patient agent to ensure that its answers are based only on the patient's own perceptions, past experiences and known examination results, and do not output medical diagnostic conclusions or professional inferences that exceed the role settings, thereby maintaining the authenticity of the patient role.
[0042] Role characteristics: These represent the personalized attributes exhibited by the patient's intelligent agent during interaction, including but not limited to age, gender, occupation, education level, personality traits, expression habits, emotional state, and communication style. These characteristics are used to generate response content and voice performance that are consistent with the patient's identity.
[0043] Information Disclosure Rules: A set of pre-defined rules used to control the order and conditions of information release by the patient's intelligent agent during the consultation process. These rules determine which information can be proactively disclosed, which requires follow-up questioning before disclosure, and which must not be disclosed, based on the semantic type of the doctor's questions, the current stage of the consultation, the sensitivity of the information itself, and the patient's cognitive level.
[0044] Patient information scope: The set of information that the patient agent is allowed to output to the doctor in a specific interaction round, as determined by the information disclosure rules. This scope is dynamically updated as the consultation progresses.
[0045] Response results: The output content generated by the patient agent based on the scope of patient information currently allowed to be disclosed and role characteristics can be text replies, voice output, or supplementary information such as examination materials and imaging data that are linked with text / voice feedback.
[0046] Target case information: The original case data used to build the patient agent, which can come from real medical records, simulated case databases or teaching case sets, and includes the patient's complete medical background and related information.
[0047] Multi-source input data: In the virtual consultation session, the system continuously receives various types of input information, including doctor's voice input, doctor's text question input, speech-to-text transcription, material viewing requests triggered by the doctor, system status events during the session, and interactive events returned by the front end (such as playback completion, interruption, and continued questioning).
[0048] Conversation-level interaction context: A comprehensive set of information reflecting the current consultation status, constructed by the system based on historically disclosed information and current multi-source input data. It includes at least the current consultation round, the current doctor's question, the current discussion topic, disclosed patient information, undisclosed but triggerable information, information that cannot be proactively disclosed, output voice content, feedback materials, and the current role's state and emotional state.
[0049] Information that can be proactively disclosed: In the information hierarchy of the patient's intelligent agent, the information types marked as those that can be proactively stated to the doctor without further questioning by the doctor usually include chief complaints, obvious discomfort, and some basic symptom information.
[0050] Information that requires further questioning to be disclosed: In the patient's intelligent agent's information hierarchy structure, the information types that need to be gradually released by the doctor through questioning are marked, such as details of past medical history, triggers, accompanying symptoms, lifestyle habits, and past examination procedures.
[0051] Material-triggered information: This type of information requires the doctor to explicitly ask for it, request it, or for the system to meet specific conditions before it can be fed back. It usually manifests as supplementary materials such as examination report details, imaging results, and previous diagnostic reports.
[0052] Patient-incomplete cognitive information: Information that the patient is unaware of but exists in their medical records, or information that the patient has a vague memory of or cannot accurately describe. For this type of information, the patient's agent can only respond with phrases such as "not quite sure" or "can't remember."
[0053] Proactive disclosure of information is prohibited: background information that has not yet reached the appropriate consultation stage, or that is only used for session control and should not be directly spoken by the patient, such as metadata used only for internal system status maintenance, and treatment clues that have not yet been triggered.
[0054] Doctor Question Sequence: After a virtual consultation session ends, the continuously received doctor input data is combined in chronological order to form a sequence, which is used to record the doctor's consultation path and questioning methods.
[0055] Patient agent response sequence: After a virtual consultation session ends, the continuously generated patient agent response results are combined in chronological order to form a sequence, which is used to record the content and expression of the patient's answers.
[0056] Information Disclosure Tracking: After a virtual consultation session ends, a record is created by combining the changes in the range of patient information that is currently allowed to be disclosed in chronological order. This record is used to track the order and pace of information release.
[0057] Material Trigger Trajectory: After a virtual consultation session ends, the call records of the target feedback materials displayed in conjunction with the session are combined in chronological order to form a record, which is used to track the timing and type of material feedback.
[0058] The virtual diagnosis and treatment interaction method based on patient intelligent agents provided in this invention is executed by a virtual diagnosis and treatment interaction system based on patient intelligent agents (hereinafter referred to as "the system"). This system can be physically implemented as an independent server, a service platform deployed in the cloud, an application integrated into a terminal device, or a distributed system composed of multiple computing nodes. The specific form can be flexibly configured according to the application scenario.
[0059] From a functional perspective, the system includes at least the following core modules: an input receiving module, a context building module, an information disclosure control module, a response generation module, and an output and update module. These modules collaborate through data interaction to complete the entire processing flow from receiving multi-source inputs, building the session context, determining the scope of information disclosure, generating response results, to updating the session state. Optionally, the system may also include a voice performance module, a material feedback module, and a continuous optimization module to support multimodal output and iterative optimization of the agent.
[0060] Those skilled in the art should understand that the specific implementation form of the above-mentioned execution subject does not affect the essence of the technical solution of the present invention. As long as the method steps described in the present invention can be executed, it should be regarded as falling within the protection scope of the present invention.
[0061] The virtual diagnosis and treatment interaction method based on patient intelligent agents provided by the present invention will be described in detail below with reference to specific embodiments. This embodiment takes a virtual respiratory medicine consultation scenario as an example, which is intended to exemplify the implementation process of the steps described in the independent claims, but does not constitute a limitation on the scope of protection of the present invention.
[0062] Before executing the method described in this embodiment, a patient agent needs to be configured in the virtual interaction environment. This patient agent is constructed based on the target case information, and internally organizes patient information in a structured hierarchical manner, with cognitive boundary constraints and role characteristics. The structured hierarchical organization refers to dividing various types of patient information according to attributes, levels, and disclosure priorities in multiple dimensions, so that it can be accessed as needed during the interaction. Cognitive boundary constraints limit the scope of the patient agent's responses, ensuring that it only outputs content that conforms to the patient's own perception and cognitive level. Role characteristics represent the patient's personalized expression habits, emotional state, and communication style. In this embodiment, the constructed patient agent corresponds to a 35-year-old male whose chief complaint is "cough with fever for three days," exhibiting a normal communication style, insensitivity to medical terminology, and only able to express himself based on his own perception and known information.
[0063] Before describing the virtual diagnosis and treatment interaction process, this paper first provides a detailed explanation of the construction and configuration process of the patient's intelligent agent in this invention, particularly the structured hierarchical information organization method and its feature template processing. This construction process is the foundation for all subsequent interaction steps, ensuring that the patient's intelligent agent can progressively disclose information according to preset rules during the consultation process.
[0064] In this embodiment, the system first acquires the target case information. This case information can come from real medical records, a simulated case database, or a teaching case set, and its content covers the patient's complete medical background and related information. Taking a virtual respiratory medicine consultation scenario as an example, the acquired target case information corresponds to a 35-year-old male patient whose chief complaint is "cough with fever for three days," accompanied by symptoms such as coughing up yellow sputum and worsening at night. He has a history of good health and no chronic diseases, and a recent chest X-ray shows increased lung markings.
[0065] After acquiring the target case information, the system does not store the case text as a whole, but performs structured hierarchical processing, dividing the information into multiple levels with clear semantic boundaries. In this embodiment, the system divides the case information into the following seven levels: The basic identity layer includes the patient's age (35 years old), gender (male), occupation (company employee), marital status (married), and lifestyle habits (smoking history of 10 years, half a pack of cigarettes per day). Chief complaint and symptom layer, including the patient's main discomfort (cough, fever), duration of symptoms (three days), accompanying symptoms (coughing up yellow sputum, worsening at night), aggravating or relieving factors (worsening after activity, slightly relieved after rest), etc. The medical history includes present illness (the course of the current illness), past medical history (previously healthy, denies hypertension or diabetes), surgical history (none), allergy history (denies drug allergies), family history (father has hypertension), and personal history (smoking), etc. The examination materials include laboratory reports (blood routine showed elevated white blood cell count), imaging data (chest X-ray), test reports (mycoplasma pneumoniae antibody negative), previous medical records (none), and other supporting information; The subjective perception layer includes subjective information such as the patient's self-assessment of the pain level (chest pain when coughing, moderate level), description of symptom sensation (feeling hot, but not taking body temperature), level of anxiety (somewhat worried), and vagueness of expression ("maybe a little short of breath, but not quite sure"). The cognitive boundary layer is used to mark information that the patient is aware of (such as their own symptoms and smoking history), information that the patient is uncertain about (such as whether they actually have a fever and the specific temperature value), information that the patient is unaware of but that exists in the records (such as the specific values of blood routine tests and X-ray conclusions), and information that can only be triggered by the doctor's follow-up questions (such as the specific number of years of smoking and previous physical examination results). The interaction strategy layer includes patient preferences such as whether they tend to express themselves proactively (yes, proactive in their chief complaint), whether they are prone to giving irrelevant answers (no), whether they need further questioning (past medical history needs further questioning), and whether they prefer brief answers (the chief complaint is answered in more detail, and further details need further questioning).
[0066] After completing the information layering, the system further organizes the information at each layer into a structured feature template. This feature template is a standardized data format used to map information scattered across different layers into tagged attributes that can be dynamically invoked by the patient's intelligent agent during interaction. In this embodiment, the formed structured feature template includes at least the following tags: Basic tags: {Age: 35, Gender: Male, Occupation: Employee, Smoking history: 10 years / half a pack}; The patient's condition is tagged with: {Chief complaint: cough and fever, duration of illness: 3 days, accompanying symptoms: yellow sputum / worsening at night}; Symptom expression style tags: {Expression style: conversational, level of detail: detailed chief complaint / details need to be followed up}; Cognitive clarity tags: {Known: Personal symptoms / smoking history, Uncertain: Body temperature value, Unknown: Blood routine value / X-ray conclusion}; Emotional and Communication Style Tags: {Emotion: Mild Anxiety, Communication Style: Cooperating but Needs Guidance}; Information disclosure priority tags: {Proactive disclosure: chief complaint / basic symptoms, disclosure after follow-up questioning: past medical history / smoking details, material trigger: examination report}; Attachment materials trigger tags: {Complete blood test: requires explicit inquiry from doctor; X-ray: requires doctor's request to view}; Voice style tags: {Tone: Adult male, Speech rate: Medium, Tone: Steady with slight worry}; Session restrictions and boundary labels: {Prohibit proactive diagnosis: Yes, prohibit proactive provision of examination values: Yes, prohibit exceeding cognitive inference: Yes}.
[0067] Through the aforementioned structured layering and feature template processing, the patient agent can access information at the appropriate level based on the specific context during subsequent interactions and express it according to preset tag attributes. For example, in the chief complaint collection stage, the system primarily accesses chief complaint and symptom layer information, and generates a conversational description by combining symptom expression tags; in the past medical history inquiry stage, the system accesses medical history layer information, and, constrained by cognitive boundaries, only discloses content known to the patient; in the document retrieval stage, the system determines whether to output specific examination reports based on the tags triggered by the attached materials. This organizational approach makes the patient agent's information retrieval more precise and controllable, laying a data foundation for achieving progressive information disclosure.
[0068] Figure 1 This is a flowchart illustrating the virtual diagnosis and treatment interaction method based on a patient intelligent agent provided by the present invention. The method includes: Step 101: Continuously receive multi-source input data in the virtual interaction session, and construct a session-level interaction context based on historically disclosed information and the multi-source input data.
[0069] After the virtual consultation session is initiated, the system enters a continuous listening state, receiving multi-source input data from the doctor in real time. In this embodiment, the multi-source input data is specifically represented by the doctor's voice input of "Where do you feel unwell?" The system receives this voice data and converts it into text for subsequent processing.
[0070] Meanwhile, the system reads historically disclosed information from the storage space of the current session. At the initial stage of the session, since no information exchange has occurred, the set of historically disclosed information is empty. The system combines the currently received doctor's question text "Where do you feel unwell?" with this empty set of historically disclosed information to construct the session-level interaction context for the current round.
[0071] The constructed session-level interaction context is a comprehensive data structure, which includes at least the following elements: the current consultation round (round 1), the current doctor's question ("Where do you feel unwell?"), the current discussion topic (collected by the system based on semantic recognition to determine the chief complaint), the set of disclosed information (empty), and the current role state of the patient's agent (35-year-old male, normal communication style, stable emotions). This context will serve as the unified decision-making basis for information disclosure determination and response generation in subsequent steps.
[0072] It's important to note that in subsequent rounds of interaction, the conversational interaction context will be continuously updated as the doctor asks questions and the patient answers. For example, in the second round of interaction, the system will receive new questions from the doctor and simultaneously read updated historical disclosed information (such as the "cough and fever" information already included in the first round's response) to construct a context containing more information, thereby ensuring the continuity and consistency of the multi-round dialogue.
[0073] Step 102: Determine the range of patient information that is currently allowed to be disclosed within the cognitive boundary constraints, based on the interaction stage of the session-level interaction context and the preset information disclosure rules.
[0074] After constructing the session-level interaction context, the system first analyzes the context to identify the current stage of the interaction. The determination of the interaction stage mainly includes the semantic type of the doctor's question, the current discussion topic, and the information already disclosed. In this embodiment, since the doctor's question "Where do you feel unwell?" is a chief complaint inquiry question, and the currently disclosed information is empty, the system determines that the current stage is "chief complaint collection stage".
[0075] After the interaction phase is completed, the system invokes preset information disclosure rules. These rules are a set of conditions and actions, designed based on the general process of medical consultation, the sensitivity of the information itself, and the patient's communication habits. In this embodiment, the information disclosure rule corresponding to the chief complaint collection phase is: for questions related to the chief complaint, the patient is allowed to disclose basic identity information and chief complaint symptom information, but deeper information such as past medical history and examination results is not disclosed at this time.
[0076] At the same time, the system must strictly limit the information disclosure results to the scope of cognitive boundary constraints. The core of cognitive boundary constraints is that the patient agent's responses should only be based on information that the patient can perceive and confirm, and should not output content beyond the scope of its role settings. In this embodiment, the 35-year-old male patient can perceive that he has symptoms of cough and fever, but cannot confirm the medical diagnosis corresponding to these symptoms. Therefore, when filtering information, the system only extracts relevant information from the chief complaint and symptom layer, and does not extract any diagnostic content from the medical history layer or examination material layer.
[0077] Based on a comprehensive assessment of the aforementioned interaction stages, information disclosure rules, and cognitive boundary constraints, the system filters out the currently permissible range of information to be disclosed from the patient's structured, hierarchical information. This range specifically includes: the patient's chief complaint of "cough with fever for three days," and related symptom details (such as basic descriptions of the nature of the cough and the degree of fever), but excludes past medical history, examination results, and other information requiring further inquiry. This range will be explicitly recorded and used as direct material for response generation in subsequent steps.
[0078] Step 103: Based on the determined range of patient information currently allowed to be disclosed and the role characteristics of the patient agent, generate a response result that matches the interaction phase.
[0079] After obtaining the scope of patient information currently permitted for disclosure, the system further extracts the role characteristics of the patient's intelligent agent. Role characteristics are the basis for the patient's personalized expression, and in this embodiment, they are specifically manifested as: 35-year-old male, ordinary communication style, tendency to use simple and straightforward everyday language, and stable emotional state.
[0080] The system combines the currently permitted range of information to be disclosed ("cough with fever for three days" and related symptom details) with role characteristics to generate response content. The generation process is not simply piecing together information items into sentences, but rather organizing and stylizing the language according to role characteristics. Specifically, the system transforms structured information items into natural language that conforms to first-person expression habits, and incorporates tone and wording that match the role characteristics.
[0081] In this embodiment, the system-generated patient text response is: "I've been coughing for the past few days, and I also have a slight fever." This response accurately reflects the scope of information currently allowed to be disclosed (chief complaint symptoms), maintains consistency with the patient's role characteristics in terms of expression (concise, conversational, and in line with the daily expression habits of a 35-year-old male), and matches the interaction requirements of the current chief complaint collection stage (the answer directly addresses the core of the doctor's question). The generated text response will serve as the main content of the response result and proceed to the next step of the output process.
[0082] Step 104: Output the response result and update the set of disclosed information and the set of information to be disclosed for the current session according to the current allowed range of patient information to be disclosed.
[0083] The system outputs the response "I've been coughing and have a slight fever for the past few days" to the doctor's end. Depending on the system configuration and the interaction scenario, the output format can be flexibly selected: in a pure text interaction scenario, the text is directly displayed on the doctor's interface; in a voice interaction scenario, the system further calls the speech synthesis module to convert the text into speech content consistent with the patient's role characteristics (such as a 35-year-old male voice and a steady speaking speed) before broadcasting it to the doctor.
[0084] Synchronizing with the output response, the system performs critical status update operations. Specifically, the system updates the set of disclosed information and the set of information to be disclosed for the current session based on the range of patient information that can be disclosed at present. In this embodiment, "Chief complaint: cough with fever for three days" and its related symptom details are added to the set of disclosed information, while this information is removed from the set of information not disclosed (if it previously existed in the set of information to be disclosed). In addition, the system will also adjust the content of the set of information to be disclosed based on the current progress of the consultation. For example, information such as past medical history and examination results may remain in the state of being to be disclosed, marked as "disclosure after follow-up questioning" or "disclosure after material triggering," so that it can be gradually released according to the rules in subsequent questions.
[0085] After the above state update operation is completed, this round of interaction ends, and the system returns to the listening state, ready to receive the next round of multi-source input. As the doctor continues to ask questions (such as "Is your cough dry or with phlegm?" "Have you had similar symptoms before?"), the system will repeat steps 101 to 104, forming multiple rounds of continuous virtual diagnosis and treatment interaction. During this process, the conversational interaction context evolves continuously with each round of input and output, the set of disclosed information is gradually enriched, and the patient agent always generates progressive responses that conform to the role characteristics under the guidance of cognitive boundary constraints and information disclosure rules, thereby systematically simulating the clinical consultation behavior of a real patient who "answers questions and gradually supplements information."
[0086] It should be noted that although the above embodiments use specific consultation content as an example, the operations described in each step are universal: step 101, "receiving multi-source input and constructing context," is applicable to any consultation round and any question type; step 102, "determining the information scope based on the interaction stage and disclosure rules," is applicable to access control of information at any level in the patient's intelligent agent; step 103, "generating a response based on the information scope and role characteristics," is applicable to any linguistic expression; and step 104, "outputting the response and updating the state," is applicable to post-processing of all interaction results. Those skilled in the art should understand that the technical concept embodied in the above steps—namely, achieving progressive disclosure of patient information and maintaining role consistency through context awareness, rule constraints, boundary limitations, and state maintenance—can be applied to any type of case and any consultation scenario, and is not limited to the specific content listed in this embodiment.
[0087] The following detailed explanation of step 101, using a virtual respiratory medicine consultation scenario, further elaborates on how the system receives multi-source input data, processes it uniformly, and constructs a session-level interactive context. See also... Figure 2 Step 101 includes: Step 201: Receive multi-source input data.
[0088] During the ongoing virtual consultation session, the system monitors and receives various types of input data from the doctor in real time. In this embodiment, the types of multi-source input data include, but are not limited to, the following: Doctor's voice input: The doctor asks questions in natural language through a microphone. For example, in the first round of consultation, the doctor says "Where do you feel uncomfortable?" The system receives the voice data, performs speech recognition, and converts it into text. Doctor text question input: If the system supports text interaction mode, doctors can directly type their questions in the input box. For example, in a later round, the doctor can type "Is the cough dry or with phlegm?" Doctor-triggered document viewing requests: When a doctor wants to view a patient's relevant examination data, they can click on buttons such as "View Blood Routine Report" or "View Chest X-ray" on the interface. The system receives this request as a special type of input. System status events during the session: These include interactive events returned from the front end, such as the completion of the previous round of voice playback, interruption by the doctor, and session timeout. These events are also received by the system as input data and used to adjust the interaction flow.
[0089] In the first round of interaction in this embodiment, the system receives multi-source input data, which is the doctor's voice input "Where do you feel unwell?", and obtains the corresponding text content after speech recognition.
[0090] Step 202: Perform unified processing on the multi-source input data to determine the current consultation round, the current doctor's question, and the current discussion topic.
[0091] The system performs unified parsing and processing of the received multi-source input data to extract key information in order to clarify the basic elements of the current interaction.
[0092] First, the system determines the current consultation round based on the session history. Since this round is the first interaction after the session started, the system marks the current round as "Round 1". In subsequent interactions, the system will automatically increment the round number based on the recorded interaction history.
[0093] Secondly, the system extracts the current doctor's question from the input data. For voice input, the system has already completed the voice-to-text conversion; for text input, it directly extracts the text string; for requests to view materials, the system parses them into standardized question statements, such as converting "view blood routine report" into "doctor requests to view blood routine report" as the question content to be recorded; for system status events, it records the event type and timestamp.
[0094] Secondly, the system uses semantic analysis to determine the current discussion topic. Taking the first round of input, "Where do you feel unwell?", as an example, the system identifies this question as a chief complaint inquiry, and therefore marks the current discussion topic as "Chief Complaint Collection". If the doctor subsequently asks, "Have you had similar symptoms before?", the system identifies this as a medical history inquiry, and marks the discussion topic as "Medical History Collection". Accurate identification of the discussion topic is crucial for the subsequent invocation of information disclosure rules.
[0095] In the first round of interaction in this embodiment, the system determines the current consultation round as 1 after unified processing, the current doctor's question is "Where do you feel unwell?", and the current discussion topic is "Chief Complaint Collection".
[0096] Step 203: Combine the determined current consultation round, current doctor's question, current discussion topic, and historically disclosed information to construct a conversation-level interaction context.
[0097] The system combines the information determined in step 202 with the historically disclosed information already present in the current session to construct a comprehensive session-level interaction context. This context is a structured data object used to comprehensively describe the current consultation status.
[0098] Historical disclosed information refers to all information that the patient agent has output to the doctor in the current session up to the current round of interaction. In the first round of interaction, since no patient response has been generated, the historical disclosed information is an empty set. However, in subsequent rounds, historical disclosed information will gradually accumulate; for example, the response in the first round, "I've been coughing for the past few days and have a slight fever," will be recorded in the historical disclosed information.
[0099] The system integrates the current consultation round, the current doctor's question, the current discussion topic, and previously disclosed information, and supplements this with the patient's agent's current role state (such as emotional state and expression style), ultimately forming a complete conversational interaction context. This context must contain at least the following fields: Session ID: A unique ID for the current session; Round number: 1; The doctor asked: "Where do you feel unwell?" Discussion topic: Chief complaint collection; Historically disclosed information: empty; Patient role status: {Emotion: Stable, Expression style: Concise}; Timestamp: Current system time.
[0100] The constructed session-level interaction context is stored in the session cache, serving as the direct basis for subsequent information disclosure scope determination steps. As the consultation progresses, each round of interaction reconstructs the context based on new inputs and updated historically disclosed information, ensuring that the system always makes decisions in the latest session state.
[0101] Through the detailed operations described in steps 201 to 203, the system can accurately capture the input information of each round of interaction and organically integrate it with historical states, laying a solid contextual foundation for subsequent progressive information disclosure and role-consistent response generation. Those skilled in the art should understand that this processing flow is not only applicable to the chief complaint collection stage, but also to any consultation stage and any type of multi-source input, possessing versatility and scalability.
[0102] Further, see Figure 3 Step 102 includes: S31. Determine the semantic type of the current doctor's question and the consultation stage of the current discussion topic.
[0103] The system first extracts the current doctor's question and the current discussion topic from the conversational interaction context. Taking the first round of interaction as an example, the doctor asks "Where do you feel unwell?" The system uses natural language understanding technology to determine that the semantic type of the question is "chief complaint inquiry." At the same time, combined with the current discussion topic "chief complaint collection" and the fact that the historical disclosed information is empty, the system determines that the current consultation stage is the "chief complaint collection stage."
[0104] In subsequent rounds, if the doctor asks, "Have you had similar experiences before?", the system identifies the semantic type as "past medical history inquiry" and, combined with the current discussion topic, shifts it to "past medical history collection," determining that the current stage is "past medical history collection." If the doctor asks, "Can I see your test results?", the system identifies the semantic type as "materials review request," determining that the current stage is "test materials retrieval."
[0105] S32. Based on the judgment result and the cognitive boundary constraints, classify the patient information of the structured hierarchical organization into multiple types of information, including at least one of the following: information that can be proactively disclosed, information that needs to be disclosed after further questioning, material-triggered information, information with incomplete patient cognition, and information that is prohibited from being proactively disclosed.
[0106] Based on the above judgment results, the system further combines the cognitive boundary constraints of the patient's intelligent agent to extract information relevant to the current question from the structured hierarchical information and classify it into five preset information types. Taking the first round of chief complaint collection as an example: Information that can be proactively disclosed: Chief complaint of "cough with fever for three days" and its basic description. This information is marked as proactively disclosed in the feature template and is within the scope of the patient's knowledge.
[0107] Information requiring further inquiry before disclosure includes: smoking history ("smoking for 10 years, half a pack per day"), past medical history details ("previously healthy"), etc. Although this information exists in the medical record, it can only be disclosed after the doctor inquires further, according to the interaction strategy layer settings.
[0108] Material-triggered information: Examination materials such as blood routine reports and chest X-rays. This information can only be generated after a doctor makes an explicit request or the system meets specific conditions.
[0109] Patient's incomplete knowledge of information: specific body temperature value (the patient only "felt feverish, but did not take their temperature"), X-ray findings (unknown to the patient). This information is unknown to the patient.
[0110] Unauthorized disclosure of information is prohibited: Medical diagnostic conclusions (such as "possibly pneumonia"). This information exceeds the patient's role and should not be proactively disclosed.
[0111] In the third round of medical history collection, the system categorized doctors' questions about "have you had similar experiences before" in a similar way: "previous health" was classified as information that could be proactively disclosed (as disclosure was allowed at this stage), details such as "smoking years" and "daily smoking amount" were classified as information that needed to be disclosed after further inquiry, and "family history" was classified as information that needed to be disclosed after further inquiry, etc.
[0112] S33. According to the preset information disclosure rules, determine the information that can be proactively disclosed and the information that needs to be disclosed after further inquiry after classification, and include the information that can be proactively disclosed and the information that needs to be disclosed after further inquiry that is determined to be allowed to be disclosed into the scope of patient information that is currently allowed to be disclosed.
[0113] The system invokes preset information disclosure rules to further determine which information can be proactively disclosed and which requires follow-up inquiry before disclosure. The information disclosure rules determine which information can be actually disclosed in the current round based on factors such as the consultation stage, semantic type, and information sensitivity.
[0114] Taking the first round of chief complaint collection as an example, the information disclosure rules stipulate that for questions with the semantic type of "chief complaint inquiry," all content related to the chief complaint marked as "information that can be proactively disclosed" is allowed to be disclosed. Therefore, the system includes information items such as "cough" and "fever" in the scope of patient information that can be disclosed at present. For information that needs to be disclosed after follow-up questioning (such as smoking history), since the current stage does not meet the conditions for follow-up questioning, the rules determine that it will not be disclosed for the time being and will be retained in the set of information to be disclosed.
[0115] Taking the third round of medical history collection as an example, the information disclosure rules stipulate that for questions with the semantic type of "medical history inquiry," content related to medical history marked as "information requiring follow-up inquiry" is allowed to be disclosed. Therefore, the system includes information such as "previously healthy" and "smoking for 10 years" within the scope of currently allowed patient information disclosure.
[0116] Based on the above assessment, the system ultimately determines the scope of patient information that can be disclosed in the current round. For example, the scope of disclosure allowed in the first round is: {Chief complaint: "Cough", "Fever", Symptom description: "Cough with fever for three days"}; the scope of disclosure allowed in the third round is: {Patient history: "Previously healthy", Smoking history: "Smoked for 10 years, half a pack per day"}. This scope will be used to generate patient responses.
[0117] It should be noted that for material-triggered information, its disclosure depends not only on the semantic type and consultation stage, but also on meeting the material triggering conditions (such as an explicit request from the doctor). This type of information is usually not directly included in the permitted disclosure scope of text responses, but is processed separately by the system.
[0118] Furthermore, following the information classification process described in step 102, this embodiment sets up a cognitive boundary constraint triggering mechanism to handle special situations when a doctor asks questions involving incomplete cognitive information of the patient or prohibits the proactive disclosure of information.
[0119] Specifically, in one round of interaction, the doctor asks, "How high is your fever?" After semantic analysis, the system determines that the question refers to a specific body temperature value. However, according to the patient agent's cognitive boundary layer, the patient only "feels feverish but has not measured their temperature." Therefore, the "specific body temperature value" is information that the patient cannot confirm, i.e., "incomplete patient cognitive information." Upon detecting that the current question falls into this category, the system immediately triggers cognitive boundary constraints. Under these constraints, the system no longer attempts to extract the body temperature value from the medical records (even if the records might show the temperature measured at the time of the visit, this value is beyond the patient's cognitive range at that time). Instead, it limits the range of patient information that can be disclosed to an uncertain expression template. Subsequently, the response generation module generates a response that matches the patient's cognitive level based on this range, such as "I'm not sure what the exact temperature is, I just feel feverish." Through this mechanism, the patient agent avoids fabricating values or outputting information beyond their perceptual range, thus maintaining the authenticity of the role.
[0120] In another round of interaction, the doctor directly asked, "Is your illness pneumonia?" in the early stages of the consultation. The system identified the semantic type of this question as a "diagnostic inference inquiry." However, according to the cognitive boundary settings of the patient's agent, the patient does not possess medical diagnostic capabilities, and proactively providing a diagnostic conclusion falls under the category of "prohibited information disclosure." Upon detecting that the current question involves prohibited information disclosure, the system also triggers cognitive boundary constraints, limiting the range of patient information that can be disclosed to uncertain expression templates. The response generation module then generates a reply accordingly, such as "I don't know what's wrong, so I came to see a doctor" or "I'm not sure, could you please take a look?" instead of outputting content beyond the patient's role settings, such as "I might have pneumonia."
[0121] Through the aforementioned triggering mechanism, the system ensures that the patient agent can respond with reasonable uncertainty when faced with questions that exceed its cognitive range or role settings. This approach not only avoids compromising the realism of virtual diagnosis and treatment due to inappropriate information disclosure, but also ensures that the patient agent maintains consistent expression boundaries with real patients throughout the consultation process, further enhancing the naturalness and credibility of the interaction.
[0122] In this embodiment, after the system completes the information disclosure scope determination in step 102, it enters the response result generation stage. The core of this stage is to combine the currently allowed scope of patient information disclosure with the personalized role characteristics of the patient's intelligent agent to generate a natural language response that conforms to the patient's identity and the current interaction context.
[0123] (1) Extract the role features representing the expression style and emotional state in the patient's agent.
[0124] The system first extracts role features from the patient agent to guide language expression. These features are pre-configured during the patient agent construction phase and determine "how" the patient speaks, rather than "what" they say. In this embodiment, the constructed patient agent has the following role features: an expression style that is "concise, straightforward, and conversational," tending to use everyday language rather than professional terminology; an emotional state that is "mildly anxious but generally stable," naturally revealing a certain degree of concern in its expression; and a communication style that is "cooperative but requires guidance," gradually supplementing information in response to the doctor's follow-up questions, but not proactively elaborating beyond the scope of the question. These features will be used in the subsequent language generation process to ensure that the generated responses are consistent with the patient's personalized settings in terms of expression.
[0125] (2) Based on the questioning method of the current doctor’s question and the role characteristics, and using the range of patient information that is currently allowed to be disclosed, generate a patient text response and use the patient text response as the response result that matches the interaction stage.
[0126] After extracting role characteristics, the system further analyzes the questioning style of the current doctor's questions and, in conjunction with the currently permitted range of patient information, generates a patient text response. The generation process follows these principles: the content is strictly limited to the permitted range of information, the expression strictly matches the role characteristics, and the style is adapted to the question type.
[0127] Taking the first round of interaction as an example, the doctor's current question is "Where do you feel unwell?", which is an open-ended question. The currently allowed patient information disclosure is limited to the primary complaint level: "cough with fever for three days." The system transforms the structured information items into natural language expressions that conform to first-person narration, while incorporating the patient's "concise and straightforward" expression style and "mild anxiety" emotional characteristics, generating the patient's text response: "I've been coughing for the past few days, and I also have a slight fever." This response accurately reflects the currently allowed information disclosure range, is highly consistent with the patient's role characteristics in terms of expression, and matches the interaction requirements of the current primary complaint collection phase.
[0128] In subsequent rounds, the system continues to follow these generation principles even when doctors ask different questions. For example, for closed-ended questions about past medical history, the system combines the patient's "insensitivity to medical terminology" with cognitive boundary constraints to generate natural language responses that include uncertain expressions, ensuring that the responses comply with both information disclosure rules and the patient's expression habits.
[0129] Through the above mechanism, the system-generated text response not only ensures the accuracy of the information content but also guarantees the matching of the expression style with the characteristics of the patient's role, thus achieving the technical effect of "answering correctly" and "answering in a manner consistent with the patient." This text response then serves as a response result matching the current interaction stage and enters the subsequent output stage.
[0130] In scenarios configured for voice interaction, the system generates a text response from the patient and then performs voice output processing to support multimodal interaction.
[0131] (1) Extract voice-related role features.
[0132] The system first extracts attributes related to speech performance from the patient agent's role characteristics. These attributes are pre-configured during the patient agent construction phase and are used to determine the personalized features of the speech output. Specifically, the features extracted by the system include: patient gender (e.g., male or female), patient age group (e.g., teenager, middle-aged, elderly), and existing voice settings (e.g., adult male voice, adult female voice, child's voice, etc.). Simultaneously, the system acquires the current emotional state (e.g., calm, anxious, relaxed, etc.) as the emotional tone for speech synthesis. These features together constitute a personalized parameter set for the speech output, ensuring that the generated speech matches the patient's identity.
[0133] (2) Convert the text reply into voice output content.
[0134] Based on the extracted speech-related features, the system invokes the speech synthesis module to convert the patient's text response into corresponding speech output. The speech synthesis module selects an appropriate voice library based on the patient's gender and age group, sets parameters such as speech rate, pitch, and volume according to existing voice settings, and incorporates appropriate emotional nuances into the speech based on the patient's current emotional state. For example, if the patient is mildly anxious, the synthesized speech will naturally convey a sense of worry; if the patient is calm, a neutral tone will be used. Through this conversion process, the original text response is transformed into a personalized speech signal, enabling the patient's AI agent to participate in virtual diagnostic interactions by "speaking."
[0135] (3) Maintain consistency in speech output within the conversation.
[0136] Throughout the entire virtual interactive session, the system maintains consistency in its voice output. This mechanism ensures that the patient's agent behaves as the same person throughout the consultation, preventing abrupt changes in timbre, speech rate, or tone from disrupting the continuity and realism of the interaction. Specifically, the system initializes a set of voice parameters at the start of the session and continues to use the same parameters for speech synthesis in all subsequent rounds. Regardless of the round of the consultation, as long as the patient's agent does not change roles, the doctor always hears the same "voice" answering questions, with consistent speech rate and tone.
[0137] Ultimately, the system synchronously outputs the generated voice content and the original text response to the doctor's end: the text content is displayed on the interactive interface, while the voice content is broadcast through the speaker. This synchronized output method, combining text, image, and audio, not only meets the needs of different interaction scenarios but also further enhances the stability of the patient's intelligent agent role and the immersive experience of the interaction. Through this mechanism, the present invention achieves consistency in patient performance in voice-based virtual diagnosis and treatment scenarios, making the virtual patient not only realistic and credible in content but also highly similar to a real patient in auditory presentation.
[0138] After outputting the response results and updating the session state, this embodiment further sets up a material linkage feedback mechanism to support the on-demand return of supplementary information such as examination data and test results during virtual diagnosis and treatment.
[0139] After each round of interaction, the system continuously monitors whether the preset conditions for triggering attachment feedback are met in the current session-level interaction context. These preset conditions are pre-configured based on clinical consultation logic and patient information disclosure rules, and mainly include the following scenarios: First, the doctor explicitly inquires about relevant materials in their question, such as "Can I see your blood routine report?" or "Please show me your chest X-ray"; second, the doctor triggers a material viewing request through interface operations, such as clicking the "View Examination Report" button; third, the current interaction stage naturally reaches a reasonable time for material disclosure, such as after completing the chief complaint collection and follow-up medical history questions, the system determines according to preset rules that it can proactively prompt the patient to provide relevant examination data. The system uses natural language understanding technology to perform semantic analysis on the doctor's questions, while simultaneously monitoring interface events and changes in the session stage to determine in real time whether any of the above conditions are met.
[0140] When the triggering conditions are met, the system immediately extracts the target feedback material associated with the current response result from the patient's structured hierarchical information. This material is stored in the examination material layer, and specific types include examination reports (such as complete blood count and urinalysis), laboratory results (such as biochemical indicators and pathogen detection), imaging data (such as X-rays and CT images), previous medical records, and other medical records. The system determines the scope of materials requiring feedback based on the current interaction context: if the doctor explicitly specifies the material type, the corresponding specific material is extracted; if the doctor does not specify but the interaction stage reaches the material disclosure time, the material most relevant to the current discussion topic is extracted. In this embodiment, when the doctor asks, "Can I see your complete blood count report?", the system extracts the pre-stored complete blood count report file and its related descriptive information from the patient's examination material layer.
[0141] The system will link the extracted target feedback materials with the response results generated in step 103 for synchronized display and output. The linked output method can be flexibly selected based on the system configuration: one method is synchronous output, where the patient's text response and the materials are presented simultaneously. For example, on the doctor's interface, a thumbnail or file link to the blood routine report is simultaneously displayed below the patient's reply, "This is a test I had done at another hospital a couple of days ago." The other method is decoupled output, where the materials are returned to the front end as an independent event, and the front end determines the timing and method of display based on the interaction logic. Regardless of the output method used, the system ensures that the material feedback and the text / voice response are consistent in content, together forming a complete patient-side response.
[0142] Through the aforementioned material linkage and feedback mechanism, the virtual patient in this invention is no longer merely a "verbal respondent," but can gradually provide supporting data during the consultation process, just like a real patient. This mechanism significantly enhances the completeness of the virtual diagnosis and treatment process, enabling doctors not only to obtain information about the patient's condition through questioning, but also to access examination data corresponding to the patient's medical record when needed, thereby more realistically recreating the complete closed loop of "consultation—data review—further judgment" in clinical consultation.
[0143] After the virtual interaction session ends, this embodiment further sets up a session archiving and continuous optimization mechanism to achieve iterative evolution of the patient's intelligent agent and continuous improvement of its interactive performance.
[0144] When a complete virtual consultation session ends, the system systematically organizes and combines the various data generated during the session. Specifically, the system combines the continuously received multi-source input data in chronological order into a doctor's question sequence. This sequence fully records the doctor's interactive behaviors, such as the content and manner of asking questions and requests for reviewing materials, in each round of the consultation. Simultaneously, the system combines the continuously generated response results in chronological order into a patient agent response sequence. This sequence records the text or voice responses output by the patient agent to each round of doctor's questions.
[0145] Furthermore, the system will continuously determine the changes in the currently permissible range of patient information to be disclosed, chronologically combining these changes into an information disclosure trajectory. This trajectory reflects how patient information is gradually released as the doctor asks questions throughout the consultation process, including which information was disclosed in each round, which information remains in the pending disclosure set, and whether the order of information disclosure conforms to preset rules. For rounds that trigger material feedback during the consultation, the system will also link and display the call records of the output target feedback materials, chronologically combining them into a material trigger trajectory, recording the triggering time, material type, and corresponding doctor's questions for each material feedback.
[0146] After completing the above combination, the system performs unified archiving of the doctor's question sequence, the patient's agent response sequence, the information disclosure trajectory, and the material trigger trajectory. The archived content is stored in the system database in the form of structured data, serving as the basis for subsequent analysis and optimization.
[0147] The system performs in-depth analysis of the archived session data and optimizes and updates the patient agent based on the analysis results. The updates include at least the following aspects: First, the system fine-tunes the role characteristics of the patient's intelligent agent. By analyzing the patient's performance in actual interactions, the system can identify potential deviations in the role characteristic configuration. For example, if archived data shows that the patient's expression style in multiple interactions differs from the preset "concise and straightforward" characteristic, the system can appropriately adjust the expression style parameters to better match the actual interaction pattern.
[0148] Secondly, the system optimizes the pre-set information disclosure rules. By analyzing the information disclosure trajectory, the system can assess whether the pace and sequence of information disclosure meet expectations. For example, if it finds that certain information requiring follow-up disclosure is released prematurely without further inquiry, the system can adjust the corresponding disclosure rules and strengthen access control over this information; conversely, if it finds that certain information is not disclosed in a timely manner at a reasonable stage, the system can appropriately relax the disclosure conditions.
[0149] Third, the preset conditions used to trigger accessory feedback are updated. By analyzing the material trigger trajectory, the system can assess whether the timing of material feedback triggering is reasonable. For example, if it is found that some materials are not triggered in a timely manner after a doctor's explicit request, the system can optimize the identification logic of the material trigger conditions; if it is found that materials are mistakenly triggered when the conditions are not met, the system can increase the constraint strength of the trigger conditions.
[0150] Through the aforementioned continuous optimization mechanism, the patient agent in this embodiment of the invention is not a fixed, statically generated role, but a dynamic agent that can continuously learn and evolve through ongoing use. As the number of interactions increases, the consistency of the patient agent's role, the rationality of information disclosure, and the accuracy of multimodal feedback will be continuously improved, thereby better adapting to the needs of different doctors, different cases, and different consultation scenarios, further enhancing the practicality and universality of the virtual diagnosis and treatment system.
[0151] The following detailed explanation of the complete implementation process of the virtual diagnosis and treatment interaction method based on patient intelligence provided by this invention, using a virtual respiratory medicine consultation scenario as an example. This embodiment uses... Figure 4 Based on the flowchart shown, the ten main steps from patient agent construction to session end archiving are fully presented.
[0152] Step 1: Construct the patient's intelligent agent.
[0153] The system first acquires the target case information and constructs a patient agent based on this information. In this embodiment, the constructed patient agent corresponds to a 35-year-old male patient whose chief complaint is "cough with fever for three days," accompanied by symptoms such as yellow sputum and worsening at night. The system sets a general communication style for this patient agent, meaning that the expression is concise, straightforward, and conversational, insensitive to medical terminology, and can only express itself based on its own perception and known information. During the construction process, the system performs structured layering of the case information, dividing it into a basic identity layer, a chief complaint and symptom layer, a medical history layer, an examination material layer, a subjective perception layer, a cognitive boundary layer, and an interaction strategy layer, forming a structured feature template that includes basic tags, disease topic tags, cognitive clarity tags, and information disclosure priority tags. This patient agent will serve as the patient-side subject in subsequent virtual diagnosis and treatment interactions, maintaining role consistency throughout the entire conversation.
[0154] Step 2: The doctor initiates the first round of questions.
[0155] After the virtual consultation session begins, the doctor initiates the first round of questions via voice input: "Where do you feel unwell?" The system receives this multi-source input data in real time and uses it as the doctor's input for the current round. Simultaneously, the system reads the previously disclosed information from the current session's history; this set is empty at the initial stage of the session.
[0156] Step 3: The system identifies the problem type.
[0157] The system processes the received doctor's questions using natural language understanding and semantic analysis to determine the semantic type of the question. It is identified that the doctor's question, "Where do you feel unwell?", is a "complaint-triggered question," corresponding to the complaint collection stage of the consultation. The system marks the current consultation round as Round 1, the current discussion topic as "complaint collection," and constructs a session-level interaction context based on previously disclosed information (an empty set), serving as the basis for subsequent information disclosure determination.
[0158] Step 4: Generate the patient's first response.
[0159] Based on the session-level interaction context constructed in step 3, the system enters the information disclosure scope determination stage. During the chief complaint collection stage, the preset information disclosure rules allow the patient agent to disclose the chief complaint and some basic symptom information. The system extracts the chief complaint layer information "cough with fever for three days" from the patient agent's structured hierarchical information and determines it as the currently permissible range of patient information to be disclosed. Subsequently, the system extracts the patient agent's role characteristics (35-year-old male, normal communication style, concise and straightforward), transforms the structured information items into natural language expressions conforming to first-person cues, and generates a patient text response: "I've been coughing for the past few days, and I also have a slight fever." This response accurately reflects the permissible range of information to be disclosed, and its expression is highly consistent with the patient's role characteristics. The system outputs this text response as a response result to the doctor's end and updates the session state based on the disclosed information, adding the chief complaint information to the disclosed information set.
[0160] Step 5: The doctor continues to ask for details about the symptoms.
[0161] After receiving the initial response, the doctor, following clinical consultation logic, continues to inquire about symptom details, raising a second round of questions: "Is the cough dry or with phlegm?" The system receives this multi-source input, identifies its semantic type as "symptom refinement inquiry," and determines that the current consultation stage has transitioned to the symptom refinement stage. The system reads the currently disclosed historical information (including the chief complaint information) and constructs a new session-level interaction context.
[0162] Step 6: Generate detailed symptom responses.
[0163] Based on the conversation context during the symptom refinement stage, the system extracts information relevant to the current question from the patient agent's symptom layer. It determines that the currently permissible information range is "I have yellow phlegm, and the cough is more pronounced at night." Combining the patient's role characteristics, the system generates a text response from the patient: "I have some yellow phlegm, and the cough is more pronounced at night." This response further refines the symptom description, consistent with the way real patients gradually add information when pressed. The system outputs this response and updates the set of disclosed information, adding more detailed symptom information.
[0164] Step 7: The doctor asks for further information about the patient's medical history.
[0165] The doctor continues the consultation, asking a third round of questions: "Have you had similar experiences before? Do you have any chronic diseases?" The system receives this input, identifies its semantic type as "past medical history inquiry," and determines that the current consultation stage has transitioned to the past medical history collection stage. The system constructs a new session-level interaction context, preparing to determine information disclosure.
[0166] Step 8: Generate medical history responses based on cognitive boundaries.
[0167] Based on the conversation context during the medical history collection phase, the system extracts relevant information from the patient's medical history layer and cognitive boundary layer. It determines that the currently permissible information disclosure range is: the patient's known "previous good health" and "10-year smoking history." Simultaneously, the system detects that the specific diagnosis of "chronic disease" mentioned in the doctor's question falls under the category of incomplete information in the patient's cognition (the patient cannot confirm whether they have a chronic disease), thus triggering cognitive boundary constraints. Based on this constraint, the system generates a composite response containing direct answers and uncertain expressions: for information the patient already knows, the direct answer is "I've coughed before, but not this badly"; for information the patient is unsure about, the response is "I'm not sure." The final generated patient text response is: "I've coughed before, but not this badly" and "I'm not sure if I have a chronic disease." This response provides information the patient can confirm while maintaining a reasonable state of unknown about content beyond their cognitive scope, demonstrating the core role of cognitive boundary constraints.
[0168] Step 9: The doctor reviews the examination materials.
[0169] After collecting the chief complaint, symptoms, and medical history, the doctor requests to view the examination materials: "Can I see your examination report?" The system receives this request, identifies it as a material-triggered interaction, and triggers the material feedback module. The system extracts the blood routine report relevant to the current case from the patient's examination material layer and uses it as the target feedback material. Simultaneously, the system generates a text response from the patient that complements the material feedback: "This is a test I had done at another hospital a couple of days ago." The system then displays the text response and the examination report in a linked manner: the text content is displayed on the interface, and the examination report is presented synchronously as a thumbnail or file link. Through this mechanism, the virtual patient can not only answer verbally during the consultation but also provide supporting documentation like a real patient, forming a complete closed loop of diagnostic and treatment information.
[0170] Step 10: Voice output and session archiving.
[0171] In scenarios configured for voice interaction, the system performs speech conversion on each round of generated text responses. Taking the first round response, "I've been coughing and have a slight fever for the past few days," as an example, the system extracts the patient's agent's voice-related role characteristics (35-year-old male voice, moderate speaking speed, mild anxiety), and calls the speech synthesis module to convert the text into the corresponding speech output. Throughout the entire conversation, the system strictly maintains consistency in voice, speaking speed, and expression style, ensuring that the doctor always hears the same "voice" answering questions.
[0172] After the entire virtual consultation session concludes, the system archives the session. The archived content includes: the doctor's question sequence arranged chronologically ("Where do you feel unwell?", "Is your cough dry or with phlegm?", "Have you had similar symptoms before? Do you have any chronic illnesses?", "Can I see your test results?"), the patient agent's response sequence (corresponding to four rounds of responses), the information disclosure trajectory (the order of disclosure of chief complaint information → symptom details → past medical history information → material feedback), and the material trigger trajectory (records of blood test report retrieval). Based on the archived results, the system optimizes and updates the patient agent, such as adjusting the granularity of information disclosure rules, correcting role consistency parameters, and optimizing material trigger strategies, thereby making the patient agent behave more stably and realistically in subsequent interactions.
[0173] As can be seen from the complete implementation process of the above 10 steps, the patient agent constructed by this invention does not reveal all case information at once, but gradually discloses the chief complaint, symptoms, medical history, and examination materials through multiple rounds of questioning and answering, while maintaining consistency between text and voice roles throughout the process. The patient answers as the doctor asks; when the doctor probes further, the patient adds information step by step; for information the patient does not yet possess, the patient maintains a reasonable level of unawareness—this series of features makes the virtual diagnosis and treatment process highly similar to the consultation patterns and interaction rules in a real outpatient clinic, significantly improving the realism and usability of the virtual diagnosis and treatment system.
[0174] The following describes the virtual diagnosis and treatment interaction system based on patient intelligent agents provided by the embodiments of the present invention. The virtual diagnosis and treatment interaction system based on patient intelligent agents described below can be referred to in correspondence with the virtual diagnosis and treatment interaction method based on patient intelligent agents described above.
[0175] This invention provides a virtual diagnosis and treatment interaction system based on a patient intelligent agent. See [link to relevant documentation]. Figure 5 The system is applied in a virtual interactive environment pre-configured with patient intelligent agents, wherein the patient intelligent agents contain structured, hierarchically organized patient information cognitive boundary constraints and role characteristics. The system includes: The construction module 510 is used to continuously receive multi-source input data in a virtual interaction session and construct a session-level interaction context based on historically disclosed information and the multi-source input data. The determination module 520 is used to determine the range of patient information that is currently allowed to be disclosed within the cognitive boundary constraints, based on the interaction stage of the session-level interaction context and the preset information disclosure rules. The generation module 530 is used to generate a response result that matches the interaction phase based on the determined range of currently allowed patient information disclosure and the role characteristics of the patient agent. The update module 540 is used to output the response result and update the set of disclosed information and the set of information to be disclosed in the current session according to the range of patient information that is currently allowed to be disclosed.
[0176] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can invoke logical instructions in the memory 830 to execute a virtual diagnosis and treatment interaction method based on a patient intelligent agent. This method includes: continuously receiving multi-source input data in a virtual interaction session, and constructing a session-level interaction context based on historically disclosed information and the multi-source input data; determining the range of patient information currently allowed to be disclosed within the cognitive boundary constraints according to the interaction stage of the session-level interaction context and preset information disclosure rules; generating a response result matching the interaction stage based on the determined range of currently allowed patient information and the role characteristics of the patient intelligent agent; outputting the response result, and updating the set of disclosed information and the set of information to be disclosed in the current session according to the range of currently allowed patient information.
[0177] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0178] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the virtual diagnosis and treatment interaction method based on the patient intelligent agent provided by the above methods. The method includes: continuously receiving multi-source input data in a virtual interaction session, and constructing a session-level interaction context based on historically disclosed information and the multi-source input data; determining the range of patient information that is currently allowed to be disclosed within the cognitive boundary constraints according to the interaction stage of the session-level interaction context and preset information disclosure rules; generating a response result that matches the interaction stage based on the determined range of patient information that is currently allowed to be disclosed and the role characteristics of the patient intelligent agent; outputting the response result, and updating the set of disclosed information and the set of information to be disclosed in the current session according to the range of patient information that is currently allowed to be disclosed.
[0179] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the virtual diagnosis and treatment interaction method based on a patient intelligent agent provided by the above methods. The method includes: continuously receiving multi-source input data in a virtual interaction session, and constructing a session-level interaction context based on historically disclosed information and the multi-source input data; determining the range of patient information currently allowed to be disclosed within the cognitive boundary constraints according to the interaction stage of the session-level interaction context and preset information disclosure rules; generating a response result matching the interaction stage based on the determined range of patient information currently allowed to be disclosed and the role characteristics of the patient intelligent agent; outputting the response result, and updating the set of disclosed information and the set of information to be disclosed in the current session according to the range of patient information currently allowed to be disclosed.
[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A virtual diagnosis and treatment interaction method based on a patient intelligent agent, characterized in that, The method, applied in a virtual interactive environment pre-configured with a patient agent, wherein the patient agent comprises structured, hierarchically organized patient information cognitive boundary constraints and role characteristics, includes: In the virtual interactive session, multi-source input data is continuously received, and a session-level interactive context is constructed based on historically disclosed information and the multi-source input data; Based on the interaction stage of the session-level interaction context and the preset information disclosure rules, the range of patient information that is currently allowed to be disclosed is determined within the scope of the cognitive boundary constraints. Based on the determined range of patient information currently allowed to be disclosed and the role characteristics of the patient agent, a response result matching the interaction phase is generated; Output the response result, and update the set of disclosed information and the set of information to be disclosed for the current session according to the currently allowed range of patient information to be disclosed.
2. The method according to claim 1, characterized in that, The patient information of the structured, hierarchical organization is pre-configured through the following steps: Obtain target case information and divide the target case information into at least one of the following layers: basic identity layer, chief complaint and symptom layer, medical history layer, examination material layer, subjective perception layer, cognitive boundary layer and interaction strategy layer; The segmented information at each level is organized into a structured feature template, wherein the structured feature template includes at least one of the following: basic tags, disease topic tags, symptom expression mode tags, cognitive clarity tags, emotion and communication style tags, information disclosure priority tags, attachment material trigger tags, voice style tags, and conversation restrictions and boundaries tags.
3. The method according to claim 1, characterized in that, The step of continuously receiving multi-source input data in a virtual interactive session and constructing a session-level interactive context based on historically disclosed information and the multi-source input data includes: The multi-source input data is received, which includes at least one of the following: doctor's voice input, doctor's text question input, material viewing request triggered by the doctor's terminal, and system status events during the session. The multi-source input data is processed in a unified manner to determine the current consultation round, the current doctor's question, and the current discussion topic; The current consultation round, the current doctor's question, the current discussion topic, and the historically disclosed information are jointly processed to construct the session-level interaction context.
4. The method according to claim 3, characterized in that, The step of determining the range of patient information currently allowed to be disclosed within the cognitive boundary constraints based on the interaction stage of the session-level interaction context and preset information disclosure rules includes: Determine the semantic type of the current doctor's question and the consultation stage of the current discussion topic; Based on the judgment results and the cognitive boundary constraints, the patient information of the structured hierarchical organization is classified into multiple categories of information, which include at least one of the following: information that can be proactively disclosed, information that needs to be disclosed after further questioning, material-triggered information, information with incomplete patient cognition, and information that is prohibited from being proactively disclosed. According to the preset information disclosure rules, the information that can be proactively disclosed and the information that needs to be disclosed after further inquiry are classified, and the information that can be proactively disclosed and the information that needs to be disclosed after further inquiry are included in the scope of patient information that is currently allowed to be disclosed.
5. The method according to claim 4, characterized in that, After classifying the patient information of the structured hierarchical organization into multiple categories based on the judgment result and the cognitive boundary constraints, the method further includes: When the current doctor's question includes information that the patient has incomplete knowledge or information that is prohibited from being disclosed proactively, the cognitive boundary constraint is triggered; Based on the triggered cognitive boundary constraint, the scope of currently allowed patient information disclosure is limited to include uncertain expressions, so as to constrain the generated response result to not exceed the scope of the patient's perception and expression.
6. The method according to claim 1, characterized in that, The step of generating a response result matching the interaction phase based on the determined range of currently permitted patient information disclosure and the role characteristics of the patient agent includes: Extract the role features representing expressive style and emotional state from the patient's intelligent agent; Based on the questioning style of the current doctor's question and the role characteristics, and using the currently allowed range of patient information to be disclosed, a patient text response is generated, and the patient text response is used as the response result that matches the interaction stage.
7. The method according to claim 6, characterized in that, Following the step of using the patient's text response as the response result matching the interaction phase, the method further includes: Extract the patient's gender, patient age group, and existing voice settings contained in the character characteristics; Based on the patient's gender, age group, existing voice settings, and current emotional state, the patient's text response is converted into corresponding speech output content; Maintain consistency in tone, speed, and expression style of the voice output content within the same virtual interactive session, and output the voice output content synchronously with the patient's text response.
8. The method according to claim 1, characterized in that, After the steps of outputting the response result and updating the set of disclosed information and the set of information to be disclosed in the current session according to the currently allowed range of patient information to be disclosed, the method further includes: Monitor whether the preset conditions for triggering attachment feedback are met in the session-level interaction context. The preset conditions include the doctor asking about relevant materials, the doctor making a request to view them, or the interaction stage reaching the point where materials are disclosed. When the preset conditions are met, the examination report, test results or imaging data associated with the response result are extracted from the patient's intelligent agent as target feedback material. The target feedback material and the generated response results are linked and displayed.
9. The method according to claim 1, characterized in that, After the steps of outputting the response result and updating the set of disclosed information and the set of information to be disclosed in the current session according to the currently allowed range of patient information to be disclosed, the method further includes: After the virtual interaction session ends, the continuously received multi-source input data is combined into a doctor's question sequence in chronological order, the continuously generated response results are combined into a patient agent response sequence in chronological order, the continuously determined changes in the range of currently allowed patient information disclosure are recorded and combined into an information disclosure trajectory, and the call records of the target feedback material displayed in conjunction with the output are combined into a material trigger trajectory. The combined doctor question sequence, patient agent response sequence, information disclosure trajectory, and material trigger trajectory are archived. Based on the results of the archiving process, the role characteristics in the patient's intelligent agent, the preset information disclosure rules, and the preset conditions for triggering attachment feedback are optimized and updated.
10. A virtual diagnosis and treatment interaction system based on a patient intelligent agent, characterized in that, The system is applied in a virtual interactive environment pre-configured with patient intelligent agents, wherein the patient intelligent agents contain structured, hierarchically organized patient information cognitive boundary constraints and role characteristics. The system includes: A construction module is used to continuously receive multi-source input data in a virtual interaction session and construct a session-level interaction context based on historically disclosed information and the multi-source input data; The determination module is used to determine the range of patient information that is currently allowed to be disclosed within the cognitive boundary constraints, based on the interaction stage of the session-level interaction context and the preset information disclosure rules. A generation module is used to generate a response result that matches the interaction phase based on the determined range of currently permitted patient information disclosure and the role characteristics of the patient agent. The update module is used to output the response result and update the set of disclosed information and the set of information to be disclosed in the current session according to the range of patient information that is currently allowed to be disclosed.