Computer-aided diagnostic methods, devices, storage media, software products and systems
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本申请提供一种计算机辅助诊断方法、设备、存储介质、程序产品及系统,以减少证据信息损失,解决诊断结论缺乏证据支持、易产生与证据不符的内容、推理过程不可审计等问题
Smart Images

Figure CN122575688A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical artificial intelligence, and in particular to a computer-aided diagnostic method, device, storage medium, program product and system. Background Technology
[0002] In the field of medical image-assisted diagnosis, deep learning-based image analysis models are typically used to combine clinical data to output disease classification results or diagnostic suggestions.
[0003] However, existing solutions often compress information such as the location, shape, and changes of lesions in images into labels or direct conclusions, resulting in a loss of evidentiary information. This leads to problems such as diagnostic conclusions lacking evidentiary support, the generation of content that is inconsistent with the evidence, and the reasoning process being unauditable. Summary of the Invention
[0004] This application provides a computer-aided diagnostic method, device, storage medium, program product, and system to reduce the loss of evidentiary information and solve problems such as diagnostic conclusions lacking evidentiary support, easily generating content inconsistent with evidence, and the reasoning process being unauditable.
[0005] In a first aspect, embodiments of this application provide a computer-aided diagnostic method, including:
[0006] In response to requests for case-assisted diagnosis, obtain medical images and structured clinical data of the case;
[0007] Using a multimodal perception model, structured text containing lesion location information, sign descriptions, and uncertainties is generated based on medical images;
[0008] The reasoning agent infers and generates auxiliary diagnostic information based on structured text and structured clinical data, combined with retrieved medical evidence related to the cases.
[0009] Output auxiliary diagnostic information.
[0010] In one possible embodiment, the multimodal perception model includes a multimodal perception module and a transformation module. The multimodal perception model generates structured text from medical images, containing lesion location information, sign descriptions, and their uncertainties, including:
[0011] Medical images are input into the multimodal perception module of the multimodal perception model to generate a text representation of structured text;
[0012] The conversion module converts the text representation of structured text into text, thus obtaining structured text.
[0013] In one possible embodiment, the training process of the multimodal perception model includes:
[0014] Obtain historical images and corresponding structured annotation text. The structured annotation text contains lesion location information, sign descriptions and their uncertainties in the historical images.
[0015] Historical images are input into the multimodal perception module of the multimodal perception model to generate a predicted representation of structured labeled text;
[0016] The text representation module converts structured labeled text into text representation.
[0017] Based on the predicted representation and the text representation of the structured labeled text, the parameters of the multimodal perception module are adjusted to obtain the trained multimodal perception module.
[0018] In one possible embodiment, the training process of the multimodal perception model further includes:
[0019] A transformation module is constructed based on the structured annotated text and its text representation.
[0020] In one possible embodiment, a multimodal perception model is used to generate structured text from medical images, including lesion localization information, descriptions of signs, and their uncertainties, including:
[0021] A multimodal perception model is used to perform multi-step iterative observation of medical images, one of which includes:
[0022] Action decisions are generated based on the current observation status of medical images. The action decisions include at least one of the following operations: adjusting image display parameters, zooming in on a specified area, switching image layers, and retrieving historical images for comparison.
[0023] The current observation state is updated based on the action decision, and a continuous representation is generated based on the updated current observation state;
[0024] When the termination condition is met, a text representation of the structured text is generated based on the continuous representations obtained from multi-step iterative observations.
[0025] In one possible embodiment, the multimodal perception model includes a multimodal perception model corresponding to at least one type of medical image, including: radiological images, pathological images, ultrasound images, nuclear medicine images, and endoscopic images.
[0026] In one possible embodiment, the inference agent has the ability to use models and invoke tools, and the inference agent supports replacing the model used. The tools that the inference agent can invoke include at least one of the following: data preprocessing and measurement tools, anomaly detection tools, retrieval tools, and artificial intelligence models.
[0027] Through a reasoning agent, based on structured text and structured clinical data, and combined with retrieved case-related medical evidence, auxiliary diagnostic information is generated, including:
[0028] Through reasoning agents, structured text and structured clinical data are used as factual sources of information for cases;
[0029] Invoke the hypothesis generation model to generate at least one diagnostic hypothesis based on factual source information;
[0030] For any diagnostic hypothesis, the hypothesis reasoning model is invoked to retrieve relevant medical evidence based on the diagnostic hypothesis, and evidence-based reasoning is performed on the diagnostic hypothesis based on factual source information and medical evidence to generate a structured evidence chain for the diagnostic hypothesis.
[0031] The verification model is invoked to verify the structured evidence chain of each diagnostic hypothesis from at least one dimension: reliability of evidence source, consistency of facts, consistency of reasoning logic, and sufficiency of evidence citation.
[0032] After the structured evidence chains of each diagnostic hypothesis are verified, the ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured evidence chains across diagnostic hypotheses, and to generate structured auxiliary diagnostic information.
[0033] In one possible embodiment, it also includes:
[0034] For any diagnostic hypothesis, if the structured chain of evidence for the diagnostic hypothesis fails to be verified, the hypothesis reasoning model corresponding to the diagnostic hypothesis is invoked to generate a new structured chain of evidence for the diagnostic hypothesis.
[0035] The verification model is invoked to verify the new structured chain of evidence from at least one dimension: reliability of evidence source, consistency of facts, consistency of reasoning logic, and sufficiency of evidence citation.
[0036] Until a new, structured chain of evidence is validated for the diagnostic hypothesis.
[0037] In one possible embodiment, before invoking the hypothesis reasoning model corresponding to the diagnostic hypothesis to generate a new structured chain of evidence for any diagnostic hypothesis if the verification of the structured chain of evidence for the diagnostic hypothesis fails, the method further includes:
[0038] Based on the verification results of the structured evidence chain of the diagnostic hypothesis, an incremental retrieval query for the diagnostic hypothesis is constructed.
[0039] Incremental search queries retrieve relevant medical evidence from multi-source heterogeneous medical knowledge bases, yielding incremental search results.
[0040] In one possible embodiment, the hypothesis reasoning model corresponding to the diagnostic hypothesis is invoked to generate a new structured chain of evidence for the diagnostic hypothesis, including:
[0041] The hypothesis reasoning model corresponding to the diagnostic hypothesis is invoked. Based on factual source information, medical evidence, and incremental retrieval results, the hypothesis is reasoned on an evidence-based basis to generate a new structured chain of evidence for the diagnostic hypothesis.
[0042] In one possible embodiment, relevant medical evidence is retrieved based on the diagnostic hypothesis, including:
[0043] Transform diagnostic hypotheses into diagnostic clinical questions;
[0044] Relevant medical evidence is retrieved from multi-source heterogeneous medical knowledge bases based on diagnostic clinical questions.
[0045] In one possible embodiment, relevant medical evidence is retrieved from a multi-source heterogeneous medical knowledge base based on a diagnostic clinical question, including:
[0046] The retrieval planning model is invoked to identify the retrieval intent based on the diagnostic clinical question and generate tool invocation information corresponding to the retrieval intent. The tool invocation information includes the identifier of the retrieval tool and the retrieval parameters.
[0047] Based on the tool call information, the corresponding search tool is invoked to retrieve relevant medical evidence from the corresponding medical knowledge base.
[0048] In one possible embodiment, after invoking the corresponding retrieval tool to retrieve relevant medical evidence from the corresponding medical knowledge base based on the tool invocation information, the process further includes:
[0049] The evidence sufficiency assessment model is invoked to determine whether the retrieved medical evidence is sufficient to respond to the diagnostic clinical question based on the diagnostic clinical question and the retrieved medical evidence. If the retrieved medical evidence is insufficient to respond to the diagnostic clinical question, a supplementary query question is generated.
[0050] Relevant medical evidence is retrieved from multi-source heterogeneous medical knowledge bases based on supplementary query questions.
[0051] In one possible embodiment, the method further includes:
[0052] For medical literature from different sources, the content of the medical literature is segmented into semantic segments, and the credibility, clinical task label and content nature label of the semantic segments are determined. The clinical task label is used to identify the clinical task type corresponding to the semantic segment, the content nature label is used to identify the nature type of the content of the semantic segment, and the credibility is determined according to the credibility of the content of the medical literature to which it belongs.
[0053] Semantic fragments are treated as medical knowledge. The semantic fragments and their attribute information are stored in the medical knowledge base, which also stores the literature information and chapter structure information of medical literature. The attribute information of the semantic fragments includes attribute information at three levels: the literature to which they belong, the chapter to which they belong, and the semantic fragment itself.
[0054] In one possible embodiment, the multi-source heterogeneous medical knowledge base includes at least one of the following: a vector knowledge base, a graph knowledge base, and a case knowledge base; the method further includes:
[0055] For any semantic fragment, a multi-granularity index of the semantic fragment is constructed in the medical knowledge base. The multi-granularity index includes at least one of the following: keyword, phrase, whole sentence, and question sentence template.
[0056] In one possible embodiment, a hypothesis generation model is invoked to generate at least one diagnostic hypothesis based on fact source information, including:
[0057] The hypothesis generation model is invoked, and a retrieval tool is used to search for relevant medical evidence in a multi-source heterogeneous medical knowledge base based on fact source information. Based on the fact source information and medical evidence, at least one diagnostic hypothesis is generated.
[0058] In one possible embodiment, after generating auxiliary diagnostic information by reasoning through a reasoning agent based on structured text and structured clinical data, combined with retrieved case-related medical evidence, the process further includes:
[0059] The examination recommendation model is invoked to determine the next candidate examination items based on auxiliary diagnostic information and existing medical evidence. The marginal information value and cost evaluation value of each candidate examination item are evaluated, and the next examination item is recommended based on the marginal information value and cost evaluation value of each candidate examination item.
[0060] Among them, the marginal information value is used to characterize the contribution of candidate examination items to diagnostic uncertainty, and the cost assessment value is used to characterize the cost of candidate examination items in at least one dimension of cost, risk, and time.
[0061] In one possible embodiment, the auxiliary diagnostic information includes:
[0062] The report summarizes the evidence-based diagnosis, identifies diagnostic bottlenecks, and presents a list of diagnostic hypotheses, their corresponding confidence levels, and the rationale behind them.
[0063] In one possible embodiment, the auxiliary diagnostic information also includes suggestions for further medical attention. This auxiliary diagnostic information is generated by a reasoning agent based on structured text and structured clinical data, combined with retrieved case-related medical evidence. The information further includes:
[0064] The suggestion generation model is invoked to generate next-step medical advice based on structured text, structured clinical data, existing medical evidence, diagnostic hypotheses and their confidence levels, structured evidence chains, and diagnostic bottleneck information.
[0065] The next steps for medical attention should include at least one of the following: recommended examinations, treatment and management recommendations, follow-up and monitoring recommendations, and precautions to be informed to the patient.
[0066] In one possible embodiment, after outputting auxiliary diagnostic information, the method further includes:
[0067] Receive feedback data for recommendations on the next medical visit, including newly added imaging and / or newly added clinical data;
[0068] The reasoning agent infers new auxiliary diagnostic information by combining the structured text of medical images, structured clinical data, and the structured text and / or new clinical data of newly added images with retrieved case-related medical evidence. The structured text of newly added images is generated by a multimodal perception model based on the newly added images.
[0069] Output new auxiliary diagnostic information.
[0070] In one possible embodiment, auxiliary diagnostic information is output, including:
[0071] Based on each diagnostic hypothesis and its corresponding confidence level in the auxiliary diagnostic information, determine the inconsistency score of each diagnostic hypothesis;
[0072] Diagnostic hypotheses with inconsistency scores less than a preset inconsistency threshold are added to the prediction set. The preset inconsistency threshold is obtained by calculating the quantile threshold of the inconsistency score distribution corresponding to the real diagnosis in the calibration dataset based on the conformal prediction method.
[0073] Based on the evidence-based diagnostic summary, diagnostic bottleneck information, and the diagnostic hypotheses contained in the prediction set, along with their corresponding confidence levels and reasoning, an auxiliary diagnostic report is generated.
[0074] Output auxiliary diagnostic reports.
[0075] In one possible embodiment, for any diagnostic hypothesis, a hypothesis reasoning model is invoked to retrieve relevant medical evidence based on the diagnostic hypothesis, and evidence-based reasoning is performed on the diagnostic hypothesis based on factual source information and medical evidence to generate a structured chain of evidence for the diagnostic hypothesis, including:
[0076] Identify the specialty to which the diagnostic hypothesis belongs;
[0077] Based on the specialty type to which the diagnostic hypothesis belongs, the hypothesis reasoning model corresponding to the specialty type is invoked. Through the hypothesis reasoning model corresponding to the specialty type, relevant specialty medical evidence is retrieved based on the diagnostic hypothesis and its specialty type. Based on the factual source information and specialty medical evidence, the diagnostic hypothesis is reasoned on an evidence-based basis to generate a structured evidence chain for the diagnostic hypothesis.
[0078] In one possible embodiment, a ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured chain of evidence across diagnostic hypotheses, and to generate structured auxiliary diagnostic information, including:
[0079] The ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured evidence chain across diagnostic hypotheses. The contradictions between the structured evidence chains of each diagnostic hypothesis and the cross-specialty system influence relationship are evaluated, and auxiliary diagnostic information for multi-specialty consultation is generated.
[0080] In one possible embodiment, the method further includes:
[0081] Using a benchmark set containing historical images and their corresponding structured reference texts, we evaluate the perceptual error of the multimodal perception model and the performance ceiling of the inference agent.
[0082] Based on the auxiliary diagnostic information output by the reasoning agent, determine the actual performance of the reasoning agent under the current perception conditions, which include the structured text output by the multimodal perception model.
[0083] The difference between the performance ceiling and the actual performance is used as the performance degradation of the inference agent under the current perception conditions.
[0084] Error attribution is performed on the multimodal perception model and the reasoning agent based on perception error and performance degradation.
[0085] Output the results of error attribution.
[0086] In one possible embodiment, the method further includes:
[0087] After the reasoning model used by the reasoning agent changes, the same benchmark set is used to evaluate the first performance limit of the original reasoning agent before the reasoning model change, and the second performance limit of the new reasoning agent after the reasoning model change.
[0088] Based on the first and second performance limits, determine whether the performance of the new inference agent has degraded, and output the result.
[0089] Secondly, embodiments of this application provide a computer-aided diagnostic device, comprising:
[0090] The data receiving module is used to acquire medical images and structured clinical data of a case in response to a request for case-assisted diagnosis.
[0091] The perception module is used to generate structured text containing lesion location information, sign descriptions, and uncertainties based on medical images through a multimodal perception model.
[0092] The reasoning module is used to generate auxiliary diagnostic information by reasoning based on structured text and structured clinical data, combined with retrieved case-related medical evidence, through a reasoning agent. The reasoning agent can replace the model used.
[0093] The output module is used to output auxiliary diagnostic information.
[0094] Thirdly, embodiments of this application provide an electronic device, including:
[0095] At least one processor; and
[0096] Memory that is communicatively connected to at least one processor;
[0097] The memory stores instructions that can be executed by at least one processor to cause the electronic device to perform the methods provided above.
[0098] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method provided above.
[0099] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the methods provided above.
[0100] Sixthly, embodiments of this application provide a computer-aided diagnostic system, comprising:
[0101] The client and the electronic devices provided by the aforementioned third party,
[0102] The client is used to send case-assisted diagnosis requests to electronic devices and to provide medical images and structured clinical data of the cases to the electronic devices.
[0103] The computer-aided diagnostic methods, devices, storage media, program products, and systems provided in this application, by responding to a case's request for auxiliary diagnosis, acquire medical images and structured clinical data of the case, and generate structured text containing lesion location information, sign descriptions, and uncertainties based on the medical images using a multimodal perception model. This reduces information compression loss during the expression of image diagnostic evidence and improves the fine-grained representation capability of diagnostic evidence. By using a reasoning agent to infer and generate auxiliary diagnostic information based on the structured text and structured clinical data, combined with retrieved case-related medical evidence, the comprehensive analysis capability of heterogeneous evidence and the flexibility and auditability of the reasoning process can be enhanced. This, in turn, improves the granularity of evidence' support for the diagnostic result (i.e., diagnostic hypothesis) in the auxiliary diagnostic information and the consistency between the diagnostic result and the evidence. In this embodiment, the assisted diagnosis process is divided into a decoupled perception stage and a reasoning stage: the multimodal perception model outputs structured text in natural language form based on medical images, rather than diagnostic conclusions; the reasoning agent performs fusion reasoning based on multi-source evidence, such as structured text from medical images, structured clinical data, and retrieved medical evidence. This reduces information loss while achieving flexibility and auditability in the reasoning process, thereby improving the granularity of evidence support for diagnostic results (i.e., diagnostic hypotheses) in assisted diagnostic information and the consistency between diagnostic results and evidence. Attached Figure Description
[0104] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0105] Figure 1 A flowchart of a computer-aided diagnostic method provided in an exemplary embodiment of this application;
[0106] Figure 2 A flowchart illustrating evidence-based reasoning by a reasoning agent provided in an exemplary embodiment of this application;
[0107] Figure 3 Example diagram of auxiliary diagnostic information provided for an exemplary embodiment of this application;
[0108] Figure 4 A framework diagram of a computer-aided diagnostic system provided in an embodiment of this application;
[0109] Figure 5 A framework diagram of the multi-specialty consultation subsystem provided in the embodiments of this application;
[0110] Figure 6 A schematic diagram of the structure of a computer-aided diagnostic device provided in an exemplary embodiment of this application;
[0111] Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0112] Figure 8 This is a structural block diagram of a computing device according to an embodiment of this application.
[0113] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0114] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0115] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0116] PICO is a standardized, structured framework for constructing answerable clinical questions in Evidence-Based Medicine (EBM) practice. This framework breaks down clinical questions into four core elements to facilitate their retrieval and evaluation. The four elements are: P (Population / Patient / Problem): This refers to the patient population, population characteristics, or clinical problem to be solved, typically including descriptions of comorbidities, risk factors, and specific disease states; I (Intervention / Exposure): This refers to the intervention, diagnostic test, exposure factor, or prognostic indicator under consideration; C (Comparison / Control): This refers to the alternative compared to the intervention, which can be a blank control, placebo, standard treatment, or other interventions; and O (Outcome): This refers to the clinical outcome that is expected to be measured, improved, or influenced, and can be quantifiable health-related results such as disease incidence, mortality, complication rate, and diagnostic accuracy indicators (sensitivity / specificity).
[0117] Large models refer to deep learning models with a large number of parameters. Also known as foundation models (FM), they are pre-trained models produced by pre-training on large-scale unlabeled corpora. These models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and Multi-modal Pre-training Models. In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before being applied to different tasks. Large models can be widely used in Natural Language Processing (NLP), Computer Vision, and other fields. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. Major application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0118] Medical image-assisted diagnosis technology belongs to the field of medical artificial intelligence and medical information processing. It is usually used in hospital radiology departments, specialist outpatient clinics and multi-specialty consultation scenarios. It is used to combine medical images such as computed tomography (CT) and magnetic resonance imaging (MRI) with clinical data such as test reports and medical history to provide doctors with disease analysis results or diagnostic suggestions.
[0119] In this type of scenario, after receiving case images and clinical information, the auxiliary diagnostic system identifies and analyzes the lesions, and then provides the analysis results to doctors in the form of classification labels, risk probabilities, or brief diagnostic conclusions.
[0120] In existing technologies, a common approach is to use deep learning models to extract features from medical images and directly output disease categories, benign or malignant probabilities, or single diagnostic suggestions. The basic idea is to use the image perception results directly as input for subsequent diagnostic judgments, thus forming a tightly coupled processing chain from perception to conclusion.
[0121] However, under this technical approach, information closely related to clinical judgment in imaging, such as the location, morphological characteristics, boundary status, and trend of changes of lesions, is often compressed into a small number of labels or probability values, resulting in insufficient granularity of evidence expression. Although doctors can see the conclusions given by the system, they find it difficult to clearly understand what specific imaging signs and clinical evidence the conclusions are based on, and it is also difficult to independently verify the perceived results and the inferred results.
[0122] For complex cases involving multiple systems and specialties, the aforementioned compressed output weakens the ability to express the correlation between heterogeneous evidence, making it difficult to form a complete and coherent chain of diagnostic evidence between imaging information and clinical data. This not only affects the comprehensiveness and auditability of the analysis but also increases system maintenance and adaptation costs when the model is upgraded or replaced. Therefore, how to preserve the expression of imaging evidence as much as possible during medical auxiliary diagnosis and enable the diagnostic generation process to combine clinical data and case-related medical evidence for traceable reasoning has become an urgent technical problem to be solved.
[0123] To address the aforementioned issues, this application provides a computer-aided diagnostic method. Upon responding to a case's request for auxiliary diagnosis, the method acquires the case's medical images and structured clinical data. A multimodal perception model generates structured text containing lesion location information, sign descriptions, and uncertainties based on the medical images. This reduces information compression loss during the expression of image diagnostic evidence and enhances the fine-grained representation of diagnostic evidence. Furthermore, an inference agent generates and outputs auxiliary diagnostic information based on the structured text, structured clinical data, and retrieved case-related medical evidence. This enhances the comprehensive analysis capability of heterogeneous evidence and the flexibility and auditability of the reasoning process, thereby improving the granularity of evidence' support for the diagnostic result (i.e., the diagnostic hypothesis) and the consistency between the diagnostic result and the evidence in the auxiliary diagnostic information.
[0124] In this embodiment, the assisted diagnosis process is divided into a decoupled perception stage and a reasoning stage: the multimodal perception model outputs structured text in natural language form based on medical images, rather than diagnostic conclusions; the reasoning agent performs fusion reasoning based on multi-source evidence, including structured text from medical images, structured clinical data, and retrieved medical evidence. Using structured text as an intermediate expression between image evidence and subsequent reasoning helps reduce information loss caused by evidence compression and improves the auditability and complex case analysis capabilities of the assisted diagnosis process.
[0125] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0126] Figure 1 This is a flowchart illustrating a computer-aided diagnostic method provided in an exemplary embodiment of this application. The executing entity in this embodiment can be an electronic device within an auxiliary diagnostic system, specifically a hospital intranet server, a dedicated inference workstation, a cloud-based medical computing node, or a diagnostic service module deployed in an imaging information system; no specific limitations are made here. Figure 1 As shown, the specific steps of this method are as follows:
[0127] S101: In response to a case-aided diagnostic request, obtain the case's medical images and structured clinical data.
[0128] The case-assisted diagnosis request in this step is used to trigger a case-oriented auxiliary diagnosis process. This request carries at least one of the following: case identifier, examination identifier, data access credentials, or session context, enabling the executing entity to locate the corresponding case's data resources. In an example scenario, the case-assisted diagnosis request may carry the user-inputted original query. When carrying the original query, both the original query and structured clinical data can serve as the source of factual information for the input.
[0129] Medical imaging can refer to the imaging examination results corresponding to a case, including but not limited to various types of medical images such as radiological images (e.g., CT scans), pathological images (e.g., pathological slides), ultrasound images, nuclear medicine images, and endoscopic images. Structured clinical data refers to the structured information of clinical data, such as clinical factual information that can be directly involved in subsequent reasoning in a field-based manner, which may include but is not limited to laboratory data and medical history data.
[0130] In practice, upon receiving a request for auxiliary diagnosis of a case, the implementing entity acquires the medical images and structured clinical data corresponding to the current case. Optionally, the acquired medical images and structured clinical data can be preprocessed (e.g., identifying CT phases) to facilitate subsequent processing.
[0131] In one possible implementation, structured clinical data is not processed through a multimodal perception model, but is instead passed to the inference agent as an independent source of clinical facts, entering the inference chain in parallel with the structured text generated from medical images. Based on the above analysis, this step establishes a unified input foundation for imaging evidence and non-imaging clinical facts, enabling subsequent processing to conduct joint analysis while preserving their original expression granularity.
[0132] S102: Using a multimodal perception model, generate structured text from medical images that includes lesion location information, sign descriptions, and their uncertainties.
[0133] The multimodal perception model in this step is used to convert the visual information contained in medical images into structured text suitable for subsequent inference processing. This structured text includes lesion localization information, sign descriptions, and uncertainties generated based on the medical images. The structured text can be a text result with a fixed field organization format, which may include lesion entry identifiers, anatomical location fields, sign fields, uncertainty fields, and necessary comparison description fields, thus serving as an intermediate representation between the multimodal perception model output and the inference agent input.
[0134] In this context, a lesion refers to a diseased or abnormal tissue or organ in a localized area of the body, such as a tumor, an inflamed lung lobe, a cyst, or a hemorrhage point. Lesion localization information is used to characterize the anatomical location, side, level, regional boundary, or spatial reference position of the lesion. A sign refers to the specific "appearance" of the lesion on imaging films (CT / MRI). A sign is an objective fact presented by medical imaging, a specific change seen on the image, such as "a round white shadow in the lung" or "a dark shadow with indistinct borders in the brain." Sign description is an objective description of the sign in medical imaging, used to characterize one or more of the following: lesion size, shape, density or signal characteristics, edge state, enhancement pattern, internal components, surrounding invasion, accompanying signs, and trends over time. Uncertainty in sign description characterizes the reliability of the sign description, and can be expressed using confidence scores, grading indicators, probability intervals, or discrete labels such as "uncertain," "to be identified," or "affected by artifacts."
[0135] In practice, medical images can be input into a multimodal perception model to generate structured text. For example, for a chest CT scan, the output could be "Lesion 1: Right upper lobe posterior segment, nodule, long diameter 8mm, lobulated, mild spiculation, pleural traction visible, uncertainty 0.18; no obvious enlargement of mediastinal lymph nodes seen, uncertainty 0.06"; for a brain MRI, the output could be "Lesion 1: Left frontal lobe subcortical, patchy long T2 signal, indistinct borders, minimal enhancement, slightly larger area compared to previous images, uncertainty 0.27". This type of structured text can preserve lesion-level evidence, without directly compressing it into a diagnostic category or outputting a diagnostic conclusion. This reduces information compression loss during the expression of imaging diagnostic evidence and improves the fine-grained representation capability of diagnostic evidence.
[0136] In one optional embodiment, the multimodal perception model may include a multimodal perception module and a conversion module. In this step, the medical image is input into the multimodal perception module of the multimodal perception model, which generates a text representation of structured text. Then, the conversion module converts the text representation of the structured text into actual text, resulting in structured text.
[0137] The multimodal perception module takes medical images as input and outputs structured text. It can be implemented using a trained ViT (Vision Transformer) or other similar multimodal models. The conversion module can be implemented using any text decoder, without any specific limitations.
[0138] In one possible implementation, the parameters of the transformation module can be jointly trained with the parameters of the multimodal perception module to ensure that the internal text representation is semantically consistent with the output structured text.
[0139] The structure of this multimodal perception model involves first forming a text representation that preserves lesion evidence information in the multimodal perception module, and then converting it into structured text by the transformation module. This allows key information from medical images to be output in structured text form, serving as input for the subsequent reasoning agent, thus forming an intermediate expression between image perception results and diagnostic reasoning. Because lesion localization information, sign descriptions, and their uncertainties are all preserved in the structured text output, subsequent diagnostic processes can obtain a more complete representation of image evidence.
[0140] With this specific implementation, the location, morphological features, and uncertainties of lesions in medical images are no longer compressed into single labels or probability values, but are preserved and output in the form of structured text. This supports subsequent reading by the inference agent, manual verification, and case tracing. At the same time, the separate design of the multimodal perception module and the conversion module decouples the internal perception representation from the external text expression, facilitating model replacement and functional expansion while maintaining output consistency.
[0141] In this embodiment, the medical images can be any applicable image data, and the fields and formats contained in the structured text can also be adapted according to the actual diagnostic scenario. Based on the above analysis, this step, by explicitly outputting lesion localization, signs, and uncertainties in the form of structured text, ensures that the imaging evidence remains readable, verifiable, and citationable before entering the reasoning stage, thereby providing a direct evidence basis for subsequent evidence-based diagnosis.
[0142] In one possible implementation, the multimodal perception model can be configured to actively examine medical images, i.e., autonomously adjust window width and level, zoom in on suspicious areas, switch slices, or retrieve historical images for comparison, forming structured text based on multi-step observation. For example, generating structured text containing lesion location information, sign descriptions, and uncertainties from medical images using a multimodal perception model can be achieved in the following way:
[0143] A multimodal perception model is used to perform multi-step iterative observation of medical images. One step of iterative observation includes: generating action decisions based on the current observation state of the medical image, updating the current observation state based on the action decisions, and generating a continuous representation based on the updated current observation state. The action decisions include at least one of the following operations: adjusting image display parameters, zooming in on a specified area, switching image layers, or retrieving historical images for comparison. When a termination condition is met, a structured text representation is generated based on the continuous representation obtained from the multi-step iterative observation. The termination condition can be that the number of iterations reaches a preset threshold, the current observation state meets a preset convergence condition, or sufficient information has been obtained to generate structured text. The specific termination condition can be set according to the actual application scenario requirements and empirical values, and is not specifically limited here.
[0144] During multi-step iterative observation, after each update of the current observation state based on action decisions, the multimodal perception model encodes new local observation results, cross-level information, and historical comparison information into a continuous representation. This continuous representation uses continuous vectors to carry lesion location clues, morphological features, boundary features, and change clues, and accumulates continuously with each observation round to form a layer-by-layer aggregated expression of the medical image. When the termination condition is met, the continuous representation obtained from multi-step iterative observation is used as the text representation of structured text. Furthermore, the continuous representation obtained from multi-step iterative observation is input into a conversion module to convert the text representation of structured text into structured text containing lesion location information, sign descriptions, and their uncertainties.
[0145] For example, in actual operation, the multimodal perception model can work collaboratively based on an encoder (i.e., the multimodal perception module), a state update unit, and a text decoding unit (i.e., the transformation module). The encoder is used to extract image features from different observation rounds, the state update unit is used to fuse the current observation state with the action decision result, and the text decoding unit is used to map the continuous representation into structured text. This approach allows the multimodal perception model to no longer output compressed conclusions all at once, but to gradually form a more complete expression of image evidence through iterative observation, and output structured text that can be used for subsequent reasoning after the iteration terminates.
[0146] By adopting this approach, the multimodal perception model can retain richer lesion evidence when generating structured text, making the descriptions of lesion localization, signs, and uncertainties in the structured text closer to the results obtained from multiple rounds of observation. This improves the integrity and traceability of medical image expression and provides a stable intermediate representation for subsequent auxiliary diagnostic reasoning.
[0147] In one optional embodiment, the multimodal perception model may include a multimodal perception model corresponding to at least one type of medical image. The medical image types include, but are not limited to: radiological images, pathological images, ultrasound images, nuclear medicine images, and endoscopic images.
[0148] For any type of medical image, a multimodal perception model specifically designed for processing that type of medical image can be trained. For example, a multimodal perception model for radiological images can employ a combination of a visual encoder (such as ViT) and a text decoder to encode the location, size, shape, edges, and enhancement features of lesions, and output a structured text representation. A multimodal perception model for pathological images can use a multi-instance learning structure or a slice-level coding structure to characterize glandular structure, cellular atypia, infiltration extent, and histological grade. A multimodal perception model for ultrasound images can combine intra-frame spatial features with necessary temporal dimensions to describe echo intensity, boundary clarity, posterior echo changes, and blood flow signals. A multimodal perception model for nuclear medicine images can model the distribution of radioactive uptake, abnormal aggregation areas, and differences in metabolic activity. A multimodal perception model for endoscopic images can perceive changes in mucosal color, surface microstructure, morphology of bulges or depressions, and signs of bleeding and erosion. The multimodal perception model can call the corresponding parameter group or the corresponding branch network according to the type of input image. In practical applications, this component can also be implemented in other ways, and this application does not limit it.
[0149] Through the above configuration, the multimodal perception model can be specifically adapted to the imaging mechanisms and visual features of different medical images. This allows the input medical images to be transformed into structured text that more closely reflects the expression habits of that image type after perception, while retaining information such as lesion localization, sign description, and uncertainties within the structured text. Consequently, subsequent diagnostic reasoning can be directly based on fine-grained image evidence from different modalities, reducing information loss caused by modal mixing or representation compression, and improving the adaptability and consistency of the structured text across various medical imaging scenarios.
[0150] S103: Through a reasoning agent, based on structured text and structured clinical data, combined with retrieved case-related medical evidence, reason to generate auxiliary diagnostic information.
[0151] The reasoning agent in this step is the core reasoning unit. It can receive structured text from S102 and structured clinical data from S101, and based on this, combine it with retrieved medical evidence related to the current case to form auxiliary diagnostic information. The auxiliary diagnostic information is the diagnostic support result generated by the reasoning, which may include a list of candidate diagnoses, the ranking of each candidate diagnose, corresponding confidence levels, diagnostic bottlenecks, structured evidence chains, supplementary information items, and explanations of multi-specialty impacts. The information items included in the auxiliary diagnostic information and the output format can be designed and adjusted according to the actual diagnostic scenario requirements; no specific limitations are made here.
[0152] In practice, the reasoning agent first integrates structured text and structured clinical data into factual source information for the case. This factual source information does not categorize or compress the original imaging evidence; instead, it retains specific information such as lesion-level descriptions, temporal changes, abnormal test results, symptoms, and signs, along with source identifiers for each field. Subsequently, the reasoning agent combines the retrieved case-related medical evidence to perform reasoning analysis on the case and generate auxiliary diagnostic information. This step uses structured text as an intermediate layer of imaging evidence and incorporates it, along with structured clinical data and external medical evidence, into the reasoning process, ensuring a traceable formation path for the auxiliary diagnostic information.
[0153] S104: Output auxiliary diagnostic information.
[0154] The output in this step refers to providing the auxiliary diagnostic information generated by S103 to the target user or target system in a readable, archiveable, and interactive form.
[0155] The auxiliary diagnostic information includes at least a sequentially arranged list of candidate diagnoses, their corresponding confidence levels or ranking scores, a summary of the structured chain of evidence associated with each candidate diagnosis, and current diagnostic bottleneck information. Diagnostic bottleneck information indicates key missing evidence, conflicting signs, or supplementary examinations that lead to difficulty in distinguishing diagnoses. The output can be to a physician workstation, multidisciplinary consultation terminal, medical record system, imaging reporting system, or clinical decision support platform; no specific limitations are imposed here.
[0156] For example, electronic devices encode auxiliary diagnostic information into two types of results: structured messages and presentation text. Structured messages can take the form of JSON (JavaScript Object Notation) objects, database records, or FHIR (FastHealthcare Interoperability Resources) extended resources, and include case identifiers, task numbers, candidate diagnosis lists, evidence citation identifiers, generation time, model version, and validation status. Presentation text is organized according to clinical reading habits as an evidence-based diagnostic summary, such as summarizing the main lesion signs, key clinical facts, supporting and opposing evidence, and points of contention to be addressed.
[0157] Optionally, when outputting auxiliary diagnostic information, structured text, corresponding image localization anchors, and links to literature sources in the evidence chain can be simultaneously displayed in the doctor's interface, allowing users to trace back from candidate diagnoses to lesion descriptions, and then back to original images and external evidence. If the sorting results contain multiple possible candidate diagnoses, the distinguishing criteria between each candidate diagnose and remaining uncertainties can also be displayed simultaneously.
[0158] The solution implemented in this application includes: in response to a case's request for auxiliary diagnosis, acquiring the case's medical images and structured clinical data; generating structured text containing lesion location information, sign descriptions, and their uncertainties based on the medical images using a multimodal perception model, which can reduce information compression loss during the expression of image diagnostic evidence and improve the fine-grained representation capability of diagnostic evidence; generating auxiliary diagnostic information through an inference agent based on the structured text and structured clinical data, combined with retrieved case-related medical evidence, and the inference agent supports alternative models; outputting auxiliary diagnostic information, which can enhance the comprehensive analysis capability of heterogeneous evidence and the flexibility and auditability of the inference process, thereby improving the granularity of evidence' support for the diagnostic result (i.e., diagnostic hypothesis) and the consistency between the diagnostic result and the evidence in the auxiliary diagnostic information.
[0159] In this embodiment, the assisted diagnosis process is divided into a decoupled perception stage and a reasoning stage: a multimodal perception model outputs structured text in natural language form, rather than diagnostic conclusions, based on medical images; and a reasoning agent performs fusion reasoning based on multi-source evidence, including structured text from medical images, structured clinical data, and retrieved medical evidence. By using structured text as an intermediate expression between image evidence and subsequent reasoning, the formation process of assisted diagnostic information can retain the granularity of image evidence and can be traced along the source of facts, evidence retrieval, and reasoning path. This can reduce information loss caused by evidence compression and improve the auditability of the assisted diagnosis process and the ability to analyze complex cases.
[0160] In one optional embodiment, the training process of the multimodal perception model includes: acquiring historical images and corresponding structured annotation text, wherein the structured annotation text contains lesion location information, sign descriptions and their uncertainties in the historical images; inputting the historical images into the multimodal perception module of the multimodal perception model to generate a predicted representation of the structured annotation text; converting the structured annotation text into a text representation through a text representation module; and adjusting the parameters of the multimodal perception module according to the predicted representation and the text representation of the structured annotation text to obtain the trained multimodal perception module.
[0161] Historical images can be medical images obtained from publicly available training sets, or medical images obtained from hospital data systems or third-party platforms that have been authorized for use by the user. Specifically, they can cover various types of images such as radiological images, pathological images, ultrasound images, nuclear medicine images, and endoscopic images.
[0162] Structured annotated text of historical images refers to annotated / reviewed text data containing lesion location information, sign descriptions, and uncertainties in historical images. Structured annotated text can be extracted from the actual diagnostic report corresponding to the historical images or obtained through expert annotation; no specific limitations are made here.
[0163] During training, the structured labeled text is encoded by the text representation module to obtain its text representation. After historical images are input into the multimodal perception module, the multimodal perception model outputs a predicted representation that resides in the same representation space as the text representation of the structured labeled text. This predicted representation characterizes the fitting results of the multimodal perception model for lesion localization information, sign description, and their uncertainties. Furthermore, based on the difference between the predicted representation and the text representation, the parameters of the multimodal perception module are updated in reverse, thus training the multimodal perception model. The difference between the predicted representation and the text representation can be characterized by cosine distance, cross-entropy loss, etc., allowing the trained multimodal perception module to gradually learn the standard expression form corresponding to the structured labeled text.
[0164] In this embodiment, the text representation module can be implemented using a pre-trained text encoder, which is any model capable of converting text into text representation; no specific limitation is made here.
[0165] The conversion module in the multimodal perception model is used to convert the text representation of structured text into text, which is the reverse process of the text representation module encoding structured labeled text into text representation. In an optional embodiment, the conversion module can be constructed based on the structured labeled text and its text representation. For example, during the training process of the multimodal perception module, the text representation of the structured labeled text generated by the text representation module is stored in correspondence with the structured labeled text, that is, the data pairs <text representation, structured labeled text> are stored, which constitute a lightweight text decoder and serve as the conversion module of the multimodal perception model.
[0166] In the trained multimodal perception model, the conversion module works collaboratively with the multimodal perception module. The text representation output by the multimodal perception module is mapped by the conversion module to form structured text, thereby ensuring that lesion localization information, sign descriptions, and their uncertainties are output in a consistent text format. This construction method maintains a one-to-one correspondence between the conversion module and the structured labeled text in the training data, reducing text output offset and ensuring the stability and reusability of the structured representation of the multimodal perception model across different cases. In other optional embodiments, the conversion module can be implemented using any text decoder. In one optional embodiment, the conversion module of the multimodal perception model can be trained on a text encoder based on the structured labeled text and its text representation.
[0167] Through the above training method, the multimodal perception module can maintain the integrity of lesion localization, sign description and uncertainty information at the output end, so that its generated results are consistent with the manual structured annotation in terms of semantics and expression, thereby improving the accuracy and stability of subsequent structured text generation, and making the intermediate representation of image evidence more convenient for subsequent diagnostic reasoning and verification.
[0168] Figure 2 This is a flowchart illustrating evidence-based reasoning performed by an inference agent provided in an exemplary embodiment of this application. Building upon the foregoing embodiments, in an optional implementation, the inference agent has the ability to use models and invoke tools, and the inference agent supports replacing the models used.
[0169] This embodiment provides a comprehensive tool ecosystem for LLM, integrating various medical tools from different dimensions. These tools interact with the inference agent engine in one step or are invoked on demand through a Model Context Protocol Server (MCP). The tools that the inference agent can invoke include at least one of the following: data preprocessing and measurement tools, anomaly detection tools, retrieval tools, and artificial intelligence (AI) models (such as LLM).
[0170] For example, data preprocessing and measurement tools may include, but are not limited to: CT phase recognition tools, and tools for automatic segmentation and precise measurement and localization of various tissues and organs (such as pancreatic organs and substructures, such as pancreatic ducts and bile ducts). Anomaly detection tools may include, but are not limited to: multi-disease comparative learning models (a classification and recognition model for multiple diseases), diagnostic models for specific diseases such as pancreas, liver, breast, and lung, and tools for retrieving abnormal medical histories from historical medical records. Retrieval tools include, but are not limited to: similar case retrieval tools, retrieval tools for various medical knowledge bases (such as internal disease knowledge bases) (used to retrieve medical knowledge such as clinical guidelines, literature, and disease classifications), tools for obtaining new external medical knowledge through internet searches, and patient data retrieval tools (such as querying clinical records). Artificial intelligence models refer to models that can be used by reasoning agents, such as large models like LLM.
[0171] In this embodiment, the inference agent can use a model (such as an LLM or other large model) through a task interface during the inference process. In an optional embodiment, the inference agent adopts a calling method that decouples the task interface from the model. Therefore, the model used can be replaced without changing the overall process, allowing the assisted diagnostic system to continuously evolve with the model used, avoiding the risk of strong coupling with a specific model.
[0172] In practical implementation, corresponding prompts and callable tools can be configured for the tasks that the LLM needs to perform during the inference process. When the LLM needs to perform different tasks, the prompts and callable tools for the current task are provided to the LLM, enabling the LLM to perform the current task based on the prompts and call tools as needed during the execution of the current task.
[0173] For example, task prompts may include the following: role setting, task description, contextual information, output format and requirements, examples, etc. The role setting tells the LLM what identity or role it should play. For example, for a diagnostic hypothesis generation task, the role setting could be "You are a hypothesis generation node in an evidence-based diagnostic system"; for a hypothesis reasoning task, the role setting could be "You are a hypothesis reasoning node in an evidence-based diagnostic system, a clinical evidence-based reasoning expert."
[0174] The task description is used to indicate the core operations or task results that LLM needs to perform. For example, for a diagnostic hypothesis generation task, the task description could be "You are responsible for generating a comprehensive and evidence-based list of diagnostic hypotheses based on the given factual sources and medical knowledge base retrieval"; for a diagnostic hypothesis generation task, the task description could be "Your core task is to use known patient information and medical knowledge retrieval tools to perform systematic evidence-based reasoning and produce a structured chain of evidence for the currently specified diagnostic hypothesis."
[0175] Contextual information provides the strategies, rules, background information, raw data, known conditions, external constraints, and available tools required to complete the task. Output format and requirements explicitly define the form, structure, data format, length, language style, and other visual or quantifiable specifications of the LLM output. Examples provide the LLM with at least one set of input-output pairings. Furthermore, task prompts may include other information required to perform the task; these can be designed according to the specific task scenario and are not specifically limited here.
[0176] like Figure 2 As shown, the reasoning agent, based on structured text and structured clinical data, and combined with retrieved case-related medical evidence, generates auxiliary diagnostic information through reasoning. The specific steps include:
[0177] S201: Through reasoning agents, structured text and structured clinical data are used as factual sources of information for cases.
[0178] In this embodiment, structured text and structured clinical data are used as factual sources of information for evidence-based diagnosis of cases, and together with relevant retrieved medical evidence, they serve as sources of evidence that can be cited for evidence-based diagnosis.
[0179] S202: Invoke the hypothesis generation model to generate at least one diagnostic hypothesis based on fact source information.
[0180] In this step, the LLM is invoked to perform the hypothesis generation task. Specifically, the hypothesis generation task prompts configured for the task are retrieved, and the fact source information is input into the LLM along with these prompts. This allows the LLM to generate at least one diagnostic hypothesis based on the fact source information and the prompts. During the hypothesis generation task, the LLM can invoke tools as needed based on the available tool information provided by the prompts to acquire external knowledge or execute corresponding processing logic.
[0181] The hypothetical task prompts include the role settings, task description, context information, output format and requirements, and examples corresponding to the hypothetical task.
[0182] For example, suppose a sample structure for generating task prompts is as follows:
[0183] Character setting and mission description: ...;
[0184] Tool usage: ...;
[0185] Search strategy: ...;
[0186] Assume the generation rule is: ...;
[0187] Output format (output a JSON array, where each element must be a diagnostic hypothesis string, and each string must strictly conform to the following format: "Diagnosis name (confidence level: high / medium / low): one-sentence summary of key supporting evidence"...).
[0188] The hypothesis generation task prompts used in this embodiment can be designed according to the needs and requirements for generating diagnostic hypotheses in actual application scenarios. The structure and specific content of the hypothesis generation task prompts are not limited here.
[0189] In one possible implementation, invoking a hypothesis generation model to generate at least one diagnostic hypothesis based on fact source information includes: invoking the hypothesis generation model, using a retrieval tool to retrieve relevant medical evidence from a multi-source heterogeneous medical knowledge base based on fact source information, and generating at least one diagnostic hypothesis based on the fact source information and the medical evidence.
[0190] In its implementation, the hypothesis generation model first constructs a retrieval request based on factual source information, and then uses a retrieval tool to obtain evidence fragments related to the current case from a multi-source heterogeneous medical knowledge base. Subsequently, the factual source information and medical evidence are jointly encoded, and by combining the correlation between evidence, the credibility of the evidence, and its consistency with the case facts, at least one diagnostic hypothesis is generated. The diagnostic hypothesis can be a single disease candidate or a set of multiple candidates with discriminative relationships. The hypothesis generation model can output the confidence level or evidence citation identifier corresponding to the diagnostic hypothesis for subsequent inference agents to use. In practical applications, this hypothesis generation model can also use generative networks with different parameter scales or different structures; this application does not impose any limitations on this.
[0191] By employing a search-then-generate approach, diagnostic hypotheses no longer rely solely on direct mappings of case facts. Instead, they are built upon a foundation of multi-source, heterogeneous medical evidence. This ensures that candidate diagnoses align with medical knowledge in the medical knowledge base and provides a clearer starting point for subsequent evidence-based reasoning. The resulting diagnostic hypotheses possess stronger evidentiary relevance and traceability, enhancing the plausibility of candidate diagnoses in complex cases and reducing biases when hypotheses are made based solely on partial facts.
[0192] S203: For any diagnostic hypothesis, by calling the hypothesis reasoning model, relevant medical evidence is retrieved based on the diagnostic hypothesis, and evidence-based reasoning is performed on the diagnostic hypothesis based on factual source information and medical evidence to generate a structured evidence chain for the diagnostic hypothesis.
[0193] In this step, the hypothetical reasoning model can be an LLM (Limited Least Metric), which is invoked to perform the hypothetical reasoning task. Specifically, the hypothetical reasoning task prompts configured for the task are obtained. The diagnostic hypothesis and factual source information are input into the LLM along with the prompts. Based on the prompts, the LLM retrieves relevant medical evidence according to the diagnostic hypothesis and performs evidence-based reasoning on the diagnostic hypothesis based on the factual source information and medical evidence, generating a structured chain of evidence for the diagnostic hypothesis.
[0194] During the hypothetical reasoning task, LLM can invoke retrieval tools on demand based on the available tool information provided by the hypothetical reasoning task prompts to obtain the evidence needed for the evidence-based reasoning of the current diagnostic hypothesis.
[0195] The hypothetical reasoning task prompts include the corresponding role settings, task description, contextual information, output format and requirements, and examples. The contextual information in the hypothetical reasoning task prompts may include, but is not limited to: core constraints of hypothetical reasoning, tool usage strategies, diagnostic reasoning dimensions, classification principles, retry scenario constraints, retrieval priority, reasoning requirements, and output format.
[0196] The core constraints of hypothetical reasoning include, but are not limited to:
[0197] Language consistency constraints: The output language must be consistent with the primary language of the input original query and patient information; the query language should conform to the tool description and usage guidelines, such as preferably using the language clearly marked in the tool description; cross-language citation of evidence, etc.
[0198] Diagnostic hypothesis focus constraint: Only reason about the diagnostic hypotheses specified in the input; do not produce analysis with unspecified hypotheses; do not perform global diagnostic comparisons.
[0199] Evidence-based constraint: All judgments must be supported by evidence—from factual sources or retrieved medical knowledge. Fabricated examination results, medical histories, or literature citations are prohibited.
[0200] Tool priority constraint: Tools must be used when they can improve the quality of reasoning. The specific order in which tools are invoked follows the tool invocation order rules given in the tool usage strategy.
[0201] Finite-step completion constraint: The hypothesis reasoning task must be completed within a finite number of rounds, and repeatedly calling the tool while waiting for better evidence is prohibited. After the tool call is completed, a structured chain of evidence (such as JSON formatted text) must be produced directly.
[0202] Based on the dimensions of diagnostic reasoning, the hypothesis reasoning model can perform evidence-based reasoning on diagnostic hypotheses from at least one of the following dimensions: diagnostic gold standard, supporting evidence, rebuttal evidence, input anomaly RAG (Retrieval-Augmented Generation) validation, medical history and risk factors, interpretation of indicator anomalies, time series changes, missing information, danger signs, and confounding factors.
[0203] The gold standard for diagnosis refers to the diagnostic reference standard (such as pathological biopsy, culture, surgical confirmation, specific imaging signs, etc.) that corresponds to the diagnostic hypothesis.
[0204] Supporting evidence refers to medical evidence that supports the diagnostic hypothesis, and can come from retrieved medical knowledge and factual information. Rebuttal evidence refers to medical evidence that contradicts the diagnostic hypothesis, and can also come from retrieved medical knowledge and factual information.
[0205] RAG validation of input anomalies refers to verifying the diagnostic association between each anomaly found in the input factual source information (including but not limited to chief complaint and admission symptoms, positive / abnormal physical examination signs, abnormal laboratory indicators, abnormal imaging findings, and abnormal perceptual model predictions) and the diagnostic hypothesis through at least one RAG search query before performing anomaly classification. The contextual information in the hypothesis reasoning task prompts can also include principles for anomaly classification. For example, if an anomaly is found to have a direct causal / pathophysiological chain with the diagnostic hypothesis, it is classified as supporting; if an anomaly is found to have a contradictory / rejective relationship with the diagnostic hypothesis, it is classified as refuting. Classification as context / background information also requires RAG evidence.
[0206] Medical history and risk factors refer to reviewing the patient's past medical history, family history, exposure history (smoking / drinking / occupational exposure, etc.), medication / surgical history, and their association with the diagnostic hypothesis, and distinguishing between high-risk factors, protective factors, and differential cues.
[0207] Interpreting abnormal indicators refers to predicting the diagnosis of a patient's existing single-point abnormality in laboratory / imaging / sign / perceptual models, and giving the diagnostic significance of the abnormal value itself (such as specificity, sensitivity, positive predictive value, threshold or critical value).
[0208] Time series changes refer to the analysis and clear recording of trends when the source information contains data from multiple time points (such as multiple imaging reports, changes in laboratory indicators, and evolution of medical history).
[0209] Missing information refers to the tests or information that still need to be performed or obtained to confirm or exclude a diagnostic hypothesis. Danger signs refer to information on any dangerous signs that require urgent attention. Confounding factors refer to factors that may lead to misdiagnosis (such as false positives / false negatives).
[0210] The tool usage strategy provides descriptions of available tools / skills, tool invocation rules, clinical tool context, tool invocation limits and stopping rules, etc. (e.g., 1-5 invocations per diagnostic hypothesis). During hypothesis reasoning, if the retrieved evidence is insufficient, LLM can generate new search queries based on the information gaps and continue to invoke tools to retrieve relevant medical evidence; tool invocation can be stopped based on the tool invocation limit and stopping rules. For example, tool invocation stops when the gold standard for the diagnostic hypothesis has been obtained, at least one piece of supporting / refuting / background knowledge evidence, and a key gap in the source information of the facts; tool invocation stops when the currently missing information can only be obtained through new examinations, pathology, follow-up, or manual supplementation of medical history, and cannot be resolved by continuing to search; and tool invocation stops when the maximum number of retrieval tool invocations is reached.
[0211] The clinical tool context refers to the review strategy (such as review dimensions, reading order, and key finding checklist) returned by executing clinical skills tools corresponding to the organ / disease for the current diagnostic hypothesis. If the clinical tool context is empty, it means that there is no matching clinical skills tool for the current diagnostic hypothesis.
[0212] The retry scenario constraint means that even if the current process is in the process of retrying the reasoning of the diagnostic hypothesis, the medical knowledge retrieval tool must be called at least once for the diagnostic hypothesis in this retry, and it is not allowed to rely on the results returned by the retrieval tool in the past reasoning process as a source of evidence.
[0213] The search priority section provides the preferred search tools for different search needs. For example, for tumor diagnosis processes and treatment pathways, guideline search tools (such as clinical guidelines or oncology clinical practice guidelines) can be used first; for disease classifications and definitions, search tools based on the International Classification of Diseases (ICD) can be used first.
[0214] Reasoning requirements refer to the requirements / rules that LLMs must follow during the hypothetical reasoning process. For example, reasoning requirements may include: reasoning should be based solely on known factual sources and retrieved medical evidence, and fabrication is prohibited; for each fact in the input data (such as the patient's basic information and medical history, vital signs and physical examination data, laboratory and pathological examination data) that has a numerical value or is marked with indicators such as increased / decreased / abnormal / positive / significant / lower / higher / not provided, it must appear in the structured chain of evidence and be explicitly labeled as supporting evidence, rebuttal evidence, neutral / background information, gap information, etc.; every sentence containing medical knowledge judgments must be labeled with at least one cited piece of evidence.
[0215] The output format refers to the format and requirements for outputting the structured chain of evidence. For example, the structured chain of evidence can be output in the format of a JSON object; a structured chain of evidence for each diagnostic hypothesis must be output; the citation markers in the structured chain of evidence for each diagnostic hypothesis should reference the retrieved medical evidence, while citations are not required for evidence from factual sources.
[0216] The hypothesis reasoning task prompts used in this embodiment can be designed according to the needs and requirements of evidence-based reasoning for diagnosing hypotheses in actual application scenarios. The structure and specific content of the hypothesis reasoning task prompts are not limited here.
[0217] In this embodiment, the structured chain of evidence for the diagnostic hypothesis includes at least one of the following: a summary of the diagnostic hypothesis, the gold standard for diagnosis, supporting evidence, opposing evidence, neutral / comorbid information (relevant factual evidence), evidence gaps and uncertainties, supplementary examinations and information, etc.
[0218] S204: Invoke the verification model and verify the structured evidence chain of each diagnostic hypothesis from at least one of the following dimensions: reliability of evidence source, consistency of facts, consistency of reasoning logic, and sufficiency of evidence citation.
[0219] In this step, the structured evidence chains of each diagnostic hypothesis are verified from different dimensions. This can be achieved using different verification models, or by using a unified verification model based on different verification rules to verify multiple dimensions.
[0220] In one optional embodiment, different verification dimensions correspond to different verification models. When verifying the structured evidence chain of each diagnostic hypothesis from any target dimension, the LLM is invoked to perform the verification task for the target dimension of the structured evidence chain of the diagnostic hypothesis. In specific implementation, the target dimension verification task prompts configured for the target dimension verification task are obtained. The factual source information, retrieved medical evidence, diagnostic hypotheses and their structured evidence chains, along with the target dimension verification task prompts, are input into the LLM. Based on the prompts from the target dimension verification task prompts, the LLM verifies the structured evidence chain of each diagnostic hypothesis from the target dimension, obtaining the verification result corresponding to each diagnostic hypothesis.
[0221] The structure of verification task prompts across different dimensions is similar, specifically including: role setting, task description, contextual information, output format and requirements, and examples. The contextual information in the verification task prompts for any target dimension may include, but is not limited to, the following information corresponding to that target dimension: verification strategy, soft questions, and output format.
[0222] The reliability of the source of evidence refers to verifying whether the source of the cited evidence is reliable.
[0223] Factual consistency means that the conclusions and descriptions in a structured chain of evidence must be faithful to the information from the source of the input facts (including input clinical data and structured text generated by multimodal perception models based on medical images), and must not deviate from, tamper with, or over-interpret the established facts.
[0224] Logical consistency in reasoning refers to the completeness of the logical chain of a structured evidence chain, the absence of self-contradiction, and the strength of the conclusion not exceeding the safe boundary that the evidence can support.
[0225] Sufficiency of evidence citation refers to the fact that all medical inferences supported by a structured chain of evidence are accompanied by traceable citation tags, so that each assertion can be verified and audited.
[0226] For example, based on the verification task prompts of the fact consistency dimension, the verification model determines that the verification fails when a statement in the structured evidence chain meets the following conditions: 1) The statement conflicts with the numerical values / discoveries in the fact source information. For example, the structured evidence chain states "the lesion is located in the pancreatic body," but the multimodal perception model's CT identification result is that the lesion is located in the pancreatic head; 2) The statement contains patient factual information that does not appear in the input fact source information, such as a fabricated test value, a fictitious past surgical history, or a family history; 3) The deterministic output is distorted: the deterministic output of the multimodal perception model is incorrectly described in a way that cannot be reconciled by clinical interpretation.
[0227] Furthermore, for cases where there are modifiable soft issues within the structured chain of evidence, the validation model can determine that the validation fails and provide actionable modification suggestions. For example, when reasoning makes an absolute assertion about a fact that is neither directly supported nor refuted by evidence (a typical scenario: stating "missing information" as a "fact"), the underlying reasoning is usually reasonable, but the wording is too definitive, leading to a validation failure and the provision of actionable modification suggestions. For instance, it might suggest rewriting "the patient has no history of chronic pancreatitis" as "no record of chronic pancreatitis is found in the existing medical history; a history of chronic pancreatitis is not currently supported."
[0228] In this embodiment, the structure and specific content of the verification task prompts for each dimension can be designed according to the needs and experience of actual application scenarios, and no specific limitations are made here.
[0229] S205: After the structured evidence chain of each diagnostic hypothesis is verified, the ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured evidence chain across diagnostic hypotheses, and generate structured auxiliary diagnostic information.
[0230] In this step, the ranking model can be an LLM (Limited Learning Model). The LLM is invoked to perform a ranking task on multiple diagnostic hypotheses. Specifically, the ranking task prompts are obtained from the configuration. The diagnostic hypotheses, their structured chains of evidence, factual source information, and retrieved medical evidence are input into the LLM along with the ranking task prompts. Based on the prompts, the LLM evaluates and ranks each diagnostic hypothesis and its structured chains of evidence across diagnostic hypotheses, according to the factual source information and retrieved medical evidence, and generates structured auxiliary diagnostic information.
[0231] The sorting task prompts include the corresponding role settings, task description, contextual information, output format and requirements, and examples. The contextual information in the sorting task prompts may include, but is not limited to: sorting principles, output format and requirements, citation conventions, and important constraints. Sorting task prompts can be designed based on specific application scenario requirements and experience; no specific limitations are imposed here.
[0232] Based on the ranking principle, the ranking model can comprehensively evaluate the strength of evidence for each diagnostic hypothesis, placing the more likely hypothesis to be true at the top. When comprehensively evaluating the strength of evidence for each diagnostic hypothesis, the ranking model can consider at least one of the following dimensions:
[0233] 1) The degree of agreement between imaging phenotype and clinical manifestation.
[0234] 2) Strength of mutual corroboration of evidence: Whether multiple independent sources (medical images / laboratory / literature / multimodal perception model) corroborate each other.
[0235] 3) Handling conflicting information: If there is rebuttal evidence that directly contradicts the diagnostic hypothesis, points should be deducted in the ranking and explained in the ranking reason.
[0236] 4) Use hypothetical reasoning model results with caution: When different hypothetical reasoning models give conflicting predictions, the decision should be made based on more robust sources such as imaging phenotypes and clinical evidence.
[0237] 5) Facts take precedence over safety considerations: Imaging phenotypes and other objective clinical facts (laboratory findings, physical signs, medical history, literature evidence) take precedence over the safety heuristic of "malignancy must be ruled out".
[0238] The core objective of the above ranking rules is to prioritize diagnostic hypotheses that are more likely to be the patient's actual diagnosis, and to provide clear reasons based on upstream evidence for each diagnostic hypothesis.
[0239] For example, structured auxiliary diagnostic information can be output as a JSON object. The core objective of this structured definition of the auxiliary diagnostic information output from evidence-based diagnostic reasoning is to present the information in a standardized and traceable format. For example, the structure can be structured around four layers of logic: "overall judgment—bottleneck identification—diagnostic hypothesis ranking—reasoning for ranking." The top layer uses the "overall_judgment" field to summarize the most likely diagnostic direction and its convergence status in one sentence; the "diagnostic_bottleneck" field clearly identifies key information gaps in the current chain of evidence, providing guidance for subsequent checks. The core part is the "ranked_candidates" structure, which ranks the list of diagnostic hypotheses from highest to lowest probability, enabling bidirectional traceability between the conclusion and the structured chain of evidence. Finally, the "ranking_rationale_summary" structure uses three texts to describe the overall diagnostic status, key evidence gaps or points of conflict, and the reasons for ranking each diagnostic hypothesis, ensuring the transparency and auditability of the ranking logic.
[0240] The citation guidelines describe the usage of citations in the output auxiliary diagnostic information. For example, the output text must not contain any citation markers. Important constraints are those that must be followed when generating and outputting auxiliary diagnostic information. For example, auxiliary diagnostic information should only output a JSON object, without a preface or any explanation outside of the JSON; the ranking must be based on evidence in a structured chain of evidence, not on one's own medical knowledge; no new diagnoses should be introduced beyond the input hypotheses; and the core rationale for ranking should be provided for each diagnostic hypothesis, along with the reasons for the ranking.
[0241] For example, Figure 3 An example diagram of auxiliary diagnostic information provided for an exemplary embodiment of this application. In one example, such as... Figure 3As shown, auxiliary diagnostic information may include: evidence-based diagnostic summaries (such as...) Figure 3 Evidence-based diagnostic conclusions shown), diagnostic bottleneck information (such as...) Figure 3 The current diagnostic bottleneck is shown in the diagram, along with the diagnostic hypotheses arranged in sequence and their corresponding confidence levels and reasoning.
[0242] In practical implementation, the evidence-based diagnostic summary can be generated by the ranking model based on the structured evidence chain of each diagnostic hypothesis, extracting common conclusions and dominant evidence. Diagnostic bottleneck information can be obtained by the validation model by summarizing unsatisfied items in the factuality, consistency, logical consistency of reasoning, and sufficiency of evidence citation. The confidence level of each diagnostic hypothesis can be quantified or graded by the ranking model based on the number of pieces of evidence, the strength of evidence, the degree of consistency of evidence, and the degree of conflict between cross-diagnostic hypotheses. The reasoning basis can be extracted from the factual source information, medical evidence, and their reasoning links. The above confidence levels and reasoning basis can be mapped one-to-one with the diagnostic hypotheses and organized according to the ranking results. In practical applications, this output format can also adopt different field mapping methods according to different clinical systems, which is not limited in this application.
[0243] In terms of working principle, the system provides an overall conclusion based on evidence-based diagnosis, then identifies the most needed evidence gaps by using diagnostic bottleneck information, and finally displays the priority of candidate diagnoses and their supporting logic by presenting diagnostic hypotheses, confidence levels, and reasoning bases in sequence. This allows users to directly review diagnostic conclusions, identify weaknesses, and decide whether to continue supplementary examinations or adjust the diagnostic direction accordingly.
[0244] After adopting this method of organizing auxiliary diagnostic information, the output content is no longer limited to a single conclusion, but simultaneously includes a structured chain of evidence expressing summary conclusions, key uncertainties, and candidate diagnostic hypotheses. This makes the diagnostic results more structured, facilitates rapid verification of diagnostic evidence by clinicians, and supports comparative review of different candidate diagnostic hypotheses.
[0245] In this embodiment, the reasoning agent first aggregates structured text and structured clinical data into factual source information for the case, and then inputs it into the hypothesis generation model to output candidate diagnostic directions. Through the hypothesis reasoning model, it retrieves relevant medical evidence for each diagnostic hypothesis and generates a structured evidence chain for each diagnostic hypothesis based on the factual source information and medical evidence. Through the verification model, it verifies the structured evidence chain of each diagnostic hypothesis from at least one dimension, including the reliability of evidence sources, consistency of facts, consistency of reasoning logic, and sufficiency of evidence citation. After the verification is passed, the ranking model compares and ranks each diagnostic hypothesis and its structured evidence chain across diagnostic hypotheses, and outputs structured auxiliary diagnostic information.
[0246] In this embodiment, factual source information serves as a unified input basis, ensuring that subsequent reasoning revolves around the true information of the case. A structured chain of evidence explicitly links diagnostic hypotheses, evidence, and reasoning relationships, facilitating item-by-item verification by the validation model. The ranking model comprehensively compares the strength and consistency of evidence in the context of multiple concurrent diagnostic hypotheses, thereby outputting verifiable auxiliary diagnostic information. Using this approach, diagnostic conclusions are no longer output as a single label, but rather form structured auxiliary diagnostic information corresponding to the factual information and medical evidence of the case, facilitating traceability and verification, and supporting continuous adaptation after model replacement.
[0247] In an optional embodiment, for any diagnostic hypothesis, if the verification of any dimension of the structured evidence chain of the diagnostic hypothesis fails, steps S203-S204 can be re-executed for that diagnostic hypothesis. The hypothesis reasoning model corresponding to the diagnostic hypothesis is invoked to regenerate a new structured evidence chain for the diagnostic hypothesis. The verification model is invoked to verify the new structured evidence chain from at least one dimension, including the reliability of the evidence source, the consistency of facts, the consistency of reasoning logic, and the sufficiency of evidence citation. The verification continues until the new structured evidence chain of the diagnostic hypothesis passes.
[0248] For example, the reasoning agent first inputs any diagnostic hypothesis and its original structured chain of evidence into the verification model. If the verification model outputs a failed verification result and returns the reason for the failure, the reasoning agent inputs this reason for failure as a constraint into the corresponding hypothesis reasoning model. The hypothesis reasoning model can reselect key evidence from the factual source information based on the same diagnostic hypothesis, adjust the order of evidence, retrieve and supplement missing medical evidence, and reconstruct the reasoning relationship, outputting a new structured chain of evidence. This new structured chain of evidence re-enters the verification model, which still verifies it based on at least one dimension of evidence source reliability, factual consistency, logical consistency of reasoning, and sufficiency of evidence citation, and outputs a pass or fail verification result. If it still fails, the corresponding hypothesis reasoning model is called again to correct the new verification feedback until the verification passes.
[0249] This iterative verification mechanism enables the reasoning results of diagnostic hypotheses to be continuously revised in terms of factual citation, evidence organization, and logical links until a structured evidence chain that has been verified is formed, thereby giving the output auxiliary diagnostic information more complete evidence support and more consistent reasoning expression.
[0250] In one optional implementation, for any diagnostic hypothesis, if the verification of the structured evidence chain of the diagnostic hypothesis fails, before calling the hypothesis reasoning model corresponding to the diagnostic hypothesis to regenerate a new structured evidence chain of the diagnostic hypothesis, the reasoning agent constructs an incremental retrieval query for the diagnostic hypothesis based on the verification result of the structured evidence chain of the diagnostic hypothesis; and retrieves relevant medical evidence in a multi-source heterogeneous medical knowledge base based on the incremental retrieval query to obtain incremental retrieval results.
[0251] In practical operation, the verification results output by the verification model can indicate situations such as insufficient evidence, inadequate citation, factual inconsistencies, or broken reasoning chains in the current structured evidence chain. Based on this, the reasoning agent extracts the corresponding evidence gaps and generates incremental search queries containing constraints such as lesion location, imaging features, laboratory indicators, past medical history, or similar cases. These incremental search queries are then sent to multi-source heterogeneous medical knowledge bases for retrieval. During retrieval, the query can be rewritten through entity standardization, synonym expansion, and constraint completion to improve the matching accuracy of relevant medical evidence. The retrieved textual evidence, structured relational evidence, or case fragment evidence are merged into incremental search results. These incremental search results, along with the original factual source information, are then input into the corresponding hypothesis reasoning model to regenerate a structured evidence chain for the new diagnostic hypothesis. In practical applications, when searching in various medical knowledge bases, a combination of vector recall and knowledge graph constraint retrieval can be used; this embodiment does not limit this approach.
[0252] This incremental retrieval result can specifically fill gaps in the chain of evidence, making the new structured evidence chain more sufficient in terms of factual support and citation completeness. Furthermore, it can re-perform the hypothesis reasoning process based on updated medical evidence to regenerate a better structured evidence chain. With this implementation, the reasoning agent can supplement medical evidence before triggering the generation of a new structured evidence chain if the structured evidence chain fails verification. This improves the completeness and consistency of the structured evidence for the diagnostic hypothesis and reduces redundant reasoning and invalid output caused by insufficient original evidence.
[0253] In one possible implementation, the hypothetical reasoning model corresponding to the diagnostic hypothesis is invoked to generate a new structured chain of evidence for the diagnostic hypothesis. This includes: invoking the hypothetical reasoning model corresponding to the diagnostic hypothesis, and performing evidence-based reasoning on the diagnostic hypothesis based on factual source information, medical evidence, and incremental search results to generate a new structured chain of evidence for the diagnostic hypothesis. This implementation combines the input of new evidence (incremental search results) with existing evidence for evidence-based reasoning, allowing the new structured chain of evidence to retain the continuity of the original facts while incorporating new evidence, thus improving the completeness, relevance, and traceability of the structured chain of evidence.
[0254] Based on the foregoing embodiments, in one possible implementation, retrieving relevant medical evidence according to a diagnostic hypothesis includes: converting the diagnostic hypothesis into a diagnostic clinical question; and retrieving relevant medical evidence from a multi-source heterogeneous medical knowledge base based on the diagnostic clinical question.
[0255] In its implementation, the hypothetical reasoning model, upon receiving a diagnostic hypothesis, extracts the disease name, symptom description, anatomical location, temporal evolution characteristics, and differential diagnosis direction from the hypothesis based on PICO paradigm constraints, and reorganizes these into a retrieval-oriented diagnostic clinical question. This diagnostic clinical question can be generated using standard medical terminology, synonym expansion, and combinations of limiting conditions. For example, the hypothesis "consider lung infection" can be transformed into a clinical question focusing on "whether adult lung infections exhibit ground-glass opacities, consolidation, and inflammatory distribution characteristics on imaging." Furthermore, the hypothetical reasoning model performs a search in a multi-source heterogeneous medical knowledge base based on the diagnostic clinical question. The search objects covered by the medical knowledge base can include clinical practice guidelines, evidence-based medicine abstracts, medical textbooks, medical records, and labeled medical case evidence. The model can also recall and rank results according to similar diseases, related symptoms, and evidence levels, thereby outputting a set of medical evidence related to the diagnostic hypothesis.
[0256] In one possible implementation, the diagnostic hypothesis can be converted into a diagnostic clinical question (PICO query) by invoking an LLM to perform a diagnostic clinical question construction task. Specifically, the diagnostic clinical question construction task prompts configured in the task are retrieved, and the diagnostic hypothesis is input into the LLM along with these prompts. The LLM then converts the diagnostic hypothesis into a diagnostic clinical question (PICO query) based on the prompts.
[0257] The prompts for constructing diagnostic clinical questions include the corresponding role settings, task description, contextual information, output format and requirements, and examples. The contextual information may include, but is not limited to: rules for constructing PICO elements, rules for constructing diagnostic clinical questions, and descriptions of PICO elements.
[0258] Based on the rules for constructing PICO elements, LLM can build PICO queries based on a checklist within the clinical tool context of the current diagnostic hypothesis, rather than constructing them out of thin air. Specifically, LLM can extract specific signs as I-elements in the PICO from the review dimensions in the clinical tool context, such as "pancreatic duct truncation" and "peripancreatic fat infiltration," instead of using generalized disease names. For multi-dimensional review scenarios, a diagnostic hypothesis can be broken down into multiple sub-queries corresponding to each dimension, with each sub-query covering different diagnostic dimensions, such as retrieving "diagnostic specificity of pancreatic duct truncation" and "staging significance of peripancreatic vascular invasion" respectively. LLM can prioritize signs marked as highly specific / highly sensitive in the clinical tool context as I-elements for the first round of retrieval. If the clinical tool context includes time-series analysis dimensions (such as "lesion size increased compared to previous images" and "CA19-9 dynamic upward trend"), specific time-varying sub-queries can be constructed.
[0259] Based on the rules for constructing diagnostic clinical questions, LLM requires the following rules when constructing diagnostic clinical questions: 1) Diagnostic clinical questions must cover both the target disease and the list of tests; 2) Diagnostic clinical questions should not be searched broadly using only disease names, but should reflect the patient's context and diagnostic elements; 3) P elements should be extracted from the patient's clinical data and cannot be fabricated; 4) Known risk factors and single abnormalities of the patient should be prioritized as I elements in the PICO; 5) If a diagnostic hypothesis has both risk factors and single abnormalities, the risk factors and single abnormalities can be separated into independent diagnostic clinical questions for separate retrieval, and then the conclusions can be combined in the inference chain; 6) If the threshold or diagnostic definition is unknown, its standard definition should be retrieved first. If there is a lack of clear individual or reference standards for tests, the conclusion should be inconclusive.
[0260] PICO is a commonly used problem-building paradigm in clinical evidence-based medicine. The elements of PICO are: P (Patient / Population): Patient or clinical scenario (automatically extracted from the patient's clinical data); I (Intervention / Indextest): Examination / symptom / sign to be evaluated (specific signs extracted from the examination checklist in the context of clinical tools); C (Comparison / Reference): Reference standard / gold standard (e.g., pathological biopsy); O (Outcome): Diagnostic value / target disease. Using the PICO paradigm, diagnostic hypotheses are transformed into structured diagnostic clinical questions, enabling LLMs to accurately retrieve the most relevant evidence-based evidence from the medical knowledge base for the current case.
[0261] In this embodiment, the structure and specific content of the task prompts for constructing diagnostic clinical questions can be designed according to the actual application scenario, and are not specifically limited here.
[0262] In this process, diagnostic hypotheses are transformed from freely expressed reasoning conclusions into standardized diagnostic clinical questions, establishing a consistent mapping between search criteria and the index fields and semantic tags of the medical knowledge base. The multi-source, heterogeneous medical knowledge base can retrieve supporting or refutational materials related to the same diagnostic hypothesis from different evidence sources. Based on this, the system can obtain more comprehensive and clearly defined medical evidence for subsequent evidence-based reasoning. With this implementation, a stable retrieval correspondence is formed between diagnostic hypotheses and knowledge base content, the retrieval of relevant evidence is more focused, and evidence from different sources can jointly participate in subsequent reasoning, thus making the evidence chain on which the auxiliary diagnosis is based more complete, and the search results easier to verify and trace.
[0263] In one possible implementation, relevant medical evidence is retrieved from a multi-source heterogeneous medical knowledge base based on a diagnostic clinical question. This includes: invoking a retrieval planning model; identifying the retrieval intent based on the diagnostic clinical question and generating tool invocation information corresponding to the retrieval intent, including the identifier of the retrieval tool and retrieval parameters; and, based on the tool invocation information, invoking the corresponding retrieval tool to retrieve relevant medical evidence from the corresponding medical knowledge base.
[0264] Diagnostic clinical questions can be generated by an inference agent or formed from user input through standardization. The retrieval planning model can employ a large language model, a dedicated sequence generation model, or a rule-enhancing model. By invoking the retrieval planning model, it executes retrieval planning tasks for diagnostic clinical questions, identifies retrieval intents, and generates tool invocation information corresponding to those intents.
[0265] For example, the search planning task prompts are obtained from the search planning task configuration; the diagnostic clinical question, factual source information, and existing medical evidence are input into the LLM along with the search planning task prompts, so that the LLM can think based on the prompts of the search planning task prompts, determine the search tools and search parameters to be called (including the identifier of the target medical knowledge base to be matched, such as the knowledge base name), and generate tool call information.
[0266] The retrieval planning task prompts include the corresponding role settings, task description, contextual information, output format and requirements, and examples. The contextual information in the retrieval planning task prompts may include, but is not limited to: a list and description of retrieval intents, a list and description of retrieval tools, descriptions of retrieval parameters, output format, and constraints. Retrieval planning task prompts can be designed based on specific application scenario requirements and experience; no specific limitations are imposed here.
[0267] Search intents can be categorized by department, disease type, organ type, etc., with different search intents corresponding to different categorization tags. During the information retrieval process of the generation tool, a medical knowledge base can be matched based on the search intent. The medical knowledge base can be configured with at least one suitable search intent.
[0268] Furthermore, based on the generated tool invocation information, the retrieval tool corresponding to the identifier of the retrieval tool is invoked to retrieve relevant medical evidence from the corresponding medical knowledge base.
[0269] The retrieval tool can be deployed as a service interface. It can take retrieval parameters as input parameters based on the tool call information, call the corresponding retrieval tool based on the identifier of the retrieval tool to access the corresponding target medical knowledge base, match, sort and extract candidate evidence in the target medical knowledge base, and output a set of medical evidence including evidence identifier, medical evidence fragment, evidence source type and confidence level.
[0270] Furthermore, the sufficiency of evidence assessment model is invoked to perform an sufficiency of evidence assessment task. Specifically, the sufficiency of evidence assessment model is invoked to determine whether the retrieved medical evidence is sufficient to respond to the diagnostic clinical question based on the diagnostic clinical question and the retrieved medical evidence. If the retrieved medical evidence is insufficient to respond to the diagnostic clinical question, a supplementary query question is generated.
[0271] The evidence sufficiency assessment model can be an LLM (Limited Ledger Model). By invoking the LLM, the sufficiency of the retrieved medical evidence is assessed to determine whether it is sufficient to address the diagnostic clinical question. In specific implementation, the evidence sufficiency assessment task prompts configured for the task are obtained. The retrieved medical evidence, the diagnostic clinical question, and the evidence sufficiency assessment task prompts are input into the LLM. Based on the prompts, the LLM determines whether the retrieved medical evidence is sufficient to address the diagnostic clinical question. If the evidence is deemed insufficient (i.e., the retrieved medical evidence is inadequate to address the diagnostic clinical question), a supplementary query question is generated. This supplementary query question is used to retrieve additional medical evidence related to the diagnostic clinical question.
[0272] Furthermore, relevant medical evidence is retrieved from a multi-source heterogeneous medical knowledge base based on supplementary query questions. The specific implementation process is consistent with that for retrieving relevant medical evidence from a multi-source heterogeneous medical knowledge base based on diagnostic clinical questions, and will not be elaborated here.
[0273] After at least one round of the Reasoning and Acting (ReAct) loop framework: diagnostic clinical question → search planning → tool invocation → evidence sufficiency assessment → supplementary query question → search planning → tool invocation → evidence sufficiency assessment → ..., until the evidence is sufficient (i.e. the retrieved medical evidence is sufficient to respond to the diagnostic clinical question) or the number of loops reaches the preset number of loops, the retrieved medical evidence is output.
[0274] In one optional implementation, for retrieved medical evidence, it can be aggregated according to the search intent, and duplicate medical evidence with the same search intent can be removed. For the retained medical evidence, a short summary, source information, clinical task tags, content nature tags, and explanations of why it is important to this case are provided for use in generating a structured chain of evidence.
[0275] In this embodiment, the source range and retrieval conditions of medical evidence are explicitly recorded. Subsequent reasoning can directly reference the retrieved medical evidence, and it is convenient to trace and verify the retrieval basis, thereby improving the targeting, consistency and auditability of multi-source heterogeneous medical knowledge retrieval.
[0276] This ReAct-based RAG retrieval process for diagnostic clinical questions achieves its core technological advantage through a dynamic closed loop of "reasoning-action-evaluation," ensuring both the reliability and security of retrieved medical evidence while maintaining operational efficiency. By leveraging an evidence sufficiency assessment mechanism to initiate supplementary queries when evidence is insufficient, the sufficiency of retrieved medical evidence is ensured, thereby improving the reliability, traceability, and verifiability of the output auxiliary diagnostic information. Simultaneously, by pre-setting the number of iterations, it improves the accuracy of handling complex problems while preventing resource waste from infinite loops, achieving a balance between high reliability, high transparency, and high computational efficiency.
[0277] In one optional embodiment, a RAG module for multimodal medical AI systems is provided, constructing a unified medical knowledge base metadata system to enable medical literature from different sources to follow the same structured process, supporting accurate semantic retrieval and evidence-based diagnostic reasoning. Medical literature from various sources is managed in three layers: "literature → chapter structure → semantic fragments". Within the medical knowledge base, data at these three levels are linked through unique identifiers.
[0278] In this embodiment, semantic fragments in the medical knowledge base are described semantically using the following two orthogonal dimensions:
[0279] 1) Clinical task dimension: Which clinical stage does this semantic fragment serve, such as diagnosis, examination, treatment, prognosis, prevention, or general knowledge; 2) Content nature dimension: What kind of content does this semantic fragment contain, such as standards / thresholds, recommendations, research findings, case descriptions, background knowledge, case narratives, or others.
[0280] The two dimensions mentioned above can be freely combined to express rich semantic scenarios, such as "diagnostic criteria," "treatment recommendations," and "prognostic findings." Based on these two dimensions, clinical task tags and content nature tags are configured for semantic fragments. Clinical task tags correspond to the clinical task dimension and are used to identify the type of clinical task to which the semantic fragment belongs. Content nature tags correspond to the content nature dimension and are used to identify the nature of the content of the semantic fragment. Furthermore, a credibility level can be set for the semantic fragment; the credibility of the semantic fragment is determined based on the credibility of the content in the medical literature to which it belongs.
[0281] For example, the credibility of medical literature can be divided into credibility levels, where different credibility levels correspond to different credibility values. The specific values can be set according to actual application needs, and no specific limitations are made here.
[0282] The retrieval anchors for semantic fragments in the medical knowledge base adopt a multi-retrieval anchor design. The retrieval anchors are not limited to disease names, entity words, or text clues, enabling full-scenario retrieval and query hits from professional doctors to ordinary patients, thereby improving the hit rate.
[0283] In one optional implementation, for any semantic fragment in the medical knowledge base, a multi-granularity index of the semantic fragment is constructed in the medical knowledge base. The multi-granularity index includes at least one granularity index among keywords, phrases, complete sentences, and question sentence templates. In this way, during retrieval, it is not limited to using an index of only one granularity; multiple indexes of different granularities can be used simultaneously to retrieve semantic fragments.
[0284] For example, for any semantic fragment, its core medical concepts, local modifiers, and complete sentence meaning can be determined first, and then corresponding keyword indexes, phrase indexes, whole sentence indexes, and interrogative sentence template indexes can be generated respectively. Keyword indexes include disease names, symptom names, location names, and examination indicator names. Phrase indexes can include medical terms and proper nouns, such as "enlarged hilar lymph nodes" or "unclear boundaries." Whole sentence indexes correspond to the original sentence text, while interrogative sentence template indexes convert the semantic fragment into interrogative expressions such as "Does it exist…?" or "How likely is it…?". Each granularity index is bound to the same semantic fragment identifier and written into the corresponding medical knowledge base. For example, a vector knowledge base can encode indexes of different granularities into vectors, a graph knowledge base can map keywords and phrases to entity nodes, and a case knowledge base can map whole sentence and interrogative sentence template indexes to similar case entries.
[0285] In practical applications, keyword indexing is suitable for quickly identifying specific medical concepts, phrase indexing is suitable for covering local semantic combinations, whole sentence indexing is suitable for locating complete evidence content, and question-template indexing is suitable for matching retrieval requests presented in question form. By adopting multi-granularity indexing, the same semantic fragment can simultaneously support retrieval inputs in multiple expression formats, and complements vector similarity retrieval, graph relation retrieval, and case matching retrieval. This approach makes the storage and retrieval of medical evidence in the medical knowledge base more consistent, and can return evidence fragments with more matching granularity in complex diagnostic scenarios, thereby improving the evidence hit rate and the usability of retrieval results.
[0286] In practical applications, multi-source heterogeneous medical knowledge bases include at least one of the following: vector knowledge bases, graph knowledge bases, and case knowledge bases. Vector knowledge bases store vector representations of semantic segments, their corresponding semantic segment identifiers, and semantic segment content; they are more suitable for semantic similarity retrieval. Graph knowledge bases store relationships between entities such as medical entities, examination indicators, and disease names in medical literature; they are more suitable for relational reasoning retrieval. Case knowledge bases store case information, such as case progression, imaging findings, and diagnostic conclusions; they are more suitable for retrieving related cases (such as similar cases or historical cases).
[0287] For any search query, at least one of these three medical knowledge bases can be selected for simultaneous retrieval, enabling heterogeneous knowledge fusion across multiple medical knowledge bases and improving the hit rate of search queries. The three bases can be deployed individually or jointly and linked through a unified fragment identifier, allowing the same semantic fragment to be retrieved by different search channels.
[0288] In this embodiment, a medical knowledge base can be constructed as follows: For medical literature from different sources, the content of the medical literature is segmented into semantic fragments, and the credibility, clinical task tags, and content nature tags of the semantic fragments are determined. The clinical task tags are used to identify the clinical task type corresponding to the semantic fragment, the content nature tags are used to identify the nature type of the semantic fragment's content, and the credibility is determined based on the credibility of the content of the medical literature to which it belongs.
[0289] Furthermore, semantic fragments are treated as medical knowledge, and their attribute information is stored in a medical knowledge base. The medical knowledge base also stores the bibliographic information and chapter structure information of medical literature. The attribute information of the semantic fragments includes three levels: the literature to which they belong, the chapter to which they belong, and the semantic fragment itself.
[0290] In this implementation, medical literature can come from various sources such as guidelines, consensus statements, clinical research papers, case reports, and testing standards. First, the original medical literature undergoes layout analysis, chapter identification, and paragraph segmentation. Then, the smallest semantic unit capable of independently expressing medical meaning is identified as a semantic fragment. Credibility is assigned based on a comprehensive evaluation of the literature source level, the journal level, the level of evidence, whether it has undergone peer review, and whether there are conflicting conclusions. When a semantic fragment is added to the database, it is simultaneously bound to its parent document identifier, its parent chapter identifier, and its own identifier. The parent document identifier is used to locate the original source, the parent chapter identifier is used to locate its contextual position within the literature, and the semantic fragment is used to form a searchable and traceable knowledge object in the medical knowledge base.
[0291] In practical implementation, medical knowledge bases can adopt a hierarchical storage structure of document index tables, chapter index tables, and fragment index tables. Semantic fragment records are associated with their parent document records through primary keys, which makes it convenient to return the semantic fragment content and its source information at the same time when retrieving medical evidence.
[0292] By segmenting medical literature from different sources into semantic fragments and assigning them credibility, clinical task tags, and content nature tags, the granularity of knowledge in the medical knowledge base is unified. During retrieval, precise positioning can be performed according to task and content nature, and during reasoning, evidence verification can be carried out by combining source hierarchical information, thereby improving the traceability of medical evidence organization and the reliability of diagnostic reasoning.
[0293] In one optional embodiment, after generating auxiliary diagnostic information by reasoning through a reasoning agent based on structured text and structured clinical data, combined with retrieved case-related medical evidence, the process further includes: invoking an examination recommendation model to determine candidate examination items for the next step based on the auxiliary diagnostic information and existing medical evidence; evaluating the marginal information value and cost assessment value of each candidate examination item; and recommending the next examination item based on the marginal information value and cost assessment value of each candidate examination item. Here, the marginal information value characterizes the contribution of the candidate examination item to the diagnostic uncertainty, and the cost assessment value characterizes the cost of the candidate examination item in at least one dimension: cost, risk, and time.
[0294] The examination recommendation model is used to generate examination suggestions based on auxiliary diagnostic information. Its input includes auxiliary diagnostic information obtained from the current reasoning and existing medical evidence related to the current case. The auxiliary diagnostic information specifically includes evidence-based diagnostic summary, diagnostic bottleneck information, and each diagnostic hypothesis arranged in order, along with its corresponding confidence level and reasoning basis. The output is a set of candidate examination items and the corresponding recommendation results.
[0295] In one possible implementation, the inspection recommendation model can encode the correspondence between candidate inspection items and current diagnostic uncertainty, forming a marginal information value assessment result for each candidate inspection item. Simultaneously, it quantifies at least one of the inspection costs, risks, and time consumption for each candidate inspection item, forming a cost assessment value. Marginal information value can be characterized by the change in the diagnostic probability distribution before and after the inspection, the change in the number of distinguishable diagnostic hypotheses, or the decrease in uncertainty entropy. The cost assessment value can be represented by a standardized score with uniform dimensions, or synthesized by weighting cost, risk, and time separately. The inspection recommendation model can employ a combination of rule constraints and a learning-based ranking model to filter out unsuitable inspection items and then output recommendation results based on the comprehensive matching relationship between marginal information value and cost assessment value.
[0296] When recommending the next steps for examination, candidate examinations with higher marginal information value and lower cost assessment values can be prioritized, and the ranking results can be used as a reference for physicians' decision-making. In cases where multiple similar candidate examinations exist, information on indications, contraindications, and restrictions for specific populations from existing medical evidence can be used to further refine the candidate examinations. In this way, the examination recommendation model can supplement the current diagnosis with examinations that are most helpful in narrowing down the diagnostic scope, while also considering implementation costs, thereby outputting actionable suggestions for the next steps.
[0297] In another possible implementation, the examination recommendation model can be an LLM (Limited Learning Model), which executes the examination recommendation task by invoking the LLM. Specifically, the examination recommendation task prompts configured for the task are obtained; auxiliary diagnostic information and existing medical evidence are input into the LLM along with the prompts. Based on the prompts, the LLM determines the next candidate examination items according to the auxiliary diagnostic information and existing medical evidence, evaluates the marginal information value and cost assessment value of each candidate examination item, and recommends the next examination item based on these values.
[0298] The task prompts for inspection recommendations include the corresponding role settings, task description, contextual information, output format and requirements, and examples. The contextual information in the task prompts may include, but is not limited to: selection rules for candidate inspection items, rules for evaluating marginal information value, cost evaluation rules, rules for recommending the next inspection item, output format and requirements, and important constraints. The details of each item in the task prompts can be designed based on specific application scenario requirements and experience; no specific limitations are imposed here.
[0299] With the above settings, after generating auxiliary diagnostic information, it is possible to further complete the examination recommendations for subsequent diagnosis and treatment decisions. This means that the examination suggestions no longer rely solely on single experience judgments, but are comprehensively evaluated by combining diagnostic uncertainty, evidence support, and examination costs, thereby improving the pertinence and interpretability of the next examination recommendation results.
[0300] Building upon the foregoing embodiments, in an optional embodiment, the auxiliary diagnostic information may further include suggestions for further medical attention. An inference agent, based on structured text and structured clinical data, and combined with retrieved case-related medical evidence, generates auxiliary diagnostic information. This further includes: invoking a suggestion generation model to generate suggestions for further medical attention based on structured text, structured clinical data, existing medical evidence, diagnostic hypotheses and their confidence levels, structured evidence chains, and diagnostic bottleneck information. These suggestions for further medical attention include at least one of the following: recommended examinations, treatment and management suggestions, follow-up and monitoring suggestions, and precautions to be informed to the patient.
[0301] In this embodiment, the suggestion generation model can be an LLM (Limited Learning Model), which executes the task of generating next-step medical advice by calling the LLM. Specifically, the suggestion generation task prompts configured for the next-step medical advice generation task are obtained. Structured text, structured clinical data, existing medical evidence, diagnostic hypotheses and their confidence levels, structured evidence chains, and diagnostic bottleneck information are input into the LLM along with the suggestion generation task prompts. Based on the prompts, the LLM generates next-step medical advice for clinicians, according to the structured text, structured clinical data, existing medical evidence, diagnostic hypotheses and their confidence levels, structured evidence chains, and diagnostic bottleneck information.
[0302] The suggested task prompts include the corresponding role settings, task description, contextual information, output format and requirements, and examples. The contextual information in the suggested task prompts may include, but is not limited to, input usage rules, constraints, and output format and requirements. The details of each item in the suggested task prompts can be designed based on specific application scenario requirements and experience; no specific limitations are imposed here.
[0303] By adopting this implementation method, the auxiliary diagnostic information can be expanded from simple auxiliary diagnostic output to suggestion output for subsequent medical behavior. The suggestion content corresponds to the structured evidence chain, diagnostic hypothesis and diagnostic bottleneck information, which can improve the completeness of the result expression and enable doctors to obtain the next examination, treatment, follow-up and patient information more directly.
[0304] Based on the foregoing embodiments, in an optional embodiment, auxiliary diagnostic information is output, which can be implemented as follows: Based on each diagnostic hypothesis and its corresponding confidence level in the auxiliary diagnostic information, the inconsistency score of each diagnostic hypothesis is determined; diagnostic hypotheses with inconsistency scores less than a preset inconsistency threshold are added to the prediction set, wherein the preset inconsistency threshold is obtained by calculating the quantile threshold of the inconsistency score distribution corresponding to the true diagnosis in the calibration dataset based on a conformal prediction method. Further, based on the evidence-based diagnostic summary, diagnostic bottleneck information, and the diagnostic hypotheses, their corresponding confidence levels, and reasoning bases contained in the prediction set, an auxiliary diagnostic report is generated; the auxiliary diagnostic report is then output.
[0305] In the specific implementation, after obtaining the auxiliary diagnostic information, the confidence level of each diagnostic hypothesis is first extracted, and then the inconsistency score of the diagnostic hypothesis is generated according to the preset inconsistency calculation function. The preset inconsistency calculation function can be implemented by inverting the confidence level, weighting the degree of evidence contradiction, or measuring the distance with historical labeled diagnoses. It can also be set according to the actual application scenario, and no specific limitation is made here.
[0306] Subsequently, the preset inconsistency threshold generated during the offline calibration phase is read. This preset inconsistency threshold can be determined as follows: The quantile threshold is obtained by calculating the quantile threshold from the inconsistency score distribution corresponding to the actual diagnoses in the calibration dataset; based on the preset inconsistency calculation function, the inconsistency score distribution corresponding to the actual diagnoses in the calibration dataset is determined, and the preset percentile of the inconsistency scores corresponding to the actual diagnoses in the calibration dataset is determined as the preset inconsistency threshold. The preset percentile can be set according to the actual application scenario requirements and experience; no specific limitations are made here.
[0307] Furthermore, diagnostic hypotheses with inconsistency scores below a preset inconsistency threshold are added to the prediction set as the diagnostic hypotheses to be output. When generating the auxiliary diagnostic report, the evidence-based diagnostic summary is used as the overall conclusion, the diagnostic bottleneck information is used as the source of uncertainty, and the diagnostic hypotheses in the prediction set are organized according to confidence level or ranking results, along with their reasoning basis, and written into the report template to form a structured auxiliary diagnostic report for clinical users / physicians.
[0308] Optionally, during the computer-aided diagnosis generation process, the structured clinical data of the current case (such as patient information, test results, etc.), the structured text output by the multimodal perception model, the diagnostic hypotheses and their structured evidence chains (including the gold standard, supporting evidence, opposing evidence, relevant patient facts (neutral / common), supplementary examinations and information) generated by the reasoning agent during the evidence-based diagnostic reasoning process, auxiliary diagnostic information (including evidence-based diagnostic summary, diagnostic bottleneck information, ranking results, confidence level and reasoning basis of diagnostic hypotheses), and next steps for patient recommendations can all be written into the report template. The output report content can reflect the credibility of the diagnostic hypotheses and the evidence supporting them, thereby improving the interpretability, auditability and clinical usability of the report.
[0309] Building upon the aforementioned embodiments, after outputting auxiliary diagnostic information, the auxiliary diagnostic system can also receive feedback data for suggestions on the next steps of medical treatment. This feedback data may include newly added images and / or newly added clinical data. Furthermore, through a reasoning agent, new auxiliary diagnostic information is generated and output based on the structured text of medical images, structured clinical data, and the structured text and / or newly added clinical data of the newly added images, combined with retrieved case-related medical evidence;
[0310] The structured text of the newly added images can be generated based on the newly added images using a multimodal perception model. For the specific implementation principle, please refer to the aforementioned embodiments, which will not be repeated here.
[0311] In one possible implementation, after providing suggestions for the next medical visit, the assisted diagnostic system continuously monitors the feedback interface or case update interface. When it receives feedback data containing newly added images and / or newly added clinical data, it inputs the newly added images into a multimodal perception model for perception analysis, and the multimodal perception model generates corresponding structured text. The structured text of the newly added images, along with the structured text of the original medical images and the structured clinical data, are then sent to the reasoning agent. The reasoning agent retrieves medical evidence related to the current case from a case-related medical knowledge base, and under the constraints of this medical evidence, reconstructs diagnostic hypotheses, a structured chain of evidence, and an evidence-based diagnostic summary, thereby generating new assisted diagnostic information for display or feedback.
[0312] If the feedback data only contains new clinical data, the new clinical data, along with the structured text of the original medical images and the structured clinical data, are sent to the reasoning agent. The reasoning agent then performs an evidence-based diagnostic reasoning process based on the new clinical data to generate new auxiliary diagnostic information, which is then displayed or returned. If the feedback data contains both new images and new clinical data, both are used together as new factual source information. The structured text of the new images, the new clinical data, along with the structured text of the original medical images and the structured clinical data, are sent to the reasoning agent. The reasoning agent then performs an evidence-based diagnostic reasoning process based on the structured text of the new images and the new clinical data to generate new auxiliary diagnostic information, which is then displayed or returned, thus forming a closed-loop update.
[0313] This mechanism enables newly added images and clinical data to enter the subsequent reasoning stages of the existing auxiliary diagnostic chain. A multimodal perception model maintains the structured representation of the imaging evidence, while an inference agent completes evidence fusion and conclusion updates. Therefore, auxiliary diagnostic information matching the latest condition can be generated as case information is continuously supplemented. Because the update process is still based on case-related medical evidence for reasoning, the new auxiliary diagnostic information is continuous and traceable, and can reflect changes after patient follow-up or re-examination.
[0314] For example, Figure 4 This is a framework diagram of a computer-aided diagnostic system provided in an embodiment of this application. Figure 4 As shown, this computer-aided diagnostic system includes a perception module, an inference module, a calibration module, a human-computer interaction module, and a RAG module. In the computer-aided diagnostic process, medical images and structured clinical data of the case to be diagnosed are first acquired. The perception module (multimodal perception model) generates structured text (including lesion localization information, sign descriptions, and their uncertainties) based on the medical images. The structured text and clinical data are then input into the inference module (inference agent).
[0315] like Figure 4 As shown, the inference module can use the retrieval enhancement capabilities provided by the RAG module to retrieve case-related medical evidence. The inference module performs evidence-based diagnostic inference by calling a replaceable large model, that is, it infers from structured text and structured clinical data, combined with the retrieved case-related medical evidence, to generate auxiliary diagnostic information.
[0316] The calibration module calibrates the auxiliary diagnostic information and generates an auxiliary diagnostic report containing cited evidence. The human-computer interaction module outputs the auxiliary diagnostic report for user viewing.
[0317] The human-computer interaction module also has a feedback function, which can receive user feedback data (including newly added images and / or newly added clinical data) and provide the feedback data to the perception module and / or reasoning module in the computer-aided diagnostic system to continue evidence-based diagnostic reasoning and update the reasoning results.
[0318] In this embodiment, the specific implementation principle of the computer-aided diagnostic system is the same as that in the previous embodiment, and will not be repeated here.
[0319] In an optional embodiment, the computer-aided diagnostic system may further include an evaluation module configured to: determine the upper bound of the performance of the inference module given a reference structured text, and determine the performance degradation of the inference module given a structured text generated by the perception module, so as to attribute errors between the perception module and the inference module.
[0320] In its specific implementation, the computer-aided diagnosis method further includes: using a benchmark evaluation set containing historical images of cases and their corresponding structured reference texts to evaluate the perception error of the multimodal perception model and the performance ceiling of the inference agent; determining the actual performance of the inference agent under the current perception conditions based on the auxiliary diagnostic information output by the inference agent, wherein the current perception conditions include the structured text output by the multimodal perception model; using the difference between the performance ceiling and the actual performance as the performance degradation of the inference agent under the current perception conditions; performing error attribution on the multimodal perception model and the inference agent based on the perception error and the performance degradation; and outputting the error attribution results.
[0321] The benchmark assessment set also includes other information such as structured clinical data of cases, which can be set according to actual application needs, and no specific limitations are made here.
[0322] In practical implementation, the structured reference text corresponding to historical images can be input into the reasoning agent. The reasoning agent generates first auxiliary diagnostic information based on the structured reference text, and the quality score of the first auxiliary diagnostic information is evaluated to characterize the performance of the reasoning agent under the reference structured text condition, serving as the upper limit of the reasoning agent's performance. The evaluation of the quality score of the first auxiliary diagnostic information can be implemented according to preset diagnostic evaluation rules or using a pre-trained evaluation model; no specific limitations are made here.
[0323] Historical imagery is input into a multimodal perception model to obtain the corresponding first structured text. Then, errors down to one dimension are calculated based on the first structured text and a structured reference text. These errors include, but are not limited to: calculating the mean squared error (MSE) and mean absolute error (MAE) based on the probability distributions of the first and structured reference texts; and calculating the cosine distance based on the corresponding vectors of the first and structured reference texts. The errors down to one dimension are then weighted and fused to obtain the perception error of the multimodal perception model. The weighting coefficients for the errors in different dimensions can be set according to the actual application scenario and experience; no specific limitations are imposed here.
[0324] In addition, the perception error of the multimodal perception model can be evaluated based on the preset comparison evaluation rules between the generated structured text and the reference structured text. The specific rules can be designed according to the actual application scenario, and no specific limitations are made here.
[0325] After generating second auxiliary diagnostic information in response to a case's auxiliary diagnostic request, the quality score of this second auxiliary diagnostic information is evaluated to characterize the actual performance of the reasoning agent under the current perception conditions. The current perception conditions include the structured text output by the multimodal perception model. Further, the difference between the upper limit of the reasoning agent's performance and its actual performance under the current perception conditions is calculated to obtain the performance degradation of the reasoning agent under these conditions. The perception error of the multimodal perception model and the performance degradation of the reasoning agent under the current perception conditions are jointly analyzed. When the perception error is large and the performance degradation increases simultaneously, the error can be mainly attributed to the multimodal perception model; when the perception error is small but the performance degradation is still significant, the error can be mainly attributed to the reasoning agent. The error attribution results can be output as a text report, structured labels, or numerical scores and written to the system log or diagnostic evaluation record.
[0326] This approach enables the perception error of a multimodal perception model and the performance degradation of the inference agent to be quantified separately on the same benchmark set, and establishes a correlation through structured text as an intermediate representation, thereby achieving targeted identification of system bottlenecks. By outputting error attribution results, it can provide a basis for model replacement, parameter updates, and system optimization, and improve the measurability and traceability of the assisted diagnostic system.
[0327] In one optional embodiment, after the inference model used by the inference agent changes, the same benchmark set can be used to evaluate the first performance limit of the original inference agent before the inference model change, and the second performance limit of the new inference agent after the inference model change; based on the first performance limit and the second performance limit, it is determined whether the performance of the new inference agent has degraded, and the determination result is output.
[0328] In practical implementation, when the inference model used by the inference agent changes, the system calls the same benchmark set to conduct closed-loop evaluations on both the original and new inference agents. During the evaluation, the input samples, output format, and scoring rules are fixed, ensuring that the two evaluations are only affected by the difference in the inference model used. The calculation methods for the first and second performance limits are described in the aforementioned embodiments and will not be repeated here.
[0329] Optionally, after obtaining two sets of performance limits, the system compares the second performance limit with the first performance limit and calculates the difference between the two. If the second performance limit is lower than the first performance limit, and the difference between the second and first performance limits is greater than a preset difference threshold, then the system determines that the performance of the new inference agent has degraded. If the second performance limit is not lower than the first performance limit, or the difference between the second and first performance limits is less than or equal to the preset difference threshold, then the system determines that the performance of the new inference agent has not degraded.
[0330] Optionally, the second performance upper limit is compared with the first performance upper limit, and the difference between the second performance upper limit and the first performance upper limit is calculated. If the second performance upper limit is lower than the first performance upper limit, and the ratio of the difference between the second performance upper limit and the first performance upper limit to the first performance upper limit is greater than a preset ratio threshold, then the performance of the new inference agent is determined to have degraded. If the second performance upper limit is not lower than the first performance upper limit, or the ratio of the difference between the second performance upper limit and the first performance upper limit to the first performance upper limit is less than or equal to the preset ratio threshold, then the performance of the new inference agent is determined not to have degraded.
[0331] Furthermore, the judgment rules for determining whether the performance of the new inference agent has degraded based on the first and second performance limits can be set and adjusted according to the actual application scenario, and are not specifically limited here.
[0332] This mechanism compares and evaluates the performance limits of the inference agent before and after the inference model replacement using the same benchmark set, ensuring that the performance before and after the replacement is under a unified standard. The evaluation results are then directly used to determine whether the new inference agent has experienced performance degradation. This approach limits the impact of model updates to a comparable range and outputs clear judgment results, thus providing a basis for inference model replacement verification, version rollback, and deployment review.
[0333] In an optional embodiment, based on the computer-aided diagnostic method / system provided in the foregoing embodiments, the reasoning module can be implemented as a multi-specialty consultation subsystem, thereby realizing a computer-aided diagnostic method / system for multi-specialty consultation.
[0334] For example, Figure 5 This is a framework diagram of the multi-specialty consultation subsystem provided in an embodiment of this application. Figure 5As shown, the multi-specialty consultation subsystem includes specialist perspective units of various specialty types and cross-specialty comprehensive modules. Each specialist perspective unit can be a hypothetical reasoning model corresponding to that specialty type. Figure 5 The document only showcases specialty perspective units for three specialties: "cardiovascular," "respiratory," and "oncology." In practical applications, more specialty perspective units for other specialties can be constructed, without specific limitations here.
[0335] In practical applications, hypothetical reasoning models corresponding to different specialty types can be implemented using a unified base model (such as LLM). The hypothetical reasoning task prompts and available search tools differ for each specialty type. These prompts require the LLM to perform hypothetical reasoning based on the diagnostic logic and medical evidence of the corresponding specialty, enabling the LLM to analyze information from a unified source of facts (including structured text and structured clinical data) from the perspectives of different specialty systems. The hypothetical reasoning task prompts and available search tools for different specialty types can be designed and configured according to actual application scenarios and experience; no specific limitations are imposed here.
[0336] Specialty perspective units of any specialty type can be implemented by loading the corresponding hypothesis reasoning task prompts and available tools for the specialty type through LLM. Cross-specialty integration modules can be implemented using ranking models, which reconcile conflicts between analyses of different specialty perspective units, identify influence relationships across different physiological systems, and generate auxiliary diagnostic information for multi-specialty consultations.
[0337] In one possible implementation, for any diagnostic hypothesis, the hypothesis reasoning model corresponding to the diagnostic hypothesis is invoked to perform evidence-based reasoning on the diagnostic hypothesis based on factual source information and medical evidence, generating a structured chain of evidence for the diagnostic hypothesis, including:
[0338] Identify the specialty type to which the diagnostic hypothesis belongs; based on the specialty type to which the diagnostic hypothesis belongs, call the hypothesis reasoning model corresponding to the specialty type; through the hypothesis reasoning model corresponding to the specialty type, retrieve relevant specialty medical evidence based on the diagnostic hypothesis and its specialty type; and perform evidence-based reasoning on the diagnostic hypothesis based on factual source information and specialty medical evidence to generate a structured evidence chain for the diagnostic hypothesis.
[0339] In one possible implementation, upon receiving any diagnostic hypothesis, the system first extracts the disease name, organ location, symptoms and signs, examination items, or diagnostic terms from the hypothesis. Then, based on preset specialty mapping rules, it performs specialty mapping to determine the specialty type to which the diagnostic hypothesis belongs. These preset specialty mapping rules can be set according to the actual application scenario requirements and are not specifically limited here. Specialty types may include, but are not limited to, cardiology, respiratory medicine, gastroenterology, neurology, orthopedics, obstetrics and gynecology, or radiology.
[0340] In another possible implementation, a pre-trained specialty classification model can be used to classify and identify the specialty type to which the diagnostic hypothesis belongs. The specialty classification model can be any text classification model, trained using a training set containing the diagnostic conclusion and its corresponding specialty classification.
[0341] Furthermore, the system constructs a retrieval constraint by combining the diagnostic hypothesis and the specialty type. It then retrieves relevant guidelines, consensus statements, clinical pathways, case evidence, or diagnostic criteria fragments related to the diagnostic hypothesis from the medical knowledge base, and organizes the retrieval results into specialty medical evidence. This process differs from the previous embodiment's retrieval of medical evidence related to the diagnostic hypothesis from the medical knowledge base in that it adds the constraint of matching the retrieval results with the corresponding specialty type. When recalling retrieval results, medical evidence corresponding to the specialty type is recalled as specialty medical evidence. Otherwise, the retrieval process is similar; please refer to the previous embodiment for details, which will not be repeated here.
[0342] Furthermore, a corresponding hypothetical reasoning model is selected based on the specialty type, such as a cardiovascular specialty model, a respiratory specialty model, or a gastroenterology specialty model. The factual source information and specialty medical evidence are then input into the hypothetical reasoning model corresponding to that specialty type to generate evidence-based reasoning results and a structured chain of evidence for the diagnostic hypothesis. The specific implementation principle is similar to the hypothetical reasoning model in the aforementioned embodiments and will not be repeated here.
[0343] In practical applications, the identification of specialty types enables the same diagnostic hypothesis to enter the matching evidence retrieval channel. The introduction of specialty medical evidence limits the reasoning basis to the diagnostic norms of that specialty, while the hypothesis reasoning model of the corresponding specialty integrates evidence and generates conclusions according to the knowledge expression mode of that specialty, thus forming a structured evidence chain with clear hierarchy and traceable source.
[0344] By adopting this approach, diagnostic hypotheses can be processed in a one-to-one correspondence with specialist evidence and specialist reasoning models, reducing evidence bias caused by general reasoning. The generated structured evidence chain is more consistent with specialist diagnostic logic and facilitates subsequent verification, auditing, and output of diagnostic results.
[0345] Furthermore, the ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured evidence chain across diagnostic hypotheses. This assesses the contradictions between the structured evidence chains of each diagnostic hypothesis and the cross-specialty physiological system influence relationships, generating auxiliary diagnostic information for multi-specialty consultations. In this embodiment, the ranking model used in the multi-specialty consultation scenario can be consistent with the ranking model used in the aforementioned embodiments. The specific implementation principle is detailed in the aforementioned embodiments and will not be repeated here.
[0346] In an optional embodiment, in a multi-specialty consultation scenario, specially configured ranking task prompts for multi-specialty consultations can be used. These prompts may differ from those used in the aforementioned embodiments and can be designed and adjusted according to the cross-specialty aggregation requirements of multi-specialty consultations; no specific limitations are imposed here. In this embodiment, for conflicts existing in the structured evidence chains of diagnostic hypotheses from different specialty types, the ranking model can further align and compare abnormal indicators, imaging signs, and medical history descriptions from different specialties within the same case to identify whether there are mutually contradictory or complementary relationships, and mark the conflict type as evidence contradiction, temporal contradiction, or systemic correlation contradiction. For cross-specialty systemic influence relationships, the ranking model can analyze the direction, intensity, and propagation path of the influence of a certain specialty abnormality on the diagnostic hypothesis of another specialty, thereby forming a cross-specialty influence chain.
[0347] The ranking model outputs auxiliary diagnostic information for multi-specialty consultations, which may include a ranking list of diagnostic hypotheses, a summary of evidence for each diagnostic hypothesis, results of conflicting evidence comparisons, and a description of cross-specialty influence relationships. This auxiliary diagnostic information is organized into a preset output format (such as Markdown) for direct display on the consultation interface. If a diagnostic hypothesis has both high strength of evidence and strong cross-specialty influence, the system can retain its priority in the output. If there are significant contradictions among multiple structured chains of evidence, the source of the contradiction and its corresponding specialty will be marked to facilitate rapid verification by the consulting physicians.
[0348] For example, in one instance, the core summary of the structured evidence chain corresponding to cardiology is: the patient has heart failure and requires cardiotonics, diuretics, and vasodilators. The core summary of the structured evidence chain corresponding to nephrology is: the patient has renal insufficiency (elevated creatinine), and requires protection of renal function, with caution in the use of certain diuretics.
[0349] Based on this example, we analyze the contradictions between the structured chains of evidence for each diagnostic hypothesis: strong diuresis may worsen kidney damage, while protecting kidney function may lead to worsening heart failure. This is a direct conflict between two specialist perspectives.
[0350] Interdisciplinary systemic effects refer to the direct or indirect functional or structural impacts of pathological changes (or therapeutic interventions) in a physiological system governed by one specialty on physiological systems governed by one or more other specialties. Based on this example, we analyze the interdisciplinary systemic effects: a complete vicious cycle of neurohumoral regulation: "decreased cardiac output → reduced renal perfusion → decreased glomerular filtration rate → water and sodium retention → increased cardiac preload."
[0351] The resulting multi-specialty consultation auxiliary diagnostic information can clearly describe this cross-system logic, allowing cardiologists and nephrologists to understand and accept a joint approach based on global homeostasis, rather than a single-specialty solution.
[0352] Through the above methods, the ranking model can establish a unified comparative framework among diagnostic hypotheses and incorporate contradictions in structured evidence chains and cross-specialty systemic influences into the same ranking process, enabling the output to cover the relevant information required for multi-specialty consultations. Therefore, auxiliary diagnostic information is no longer limited to a single diagnostic conclusion, but presents hypothesis priorities, conflicting relationships, and systemic linkages in a structured form, facilitating collaborative interpretation by multiple specialists and the formation of a consistent diagnostic opinion.
[0353] Figure 6 This is a schematic diagram of the structure of a computer-aided diagnostic device provided as an exemplary embodiment of this application. Figure 6 As shown, the computer-aided diagnostic device 60 includes: a data receiving module 601, a sensing module 602, an inference module 603, and an output module 604.
[0354] The data receiving module 601 is used to acquire medical images and structured clinical data of a case in response to a request for case-assisted diagnosis.
[0355] The perception module 602 is used to generate structured text containing lesion location information, sign descriptions and their uncertainties based on medical images through a multimodal perception model.
[0356] The reasoning module 603 is used to generate auxiliary diagnostic information by reasoning based on structured text and structured clinical data, combined with retrieved medical evidence related to the case, through a reasoning agent. The reasoning agent supports the replacement of the model.
[0357] Output module 604 is used to output auxiliary diagnostic information.
[0358] In one optional embodiment, the multimodal perception model includes a multimodal perception module and a transformation module. When generating structured text containing lesion location information, sign descriptions, and their uncertainties based on medical images using the multimodal perception model, the perception module is further configured to:
[0359] Medical images are input into the multimodal perception module of the multimodal perception model to generate a text representation of structured text;
[0360] The conversion module converts the text representation of structured text into text, thus obtaining structured text.
[0361] In an optional implementation, the computer-aided diagnostic device further includes a multimodal perception model training module for implementing the training process of the multimodal perception model, specifically including:
[0362] Obtain historical images and corresponding structured annotation text. The structured annotation text contains lesion location information, sign descriptions and their uncertainties in the historical images.
[0363] Historical images are input into the multimodal perception module of the multimodal perception model to generate a predicted representation of structured labeled text;
[0364] The text representation module converts structured labeled text into text representation.
[0365] Based on the predicted representation and the text representation of the structured labeled text, the parameters of the multimodal perception module are adjusted to obtain the trained multimodal perception module.
[0366] In an optional implementation, when generating structured text containing lesion location information, sign descriptions, and their uncertainties based on medical images using a multimodal perception model, the perception module is also used to:
[0367] A multimodal perception model is used to perform multi-step iterative observation of medical images, one of which includes:
[0368] Action decisions are generated based on the current observation status of medical images. The action decisions include at least one of the following operations: adjusting image display parameters, zooming in on a specified area, switching image layers, and retrieving historical images for comparison.
[0369] The current observation state is updated based on the action decision, and a continuous representation is generated based on the updated current observation state;
[0370] When the termination condition is met, a text representation of the structured text is generated based on the continuous representations obtained from multi-step iterative observations.
[0371] In one optional embodiment, the multimodal perception model includes a multimodal perception model corresponding to at least one type of medical image, including: radiological images, pathological images, ultrasound images, nuclear medicine images, and endoscopic images.
[0372] In one alternative embodiment, the inference agent has the ability to use models and invoke tools, and the inference agent supports replacing the models used. The tools that the inference agent can invoke include at least one of the following: data preprocessing and measurement tools, anomaly detection tools, retrieval tools, and artificial intelligence models.
[0373] When the reasoning agent infers and generates auxiliary diagnostic information based on structured text and structured clinical data, combined with retrieved case-related medical evidence, the reasoning module is also used for:
[0374] Through reasoning agents, structured text and structured clinical data are used as factual sources of information for cases;
[0375] Invoke the hypothesis generation model to generate at least one diagnostic hypothesis based on factual source information;
[0376] For any diagnostic hypothesis, the hypothesis reasoning model is invoked to retrieve relevant medical evidence based on the diagnostic hypothesis, and evidence-based reasoning is performed on the diagnostic hypothesis based on factual source information and medical evidence to generate a structured evidence chain for the diagnostic hypothesis.
[0377] The verification model is invoked to verify the structured evidence chain of each diagnostic hypothesis from at least one dimension: reliability of evidence source, consistency of facts, consistency of reasoning logic, and sufficiency of evidence citation.
[0378] After the structured evidence chains of each diagnostic hypothesis are verified, the ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured evidence chains across diagnostic hypotheses, and to generate structured auxiliary diagnostic information.
[0379] In an optional embodiment, the computer-aided diagnostic device further includes a verification module for:
[0380] For any diagnostic hypothesis, if the structured chain of evidence for the diagnostic hypothesis fails to be verified, the hypothesis reasoning model corresponding to the diagnostic hypothesis is invoked to generate a new structured chain of evidence for the diagnostic hypothesis.
[0381] The verification model is invoked to verify the new structured chain of evidence from at least one dimension: reliability of evidence source, consistency of facts, consistency of reasoning logic, and sufficiency of evidence citation.
[0382] Until a new, structured chain of evidence is validated for the diagnostic hypothesis.
[0383] In an optional embodiment, the verification module is further configured to: for any diagnostic hypothesis, if the verification of the structured evidence chain of the diagnostic hypothesis fails, call the hypothesis reasoning model corresponding to the diagnostic hypothesis to generate a new structured evidence chain of the diagnostic hypothesis, before constructing an incremental retrieval query for the diagnostic hypothesis based on the verification result of the structured evidence chain of the diagnostic hypothesis; and retrieve relevant medical evidence in a multi-source heterogeneous medical knowledge base based on the incremental retrieval query to obtain incremental retrieval results.
[0384] In an optional embodiment, the computer-aided diagnostic device further includes a RAG module for: retrieving relevant medical evidence based on diagnostic hypotheses, specifically including:
[0385] Transform diagnostic hypotheses into diagnostic clinical questions;
[0386] Relevant medical evidence is retrieved from multi-source heterogeneous medical knowledge bases based on diagnostic clinical questions.
[0387] In an optional embodiment, when retrieving relevant medical evidence from a multi-source heterogeneous medical knowledge base based on a diagnostic clinical question, the RAG module is further configured to:
[0388] The retrieval planning model is invoked to identify the retrieval intent based on the diagnostic clinical question and generate tool invocation information corresponding to the retrieval intent. The tool invocation information includes the identifier of the retrieval tool and the retrieval parameters.
[0389] Based on the tool call information, the corresponding search tool is invoked to retrieve relevant medical evidence from the corresponding medical knowledge base.
[0390] In an optional embodiment, the RAG module is further configured to:
[0391] Based on the tool call information, the corresponding retrieval tool is invoked to retrieve relevant medical evidence from the corresponding medical knowledge base. Then, the evidence sufficiency assessment model is invoked to determine whether the retrieved medical evidence is sufficient to respond to the diagnostic clinical question. If the retrieved medical evidence is insufficient to respond to the diagnostic clinical question, a supplementary query question is generated. Based on the supplementary query question, relevant medical evidence is retrieved from multi-source heterogeneous medical knowledge bases.
[0392] In an optional embodiment, when invoking the hypothesis generation model to generate at least one diagnostic hypothesis based on fact source information, the inference module is further configured to:
[0393] The hypothesis generation model is invoked, and a retrieval tool is used to search for relevant medical evidence in a multi-source heterogeneous medical knowledge base based on fact source information. Based on the fact source information and medical evidence, at least one diagnostic hypothesis is generated.
[0394] In an optional embodiment, the inference module is further configured to:
[0395] The examination recommendation model is invoked to determine the next candidate examination items based on auxiliary diagnostic information and existing medical evidence. The marginal information value and cost evaluation value of each candidate examination item are evaluated, and the next examination item is recommended based on the marginal information value and cost evaluation value of each candidate examination item.
[0396] Among them, the marginal information value is used to characterize the contribution of candidate examination items to diagnostic uncertainty, and the cost assessment value is used to characterize the cost of candidate examination items in at least one dimension of cost, risk, and time.
[0397] In one optional embodiment, the auxiliary diagnostic information includes:
[0398] The report summarizes the evidence-based diagnosis, identifies diagnostic bottlenecks, and presents a list of diagnostic hypotheses, their corresponding confidence levels, and the rationale behind them.
[0399] In an optional embodiment, the auxiliary diagnostic information may further include suggestions for further medical attention. When generating auxiliary diagnostic information by reasoning through a reasoning agent based on structured text and structured clinical data, combined with retrieved case-related medical evidence, the reasoning module is also used for:
[0400] The suggestion generation model is invoked to generate next-step medical advice based on structured text, structured clinical data, existing medical evidence, diagnostic hypotheses and their confidence levels, structured evidence chains, and diagnostic bottleneck information.
[0401] The next steps for medical attention should include at least one of the following: recommended examinations, treatment and management recommendations, follow-up and monitoring recommendations, and precautions to be informed to the patient.
[0402] In an optional embodiment, the data receiving module is further configured to: receive feedback data for next-step medical advice, the feedback data including newly added images and / or newly added clinical data.
[0403] The reasoning module is also used to: generate new auxiliary diagnostic information by reasoning through a reasoning agent based on the structured text of medical images, structured clinical data, and the structured text and / or new clinical data of newly added images, combined with retrieved case-related medical evidence; wherein, the structured text of newly added images is generated based on the newly added images through a multimodal perception model.
[0404] The output module is also used to output new auxiliary diagnostic information.
[0405] In an optional embodiment, the output module is further configured to:
[0406] Based on each diagnostic hypothesis and its corresponding confidence level in the auxiliary diagnostic information, determine the inconsistency score of each diagnostic hypothesis;
[0407] Diagnostic hypotheses with inconsistency scores less than a preset inconsistency threshold are added to the prediction set. The preset inconsistency threshold is obtained by calculating the quantile threshold of the inconsistency score distribution corresponding to the real diagnosis in the calibration dataset based on the conformal prediction method.
[0408] Based on the evidence-based diagnostic summary, diagnostic bottleneck information, and the diagnostic hypotheses contained in the prediction set, along with their corresponding confidence levels and reasoning, an auxiliary diagnostic report is generated.
[0409] Output auxiliary diagnostic reports.
[0410] In an optional embodiment, when, for any diagnostic hypothesis, the reasoning module is used to generate a structured chain of evidence for the diagnostic hypothesis by invoking the hypothesis reasoning model, retrieving relevant medical evidence based on the diagnostic hypothesis, and performing evidence-based reasoning on the diagnostic hypothesis based on factual source information and medical evidence, the reasoning module is further configured to:
[0411] Identify the specialty to which the diagnostic hypothesis belongs;
[0412] Based on the specialty type to which the diagnostic hypothesis belongs, the hypothesis reasoning model corresponding to the specialty type is invoked. Through the hypothesis reasoning model corresponding to the specialty type, relevant specialty medical evidence is retrieved based on the diagnostic hypothesis and its specialty type. Based on the factual source information and specialty medical evidence, the diagnostic hypothesis is reasoned on an evidence-based basis to generate a structured evidence chain for the diagnostic hypothesis.
[0413] In an optional embodiment, when invoking the ranking model to evaluate and rank each diagnostic hypothesis and its structured chain of evidence across diagnostic hypotheses, and to generate structured auxiliary diagnostic information, the inference module is further configured to:
[0414] The ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured evidence chain across diagnostic hypotheses. The contradictions between the structured evidence chains of each diagnostic hypothesis and the cross-specialty system influence relationship are evaluated, and auxiliary diagnostic information for multi-specialty consultation is generated.
[0415] In an optional embodiment, an evaluation module is further included, for:
[0416] Using a benchmark set containing historical images and their corresponding structured reference texts, we evaluate the perceptual error of the multimodal perception model and the performance ceiling of the inference agent.
[0417] Based on the auxiliary diagnostic information output by the reasoning agent, determine the actual performance of the reasoning agent under the current perception conditions, which include the structured text output by the multimodal perception model.
[0418] The difference between the performance ceiling and the actual performance is used as the performance degradation of the inference agent under the current perception conditions.
[0419] Error attribution is performed on the multimodal perception model and the reasoning agent based on perception error and performance degradation.
[0420] Output the results of error attribution.
[0421] In an optional embodiment, the evaluation module is further configured to:
[0422] After the reasoning model used by the reasoning agent changes, the same benchmark set is used to evaluate the first performance limit of the original reasoning agent before the reasoning model change, and the second performance limit of the new reasoning agent after the reasoning model change.
[0423] Based on the first and second performance limits, determine whether the performance of the new inference agent has degraded, and output the result.
[0424] The computer-aided diagnostic device provided in this application specifically implements the computer-aided diagnostic method provided in the foregoing embodiments. For the specific implementation principle and technical effects, please refer to the relevant content of the foregoing embodiments, which will not be repeated here.
[0425] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device of this embodiment may include: at least one processor 1001; and a memory 1002 communicatively connected to the at least one processor. Optionally, the electronic device may also include a communication component 1003. The processor 1001, memory 1002, and communication component 1003 are connected via a bus 1004.
[0426] The memory 1002 stores instructions that can be executed by at least one processor 1001 to cause the electronic device to perform the method as described in any of the above embodiments.
[0427] Alternatively, the memory 1002 can be either standalone or integrated with the processor 1001.
[0428] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the foregoing embodiments, and will not be repeated here.
[0429] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0430] This application provides a computer-aided diagnostic system, including a client and the aforementioned electronic device. The client is used to send a case-aided diagnostic request to the electronic device and to provide the electronic device with medical images and structured clinical data of the case. The client can be a browser, application software, or a smart client, diagnostic terminal, etc., and is not specifically limited herein.
[0431] Figure 8 This is a structural block diagram of a computing device according to an embodiment of this application. Figure 8 As shown, the computing device 2000 may include one or more (only one is shown in the figure) processors 2001 and memory 2002. The memory 2002 stores computer programs / instructions, and the processor 2001 executes the computer programs / instructions. When the computer programs / instructions are executed by the processor 2001, they implement the technical solutions provided in any of the aforementioned method embodiments. Their specific functions and the technical effects they can achieve are similar and will not be repeated here.
[0432] The aforementioned computing device can be understood as an integrated smart terminal, including but not limited to servers, desktop computers, PCs (Personal Computers), all-in-one model machines, mobile phones, tablets, or other portable smart terminals, and the computing device may have the model in the above embodiments of this application pre-installed.
[0433] Specifically, this computing device can pre-install various types of models, including but not limited to models in natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thus providing diverse model selection. In different product forms, this computing device can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference, and application. In some product forms, this computing device also supports model management, including but not limited to multi-type model management (supporting the management of discriminative, generative, and other model types), model version control (supporting the control of different model versions), and model evaluation (evaluating model performance and effectiveness based on model evaluation tools). In other product forms, this computing device can also create applications based on models, providing API (Application Programming Interface) calling capabilities. Users can call models into created applications through the API interface, and application management tools are also provided to manage and monitor the applications.
[0434] Furthermore, the computing device can also include data management (supporting the creation and management of model tuning datasets), a training center (providing abundant training resources to help users learn and master AI technology), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, it provides a comprehensive and integrated device for AI development, training, deployment, and application.
[0435] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the method of any of the foregoing embodiments. The specific functions and technical effects to be achieved are not described here.
[0436] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments. The computer program is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium. The execution of the computer program by the at least one processor causes the electronic device to perform the technical solutions provided in any of the above method embodiments. Specific functions and achievable technical effects are not elaborated here.
[0437] This application provides a chip, including a processing module and a communication interface. The processing module is capable of executing the technical solutions of the electronic devices described in the foregoing method embodiments. Optionally, the chip further includes a storage module (e.g., a memory), which stores instructions. The processing module executes the instructions stored in the storage module, and the execution of the instructions stored in the storage module causes the processing module to execute the technical solutions provided in any of the foregoing method embodiments.
[0438] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0439] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), a graphics processing unit (GPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules in at least one processor.
[0440] The memory may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0441] The aforementioned memory can be object storage (OSS). This memory can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0442] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as mobile hotspots (WiFi), second-generation (2G), third-generation (3G), fourth-generation (4G) / Long Term Evolution (LTE), fifth-generation (5G), or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), infrared, Ultra Wide Band (UWB), Bluetooth, and other technologies.
[0443] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0444] The aforementioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0445] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside within an application-specific integrated circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components within an electronic device or host device.
[0446] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0447] The order of the embodiments described above is merely for illustrative purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The sequence numbers are merely used to distinguish different operations, and the sequence numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types. "Multiple" means two or more, unless otherwise explicitly specified.
[0448] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0449] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0450] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A computer-aided diagnostic method, characterized in that, include: In response to requests for case-assisted diagnosis, obtain medical images and structured clinical data of the case; Using a multimodal perception model, structured text containing lesion location information, sign descriptions, and their uncertainties is generated based on the medical images; The reasoning agent, based on the structured text and structured clinical data, and combined with the retrieved medical evidence related to the case, performs reasoning to generate auxiliary diagnostic information. Output the auxiliary diagnostic information.
2. The method according to claim 1, characterized in that, The multimodal perception model includes a multimodal perception module and a transformation module. The step of generating structured text containing lesion localization information, sign descriptions, and uncertainties based on the medical images using the multimodal perception model includes: The medical image is input into the multimodal perception module of the multimodal perception model to generate a text representation of the structured text; The conversion module converts the text representation of the structured text into text, thus obtaining the structured text.
3. The method according to claim 2, characterized in that, The training process of the multimodal perception model includes: Obtain historical images and corresponding structured annotation text, wherein the structured annotation text contains lesion location information, sign description and its uncertainty in the historical images; The historical images are input into the multimodal perception module of the multimodal perception model to generate a predicted representation of the structured labeled text; The structured labeled text is converted into a text representation using the text representation module; Based on the predicted representation and the text representation of the structured labeled text, the parameters of the multimodal perception module are adjusted to obtain the trained multimodal perception module.
4. The method according to claim 1, characterized in that, The process of generating structured text containing lesion location information, sign descriptions, and uncertainties based on the medical images using a multimodal perception model includes: The medical image is subjected to multi-step iterative observation using the multimodal perception model, wherein one step of the iterative observation includes: Action decisions are generated based on the current observation status of the medical images. The action decisions include at least one of the following operations: adjusting image display parameters, zooming in on a specified area, switching image layers, and retrieving historical images for comparison. The current observation state is updated based on the action decision, and a continuous representation is generated based on the updated current observation state; When the termination condition is met, the text representation of the structured text is generated based on the continuous representation obtained from the multi-step iterative observation.
5. The method according to claim 1, characterized in that, The multimodal perception model includes a multimodal perception model corresponding to at least one type of medical image, including: radiological images, pathological images, ultrasound images, nuclear medicine images, and endoscopic images.
6. The method according to claim 1, characterized in that, The inference agent has the ability to use models and call tools, and the inference agent supports replacing the model used. The tools that the inference agent can call include at least one of the following: data preprocessing and measurement tools, anomaly detection tools, retrieval tools, and artificial intelligence models. The process involves a reasoning agent that, based on the structured text and structured clinical data, and combined with retrieved medical evidence related to the case, performs reasoning to generate auxiliary diagnostic information, including: The reasoning agent uses the structured text and the structured clinical data as factual source information for the case. Invoke the hypothesis generation model to generate at least one diagnostic hypothesis based on the factual source information; For any of the diagnostic hypotheses, the hypothetical reasoning model is invoked to retrieve relevant medical evidence based on the diagnostic hypotheses, and evidence-based reasoning is performed on the diagnostic hypotheses based on the factual source information and the medical evidence to generate a structured evidence chain for the diagnostic hypotheses. The verification model is invoked to verify the structured evidence chain of each diagnostic hypothesis from at least one dimension: reliability of evidence source, consistency of facts, consistency of reasoning logic, and sufficiency of evidence citation. After the structured evidence chains of each diagnostic hypothesis are verified, the ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured evidence chains across diagnostic hypotheses, and to generate structured auxiliary diagnostic information.
7. The method according to claim 6, characterized in that, Also includes: For any of the diagnostic hypotheses, if the verification of the structured evidence chain of the diagnostic hypothesis fails, the hypothesis reasoning model corresponding to the diagnostic hypothesis is invoked to generate a new structured evidence chain for the diagnostic hypothesis. The verification model is invoked to verify the new structured chain of evidence from at least one dimension: reliability of evidence source, consistency of facts, consistency of reasoning logic, and sufficiency of evidence citation. Until a new structured chain of evidence is verified for the diagnostic hypothesis.
8. The method according to claim 7, characterized in that, For any of the aforementioned diagnostic hypotheses, if the verification of the structured chain of evidence for the diagnostic hypothesis fails, before invoking the hypothesis reasoning model corresponding to the diagnostic hypothesis to generate a new structured chain of evidence for the diagnostic hypothesis, the method further includes: Based on the verification results of the structured evidence chain of the diagnostic hypothesis, an incremental retrieval query for the diagnostic hypothesis is constructed; Based on the incremental search query, relevant medical evidence is retrieved from a multi-source heterogeneous medical knowledge base to obtain incremental search results.
9. The method according to claim 6, characterized in that, The step of retrieving relevant medical evidence based on the diagnostic hypothesis includes: Transform the diagnostic hypotheses into diagnostic clinical questions; Relevant medical evidence is retrieved from a multi-source heterogeneous medical knowledge base based on the diagnostic clinical question.
10. The method according to claim 9, characterized in that, The process of retrieving relevant medical evidence from a multi-source heterogeneous medical knowledge base based on the diagnostic clinical question includes: The retrieval planning model is invoked to identify the retrieval intent based on the diagnostic clinical question and generate tool invocation information corresponding to the retrieval intent. The tool invocation information includes the identifier of the retrieval tool and retrieval parameters. Based on the tool call information, the corresponding retrieval tool is invoked to retrieve relevant medical evidence from the corresponding medical knowledge base.
11. The method according to claim 10, characterized in that, After invoking the corresponding retrieval tool based on the tool invocation information to retrieve relevant medical evidence from the corresponding medical knowledge base, the process further includes: The evidence sufficiency assessment model is invoked to determine whether the retrieved medical evidence is sufficient to respond to the diagnostic clinical question based on the diagnostic clinical question and the retrieved medical evidence. If the retrieved medical evidence is insufficient to respond to the diagnostic clinical question, a supplementary query question is generated. Relevant medical evidence is retrieved from a multi-source heterogeneous medical knowledge base based on the supplementary query question.
12. The method according to any one of claims 6-11, characterized in that, The hypothesis generation model generates at least one diagnostic hypothesis based on the fact source information, including: The hypothesis generation model is invoked, and a retrieval tool is used to retrieve relevant medical evidence from a multi-source heterogeneous medical knowledge base based on the fact source information. Based on the fact source information and the medical evidence, at least one diagnostic hypothesis is generated.
13. The method according to claim 1, characterized in that, After the step of generating auxiliary diagnostic information by reasoning through the reasoning agent based on the structured text and the structured clinical data, combined with the retrieved medical evidence related to the case, the method further includes: The examination recommendation model is invoked to determine the next candidate examination items based on the auxiliary diagnostic information and existing medical evidence. The marginal information value and cost evaluation value of each candidate examination item are evaluated, and the next examination item is recommended based on the marginal information value and cost evaluation value of each candidate examination item. Wherein, the marginal information value is used to characterize the degree of contribution of the candidate examination item to the diagnostic uncertainty, and the cost assessment value is used to characterize the cost of the candidate examination item in at least one dimension of cost, risk, and time.
14. The method according to any one of claims 6-11, characterized in that, The auxiliary diagnostic information includes: The report summarizes the evidence-based diagnosis, identifies diagnostic bottlenecks, and lists the diagnostic hypotheses, their confidence levels, and reasoning bases in sequence.
15. The method according to claim 14, characterized in that, The auxiliary diagnostic information also includes suggestions for further medical treatment. The auxiliary diagnostic information is generated by the reasoning agent based on the structured text and structured clinical data, combined with retrieved medical evidence related to the case. This includes: The suggestion generation model is invoked to generate next-step medical advice based on the structured text, the structured clinical data, existing medical evidence, each diagnostic hypothesis and its confidence level, the structured evidence chain, and diagnostic bottleneck information. The next-step medical advice includes at least one of the following: recommended examination items, treatment and management suggestions, follow-up and monitoring suggestions, and precautions to be informed to the patient.
16. The method according to claim 15, characterized in that, After outputting the auxiliary diagnostic information, the method further includes: Receive feedback data regarding the next medical visit recommendations, including newly added images and / or newly added clinical data; The reasoning agent, based on the structured text of the medical images, the structured clinical data, and the structured text and / or new clinical data of the newly added images, and combined with the retrieved medical evidence related to the cases, performs reasoning to generate new auxiliary diagnostic information; wherein, the structured text of the newly added images is generated by the multimodal perception model based on the newly added images; Output the new auxiliary diagnostic information.
17. The method according to claim 14, characterized in that, The output of the auxiliary diagnostic information includes: Based on each diagnostic hypothesis and its corresponding confidence level in the auxiliary diagnostic information, determine the inconsistency score of each diagnostic hypothesis; The diagnostic hypothesis that the non-consistency score is less than a preset non-consistency threshold is added to the prediction set, wherein the preset non-consistency threshold is obtained by calculating the quantile threshold of the non-consistency score distribution corresponding to the real diagnosis in the calibration dataset based on the conformal prediction method. Based on the evidence-based diagnostic summary, the diagnostic bottleneck information, and the diagnostic hypotheses contained in the prediction set, along with their corresponding confidence levels and reasoning bases, an auxiliary diagnostic report is generated. Output the auxiliary diagnostic report.
18. The method according to any one of claims 6-11, characterized in that, For any of the aforementioned diagnostic hypotheses, the process involves invoking a hypothesis reasoning model to retrieve relevant medical evidence based on the diagnostic hypothesis, and performing evidence-based reasoning on the diagnostic hypothesis based on the factual source information and the medical evidence to generate a structured evidence chain for the diagnostic hypothesis, including: Identify the specialty to which the diagnostic hypothesis belongs; Based on the specialty type to which the diagnostic hypothesis belongs, the hypothesis reasoning model corresponding to the specialty type is invoked. Through the hypothesis reasoning model corresponding to the specialty type, relevant specialty medical evidence is retrieved based on the diagnostic hypothesis and its specialty type. Based on the factual source information and the specialty medical evidence, the diagnostic hypothesis is reasoned on an evidence-based basis to generate a structured evidence chain for the diagnostic hypothesis.
19. The method according to claim 18, characterized in that, The invocation ranking model evaluates and ranks each diagnostic hypothesis and its structured evidence chain across diagnostic hypotheses, and generates structured auxiliary diagnostic information, including: The ranking model is invoked to evaluate and rank each diagnostic hypothesis and its structured evidence chain across diagnostic hypotheses, assess the contradictions between the structured evidence chains of each diagnostic hypothesis and the cross-specialty system influence relationship, and generate auxiliary diagnostic information for multi-specialty consultation.
20. The method according to any one of claims 1-11, characterized in that, The method further includes: The perceptual error of the multimodal perception model and the performance ceiling of the inference agent are evaluated using a benchmark set containing historical images and their corresponding structured reference text. Based on the auxiliary diagnostic information output by the reasoning agent, the actual performance of the reasoning agent under the current perception conditions is determined, wherein the current perception conditions include the structured text output by the multimodal perception model; The difference between the performance upper limit and the actual performance is used as the performance degradation of the reasoning agent under the current perception conditions. Based on the perception error and the performance degradation, error attribution is performed on the multimodal perception model and the reasoning agent; Output the results of the error attribution.
21. The method according to claim 20, characterized in that, The method further includes: After the reasoning model used by the reasoning agent changes, the same benchmark evaluation set is used to evaluate the first performance limit of the original reasoning agent before the reasoning model change, and the second performance limit of the new reasoning agent after the reasoning model change. Based on the first performance limit and the second performance limit, determine whether the performance of the new inference agent has degraded, and output the determination result.
22. A computer-aided diagnostic device, characterized in that, include: The data receiving module is used to acquire medical images and structured clinical data of a case in response to a request for case-assisted diagnosis. The perception module is used to generate structured text containing lesion location information, sign descriptions and their uncertainties based on the medical images using a multimodal perception model. The reasoning module is used to generate auxiliary diagnostic information by reasoning through a reasoning agent based on the structured text and the structured clinical data, combined with the retrieved medical evidence related to the case. The reasoning agent supports replacing the model used. The output module is used to output the auxiliary diagnostic information.
23. An electronic device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-21.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-21.
25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-21.
26. A computer-aided diagnostic system, characterized in that, include: The client and the electronic device as described in claim 23, The client is used to send a case-assisted diagnosis request to the electronic device and provide the electronic device with medical images and structured clinical data of the case.