A cerebral vascular lesion auxiliary evaluation system and method based on a multi-modal intelligent agent
By deeply integrating 3D medical images and clinical text through a multimodal intelligent agent system, the problem of insufficient image-text fusion in existing technologies has been solved, achieving highly sensitive and safe auxiliary diagnosis of cerebrovascular lesions, which is particularly suitable for emergency stroke scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING TSINGHUA CHANGGUNG HOSPITAL
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies cannot effectively integrate 3D medical images with clinical text in stroke emergency scenarios, lack cross-modal interaction, resulting in insufficient diagnostic sensitivity, high risk of misjudgment, and lack of medical safety guarantees.
A multimodal intelligent agent system is adopted, including a data processing module, a cross-modal consistency verification agent, a visual refocusing and feature reverse lookup agent, and a multi-factor collaborative reasoning and decision-making agent, to achieve deep structured analysis of images and text, logical consistency assessment, and security decision support.
It significantly improves the sensitivity and accuracy of detecting cerebrovascular lesions, reduces the risk of missed or misdiagnosed cases, provides reliable auxiliary diagnostic support, and is suitable for auxiliary diagnosis and thrombolysis suitability assessment in emergency stroke scenarios.
Smart Images

Figure CN121983293B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence and smart medical data processing technology, and in particular to an auxiliary assessment system and method for cerebrovascular lesions based on multimodal intelligent agents. Background Technology
[0002] Stroke is an acute cerebrovascular disease caused by the sudden rupture or blockage of a blood vessel in the brain, resulting in an interruption of blood flow. This disease is extremely time-sensitive, and in clinical treatment, "time is brain." In the emergency room, doctors must, within a very narrow time window (usually within 4.5 hours of onset), rapidly combine the patient's imaging examinations (primarily non-contrast CT) with clinical history to accurately differentiate between hemorrhagic and ischemic stroke, thereby developing a targeted treatment plan (such as intravenous thrombolysis). This process places stringent demands on both diagnostic accuracy and efficiency.
[0003] Current diagnostic aids primarily revolve around single-modal data. In medical image analysis, most studies rely on convolutional neural networks (such as U-Net and ResNet) or visual Transformers for lesion segmentation or classification in NCCT images, for example, using 3D CNNs to identify large vessel occlusions. In clinical text analysis, natural language processing techniques (such as BERT and RNNs) are mainly used to extract information such as symptoms, chief complaints, and neurological function scores from electronic medical records. To overcome the limitations of single-modality approaches, multimodal fusion methods are gaining attention. Early techniques often used feature stitching to achieve static fusion of images and text, but lacked deep cross-modal interaction. Furthermore, with the development of multimodal large language models, attempts have been made to use general models (such as GPT-4V) for diagnostic aids, but these are mainly designed for natural images and have limited support for medical 3D images. To adapt to these models, 3D images are often dimensionality-reduced (e.g., selecting intermediate layers or maximum density projection), resulting in the loss of spatial information along the scanning direction.
[0004] Existing methods in stroke emergency scenarios that rely solely on image-based models lack sensitivity to subtle early signs (such as blurred gray-white matter boundaries and insular signs), easily leading to missed diagnoses. Text-based models, on the other hand, have a high risk of misdiagnosis due to the similarity of hemorrhagic and ischemic stroke symptoms, potentially causing incorrect treatment. Secondly, traditional feature stitching methods lack dynamic cross-modal validation capabilities. When image presentations significantly conflict with clinical symptoms, the model struggles to achieve a closed-loop reasoning process similar to a clinician's "re-examination," limiting diagnostic robustness. Furthermore, existing general-purpose multimodal models are primarily optimized for 2D images, offering weak support for the spatial continuity of 3D medical images. Dimensionality reduction easily loses features of small, cross-layer lesions, restricting their application in refined diagnosis. Moreover, directly using general-purpose large-scale language models for assisted diagnosis often lacks strict medical safety gating, fails to systematically screen for contraindications, and the reasoning process does not conform to clinical guidelines, posing a risk of "illusory" outputs and poor interpretability, resulting in insufficient clinical credibility. Summary of the Invention
[0005] In view of the above analysis, the embodiments of the present invention aim to provide a cerebrovascular lesion auxiliary assessment system and method based on multimodal intelligent agents, in order to solve the problems that existing methods cannot perform deep integration and conflict verification with clinical text while maintaining the complete spatial information of three-dimensional medical images and lack deterministic security guarantees that comply with medical standards.
[0006] On one hand, embodiments of the present invention provide an auxiliary assessment system for cerebrovascular lesions based on multimodal intelligent agents, comprising: The data processing module is used to acquire three-dimensional non-enhanced CT image data of the patient's cerebral blood vessels and the corresponding electronic medical record text data, and generate structured image feature data and structured clinical feature data. A cross-modal consistency verification agent is used to perform medical logical consistency assessment on the structured image feature data and structured clinical feature data to obtain consistency assessment results; if the consistency assessment results meet preset conditions, a visual query instruction is generated and output to the visual refocusing and feature reverse lookup agent. A visual refocusing and feature reverse lookup intelligent agent is used to perform directional verification of the three-dimensional non-enhanced CT image according to the visual query instruction and generate structured reverse lookup results; A multi-factor collaborative reasoning and decision-making intelligent agent is used to fuse the structured image feature data, structured clinical feature data, and / or structured reverse lookup results, and generate the final auxiliary evaluation result through counterfactual reasoning and security gating mechanisms.
[0007] On the other hand, this invention discloses an auxiliary assessment method for cerebrovascular lesions based on multimodal intelligent agents, comprising the following steps: Acquire three-dimensional non-contrast CT image data of the patient's cerebral blood vessels and corresponding electronic medical record text data, and generate structured image feature data and structured clinical feature data. A medical logic consistency assessment is performed on the structured image feature data and the structured clinical feature data to obtain a consistency assessment result; if the consistency assessment result meets the preset conditions, a visual query instruction is generated; the three-dimensional non-enhanced CT image is then re-examined according to the visual query instruction to generate a structured reverse query result; By integrating the structured image feature data, structured clinical feature data, and / or structured reverse lookup results, and through counterfactual reasoning and security gating mechanisms, the final auxiliary assessment result is generated.
[0008] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: This invention provides a multimodal intelligent agent-based auxiliary assessment system and method for cerebrovascular lesions. Through parallel processing of a primary image interpretation agent and a clinical information structuring agent, it achieves deep structured analysis of multimodal data, providing a standardized evidentiary basis for subsequent reasoning. A cross-modal consistency verification agent evaluates the logical matching degree between images and clinical features using natural language reasoning. Upon detecting conflicts or uncertainties, it automatically generates visual query instructions, driving a visual refocusing and feature reverse lookup agent to perform targeted verification of suspicious areas. This effectively uncovers hidden signs of ultra-early and minute lesions, significantly improving the sensitivity and accuracy of lesion detection. A multi-factor collaborative reasoning and decision-making agent verifies the logical robustness of conclusions through counterfactual reasoning and automatically screens for treatment contraindications using a security gating mechanism. While ensuring clinical safety, it generates auxiliary assessment results including diagnostic conclusions, treatment recommendations, and interpretable reasoning paths. This effectively reduces the risk of missed or misdiagnosed cases in complex and high-risk cases, providing reliable, transparent, and traceable intelligent support for clinical decision-making. It is particularly suitable for auxiliary diagnosis, classification, and thrombolysis suitability assessment in emergency stroke scenarios.
[0009] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0010] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0011] Figure 1This is a schematic diagram of the structure of the cerebrovascular lesion auxiliary assessment system based on multi-agent collaboration provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart illustrating the auxiliary assessment method for cerebrovascular lesions based on multi-agent collaboration provided in Embodiment 2 of the present invention. Detailed Implementation
[0012] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0013] Example 1 A specific embodiment of the present invention discloses an auxiliary assessment system for cerebrovascular lesions based on multi-agent collaboration, such as... Figure 1 As shown, it includes: The data processing module is used to acquire three-dimensional non-enhanced CT image data of the patient's cerebral blood vessels and the corresponding electronic medical record text data, and generate structured image feature data and structured clinical feature data. A cross-modal consistency verification agent is used to perform medical logical consistency assessment on the structured image feature data and structured clinical feature data to obtain consistency assessment results; if the consistency assessment results meet preset conditions, a visual query instruction is generated and output to the visual refocusing and feature reverse lookup agent. A visual refocusing and feature reverse lookup intelligent agent is used to perform directional verification of the three-dimensional non-enhanced CT image according to the visual query instruction and generate structured reverse lookup results; A multi-factor collaborative reasoning and decision-making intelligent agent is used to fuse the structured image feature data, structured clinical feature data, and / or structured reverse lookup results, and generate the final auxiliary evaluation result through counterfactual reasoning and security gating mechanisms.
[0014] In implementation, the data processing module includes: The data receiving module is used to acquire three-dimensional non-enhanced CT image data of the patient's cerebral blood vessels and the corresponding electronic medical record text data; A primary image interpretation agent is used to perform structured analysis on the three-dimensional non-enhanced CT images, extract image features related to cerebrovascular lesions, and generate structured image feature data. A clinical information structured intelligent agent is used to perform semantic parsing on the electronic medical record text, converting unstructured medical record information into structured clinical feature data with predefined fields.
[0015] In specific implementation, the primary image interpretation agent includes: The preprocessing unit is used to perform physical value conversion, spatial standardization, and 3D data construction on 3D non-enhanced CT images to generate standardized 3D image tensors. The anatomical partitioning unit is used to divide a standardized three-dimensional image tensor into multiple anatomical regions based on predefined brain anatomical partitioning rules, generating regional feature data with anatomical region identifiers. The regional feature data with anatomical region identifiers is a three-dimensional data volume of the anatomical region, wherein each voxel in the three-dimensional data volume includes the voxel value of the voxel and the anatomical region identifier to which the voxel belongs. The voxel value of each voxel in the three-dimensional data volume is obtained based on the standardized three-dimensional image tensor, i.e., Henlein units (HU). The multimodal large model inference unit is used to extract features from each anatomical region based on regional feature data with anatomical region identifiers and generate structured image feature data.
[0016] Specifically, the preprocessing unit generates the normalized 3D image tensor in the following manner: The pixel values of the three-dimensional non-enhanced CT images are calibrated to Henle units (HU) using device metadata, and contrast enhancement and intensity normalization are performed based on window width and window level (WW / WC) to achieve physical value conversion, resulting in the three-dimensional non-enhanced CT images after physical value conversion; wherein, intensity normalization is performed in the range of [0, 255]; it can be understood that the problem of brightness distribution differences between different CT devices is solved by physical value conversion. By using slice thickness consistency screening, instance number rearrangement, centering cropping, and zero-value filling, the planar resolution of the 3D non-enhanced CT images after physical value conversion is standardized, and the Z-axis scanning gap is eliminated to achieve spatial standardization, resulting in each 2D slice after spatial standardization. Among them, the planar resolution is standardized to 512×512. It can be understood that spatial standardization ensures the geometric stability of global feature extraction. The two-dimensional slices after spatial standardization are stacked into a three-dimensional NumPy array, which serves as a standardized three-dimensional image tensor to realize the construction of three-dimensional data. It can be understood that the data format of the standardized three-dimensional image tensor can directly support the video-based 3D visual encoding in the backend of the system, thereby ensuring the system's sensitivity in capturing small lesions across different planes from the root.
[0017] Specifically, the anatomical partition unit generates regional feature data with anatomical region identifiers in the following manner: Based on predefined brain anatomical partitioning rules, three-dimensional voxel masks for each anatomical region are generated through anatomical atlas registration or lightweight segmentation algorithms. Based on the three-dimensional voxel masks for each anatomical region and the standardized three-dimensional image tensor, regional feature data with anatomical region identifiers are obtained.
[0018] More specifically, the brain anatomical zoning rules are predefined based on standard brain anatomical atlases and clinical imaging consensus. In this embodiment, the anatomical regions include brain parenchyma, ventricular system, subarachnoid space, cerebellum, basal ganglia, internal capsule, insular cortex, thalamus, septum pellucidum, dura mater folds, and venous sinuses. The spatial extent of each anatomical region is calibrated according to the standard spatial coordinates of the Montreal Neuroscience Institute (MNI) to form a three-dimensional region template mask. This template mask assigns an anatomical region identifier ID to each voxel position in the MNI space (e.g., 0 represents the background, and 1~11 represent the aforementioned anatomical regions).
[0019] More specifically, the atlas mapping based on image registration involves nonlinearly registering the patient's standardized 3D image tensor with the MNI standard template to obtain a deformation field from the patient image space to the MNI space. This deformation field is then used to reverse-map a predefined MNI space region template mask back to the patient image space, thereby assigning a corresponding anatomical region identifier to each voxel of the patient image. The registration can be implemented using classical algorithms (such as SyN and ANTs) or deep learning registration networks (such as VoxelMorph).
[0020] More specifically, deep learning-based semantic segmentation trains a 3D segmentation network (such as U-Net or V-Net) that takes a standardized 3D image tensor as input and directly outputs the anatomical region probability map for each voxel. The voxel-level region identifier ID is obtained through the argmax operation. The training data uses public datasets (such as CANDI or LPBA40) or private labeled data, where each voxel is labeled with an anatomical region identifier ID. The advantage of this method is that it does not require explicit registration and can directly partition the original image space.
[0021] Understandably, by dividing the standardized three-dimensional image tensor into specific anatomical regions according to predefined brain anatomical partitioning rules, the multimodal large model inference unit can perform feature analysis using anatomical regions as the basic unit. This effectively avoids the problem of local subtle signs (such as signs of ultra-early ischemia) being "diluted" or masked during the global feature extraction process, thereby improving the system's sensitivity to capturing early subtle lesions.
[0022] Specifically, the multimodal large model inference unit, obtained through training on the multimodal large model, includes: The 3D visual encoder is used to generate an enhanced visual feature sequence based on region feature data with anatomical region identifiers. The 3D visual encoder adopts a 3D ViT architecture, which segments the input region feature data with anatomical region identifiers into a sequence of three-dimensional image patches. Each three-dimensional image patch is mapped to a visual word through a linear projection layer. The anatomical region identifier corresponding to each three-dimensional image patch is encoded into a learnable region embedding vector. After being added to the visual word, it is then passed through multiple Transformer layers for deep feature extraction, and the enhanced visual feature sequence is output. The modality alignment layer is used to align the visual feature sequence of each anatomical region with the text semantic space. It includes: using a cross-attention mechanism, with a preset query term as the query and a high-dimensional visual feature vector as the key and value, and calculating the attention output; then using a learnable modality adapter to map the visual features to the semantic space of the language model. The adapter consists of two fully connected layers with the GELU activation function in between, and outputs the aligned feature vector sequence. The large language model decoder is used to decode aligned feature vectors into structured text and generate structured image feature data. It includes: a pre-trained large language model (such as LLaMA, Qwen, etc.) that uses an autoregressive generation method, takes the aligned feature vector sequence as a prefix input, and guides the model to generate text according to a predefined structured format. At each generation step, the model predicts the next word based on the generated content and semantic features. The generation process is constrained by the structured format to ensure that the output conforms to JSON or other predefined format specifications.
[0023] Specifically, the structured image feature data includes at least one anomaly, and each anomaly includes structured image feature data such as anomaly indication, anatomical location, anomaly coordinate range, density features, and morphological features; wherein, Abnormal items are indicated by the qualitative lesion type or sign name described in medical semantics, and the pathological changes presented by the images are characterized by standard medical terminology, such as "high-density arterial sign", "shallow sulci", "blurred gray and white matter boundary", "intracranial hematoma", "local density decrease" etc. Anatomical location: Identify the specific neuroanatomical region where the abnormality is located, and describe it using the anatomical region name obtained from the anatomical zoning rules and the lateral information (e.g., "left", "right", "bilateral" or "midline"). Anomaly coordinate range: the coordinates of the bounding rectangle indicating the area where the anomaly is located; Density characteristics describe the physical properties of abnormal areas based on Henle units (HU). They can include quantitative values (such as HU range, average HU) or qualitative descriptions (such as "high density", "low density", "isodensity", "mixed density"). They are the core criteria for distinguishing between hemorrhagic and ischemic lesions. Morphological features, which describe the spatial morphology of the lesion using structured parameters, include at least one or a combination of the following subfields: volume (three-dimensional dimensions of the abnormality), maximum diameter (maximum linear dimension of the abnormality in any direction), boundary (description of boundary clarity (e.g., "clear", "fuzzy", "irregular")), shape (description of geometric morphology (e.g., "circular", "tubular", "sheet-like")), and occupancy effect (whether it causes compression or displacement of surrounding structures (e.g., "yes", "no")).
[0024] Preferably, in addition to anomalies, the structured imaging feature data may also include: statistical information of each anatomical region (such as average HU, symmetry differences), global imaging features (such as midline shift, ventricular compression status), etc., to provide more comprehensive imaging factual evidence.
[0025] More specifically, the multimodal large model inference unit is trained using a hybrid source multi-task training mechanism; wherein the training tasks are report generation, visual question answering, target localization, and cross-modal alignment.
[0026] Furthermore, the training sample dataset in the training mechanism is generated using publicly available medical image datasets, private clinical image data, and pseudo-labeled data. Each training sample includes 3D image data (i.e., regional feature data with anatomical region identifiers) and corresponding structured labels. The structured labels include report generation labels (including anomaly indications, anatomical location, density features, and morphological features), target localization labels (including anomaly coordinate ranges), and visual question-answering labels. Public medical imaging datasets, such as the RSNA intracranial hemorrhage detection dataset and the CANDI brain structure dataset, provide basic lesion identification capabilities; private clinical imaging data, including privacy-free 3D NCCT images and corresponding radiology reports obtained from partner hospitals; and pseudo-labeled data, which uses a basic model to generate preliminary labels for unlabeled private clinical images, and expands the training samples after rule correction. Visual question-and-answer tags are question-and-answer pairs constructed based on medical rules. Question-and-answer types include existence judgments (e.g., "Is there a low-density lesion in the right basal ganglia?" Answer "Yes / No"), classification judgments (e.g., "Is this lesion hemorrhagic or ischemic?" Answer "Ischemic"), counting questions (e.g., "How many types of hemorrhage are involved in the imaging?" Answer "2 types"), multi-label judgments (e.g., "Which of the following structures have abnormal density? Candidates: basal ganglia, internal capsule, insular cortex" Answer "basal ganglia, insular cortex"), and logical reasoning (e.g., "Comparing the left and right ventricular systems, is there midline shift?" Answer "Yes, shifted 3mm to the left").
[0027] For example, the visual question-answering label is: "1. Based on the rules of neuroanatomy, please focus on examining the right basal ganglia region (especially the lentiform nucleus and internal capsule region) to assess whether there is a slight decrease in density or a slight mass effect (such as shallowing of the sulci)." 2. Please perform a full CT scan of the patient's images and determine the types of bleeding involved. 3. Based on predefined intracranial anatomical regions, please determine which of the following structures in the right hemisphere exhibit abnormal decreased density? (Candidates: brainstem, cerebellum, basal ganglia, internal capsule, insular cortex, central sulcus).
[0028] 4. By comparing the volume and morphology of the left and right ventricular systems, and considering the positional changes of the septum pellucidum, infer whether the current imaging indicates midline shift in the brain tissue due to a unilateral mass effect. Furthermore, the training mechanism employs SFT fine-tuning technology for model fine-tuning, using a loss function... Represented as: ; In the formula, , , , These represent the report generation loss, visual question answering loss, target localization loss, and cross-modal alignment loss, respectively. , , , These represent the weight coefficients of the report generation loss, visual question answering loss, target localization loss, and cross-modal alignment loss, respectively, used to balance the contributions of different tasks. The weight coefficients are usually dynamically adjusted according to the convergence difficulty and importance of the tasks, and can be determined by grid search or adaptive adjustment methods (such as uncertainty weighting). In the early stage of training, they can be set to equal weights (such as 0.25 for each), and then fine-tuned according to the performance on the validation set.
[0029] Furthermore, the report generates losses. Represented as: ; In the formula, This indicates the total number of tokens in the generated labels reported in the sample. Indicates the number of labels generated in the report of the sample. The real word element at each position, The report generates labels indicating the forecast. A sequence of lexical units at positions, This represents the sequence of visual features extracted by the 3D visual encoder. This represents a predefined sequence of query terms. The model predicts the first... In the probability distribution of lexical elements at each position, the corresponding real lexical elements The probability value.
[0030] Understandably, the report generated a loss. During training, the model uses real report-generated labels as the target. It predicts each word in an autoregressive manner, calculating the cross-entropy between the predicted probability and the real word at each step, and then averaging the results over all positions. This loss forces the model to generate text that closely approximates real reports in terms of vocabulary and grammar, while ensuring that the content is consistent with visual features. This enables the model to learn to transform visual features into structured text that conforms to medical semantics, accurately describing anomalies, anatomical locations, density features, and morphological features.
[0031] Furthermore, visual question answering loss Represented as: ; For classification problems, For category indexes of the correct answers; for generative questions, The sequence of words for the answer is given, and the loss is the average of the sequence cross-entropy. In this embodiment, a classification problem is used as an example. In the formula, This represents the total number of question-answer pairs in the visual question-answering labels of the sample. This indicates the first visual question-answering label in the sample. Textual representation of the questions in a question-and-answer pair; This indicates the first visual question-answering label in the sample. The standard answer (category label or text sequence) to each question in the question-and-answer pair; The model represents the solution to a given problem. and visual feature sequences The probability of outputting the correct answer.
[0032] Understandably, the model is fed both the question text and visual features, and the model outputs the probability distribution of the answers, calculating the cross-entropy loss between the output and the true answers. For multiple-choice or classification questions, the loss is the category cross-entropy; for open-ended generation questions, the same word-by-word meta-cross-entropy as in report generation is used. This loss forces the model to understand image content at a fine-grained level, enabling it to answer specific questions about the existence, classification, counting, multi-label judgment, and logical reasoning of lesions (anomalies).
[0033] Furthermore, target positioning loss Represented as: ; In the formula, This indicates the range of abnormal coordinates in the target location labels within the sample. This indicates the range of anomaly coordinates predicted by the model. Indicates Smooth L1 loss, Indicates the 3D GIoU loss. , These represent the first balance coefficient and the second balance coefficient, respectively; where, Used to measure the degree of overlap and containment relationship between predicted bounding boxes and ground truth bounding boxes. For example, , Set them to 1 and 0.5 respectively, or fine-tune them based on the validation set.
[0034] Understandably, the model outputs multiple candidate bounding boxes and their confidence scores through the object detection head, and only calculates the loss for the predicted boxes that match the ground truth boxes. The Smooth L1 loss focuses on the absolute error of the coordinates, while the GIoU loss focuses on the overall overlap quality of the boxes. The combination of the two ensures that the predicted boxes are both accurate and closely match the target. This loss enables the model to accurately pinpoint the spatial location of lesions (anomalies) and output the 3D bounding box coordinates, i.e., the range of anomaly coordinates.
[0035] Furthermore, cross-modal alignment loss Represented as: ; In the formula, Indicates the first Visual feature sequences of a sample; Indicates the first The positive sample feature sequence obtained by data augmentation (such as rotation, cropping, intensity adjustment) of the three-dimensional image data of each sample; Indicates the first in the same training batch A sequence of visual features for each sample, including positive and negative samples; sim() represents the cosine similarity function; This represents the temperature coefficient, which controls the smoothness of the contrast loss, and is usually set to 0.07. This indicates the training batch size.
[0036] Understandably, for each training batch, an augmented view is generated for each image, and the feature vectors of all samples are calculated. For each sample... Its positive samples are the corresponding augmented view features. ,the remaining Each sample (including other images and their augmented views) is used as a negative sample. This loss function encourages... and High similarity with other samples and low similarity enhances the robustness and discriminativeness of visual features, ensuring the model's semantic understanding of images is unaffected by enhancement transformations. It also spatially separates features from different images, facilitating differentiation in downstream tasks. Furthermore, through contrastive learning, the features extracted by the visual encoder become more stable and insensitive to noise and grayscale changes, contributing to improved accuracy in report generation and question-answering pairs. Simultaneously, this loss promotes initial alignment between visual features and the semantic space of the language model, providing better initialization for subsequent modality alignment layers.
[0037] In practice, the structured intelligent agent for clinical information is built on a large language model. Through prompt word engineering or supervised fine-tuning, it extracts key information from electronic medical record text and converts it into standardized structured clinical feature data.
[0038] Specifically, the electronic medical record text input into the structured clinical information agent typically includes paragraphs such as chief complaint, present illness, past medical history, physical examination, and auxiliary examinations. The output structured clinical feature data includes the patient's basic demographic information (such as gender and age), stroke-related key timelines (such as onset time, consultation time, and the time interval between the two), clinical manifestations and neurological assessments (such as the original chief complaint, lateralization of symptoms (left, right, or bilateral), specific symptom types (such as limb weakness, speech disorders, etc.), and a detailed description of neurological examinations, including quantified NIHSS scores when necessary), past medical history and risk factors (such as hypertension, diabetes, atrial fibrillation, and other cerebrovascular risk factors), contraindications and medication history (such as recent surgery, intracranial hemorrhage, anticoagulants and antiplatelet drugs), laboratory and imaging results (blood glucose, coagulation function), and logical consistency and data quality evaluation (indicating which fields are not mentioned in the medical record or cannot be extracted).
[0039] For example, the fields of the output structured clinical feature data include gender, age, onset time, consultation time, interval between onset and consultation, original chief complaint, symptom lateralization, list of symptom types, description of neurological examination, NIHSS score, history of hypertension, history of diabetes, and list of missing fields.
[0040] Specifically, if the structured intelligent agent for clinical information is implemented through prompt word engineering, the large language model calls the general large language model API (such as GPT-4, Qwen-max, etc.).
[0041] For example, the prompt word template is: [System Command] You are a professional medical information extraction assistant. Please extract information from the following electronic medical record text and return it strictly in JSON format.
[0042] [Medical Record Text] {Electronic Medical Record Text} [Field Definition] Gender, string Age, integer Onset time, in ISO format (if described as "X hours ago", please estimate based on the time of medical visit). Consultation time, ISO format The interval between the onset of illness and the visit to a doctor is an integer. Chief complaint text, string Symptom lateralization, enumeration values: left / right / bilateral / none Neurological examination description, string NIHSS score, or null if none. History of hypertension, Boolean value ... (Complete field definition) [Output Requirements] Returns valid JSON, with keys using the English names mentioned above.
[0043] Fields that cannot be determined should be set to null.
[0044] Do not include any extra text. Specifically, if the clinical information structured intelligent agent is implemented through supervised fine-tuning, an open-source large language model (such as Qwen-7B, LLaMA-2-7B) is selected as the base model for training. This involves collecting de-privacy electronic medical record texts, which are then annotated with corresponding structured JSON by neurologists. Each sample's input is the original medical record text, and the output is the target JSON string. Data augmentation, such as synonym replacement and paragraph rearrangement, is performed on the original text, expanding it to over 5000 entries. The task description and medical record text are concatenated as input, and the target JSON is used as the output, with training conducted using a standard instruction fine-tuning paradigm.
[0045] In implementation, the cross-modal consistency verification agent is constructed based on a large language model and adopts a natural language reasoning framework. It uses the structured clinical feature data as a reasoning premise and the structured image feature data as a reasoning hypothesis to evaluate the degree of logical matching in neuroanatomy and output the consistency assessment result. When the consistency assessment result meets the preset conditions, the cross-modal consistency verification agent generates a visual query instruction based on neuroanatomy prior rules and structured clinical feature data. The visual query instruction includes the target side, the target anatomical region identifier, and the type of image sign to be verified.
[0046] Specifically, the consistency assessment result is a quantified consistency score. The cross-modal consistency verification agent maps the consistency score to multiple predefined risk level intervals. Each risk level interval includes at least: a first interval, indicating a high degree of consistency between structured imaging feature data and structured clinical feature data; a second interval, indicating uncertainty or insufficient evidence between structured imaging feature data and structured clinical feature data; and a third interval, indicating a logical conflict between structured imaging feature data and structured clinical feature data. A preset condition is that the consistency score falls within either the second or third interval. For example, the consistency score ranges from 0 to 1; the first interval ranges from 0.8 to 1 (inclusive), the second interval ranges from 0.4 to 0.8 (inclusive), and the third interval ranges from 0 to 0.4 (inclusive).
[0047] Specifically, the neuroanatomical a priori rules include: a lesion lateral mapping rule, which maps symptoms in clinical features to the contralateral cerebral hemisphere based on the principle of cross-innervation of the nervous system. For example, if the clinical feature is "left limb dysfunction", the rule automatically maps it to "right cerebral hemisphere"; a functional impairment-anatomical region mapping rule, which maps the symptom type in clinical features to the corresponding neurofunctional anatomical region. For example, "limb motor dysfunction" is mapped to "basal ganglia" or "motor cortex", and "speech dysfunction" is mapped to "dominant hemisphere language area"; and an imaging sign type mapping rule, which determines the type of imaging sign to be verified based on the onset time and severity of the clinical features. For example, if the onset is within 2 hours and the NIHSS score is high, it is mapped to verify "high-density middle cerebral artery sign (HMCAS)" or "south insula sign"; if the NIHSS score indicates typical infarction, it is mapped to verify "shallow sulci" or "localized density reduction".
[0048] Specifically, in the visual query instruction, the target lateralization is used to identify the lateralization of the cerebral hemisphere that needs to be reviewed, the target anatomical region identifier is used to identify the specific neuroanatomical region that needs to be reviewed, and the image sign type to be verified is used to identify the subtle image features that need to be checked in detail.
[0049] In specific implementation, the cross-modal consistency verification agent is implemented using prompt word engineering. It encapsulates structured clinical feature data and structured image feature data into a natural language reasoning task through pre-designed prompt word templates, calls a general large language model API to perform consistency assessment, and parses the consistency score and visual query instructions from the model output.
[0050] Specifically, the prompt word template for the cross-modal consistency verification agent's output consistency score includes: a system role setting, used to instruct the model to perform the assessment as a senior stroke intervention expert; an inference premise field, used to embed the structured clinical feature data; an inference hypothesis field, used to embed the structured imaging feature data; a judgment task description, used to require the model to compare the logical matching degree of the structured clinical feature data and the structured imaging feature data in neuroanatomy; and output constraints, used to specify the consistency score format and risk level interval division criteria of the model output.
[0051] For example, the prompt word template for the cross-modal consistency verification agent to obtain a consistency score includes: "[System Role]: You are a senior stroke intervention expert, responsible for verifying the consistency between clinical signs and imaging characteristics."
[0052] [Inference Premise - Clinical Signs]: {Structured Medical Record Text} [Inference Hypothesis - Image Features]: {Primary Image Agent Recognition Results} [Judgment Task]: Please compare clinical signs and imaging features to assess the degree of logical match between the two in neuroanatomy.
[0053] [Output Constraints]: Consistency score: Please give a score between 0 and 1.
[0054] Risk level range: [0.8, 1.0]: High consistency (imaging strongly supports vital signs).
[0055] [0.4, 0.8): Uncertain or insufficient information (imaging signs are not obvious, observation is recommended).
[0056] [0,0.4): Logical conflict (severe clinical symptoms but negative or contradictory imaging results).
[0057] Recommended approach: If the score is below 0.8, please generate specific visual query instructions, specifying the anatomical areas requiring review and any suspected signs. It should be noted that when the processing suggestions output by the consistency score prompt word template include visual query instructions, the system outputs visual query instructions based on the prompt word template of the visual query instructions.
[0058] Specifically, the visual query instruction prompt template generated by the cross-modal consistency verification agent includes: system role settings, used to instruct the model to perform assessments as a senior stroke intervention expert; input information, used to provide the clinical evidence and preliminary imaging conclusions required for reasoning, including abnormal clinical signs (core symptom descriptions extracted from structured clinical features (e.g., laterality, type, severity)) and preliminary imaging conclusions (preliminary findings extracted from structured imaging feature data (e.g., abnormal item prompts, anatomical location)); neuroanatomical prior rules, encapsulating neuroanatomical prior knowledge as the rule basis for model reasoning; and output task descriptions, used to specify the structured visual query instructions and output format constraints of the model output. It should be noted that the neuroanatomical prior rules can be dynamically adjusted according to clinical practice without modifying the model parameters.
[0059] For example, the visual query instruction generated by the cross-modal consistency verification agent is: "[System Role]: You are a senior expert in neuroanatomy and neuroimaging. The system has detected a "clinical-imaging" logical conflict (consistency score <0.8), and you are required to use neuroanatomical principles to reverse-map clinical signs into specific imaging-oriented review instructions."
[0060] [Input Information]: Abnormal clinical signs: {Core symptoms in structured clinical features, such as left-sided hemiplegia} Primary image conclusion: {Conclusions output by the primary image agent, such as: No obvious abnormal density shadows observed globally} [Preset Mapping Rule Base]: Lateral crossover rule: signs on the left side correspond to lesions in the right cerebral hemisphere; signs on the right side correspond to lesions in the left cerebral hemisphere.
[0061] Blood supply area and function mapping rules: For hemiplegia / motor disorders, focus on the blood supply area of the middle cerebral artery (MCA) (including the basal ganglia, internal capsule, and insular cortex); for aphasia, focus on the dominant hemisphere (usually the left) frontotemporal-parietal junction; for ataxia, focus on the cerebellum and brainstem.
[0062] [Output Task]: Please strictly follow the above preset mapping rules to generate a structured "visual query instruction" to drive the visual refocusing agent.
[0063] [Output format constraints]: Target lateral: {left / right / bilateral} Target anatomical region identification: {Specific brain region identified} Image feature type to be verified: {subtle anomalous features requiring targeted magnification by a visual agent} Mapping logic description: {Briefly describe the anatomical rules upon which this instruction is derived} For example, regarding ischemic stroke: Clinical trigger: The patient has complete right-sided hemiplegia, 1 hour after onset, and the initial imaging assessment is "no obvious hemorrhage"; Rule mapping result: Target side: Left cerebral hemisphere; Anatomical landmark: Blood supply area of the left middle cerebral artery (MCA). Verification signs: Verify whether there is a high-density arterial sign in the M1 segment of the left MCA, and whether there is blurred margin in the left basal ganglia; Generated visual query instruction: "Is there a high-density vascular sign in the left middle cerebral artery area of the image, and a decrease in gray-white matter contrast caused by early ischemia?". For hemorrhage screening: Clinical trigger: The patient suddenly experiences severe headache accompanied by confusion, with an NIHSS score of 20; Rule mapping result: Target side: Whole brain / midline region; Anatomical landmark: Ventricular system, subarachnoid space; Verification signs: High-density shadow, midline shift; Generated visual query instruction: "Clinical indication is high risk. Please focus on checking whether there are subtle high-density hemorrhage signs in the ventricles and subarachnoid space, and assess whether the midline structure is shifted."
[0064] In implementation, the visual refocusing and feature reverse lookup agent includes: The preceding data processing module is used to locate and extract the three-dimensional region of interest to be reviewed from the three-dimensional non-enhanced CT image based on the visual query command, and generate a sub-tensor of the region; The fine-grained multimodal reasoning module is used to perform targeted feature extraction and semantic understanding of the three-dimensional region of interest to be reviewed based on the visual query command and the sub-tensor of the three-dimensional region of interest, and generate structured reverse query results.
[0065] In specific implementation, the preceding data processing module generates the sub-tensor of the 3D region of interest to be reviewed in the following manner: Extract the target laterality, target anatomical region identifier, and image sign type to be verified from the visual query command, and obtain a standardized three-dimensional image tensor based on the three-dimensional non-enhanced CT image; Based on the target anatomical region identifier, determine its three-dimensional coordinate range in the standardized three-dimensional image tensor; Based on the three-dimensional coordinate range, the corresponding sub-tensor is cropped from the standardized three-dimensional image tensor, and the sub-tensor is resampled at high resolution to output the enhanced sub-tensor as the sub-tensor of the three-dimensional region of interest to be verified.
[0066] Specifically, based on the target anatomical region identifier, its three-dimensional coordinate range in the standardized three-dimensional image tensor is determined in the following manner: From the structured image feature data output by the primary image interpretation agent, find outliers that match the target anatomical region identifier, and reuse the outlier coordinate range contained in the outlier as the three-dimensional coordinate range. or, A lightweight 3D segmentation model is invoked to segment the voxel region corresponding to the target anatomical region identifier in real time from the standardized 3D image tensor, and the coordinates of the bounding rectangle of the region are calculated to obtain the 3D coordinate range.
[0067] In practical implementation, the fine-grained multimodal inference module is built on a large multimodal model. Its model architecture is the same as that of the multimodal inference unit in the primary image interpretation agent, consisting of three parts: a 3D visual encoder, a modal alignment layer, and a large language model decoder. The 3D visual encoder is used to extract features from the sub-tensors of the region of interest and generate a sequence of visual features of the region. The modal alignment layer is used to perform cross-modal alignment between the textual semantics of the visual query instruction and the sequence of visual features of the region to generate aligned features. The large language model decoder is used to generate structured reverse lookup results in an autoregressive manner based on the aligned features.
[0068] Specifically, the structured reverse lookup results include refined image features and local condition assessment of the verification area; refined image features involve subdividing the target anatomical region according to predefined sub-region division rules, and providing density and morphological features for each sub-region; local condition assessment includes the existence judgment of the image features to be verified, a detailed description of the features, and a confidence score. The sub-region division of the target anatomical region is based on predefined detailed anatomical atlases or expert knowledge, such as the AAL atlas or the Harvard-Oxford atlas.
[0069] For example, for the "right basal ganglia region", the refined imaging features can be further subdivided into the lentiform nucleus, the head of the caudate nucleus, and the anterior limb of the internal capsule; the local condition assessment indicates the presence of the target sign, described as "slightly decreased density is visible in the right lentiform nucleus and the anterior limb of the internal capsule, with blurred gray-white matter boundaries, consistent with ultra-early ischemic changes", with a confidence level of 0.85.
[0070] Specifically, the training method for the fine-grained multimodal inference module is the same as that for the multimodal large-model inference unit in the primary image interpretation agent, employing a hybrid-source multi-task training mechanism, including four tasks: report generation, visual question answering, target localization, and cross-modal alignment. The difference lies in the fact that the training samples for the fine-grained multimodal inference module are designed for targeted review scenarios, with more refined question-answer pairs, focusing on the recognition of subtle signs in specific anatomical regions.
[0071] More specifically, the training samples in the fine-grained multimodal inference module include region-of-interest sub-tensors, visual query commands, question-answer pair labels, and reverse query result labels. Each training sample's question-answer pair label contains multiple fine-grained question-answer pairs, each targeting a sub-region. For example, the question "Is there a subtle decrease in density in the right basal ganglia?" corresponds to an answer of "yes" or "no." These question-answer pairs are either expert-annotated or automatically generated based on rules. The sample data sources include typical cases from publicly available datasets (cases with clear lesions but concealed early signs), private clinical data mining (screening cases with "typical symptoms but initially negative" from historical medical records, with expert annotation of verification areas and refined features), and the automatic generation of query-answer pairs from existing labeled data using neuroanatomical rules.
[0072] More specifically, the loss function in the training of the fine-grained multimodal inference module. Represented as: ; In the formula, , , , These represent the local report generation loss, local visual question answering loss, local target localization loss, and local cross-modal alignment loss, respectively. , , , These represent the weight coefficients of the local report generation loss, local visual question answering loss, local target localization loss, and local cross-modal alignment loss, respectively, used to balance the contributions of different tasks. The weight coefficients are usually dynamically adjusted according to the convergence difficulty and importance of the tasks, and can be determined by grid search or adaptive adjustment methods (such as uncertainty weighting). In the early stage of training, they can be set to equal weights (such as 0.25 for each), and then fine-tuned according to the performance on the validation set.
[0073] Furthermore, localized reporting generates losses. Represented as: ; In the formula, This indicates a sub-region index, which refers to the functional or structural sub-regions (e.g., lentiform nucleus, caudate nucleus head, anterior limb of internal capsule, etc.) of a target anatomical region (e.g., basal ganglia region). Subregion Whether it is the focus of the current visual query command. If the command requires verification of this sub-region (e.g., the command explicitly mentions "lentiform nucleus and anterior limb of internal capsule"), then ,otherwise , ; This represents the sum of all subregions within the target anatomical region; Subregion The total number of text tokens, that is, the number of tokens contained in the text description generated for this sub-region in the structured reverse lookup results; Subregion The first in the description text The true value of each word element; Subregion The description text before A sequence of lexical units; This represents a sequence of visual features extracted from the subtensor of the region of interest; This represents a predefined sequence of query terms; This indicates that the model, given preorder terms, visual features, and query terms, predicts the first term. Each lexical element is the true value. The probability of.
[0074] Understandably, this loss forces the model to accurately reflect the true image features of a specific sub-region when generating textual descriptions of that sub-region. By only considering... The loss is calculated for sub-regions of 1. The model focuses its learning on the subtle structures that the instructions require, ignoring irrelevant regions, thereby improving the accuracy of describing subtle lesion signs. At the same time, the autoregressive mechanism ensures that the generated text conforms to medical semantics and structured format.
[0075] Furthermore, local visual question-answering loss Represented as: ; in, Indicates about sub-regions The correct answer to the question, where, for classification questions, For category indexes of the correct answers; for generative questions, Let be the word sequence of the answer. In this case, the loss is the average of the sequence cross-entropy. Here, the local visual question answering loss is illustrated using a classification problem as an example. In the formula, The question portion of the visual query instruction (e.g., "Is there a subtle decrease in density in the right basal ganglia region?") is usually represented in text form, encoded, and then input into the model. The model represents the solution to a given problem. and visual feature sequences The probability of outputting the correct answer.
[0076] Understandably, this loss function enables the model to answer medical questions targeting specific sub-regions, enhancing its discriminative understanding of subtle signs. Through this loss function, the model only needs to focus on the sub-regions relevant to the question, avoiding interference from irrelevant information and thus improving the accuracy of question answering. This fine-grained question-answering capability helps verify specific hypotheses in subsequent inference (such as "whether there are early signs of ischemia").
[0077] Furthermore, local target localization loss Represented as: ; In the formula, Represents the true bounding box coordinates of the sub-region, that is, the bounding box coordinates of the sub-region (such as the bean-shaped kernel) in the normalized 3D image tensor; This represents the coordinates of the bounding box of the sub-region predicted by the model. Indicates Smooth L1 loss, Indicates the 3D GIoU loss. , These represent the first balance coefficient and the second balance coefficient, respectively; where, Used to measure the degree of overlap and containment relationship between predicted bounding boxes and ground truth bounding boxes. For example, , Set them to 1 and 0.5 respectively, or fine-tune them based on the validation set.
[0078] Understandably, this loss enables the model to accurately locate the spatial position of each sub-region, providing a foundation for subsequent morphological feature calculations (such as volume and maximum diameter). The Smooth L1 loss ensures the accuracy of coordinate regression, while the GIoU loss ensures the overall fit between the predicted bounding box and the ground truth bounding box.
[0079] Furthermore, cross-modal alignment loss Represented as: ; In the formula, A sequence of visual features representing the region of interest in the current sample; This represents the positive sample feature sequence obtained after data augmentation (such as rotation, cropping, and intensity adjustment) of the region of interest of the current sample; Indicates the first in the same training batch A sequence of visual features for each sample, including positive and negative samples; sim() represents the cosine similarity function; This represents the temperature coefficient, which controls the smoothness of the contrast loss, and is usually set to 0.07. This indicates the training batch size.
[0080] Understandably, this loss employs a contrastive learning mechanism, bringing features from different enhanced views of the same region of interest closer together and pushing features from different regions of interest further apart. This makes the model less sensitive to grayscale changes and geometric perturbations in the image, resulting in more stable extracted features. Spacing the features of different regions of interest apart in space facilitates downstream tasks in distinguishing different regions from different lesions. Through contrastive learning, the visual feature space and the semantic space of the language model are initially aligned, providing better initialization for the modality alignment layer and thus improving the accuracy of text generation.
[0081] It is understandable that visual query commands and the triggered image reverse lookup process exist as an independent information processing stage in the system. Their output, in the form of structured features, participates in subsequent collaborative reasoning, rather than directly replacing or covering the primary image interpretation results. By introducing image reverse lookup results as supplementary evidence, targeted information enhancement for specific medical questions can be achieved while maintaining the integrity of the primary image analysis results. This approach avoids the instability caused by repeatedly adjusting attention weights during a single reasoning process, enabling the system to maintain a controllable and traceable analysis path even in complex or high-risk scenarios.
[0082] In implementation, the multi-factor collaborative reasoning and decision-making intelligent agent includes: The information fusion unit is used to splice together the structured image feature data, structured clinical feature data, and / or structured reverse lookup results to form multi-source heterogeneous information evidence; The counterfactual reasoning unit is used to generate preliminary diagnostic conclusions based on multi-source heterogeneous information evidence and built-in clinical diagnosis and treatment guidelines; it is also used to construct counterfactual hypotheses that contradict the preliminary diagnostic conclusions, and to test the rationality of multi-source heterogeneous information evidence under these hypotheses, thereby obtaining verification results. The security gating unit is used to traverse the contraindication-related fields in the structured clinical feature data, compare them with preset clinical treatment safety rules, and generate structured contraindication verification results. The results generation unit is used to generate the final auxiliary evaluation results based on multi-source heterogeneous information evidence, preliminary diagnostic conclusions and their verification results, and structured contraindication verification results.
[0083] Specifically, the information fusion unit concatenates the input structured image feature data, structured clinical feature data, and / or structured reverse lookup results into a unified text context, which serves as the input for subsequent reasoning. The concatenation method uses a predefined template to organize data from different sources according to clinical logic, ensuring the integrity and readability of the information.
[0084] For example, the splicing template is as follows: "
Clinical Information
[0085] Specifically, the preliminary diagnostic conclusion includes the type of lesion, its possible location, severity, and initial confidence level; the verification results include whether the verification is passed, conflicting evidence, and interpretation.
[0086] Specifically, the counterfactual reasoning unit is implemented using a large language model and driven by a set of prompt word templates. Clinical treatment guidelines are embedded in the prompt words in text form as the basis for reasoning. The construction of the opposing hypothesis can be based on a preset template (for example, if the preliminary diagnosis is ischemic stroke, the opposing hypothesis is hemorrhagic stroke or stroke mimicry; if it is cerebral hemorrhage, the opposing hypothesis is ischemic stroke with hemorrhagic transformation), or it can be dynamically generated by the model based on preliminary conclusions and evidence.
[0087] For example, for prompt words indicating a preliminary diagnostic conclusion: [System Role] You are a senior neurology and neuroimaging expert, responsible for synthesizing all clinical and imaging evidence to generate a preliminary diagnosis of cerebrovascular disease.
[0088] [Enter information] Clinical features: {Natural language description of structured clinical features} Image features: {Natural language description of structured image features} If there are reverse lookup results, add: Reverse lookup results: {Natural language description of the reverse lookup results}} [Treatment Guidelines] Refer to the following guidelines: For acute ischemic stroke: intravenous thrombolysis can be considered for patients with onset time <4.5 hours and no contraindications; thrombectomy can be considered for large vessel occlusion.
[0089] Cerebral hemorrhage: Urgent blood pressure control is required, surgical indications should be assessed, and anticoagulation-related bleeding should be reversed.
[0090] Stroke mimicry: Hypoglycemia, epilepsy, etc. need to be differentiated.
[0091] [Task] Based on the above evidence, please generate a preliminary diagnostic conclusion. The output format is as follows: { "preliminary_diagnosis": "Lesion type and location", "severity": "severity level" "preliminary_confidence": "high / medium / low", "reasoning_summary": Brief reasoning basis }".
[0092] For example, for the prompt words used to obtain the verification result in counterfactual verification: [System Role] You continue in the role of expert and now need to perform counterfactual reasoning to verify the robustness of the initial conclusions.
[0093] [Preliminary Conclusion] {Preliminary Diagnosis} [Opposing hypothesis] Please construct an alternative diagnosis that is most likely to be confused with the initial conclusion (e.g., if the initial diagnosis is ischemic, the alternative hypothesis is hemorrhagic or stroke mimicry). The alternative hypothesis is: {Automatically generated alternative hypothesis} [Evidence Review] Clinical features: {structured imaging feature data} Imaging features: {Structured clinical feature data} Reverse lookup results: {Structured reverse lookup results} [Task] Please assess: If the opposing hypothesis holds true, does the existing evidence support or strongly contradict it? Specifically consider: Do the clinical signs and symptoms appear typical and support the opposing hypothesis? Do the image features conform to the typical manifestations of the opposing hypothesis? Is there any evidence that completely rules out the opposing hypothesis? Please output: { "counterfactual_plausibility": "low / medium / high", / / The plausibility of the opposing assumptions. "conflicting_evidence": "List the key pieces of evidence that conflict with the opposing hypothesis", "supports_original": "true / false", / / Whether to pass the counterfactual reasoning check "explanation": Detailed explanation }".
[0094] It should be noted that if the counterfactual reasoning unit fails the verification, the initial confidence level of the preliminary diagnosis conclusion will be marked as "low", and the final output will indicate that there is a logical conflict, while prompting manual intervention for review.
[0095] In practice, the security gating unit generates structured contraindication verification results in the following manner: Based on the lesion type in the preliminary diagnosis and in conjunction with the built-in clinical guidelines, the corresponding standard treatment direction is determined. For example, if the diagnosis is "acute ischemic stroke", the treatment direction is "intravenous thrombolysis" or "mechanical thrombectomy"; if the diagnosis is "cerebral hemorrhage", the treatment direction is "controlling blood pressure, evaluating surgery", etc. The contraindication-related fields in the structured clinical feature data are traversed and compared with preset clinical treatment safety rules to determine whether there are contraindications for the current treatment direction, thereby generating a structured contraindication verification result. The structured contraindication verification result includes whether there is a conflict, the type of conflict (absolute or relative), the conflicting fields, and a specific description.
[0096] Specifically, the contraindication fields should include at least: recent major surgery history, history of intracranial hemorrhage, history of anticoagulant use, coagulation dysfunction, and uncontrolled severe hypertension. Clinical treatment safety rules are based on authoritative stroke diagnosis and treatment guidelines and should include at least the logic for determining absolute and relative contraindications for thrombolytic therapy. For example: absolute contraindications include: history of intracranial hemorrhage, recent major surgery, active bleeding, etc.; relative contraindications include: minor stroke, rapid recovery, seizures, etc.
[0097] In practice, the final auxiliary assessment results generated by the results generation unit include the final diagnosis, severity classification, final treatment recommendations, reasoning path description, logical confidence level, safety warning list, and whether manual review is required.
[0098] Specifically, the final diagnosis is based on the preliminary diagnosis and the counterfactual verification result. If the counterfactual verification result is passed, the final diagnosis is consistent with the preliminary diagnosis. If it fails, the confidence level of the preliminary diagnosis is reduced to "low", and the preliminary diagnosis is retained in the final diagnosis but a manual review is requested. Alternatively, if the evidence is sufficient, the conclusion may be adjusted and explained in the reasoning path.
[0099] Specifically, the severity grading is based on the severity at the time of the initial diagnosis. The severity value in the initial diagnosis is determined based on morphological parameters in imaging features (such as lesion volume, mass effect, midline shift, etc.) and neurological deficit scores in clinical features (such as NIHSS scores), combined with the grading criteria in authoritative stroke guidelines.
[0100] Specifically, the final treatment recommendation safety gating unit is obtained based on the standard treatment direction and the structured contraindication verification results. If there is an absolute contraindication, it clearly states "not recommended" and gives the reason; if there is a relative contraindication, it explains that clinical consideration is required; if there is no contraindication, it gives the standard treatment recommendation.
[0101] Specifically, the reasoning path description is automatically generated by the result generation unit based on the key nodes of the entire reasoning process, using natural language to connect the following information: core clinical evidence (such as symptom lateralization and onset time), key imaging findings (such as abnormalities and reverse lookup results), counterfactual verification conclusions (which opposing hypotheses were ruled out), and safety gating results (whether contraindications exist). The generation process follows a preset template, converting structured data into coherent text descriptions, aiming to provide clinicians with transparent and traceable decision-making basis, such as "clinical presentation of right hemiplegia, initial imaging suspicion of low density in the left basal ganglia, confirmed by reverse lookup as signs of very early ischemia; counterfactual hypothesis of hemorrhage was ruled out; no contraindications for thrombolysis."
[0102] Specifically, the logical confidence level reflects the assessment of the reliability of the final conclusion, integrating information from multiple dimensions. The initial value is taken from the preliminary confidence level in the preliminary diagnosis, and then dynamically adjusted by the result generation unit based on the following factors: if the counterfactual verification fails, the confidence level is forcibly reduced to "low"; if key clinical or imaging fields are missing (such as unknown onset time), the confidence level is appropriately reduced; if the reverse lookup results provide strong evidence, the confidence level is maintained or increased; the final output is in three levels: "high", "medium", and "low".
[0103] Specifically, the safety alert list is the result of structured contraindication verification. The alert information is intended to remind clinicians of potential treatment risks and ensure the safety of the final recommendation.
[0104] Specifically, the manual review flag is automatically set by the result generation unit based on various high-risk conditions. Triggering conditions include: failure of counterfactual verification, existence of complex relative contraindications requiring clinical consideration, lack of key information leading to low confidence, or detection of unresolved logical conflicts within the system. If any of these conditions are met, manual review is required, prompting the physician to intervene; otherwise, manual review is not necessary. This mechanism ensures that in situations of high uncertainty, the system does not replace clinical decision-making, but rather leaves the final judgment to the physician.
[0105] Compared with existing technologies, this embodiment provides a cerebrovascular lesion auxiliary assessment system based on multimodal intelligent agents. Through the parallel processing of primary image interpretation intelligent agents and clinical information structuring intelligent agents, it achieves deep structured analysis of multimodal data, providing a standardized evidentiary basis for subsequent reasoning. The cross-modal consistency verification intelligent agent evaluates the logical matching degree between images and clinical features through natural language reasoning, and automatically generates visual query instructions when conflicts or uncertainties are detected. This drives the visual refocusing and feature reverse lookup intelligent agents to conduct targeted verification of suspicious areas, effectively uncovering hidden signs of ultra-early and micro lesions, and significantly improving the sensitivity and accuracy of lesion detection. The multi-factor collaborative reasoning and decision-making intelligent agent verifies the logical robustness of conclusions through counterfactual reasoning and automatically screens for treatment contraindications with a security gating mechanism. Under the premise of ensuring clinical safety, it generates auxiliary assessment results containing diagnostic conclusions, treatment suggestions, and interpretable reasoning paths, thereby effectively reducing the risk of missed diagnosis and misdiagnosis in complex and high-risk cases. It provides reliable, transparent, and traceable intelligent support for clinical decision-making, and is particularly suitable for auxiliary diagnosis, classification, and thrombolysis suitability assessment in emergency stroke scenarios.
[0106] Example 2 A specific embodiment of the present invention discloses an auxiliary assessment method for cerebrovascular lesions based on multimodal intelligent agents, such as... Figure 2 As shown, it includes: Acquire three-dimensional non-contrast CT image data of the patient's cerebral blood vessels and corresponding electronic medical record text data, and generate structured image feature data and structured clinical feature data. A medical logic consistency assessment is performed on the structured image feature data and the structured clinical feature data to obtain a consistency assessment result; if the consistency assessment result meets the preset conditions, a visual query instruction is generated; the three-dimensional non-enhanced CT image is then re-examined according to the visual query instruction to generate a structured reverse query result; By integrating the structured image feature data, structured clinical feature data, and / or structured reverse lookup results, and through counterfactual reasoning and security gating mechanisms, the final auxiliary assessment result is generated.
[0107] The specific implementation process of this invention can be found in the above system embodiments, and will not be repeated here.
[0108] Since this embodiment is based on the same principle as the above system embodiment, this method also has the corresponding technical effects of the above method embodiment.
[0109] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0110] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-modal agent-based cerebral vascular lesion assisted evaluation system, characterized in that, include: The data processing module is used to acquire three-dimensional non-enhanced CT image data of the patient's cerebral blood vessels and the corresponding electronic medical record text data, and generate structured image feature data and structured clinical feature data. A cross-modal consistency verification agent is used to perform medical logical consistency assessment on the structured image feature data and structured clinical feature data to obtain consistency assessment results; if the consistency assessment results meet preset conditions, a visual query instruction is generated and output to the visual refocusing and feature reverse lookup agent. The cross-modal consistency verification agent is built based on a large language model and employs a natural language reasoning framework. It uses structured clinical feature data as a reasoning premise and structured imaging feature data as a reasoning hypothesis to assess the degree of logical matching in neuroanatomy and output a consistency assessment result. When the consistency assessment result meets preset conditions, the cross-modal consistency verification agent generates a visual query instruction based on neuroanatomical prior rules and structured clinical feature data. The visual query instruction includes the target side, the target anatomical region identifier, and the type of imaging sign to be verified. The consistency assessment result is a quantified consistency score. The consistency score is mapped to multiple predefined risk level intervals, each including at least: a first interval indicating a high degree of consistency between structured imaging feature data and structured clinical feature data; a second interval indicating uncertainty or insufficient evidence between structured imaging feature data and structured clinical feature data; and a third interval indicating a logical conflict between structured imaging feature data and structured clinical feature data. The preset condition is that the consistency score falls within the second or third interval. A visual refocusing and feature reverse lookup intelligent agent is used to perform directional verification of the three-dimensional non-enhanced CT image according to the visual query instruction, and generate structured reverse lookup results; the visual refocusing and feature reverse lookup intelligent agent includes: The preceding data processing module is used to locate and extract the three-dimensional region of interest to be reviewed from the three-dimensional non-enhanced CT image based on the visual query command, and generate a sub-tensor of the region; The fine-grained multimodal reasoning module is used to perform targeted feature extraction and semantic understanding of the three-dimensional region of interest to be verified based on the visual query command and the sub-tensor of the three-dimensional region of interest, and generate structured reverse lookup results; wherein, the structured reverse lookup results include the refined image features and local state evaluation of the verified region; A multi-factor collaborative reasoning and decision-making intelligent agent is used to fuse the structured image feature data, structured clinical feature data, and / or structured reverse lookup results, and generate the final auxiliary evaluation result through counterfactual reasoning and security gating mechanisms.
2. The cerebrovascular lesion auxiliary assessment system based on multimodal intelligent agents according to claim 1, characterized in that, The data processing module includes: The data receiving module is used to acquire three-dimensional non-enhanced CT image data of the patient's cerebral blood vessels and the corresponding electronic medical record text data; A primary image interpretation agent is used to perform structured analysis on the three-dimensional non-enhanced CT images, extract image features related to cerebrovascular lesions, and generate structured image feature data. A clinical information structured intelligent agent is used to perform semantic parsing on the electronic medical record text, converting unstructured medical record information into structured clinical feature data with predefined fields.
3. The cerebrovascular lesion auxiliary assessment system based on multimodal intelligent agents according to claim 2, characterized in that, The primary image interpretation agent includes: The preprocessing unit is used to perform physical value conversion, spatial standardization, and 3D data construction on 3D non-enhanced CT images to generate standardized 3D image tensors. Anatomical partitioning unit is used to divide a standardized 3D image tensor into multiple anatomical regions based on predefined brain anatomical partitioning rules, generating regional feature data with anatomical region identifiers; The multimodal large model inference unit is used to extract features from each anatomical region based on regional feature data with anatomical region identifiers and generate structured image feature data.
4. The cerebrovascular lesion auxiliary assessment system based on multimodal intelligent agents according to claim 3, characterized in that, The structured image feature data includes at least one anomaly, and each anomaly includes structured image feature data such as anomaly indication, anatomical location, anomaly coordinate range, density features, and morphological features; The multimodal large model inference unit is obtained by training a multimodal large model; wherein, a hybrid source multi-task training mechanism is used for training, and the training tasks are report generation, visual question answering, target localization and cross-modal alignment.
5. The cerebrovascular lesion auxiliary assessment system based on multimodal intelligent agents according to claim 2, characterized in that, The structured intelligent agent for clinical information is built based on a large language model. Through prompt word engineering or supervised fine-tuning, it extracts key information from electronic medical record text and converts it into standardized structured clinical feature data. The structured clinical feature data includes the patient's basic demographic information, stroke-related key timelines, clinical manifestations and neurological assessments, past medical history and risk factors, contraindications and medication history, laboratory and imaging examination results, logical consistency and data quality evaluation.
6. The cerebrovascular lesion auxiliary assessment system based on multimodal intelligent agents according to claim 1, characterized in that, The preceding data processing module generates the sub-tensor of the 3D region of interest to be reviewed in the following manner: Extract the target laterality, target anatomical region identifier, and image sign type to be verified from the visual query command, and obtain a standardized three-dimensional image tensor based on the three-dimensional non-enhanced CT image; Based on the target anatomical region identifier, determine its three-dimensional coordinate range in the standardized three-dimensional image tensor; Based on the three-dimensional coordinate range, the corresponding sub-tensor is cropped from the standardized three-dimensional image tensor, and the sub-tensor is resampled at high resolution to output the enhanced sub-tensor as the sub-tensor of the three-dimensional region of interest to be verified.
7. The cerebrovascular lesion auxiliary assessment system based on multimodal intelligent agents according to claim 1, characterized in that, The multi-factor collaborative reasoning and decision-making intelligent agent includes: The information fusion unit is used to splice together the structured image feature data, structured clinical feature data, and / or structured reverse lookup results to form multi-source heterogeneous information evidence; The counterfactual reasoning unit is used to generate preliminary diagnostic conclusions based on multi-source heterogeneous information evidence and built-in clinical diagnosis and treatment guidelines; it is also used to construct counterfactual hypotheses that contradict the preliminary diagnostic conclusions, and to test the rationality of multi-source heterogeneous information evidence under these hypotheses, thereby obtaining verification results. The security gating unit is used to traverse the contraindication-related fields in the structured clinical feature data, compare them with preset clinical treatment safety rules, and generate structured contraindication verification results. The results generation unit is used to generate the final auxiliary assessment results based on multi-source heterogeneous information evidence, preliminary diagnostic conclusions and their verification results, and structured contraindication verification results. The final auxiliary assessment results include the final diagnostic conclusion, severity classification, final treatment recommendations, explanation of reasoning path, logical confidence level, safety warning list, and whether manual review is required.
8. A method for auxiliary assessment of cerebrovascular lesions based on multimodal intelligent agents, characterized in that, Includes the following steps: Acquire three-dimensional non-contrast CT image data of the patient's cerebral blood vessels and corresponding electronic medical record text data, and generate structured image feature data and structured clinical feature data. A medical logical consistency assessment is performed on the structured image feature data and the structured clinical feature data to obtain a consistency assessment result; if the consistency assessment result meets the preset conditions, a visual query instruction is generated. Based on a large language model, a natural language reasoning framework is adopted. Structured clinical feature data is used as a reasoning premise, and structured imaging feature data is used as a reasoning hypothesis. The degree of logical matching in neuroanatomy is evaluated, and a consistency assessment result is output. When the consistency assessment result meets preset conditions, a visual query instruction is generated based on neuroanatomical prior rules and structured clinical feature data. The visual query instruction includes the target side, the target anatomical region identifier, and the type of imaging sign to be verified. The consistency assessment result is a quantified consistency score. The consistency score is mapped to multiple predefined risk level intervals, each including at least: a first interval indicating a high degree of consistency between structured imaging feature data and structured clinical feature data; a second interval indicating uncertainty or insufficient evidence between structured imaging feature data and structured clinical feature data; and a third interval indicating a logical conflict between structured imaging feature data and structured clinical feature data. The preset condition is that the consistency score falls into either the second or third interval. The three-dimensional non-enhanced CT image is subjected to targeted verification according to the visual query command to generate a structured reverse lookup result; this includes: locating and extracting the three-dimensional region of interest to be verified from the three-dimensional non-enhanced CT image based on the visual query command, and generating a sub-tensor of the region; based on the visual query command and the sub-tensor of the three-dimensional region of interest, targeted feature extraction and semantic understanding are performed on the three-dimensional region of interest to be verified to generate a structured reverse lookup result; wherein, the structured reverse lookup result includes refined image features and local state assessment of the verified region; By integrating the structured image feature data, structured clinical feature data, and / or structured reverse lookup results, and through counterfactual reasoning and security gating mechanisms, the final auxiliary assessment result is generated.