Intelligent medical image diagnosis system and method based on hierarchical cross-modal conversion and dynamic feature tracking

The intelligent medical image diagnostic system, which utilizes hierarchical cross-modal transformation and dynamic feature tracking, solves the problems of static analysis limitations, insufficient credibility, and response delays in existing technologies. It achieves multi-level verification and rapid response, improves diagnostic accuracy and compliance, and supports multimodal fusion and teaching applications.

CN120998469APending Publication Date: 2025-11-21HANGZHOU MAGIC BYTE TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511170789.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing medical imaging AI systems suffer from limitations in static analysis, insufficient clinical credibility, multimodal fragmentation, legal and compliance risks, and response delays. They are unable to effectively and dynamically track lesion evolution, simulate doctors' thinking logic, establish cross-modal associations, meet medical regulatory requirements, and respond quickly to emergencies.

Method used

An intelligent medical image diagnostic system employing hierarchical cross-modal transformation and dynamic feature tracking achieves multi-level verification and optimized diagnosis of image and text features through lightweight text preprocessing, multimodal analysis, and collaborative decision-making using a large language model, combined with a dynamic lesion tracking engine, multi-expert collaborative decision-making, and cross-modal semantic alignment mechanism.

Benefits of technology

It significantly improves diagnostic accuracy and reliability, optimizes system efficiency and response speed, enhances medical compliance, supports multimodal integration and teaching applications, and meets clinical needs and legal requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998469A_ABST
    Figure CN120998469A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent medical image diagnosis system and method based on hierarchical cross-modal conversion and dynamic feature tracking. The system adopts three-step cross-modal conversion: a first-layer small model for converting user questions to realize medical ontology matching; the second-layer multi-modal model extracts image features, and outputs text states such as JSON data with focus coordinates, density and other features; and the third-layer large model fuses the medical history and the image features to generate diagnosis suggestions, and credibility verification is carried out. A dynamic focus tracking engine is introduced, a focus evolution rule of multiple scanning is analyzed through a convolutional network, and an optical flow field is adopted to compensate artifacts. The system also integrates a multi-expert voting mechanism to simulate a clinical consultation process, and outputs consensus diagnosis and objection viewpoints. A hierarchical routing algorithm is designed for emergency treatment scenes, so that the recognition response time of emergencies such as pneumothorax is shortened. Further, the system automatically generates a full chain of evidence report that conforms to medical regulations, including a model version, a guide reference, and a data hash value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image artificial intelligence analysis technology, specifically relating to an intelligent medical image diagnostic system and method based on hierarchical cross-modal transformation and dynamic feature tracking. This system is applicable to the automated analysis of multimodal medical images such as CT, MRI, and X-rays. By integrating temporal lesion evolution tracking, multi-level cross-modal transformation, credibility verification, and emergency response optimization, it achieves an integrated solution for assisting clinical diagnosis, teaching and training, and generating compliant reports. Background Technology

[0002] Current medical imaging AI systems suffer from the following urgent technical deficiencies that need to be addressed: 1. Limitations of static analysis: Existing technologies (such as single-model architectures like U-Net and ResNet) can only perform planar or 3D analysis on single scan images, and cannot dynamically track the evolution of lesions over time (such as tumor growth rate and treatment effect evaluation).

[0003] 2. Insufficient clinical credibility: (1) Weak explanatory power: It is difficult to simulate the doctor's thinking logic by relying on visualization tools such as heat maps, and there is a lack of visualization of multi-level reasoning processes; (2) Multimodal fragmentation: Imaging features are processed separately from textual data such as patient history and pathology reports, and no cross-modal associations are established (such as the quantitative association between smoking history and the probability of malignancy of lung nodules). (3) Lack of confidence: The diagnostic results output by a single model have not been validated at multiple levels, resulting in a significant risk of misdiagnosis. 3. Legal and Compliance Risks: The diagnostic process lacks end-to-end evidence traceability, failing to meet the core auditability requirements of medical regulations. (1) The medical guideline version on which the model reasoning path and the basis were not recorded; (2) Lack of data integrity verification (such as DICOM file hash value); (3) No dynamic disclaimer was embedded to avoid clinical decision-making risks.

[0004] 4. Response delay: Unified computing architectures cannot differentiate between cases of urgency. Existing systems (such as conventional CNN architectures) have high average latency in recognizing emergencies like pneumothorax and cerebral hemorrhage, far exceeding the golden window for clinical treatment. The main reasons include: (1) A lightweight emergency-specific reasoning channel was not designed; (2) Respiratory motion artifacts cause feature extraction bias.

[0005] Therefore, existing technologies are insufficient to support real-world application scenarios. Summary of the Invention

[0006] This invention provides an intelligent medical image diagnostic system and method for hierarchical cross-modal transformation and dynamic feature tracking, the technical solution of which is as follows (see attached). Figure 1 (as shown) 1. System Architecture and Diagnostic Process An intelligent medical imaging diagnostic system based on hierarchical cross-modal transformation performs the following steps: Step A: Receive end-user requests (the requests themselves may contain other text-based information) and medical image data; Step B: Perform medical intent parsing and standardized rewriting of the user request using the first-layer lightweight text preprocessing model; Step C: Input the rewritten request and medical images into the second-layer multimodal model, extract image features, and generate intermediate results readable by the large model (including JSON data of lesion location coordinates, image feature descriptions, and initial confidence scores). Step D: Input the intermediate results into the third-layer large language model (including but not limited to a logic model with a larger number of parameters or a dedicated MOE inference model) for logical verification and enhanced inference, and output the final diagnostic report.

[0007] Solving technical problems: Traditional diagnostic software has low accuracy and fragmented cross-modal information.

[0008] Beneficial effects: Achieves accurate understanding of intent and multimodal collaborative decision-making through a three-step process.

[0009] 2. Core Functional Modules The system includes the following modules: Text preprocessing module: Enhances the medical standardization of user-submitted questions based on a lightweight model; Multimodal analysis module: fuses image and text features using a medium-sized model to output structured intermediate results; Result verification module: Based on a large language model, the intermediate results are verified for credibility and optimized through deep reasoning.

[0010] Dynamic lesion tracking engine (as attached) Figure 2 (As shown) For time-series images with multiple scans, perform: (A) Spatiotemporal feature extraction: (a) Construct a lesion evolution model, for example, the input is a registered temporal DICOM sequence; (b) Compensating for breathing motion artifacts using optical flow fields; (B) Quantitative output: Generate a malignancy probability curve with a time dimension (e.g., "Lung nodule volume increased by 21% in 3 months, annualized malignancy risk +22%").

[0011] 3. Standardization of Intermediate Results The intermediate results output by the multimodal analysis module can be text or various conventional formats, and in this specific implementation case, they are in JSON format, including: Lesion localization: description of spatial coordinates or anatomical location; Image features: quantified parameters such as density, edge morphology, and enhancement mode; Initial confidence assessment: confidence score based on image features (range 0–1).

[0012] 4. Query intelligent modification mechanism Before submitting a query to the large language model, the system automatically detects medical keywords, sensitive words, or intent types in the user's question and dynamically inserts appropriate follow-up questions (e.g., when "cancer" is detected, add "please provide past medical history").

[0013] 5. Credibility Iterative Optimization Mechanism The second-layer model outputs an initial confidence score for the preliminary diagnosis; The third-layer model modifies the score based on medical logic rules, and finally outputs a probabilistic diagnostic distribution (e.g., [benign: 82%, malignant: 18%]).

[0014] 6. Multi-expert collaborative decision-making subsystem (as attached) Figure 3 (As shown) The system integrates a multi-model voting mechanism to simulate the medical consultation process: Task scheduling unit: Allocates data to three specialized models: Image analysis model (extracting detailed features from images) Clinical correlation model (integrating medical history and physiological indicators) Textual reasoning model (analyzing the semantics of pathology reports) Consensus Decision Unit: Aggregates the outputs of the three models and makes voting decisions according to preset weights (e.g., if two votes are in agreement, the decision is adopted).

[0015] 7. Dynamic Multimodal Fusion Mechanism The second layer calculates the confidence scores for each modality in real time (e.g., Dice coefficients for image segmentation and F1 scores for entity recognition). Dynamically generate feature fusion weight matrices for each modality; The third-layer large model optimizes the weights a second time through an attention mechanism, outputting a diagnosis with strong noise robustness.

[0016] Emergency triage routing is attached. Figure 4 As shown.

[0017] 8. Cross-modal semantic alignment mechanism By using a contrastive learning loss function (such as InfoNCE), image feature vectors and medical text descriptions are mapped to the same semantic space, enabling cross-modal feature comparison.

[0018] 9. Dynamic Diagnostic Tree Generation When the third-layer large model receives suspected complex case data, it automatically constructs a dynamic diagnostic tree and gradually eliminates interfering factors through multiple rounds of decision branches (e.g., first distinguishing between inflammation and tumors, and then further subdividing the subtypes).

[0019] 10. Adaptive Consultation Path Method Dynamically generate consultation paths based on real-time physiological data: Hybrid decision engine: Combining language model reasoning (weight α), clinical rule base (weight β), and physiological indicators (weight γ), to calculate the optimal question sequence; Context-aware mechanism: When the monitored parameters are abnormal (such as blood oxygen <95%), relevant follow-up questions ("Are you having difficulty breathing?") are triggered first.

[0020] The overall system deployment architecture diagram is attached. Figure 5 As shown. In a specific implementation case, the system includes a cloud-based SaaS service module (containing a layered AI engine) ←→ edge devices (embedded deployment for CT / MRI) ←→ user terminals (see report).

[0021] Beneficial effects This invention achieves the following significant advancements in the field of medical image diagnosis through the synergistic innovation of a hierarchical AI architecture and dynamic feature tracking: 1. Improve diagnostic accuracy and reliability (1) Multi-level verification mechanism: Through the division of labor and cooperation of the small-medium-large three-stage model, the progressive analysis from feature extraction to logical verification is realized, which significantly reduces the risk of misdiagnosis; (2) Multi-expert consensus support: cross-validation of heterogeneous models such as CNN, GNN and Transformer is integrated to simulate the multidisciplinary consultation process in clinical practice, and the clinical acceptability of the output diagnostic conclusions is significantly improved. (3) Dynamic evolution analysis: The spatiotemporal 4D convolutional network has the ability to track lesions across cycles, which solves the problem of misjudgment of evolution patterns caused by traditional static analysis.

[0022] 2. Optimize system efficiency and response speed (1) Efficient allocation of computing resources: The layered architecture avoids large models from directly processing raw image data. The intermediate cross-modal processing transformation results are transmitted through structured text or JSON, which greatly reduces redundant computing. (2) Emergency priority routing: a dedicated lightweight channel for emergencies such as pneumothorax and cerebral hemorrhage, enabling rapid response to critical cases and meeting the needs of the clinical golden time window; (3) Respiratory artifact suppression: Optical flow field compensation technology effectively reduces the interference of motion artifacts on dynamic analysis and improves the reliability of feature extraction.

[0023] 3. Strengthen medical compliance and legal protection (1) Full-chain evidence storage: Automatically record the model reasoning path, data hash value and medical guideline basis, and generate an audit package that complies with international medical regulations (such as FDA / NMPA); (2) Dynamic compliance protection: The disclaimer injection mechanism triggered by sensitive words ensures system availability while avoiding legal risks in clinical decision-making.

[0024] 4. Expanding Clinical Application Value (1) Multimodal fusion capability: Simultaneously analyze the correlation between image features and text medical history, and output enhanced suggestions that conform to the doctor's diagnostic logic; (2) Standardized teaching tools: Automatically generated 3D teaching models and reasoning chain visualization provide medical students with traceable learning cases; (3) System compatibility: Supports embedded deployment (such as CT / MRI equipment) or SaaS service mode to adapt to different medical scenario needs.

[0025] Actual testing has shown that it can effectively identify fractures, traumatic injuries, and brain lesions that were previously undetectable by the software. The results using only steps a and the first half of step c are shown in the appendix. Figure 6 and 7 As shown. The effects of performing the steps of this invention are shown in the appendix. Figure 8 and 9 As shown, there are obvious improvements. Attached Figure Description

[0026] Figure 1 Layered AI architecture flowchart.

[0027] User input layer (text question + medical image) → First layer small model (question expansion + medical ontology matching) → Second layer multimodal model (medical feature extraction + image feature extraction) → Third layer large model (credibility verification + guideline application) → Output layer (diagnostic report).

[0028] Figure 2 : Dynamic lesion tracking algorithm framework.

[0029] (1) Input: Registered time-series DICOM sequence (t1, t2, t3 scans).

[0030] (2) Processing core: Convolutional network; extraction of spatiotemporal features; optical flow field compensation engine.

[0031] (3) Output: Lesion evolution curve.

[0032] Figure 3 Logic diagram of a multi-expert voting system.

[0033] (1) Parallel input of three models: image slices; medical history relationship diagram; text report.

[0034] (2) Weighted voting layer: For example, the weight allocation of radiology (40%) / clinical (35%) / pathology (25%).

[0035] (3) Output: consensus diagnosis (e.g., "92% probability of malignancy") and objection label (e.g., "possible inflammatory pseudotumor").

[0036] Figure 4 Flowchart of the emergency triage and routing algorithm.

[0037] Emergency Feature Detector → Emergency Feature Database (templates for pneumothorax / cerebral hemorrhage, etc.) → Lightweight Channel Activation Judgment (Yes / No) → Dedicated Inference Path.

[0038] Figure 5 : System deployment architecture diagram.

[0039] Cloud SaaS service module (including layered AI engine) ←→ Edge device (CT / MRI embedded deployment) ←→ User terminal (report viewing / 3D teaching interaction). Figure 6: Excerpt of hydrocephalus identification images and analysis results from an implementation case. Figure 7: Image recognition excerpt results. Figure 8: Excerpt of image diagnosis results based on the present invention. Figure 9: Excerpt of image diagnosis results based on the present invention. Detailed Implementation

[0040] Specific Implementation Case 1: Implementation Case of Brain and Spinal Cord Diagnosis (A Hospital in Eastern Zhejiang) Implementation Case Background The patient is a 58-year-old male who presented with progressive lower limb weakness and dysuria for 3 months. Spinal cord compression is clinically suspected, and spinal MRI is needed to determine the nature and location of the lesion.

[0041] The following implementation process will be adopted. 1. Input Stage (1) Terminal requests a question The doctor entered: "Large-space lesion in the T12-L1 spinal canal. Please analyze the nature of the lesion and the degree of nerve compression."

[0042] (2) Image data Upload the DICOM sequences of the patient's three spinal MRI scans (1 month apart).

[0043] 2. Layered AI Processing Flow (1) First layer (Qwen-3's 8B miniature model) (A) Standardized Questions (a) Rearrange user requests according to medical ontology knowledge and output structured query: "Please analyze the <boundary features>, <signal intensity>, <enhancement pattern>, and <spinal cord compression displacement> of intraspinal lesions in T12-L1 and assess the probability of malignancy".

[0044] (B) Smart Modification Injection Request (a) Automatic model addition prompt: "This analysis is primarily for educational purposes and may have some inaccuracies. Clinical decisions will be made in conjunction with other examinations."

[0045] (2) The second layer (a multimodal model specifically trained for the corresponding image scene).

[0046] (A) Feature extraction results (JSON format) { "location": {"z_range": [T12-L1], "coordinates": [x:45,y:120,z:33]}, "features": { "signal_T1": "Low signal", "signal_T2": "High signal", "enhancement": "ring reinforcement", "cord_compression": "3.2mm" }, "confidence_score": 82% }

[0047] (3) The third layer (DeepSeek-671B large model).

[0048] (A) Dynamic tracking analysis Comparing the data from the three scans, we found that: (a) Monthly growth rate of lesion volume: 18% (b) Increased spinal cord compression displacement: 1.8mm → 3.2mm → 4.7mm.

[0049] (B) Results of multi-expert voting: (a) CNN radiology model, the diagnosis is highly likely to be metastatic tumor (85%), accounting for 40% of the weight; (b) GNN clinical model, with a diagnostic opinion that supports metastatic tumors based on lung cancer history, accounting for 35% of the weight; (c) Transformer model, diagnosis opinion requiring differentiation of meningioma (15%), accounting for 25% of the total; (d) Final consensus: 89% probability of metastasis, grade III (severe) spinal cord compression.

[0050] 3. Compliant Output (1) The report includes: (A) 3D interactive model: showing the spatial relationship between the lesion and the spinal cord; (B) Chain of Evidence Tracing: (a) Referenced Guideline: NCCN Guidelines for the Diagnosis and Treatment of Spinal Metastases 2025; (b) Data hash: DICOM_sha256...

[0051] 5. Clinical validation Postoperative pathology confirmed lung adenocarcinoma with spinal cord metastasis, consistent with the systematic analysis. Dynamic tracking data was used to evaluate the effectiveness of radiotherapy, and the system-generated evolutionary model helped to develop precise radiotherapy plans.

[0052] Specific Implementation Case 2: Dynamic Evolution Diagnosis of Lung Nodules (Time-Series Analysis Scenario) 1. Input data: Patient's three chest CT scan images (3 months apart); Medical history: "Male, 51 years old, 30-year smoking history, cough for 2 months".

[0053] 2. System workflow: (1) Layered AI processing: (A) First layer: Change "Check if the lung nodules have changed" to → "Please analyze the <diameter change rate>, <density change>, and <margin feature evolution> of the nodules in the lower right lobe".

[0054] (B) Second layer: Output JSON features (including coordinates of 3 scans, HU value, and spur score).

[0055] { "scan_t1": {"location": [x1,y1,z1], "density_HU": 420, "spiculation_score": 0.72}, "scan_t2": {"location": [x2,y2,z2], "density_HU": 455, "spiculation_score": 0.81}, "scan_t3": {"location": [x3,y3,z3], "density_HU": 490, "spiculation_score": 0.89}} (C) Third layer (DeepSeek-671B): Combining smoking history, output "21% increase in volume in 3 months, annualized malignancy risk +22%".

[0056] (2) Dynamic lesion tracking: (A) 4D convolutional network detection of breathing artifacts (diaphragm displacement > 8 mm), coordinate error < 0.3 mm after optical flow field compensation. (B) Generate the malignancy probability time curve (t1:65% → t3:87%).

[0057] Specific Implementation Case 3: Hydrocephalus Identification 1. Input data: (A) Emergency head CT scan; (B) Text description: "Sudden onset of headache with vomiting for 1 hour".

[0058] The patient's imaging is attached. Figure 6 As shown, sensitive information has been removed. In the absence of application of this invention, the preliminary diagnosis is as follows: Figure 6 and 7 As shown. In the application of this invention, its imaging diagnosis is as attached. Figure 8 and 9 As shown.

[0059] 2. The first-layer mini-model transcribed the request intent from the following six dimensions to guide more accurate imaging analysis (medical logic explanation in parentheses): a. Structural integrity "Please observe whether the inner and outer tables of the skull are continuous? Is there any shift in the midline structures of the brain parenchyma?" (Triggering model to detect fracture / space-occupying effect).

[0060] b. Location of density anomalies "Could you label abnormal areas with CT values ​​>80 HU or <20 HU?" (Rapid screening for calcification / hemorrhagic / fatty lesions).

[0061] c. Multimodal analysis invitation "If there is a need for enhanced CT or MRI registration, please specify which sequences need to be compared?" (Analysis of induced vascular / white matter lesions).

[0062] d. List of key anatomical structures Each item needs to be evaluated individually: 1. Supratentorial region: basal ganglia, periventricular white matter 2. Infratentorial region: Cerebellar vermis, fourth ventricle 3. Skull base: pituitary fossa, cavernous sinus (To avoid AI missing diagnoses of hidden lesions such as pituitary microadenomas).

[0063] e. Emergency sign screening Prioritize screening for the following critical conditions: Subarachnoid hemorrhage (high density in the suprasellar cistern) Brain herniation (disappearance of the ambient cistern) "Venous sinus thrombosis (empty triangle sign)".

[0064] f. Quantitative analysis request "Measure the ratio of the anterior horn distance of the lateral ventricles to the internal diameter of the skull (to assess hydrocephalus)" "Calculate the volume of the lesion (e.g., using the ABC / 2 method)."

[0065] The first layer (Qwen-3.8B) outputs structured instructions accordingly.

[0066] 3. The second-layer model performs image analysis based on this: see appendix. Figure 8 and 9 As shown, it further generates formatted information. It is evident that compared to... Figure 7 and 8 Significant improvement has been achieved.

[0067] 4. The third-layer large model parses the medical history text and performs multimodal decision-making and multi-expert consultation simulations based on the formatted information after modality transformation in the second layer. Due to the involvement of medical history, this information is not uploaded here.

Claims

1. An intelligent medical image diagnostic system and method based on hierarchical cross-modal transformation and dynamic feature tracking, characterized in that... Includes the following steps: A. Receive terminal requests (the requests may also contain some text-based information) and medical images; B. Use the first-level small model to expand the intent of the request and translate it into medical standards; C. Input the transformed questions and medical images into the second-layer multimodal model, extract the features of the images, analyze and generate intermediate text-based results (including but not limited to text, JSON data, etc.) that can be read by the large model. D. Input the intermediate results into the large model of the third layer (including but not limited to a logic model with a larger number of parameters or a dedicated MOE inference model) for credibility verification or enhancement analysis.

2. The system according to any of the preceding claims, characterized in that... Include: A text preprocessing module (small model) is used to enhance user queries; The multimodal analysis module (intermediate model) is used for medical image feature extraction, realizing cross-modal extraction and converting it into machine-readable intermediate results; The results verification module (large model) is used to perform logical verification and enhanced reasoning on intermediate results.

3. The system according to any of the preceding claims, characterized in that, The output of the second-layer multimodal model is in text information format, including but not limited to JSON format, including: Lesion localization (spatial coordinates / anatomical description); Imaging features (density, margins, enhancement patterns); Initial credibility score.

4. The system according to any of the preceding claims, characterized in that, For multiple time-series images, execute: (A) Spatiotemporal feature extraction: (a) Constructing a model of lesion evolution; (b) Compensating for artifacts using optical flow; (B) Quantization output: Generates probability curves with a time dimension.

5. The system according to any of the preceding claims, characterized in that, in: The second-layer model provides a confidence score for the initial diagnosis; The third-layer model corrects and enhances the score, outputting the final probability distribution analysis.

6. The system according to any of the preceding claims, characterized in that... include: The multi-expert collaborative reasoning subsystem is configured to simulate a multi-departmental consultation process through a multi-model voting mechanism, including but not limited to: The task scheduling module is used to allocate input data to other independent inference models: Image analysis models specifically designed to analyze fine-grained details in medical images; Clinical association model, specifically designed for integrating patient medical history for cross-node reasoning; A text-based reasoning model specifically designed for parsing semantic information in pathology reports; The expert consensus module is used to aggregate the output results of the three independent reasoning models and generate a final diagnostic decision according to preset voting rules.

7. The system according to any of the preceding claims, characterized in that... include: The second-layer multimodal processing module calculates the confidence level of the multimodal input in real time, wherein the confidence level is based on at least one index definition, including but not limited to the Dice coefficient for image lesion segmentation or the F1 value for text entity recognition. A dynamic weight generation unit is used to dynamically construct a fusion weight matrix based on the confidence level; The third layer of the large language model is used to receive feature inputs weighted based on the fusion weight matrix, and to perform secondary optimization of the weights through an attention mechanism to generate the final diagnostic result.

8. The system according to any of the preceding claims, characterized in that... include: A feature mapping module is used to align image features with text descriptions to the same semantic space through contrastive learning.

9. The system according to any of the preceding claims, characterized in that: The third-layer large language model is configured to perform dynamic diagnostic tree generation.

10. The system according to any of the preceding claims, characterized in that... The final result is not a report, but an adaptive consultation path based on the patient's physiological data, including: A hybrid decision model was used to determine the consultation path, in which the decision weights were the weights α of the language model inference output, the weights β of clinical guideline rules, and the weights γ of real-time physiological indicators. Implement a context-aware follow-up question suggestion mechanism, which prioritizes follow-up questions when preset conditions are met.

Citation Information

Cited By

  • Medical image processing method and system based on machine vision and electronic equipment

    CN121236066A

  • A medical image processing method and system based on machine vision and electronic equipment

    CN121236066B

  • A liver cancer multi-modal image whole course intelligent tracking and evaluation system

    CN122436195A