Radiographic image intelligent auxiliary diagnosis system

By constructing an intelligent assisted diagnostic system for radiological images based on multimodal image fusion analysis and dynamic knowledge reasoning, the system addresses the problems of insufficient identification of lesions in small samples and lack of transparency in decision-making in existing technologies. It achieves efficient and interpretable diagnostic support, thereby improving diagnostic accuracy and work efficiency.

CN121885155AInactive Publication Date: 2026-04-17SHAANXI ZHIHUI MEDICAL INNOVATION MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI ZHIHUI MEDICAL INNOVATION MEDICAL TECH CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing intelligent assisted diagnostic systems for radiological imaging lack robustness in identifying rare lesions in small samples, struggle to adapt to differences in equipment parameters across different hospitals, lack the ability to model contextual information, and have opaque decision-making processes, leading to missed diagnoses, misdiagnoses, and low doctor trust. They also cannot be deeply integrated with clinical diagnosis and treatment processes, and their application is limited, especially in scenarios with high timeliness or limited resources.

Method used

We construct an intelligent auxiliary diagnostic system for radiological imaging, which integrates multimodal image fusion analysis, dynamic knowledge reasoning, human-machine collaborative decision-making, and automatic generation of structured reports. Through a multi-scale feature extraction engine, a cross-modal alignment network, a pathological semantic reasoning unit, a diagnostic consensus generator, and an interpretable visualization module, we achieve unified representation and efficient diagnostic support for multi-source heterogeneous image data.

Benefits of technology

It has improved the comprehensive ability to distinguish complex cases, enhanced the transparency and medical rationality of model decision-making, reduced the risk of misjudgment, reduced the workload of physicians, shortened the report issuance time, improved the robustness and stability of diagnostic results, and enhanced doctors' trust in AI judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121885155A_ABST
    Figure CN121885155A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and medical image processing, particularly relates to an intelligent auxiliary diagnosis system for radiographic images, and aims to solve the problems of missed diagnosis and misdiagnosis caused by overload of doctors, large subjective difference and low early focus recognition sensitivity in the prior art. The system integrates a multi-scale feature extraction engine, a cross-modal alignment network, a pathological semantic reasoning unit and a diagnosis consensus generator, and realizes multi-modal image fusion analysis, dynamic knowledge reasoning and man-machine collaborative decision. A thermodynamic diagram and a feature contribution degree map are output through an interpretability visualization module, and a structured report conforming to clinical specifications is generated by a report automation engine, so that the diagnosis accuracy and efficiency are improved, the report period is shortened, and credible deployment of AI auxiliary diagnosis in a complex medical scene is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and medical image processing technology, and specifically relates to an intelligent auxiliary diagnostic system for radiological imaging. Background Technology

[0002] The deep application of artificial intelligence (AI) technology in the healthcare field is accelerating the evolution of medical image analysis towards intelligence and automation. As a crucial link in clinical diagnosis, the quality of medical image interpretation directly affects disease detection rates and the efficiency of treatment decisions. In recent years, deep learning models, especially convolutional neural networks, have made breakthrough progress in image recognition tasks, providing new technical pathways for lesion detection, segmentation, and classification in radiological images. Against this backdrop, building intelligent auxiliary diagnostic systems with high precision and strong generalization capabilities has become an important research direction for improving the efficiency of medical services and alleviating the workload of physicians.

[0003] Among them, intelligent assisted diagnostic systems for radiological imaging focus on using artificial intelligence algorithms to automatically analyze medical images from modalities such as X-rays, CT scans, and MRIs to assist doctors in quickly locating lesions, assessing the severity of conditions, and providing diagnostic suggestions. These systems typically train deep neural network models based on large-scale labeled image datasets, aiming to simulate the visual interpretation capabilities of experienced radiologists. Ideally, the system should achieve multi-disease coverage, cross-device compatibility, and real-time response, thereby embedding itself into existing clinical workflows and forming an effective human-machine collaboration mechanism.

[0004] While existing technologies have achieved basic lesion identification, they still face multiple bottlenecks in practical deployment: the models lack robustness in identifying rare lesions in small samples, making it difficult to adapt to image quality fluctuations caused by differences in equipment parameters across different hospitals; most systems employ static inference architectures, lacking effective modeling capabilities for contextual information (such as patient history and before-and-after image comparisons); simultaneously, the black-box decision-making process weakens clinical interpretability, leading to limited physician trust; furthermore, insufficient integration depth with hospital PACS / RIS systems results in data lag, missing feedback loops, and an inability to support continuous iterative optimization of the model. These problems are particularly prominent in time-sensitive or resource-constrained scenarios such as emergency screening and primary healthcare, severely restricting the practical value and promotion potential of intelligent assisted diagnostic systems. Therefore, there is an urgent need to construct a radiological imaging intelligent assisted diagnostic system with strong adaptability, interpretability, and system integration capabilities to overcome current technological application bottlenecks. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent assisted diagnostic system for radiological imaging, addressing the problems of missed and misdiagnosed diagnoses caused by excessive workload for physicians, significant differences in subjective judgment, and insufficient sensitivity in early lesion identification in the current field of medical imaging diagnosis. Currently, with the continuous improvement of medical imaging equipment resolution and the exponential growth in the number of examinations, radiologists face massive image analysis tasks, and the traditional model relying on manual image interpretation is no longer sufficient to meet the needs of efficient and accurate clinical practice. Although deep learning-based assisted detection tools have been applied to the identification of specific lesions such as lung nodules and breast masses, they are generally limited to a single modality, single location, or single disease type, lacking the ability to uniformly represent multi-source heterogeneous image data. Furthermore, most systems only provide isolated detection results output, failing to deeply integrate with the clinical diagnosis and treatment process, and cannot achieve end-to-end closed-loop support from image recognition to diagnostic suggestion generation to structured report output. In addition, existing technologies have significant shortcomings in model interpretability, making it difficult to clearly present the decision-making basis to physicians, limiting their reliable deployment in high-risk medical scenarios.

[0006] The technical solution of this invention is to construct an intelligent assisted diagnostic system for radiological imaging that integrates multimodal image fusion analysis, dynamic knowledge reasoning, human-machine collaborative decision-making, and automatic generation of structured reports. This system consists of a front-end image access module, a multi-scale feature extraction engine, a cross-modal alignment network, a pathological semantic reasoning unit, a diagnostic consensus generator, an interpretability visualization module, and a report automation engine. The front-end image access module receives raw data in DICOM format from various imaging devices such as CT, MRI, X-ray, and PET-CT, and performs patient information desensitization and image sequence classification according to a preset protocol. The multi-scale feature extraction engine adopts a cascaded three-dimensional convolutional neural network architecture, performing localized refined feature capture for different organ regions while preserving global anatomical contextual information; for multi-phase scan data of the same patient, a time-dimensional modeling mechanism is introduced, capturing the dynamic evolution trend of lesions through temporal differential convolution. The cross-modal alignment network, based on the principle of shared latent space mapping, projects feature representations from different imaging modalities onto a unified semantic space, using the mutual information maximization criterion to achieve consistent matching between anatomical location and pathological features, thereby supporting the joint interpretation of PET metabolic activity and MRI soft tissue contrast.

[0007] The pathological semantic reasoning unit is based on a built-in medical ontology knowledge graph, which covers the International Classification of Diseases (ICD), RadLex (Radiological Terminology Standard), common lesion evolution paths, and their associated clinical manifestations. After the multi-scale feature extraction engine outputs a candidate set of potential abnormal regions, the pathological semantic reasoning unit initiates a bidirectional reasoning process: forward reasoning activates possible diagnostic nodes in the knowledge graph based on image features, while backward reasoning predicts the expected radiological manifestations based on a prior disease model and compares the residuals with actual observed features; the two work together to generate a preliminary set of diagnostic hypotheses and their confidence scores. The diagnostic consensus generator further integrates the independent judgments of multiple virtual expert agents. Each agent is constructed based on a differentiated training dataset and parameter initialization strategy, possessing a specific subspecialty bias; each agent outputs a diagnostic ranking list for the same case, and the consensus generator uses an improved Borda scoring method to fuse multiple ranking results, dynamically optimizing the final diagnostic priority sequence with a weight adjustment factor.

[0008] The interpretability visualization module simultaneously generates heatmap overlay masks, key slice localization sequences, and feature contribution decomposition maps to visually display the model's focus areas and decision-making basis. The heatmaps are generated using a gradient-weighted activation mapping algorithm to accurately label the core response areas of lesions. The feature contribution decomposition maps use the Shapley value approximation method to quantify the influence of each imaging feature on the final diagnostic conclusion. The report automation engine receives the primary and differential diagnosis lists output by the diagnostic consensus generator, calls a predefined template library to generate a structured diagnostic report that meets ACR guidelines requirements, including three main parts: examination description, findings summary, and impression conclusion. The impression conclusion section automatically inserts confidence intervals and recommended follow-up periods. All generated content supports online editing and confirmation by radiologists, and after digital signature, it is archived in the hospital information system.

[0009] As one embodiment of the present invention, the multi-scale feature extraction engine, when processing chest CT volumetric data, first applies a sliding window strategy to divide the lung parenchyma sub-regions. Then, it captures subtle signs such as spiculation at the edges of small nodules and pleural traction through a three-level cavitation convolutional layer with increasing expansion rates. Simultaneously, it optimizes the continuity of the segmentation boundaries using a fully connected conditional random field. Furthermore, the engine integrates a respiratory motion artifact suppression module, using a threshold for the rate of change of pixel intensity between adjacent layers to determine motion-interference slices, and employs a spatiotemporal interpolation algorithm to reconstruct information about damaged layers, ensuring the integrity of subsequent analysis inputs.

[0010] As one embodiment of the present invention, when fusing brain MRI T1-enhanced and FDG-PET data, the cross-modal alignment network constructs a dual-branch encoder structure to extract spatial texture features and metabolic hotspot distributions of the two modalities respectively; an attention gating mechanism is introduced in the bottleneck layer to guide MRI features to selectively focus on the corresponding anatomical sites in high metabolic regions; then, the consistency of the embedding vector distributions of the two modalities is constrained by the loss function of minimizing the maximum mean difference, thereby achieving semantic-level alignment without pixel-level registration.

[0011] As one embodiment of the present invention, the knowledge graph update mechanism of the pathological semantic reasoning unit is an incremental learning framework. It regularly captures the latest published clinical research literature and authoritative guideline revisions, extracts entity relationship triples through the natural language processing module, and then injects them into the knowledge graph database after expert review and interface verification. Newly added nodes automatically establish hierarchical links with existing concepts to ensure the timeliness and logical coherence of the knowledge system.

[0012] As one embodiment of the present invention, the diagnostic consensus generator sets conflict arbitration rules: when the Jaccard similarity of the preferred diagnoses of any two agents is less than a set threshold of 40%, a secondary review process is triggered, three additional backup agents are started to conduct independent evaluations, and the newly added voting results are included in the fusion calculation; if the final diagnostic score difference is still less than 5 percentage points, the case is marked as "high uncertainty" and is forcibly transferred to a senior physician for manual review.

[0013] As one embodiment of the present invention, the report automation engine has a built-in compliance verification component that automatically checks for potential legal risks such as omissions of contraindications, failure to indicate pregnancy status, and lack of warnings about the use of contrast agents before generating the report. Once a relevant image is detected that is not mentioned in the report, a red alert is immediately issued and the submission process is suspended until the user confirms the processing opinion.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This system achieves unified modeling and joint interpretation of multiple imaging modalities and organ systems by constructing a collaborative analysis architecture of a multi-scale feature extraction engine and a cross-modal alignment network. This breaks through the limitations of traditional auxiliary diagnostic tools that are limited to single tasks, significantly improving the comprehensive discrimination ability of complex cases. The pathological semantic reasoning unit integrates symbolic knowledge reasoning and connectionist deep learning, enhancing the transparency and medical rationality of model decisions while maintaining high accuracy, and solving the core obstacle of low trust in black-box models in clinical practice. The diagnostic consensus generator simulates a multi-expert consultation mechanism, effectively reducing the risk of misjudgment caused by individual model bias through a distributed proxy voting fusion strategy, and improving the robustness and stability of diagnostic results. The entire system forms a complete closed loop from raw image input to structured report output, seamlessly embedding into the existing PACS workflow, greatly reducing the paperwork burden on physicians and shortening report issuance time, with an average speedup of 63% in actual tests. The interpretability visualization module provides multi-level evidence support, which not only facilitates doctors to quickly verify the reliability of AI judgments, but also provides objective evidence for doctor-patient communication, helping to promote the large-scale and standardized application of artificial intelligence in high-end medical fields. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the overall technical architecture of the intelligent assisted diagnostic system for radiological imaging proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of the collaborative mechanism for pathological semantic reasoning and diagnostic consensus generation in this invention; Figure 3 This is a flowchart illustrating the main stages of the multi-scale feature extraction engine and cross-modal alignment network in this invention. Figure 4 This is a schematic diagram of the multi-level interaction relationship and data flow between the front-end access, AI analysis engine and report automation engine in this invention. Detailed Implementation

[0016] Please refer to Figures 1 to 4This invention relates to an intelligent assisted diagnostic system for radiological imaging, aiming to construct a closed-loop technical architecture covering multimodal medical image input, deep feature analysis, cross-modal semantic fusion, knowledge-driven reasoning, multi-agent consensus decision-making, interpretable visualization support, and automatic generation of structured reports. By deeply integrating cutting-edge artificial intelligence algorithms with clinical diagnostic logic, the system achieves intelligent, standardized, and efficient processing of complex radiological imaging data, significantly improving diagnostic accuracy, consistency, and work efficiency. The overall system operation begins with the access and preprocessing of raw image data, followed by lesion detection and feature extraction by a multi-layered AI analysis engine. Subsequently, hypothesis generation is performed in the pathological semantic reasoning unit using medical ontology knowledge. Multi-perspective judgment fusion is completed in the diagnostic consensus generator, and finally, the automated report engine outputs structured diagnostic opinions that conform to industry standards, supplemented by visual evidence for physician review and confirmation.

[0017] The front-end image access module, as the system's entry point, is responsible for the unified reception and standardized processing of multi-source heterogeneous image data. The module establishes a stable connection with the hospital's PACS system, imaging equipment, and remote transmission platform via the standard DICOM communication protocol, monitoring and capturing newly generated examination data streams in real time. Upon receiving the raw images, the module first performs patient identity information desensitization, automatically removing or encrypting metadata containing sensitive fields such as name, ID number, and contact information according to preset rules, retaining only necessary non-identifying information such as timestamps for case association, equipment model, and examination number, ensuring the entire processing complies with medical privacy protection regulations. Subsequently, based on attributes such as anatomical location, imaging sequence type, and scan phase in the DICOM tags, the module automatically classifies image slices and performs three-dimensional tissue reconstruction, forming volumetric datasets divided by organ region. For enhanced scan cases, the module further identifies arterial, venous, and delayed phase sequences and establishes temporal correspondences, providing a foundation for subsequent dynamic evolution analysis. All processed data is encapsulated in a uniform data packet format and accompanied by quality control markers, including image noise level assessment values, slice thickness consistency detection results, and preliminary motion artifact assessment status, to ensure that the input received by downstream analysis modules has reliable imaging quality.

[0018] The multi-scale feature extraction engine is the core perceptual component of the system, responsible for accurately capturing the spatial morphology, texture distribution, and dynamic changes of potential abnormal regions from high-dimensional medical images. The engine employs a cascaded 3D convolutional neural network architecture, with its backbone network based on an improved U-Net++ topology design, comprising two main functional branches: an encoding path and a decoding path. The encoding path consists of four cascaded residual convolutional blocks, each integrating a batch normalization layer and a modified linear unit activation function. It progressively downsamples the input volume data, gradually compressing the spatial resolution while expanding the receptive field to capture a wide range of anatomical contextual information. At each downsampling stage, a dense skip connection mechanism is introduced, concatenating the high-level semantic feature map of the current layer with all shallower low-level detail feature maps, effectively mitigating the gradient vanishing problem in deep networks and preserving the edge sharpness of small lesions. The decoding path achieves upsampling through transposed convolution operations, gradually restoring the spatial dimension, and incorporates multi-scale detail information from the encoding path using dense skip connections, ultimately outputting a pixel-level anomaly probability heatmap.

[0019] To address the specific needs of different anatomical regions, the multi-scale feature extraction engine incorporates a differentiated sub-network configuration strategy. Taking chest CT volume data as an example, the engine initiates a dedicated lung parenchyma analysis workflow. First, a sliding window strategy is applied to process the entire lung region into non-overlapping blocks, with each block set to a size of 128×128×64 voxels to ensure that a single sub-region can completely contain typical lung nodules and their surrounding tissue structures. Subsequently, the local sub-network employs a stacked structure of cavitary convolutional layers with three levels of increasing expansion rates, set to expansion rates of 1, 2, and 3 respectively, to capture key imaging features such as fine spiculation, lobulation, and pleural traction. The advantage of cavitary convolution is that it expands the receptive field without increasing the number of parameters, avoiding the loss of small target information caused by traditional pooling operations. At the end of the network, a fully connected conditional random field optimization module is integrated to finely adjust the segmentation boundaries using pixel affinity relationships, forcing adjacent voxels with similar gray levels to tend towards the same category label, thereby obtaining continuous and smooth lesion contour segmentation results.

[0020] To further enhance robustness against respiratory motion artifacts, the engine integrates a respiratory motion artifact suppression module. This module operates based on the principle of analyzing the rate of change of pixel intensity between adjacent layers. Specifically, the system calculates the mean absolute difference (MAD) between each axial slice and its adjacent slices above and below. If the MAD value of a slice exceeds a preset threshold of 0.15, it is determined to be significantly affected by respiratory motion. For such damaged slices, the system employs a spatiotemporal interpolation reconstruction algorithm: selecting two unaffected normal slices before and after the damaged slices, and using the voxel intensity at their corresponding locations as a basis, a weighted trilinear interpolation method is used to estimate the grayscale distribution of the missing slice, with the weighting coefficients set exponentially based on temporal distance. After reconstruction, the system re-evaluates the quality control marker of the slice, updates it to "repaired," and incorporates it into subsequent analysis processes. This mechanism effectively ensures the visual coherence of multi-planar reconstructed images of the mediastinal window and lung window, preventing false positives or false negatives due to distortion in individual slices.

[0021] When processing multiple scan sequences from the same patient, the multi-scale feature extraction engine activates a temporal dimension modeling mechanism. The system arranges the CT volume data from the baseline, short-term follow-up, and long-term follow-up periods in chronological order to construct a four-dimensional tensor input. Based on this, a temporal difference convolution module is introduced, whose core operations are defined as follows:

[0022] in, Indicates at a point in time The time-series differential response diagram at the location; and Representing the first Period and the The original images from that period; For learnable time-series filter weights, satisfying To ensure a zero-mean response; The temporal window radius is set to 2. This formula highlights regions where significant density or volume changes occur between consecutive time points, such as dynamic evolution patterns like the transformation of ground-glass nodules into solid components, enlargement of calcifications, or lymph node enlargement. The differential response map is then fed into a separate 3D convolutional branch for feature extraction and fused with the static morphological feature map at the channel level to form a joint representation vector that combines spatial detail and temporal dynamics, serving as the input basis for subsequent cross-modal alignment and pathological inference.

[0023] Cross-modal alignment networks address the heterogeneity of image data acquired from devices using different imaging principles in terms of spatial resolution, contrast characteristics, and signal representation, achieving semantically consistent fusion of multimodal information. The network is built upon a shared latent space mapping framework. Its core idea is to project the original feature representations of different modalities onto a unified high-dimensional semantic embedding space through nonlinear transformation. In this space, the different modal representations of the same anatomical structure or pathological entity should be as close as possible, while maintaining sufficient distinguishability between different entities. The network employs a dual-branch encoder-bottleneck-decoder architecture, with the two branches corresponding to two imaging modalities to be fused, such as brain MRI T1-weighted enhanced images and FDG-PET metabolic imaging.

[0024] Taking the combined interpretation of brain tumors as an example, the first branch processes T1-enhanced MRI data, and its encoder extracts spatial structural features such as tumor boundary clarity, enhancement pattern, and edema extent. The second branch processes PET data and extracts functional metabolic parameters such as maximum standard uptake value (SUVmax), metabolic volume, and heterogeneity index. Both encoder branches employ a lightweight 3D ResNet18 variant, outputting their respective low-dimensional embedding vectors. and At the bottleneck layer, the network introduces an attention gating mechanism. This is implemented by using the global average pooling features output from the PET branch as the query vector, performing a dot product operation with the feature maps of each spatial location in the MRI branch to generate a spatial attention weight map. This weight map is then multiplied with the original MRI feature map, guiding the attention of high-metabolic regions to corresponding anatomical locations. This mechanism allows the model to prioritize MRI structures corresponding to metabolically active areas during analysis, improving the relevance of joint interpretations.

[0025] To measure and optimize the consistency of the two modality embedding vector distributions, the network employs a loss function that minimizes the maximum mean difference. This loss function calculates the mean embedding distance between the two sample sets in the reproducing kernel Hilbert space, mathematically expressed as:

[0026] in, The maximum mean difference loss function, and For a single sample, the embedding vector. Let be the square norm in the RKHS space. and These represent the embedding vector sets for mode A and mode B, respectively. For kernel mapping functions; For the regenerating nucleus Hilbert space; , The number of samples is represented by this term. By minimizing this loss term during training, the network forces the embedding distributions of MRI and PET to converge, thereby achieving semantic-level alignment without requiring precise pixel-level registration. Finally, the fused joint embedding vector is fed into the downstream pathological semantic reasoning unit as input for comprehensive judgment.

[0027] The pathological semantic reasoning unit is a key module for the system to leap from "image recognition" to "medical understanding." Its function is to transform low-level image features into high-level clinical diagnostic hypotheses. The unit uses a built-in medical ontology knowledge graph as its reasoning foundation. This knowledge graph is stored in RDF triples and covers the International Classification of Diseases 11th Revision (ICD-11) coding system, the RadLex dictionary of radiological terminology, common lesion evolution pathway maps and their associated clinical manifestations, laboratory indicators, and treatment response data. Knowledge graph nodes include five major entity categories: "disease," "symptom," "imaging sign," "anatomical location," and "pathological type." Edges represent semantic relationships such as "belongs to," "leads to," "manifests as," "located in," and "accompanied by," forming a large semantic network covering over 120,000 nodes and 350,000 relationships.

[0028] After the multi-scale feature extraction engine outputs a candidate set of potential abnormal regions, the pathological semantic reasoning unit initiates a bidirectional reasoning process. The forward reasoning path activates potential diagnostic nodes in the knowledge graph based on observed image features. Specifically, the detected lesion feature vectors (such as size, density, edge, growth rate, etc.) are first encoded into a set of standardized predicates, such as "diameter greater than 8 mm," "edge with spiky appearance," and "SUVmax greater than 2.5." Then, all "disease" nodes in the knowledge graph that have a "manifestation as" relationship with these predicates are retrieved, resulting in a preliminary candidate diagnosis list. The reverse reasoning path predicts the expected imaging manifestations based on a prior disease model and compares the residuals with the actual observed features. The system traverses each disease node in the candidate diagnosis list, expanding its typical imaging manifestation sub-graphs along the "cause" relationship to generate a set of expected features. Then, the matching score between the expected features and the actual observed features is calculated, and the Jaccard similarity coefficient is used to measure the intersection ratio. Diagnoses with scores below a set threshold of 70% are downweighted or eliminated.

[0029] Through the collaborative action of bidirectional inference, the system generates a preliminary set of diagnostic hypotheses and their confidence scores. The scoring mechanism comprehensively considers three dimensions: first, the activation strength of forward inference, i.e., the sum of connection weights between observed features and diagnostic nodes; second, the residual matching degree of backward inference, reflecting the consistency between the diagnostic model and real-world observations; and third, the epidemiological prior probability, which is weighted by calling conditional incidence rates from the local database based on demographic factors such as age, gender, and smoking history. Finally, each diagnostic hypothesis receives a normalized confidence value, ranging from 0 to 1, which is submitted to the diagnostic consensus generator as the ranking criterion.

[0030] To ensure the timeliness and authority of the knowledge graph, the pathological semantic reasoning unit is equipped with an incremental learning and updating mechanism. The system regularly starts a natural language processing crawling program daily, accessing authoritative medical literature databases such as PubMed, Cochrane Library, and UpToDate to search for the latest research papers, guideline revision announcements, and expert consensus documents related to radiological diagnosis. After deduplication and cleaning, the crawled content is fed into a medical text mining model based on the BERT architecture to extract entity relation triples, such as "novel ADC value threshold <1.1×10". -3 mm 2 The message " / s indicates an increased likelihood of prostate cancer" appears. Newly generated triples are pushed to an expert review interface, where registered radiologists review their scientific validity and applicability online. Once confirmed, they are injected into the main knowledge graph database. New nodes automatically trigger a hierarchical linking process, using semantic similarity calculations to find the closest parent concept (e.g., parent disease) and differential diagnoses at the same level, maintaining the logical integrity of the knowledge system. The entire update process is fully synchronized monthly to ensure the system always bases its reasoning on the latest evidence-based medicine.

[0031] The diagnostic consensus generator integrates the independent judgments of multiple virtual expert agents, simulating the collective wisdom decision-making process in a multidisciplinary consultation scenario to reduce the risk of misjudgment caused by the bias of a single model. The system deploys seven virtual expert agents, each an independently trained deep learning model instance with a specific subspecialty focus, specializing in thoracic, abdominal, neurological, musculoskeletal, breast, cardiovascular, and pediatric imaging domains, respectively. Each agent receives the same input data package, including raw images, multi-scale feature maps, cross-modal fusion vectors, and basic patient clinical information. However, due to differences in the source of their training datasets, variations in sample distribution, and random initialization of network parameters, they exhibit differentiated diagnostic preferences when facing cases with ambiguous boundaries.

[0032] Each virtual expert agent runs an independent diagnostic reasoning process, outputting an ordered list of diagnostic rankings, including primary diagnoses, secondary diagnoses, and excluded diagnoses, each with a confidence score. The diagnostic consensus generator uses a modified Borda scoring method to merge multiple ranking results. Traditional Borda scoring assigns a fixed score to each ranking position (e.g., first place gets 7 points, second place gets 6 points, and so on). This system introduces a weighting factor to give higher-confidence diagnoses more weight. The specific scoring rule is: if a diagnosis ranks in a certain agent's list... And its confidence level is Then its contribution score is The scores of all agents for this diagnosis are summed to obtain the total Borda score, which is then arranged in descending order of total score to form the final diagnosis priority sequence.

[0033] To address highly uncertain cases, the diagnostic consensus generator employs conflict arbitration rules. When the Jaccard similarity between the preferred diagnoses of any two primary agents falls below a set threshold of 40%, a significant disagreement is identified, triggering a secondary review process. The system immediately activates three additional backup agents for independent evaluation. These agents are constructed using stronger data augmentation strategies and adversarial training methods, resulting in greater robustness. The newly added voting results are incorporated into the fusion calculation, and weighted Borda scoring is re-executed. If the final score difference between the top two diagnoses is still less than 5 percentage points, the system automatically marks the case as "highly uncertain," generates a red warning label, and forcibly transfers it to the manual review queue, suspending automatic report generation until a senior radiologist intervenes. This mechanism effectively prevents the risk of misleading conclusions on key diagnoses, ensuring patient safety.

[0034] The interpretability visualization module simultaneously generates various types of auxiliary views to intuitively demonstrate the model's decision-making basis and focus, enhancing doctors' trust in AI judgments. The module output includes three core visualization products: heatmap overlay masks, key slice localization sequences, and feature contribution decomposition maps. The heatmap is generated using a gradient-weighted class activation mapping algorithm. Its principle is to backpropagate the gradient signal of the final diagnostic category relative to the feature map of the last convolutional layer, calculate the importance weight of each channel, then perform a weighted sum and upsample to the original image resolution, forming a continuous color heatmap distribution. High-temperature areas (red) represent the core lesion areas of high model focus, while low-temperature areas (blue) represent background-irrelevant areas. This heatmap is presented transparently overlaid on the original DICOM image, allowing doctors to quickly verify the accuracy of AI localization.

[0035] The key slice localization sequence automatically extracts several axial, coronal, and sagittal slices containing the most representative lesion manifestations, arranges them in spatial continuity, and labels measurement values ​​(such as major axis, minor axis, CT value) and arrows indicating typical signs (such as spiculation, pleural indentation), forming a complete chain of evidence. This sequence can be directly inserted into the final diagnostic report as graphic evidence.

[0036] The feature contribution decomposition map is implemented using the Shapley value approximation method to quantify the influence of each imaging feature on the final diagnostic conclusion. The system treats diagnostic decisions as a payoff function in a cooperative game, with each imaging feature (such as size, density, edge, and growth rate) considered a participant. Monte Carlo sampling is used to estimate the expected value of the feature's marginal contribution across various feature combinations, which is its Shapley value. Finally, the top ten positive and negative contributing features and their values ​​are displayed in a bar chart, enabling physicians to clearly understand which factors drive the malignancy diagnosis and which support the possibility of a benign diagnosis, providing traceable logical support for clinical decision-making.

[0037] The report automation engine is responsible for transforming the structured conclusions output by the diagnostic consensus generator into professional documents that conform to clinical standards. The engine has a built-in template library that meets the requirements of the American College of Radiology (ACR) guidelines, covering common examination types in all major systems of the body. Each template contains three logical paragraphs: examination description, findings summary, and initial conclusion. The engine automatically matches the optimal template based on the anatomical location of the analysis and fills in the specific content. The examination description section automatically generates technical details such as scanning parameters, reconstruction methods, and contrast agent usage; the findings summary section lists all detected abnormalities by anatomical region, cites key slice location sequence numbers, and describes their location, size, shape, density / signal characteristics, and dynamic trends; the initial conclusion section lists the primary diagnosis and differential diagnosis, with a confidence interval (e.g., 87%-93% for a 95% confidence interval) after each diagnosis, and automatically recommends follow-up periods based on the nature of the lesion (e.g., "Recommend a follow-up chest CT scan in 6 months" or "Recommend further biopsy to clarify pathology").

[0038] To mitigate medical legal risks, the automated reporting engine integrates a compliance verification component. This component automatically checks for potential omissions before submission, such as missing contraindications, unreported pregnancy status, or missing risk warnings regarding the use of iodine-containing contrast agents in patients with renal insufficiency. The verification logic is based on a rule engine; for example, if imaging reveals a mass in the uterus and the patient's age is between 12 and 55 years, the possibility of pregnancy must be mentioned in the report; if the eGFR value is below 45 mL / min / 1.73 m... 2Furthermore, if intravenous contrast agents are used, a warning statement regarding contrast-induced nephropathy must be added. If a relevant imaging finding is detected but not mentioned in the report, the system will immediately issue a red alert pop-up, suspend the submission process, and force the user to confirm their opinion before continuing. All generated content can be edited online by radiologists, and all modifications are fully recorded. The final version, after being digitally signed by the physician, is uploaded back to the hospital information system for archiving, forming an immutable electronic medical record.

[0039] The intelligent assisted diagnostic system for radiological imaging constructed in this embodiment achieves fully automated support from raw image input to structured report output through the organic synergy of the aforementioned functional modules. The system not only improves the sensitivity and specificity of lesion detection, but more importantly, establishes an interpretable reasoning chain based on medical knowledge. This ensures that AI judgments are not isolated outputs, but rather auxiliary opinions supported by clinical logic. The multi-agent consensus mechanism effectively balances the biases of individual models, enhancing diagnostic stability; while compliance verification and mandatory review rules construct a robust security defense, ensuring reliable operation of the system in real medical environments. Real-world test data shows that the system increases the early nodule detection rate by 28% and reduces the false alarm rate by 41% in lung cancer screening; achieves an accuracy rate of 89.7% in glioma grading, approaching the level of high-level experts; and reduces the average report writing time from 25 minutes to 9 minutes, a speedup of 64%, significantly alleviating the workload of radiologists and providing a practical technical path for the development of smart healthcare.

[0040] Current radiological imaging-assisted tools generally remain at the level of "single-point breakthroughs," focusing on optimizing the detection performance of specific diseases and lacking systematic process integration capabilities. Most products only provide detection frames and probability values, failing to generate directly usable diagnostic reports, requiring physicians to spend considerable time on text editing and formatting. Furthermore, these tools often overlook the combined value of multimodal data, failing to fully utilize the complementary analysis of PET's functional metabolic information and MRI's soft tissue resolution advantages. More critically, existing models generally suffer from interpretability deficiencies, making it difficult for physicians to judge the rationality of their judgments, leading to reservations about their conclusions in practice and limiting clinical adoption rates.

[0041] The core difference of this invention lies in constructing an end-to-end, closed-loop intelligent diagnostic workflow that seamlessly connects the four major stages of image analysis, knowledge reasoning, consensus decision-making, and report generation. The system not only focuses on "whether it can detect," but also emphasizes "why this judgment is made" and "how it is delivered and used." By introducing a pathological semantic reasoning unit, the system can perform bidirectional verification based on an authoritative medical knowledge graph, generating well-founded diagnostic hypotheses; through a diagnostic consensus generator simulating multi-expert consultations, it effectively suppresses individual bias and improves the robustness of results; through an interpretability visualization module, it provides multi-level evidence support, greatly enhancing doctors' trust in the AI ​​output; and finally, through an automated report engine, it completes the automatic generation and compliance verification of professional documents, truly achieving reduced workload and increased efficiency. This entire technical architecture design concept marks a fundamental shift in medical artificial intelligence from "tool-based assistance" to "process-based restructuring."

[0042] The knowledge graph of the pathological semantic reasoning unit possesses dynamic evolution capabilities, with a monthly full synchronization frequency. Each synchronization includes approximately 300 to 500 new knowledge entries reviewed by experts. The knowledge injection process employs transactional database operations to ensure atomicity and consistency, preventing graph corruption due to power outages. The establishment of hierarchical links for newly added nodes relies on a semantic similarity calculation module. This module generates contextual embedding vectors for node names based on the pre-trained biomedical language model BioBERT, then calculates cosine similarity to find nearest neighbor nodes as parent class candidates. Finally, the rule engine, combined with ontology hierarchical constraints, performs legality verification before attachment. The entire process requires no manual intervention, enabling the knowledge system to evolve autonomously.

[0043] The virtual expert agent model in the diagnostic consensus generator is trained using a federated learning framework. Each participating medical institution independently trains its model on its local server using its own data, uploading only the model parameters, not the raw data, to the central server for aggregation and updates. This approach ensures both the broad representativeness of the model and adheres to the privacy principle of keeping data within its domain. Each agent model undergoes a six-month iterative update cycle to ensure the continuous evolution of its knowledge base.

[0044] The heatmaps generated by the interpretability visualization module maintain strict consistency with the resolution of the original images, with scaling errors controlled within ±0.5%, ensuring spatial positioning accuracy meets clinical needs. The heatmap color mapping employs the Viridis chromatographic scheme, whose monotonically increasing brightness avoids visual misleading and is suitable for black-and-white printing. The Shapley values ​​of the feature contribution decomposition map are calculated using the KernelSHAP approximation algorithm, with a sampling count of 2048, achieving a balance between computational efficiency and estimation accuracy, with a relative error of less than 5%.

[0045] The report automation engine's template library supports a dynamic expansion mechanism, allowing hospitals to upload customized report templates according to their own management standards. These templates take effect after review by the system administrator. Template version numbers are automatically incremented, and historical versions are permanently retained for easy auditing and traceability. The engine also supports a voice command input interface, allowing physicians to supplement personalized descriptions using natural language. The system automatically converts these descriptions into structured text and inserts them into specified paragraphs, enhancing the flexibility of human-computer interaction.

[0046] This embodiment addresses the fundamental shortcomings of existing technologies—fragmentation, isolation, and black-box architecture—by constructing a highly engineered, full-chain-coverage intelligent diagnostic system. The system's modules have tight data dependencies and control flow relationships, forming a closed loop from front-end input to final output. The output of each stage becomes the input of the next, and the system possesses robust error handling and rollback mechanisms. For example, when cross-modal alignment fails, the system automatically degrades to single-modal analysis mode and notes "multimodal fusion incomplete" in the report; when diagnostic consensus cannot be reached, a manual review process is automatically triggered and the event log is recorded. This failure-oriented design philosophy ensures high availability and security of the system in complex real-world environments, providing a solid technical paradigm for the large-scale deployment of artificial intelligence in high-end medical fields.

Claims

1. A system for intelligent assistant diagnosis of radiological images, characterized in that it comprises: include: The front-end image access module is used to receive raw medical image data generated by various imaging devices and to perform patient information desensitization and image sequence classification. A multi-scale feature extraction engine is used to perform local fine-grained feature capture and global anatomical context information preservation on the classified raw image data, and to introduce time dimension modeling for multi-phase scan data of the same patient to capture the dynamic evolution trend of lesions. Cross-modal alignment networks are used to project feature representations from different imaging modalities onto a unified semantic space to achieve consistent matching between anatomical locations and pathological features; The pathological semantic reasoning unit is used to initiate a bidirectional reasoning process to generate a preliminary set of diagnostic hypotheses and their confidence scores after receiving the potential abnormal region candidate set output by the multi-scale feature extraction engine, based on the built-in medical ontology knowledge graph. A diagnostic consensus generator is used to integrate independent diagnostic ranking lists from multiple virtual expert agents with specific subspecialty preferences, and to generate a final diagnostic priority sequence using a weighted fusion strategy. An interpretable visualization module is used to simultaneously generate heatmap overlay masks, key slice location sequences, and feature contribution decomposition maps; The report automation engine receives the final diagnostic priority sequence, calls a predefined template library to generate a structured diagnostic report, and performs compliance checks before submission.

2. The system for intelligent diagnostic support in radiological imaging according to claim 1, characterized in that, The multi-scale feature extraction engine includes: The backbone of the cascaded three-dimensional convolutional neural network includes an encoding path and a decoding path. The encoding path captures high-level semantic information through stepwise downsampling, while the decoding path recovers spatial details through upsampling and introduces multi-scale features through dense skip connections. Differentiated subnetwork configuration units are used to activate corresponding dedicated analysis subnetworks based on the anatomical regions of the image to be processed; The respiratory motion artifact suppression unit is used to identify image slices affected by motion interference and to reconstruct information of the damaged layers using a spatiotemporal interpolation algorithm.

3. The system for intelligent diagnostic support in radiology of claim 2, wherein, The differentiated subnetwork configuration unit is configured as follows when processing chest CT volume data: A sliding window strategy was applied to divide lung parenchyma into sub-regions. The stacked structure of void convolutional layers with increasing expansion rate captures the edge spiculation, lobulation, and pleural traction signs of micronodules. An integrated fully connected conditional random field optimization module is used to obtain continuous and smooth lesion contour segmentation results.

4. The system for intelligent diagnostic support in radiological imaging according to claim 1, characterized in that, The cross-modal alignment network includes: A dual-branch encoder structure is used to extract spatial features and functional metabolic features of two different imaging modalities respectively; Attention gating mechanisms, set at the bottleneck layer, are used to selectively focus on the spatial features of another modality by utilizing the global features of one modality. Distributional consistency constraint unit is used to achieve semantic-level alignment by minimizing the mean embedding distance between two modal embedding vectors in the regenerating kernel Hilbert space.

5. The system for intelligent diagnostic support in radiological imaging according to claim 1, characterized in that, The bidirectional reasoning process of the pathological semantic reasoning unit includes: The forward reasoning subunit is used to retrieve and activate possible diagnostic nodes in the medical ontology knowledge graph based on the observed image features; The backward reasoning subunit is used to predict the radiological manifestations that should appear based on the prior disease model and compare the residuals with the actual observed features to verify the diagnostic hypothesis. The confidence score subunit is used to integrate the activation strength of forward inference, the residual matching degree of backward inference, and the epidemiological prior probability to generate a normalized confidence score for each diagnostic hypothesis.

6. The intelligent assisted diagnostic system for radiological imaging according to claim 5, characterized in that, The pathological semantic reasoning unit further includes: The incremental knowledge update module is used to periodically retrieve new knowledge from authoritative medical literature databases, extract entity relation triples through natural language processing, and inject them into the medical ontology knowledge graph after verification through an expert review interface. The hierarchical link building module is used to automatically establish hierarchical links between new knowledge nodes and existing concepts in order to maintain the logical coherence of the knowledge system.

7. The intelligent assisted diagnostic system for radiological imaging according to claim 1, characterized in that, The diagnostic consensus generator includes: The virtual expert agent pool contains multiple independent deep learning models with specific subspecialty biases, built on differentiated training datasets. The weighted Borda scoring fusion unit is used to calculate the weighted total score for each diagnosis based on the diagnosis ranking list and confidence score output by each virtual expert agent to generate the final diagnosis priority sequence. The conflict arbitration unit is used to trigger a secondary review process and start a backup agent for evaluation when the diagnostic disagreement between agents exceeds a preset threshold.

8. The intelligent assisted diagnostic system for radiological imaging according to claim 7, characterized in that, The conflict arbitration unit is also configured to: If, after a second review, the score difference between the top two diagnoses in the final diagnostic priority sequence is still less than the preset percentage point, the case is marked as highly uncertain and is forcibly transferred to manual review.

9. The intelligent assisted diagnostic system for radiological imaging according to claim 1, characterized in that, The interpretable visualization module includes: The heatmap generation subunit is used to generate a heatmap of the core response area of ​​the lesion through a gradient-weighted class activation mapping algorithm and present it on the original image in a transparent overlay manner. The key slice extraction subunit is used to automatically extract and arrange the most representative lesion slices, and annotate the measurement values ​​and typical signs; The feature contribution quantification subunit is used to quantify the influence of each imaging feature on the final diagnostic conclusion using the Shapley value approximation method, and is displayed in the form of a bar chart.

10. The intelligent assisted diagnostic system for radiological imaging according to claim 1, characterized in that, The report automation engine includes: The template matching and filling unit is used to automatically match the optimal report template based on the anatomical site being analyzed, and fill in the examination description, findings summary, and impression conclusions. The compliance rule verification unit is used to check whether there are any omissions of contraindications, failure to indicate pregnancy status, or missing risk warnings for contrast agent use before the report is submitted, based on preset rules. When a risk is detected, an early warning is issued and the submission process is suspended. The digital signature and return unit is used to digitally sign the final report and return it to the hospital information system for archiving after the radiologist has edited and confirmed it online.