Medical imaging diagnostic methods and devices, computing equipment
Patent Information
- Application Number
- CN202610604329.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-06
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-05-06
AI Technical Summary
特定扫描图像的获取往往受到实际应用场景的限制,例如注射造影剂的增强CT扫描,不仅需要进行额外检查,增加检查成本与时间,还可能为患者带来过敏及肾功能损伤的风险;而超声心动图的获取则需要特定的扫描协议,对扫描设备有额外要求,在医疗不发达地区受到很大限制;同时,现有的诊断方法多采用单一尺度的影像分析,难以同时兼顾整体解剖结构与局部病变细节,导致特征信息利用不充分,且无法对大量已有的常规非增强CT影像进行机会性筛查,造成潜在诊断价值的浪费,从而导致疾病诊断的效率降低,准确性较差
[0012]通过获取未增强扫描图像并从中确定全局图像与多个局部图像,构建了多尺度输入基础;利用编码层分别对全局与局部图像进行特征编码,获得涵盖整体解剖结构概貌与关键局部细节的全局特征向量与多个局部特征向量;进而通过注意力层对多个局部特征向量进行注意力计算,能够自适应地赋予与诊断相关的关键区域更高权重,从而抑制无关区域的干扰,获得更具代表性的局部聚合特征向量;通过特征融合层将全局特征与局部聚合特征进行融合,生成综合了宏观与微观信息的融合特征向量;通过解码层基于该融合特征输出准确的医疗诊断结果;该方法在不依赖造影剂和其他额外扫描检查的前提下,充分挖掘常规扫描图像的潜在诊断价值,避免了造影剂或其他额外检查带来的检查风险与成本,同时通过多尺度特征提取与注意力机制的有效结合,提升了诊断的准确性与鲁棒性,为医疗影像的全面诊断分析提供了高效且通用的技术方案。
Smart Images

Figure CN122134724B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of artificial intelligence technology, and in particular to a medical image diagnosis method, a medical image diagnosis device, and a computing device. Background Technology
[0002] With the widespread application of artificial intelligence technology in the field of medical image analysis, the use of diagnostic models to assist in the diagnosis of diseases has become a research hotspot.
[0003] Currently, methods for assisted diagnosis based on medical imaging typically rely on high-quality, specialized scan images as model input. These images, such as those obtained through enhanced CT scans or echocardiography, often require specific examination methods. For instance, enhanced CT scans involve injecting contrast agents and using specialized scanning equipment, while echocardiography requires additional examination using specific scanning probes. By acquiring these specialized scan images, high-contrast, high-resolution image data of the area to be diagnosed can be provided to the diagnostic model, thereby enabling assisted diagnosis of diseases.
[0004] However, the aforementioned technical solutions face certain technical challenges in practical applications. The acquisition of specific scan images is often limited by the specific application scenario. For example, contrast-enhanced CT scans requiring contrast agents not only necessitate additional examinations, increasing costs and time, but also potentially posing risks to patients such as allergies and kidney damage. Echocardiography, on the other hand, requires specific scanning protocols and places additional demands on the scanning equipment, severely limiting its application in underdeveloped regions. Furthermore, existing diagnostic methods often employ single-scale image analysis, making it difficult to simultaneously consider both overall anatomical structures and local lesion details. This leads to insufficient utilization of feature information and an inability to opportunistically screen a large number of existing conventional non-contrast CT images, resulting in a waste of potential diagnostic value and consequently reduced efficiency and accuracy in disease diagnosis. Therefore, there is an urgent need for a method that can achieve efficient and accurate disease diagnosis and classification based on conventional non-contrast CT images. Summary of the Invention
[0005] In view of the above, embodiments of this specification provide a medical image diagnostic method. One or more embodiments of this specification also relate to a medical image diagnostic device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0006] According to a first aspect of the embodiments of this specification, a medical image diagnosis method is provided, comprising: Obtain an unenhanced scan image of the site to be diagnosed, wherein the unenhanced scan image is obtained by scanning the site to be diagnosed without injecting contrast agent; Based on the unenhanced scan image, a global image and multiple local images of the area to be diagnosed are determined; The global image and the multiple local images are input into the encoding layer of the medical image diagnosis model for feature encoding, thereby obtaining a global feature vector and multiple local feature vectors. The medical image diagnosis model also includes an attention layer, a feature fusion layer, and a decoding layer. The multiple local feature vectors are input into the attention layer for attention calculation to obtain local aggregated feature vectors; The global feature vector and the local aggregated feature vector are input into the feature fusion layer for feature fusion calculation to obtain the fused feature vector; The fused feature vector is input into the decoding layer for decoding to obtain the medical diagnostic result of the site to be diagnosed.
[0007] According to a second aspect of the embodiments of this specification, a medical imaging diagnostic device is provided, comprising: The acquisition module is configured to acquire an unenhanced scan image of the site to be diagnosed, wherein the unenhanced scan image is obtained by scanning the site to be diagnosed without injecting contrast agent; The determination module is configured to determine a global image and multiple local images of the site to be diagnosed based on the unenhanced scan image; The encoding module is configured to input the global image and the multiple local images into the encoding layer of the medical image diagnosis model for feature encoding, thereby obtaining a global feature vector and multiple local feature vectors. The medical image diagnosis model further includes an attention layer, a feature fusion layer, and a decoding layer. The attention module is configured to input the multiple local feature vectors into the attention layer for attention calculation to obtain a local aggregated feature vector; The fusion module is configured to input the global feature vector and the local aggregated feature vector into the feature fusion layer to perform feature fusion calculation and obtain a fused feature vector. The decoding module is configured to input the fused feature vector into the decoding layer for decoding to obtain the medical diagnostic result of the site to be diagnosed.
[0008] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer program / instructions, which, when executed by the processor, implement the steps of the above-described medical image diagnosis method.
[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described medical image diagnosis method.
[0010] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described medical image diagnosis method.
[0011] One embodiment of this specification implements a medical imaging diagnostic method, comprising: acquiring an unenhanced scan image of a site to be diagnosed, wherein the unenhanced scan image is obtained by scanning the site to be diagnosed without injecting contrast agent; determining a global image and multiple local images of the site to be diagnosed based on the unenhanced scan image; inputting the global image and the multiple local images into the encoding layer of a medical imaging diagnostic model for feature encoding, thereby obtaining a global feature vector and multiple local feature vectors, wherein the medical imaging diagnostic model further includes an attention layer, a feature fusion layer, and a decoding layer; inputting the multiple local feature vectors into the attention layer for attention calculation, thereby obtaining a local aggregated feature vector; inputting the global feature vector and the local aggregated feature vector into the feature fusion layer for feature fusion calculation, thereby obtaining a fused feature vector; and inputting the fused feature vector into the decoding layer for decoding, thereby obtaining a medical diagnostic result for the site to be diagnosed.
[0012] By acquiring unenhanced scan images and identifying global and multiple local images, a multi-scale input foundation is constructed. An encoding layer encodes the features of both global and local images, obtaining global feature vectors and multiple local feature vectors that encompass the overall anatomical structure and key local details. An attention layer then performs attention calculations on these local feature vectors, adaptively assigning higher weights to diagnostically relevant key regions, thereby suppressing interference from irrelevant regions and obtaining more representative local aggregated feature vectors. A feature fusion layer fuses the global and local aggregated features, generating a fused feature vector that integrates macroscopic and microscopic information. A decoding layer outputs accurate medical diagnostic results based on this fused feature. This method fully leverages the potential diagnostic value of conventional scan images without relying on contrast agents or other additional scanning examinations, avoiding the risks and costs associated with contrast agents or other additional examinations. Furthermore, the effective combination of multi-scale feature extraction and attention mechanisms improves the accuracy and robustness of the diagnosis, providing an efficient and universal technical solution for comprehensive diagnostic analysis of medical images. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating a medical imaging diagnostic method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of attention layer calculation provided in one embodiment of this specification; Figure 3 This is a schematic diagram of a two-stage model training provided in one embodiment of this specification; Figure 4 This is a graph showing the prediction results of a medical image diagnosis model provided in one embodiment of this specification; Figure 5 This is a flowchart illustrating the processing procedure of a medical image diagnosis method provided in one embodiment of this specification. Figure 6 This is a schematic diagram of the structure of a medical imaging diagnostic device provided in one embodiment of this specification; Figure 7 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in one or more embodiments of this specification are obtained through open-source datasets or public datasets that comply with their license agreements, or are obtained with full authorization from the relevant parties. Moreover, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0018] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically including hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0019] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0020] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0021] Computed tomography (CT) is a technique that uses X-rays to scan an object across multiple layers and then reconstructs the internal structure of the human body using a computer. It provides high-resolution anatomical and morphological information. CT scans include plain and contrast-enhanced scans. Plain scans can display differences in tissue density without the need for contrast agents and are widely used in medical imaging diagnosis, such as for lung cancer screening and abdominal examinations. The images are composed of voxels and tissue density is quantified in Henle units.
[0022] Computed Tomography Angiography (CTA) is a technique that uses intravenous injection of iodine contrast agent followed by CT scans at appropriate times to create high-density images of blood vessels, clearly revealing their morphology and lesions. CTA can be used to assess vascular structures such as the aorta, coronary arteries, and cerebral arteries, and to diagnose diseases such as vascular stenosis, aneurysms, and aortic dissections. However, it requires the injection of contrast agents and carries risks of allergies and kidney damage.
[0023] Magnetic Resonance Imaging (MRI) uses strong magnetic fields and radio frequency pulses to excite hydrogen protons in human tissues to generate signals, and then reconstructs images with high soft tissue contrast. MRI can be used for detailed examinations of the nervous system, musculoskeletal system, cardiovascular system, and other areas. Its multi-parameter imaging capabilities provide rich diagnostic information without the need for ionizing radiation, but the scan time is relatively long and it is sensitive to motion.
[0024] Bicuspid aortic valve (BAV) is a common congenital heart valve malformation, where the aortic valve consists of only two leaflets instead of the normal three. Patients with BAV experience increased mechanical stress on the valve, making them more susceptible to serious cardiovascular diseases such as aortic stenosis, regurgitation, aortic dissection, and aortic aneurysm. The incidence rate in the general population is approximately 1-2%. Early identification of BAV is crucial for preventing complications and developing intervention strategies.
[0025] The tricuspid aortic valve (TAV) is the normal aortic valve morphology, consisting of three symmetrical semilunar leaflets, named the left coronary leaflet, right coronary leaflet, and non-coronary leaflet. The TAV opens fully during systole and closes tightly during diastole, ensuring unidirectional blood flow and maintaining normal cardiac function. As the normal anatomical type of aortic valve, the TAV is often used as a reference standard for differentiating bicuspid aortic valves in medical imaging diagnosis.
[0026] Picture Archiving and Communication Systems (PACS) is a comprehensive information system that integrates medical image acquisition, storage, management, transmission, and display. PACS enables film-free management, supports the sharing of image data among departments within a hospital and remote consultations, and provides doctors with convenient image access and diagnostic tools. It is an important component of hospital information technology infrastructure.
[0027] MedicalNet is a large-scale pre-trained model library specifically designed for medical image analysis, providing weights for deep learning models pre-trained on various 3D medical image datasets. Based on architectures such as 3D ResNet, MedicalNet can significantly improve the performance of small-sample medical image tasks through transfer learning, reducing the dependence on massive amounts of labeled data. It is widely used in tasks such as 3D medical image segmentation, classification, and detection.
[0028] 3D DenseNet is a 3D convolutional neural network architecture that uses a dense connection mechanism to concatenate each layer with all preceding layers along the channel dimension, enabling feature reuse and gradient propagation. 3D DenseNet efficiently utilizes parameters and mitigates the vanishing gradient problem, making it suitable for processing 3D medical images. It effectively extracts spatial contextual features and performs exceptionally well in tasks such as organ segmentation and disease classification.
[0029] The 3D Vision Transformer is a deep learning model that extends the vision transformer architecture to three-dimensional space. It captures global feature dependencies by dividing a 3D image into 3D blocks and projecting them into a sequence, and by using a multi-head self-attention mechanism.
[0030] 3D EfficientNet is a 3D convolutional neural network based on neural architecture search. By balancing the network's depth, width, and resolution, it significantly reduces computational overhead while maintaining high performance.
[0031] The Hounsfield Unit (HHU) is a standardized unit used in CT images to quantify the attenuation of X-rays by tissues. The HHU value is based on the attenuation of water (0 HU), air at -1000 HU, and bone typically exceeding 400 HU. Different tissues have specific HHU value ranges, and doctors can distinguish tissue types and identify lesions by observing HHU values.
[0032] Window level and window width are parameters used in medical image display to adjust the grayscale mapping range. The window level determines the central grayscale value displayed, while the window width determines the width of the displayed grayscale range. By adjusting the window level and width, details of specific tissues can be highlighted.
[0033] 3D ResNet is a three-dimensional residual neural network that addresses the vanishing gradient problem in deep networks by introducing residual connections (skip connections), enabling efficient training of deep structures. Typically composed of stacked residual blocks, 3D ResNet can extract hierarchical features from 3D images and is widely used in medical image segmentation, classification, and detection tasks. Its pre-trained weights are frequently used for transfer learning.
[0034] A region of interest (AA-ROI) is a localized area within a medical image that contains specific anatomical structures, used for focused analysis and processing. AA-ROIs can be determined using pre-defined masks or automatic segmentation algorithms to extract local image patches for refined feature analysis and reduce interference from irrelevant backgrounds.
[0035] The ReLU (Rectified Linear Unit) activation function is a commonly used activation function for neural networks. Its mathematical expression is f(x) = max(0, x), meaning it outputs 0 when the input is negative and a linear value when the input is positive. The ReLU function introduces non-linearity and accelerates network convergence. Due to its computational simplicity and ability to alleviate the vanishing gradient problem, it has become one of the most widely used activation functions in deep learning.
[0036] The Softmax function is a normalization function that maps a vector to a probability distribution, often used in the output layer of multi-class classification tasks. Given a K-dimensional vector z, the Softmax function calculates the ratio of the exponent of each element to the sum of the exponents of all elements. The output value is between 0 and 1, with a total sum of 1, representing the probability that the sample belongs to each class. Softmax can amplify the difference between high and low scores, highlighting the most likely class.
[0037] Sparsemax normalization is a sparse probability normalization method that replaces the exponential function of Softmax with Euclidean projection, making some class weights in the output probability distribution exactly 0, thus achieving hard selection. Sparsemax is suitable for scenarios requiring feature selection or sparse output, and can produce a more compact weight allocation in attention mechanisms, enhancing the interpretability of the model.
[0038] One-hot encoding is a method of representing discrete class labels as binary vectors, where the length of the vector equals the number of classes, and the elements corresponding to the true class are 1s, with the rest being 0s. In deep learning classification tasks, one-hot encoding is often used to calculate cross-entropy loss, making the difference between the model's output probability distribution and the true distribution quantifiable, which facilitates gradient backpropagation optimization.
[0039] The Adam optimizer is a gradient descent optimization algorithm with an adaptive learning rate. It dynamically adjusts the learning rate by calculating the first and second moment estimates of the gradient. The Adam optimizer is characterized by fast convergence, minimal parameter tuning, and applicability to large-scale data, making it one of the most popular optimizers for training deep learning models.
[0040] The area under the receiver operating characteristic (AUC) curve is a comprehensive metric used to evaluate the performance of binary classification models. AUC values range from 0 to 1; the closer to 1, the stronger the model's ability to distinguish between positive and negative samples, while 0.5 indicates no discriminative ability. AUC is unaffected by the classification threshold and reflects the model's overall performance at different thresholds, making it widely used in the evaluation of medical diagnostic models.
[0041] The Receiver Operating Characteristic Curve (ROC curve) is a curve plotted with the false positive rate (1-specificity) on the x-axis and the true positive rate (sensitivity) on the y-axis. The ROC curve shows the performance change of a classification model at different thresholds, and the area under the curve (AUC) is used to quantify the model's performance. The ROC curve provides a direct comparison of the diagnostic capabilities of different models and is a commonly used evaluation tool in medical image analysis.
[0042] With the widespread application of artificial intelligence technology in the field of medical image analysis, the use of diagnostic models to assist in the diagnosis of diseases has become a research hotspot.
[0043] Currently, methods for assisted diagnosis based on medical imaging typically rely on high-quality, specialized scan images as model input. These images, such as those obtained through enhanced CT scans or echocardiography, often require specific examination methods. For instance, enhanced CT scans involve injecting contrast agents and using specialized scanning equipment, while echocardiography requires additional examination using specific scanning probes. By acquiring these specialized scan images, high-contrast, high-resolution image data of the area to be diagnosed can be provided to the diagnostic model, thereby enabling assisted diagnosis of diseases.
[0044] However, the aforementioned technical solutions have certain technical problems in practical applications. The acquisition of specific scan images is often limited by the actual application scenario. For example, contrast-enhanced CT scans not only require additional examinations, increasing examination costs and time, but may also pose risks to patients such as allergies and kidney damage. Echocardiography acquisition requires specific scanning protocols and places additional demands on scanning equipment, which is greatly limited in areas with underdeveloped medical facilities. At the same time, existing diagnostic methods mostly use single-scale image analysis, making it difficult to simultaneously consider both overall anatomical structures and local lesion details. This results in insufficient utilization of feature information and an inability to opportunistically screen a large number of existing conventional non-contrast CT images, leading to a waste of potential diagnostic value and thus reducing the efficiency and accuracy of disease diagnosis.
[0045] In view of this, this specification provides a medical image diagnosis method, and also relates to a medical image diagnosis device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0046] See Figure 1 , Figure 1 A flowchart of a medical imaging diagnostic method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0047] Step 102: Obtain an unenhanced scan image of the site to be diagnosed. The unenhanced scan image is obtained by scanning the site to be diagnosed without injecting contrast agent.
[0048] The medical imaging diagnostic methods provided in one or more embodiments of this specification can be widely applied to various medical imaging diagnostic scenarios. Specifically, they can be applied to various medical imaging-assisted diagnostic systems, including sub-scenarios such as cardiac disease screening, aortic valve morphology analysis, and opportunistic screening. Cardiac disease screening is further subdivided into bicuspid aortic valve (BAV) screening and tricuspid aortic valve (TAV) classification diagnosis. Aortic valve morphology analysis can include leaflet number assessment and valve structural abnormality detection. Opportunistic screening can cover routine chest CT examinations. Specifically, it can be deployed in a medical institution's Picture Archiving and Communication System (PACS) as a background service to automatically analyze massive amounts of routine examination images, enabling auxiliary diagnosis and opportunistic screening of diseases. It can also be seamlessly integrated with the hospital's existing imaging system as an additional functional module for routine examinations. Furthermore, this method can be directly embedded into various medical imaging acquisition devices, such as computed tomography (CT) and magnetic resonance imaging (MRI) devices, enabling the device to output diagnostic prompts immediately after scanning, assisting radiologists in rapid interpretation. Unenhanced scan images can be used as input for various applications without relying on contrast agents or special scanning protocols, offering broad deployment flexibility and clinical applicability.
[0049] The site of diagnosis is an anatomical region in medical imaging diagnosis that needs to be analyzed to determine the presence of specific diseases or abnormalities. It can be used to define the scope and objectives of diagnostic analysis, and may specifically include the aortic valve region, pulmonary nodule regions, and cerebral vascular regions. Specifically, the determination of the site of diagnosis is based on anatomical knowledge and clinical diagnostic needs. It is located using pre-defined anatomical structural features and can be used to identify regions of interest in scanned images. This allows for the subsequent cropping of global and multiple local images, thus providing a foundation for multi-scale feature extraction.
[0050] Unenhanced scan images refer to raw medical imaging data obtained when scanning a target anatomical region without the injection of contrast agents. They can provide an imaging basis independent of contrast agents. Specifically, they can include non-ECG-gated conventional CT scan images, contrast-free MRI images, and ordinary X-ray films. The advantages of unenhanced scan images are that they do not require additional contrast agents, reducing the risk of allergies and kidney burden on patients, while also reducing examination costs and complexity. Specifically, unenhanced scan images are acquired using conventional CT scanning equipment, requiring no special scanning protocols, facilitating subsequent processing and analysis. Unenhanced scan images are fundamental data for disease diagnosis. Through preprocessing and cropping, global images and multiple local images can be obtained, which can then be input into medical imaging diagnostic models for analysis.
[0051] Contrast agents, also known as contrast agents, are chemical substances used to enhance the contrast of specific tissues or organs in medical images. They can be introduced into the body through injection or oral administration to enhance the visibility of specific tissues or organs in images, allowing doctors to observe anatomical structures more clearly. Common contrast agents include iodinated contrast agents (commonly used in CT scans), gadolinium-based contrast agents (commonly used in MRI examinations), and oral barium contrast agents (commonly used in gastrointestinal imaging). Specifically, after entering the body through injection or oral administration, contrast agents accumulate in specific tissues or organs, improving image contrast and making the tissues in the area to be diagnosed more clearly visible. However, they may carry risks such as allergic reactions and kidney damage.
[0052] In practical applications, unenhanced scan images can be obtained through various means or methods.
[0053] One alternative approach is to acquire data directly from medical imaging equipment. For example, a routine chest CT scan can be performed on the patient using a computed tomography (CT) scanner. Data is acquired according to preset scanning parameters (such as tube voltage, tube current, slice thickness, pitch, etc.) and three-dimensional volumetric data is reconstructed. This process does not involve the injection of any contrast agent. Optionally, the preset scanning parameters can be: slice thickness 0.625-5mm (preferably 1-1.5mm), tube voltage 100-140kVp, tube current 100-300mAs, without using ECG gating technology, and without injecting iodine contrast agent.
[0054] Another option is to retrieve existing historical image data from the hospital's image storage system (such as a PACS system). This data could be images generated from routine examinations performed on patients for other clinical needs (such as lung cancer screening, physical examinations, etc.).
[0055] In addition, unenhanced scan images can also be read from pre-saved image files on mobile storage devices or cloud storage. Regardless of the acquisition method, the obtained images are unenhanced scan images without contrast agent injection and can be directly used for subsequent processing.
[0056] This step avoids the risks of allergies and kidney damage associated with contrast agents by acquiring unenhanced scan images. It also eliminates the need for additional examination procedures, reducing examination costs and time. This allows subsequent diagnostic processes to directly utilize existing routine CT image data, providing a data foundation for opportunistic screening. It avoids the reliance on special scan images in traditional methods, expands the applicability of diagnostic methods, and improves the efficiency and safety of medical diagnosis.
[0057] Step 104: Based on the unenhanced scan image, determine the global image and multiple local images of the area to be diagnosed.
[0058] In medical imaging diagnosis, a global image is a complete image region encompassing the overall anatomical structure of the site to be diagnosed. It can be used to extract macroscopic anatomical features, overall morphological information, and control context information of the site. Specifically, a global image typically includes the complete organ outline, adjacent structures, and their relative positions. For example, in aortic valve diagnosis, the global image may include the entire heart, the ascending aorta, and some surrounding tissues. Global images can be obtained through preprocessing of unenhanced scan images and can be cropped to a preset size and range. The specific size can be set according to the anatomical features of the site to be diagnosed and subsequent processing requirements, such as cropping to a 128×128×128 voxel three-dimensional image block.
[0059] Local images, in medical imaging diagnosis, are multiple smaller 3D image patches cropped from a global image, focusing on key regions of the site to be diagnosed. They can be used to extract fine structural details and local lesion details of the site. Local images typically cover anatomical sub-regions crucial to diagnosis, such as the leaflets, orifice, and adjacent aortic root of the aortic valve. The number of local images can be predetermined, such as seven, and each local image has a small size, such as 64×64×64 voxels. Specifically, local images are obtained by cropping the global image according to a preset size based on target segmentation points, and are used to capture key detailed features of the site to be diagnosed.
[0060] In practical applications, the global image and multiple local images can be determined based on the unenhanced scanned image.
[0061] Specifically, the acquired unenhanced scan image can be preprocessed, such as by resampling, grayscale adjustment, and normalization, to obtain a standardized global image; the region to be diagnosed is determined based on a preset mask of the region to be diagnosed, and then the target segmentation point is determined based on the anatomical features of the region; based on the determined target segmentation point, the global image is cropped according to a preset size to obtain multiple local images.
[0062] The preset mask can be determined based on an automatic segmentation algorithm; the target segmentation point can be determined based on anatomical knowledge, such as using the calculated center of mass of the left ventricle as a reference point. The cropping size and number of local images can be adjusted according to specific application requirements.
[0063] In this step, by acquiring unenhanced scan images and identifying global and multiple local images, a multi-scale input foundation is constructed. This enables medical imaging diagnostic methods to simultaneously utilize global anatomical structure information and local lesion details, avoiding the limitations of single-scale analysis. Preprocessing steps standardize image data, improving the consistency and stability of model input. Anatomical segmentation point determination and local image cropping ensure coverage of key regions, providing high-quality input data for subsequent feature extraction and diagnosis. This enhances diagnostic accuracy and robustness, providing reliable technical support for opportunistic screening.
[0064] Step 106: Input the global image and multiple local images into the encoding layer of the medical image diagnosis model for feature encoding, and obtain the global feature vector and multiple local feature vectors accordingly. The medical image diagnosis model also includes an attention layer, a feature fusion layer and a decoding layer.
[0065] The medical image diagnostic model is a computational model based on a deep learning neural network architecture. It can process medical image data and output diagnostic results, performing feature extraction and classification of medical images to assist in disease diagnosis. Specifically, the model includes an encoding layer, an attention layer, a feature fusion layer, and a decoding layer, achieving automatic analysis of medical images through multi-stage processing. The medical image diagnostic model, together with the encoding, attention, feature fusion, and decoding layers, constitutes a complete diagnostic process. The encoding layer is responsible for extracting feature representations of the input image; the attention layer adaptively aggregates local features; the feature fusion layer integrates global and local features; and the decoding layer generates the final diagnostic result. Optionally, the medical image diagnostic model can employ network architectures such as 3D ResNet18, initialized with pre-trained weights from MedicalNet, thereby improving performance under small sample conditions.
[0066] The encoding layer in a medical image diagnostic model is responsible for converting the input image into a feature representation. It can be used to extract features from both the global and local input images, generating corresponding feature vectors. Specifically, the encoding layer can include multiple convolutional layers, pooling layers, and non-linear activation functions, comprising two parallel branches: a global encoder and a local encoder, which process the global image and multiple local images, respectively. The global encoder is responsible for extracting the overall anatomical structure features of the area to be diagnosed, while the local encoder is responsible for extracting fine features from key regions.
[0067] Feature encoding is the process of nonlinearly transforming raw image data and mapping it to feature vectors in a high-dimensional feature space. It can be used to convert the visual information of an image into a mathematical representation, facilitating subsequent analysis and classification. Specifically, feature encoding processes the input image using a convolutional neural network to generate high-dimensional feature vectors representing the feature information contained in the image, such as the shape, texture, and boundaries of organs, as well as abnormal manifestations in lesion areas. The specific implementation of feature encoding depends on the network structure of the encoding layer and the pre-trained weights; for example, a three-dimensional residual network pre-trained on a large-scale medical image dataset can be used for initialization. Through feature encoding, the input global image can be transformed into a compact global feature vector, and each local image can be transformed into a corresponding local feature vector. Different feature vectors encapsulate image information at different scales and levels, providing a foundation for subsequent attention weighting and feature fusion.
[0068] A global feature vector (GMV) is a feature vector representing the overall anatomical structure of the area to be diagnosed. It can be used to capture the overall morphology and spatial relationships of the area. Specifically, GMVs typically have high dimensionality, such as 512 dimensions, where each dimension corresponds to a certain abstract feature learned by the network. GMVs integrate key information from the entire image region, such as the overall shape, size, location of organs, and their relationship with neighboring structures, providing a macroscopic feature representation for diagnostic results.
[0069] Local feature vectors (LANs) represent the features of key local regions within a diagnostic site. They are used to capture the fine structure and local lesion features of the site. Specifically, each local image corresponds to a LAN, for example, a 512-dimensional LAN; multiple local images generate multiple LANs. LANs focus on sub-regions crucial to diagnosis, such as the morphology of valve leaflets, calcifications, and apex size. These LANs can be further input into an attention layer for weighted aggregation, allowing the model to adaptively emphasize the local regions most relevant to the diagnosis, providing a detailed feature representation for the diagnostic results.
[0070] In practical applications, feature encoding can be performed on both the global image and the local image.
[0071] Specifically, the global image can be input into the global encoder of the medical image diagnosis model. This global encoder can be built based on a three-dimensional residual network (such as 3D ResNet18), which performs downsampling and feature extraction step by step through multiple three-dimensional convolutional layers, and outputs a fixed-dimensional global feature vector through a global average pooling layer.
[0072] Similarly, multiple local images can be input into a local encoder, which can use the same or different network structure as the global encoder. For example, it can also use 3D ResNet18. By performing independent feature encoding on each local image, it outputs multiple corresponding local feature vectors.
[0073] Optionally, during feature encoding, the encoder can be initialized using pre-trained weights, such as parameters provided by large-scale medical image pre-trained models like MedicalNet, to accelerate convergence and improve feature extraction quality.
[0074] Furthermore, there are various alternatives to the network structure of the coding layer. For example, different deep learning architectures such as 3D DenseNet, 3D Vision Transformer, and 3DEfficientNet can be used. Alternatively, a hybrid 2D and 3D network can be used. For instance, 2D convolutions can be used to extract intra-layer features, and then 3D convolutions can be combined to extract cross-layer features. Residual networks of different depths, such as 3D ResNet50 and 3D ResNet101, can also be used.
[0075] Global encoders and local encoders can also employ different network structures to adapt to input images of different scales.
[0076] In this step, by inputting the global image and multiple local images into the encoding layer of the medical image diagnosis model for feature encoding, the global feature vector and multiple local feature vectors are obtained. This enables the medical image diagnosis method to simultaneously extract the overall anatomical structure information and key local detail features of the site to be diagnosed, avoiding the limitations of single-scale analysis and making full use of the complementary information of images at different scales. This provides high-quality feature input for subsequent attention calculation and feature fusion, thereby improving the accuracy and robustness of diagnosis.
[0077] Step 108: Input multiple local feature vectors into the attention layer for attention calculation to obtain local aggregated feature vectors.
[0078] The attention layer is a network layer in medical image diagnostic models used to adaptively aggregate local features. It can weight multiple local feature vectors, enabling the model to adaptively focus on key regions relevant to diagnosis. The attention layer typically consists of a series of learnable neural network layers, such as fully connected layers, activation function layers, and normalization layers, dynamically assigning weights based on the importance of the input features. The attention layer receives multiple local feature vectors from the output of the encoding layer, generates an attention weight for each local feature vector through attention calculation, and aggregates the local feature vectors accordingly, outputting a local aggregated feature vector that integrates information from all local regions but focuses more on the key regions.
[0079] Optionally, the attention layer may include a first fully connected layer and a second fully connected layer. A scalar weight score is calculated for each local feature vector through the two fully connected network layers, and then weighted aggregation is performed after Softmax normalization.
[0080] Attention computation is a method used in the attention layer to determine the importance of local features. It can be used to assign corresponding weights to each local feature vector, thereby achieving adaptive attention to key local regions. Specifically, attention computation may include calculating an initial attention score for each input local feature vector through a learnable mapping function; converting these attention scores into attention weights that sum to 1 through a normalization function; further, multiplying each feature vector by its corresponding attention weight and summing the results to obtain a weighted aggregated feature vector.
[0081] Local aggregated feature vectors (LAMs) are feature vectors obtained by performing attention calculations on multiple local feature vectors through an attention layer. They can be used to represent comprehensive information about key local regions of the site to be diagnosed. Specifically, the dimensionality of the LAM is usually consistent with that of a single local feature vector, such as 512 dimensions. Its value is obtained by weighted summation of the local feature vectors. Each element in the LAM reflects a weighted combination of features from multiple local regions; higher weights indicate that the corresponding local region is more important for diagnosis. LAMs can more accurately capture disease-related microstructural features, such as the degree of valve calcification and abnormal valve leaflet morphology.
[0082] In practical applications, attention calculation can be achieved in a variety of ways.
[0083] One alternative approach is to input multiple local feature vectors into an attention layer, which may include a multilayer perceptron network. For each local feature vector, a fully connected layer can be used to reduce its dimensionality to an intermediate representation, and then another fully connected layer can be used to map it to a scalar attention score. The attention scores of all local feature vectors are normalized using a predefined activation function so that the sum of all weights is 1. Then, each local feature vector is multiplied by its corresponding normalized attention weight, and all weighted vectors are summed element-wise to obtain a local aggregated feature vector with the same dimension as a single local feature vector.
[0084] Another option is to use a self-attention mechanism, which can generate an attention matrix by calculating the pairwise similarity between all local feature vectors, and then use this matrix to weight and combine the local feature vectors.
[0085] In addition, a multi-head attention mechanism can be used to concatenate or average the outputs of multiple attention heads to obtain richer feature representations.
[0086] Furthermore, the network structure of the attention layer can be flexibly adjusted. For example, the first fully connected layer can map the 512-dimensional features to an intermediate dimension, such as 256, 128, or 64 dimensions, and then the second fully connected layer maps the intermediate dimension to 1 dimension. Different activation functions such as ReLU and Tanh can be used.
[0087] In this step, by inputting multiple local feature vectors into the attention layer for attention calculation, higher weights can be adaptively assigned to key regions relevant to diagnosis, while suppressing interference from irrelevant regions. This results in more representative local aggregated feature vectors, which can focus on local regions crucial to diagnosis. This avoids the limitations of simple averaging or max pooling of local features, providing high-quality local feature representations for subsequent feature fusion and diagnosis. Consequently, the accuracy and robustness of the entire diagnostic method are improved, providing key technical support for comprehensive diagnostic analysis of medical images.
[0088] Step 110: Input the global feature vector and the local aggregated feature vector into the feature fusion layer for feature fusion calculation to obtain the fused feature vector.
[0089] The feature fusion layer is a network component in medical image diagnostic models used to integrate global and local features. It fuses global feature vectors and locally aggregated feature vectors to generate a feature representation that integrates macroscopic and microscopic information. Specifically, the feature fusion layer receives global and locally aggregated feature vectors as input, integrates them through a pre-defined fusion strategy, and outputs a fused feature vector. Closely related to the attention and decoding layers of the medical image diagnostic model, the feature fusion layer is a crucial link connecting local feature aggregation and the final diagnosis; its fusion strategy directly affects the accuracy of the diagnostic results.
[0090] Feature fusion computation is a mathematical method used in the feature fusion layer to integrate different feature information. It can be used to perform mathematical operations on global feature vectors and local aggregated feature vectors to generate fused feature vectors. Specifically, feature fusion computation includes weighted summation, gated fusion, or bilinear fusion of global and local aggregated feature vectors. It can also concatenate global and local aggregated feature vectors in terms of dimension to form higher-dimensional vectors. For example, directly concatenating a 512-dimensional global feature vector and a 512-dimensional local aggregated feature vector in terms of feature dimension yields a 1024-dimensional fused feature vector. The optimal combination of features is achieved through a preset fusion strategy.
[0091] A fused feature vector is a feature representation obtained by fusing global and local aggregated feature vectors through feature fusion calculation. It can be used to represent the comprehensive feature information of the site to be diagnosed, providing input for subsequent diagnosis. Specifically, the dimension of the fused feature vector depends on the fusion method used. For example, if vector concatenation is used, the dimension of the fused feature vector is equal to the sum of the dimensions of the global and local aggregated feature vectors (e.g., 1024 dimensions); if weighted summation is used, the dimension remains consistent with that of a single feature vector (e.g., 512 dimensions). The fused feature vector integrates the overall anatomical structure information of the site to be diagnosed and the fine features of key local regions. It can include both background information such as the macroscopic morphology and spatial location of organs, and key clues such as microscopic details and abnormal manifestations of lesions. Therefore, it can more comprehensively and accurately characterize the pathological state of the site to be diagnosed, providing a more comprehensive feature representation for diagnosis, enabling the model to simultaneously focus on the overall morphology and key local features of the site to be diagnosed.
[0092] In practical applications, feature fusion calculations to obtain fused feature vectors can be achieved in various ways.
[0093] One alternative approach is to use vector concatenation, which connects the global feature vector and the local aggregated feature vector along their feature dimensions to form a new, higher-dimensional fused feature vector. For example, if both the global and local aggregated feature vectors are 512-dimensional, the concatenated fused feature vector will be 1024-dimensional. Vector concatenation is simple and effective, and it completely preserves all the original feature information.
[0094] Another alternative approach is to use weighted summation fusion. This involves assigning a learnable weight parameter to both the global feature vector and the local aggregated feature vector, and then performing an element-wise weighted sum to obtain a fused feature vector with the same dimension as the input. The weight parameters can be automatically optimized during model training to dynamically balance the contributions of global and local information to the diagnostic results.
[0095] In addition, a gated fusion mechanism can be used, where a gated network (such as a small fully connected network) dynamically generates weights based on the input features, enabling adaptive fusion of global and local features. More complex fusion methods can include bilinear fusion, which takes two feature vectors and performs an outer product operation to obtain a feature matrix, then flattens or reduces its dimensionality to obtain the fused feature vector. This approach can capture high-order interaction information between features.
[0096] Furthermore, the feature fusion layer can also employ a multilayer perceptron to perform further nonlinear transformations on the concatenated fused feature vectors to extract higher-level fused features. Specific feature fusion calculation methods can be flexibly selected or combined according to actual task requirements and data characteristics; the embodiments in this specification do not impose specific limitations on this.
[0097] In this step, the global feature vector and the local aggregated feature vector are input into the feature fusion layer for feature fusion calculation to obtain the fused feature vector. This enables the medical image diagnosis method to integrate global anatomical structure and local lesion details, avoiding the limitations of single feature representation. Through feature fusion calculation, macroscopic and microscopic feature representations are integrated, improving the richness and discriminativeness of feature representation, thereby enhancing the accuracy and robustness of diagnosis and providing key technical support for comprehensive diagnostic analysis of medical images.
[0098] Step 112: Input the fused feature vector into the decoding layer for decoding to obtain the medical diagnosis result of the area to be diagnosed.
[0099] The decoding layer, located after the feature fusion layer in a medical image diagnostic model, is a network component responsible for converting fused feature vectors into diagnostic categories. It maps high-dimensional feature representations to specific disease classification results, achieving the transformation from feature vectors to diagnostic outcomes. Specifically, the decoding layer typically includes one or more fully connected layers and can incorporate activation functions (such as the Softmax function) to output class probabilities. For example, in a binary classification task, the decoding layer might include a fully connected layer with an output dimension of 2, mapping the fused feature vectors to the raw scores of the two classes, and then converting them to predicted probabilities using the Softmax function. The specific structure of the decoding layer can be adjusted according to the needs of the diagnostic task; for example, for multi-class tasks, the output dimension can be increased; for regression tasks, a linear layer can be used to directly output numerical values.
[0100] Decoding is the mathematical process by which the decoding layer computes the fused feature vector and maps it to the diagnostic result space. It can be used to map representations in the feature space to specific diagnostic category probability distributions, thus achieving disease classification. Specifically, decoding performs linear transformations and non-linear activations on the fused feature vector through fully connected layers to generate a predicted score for each diagnostic category. Then, a normalized activation function converts the scores into probability values. Decoding transforms high-order abstract features extracted by the model into clinically understandable information, such as disease category and abnormality level. Its specific implementation depends on the design of the decoding layer and the task type. For classification tasks, decoding typically includes fully connected computation and probability normalization; for segmentation tasks, decoding may also include upsampling and pixel-level classification. Through decoding, the macroscopic and microscopic feature information contained in the fused feature vector is transformed into interpretable diagnostic results, enabling the model to output a clear disease classification.
[0101] Medical diagnostic results are the final output of a medical imaging diagnostic model, providing information about the state or category of the site to be diagnosed. These results can assist physicians in making clinical decisions. Medical diagnostic results may include category labels (e.g., "bicuspid aortic valve" or "tricuspid aortic valve"), and may further include confidence probabilities for each category (e.g., a probability of 0.95 for a bicuspid aortic valve and 0.05 for a tricuspid aortic valve), as well as more granular quantitative indicators (such as lesion severity scores). Medical diagnostic results can be directly displayed on the system's user interface or stored along with medical images for physicians to review and diagnose.
[0102] In practical applications, the fused feature vectors are decoded to obtain medical diagnostic results, and the corresponding methods can be adopted according to the actual diagnostic task.
[0103] One possible approach is to input the fused feature vector into a fully connected layer, where the number of neurons equals the preset number of diagnostic categories, for example, 2 for binary classification tasks. This fully connected layer can perform a linear transformation on the fused feature vector, outputting the raw score for each category. The raw scores can then be input into an activation function for normalization, yielding the predicted probability for each category, with the sum of all probabilities equal to 1. Finally, based on the predicted probabilities, the category with the highest probability can be selected as the final medical diagnosis, and this probability value can be output as the confidence level. Alternatively, a diagnosis can be obtained based on a preset probability threshold, where the predicted probability reaches the threshold. For example, the probability threshold for binary classification tasks could be 0.5, while for multi-class classification tasks, it can be flexibly set based on the actual needs of different categories.
[0104] Another option is to use multiple fully connected layers stacked together to form the decoder, thereby enhancing non-linear expressive power. Alternatively, the decoder layer can use traditional machine learning classifiers such as support vector machines or random forests instead of fully connected layers, taking the fused feature vectors as input and directly outputting the classification result.
[0105] Furthermore, for diagnostic tasks requiring continuous values (such as aortic valve calcification scores), the decoding layer can employ a linear regression layer to output a single real value. The decoding layer can also be designed for multi-task output, simultaneously providing classification results and auxiliary information.
[0106] In this step, the fused feature vector is input into the decoding layer for decoding to obtain the medical diagnostic result of the site to be diagnosed. This enables the medical imaging diagnostic method to output a clear diagnostic category based on comprehensive feature information, avoiding fuzzy judgments. Through normalization processing, the probability distribution of the diagnostic results is ensured to be reasonable, which is convenient for clinical interpretation. Through the output of the decoding layer, a complete conversion from feature representation to clinical diagnosis is realized, providing doctors with reliable auxiliary diagnostic basis and improving the accuracy and practicality of diagnosis.
[0107] In this embodiment, a multi-scale input foundation is constructed by acquiring unenhanced scan images and determining the global image and multiple local images from them. An encoding layer is used to encode the features of the global and local images respectively, obtaining a global feature vector and multiple local feature vectors covering the overall anatomical structure and key local details. Then, an attention layer performs attention calculations on the multiple local feature vectors, adaptively assigning higher weights to diagnostically relevant key regions, thereby suppressing interference from irrelevant regions and obtaining more representative local aggregated feature vectors. A feature fusion layer fuses the global features and local aggregated features to generate a fused feature vector that integrates macroscopic and microscopic information. A decoding layer outputs accurate medical diagnostic results based on this fused feature. This method fully explores the potential diagnostic value of conventional scan images without relying on contrast agents or other additional scanning examinations, avoiding the examination risks and costs associated with contrast agents or other additional examinations. Simultaneously, the effective combination of multi-scale feature extraction and attention mechanisms improves the accuracy and robustness of the diagnosis, providing an efficient and universal technical solution for comprehensive diagnostic analysis of medical images.
[0108] In one optional embodiment of this specification, based on unenhanced scan images, a global image and multiple local images of the site to be diagnosed are determined, including: Preprocess the unenhanced scan images to obtain a global image of the area to be diagnosed; Based on the preset mask of the area to be diagnosed, the target segmentation point is determined; Based on the target segmentation points, the global image is cropped according to a preset size to obtain multiple local images.
[0109] Preprocessing is a step of standardizing and normalizing the raw, unenhanced scan images. It can improve the consistency and quality of image data, facilitating subsequent feature extraction and analysis. Specifically, preprocessing includes operations such as resampling, grayscale adjustment, and normalization to bring the image data into a uniform format required by diagnostic model inputs. For example, the voxel spacing can be standardized to 1.0 × 1.0 × 1.0 mm³, and image values can be normalized to the [0, 1] interval. Furthermore, for CT images, the window width and level of the Hu values can be adjusted; for example, a grayscale mapping can be performed using a window width of 800-1200 Hu and a window level of 200-400 Hu. The preprocessed image data will serve as the basis for the global image and will be further used for local image cropping.
[0110] A preset mask is a predefined binary image template or region marker used to identify a region to be diagnosed, indicating the extent of the area. Specifically, preset masks can be generated based on automatic segmentation algorithms, such as using a deep learning-based organ segmentation model to learn from training data and automatically predict the mask region of the area to be diagnosed on a new image. Preset masks can also be manually drawn based on prior anatomical knowledge; for example, for the aortic valve region, a standard template including the left ventricle and the aortic root can be predefined. Preset masks can be used to determine the boundaries of the region to be diagnosed; for example, in the diagnosis of aortic valve disease, a preset mask can mark the areas of the heart and aorta.
[0111] Target segmentation points are key locations determined based on the anatomical structures of the area to be diagnosed, and can be used to determine the cropping center of a local image. Specifically, target segmentation points can be obtained by calculating the geometric center, center of mass, or reference points of specific anatomical structures of the region to be diagnosed. For example, in the diagnosis of aortic valve disease, the center of mass of the left ventricle can be calculated as the target segmentation point. They can also be determined based on image features such as the maximum density projection point or the location of the largest cross-section; or by selecting specific anatomical landmarks based on prior clinical knowledge, such as the junction of the aortic valve leaflets or the center of the valve orifice. Target segmentation points form the basis for cropping the global image to a preset size, ensuring that the local image focuses on the anatomical regions most important for diagnosis.
[0112] Preset sizes are the pre-defined sizes of 3D image patches when cropping local images. They can be used to standardize the input dimensions of all local images, making them suitable for the input requirements of medical image diagnostic models. Preset sizes can be set according to the anatomical characteristics of the area to be diagnosed, computational resources, and network structure. For example, they can be set to different specifications such as 32×32×32 voxels, 40×40×40 voxels, 48×48×48 voxels, 56×56×56 voxels, and 64×64×64 voxels. Smaller preset sizes can capture more refined local features, while larger preset sizes can include more contextual information, providing a suitable spatial range for subsequent feature extraction.
[0113] Cropping is the process of extracting multiple local image patches from a global image according to a preset sampling strategy. It can be used to generate multiple local images focusing on key areas of the site to be diagnosed. Specifically, cropping can be based on a target segmentation point. For example, an image patch of a preset size can be extracted from the global image with the target segmentation point as the center; or multiple image patches can be extracted around the target segmentation point with a fixed step size and grid layout. Optionally, cropping methods can include different types such as non-overlapping cropping and overlapping cropping. Non-overlapping cropping means that there is no overlapping area between the obtained local image patches, and the boundaries are strictly adjacent, which can maximize the coverage of the site to be diagnosed and reduce redundancy. Overlapping cropping means that there is a partial overlap between adjacent local image patches. That is, a step size smaller than the preset size is set during grid sampling, so that adjacent blocks share some voxels, which can increase the continuity of feature coverage and avoid the interruption of continuous anatomical structures due to block boundary division.
[0114] In practical applications, when an unenhanced scanned image is obtained, preprocessing can be performed. This can include steps such as resampling, grayscale adjustment, and normalization to obtain a standardized global image.
[0115] Furthermore, the target segmentation point can be determined based on a preset mask of the area to be diagnosed. Specifically, the geometric center or mass center of the area to be diagnosed, determined by the preset mask, can be calculated as the target segmentation point. For example, in the diagnosis of aortic valve disease, the mass center of the left ventricle can be calculated as the target segmentation point. This method, based on anatomical knowledge, ensures that the target segmentation point is located at a key position in the area to be diagnosed, providing an accurate center point for cropping the local image.
[0116] Once the target segmentation points are determined, the global image can be cropped according to a preset size based on the target segmentation points to obtain multiple local images.
[0117] Specifically, multiple local images can be formed by cropping around the target segmentation point according to a preset size. The cropping can be done using overlapping cropping (e.g., local images can have a 10% overlap area) or non-overlapping cropping. Overlapping cropping ensures that key regions are covered multiple times, improving the completeness of feature extraction, while non-overlapping cropping avoids feature duplication and improves computational efficiency. Different cropping methods can be flexibly selected according to specific application scenarios and diagnostic needs.
[0118] For example, an image patch representing the global image of the area to be diagnosed can be cropped from an unenhanced scan image. The patch can be 192×224×160 pixels in size and then resampled to 128×128×128 pixels using trilinear interpolation. This patch can cover the entire heart and ascending aorta region and be used for global feature extraction. From the global image, K local image patches of 48×48×48 voxels can be further cropped and resampled to 64×64×64 pixels to cover the entire heart and ascending aorta region and be used for local fine feature extraction.
[0119] In the embodiments of this specification, a standardized global image is obtained through preprocessing, which improves the consistency and quality of image data; the region to be diagnosed is determined by a preset mask, which ensures the accuracy of the target segmentation points; cropping is performed based on the target segmentation points and preset sizes, which ensures that the local image covers the key areas of the region to be diagnosed. This allows the medical imaging diagnostic method to utilize both global anatomical structure information and local lesion details, avoiding the limitations of single-scale analysis, constructing multi-scale input data, improving the accuracy and robustness of diagnosis, and providing reliable technical support for opportunistic screening.
[0120] In one optional embodiment of this specification, preprocessing is performed on the unenhanced scan image to obtain a global image of the area to be diagnosed, including: The unenhanced scan image is resampled to obtain a resampled unenhanced scan image; The grayscale of the resampled, unenhanced scanned image is adjusted to obtain the adjusted unenhanced scanned image. The adjusted, unenhanced scan image is normalized to obtain a global image of the area to be diagnosed.
[0121] Resampling is the process of adjusting the spatial resolution of an unenhanced scanned image. It can be used to unify the resolution differences between images from different scanning devices and protocols, facilitating subsequent feature extraction and model input. Specifically, resampling can adjust the voxel spacing of the original image to a preset uniform value using interpolation algorithms such as trilinear interpolation, nearest neighbor interpolation, or spline interpolation. For example, unifying the voxel spacing of images acquired by different CT devices to 1.0 × 1.0 × 1.0 mm³ involves calculating the pixel value at the new voxel location. This is typically done using trilinear interpolation, which determines the new voxel value by weighted averaging of the surrounding eight voxels, ensuring the image quality after spatial resolution unification.
[0122] The resampled unenhanced scan image is an unenhanced scan image that has undergone resampling processing. It can be used to ensure the consistency of image spatial resolution, facilitating subsequent feature extraction. Specifically, the resampled unenhanced scan image has a uniform voxel spacing, eliminating the inconsistency in spatial resolution of the original image, allowing images from different devices and scanning protocols to be compared and analyzed at the same scale.
[0123] Grayscale adjustment is the process of applying non-linear or linear transformations to the pixel or voxel values of medical images to enhance the contrast of specific anatomical structures and highlight the area to be diagnosed or the region of interest. It can be used to make target tissues in images easier to identify and analyze. Specifically, for CT images, grayscale adjustment typically employs window width and window level adjustment techniques. By setting specific window widths (displaying the grayscale range) and window levels (center values), the Hu values of the original CT image are mapped to a specific grayscale range, such as 0-255, thereby highlighting specific tissues. For example, for mediastinal structures, a window width of 800-1200 HU and a window level of 200-400 HU can be used for mapping, allowing soft tissues such as the heart and major blood vessels to be clearly displayed. Optionally, grayscale adjustment can use different combinations of window width and window level parameters, such as a window width of 600 HU and a window level of 100 HU, a window width of 1500 HU and a window level of 500 HU, etc. Different mapping functions, such as logarithmic mapping and exponential mapping, can also be used to adapt to the anatomical structures and diagnostic needs of different areas to be diagnosed. Grayscale adjustment can affect the visibility of target structures in an image, providing clearer visual information for subsequent feature extraction.
[0124] The adjusted unenhanced scan image is a grayscale-adjusted unenhanced scan image that can be used to improve the contrast of target structures in the image, facilitating subsequent analysis. Specifically, the grayscale values of the adjusted unenhanced scan image have been mapped to a new range (e.g., 0-255), highlighting key anatomical structures in the area to be diagnosed. This image has better visual contrast and tissue differentiation than the original image, facilitating subsequent normalization processing and feature extraction. For example, with a window width of 800 HU and a window level of 200 HU, the structures of the heart and aorta will be clearer, while low-contrast areas will be compressed to black or near black.
[0125] Normalization is the process of linearly scaling the grayscale values of a scanned image to a standard range, such as [0, 1] or [-1, 1]. It can be used to eliminate dimensional differences in grayscale values between different images and improve numerical stability. Specifically, normalization can employ the min-max normalization formula, subtracting the minimum value from each pixel value and then dividing by the difference between the maximum and minimum values, ensuring all pixel values fall within the [0, 1] range. Alternatively, Z-score normalization can be used to give the data zero mean and unit variance. The normalized image becomes the global image of the area to be diagnosed, preserving both the contrast information after grayscale adjustment and a uniform numerical range, making it suitable as input data for medical image diagnostic models.
[0126] In practical applications, resampling can be achieved in a variety of ways.
[0127] One alternative approach is to use a trilinear interpolation algorithm to resample the voxel grid of the original image based on the target voxel spacing, and then calculate the gray value of each voxel in the new grid. For example, if the original image has a layer spacing of 2.5 mm and an in-plane resolution of 0.6 mm × 0.6 mm, it can be resampled to an isotropic resolution of 1.0 mm × 1.0 mm × 1.0 mm using trilinear interpolation.
[0128] Another option is to use nearest neighbor interpolation. This method has lower computational cost and is suitable for diagnostic tasks that need to preserve the original grayscale values, such as avoiding the generation of new grayscale values in segmentation tasks. Alternatively, higher-order interpolation methods such as spline interpolation can be used to obtain smoother resampling results.
[0129] Furthermore, grayscale adjustment of the resampled, unenhanced scan images can be performed using window width and window level adjustment techniques. Specifically, preset window width and window level parameters can be selected based on the tissue type of the site to be diagnosed. For example, a window width of 800 HU and a window level of 200 HU can be used for mediastinal structures, while a window width of 1500 HU and a window level of -500 HU can be used for lung tissue. The original HU values can be linearly mapped using corresponding calculation formulas, linearly stretching the values within the window width range to 0-255. Values below the window level minus half the window width are truncated to 0, and values above the window level plus half the window width are truncated to 255.
[0130] The mapping formula can be expressed as:
[0131] in, The adjusted HU value, i.e., the grayscale value. The original HU value, For window position, For window width.
[0132] In addition, grayscale adjustment can also employ non-linear methods such as histogram equalization and adaptive histogram equalization to enhance local image contrast. Specific grayscale adjustment parameters can be flexibly adjusted based on factors such as scanning protocol and patient body size to adaptively display the anatomical structures of the area to be diagnosed.
[0133] Furthermore, the adjusted unenhanced scanned image can be normalized using a minimum-maximum normalization method. Specifically, for the unenhanced scanned image after grayscale adjustment, the minimum and maximum values of its pixel values can be calculated, and then a preset normalization formula can be applied to each pixel to linearly map the pixel values to the target interval, such as the [0, 1] interval.
[0134] The normalization formula can be expressed as:
[0135] The normalized image will be used as the global image of the area to be diagnosed for subsequent cropping and feature extraction.
[0136] In the embodiments of this specification, images with uniform resolution are obtained through resampling, which improves the spatial consistency of image data; the contrast of the target structure is highlighted by grayscale adjustment, which enhances the visibility of feature extraction; image values are set to a standardized range through normalization processing, which improves numerical stability; through a series of preprocessing steps, a high-quality global image of the area to be diagnosed is obtained, which provides an input basis for subsequent local image cropping and feature extraction, thereby improving the accuracy and robustness of medical image diagnostic methods. This enables diagnostic methods to more effectively utilize the potential diagnostic value of conventional unenhanced scan images, providing reliable technical support for comprehensive diagnostic analysis of medical images.
[0137] In one optional embodiment of this specification, determining the target segmentation point based on a preset mask of the area to be diagnosed includes: The region to be diagnosed is determined based on a preset mask of the area to be diagnosed. Based on the anatomical structure of the area to be diagnosed, target segmentation points are determined from the area to be diagnosed.
[0138] The region to be diagnosed is a specific spatial area encompassing the anatomical structure to be diagnosed, delineated from a medical image based on a predefined mask. It can be used to define the target area for subsequent analysis and reduce interference from irrelevant background. Specifically, the region to be diagnosed can typically be represented as a set of voxels covered by a binary mask. For example, in the diagnosis of aortic valve disease, the region to be diagnosed may include the entire heart, the ascending aorta, and the voxel region of the aortic root. The determination of the region to be diagnosed can affect the degree to which the local image covers key anatomical structures.
[0139] The anatomical structure of the site to be diagnosed refers to its morphology, location, and tissue relationships, which can guide the determination of target segmentation points and the cropping of local images. Specifically, the anatomical structure of the site to be diagnosed can include organ geometry, key anatomical landmarks, and relative positions of tissues. For example, in aortic valve diagnosis, the anatomical structure of the site to be diagnosed can include substructures such as the left ventricle, aortic valve leaflets, valve orifice, aortic sinus, and aortic root, as well as detailed structures such as the arrangement of the leaflets, the location of the valve orifice, and the shape of the valve annulus. Knowledge of the anatomical structure of the site to be diagnosed can be derived from anatomy textbooks, medical imaging data, or clinical experience. Determining target segmentation points based on the anatomical structure of the site to be diagnosed ensures that the segmentation points are located at sites with clear anatomical significance, thereby ensuring that the cropping of the local image covers key areas and providing accurate anatomical guidance for feature extraction.
[0140] In practical applications, the area to be diagnosed can be determined based on a preset mask of the area to be diagnosed.
[0141] Specifically, a voxel-by-voxel logical operation can be performed between the preset mask and the preprocessed global image, that is, the image region corresponding to the voxel with a mask value of 1 is retained, and the remaining voxels are set to the background value. The preset mask can be learned in advance by a deep learning segmentation model on a large amount of training data, which can automatically segment the contour of the area to be diagnosed; the preset mask can also be a standard template drawn manually based on anatomical prior knowledge, which is registered with the current global image to obtain an individualized mask region.
[0142] Furthermore, methods such as connected component analysis can be used to post-process the pre-defined mask, removing small, isolated noise regions and retaining the largest connected component as the region to be diagnosed, thereby improving the robustness of region determination. In addition, morphological operations, such as dilation or erosion, can be performed on the region to be diagnosed to fine-tune its boundaries, making it better conform to the actual anatomical structure.
[0143] Once the region to be diagnosed has been identified, target segmentation points can be determined from that region based on anatomical structures. The specific method for determining these points can be selected based on the anatomical features of the region to be diagnosed.
[0144] One alternative approach is to use the geometric center or mass center of the region to be diagnosed as the target segmentation point. For example, for the left ventricular region, the geometric center can be obtained by calculating the average of all voxel coordinates, or the mass center can be obtained by the weighted average of voxel gray values. This center point is usually located inside the heart chamber and can represent the central location of the entire region well.
[0145] Another alternative approach is to determine the target segmentation point based on specific anatomical landmarks. For example, in aortic valve diagnosis, the junction of the aortic valve leaflets or the center of the valve orifice can be used as the target segmentation point. These junctions or center points can be obtained by analyzing the local grayscale features of the image or by predicting them through a key point detection network.
[0146] In addition, target segmentation points can be determined based on the maximum density projection point. For example, in areas with obvious calcification, the voxel point with the largest CT value can be selected as the segmentation point.
[0147] In practical applications, a combination of various methods can be used to determine the location. For example, the geometric center can be determined first, and then specific anatomical feature points can be searched within the neighborhood of that center to obtain a more precise location. The specific methods can be flexibly selected or combined based on the anatomical characteristics of the site to be diagnosed and clinical needs.
[0148] For example, for aortic valve diagnosis, a preset automatic segmentation algorithm can be used to obtain cardiac and aortic masks. Based on the cardiac centroid, root seed points are determined in the aortic mask, and a centerline is calculated on the aortic mask voxel map. The ascending aorta segment is determined based on the centerline, and the Anatomical Area Region of Interest (AA-ROI), containing the aortic valve and the root of the ascending aorta, is extracted as the region to be diagnosed. Furthermore, the determined region to be diagnosed can be supplemented using methods such as region growing. Once the region to be diagnosed is determined, principal component analysis can be performed on it to obtain the principal axis direction, and K=7 center points are uniformly sampled along the principal axis. K local image blocks are obtained by cropping around each center point, and each local image block can be 48×48×48 pixels.
[0149] In the embodiments of this specification, the region to be diagnosed is determined by a preset mask based on the region to be diagnosed, ensuring the accuracy and consistency of the region boundary; the target segmentation point is determined from the region to be diagnosed based on the anatomical structure of the region to be diagnosed, ensuring the precise positioning of key positions. This allows the cropped local image to focus on the most critical area for diagnosis, enabling the medical image diagnostic method to accurately determine the key area of the region to be diagnosed. This provides a precise center point for subsequent local image cropping, avoids deviations in local feature extraction, improves the accuracy and robustness of diagnosis, and provides reliable technical support for comprehensive diagnostic analysis of medical images.
[0150] In one optional embodiment of this specification, the encoding layer includes a global encoder and a local encoder; The global image and multiple local images are input into the encoding layer of the medical image diagnostic model for feature encoding, resulting in a global feature vector and multiple local feature vectors, including: The global image is input into the global encoder, and downsampled based on a preset global step size to obtain the global feature vector; Multiple local images are input into a local encoder, and downsampled based on a preset local step size to obtain multiple local feature vectors.
[0151] A global encoder is a neural network component in medical image diagnostic models used to process global images and extract overall anatomical features of the area to be diagnosed. Specifically, a global encoder can employ a three-dimensional convolutional neural network architecture, such as 3D ResNet18 or 3D DenseNet. It typically consists of multiple convolutional layers, pooling layers, and non-linear activation functions, and can adapt to the size characteristics of the global image. For example, it can take a 128×128×128 three-dimensional image block as input and further reduce the spatial size of the image and increase the number of channels by stacking multiple convolutional and pooling layers, ultimately outputting a global feature vector of fixed dimensions.
[0152] A local encoder is a neural network component in medical image diagnostic models used to process local images. It can be used to extract fine structural features of key local regions of the site to be diagnosed. Specifically, a local encoder typically has the same or similar network structure as a global encoder, but with a smaller input size, such as a 64×64×64 three-dimensional local image patch. A local encoder can progressively extract features of the local region through multi-level downsampling, outputting multiple corresponding local feature vectors, such as a 512-dimensional local feature vector.
[0153] The global encoder and local encoder can work in parallel, processing the global image and the local image respectively, enabling the model to capture macroscopic and microscopic feature information simultaneously, providing high-quality input for subsequent attention calculation and feature fusion.
[0154] The preset global stride is a parameter in a global encoder used to control the variation of feature map size. It can be used to standardize the downsampling stride during the global image feature extraction process. Specifically, the selection of the preset global stride needs to consider the size of the input image, the depth of the model, and computational resources to ensure that the feature map can be adapted to subsequent global average pooling layers. That is, the preset global stride can be set separately for different network layers. For example, in 3D ResNet18, the stride of the first convolutional layer (Conv1) can be set to 2×2×2, the stride of the first pooling layer (MaxPool) can be set to 2×2×2, and in subsequent residual blocks, except for the first residual block with a stride of 1×1×1, the stride of the remaining residual blocks can be set to 2×2×2. This allows the global image with an input size of 128×128×128 to be progressively downsampled to a 4×4×4 feature map, which is then subjected to global average pooling to obtain a 512-dimensional global feature vector.
[0155] The preset local stride is a parameter in a local encoder used to control the variation of the feature map size. It can be used to standardize the downsampling step size during the local image feature extraction process. Specifically, the selection of the preset local stride also needs to consider the size of the local image, the depth of the model, and computational resources to ensure that the feature map can fit the subsequent global average pooling layer. That is, the preset local stride can use the same setting as the preset global stride. For example, in 3D ResNet18, by combining the preset local strides of different layers, a local image with an input size of 64×64×64 is progressively downsampled to a 2×2×2 feature map, and then global average pooling is performed to obtain a 512-dimensional local feature vector. In addition, the preset local stride can also be adjusted according to the smaller input size of the local image, for example, by reducing the number of downsampling operations or using a smaller stride to retain more spatial detail information.
[0156] Downsampling is the process of reducing the size of feature maps during encoding through convolution or pooling operations. It can be used to expand the receptive field, extract higher-level semantic features, and reduce the computational cost of subsequent network layers. Specifically, downsampling can be achieved by setting convolutional layers with a stride greater than 1 or dedicated pooling layers. For example, a convolutional layer with a stride of 2×2×2 can halve the size of the input feature map in each dimension. During downsampling, the network gradually abstracts higher-order feature representations from local details. For example, low-level features such as edges and textures are gradually combined to form mid-level features such as organ contours and structural relationships, ultimately forming high-level semantic features with discriminative capabilities.
[0157] In practical applications, the downsampling process of global and local images can be implemented by global encoders and local encoders, respectively.
[0158] Specifically, taking 3D ResNet18 as an example, the global encoder receives a global image of size 128×128×128 as input. After passing through a 7×7×7 convolutional layer (Conv1) with a stride of 2×2×2, the feature map size is halved to 64×64×64, and the number of channels is increased to 64. Then, after passing through a 3×3×3 max pooling layer with a stride of 2×2×2, the size is further halved to 32×32×32. Furthermore, the feature map can be sequentially passed through four residual blocks. The first residual block has a stride of 1×1×1, maintaining the size of 32×32×32. The second residual block has a stride of 2×2×2, halving the size to 16×16×16 and increasing the number of channels to 128. The third residual block has a stride of 2×2×2, halving the size to 8×8×8 and increasing the number of channels to 256. The fourth residual block has a stride of 2×2×2, halving the size to 4×4×4 and increasing the number of channels to 512. Finally, global average pooling yields a 512-dimensional global feature vector. Similarly, the local encoder receives a 64×64×64 local image as input. After passing through the same network structure, the final feature map size is 2×2×2, and global average pooling also yields a 512-dimensional local feature vector. The encoder structure described above can be specifically represented as follows: Input layer: Processes global images of 128×128×128 or local images of 64×64×64; Convolutional layer Conv1: 7×7×7 convolution, 64 channels, stride 2×2×2; Max pooling: 3×3×3, step size 2×2×2; Residual block 1: [3×3×3, 64 channels]×2; Residual block 2: [3×3×3, 128 channels]×2, step size 2×2×2; Residual block 3: [3×3×3, 256 channels]×2, step size 2×2×2; Residual block 4: [3×3×3, 512 channels]×2, step size 2×2×2; Global average pooling: outputs a 512-dimensional feature vector.
[0159] In the downsampling process described above, different residual blocks can extract features at different levels. For example, the first residual block can extract low-level features, such as edges, textures, and corners in the image; the second and third residual blocks can extract mid-level features, such as ventricular contours, blood vessel orientation, and valve morphology; the fourth residual block can further extract high-level semantic features, such as key discriminative features that distinguish between bicuspid and tricuspid aortic valves. These features at different levels are abstracted and combined layer by layer to ultimately form a feature vector with strong discriminative power.
[0160] Furthermore, the preset global step size and preset local step size can be flexibly adjusted according to the actual size of the area to be diagnosed, the input image size, and computing resources. For example, for larger areas to be diagnosed, the number of downsampling operations can be increased or a larger initial step size can be used; for tasks that require retaining more details, the number of downsampling operations can be reduced or a smaller step size can be used.
[0161] The global feature vectors and multiple local feature vectors output by the global encoder and local encoder can be directly concatenated along the feature dimension. For example, a 512-dimensional global feature vector and a 512-dimensional local aggregated feature vector can be concatenated to form a 1024-dimensional fused feature vector, so as to fully preserve the feature information under the two different semantic scales and provide a richer feature basis for subsequent diagnostic decisions.
[0162] In the embodiments of this specification, a parallel structure of global and local encoders is used to process global and local images respectively, enabling the medical image diagnostic method to simultaneously extract the overall anatomical features of the site to be diagnosed and the fine features of key local regions. By reasonably setting the preset global and local step sizes, the stability and consistency of the feature extraction process are ensured, allowing the feature map to adapt to the global average pooling layer and obtain a fixed-dimensional feature vector. Through the multi-level design of the downsampling process, the model can extract features at different levels, achieving complementarity and enhancement of multi-scale features. This provides high-quality input for subsequent attention calculations and feature fusion, thereby improving the accuracy and robustness of the diagnosis and providing reliable technical support for comprehensive diagnostic analysis of medical images.
[0163] In one optional embodiment of this specification, multiple local feature vectors are input into an attention layer for attention calculation to obtain a local aggregated feature vector, including: Multiple local feature vectors are input into the attention layer, and the attention score corresponding to each local feature vector is calculated. The attention scores are normalized to obtain the attention weights corresponding to each local feature vector; Based on the attention weights, multiple local feature vectors are weighted and aggregated to obtain a local aggregated feature vector.
[0164] Attention score is a numerical metric used in the attention layer to measure the importance of local features. It quantifies the contribution or importance of each local feature vector to the final diagnosis. Specifically, attention score is obtained by performing a non-linear transformation on local feature vectors through neural network layers, representing the relative importance of local feature vectors. For example, a two-layer fully connected network can be used: the first layer maps 512-dimensional features to 128-dimensional features, and the second layer maps 128-dimensional features to 1-dimensional features to obtain the attention score.
[0165] Normalization is the process of converting a set of attention scores into a probability distribution that sums to 1. It can be used to make the weights of different local feature vectors comparable and to ensure that the scale of the weighted aggregated feature vectors is reasonable. Specifically, normalization typically uses the Softmax function, which exponentializes each attention score and divides it by the sum of the exponents of all attention scores, resulting in output attention weights that are between 0 and 1 and sum to 1.
[0166] Attention weights are weight coefficients obtained by normalizing attention scores and can be used to represent the relative importance of each local feature vector in the aggregation process. Specifically, each local feature vector corresponds to an attention weight, and the sum of all weights is 1. The higher the attention weight, the more important the local region corresponding to that feature vector is to the diagnosis, and the greater its contribution will be in the weighted aggregation.
[0167] Weighted aggregation is the process of summing multiple local feature vectors based on attention weights. It can be used to fuse multiple local feature vectors into a comprehensive local aggregated feature vector. Specifically, for each local feature vector, it is multiplied by its corresponding attention weight to obtain a weighted feature vector. Then, all weighted feature vectors are summed element-wise to obtain an aggregated vector with the same dimensions as the individual local feature vectors. Weighted aggregation can also be performed in other ways, such as weighted max pooling, which takes the maximum value after weighting in each dimension, or weighted concatenation followed by dimensionality reduction.
[0168] In practical applications, the attention score corresponding to each local feature vector can be calculated in a variety of ways.
[0169] One alternative approach is to use a multilayer perceptron, where each local feature vector of dimension D (e.g., 512-dimensional) is input into a two-layer fully connected network. The first fully connected layer maps it to an intermediate dimension (e.g., 128-dimensional), and after passing through the ReLU activation function, it is mapped to a scalar value through the second fully connected layer, thus obtaining the attention score of the local feature vector.
[0170] Another option is to use dot-product attention, which calculates the attention score by multiplying the local feature vector with a learnable global query vector. Additive attention can also be used, where the local feature vector and the query vector are passed through linear layers separately, then added together, and finally a non-linear activation is applied to obtain the score.
[0171] In addition, a self-attention mechanism can be employed, which calculates a similarity matrix between local feature vectors and uses this similarity as the basis for the attention score. The specific method for calculating the attention score can be flexibly chosen based on the actual diagnostic task requirements and model complexity.
[0172] Furthermore, the attention score can be normalized using a normalization function such as Softmax.
[0173] Specifically, for attention scores corresponding to multiple local feature vectors, the attention weights for each score can be calculated using the Softmax function, ensuring that all attention weights are between 0 and 1 and sum to 1. The Softmax function amplifies the difference between high and low scores, making the model focus more on high-scoring regions. Furthermore, Sparsemax normalization can be used, which produces a sparse weight distribution, making some attention weights exactly 0, thus achieving hard selection of local regions. Additionally, temperature-based Softmax can be employed, introducing a temperature parameter to adjust the smoothness of the weight distribution.
[0174] For example, the attention scores are normalized to obtain the attention weights, which can be expressed as:
[0175] in, For unnormalized attention weights, For the first There are several local feature vectors. Furthermore, the calculation of the attention weights can be expressed as:
[0176] in, For the first Local feature vectors The corresponding attention weights.
[0177] Once the attention weights are determined, multiple local feature vectors can be weighted and aggregated according to the attention weights to obtain a local aggregated feature vector.
[0178] Specifically, for each local feature vector among multiple local feature vectors, it is multiplied by its corresponding attention weight to obtain a weighted feature vector. Then, all weighted feature vectors are summed element-wise to obtain an aggregated vector with the same dimension as the individual local feature vectors. Alternatively, a weighted max pooling method can be used, where for each dimension of the feature vector, the maximum weighted value is taken as the value of that dimension. This method can preserve the most salient feature responses.
[0179] For example, the weighted aggregation of multiple local feature vectors to obtain the local aggregated feature vector can be represented as:
[0180] in, For locally aggregated feature vectors, For the first Local feature vectors, For the first Local feature vectors The corresponding attention weights.
[0181] In the embodiments of this specification, attention scores are calculated by inputting multiple local feature vectors into the attention layer. The attention scores are then normalized to obtain attention weights. These weights are then used to weighted aggregate the multiple local feature vectors, enabling the medical image diagnostic method to adaptively assign higher weights to key diagnostic regions, suppress interference from irrelevant regions, and obtain more representative aggregated local feature vectors. The nonlinear transformation and normalization of the attention scores ensure the rationality and accuracy of the weight allocation. The weighted aggregation process allows the model to focus on local regions crucial to diagnosis, avoiding the limitations of simple averaging or max pooling of local features. This provides high-quality local feature representations for subsequent feature fusion and diagnosis, thereby improving the accuracy and robustness of the entire diagnostic method and providing crucial technical support for comprehensive diagnostic analysis of medical images.
[0182] In one optional embodiment of this specification, the attention layer includes a first fully connected layer and a second fully connected layer; Multiple local feature vectors are input into the attention layer, and the attention score corresponding to each local feature vector is calculated, including: Multiple local feature vectors are input into the first fully connected layer for the first dimensionality reduction feature mapping to obtain the intermediate feature vector corresponding to each local feature vector. The intermediate feature vectors are input into the second fully connected layer for the second dimensionality reduction feature mapping, and the attention score corresponding to each local feature vector is obtained.
[0183] The first fully connected layer is a neural network layer in the attention layer used to reduce the dimensionality of the input features. It maps high-dimensional feature vectors to a lower-dimensional intermediate representation, providing a more compact feature representation for subsequent attention calculations. Specifically, the first fully connected layer typically receives high-dimensional (e.g., 512-dimensional) local feature vectors as input and maps them to a lower-dimensional intermediate space, such as 128-dimensional, 64-dimensional, or 256-dimensional, through a learnable weight matrix. The output of the first fully connected layer is the intermediate feature vector, which has a lower dimension than the input feature vector, reducing computational complexity while preserving key feature information.
[0184] The second fully connected layer, within the attention layer, is a neural network layer used to map intermediate feature vectors to scalar attention scores. It transforms the dimensionality-reduced feature representation into a single numerical value, indicating the importance of local features. Specifically, the second fully connected layer receives the intermediate feature vector (e.g., 128-dimensional) output from the first fully connected layer as input, and maps it to a scalar value (i.e., 1-dimensional), the attention score, through a learnable weight vector. The second fully connected layer typically does not require a non-linear activation function and can directly output a scalar value reflecting the contribution of the corresponding local feature vector to the final diagnosis.
[0185] The first dimensionality reduction feature mapping is a process of transforming local feature vectors through the first fully connected layer. It can be used to map high-dimensional feature vectors to lower-dimensional intermediate representations, preserving key feature information while reducing computational complexity. Specifically, the first dimensionality reduction feature mapping can be achieved through linear transformations of the first fully connected layer, for example, mapping a 512-dimensional local feature vector to a 128-dimensional intermediate feature vector. During the first dimensionality reduction mapping process, non-linear transformations can be performed using activation functions such as ReLU. The purpose of the first dimensionality reduction feature mapping is to reduce the feature dimension, making subsequent attention score calculation more efficient, while retaining feature information important for the diagnostic task.
[0186] The second dimensionality reduction feature mapping is a process of transforming the intermediate feature vector through the second fully connected layer. It can be used to map the intermediate feature vector to a scalar attention score, representing the importance of local features. Specifically, the second dimensionality reduction feature mapping can be achieved through a linear transformation of the second fully connected layer, for example, mapping a 128-dimensional intermediate feature vector to a 1-dimensional scalar attention score. The output of the second dimensionality reduction feature mapping can be used as an attention score for subsequent normalization processing to determine attention weights.
[0187] The intermediate feature vector is the feature representation obtained by the first fully connected layer after performing a first dimensionality reduction feature mapping on the local feature vectors. It can be used as input to the second fully connected layer to further calculate the attention score. Specifically, the dimensionality of the intermediate feature vector is lower than that of the input local feature vectors, for example, from 512 dimensions to 128 dimensions, which preserves the feature information important for the diagnostic task while reducing computational complexity.
[0188] In practical applications, the intermediate feature vectors can be obtained by performing the first dimensionality reduction, which can be achieved by inputting each local feature vector into the first fully connected layer.
[0189] For example, for each local feature vector with a dimension of 512, it can be multiplied by the weight matrix of the first fully connected layer, for example, 512×128, and further multiplied by a bias vector to obtain a linear transformation result. This result can then be input into a non-linear activation function, such as the ReLU function, to introduce non-linear expressive power and obtain an intermediate feature vector. The dimension of the intermediate feature vector can be flexibly set according to computational resources and task requirements, for example, 128-dimensional, 64-dimensional, or 256-dimensional vectors can be selected. The activation function can also be replaced with Tanh, Sigmoid, or other activation functions.
[0190] Specifically, see Figure 2 , Figure 2 This specification illustrates a schematic diagram of attention layer calculation according to an embodiment, as shown below. Figure 2 As shown.
[0191] In the input layer, input can be made. Each local image corresponds to a local feature vector. , can be represented as:
[0192] The intermediate feature vector obtained after the first dimensionality reduction feature mapping of the first fully connected layer can be represented as:
[0193] in, For the first Local feature vectors The corresponding intermediate feature vector, This indicates the first fully connected layer.
[0194] Taking the ReLU activation function as an example, the attention score obtained after nonlinear activation and second dimensionality reduction feature mapping can be expressed as:
[0195] in, For the first Local feature vectors The corresponding attention score, This indicates the second fully connected layer.
[0196] Furthermore, the calculation of attention weights can be expressed as:
[0197] in, For the first Local feature vectors The corresponding attention weights.
[0198] After determining the local eigenvectors Corresponding attention weights In this case, weighted aggregation can be further performed to obtain local aggregated feature vectors.
[0199] In this embodiment, multiple local feature vectors are input into a first fully connected layer for a first dimensionality reduction feature mapping to obtain an intermediate feature vector. This intermediate feature vector is then input into a second fully connected layer for a second dimensionality reduction feature mapping to obtain an attention score. This allows the medical image diagnostic method to accurately calculate the importance of local features through two fully connected layers, achieving efficient computation of the attention mechanism. The first dimensionality reduction feature mapping reduces feature dimensions and computational complexity while retaining key diagnostic features. The second dimensionality reduction feature mapping maps the intermediate features to a scalar attention score, providing accurate input for subsequent normalization processing. The combination of the first and second fully connected layers enables the attention mechanism to adaptively focus on diagnostically relevant key regions, suppressing interference from irrelevant regions and obtaining more representative local aggregated feature vectors. This improves the accuracy and robustness of the entire diagnostic method, providing crucial technical support for comprehensive diagnostic analysis of medical images.
[0200] In one optional embodiment of this specification, the decoding layer includes a fully connected decoding layer; The fused feature vector is input into the decoding layer for decoding to obtain the medical diagnostic results for the area to be diagnosed, including: The fused feature vectors are input into the decoding fully connected layer to calculate the prediction scores for different prediction categories; The predicted scores are normalized to obtain the predicted probability that the site to be diagnosed belongs to each predicted category. Based on predicted probabilities, the medical diagnostic result for the site to be diagnosed is determined.
[0201] The decoding fully connected layer is a neural network component in the decoding layer used to map the fused feature vector to the prediction class space. It can be used to perform linear transformations on the fused feature vector to generate raw scores for different predicted classes. Specifically, the decoding fully connected layer typically consists of one or more fully connected layers, depending on the prediction class of the diagnostic task. For example, in a binary classification task, there is one decoding fully connected layer with an output dimension of 2, which can generate raw scores for both classes.
[0202] Predicted categories are pre-defined diagnostic result types or discrete classification labels in medical imaging diagnostic models, representing potential disease states or anatomical structures at the site of diagnosis. They serve to represent the category information of the diagnostic result, and their number depends on the complexity of the diagnostic task. For example, a binary classification task may have two predicted categories, while a multi-class classification task may have multiple predicted categories. For instance, in an aortic valve diagnosis task, predicted categories could include "bicuspid aortic valve" and "tricuspid aortic valve"; in a pulmonary nodule diagnosis, predicted categories could include "benign nodule" and "malignant nodule," etc. Predicted categories are a direct representation of the medical diagnostic result and, together with predicted probabilities, constitute the diagnostic output, providing a clear classification basis for clinical decision-making.
[0203] The prediction score is the unnormalized raw numerical value output by the decoding fully connected layer, which can be used to represent the model's support or confidence level for each predicted category. Specifically, for each predicted category, the decoding fully connected layer calculates a scalar score; a higher score indicates that the model believes the input image belongs to that category more likely. The prediction score can be any real number, and its absolute value does not have a direct probabilistic interpretation, but its relative magnitude reflects the bias between categories. For example, in a binary classification task, the prediction score might be [2.5, -1.2], indicating that the score for the first category is higher than that for the second category. The calculation method of the prediction score is directly related to the weight parameters of the decoding fully connected layer and the input feature vector, and a predictive probability representation of the diagnostic result can be formed through a normalization function.
[0204] Predictive probability, a value between 0 and 1, is obtained by normalizing the predicted scores. It represents the likelihood that the site to be diagnosed belongs to each predicted category. Specifically, the sum of the predicted probabilities for all categories is 1; a higher probability value indicates a higher level of confidence in that category. For example, if the predicted probability for a bicuspid aortic valve is 0.95 and the predicted probability for a tricuspid aortic valve is 0.05, the model tends to diagnose a bicuspid aortic valve. Predictive probability is a direct basis for medical diagnosis and the final diagnostic category can be determined using preset thresholds or maximum value criteria.
[0205] In practical applications, the predicted score can be calculated by inputting the fused feature vector into the decoding fully connected layer.
[0206] One alternative approach is to perform matrix multiplication between the fused feature vector and the weight matrix of the decoding fully connected layer, and add a bias term to obtain the raw score for each predicted class. For example, for a binary classification task, the decoding fully connected layer could be a 2×1024 weight matrix, outputting two scalar scores.
[0207] Another option is to use multiple fully connected layers stacked together to form a decoder, which first undergoes a non-linear transformation through one or more hidden layers, and then obtains a score through the output layer, thereby enhancing the model's expressive power.
[0208] For example, the calculation of the prediction score can be expressed as:
[0209] in, To predict the score, To fuse feature vectors, This is the weight matrix. This is a bias term. Optionally, It can be a 2×512 weight matrix. It can be a 2D bias term.
[0210] Taking binary classification tasks (such as TAV and BAV diagnostic tasks) as an example, the prediction scores for the two predicted categories can be expressed as follows:
[0211] Furthermore, the predicted scores are normalized to obtain the predicted probabilities. This can be calculated using a normalization function, such as the Softmax function. By converting the predicted scores into a probability distribution, the Softmax function is used to exponentialize each element of the predicted score and then divide it by the sum of all exponents to obtain the probability value. Normalization ensures that the predicted probabilities are between 0 and 1 and sum to 1, facilitating interpretation and comparison. In addition, other normalization functions can be used, such as the Sigmoid function (for binary classification tasks) or normalized exponential functions, to adapt to different diagnostic task requirements.
[0212] Continuing with the previous example and using the Softmax function, the prediction probability of each prediction category obtained by normalizing the prediction scores can be expressed as:
[0213] in, To predict probabilities.
[0214] Taking binary classification tasks (such as TAV and BAV diagnostic tasks) as an example, the prediction probabilities of the two predicted categories can be expressed as follows:
[0215] Once the predicted probabilities of each prediction category are determined, the medical diagnosis of the site to be diagnosed can be determined based on these probabilities.
[0216] In practical applications, the diagnostic result can be determined by selecting the predicted category with the highest probability based on the predicted probability, or by determining the diagnostic result when the predicted probability reaches a preset threshold.
[0217] Specifically, one option is to use the maximum probability criterion, that is, to select the category with the highest predicted probability as the final diagnosis. For example, if the probability of a bicuspid aortic valve is 0.95 and the probability of a tricuspid aortic valve is 0.05, then the diagnosis is a bicuspid aortic valve.
[0218] Another option is to set a probability threshold, and only output a definite diagnosis result when the maximum probability exceeds the preset threshold (such as 0.9). Otherwise, output "uncertain" or prompt that manual review is required. This is suitable for scenarios with high requirements for diagnostic accuracy.
[0219] In addition, it can directly output the predicted probabilities of all categories as supplementary information such as confidence levels for doctors' reference, instead of automatically giving a single conclusion.
[0220] In the embodiments described in this specification, a predicted score is calculated by inputting the fused feature vector into the decoding fully connected layer. The predicted score is then normalized to obtain the predicted probability. Based on the predicted probability, the medical diagnosis result is determined, enabling the medical image diagnosis method to transform high-dimensional feature representations into clear diagnostic categories and confidence levels. Through the linear transformation of the decoding fully connected layer, a mapping from the feature space to the category space is achieved, ensuring the interpretability of the diagnostic results. The normalization process ensures the rationality of the probability distribution, facilitating clinical understanding and application. By correlating the predicted probability with the diagnostic result, the model can provide the confidence level of the diagnosis, assisting doctors in making clinical decisions, thereby improving the accuracy and practicality of the diagnosis and providing reliable technical support for comprehensive diagnostic analysis of medical images.
[0221] In an optional embodiment of this specification, after the fused feature vector is input into the decoding layer for decoding to obtain the medical diagnostic result of the site to be diagnosed, the method further includes: Determine whether the medical diagnosis results meet the preset alarm conditions; When the medical diagnosis results meet the preset alarm conditions, an alarm message is generated and output.
[0222] Preset alarm conditions are pre-defined rules or thresholds in a medical imaging diagnostic system used to trigger alarm mechanisms. They generate alarm information when a diagnostic result reaches a specific risk level, improving the timeliness and relevance of clinical diagnosis. Specifically, preset alarm conditions can be set based on the category of the diagnostic result, such as triggering an alarm when the diagnosis is "bicuspid aortic valve" because this disease requires further clinical intervention. Preset alarm conditions can also be set based on the numerical value of the predicted probability, such as triggering an alarm when the predicted probability of bicuspid aortic valve exceeds 0.9, indicating a high degree of suspicion. Alternatively, multiple factors can be combined, such as a positive diagnosis with a confidence level higher than a certain threshold.
[0223] Alarm messages are warning messages generated by medical imaging diagnostic systems when preset alarm conditions are met. They are used to alert physicians to specific diagnostic results, providing additional clinical attention when a diagnosis reaches a high-risk level, thus assisting physicians in making more accurate clinical decisions. Specifically, alarm messages may include the diagnostic category, confidence level, risk level, and recommended further investigations, and can be generated in text, visual representation, or audio alerts. For example, a text alarm message such as "Bicuspid aortic valve detected, please verify immediately" can be generated, or the lesion area can be highlighted on the imaging interface. The specific content and format of the alarm messages can be set according to the actual diagnostic task or needs to ensure the accuracy and clinical applicability of the alarm information.
[0224] In practical applications, once the medical diagnosis results for the area to be diagnosed are obtained, it can be further determined whether the medical diagnosis results meet the preset alarm conditions.
[0225] One possible approach is to make a judgment based on the category of the diagnostic result. For example, the preset alarm condition is "diagnostic result is bicuspid aortic valve". That is, when the category label output by the decoding layer is bicuspid aortic valve, the alarm condition is determined to be met.
[0226] Another option is to make a judgment based on the predicted probability. For example, the preset alarm condition is "the predicted probability of bicuspid aortic valve is greater than 0.9". That is, the predicted probability value output by the model is used as the confidence level and compared with the preset alarm threshold. If it exceeds the threshold, the condition is determined to be met.
[0227] In addition, composite conditions can be used, such as "the diagnosis result is bicuspid aortic valve and the confidence level is greater than 0.8", which means that both category and probability requirements need to be met at the same time.
[0228] Furthermore, the preset alarm conditions can be dynamically adjusted based on the patient's historical data, age, gender, and other risk factors, such as setting a more sensitive threshold for younger patients.
[0229] If the medical diagnosis results meet the preset alarm conditions, alarm information can be generated and output.
[0230] Specifically, text-based alarm messages can be generated and displayed in the image diagnosis report interface or system pop-ups, such as "Suspected bicuspid aortic valve detected, please confirm clinically." Sound alarm messages can also be triggered, emitting specific alert tones through a speaker to attract the operator's attention. Notifications can also be sent via the hospital information system to designated doctors' mobile devices or workstations, for example, via SMS, email, or in-system messages.
[0231] Optionally, for diagnostic modules integrated into a PACS system, alarm information can be directly overlaid on the medical image in the form of highlighted markers or additional annotations to guide doctors to view key areas.
[0232] The timing of alarm information output can be real-time, meaning immediately after diagnosis, or it can be batch-aggregated and output periodically. The specific content of the alarm information may include the patient's basic information, diagnosis results, confidence level, and suggested actions, enabling medical staff to take appropriate measures quickly.
[0233] In the embodiments described in this specification, after obtaining the medical diagnostic results of the site to be diagnosed, it is determined whether the medical diagnostic results meet the preset alarm conditions. If the medical diagnostic results meet the preset alarm conditions, alarm information is generated and output, enabling the medical imaging diagnostic system to proactively identify and alert to high-risk diagnostic results, thereby improving the timeliness and accuracy of clinical diagnosis. Through the flexible setting of preset alarm conditions, it can adapt to the clinical needs and risk preferences of different medical institutions, enhancing the practicality and adaptability of the system. Through the structured generation and multi-channel output of alarm information, it ensures that doctors can obtain key diagnostic information in a timely manner, assisting clinical decision-making, reducing the risk of missed diagnoses and misdiagnoses, improving medical quality and patient safety, and providing more comprehensive technical support for the comprehensive diagnostic analysis of medical images.
[0234] In one optional embodiment of this specification, the medical image diagnostic model is a pre-trained medical image diagnostic model, and the training method of the medical image diagnostic model includes: Acquire sample data, which includes global images of the sample site to be diagnosed, multiple local images of the sample site, and medical diagnostic results of the sample site to be diagnosed. The global image of the sample and multiple local images of the sample are input into the encoding layer of the medical image diagnosis model to be trained for feature encoding, thereby obtaining the predicted global feature vector and multiple predicted local feature vectors. The medical image diagnosis model to be trained also includes an attention layer, a feature fusion layer and a decoding layer. Multiple predicted local feature vectors are input into the attention layer for attention calculation to obtain the predicted local aggregated feature vector. The predicted global feature vector and the predicted local aggregated feature vector are input into the feature fusion layer for feature fusion calculation to obtain the predicted fused feature vector. The predicted fusion feature vector is input into the decoding layer for decoding to obtain the predicted medical diagnosis result of the sample's diagnostic site. Based on the sample medical diagnosis results and the predicted medical diagnosis results, determine the diagnostic loss; Based on the diagnostic loss, the medical image diagnostic model to be trained is trained to obtain a fully trained medical image diagnostic model.
[0235] Sample data is a dataset used to train medical image diagnostic models. It provides the input images and corresponding ground truth labels required for model learning. Specifically, sample data includes multiple samples, each consisting of a global image of the area to be diagnosed, multiple local images, and the medical diagnostic result. Sample data is typically derived from historical clinical image data, after annotation and preprocessing.
[0236] The anatomical region to be diagnosed in the sample data corresponds to the medical imaging diagnostic task, such as the aortic valve region or pulmonary nodule region. The anatomical region to be diagnosed is used to determine the cropping range of the global and local images of the sample, ensuring that the model learns diagnostically relevant features. The anatomical region to be diagnosed, together with the global image, local image, and medical diagnostic results, constitutes a complete training sample.
[0237] The global sample image is a 3D image patch containing the overall anatomical structure, cropped from the unenhanced scan image of the area to be diagnosed in the sample. It can be used as input for the global branch during model training. The size and preprocessing method of the global sample image are consistent with those in the inference stage; for example, it can be resampled to 128×128×128 voxels.
[0238] Sample local images are multiple smaller 3D image patches further cropped from the global sample image, focusing on key regions of the site to be diagnosed. These can be used as input for local branches during model training. The number and size of the sample local images remain consistent with the inference phase; for example, they can be seven 64×64×64 voxel image patches. These sample local images are used to train the model to extract fine-grained local lesion features. The global sample image and the corresponding sample local images together constitute a multi-scale input, enabling the model to learn both macroscopic and microscopic features.
[0239] The sample medical diagnosis results are the true diagnostic labels corresponding to the sites to be diagnosed in the samples, which can be used as supervisory signals to guide model training. Specifically, the sample medical diagnosis results can be category labels (e.g., "bicuspid aortic valve" or "tricuspid aortic valve") or more fine-grained quantitative indicators. The sample medical diagnosis results and the model's predicted medical diagnosis results are used together to calculate the diagnostic loss, in order to optimize the model parameters.
[0240] The medical image diagnostic model to be trained refers to a medical image diagnostic model whose parameters have not yet been optimized. Its network structure can also include encoding layers, attention layers, feature fusion layers, and decoding layers, but the parameter weights of each layer are in an initial state or a state to be adjusted. During the training process, the medical image diagnostic model to be trained gradually updates its parameters through backpropagation, and eventually converges into a fully trained medical image diagnostic model.
[0241] The predicted global feature vector is the feature vector output by the global encoder of the medical image diagnostic model to be trained after encoding features of the global image of the sample. It can be used as one of the inputs for subsequent feature fusion. The predicted global feature vector can be fused with the predicted local aggregated feature vector to jointly generate the predicted fused feature vector.
[0242] The predicted local feature vectors are the local encoders of the medical image diagnostic model to be trained. They are multiple feature vectors output after feature encoding of local images for each sample, and can be used as input to the attention layer. The predicted local feature vectors can be weighted and aggregated by the attention layer to obtain the predicted local aggregated feature vectors.
[0243] The predicted local aggregated feature vector is a comprehensive feature vector obtained by performing attention calculations on multiple predicted local feature vectors through an attention layer. It can be used as one of the inputs to the feature fusion layer. The dimension of the predicted local aggregated feature vector is the same as that of a single predicted local feature vector, and its value can be obtained by weighted summation, reflecting the comprehensive information of key local regions.
[0244] The predictive fusion feature vector is a feature vector obtained by fusing the predicted global feature vector and the predicted local aggregated feature vector through a feature fusion layer. It can be used as input to the decoding layer. The dimensionality of the predictive fusion feature vector depends on the fusion method. The predictive fusion feature vector integrates macroscopic and microscopic features, providing rich input information to the decoding layer.
[0245] Predicted medical diagnostic results are the diagnostic outcomes output by the medical imaging diagnostic model to be trained on the sample sites to be diagnosed. These results can be compared with the sample medical diagnostic results to calculate the loss. Specifically, predicted medical diagnostic results can include predicted categories and predicted probabilities; for example, the probability of a bicuspid aortic valve is 0.85, and the probability of a tricuspid aortic valve is 0.15. The quality of the predicted medical diagnostic results reflects the performance of the current model.
[0246] Diagnostic loss is a quantitative metric that measures the difference between the predicted results of a medical image diagnostic model under training and the actual medical diagnostic results of samples. It can be used to guide the optimization of model parameters. Specifically, diagnostic loss typically uses the standard cross-entropy loss function to calculate the difference between the predicted medical diagnostic result and the actual medical diagnostic result of samples. The smaller the loss value, the closer the model prediction is to the true label. The specific calculation method of diagnostic loss depends on the type of diagnostic task. For example, for binary classification tasks, binary cross-entropy loss can be used; for multi-class classification tasks, multi-class cross-entropy loss can be used, etc.
[0247] The trained medical image diagnostic model is an optimized model developed through the training process, capable of diagnosing diseases from new medical images. Specifically, the trained model includes optimized parameters for the encoding, attention, feature fusion, and decoding layers, enabling accurate feature extraction, aggregation, fusion, and diagnostic classification of input medical images. The trained model exhibits good generalization ability and can accurately identify the disease state of the site to be diagnosed.
[0248] In the actual training process, the calculation process and technical solutions for generating predicted global feature vectors, predicted local feature vectors, predicted local aggregated feature vectors, predicted fused feature vectors, and predicted medical diagnostic results based on sample data, global sample images, and local sample images using the medical image diagnostic model to be trained can all be found in the technical solutions in the foregoing embodiments, and will not be repeated here.
[0249] Once the predicted medical diagnosis results are generated, the diagnostic loss can be determined by combining the sample medical diagnosis results.
[0250] Specifically, for classification tasks, a cross-entropy loss function can be used, which determines the loss by calculating the cross-entropy between the predicted probability distribution and the true label encoding, such as one-hot encoding. For other types of tasks, such as regression tasks, mean squared error loss can be used. Furthermore, other regularization terms, such as L2 regularization, can be combined to prevent overfitting.
[0251] For example, for binary classification tasks, such as TAV and BAV diagnostic tasks, the diagnostic loss can be calculated using cross-entropy, which can be specifically expressed as:
[0252] in, To diagnose the loss, For the unique hot encoding of the medical diagnosis results of the sample, The predicted medical diagnosis results output by the medical image diagnosis model to be trained.
[0253] Furthermore, once the diagnostic loss is determined, the medical image diagnostic model to be trained can be trained based on this loss. Specifically, gradient descent and its variants can be used. One option is to use stochastic gradient descent, calculating the loss and gradient using a mini-batch of samples in each training iteration and updating the model parameters. Another option is to use a pre-defined optimizer such as the Adam optimizer, which adaptively adjusts the learning rate to accelerate the convergence of the model's training process.
[0254] Optionally, a two-stage strategy can be adopted for the training process of the medical image diagnostic model to be trained. Specifically, in the first stage of training, the encoding layer can be frozen, and only the attention layer and decoding layer can be trained, allowing these modules to quickly adapt to the feature representation. In the second stage of training, all model parameters can be unfrozen for end-to-end fine-tuning to further improve performance. Simultaneously, hyperparameters such as learning rate, weight decay, and early stopping can be set throughout the training process to prevent overfitting and ensure model convergence. Through multiple rounds of iterative training, the training process is completed when the diagnostic loss decreases to a preset threshold or stops decreasing, resulting in a trained medical image diagnostic model.
[0255] In this embodiment, a training dataset is constructed by acquiring sample data including global images of samples, multiple local images of samples, and medical diagnostic results of samples. The sample data is then input into the medical image diagnostic model to be trained, passing through an encoding layer, an attention layer, a feature fusion layer, and a decoding layer in sequence to obtain predicted medical diagnostic results. The diagnostic loss is determined based on the sample medical diagnostic results and the predicted medical diagnostic results, thus quantifying the difference between the model output and the true label. By training the model based on the diagnostic loss and continuously optimizing the model parameters, the model gradually learns the mapping relationship from multi-scale images to accurate diagnostic results. This enables the medical image diagnostic model to effectively extract global and local features from conventional unenhanced scan images and adaptively focus on key regions, ultimately achieving high accuracy and robustness in diagnosis.
[0256] In one optional embodiment of this specification, the medical image diagnostic model to be trained is trained based on diagnostic loss to obtain a trained medical image diagnostic model, including: Freeze the encoding layer and adjust the parameters of the attention and decoding layers based on the diagnostic loss; Unfreeze the coding layer, and based on the diagnostic loss, jointly optimize the parameters of the coding layer, attention layer, and decoding layer to obtain the trained medical image diagnostic model.
[0257] Freezing the encoding layer is a training method that sets the parameters of the encoding layer to be non-updatable during training. This maintains the stability of the encoding layer's feature extraction capability and prevents fluctuations in feature representation caused by large parameter adjustments in the early stages of training. Specifically, freezing the encoding layer fixes the weight parameters of the encoding layer, keeping it in its initial state during training. This ensures that the attention layer and decoding layer can learn based on stable feature representations, avoiding instability caused by changes in the encoding layer parameters. In particular, the training process of freezing the encoding layer can be used as the first training stage for model training.
[0258] The first training phase is the initial training stage in the medical image diagnostic model training process. It involves freezing the encoding layer and updating the parameters of only the attention and decoding layers. This phase can accelerate the model's learning of the attention mechanism and the construction of classification decisions based on feature representations. Specifically, by fixing the parameters of the encoding layer, the first training phase allows for rapid training of the attention and decoding layers to adapt to the feature representations output by the encoding layer, while maintaining the encoder's feature extraction capabilities. Typically, the learning rate in the first training phase can be set relatively high, for example, 1×10⁻⁶. -3 However, the number of training epochs can be set to be fewer, such as 40-50 epochs.
[0259] Unfreezing the encoding layer is a training method that sets the parameters of the encoding layer to be updatable during training. This allows the encoding layer to adaptively adjust its feature extraction strategy according to the needs of the diagnostic task, co-optimizing with the attention and decoding layers. Specifically, unfreezing the encoding layer removes the frozen state of its parameters, enabling them to be updated according to task requirements during training, thereby optimizing feature extraction capabilities and achieving co-optimization with the attention and decoding layers. The training process of unfreezing the encoding layer can be considered a second training phase for model training.
[0260] The second training phase, in the training process of the medical image diagnostic model, involves unfreezing the encoding layers to perform end-to-end optimization of the entire network parameters. This phase can be used to achieve collaborative optimization of global feature extraction, local feature aggregation, and diagnostic classification. Specifically, by unfreezing the parameters of all network layers, the second training phase allows for joint optimization of all parameters of the entire medical image diagnostic model, building upon the initial convergence of the attention and decoding layers, further improving model performance. Typically, the learning rate in the first training phase can be set relatively small, for example, 1×10⁻⁶. -4 However, the number of training rounds can be set to be more, such as 60-100 epochs.
[0261] The second training phase, together with the first training phase, forms a continuous training process, constituting a two-stage training strategy. This allows the model to maintain the initial feature representation capability of the encoding layer while achieving synergistic improvement in global feature extraction, local feature aggregation, and diagnostic classification through end-to-end optimization, thus avoiding the limitations of single-stage training.
[0262] Joint optimization is an optimization strategy that simultaneously adjusts the parameters of the encoding, attention, and decoding layers during the training of a medical image diagnostic model. It can be used to achieve synergistic improvements in feature extraction, feature aggregation, and diagnostic classification. Specifically, joint optimization synchronously adjusts the parameters of all network layers in an end-to-end manner, enabling the model to adaptively optimize the entire process from input to output, thereby improving diagnostic accuracy and robustness. Two-stage training involves first freezing the encoding layer for rapid adaptive training of the attention and decoding layers, and then unfreezing the encoding layer for end-to-end fine-tuning of the entire network.
[0263] Specifically, see Figure 3 , Figure 3 This specification illustrates a schematic diagram of a two-stage model training method according to an embodiment, as shown below. Figure 3 As shown, the specific training settings for the two-stage training can be as follows: First, the model is initialized, which means pre-training using MedicalNet and initializing the parameter weights.
[0264] Furthermore, the training data can be divided into a training set (70%), a validation set (15%), and a test set (15%).
[0265] In the first training phase: the encoder parameters are frozen, and the attention module and classifier are trained. Specifically: Encoder parameters: frozen (not participating in gradient updates); trainable parameters: attention fusion module + classifier; optimizer: Adam; learning rate: 1×10⁻⁶ -³; Batch size: 8; Number of training epochs: 50; Training objective: To enable the attention network and classifier to converge quickly and adapt to the encoder output features.
[0266] After completing the first phase of training, the second training phase is conducted: all parameters are unfrozen, end-to-end fine-tuning is performed, and all parameters are globally optimized to further improve model performance. Specifically: Encoder parameters: unfrozen (participating in gradient update); Trainable parameters: all network parameters; Optimizer: Adam; Learning rate: 1×10⁻⁶ -4 (Reduce the learning rate to avoid corrupting the pre-trained weights); Weight decay: 1×10 -5 (L2 regularization); Batch size: 8; Number of training epochs: 100. Data augmentation: random rotation (±15°), random flipping, intensity adjustment (±10%); The early stopping strategy can be expressed as: stop training if there is no improvement for 10 consecutive epochs on the validation set.
[0267] After the second training phase, model validation and selection can be performed, and the model with the highest AUC on the validation set can be selected, thus obtaining the completed medical image diagnosis model.
[0268] Specifically, see Figure 4 , Figure 4 This specification illustrates a prediction result graph of a medical image diagnostic model provided in one embodiment, as shown below. Figure 4 As shown.
[0269] By using a test set to test and evaluate the medical image diagnostic method and random classifier provided in this manual, the receiver operating characteristic (ROC) curve can be obtained based on the true positive rate and false positive rate. Furthermore, the performance of the model can be represented by the area under the receiver operating characteristic curve (AUC), where the value of AUC can be in the range of (0, 1), and the closer it is to 1, the stronger the model's ability to distinguish between positive and negative samples.
[0270] Based on Figure 4 It can be concluded that the medical imaging diagnostic method provided in this manual can obtain high test results, with an AUC value of approximately 0.941, which is higher than the test results obtained by models such as random classifiers.
[0271] In the embodiments described in this specification, the encoding layer is frozen in the first training stage, and only the parameters of the attention layer and the decoding layer are adjusted, enabling the model to quickly adapt to the feature representation and learn the attention mechanism. In the second training stage, the encoding layer is unfrozen, and the parameters of the encoding layer, attention layer, and decoding layer are jointly optimized, enabling the model to achieve synergistic improvement in global feature extraction, local feature aggregation, and diagnostic classification. This two-stage training strategy avoids the limitations of single-stage training, balances training efficiency and model accuracy, and allows the medical image diagnostic model to efficiently utilize the advantages of pre-trained weights while adapting to the needs of specific diagnostic tasks, improving the accuracy and robustness of diagnosis. This provides an efficient and universal technical solution for comprehensive diagnostic analysis of medical images.
[0272] In one optional embodiment of this specification, a process flow description of a medical image diagnostic method is provided. Specifically, see [link to documentation]. Figure 5 , Figure 5 This specification illustrates a flowchart of the processing procedure for a medical image diagnostic method according to one embodiment. Figure 5 As shown.
[0273] First, obtain unenhanced scan images of the area to be diagnosed, which can be obtained from low-dose CT scans for lung cancer screening, or routine plain CT scans of the chest or lungs.
[0274] Furthermore, the unenhanced scanned image can be preprocessed, including resampling, grayscale adjustment, normalization, etc., to obtain a global image, and then cropped to obtain multiple local images.
[0275] The global image and local images are input into the medical image diagnosis model. The global image can be a 192×224×160 image block, and the local images can be K 48×48×48 image blocks.
[0276] The global and local images are encoded using a global encoder and a local encoder, respectively. Both the global encoder and the local encoder can adopt a 3D ResNet18 structure. The global encoder can output a 512-dimensional global feature vector, and the local encoder can output K 512-dimensional local feature vectors.
[0277] An attention layer is used to perform attention calculations and Softmax normalization on K local feature vectors to obtain a weighted aggregated 512-dimensional local aggregated feature vector.
[0278] Feature fusion calculations are performed on the global feature vector and the local aggregated feature vector, and the resulting fusion feature vector is obtained by concatenating them.
[0279] The fused feature vector is decoded using the fully connected decoding layer included in the decoding layer, and then Softmax normalization is performed to obtain the medical diagnosis result, which may specifically include the binary classification result of BAV or TAV.
[0280] The system judges the medical diagnosis results based on preset alarm conditions. If the preset alarm conditions are met, an alarm message is generated; otherwise, no alarm is generated.
[0281] In one specific application scenario of the embodiments of this specification, the medical imaging diagnostic method can be applied to the auxiliary diagnosis of aortic valve type.
[0282] The aortic valve is a vital structure of the heart, located between the left ventricle and the ascending aorta. Based on the number of leaflets, the aortic valve can be classified into a bicuspid aortic valve (BAV) and a tricuspid aortic valve (TAV). The tricuspid aortic valve is the normal type, while the bicuspid aortic valve (BAV) is one of the most common congenital heart defects, with an incidence of approximately 1-2% in the population. Because BAV patients are prone to serious cardiovascular diseases such as aortic stenosis, aortic dissection, and aortic aneurysm, early identification of BAV is crucial for the prevention and treatment of major fatal cardiovascular diseases.
[0283] Currently, the diagnosis of aortic valve type mainly relies on methods such as echocardiography, CT angiography (CTA), and ECG-gated CT. Among these, echocardiography requires professional operation and interpretation by a qualified echocardiologist, making it highly dependent on the operator; CT angiography (CTA) requires the injection of iodine contrast agents, increasing costs and the burden on the kidneys; ECG-gated CT requires additional ECG-gated equipment and scanning protocols, resulting in high examination costs and high radiation levels.
[0284] Meanwhile, a large number of non-ECG-gated, non-contrast-enhanced chest CT images (i.e., non-contrast CT images) exist in clinical practice. These CT images can specifically include low-dose CT for lung cancer screening, routine plain CT scans of the chest or lungs, low-dose CT scans of the chest or lungs, and plain CT scans of the chest or lungs.
[0285] Non-contrast CT images are widely used in clinical practice, with tens of millions of people undergoing such examinations each year. However, for several reasons, it is difficult to accurately identify the aortic valve (BAV) on these non-contrast CT images with the naked eye. Specifically, non-contrast CT images lack contrast enhancement, resulting in low contrast between the aortic valve and surrounding tissues; there is no ECG gating, causing cardiac motion artifacts to interfere with the observation of valve structure; the slice thickness is relatively thick (usually 1-5 mm), making it difficult to clearly display the valve leaflet structure; and physicians often focus on lung lesions when interpreting CT scans, easily overlooking the aortic valve.
[0286] Therefore, a large amount of aortic valve information in non-contrast CT images is wasted, thus missing the opportunity for early detection of BAV. However, if BAV could be automatically identified and diagnosed from these existing non-contrast CT images, it would achieve universal and widespread "opportunistic screening," enabling early detection of BAV patients without increasing additional examinations, economic burden, or radiation dose, which has significant clinical value and social benefits.
[0287] The medical imaging diagnostic method provided in the embodiments of this specification can effectively solve the problems faced by current technical solutions. Specifically, in the scenario of auxiliary diagnosis of aortic valve type, the above-mentioned medical imaging method can first obtain non-enhanced scan images of the area to be diagnosed, such as retrieving previous conventional chest CT plain scan data from a hospital image archive and communication system (PACS), or lung cancer screening images directly acquired from CT equipment.
[0288] Subsequently, the CT data is preprocessed, including resampling, grayscale adjustment and normalization, to obtain a standardized global image. The left ventricular centroid is located as the target segmentation point based on a preset heart and aortic mask, and then multiple local images focusing on the aortic valve region are cropped according to a preset size.
[0289] Furthermore, the global image and multiple local images are input into the encoding layer of the medical image diagnostic model. The global encoder and local encoders extract macroscopic anatomical features and microscopic fine features, respectively, to obtain a global feature vector and multiple local feature vectors. An attention layer adaptively weights and aggregates these local feature vectors, highlighting regions related to key information such as leaflet morphology and calcification, generating a local aggregated feature vector. A feature fusion layer then concatenates or weights and fuses the global and local aggregated features to obtain a fused feature vector that integrates the overall structure and local details.
[0290] Finally, the decoding layer, through a fully connected layer and a Softmax function, outputs the predicted probability that the site to be diagnosed belongs to a bicuspid or tricuspid aortic valve, and determines the final medical diagnosis result based on the maximum probability criterion.
[0291] Compared with existing diagnostic technologies, the medical imaging diagnostic method provided in the embodiments of this specification can achieve opportunistic screening by utilizing massive amounts of existing non-contrast CT images for BAV screening without additional examinations, fully leveraging the value of existing image data; it does not increase the economic burden, requiring no additional CT scans or CTA examinations, no additional contrast agent costs, and no additional financial burden on patients; it does not increase the medical burden, requiring no additional scanning time, no additional radiology resources, and no additional physician interpretation time; it does not increase radiation dose, utilizing existing CT images without additional radiation exposure; it enables early detection of BAV, allowing for "incidental" screening of BAV during routine examinations, identifying high-risk patients early and creating conditions for subsequent monitoring and intervention; it utilizes multi-scale feature fusion, simultaneously extracting global and local features through a dual-path architecture, fully utilizing complementary information at different scales to improve classification accuracy; it improves performance through pre-trained models, using MedicalNet pre-trained weights to fully utilize the knowledge pre-trained from large-scale medical imaging data, improving performance on small samples; and it has broad applicability, suitable for various clinical scenarios such as lung cancer screening and routine chest examinations, with broad application prospects.
[0292] Corresponding to the above method embodiments, this specification also provides embodiments of medical imaging diagnostic devices. Figure 6 A schematic diagram of a medical imaging diagnostic device according to one embodiment of this specification is shown. Figure 6 As shown, the device includes: The acquisition module 602 is configured to acquire an unenhanced scan image of the site to be diagnosed, wherein the unenhanced scan image is obtained by scanning the site to be diagnosed without injecting contrast agent; The determination module 604 is configured to determine a global image and multiple local images of the site to be diagnosed based on the unenhanced scan image; The encoding module 606 is configured to input the global image and multiple local images into the encoding layer of the medical image diagnosis model for feature encoding, thereby obtaining a global feature vector and multiple local feature vectors. The medical image diagnosis model also includes an attention layer, a feature fusion layer and a decoding layer. Attention module 608 is configured to input multiple local feature vectors into the attention layer for attention calculation to obtain local aggregated feature vectors; The fusion module 610 is configured to input the global feature vector and the local aggregated feature vector into the feature fusion layer for feature fusion calculation to obtain the fused feature vector; The decoding module 612 is configured to input the fused feature vector into the decoding layer for decoding to obtain the medical diagnostic results of the site to be diagnosed.
[0293] The acquisition module is a unit in a medical imaging diagnostic device responsible for receiving or reading medical image data of the area to be diagnosed. It can be used to acquire unenhanced scan images from external devices, storage systems, or networks. Specifically, the acquisition module may include an interface with medical imaging equipment (such as computed tomography (CT) scanners or magnetic resonance imaging (MRI) scanners) for real-time acquisition of scan data; it may also include a connection to a hospital image archiving and communication system (PACS) for retrieving historical image data; and it may include a data import interface for reading image files from removable storage devices or the cloud.
[0294] The determination module is a functional unit in medical imaging diagnostic devices responsible for preprocessing unenhanced scan images and extracting the global image and multiple local images from them. It can be used to construct multi-scale input data. Specifically, the determination module can perform preprocessing operations such as resampling, grayscale adjustment, and normalization on the unenhanced scan image to obtain a standardized global image; and determine the target segmentation points based on a preset mask and anatomical structural features, and then crop multiple local images from the global image according to preset sizes.
[0295] The encoding module is a functional unit in a medical imaging diagnostic device responsible for extracting features from the global image and multiple local images. It can generate global feature vectors and multiple local feature vectors. Specifically, the encoding module contains a global encoder and local encoders that work in parallel. The global encoder downsamples and performs feature mapping on the global image, outputting a fixed-dimensional global feature vector; the local encoders independently downsample and perform feature mapping on each local image, outputting multiple local feature vectors.
[0296] The attention module is a functional unit in medical imaging diagnostic devices responsible for adaptively weighting and aggregating multiple local feature vectors. It can be used to generate local aggregated feature vectors focused on key regions. Specifically, the attention module receives multiple local feature vectors output by the encoding module, calculates an attention score for each local feature vector through an attention mechanism, obtains attention weights after normalization, and finally performs a weighted sum of the local feature vectors according to the weights to obtain the local aggregated feature vector.
[0297] The fusion module is a functional unit in medical imaging diagnostic devices responsible for fusing global feature vectors and local aggregated feature vectors. It can be used to generate fused feature vectors that integrate macroscopic and microscopic information. Specifically, the fusion module can use methods such as vector concatenation, weighted summation, gated fusion, or bilinear fusion to integrate global feature vectors and local aggregated feature vectors into a higher-dimensional or richer feature representation.
[0298] The decoding module is a functional unit in a medical imaging diagnostic device responsible for converting fused feature vectors into medical diagnostic results. It can be used to output the disease category or state of the site to be diagnosed. Specifically, the decoding module may contain one or more fully connected layers that map the fused feature vectors to prediction scores for each predicted category, then obtain the prediction probabilities through a normalization function (such as Softmax), and determine the final medical diagnostic result based on a preset criterion (such as maximum probability).
[0299] Optionally, the determining module 604 is further configured to: preprocess the unenhanced scan image to obtain a global image of the area to be diagnosed; determine target segmentation points based on a preset mask of the area to be diagnosed; and crop the global image according to a preset size based on the target segmentation points to obtain multiple local images.
[0300] Optionally, the determining module 604 is further configured to: resample the unenhanced scan image to obtain a resampled unenhanced scan image; perform grayscale adjustment on the resampled unenhanced scan image to obtain an adjusted unenhanced scan image; and perform normalization processing on the adjusted unenhanced scan image to obtain a global image of the area to be diagnosed.
[0301] Optionally, the determining module 604 is further configured to: determine the region of the region to be diagnosed based on a preset mask of the region to be diagnosed; and determine target segmentation points from the region of the region to be diagnosed based on the anatomical structure of the region to be diagnosed.
[0302] Optionally, the coding layer includes a global encoder and a local encoder; the coding module 606 is further configured to: input the global image into the global encoder, downsample it based on a preset global step size to obtain a global feature vector; input multiple local images into the local encoder respectively, downsample them based on a preset local step size to obtain multiple local feature vectors.
[0303] Optionally, the attention module 608 is further configured to: input multiple local feature vectors into the attention layer, calculate the attention score corresponding to each local feature vector; normalize the attention scores to obtain the attention weight corresponding to each local feature vector; and perform weighted aggregation on the multiple local feature vectors according to the attention weights to obtain a local aggregated feature vector.
[0304] Optionally, the attention layer includes a first fully connected layer and a second fully connected layer; the attention module 608 is further configured to: input multiple local feature vectors into the first fully connected layer for a first dimensionality reduction feature mapping to obtain an intermediate feature vector corresponding to each local feature vector; input the intermediate feature vectors into the second fully connected layer for a second dimensionality reduction feature mapping to obtain an attention score corresponding to each local feature vector.
[0305] Optionally, the decoding layer includes a fully connected decoding layer; the decoding module 612 is further configured to: input the fused feature vector into the fully connected decoding layer, calculate the prediction scores for different prediction categories; normalize the prediction scores to obtain the prediction probability that the site to be diagnosed belongs to each prediction category; and determine the medical diagnosis result of the site to be diagnosed based on the prediction probability.
[0306] Optionally, the device also includes an alarm module configured to: determine whether the medical diagnosis result meets preset alarm conditions; and generate and output alarm information if the medical diagnosis result meets the preset alarm conditions.
[0307] Optionally, the medical image diagnostic model is a pre-trained medical image diagnostic model. The device further includes a training module configured to: acquire sample data, including a global image of the sample site to be diagnosed, multiple local images of the sample site, and the sample medical diagnostic results of the sample site; input the global image and multiple local images of the sample into the encoding layer of the medical image diagnostic model to be trained for feature encoding, correspondingly obtaining a predicted global feature vector and multiple predicted local feature vectors, wherein the medical image diagnostic model to be trained further includes an attention layer, a feature fusion layer, and a decoding layer; input the multiple predicted local feature vectors into the attention layer for attention calculation, obtaining a predicted local aggregated feature vector; input the predicted global feature vector and the predicted local aggregated feature vector into the feature fusion layer for feature fusion calculation, obtaining a predicted fused feature vector; input the predicted fused feature vector into the decoding layer for decoding, obtaining the predicted medical diagnostic results of the sample site to be diagnosed; determine the diagnostic loss based on the sample medical diagnostic results and the predicted medical diagnostic results; and train the medical image diagnostic model to be trained based on the diagnostic loss, obtaining the trained medical image diagnostic model.
[0308] Optionally, the training module is further configured to: freeze the encoding layer, adjust the parameters of the attention layer and the decoding layer based on the diagnostic loss; unfreeze the encoding layer, and jointly optimize the parameters of the encoding layer, the attention layer and the decoding layer based on the diagnostic loss to obtain the trained medical image diagnostic model.
[0309] The medical imaging diagnostic device provided in the embodiments of this specification, through the collaborative work of the acquisition module, encoding module, attention module, fusion module, and decoding module, enables medical imaging diagnostic methods to efficiently extract global anatomical structure information and local lesion details from conventional unenhanced scan images. Without relying on contrast agents or other additional scanning examinations, it fully explores the potential diagnostic value of conventional scan images, avoiding the examination risks and costs associated with contrast agents or other additional examinations. Simultaneously, through the effective combination of multi-scale feature extraction and attention mechanisms, it improves the accuracy and robustness of diagnosis, providing efficient and universal technical support for comprehensive diagnostic analysis of medical images.
[0310] The above is an illustrative scheme of a medical imaging diagnostic device according to this embodiment. It should be noted that the technical solution of this medical imaging diagnostic device and the technical solution of the above-described medical imaging diagnostic method belong to the same concept. For details not described in detail in the technical solution of the medical imaging diagnostic device, please refer to the description of the technical solution of the above-described medical imaging diagnostic method.
[0311] Figure 7 A structural block diagram of a computing device 700 according to one embodiment of this specification is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.
[0312] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0313] In one embodiment of this specification, the above-described components of the computing device 700 and Figure 7 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 7The illustrated block diagram of the computing device is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed. The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.
[0314] The processor 720 executes a computer program / instruction that, when executed by the processor, implements the steps of the aforementioned medical image diagnosis method. The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the aforementioned medical image diagnosis method belong to the same concept. Details not described in detail in the technical solution of the computing device can be found in the description of the technical solution of the aforementioned medical image diagnosis method.
[0315] This specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the aforementioned medical image diagnosis method. The above is an illustrative embodiment of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the aforementioned medical image diagnosis method belong to the same concept; details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the aforementioned medical image diagnosis method.
[0316] This specification also provides a computer program product in one embodiment, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described medical image diagnosis method. The above is an illustrative scheme of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above-described medical image diagnosis method belong to the same concept; details not described in detail in the technical solution of the computer program can be found in the description of the technical solution of the above-described medical image diagnosis method.
[0317] The foregoing describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous. The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0318] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0319] The preferred embodiments disclosed above are merely illustrative of this specification. Optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A medical image diagnosis method characterized by comprising: Applications include: diagnosis of aortic valve type in the heart, including: Obtain an unenhanced scan image of the site to be diagnosed, wherein the unenhanced scan image is obtained by scanning the site to be diagnosed without injecting contrast agent, and the site to be diagnosed is the heart. Based on the unenhanced scan image, a global image and multiple local images of the site to be diagnosed are determined. The global image includes the entire heart, the ascending aorta, and surrounding tissues. The multiple local images include the leaflets, orifice, and adjacent aortic root of the aortic valve. The multiple local images are determined by locating the left ventricular mass as the target segmentation point based on a preset mask of the heart and aorta, and cropping multiple regions focused on the aortic valve according to a preset size. The global image and the multiple local images are input into the encoding layer of the medical image diagnosis model for feature encoding, thereby obtaining a global feature vector and multiple local feature vectors. The medical image diagnosis model also includes an attention layer, a feature fusion layer, and a decoding layer. The multiple local feature vectors are input into the attention layer for attention calculation to obtain a local aggregated feature vector. The process of inputting the multiple local feature vectors into the attention layer for attention calculation to obtain a local aggregated feature vector includes: inputting the multiple local feature vectors into the attention layer for attention calculation to obtain the attention weight corresponding to each local feature vector, and performing weighted aggregation on the multiple local feature vectors based on the attention weights to obtain a local aggregated feature vector. The local aggregated feature vector characterizes the degree of calcification of the aortic valve and the morphological abnormalities of the aortic valve leaflets. The global feature vector and the local aggregated feature vector are input into the feature fusion layer for feature fusion calculation to obtain the fused feature vector; The fused feature vector is input into the decoding layer for decoding, and the predicted probability of whether the site to be diagnosed belongs to a bicuspid or tricuspid aortic valve is output.
2. The method according to claim 1, characterized in that, The step of determining the global image and multiple local images of the area to be diagnosed based on the unenhanced scan image includes: The unenhanced scan image is preprocessed to obtain a global image of the area to be diagnosed; Based on the preset mask of the area to be diagnosed, the target segmentation point is determined; Based on the target segmentation point, the global image is cropped according to a preset size to obtain multiple local images.
3. The method according to claim 2, characterized in that, The step of preprocessing the unenhanced scan image to obtain a global image of the area to be diagnosed includes: The unenhanced scan image is resampled to obtain a resampled unenhanced scan image; The grayscale of the resampled unenhanced scan image is adjusted to obtain the adjusted unenhanced scan image; The adjusted, unenhanced scan image is normalized to obtain a global image of the area to be diagnosed.
4. The method according to claim 2, characterized in that, The step of determining the target segmentation point based on the preset mask of the area to be diagnosed includes: Based on the preset mask of the area to be diagnosed, the region of the area to be diagnosed is determined; Based on the anatomical structure of the site to be diagnosed, target segmentation points are determined from the region of the site to be diagnosed.
5. The method according to claim 1, characterized in that, The encoding layer includes a global encoder and a local encoder; The step of inputting the global image and the multiple local images into the encoding layer of the medical image diagnosis model for feature encoding, thereby obtaining a global feature vector and multiple local feature vectors, includes: The global image is input into the global encoder, and downsampled based on a preset global step size to obtain a global feature vector. The multiple local images are input into the local encoder, and multiple local feature vectors are obtained by downsampling based on a preset local step size.
6. The method according to any one of claims 1-5, characterized in that, The step of inputting the multiple local feature vectors into the attention layer for attention calculation to obtain a local aggregated feature vector includes: The multiple local feature vectors are input into the attention layer, and the attention score corresponding to each local feature vector is calculated. The attention scores are normalized to obtain the attention weights corresponding to each local feature vector; Based on the attention weights, the multiple local feature vectors are weighted and aggregated to obtain a local aggregated feature vector.
7. The method according to claim 6, characterized in that, The attention layer includes a first fully connected layer and a second fully connected layer; The step of inputting the multiple local feature vectors into the attention layer and calculating the attention score corresponding to each local feature vector includes: The multiple local feature vectors are input into the first fully connected layer to perform the first dimensionality reduction feature mapping, thereby obtaining the intermediate feature vector corresponding to each local feature vector. The intermediate feature vector is input into the second fully connected layer for the second dimensionality reduction feature mapping to obtain the attention score corresponding to each local feature vector.
8. The method according to claim 1, characterized in that, The decoding layer includes a fully connected decoding layer; The step of inputting the fused feature vector into the decoding layer for decoding and outputting the predicted probability that the site to be diagnosed belongs to a bicuspid or tricuspid aortic valve includes: The fused feature vector is input into the fully connected decoding layer to calculate the prediction scores for different prediction categories; The predicted scores are normalized to obtain the predicted probability that the site to be diagnosed belongs to a bicuspid or tricuspid aortic valve.
9. The method according to claim 1, characterized in that, After inputting the fused feature vector into the decoding layer for decoding to obtain the predicted probability that the site to be diagnosed belongs to a bicuspid or tricuspid aortic valve, the method further includes: Determine whether the predicted probability meets the preset alarm conditions; When the predicted probability meets the preset alarm conditions, an alarm message is generated and output.
10. The method according to claim 1, characterized in that, The medical image diagnostic model is a pre-trained medical image diagnostic model, and the training method of the medical image diagnostic model includes: Acquire sample data, which includes a global image of the sample site to be diagnosed, multiple local images of the sample, and the sample prediction probability of the sample site to be diagnosed. The global image of the sample and the multiple local images of the sample are input into the encoding layer of the medical image diagnosis model to be trained for feature encoding, thereby obtaining a predicted global feature vector and multiple predicted local feature vectors. The medical image diagnosis model to be trained also includes an attention layer, a feature fusion layer and a decoding layer. The multiple predicted local feature vectors are input into the attention layer for attention calculation to obtain the predicted local aggregated feature vector; The predicted global feature vector and the predicted local aggregated feature vector are input into the feature fusion layer for feature fusion calculation to obtain the predicted fused feature vector. The predicted fusion feature vector is input into the decoding layer for decoding to obtain the predicted probability of the part of the sample to be diagnosed. The diagnostic loss is determined based on the sample prediction probability and the predicted probability. Based on the diagnostic loss, the medical image diagnostic model to be trained is trained to obtain a trained medical image diagnostic model.
11. The method according to claim 10, characterized in that, The step of training the medical image diagnostic model to be trained based on the diagnostic loss to obtain the trained medical image diagnostic model includes: Freeze the encoding layer, and adjust the parameters of the attention layer and the decoding layer based on the diagnostic loss; Unfreeze the encoding layer, and based on the diagnostic loss, jointly optimize the parameters of the encoding layer, the attention layer, and the decoding layer to obtain the trained medical image diagnostic model.
12. A medical imaging diagnostic device, characterized in that, Applications include: diagnosis of aortic valve type in the heart, including: The acquisition module is configured to acquire an unenhanced scan image of the site to be diagnosed, wherein the unenhanced scan image is obtained by scanning the site to be diagnosed without injecting contrast agent, and the site to be diagnosed is the heart. The determination module is configured to determine a global image and multiple local images of the site to be diagnosed based on the unenhanced scan image. The global image includes the entire heart, the ascending aorta and surrounding tissues. The multiple local images include the leaflets, orifice and adjacent aortic root of the aortic valve. The multiple local images are determined by locating the left ventricular centroid as the target segmentation point based on a preset mask of the heart and aorta, and cropping multiple regions focused on the aortic valve according to a preset size. The encoding module is configured to input the global image and the multiple local images into the encoding layer of the medical image diagnosis model for feature encoding, thereby obtaining a global feature vector and multiple local feature vectors. The medical image diagnosis model further includes an attention layer, a feature fusion layer, and a decoding layer. An attention module is configured to input the plurality of local feature vectors into the attention layer for attention calculation to obtain a local aggregated feature vector. The attention module is further configured to: input the plurality of local feature vectors into the attention layer for attention calculation to obtain an attention weight corresponding to each local feature vector, and perform weighted aggregation on the plurality of local feature vectors based on the attention weights to obtain a local aggregated feature vector. The local aggregated feature vector characterizes the degree of calcification of the aortic valve and the morphological abnormalities of the aortic valve leaflets. The fusion module is configured to input the global feature vector and the local aggregated feature vector into the feature fusion layer to perform feature fusion calculation and obtain a fused feature vector. The decoding module is configured to input the fused feature vector into the decoding layer for decoding and output the predicted probability that the site to be diagnosed belongs to a bicuspid aortic valve or a tricuspid aortic valve.
13. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the medical image diagnosis method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, It stores a computer program / instruction that, when executed by a processor, implements the steps of the medical image diagnosis method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the steps of the medical image diagnosis method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Oral and maxillofacial surgery image recognition and diagnosis method and system based on deep learning
CN120280132A