Method for automatically discriminating and enhancing phases in CT scanning based on attention mechanism of multi-instance learning

By combining multi-instance learning and attention mechanism, the problems of irregular labeling, large computational complexity and weak interpretability in automatic identification of enhanced phases in CT scans are solved, and efficient and interpretable enhanced phase classification is achieved, which improves the classification accuracy and processing speed of CT scans and supports large-scale clinical applications.

CN120673119AActive Publication Date: 2025-09-19深圳市一苣技术有限公司

Patent Information

Application Number
CN202510657163.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-19
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

Existing CT scans have problems with automatic enhanced phase identification, such as irregular labeling, large computational complexity, weak model interpretability, and category imbalance, resulting in low automation, low efficiency, and difficulty in meeting large-scale clinical needs.

Method used

By adopting an attention mechanism based on multi-instance learning, combined with a two-dimensional convolutional neural network and multiple slice sampling strategies, efficient and interpretable enhanced phase classification is achieved through a multi-instance attention aggregation mechanism and saliency visualization.

Benefits of technology

It improves the classification accuracy and processing speed of CT scans, enhances the interpretability and flexibility of the model, and can maintain high performance under complex data structures, supporting large-scale clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673119A_ABST
    Figure CN120673119A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical imaging, and discloses a method for automatically discriminating and enhancing phases in CT scanning by an attention mechanism based on multi-instance learning, and the method comprises the steps: extracting a CT scanning image and DICOM information from a data set, and unifying the image size after HU conversion, truncation and linear zooming; obtaining a fixed number of slices by using direct or indirect random sampling and a cubic interpolation strategy; then, a pre-trained model is adopted to extract slice features, scanning level feature vectors are formed through attention aggregation, and finally prediction probabilities of four types of phases are output through a full connection layer and a Softmax module; the model is trained by adopting cross entropy loss and an Adam optimizer, the performance is evaluated through indexes such as precision, recall rate and F1-score, and a thermodynamic diagram is generated by combining gradient back propagation, so that visual visualization of a result is realized, and clinical judgment is assisted. According to the method, the CT scanning multi-phase can be accurately, rapidly and interpretably classified in a real clinical environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical imaging technology, and in particular to a method for automatically distinguishing enhanced phases in CT scans based on an attention mechanism of multi-instance learning. Background Art

[0002] In the field of medical imaging, especially in the automatic identification of enhancement phase in contrast-enhanced CT (CE-CT) scans, there are many problems that need to be solved.

[0003] With the widespread use of CT and contrast agents in medical imaging, automatic discrimination of different enhancement phases, such as the non-contrast phase, arterial phase, and venous phase, is crucial for subsequent tasks, including organ segmentation, radiomics feature extraction and analysis, and disease prognosis prediction. However, current CT image metadata often lacks standardized annotation of enhancement phases. For example, DICOM tags are incomplete or non-standard, and manual input can be erroneous, resulting in a low level of automation. Furthermore, distinguishing between different enhancement phases presents challenges such as image intensity variations, temporal characteristics, and the sheer size of the data.

[0004] Traditional methods have many shortcomings. First, the labeling is not standardized. Phase labels in DICOM are often entered manually, which is prone to errors or omissions, and thus requires a lot of manual rechecking and staging. Secondly, the three-dimensional model is computationally intensive. Although early 3D convolutional networks such as C3D and 3D-ResNet can process the entire 3D volume data at one time, they have a large number of parameters and slow inference speed, making it difficult to meet large-scale clinical needs. Furthermore, although deep learning models such as CNN have a certain accuracy in classification results, they lack interpretability and cannot intuitively present the key anatomical areas that the model focuses on, resulting in a lack of physician trust and traceability of results in clinical applications. Finally, the problem of class imbalance is prominent. In phase classification, some phases, such as the non-contrast phase, have too much data, while the "other" phase data is relatively small. If not processed, the model will easily be biased towards the mainstream category.

[0005] Therefore, an innovative and efficient method is urgently needed to solve the above problems. Summary of the Invention

[0006] In the current state of the art, automated multi-phase classification methods for computed tomography (CT) scans face several challenges: 1. Lack of a unified and efficient phase identification mechanism: Current methods are inefficient when processing complete three-dimensional datasets or rely on manual phase labeling; in large-scale applications, these methods are not flexible or cost-effective enough. 2. Insufficient adaptability to data imbalance and incompleteness: In actual hospital scanning processes, some phases may be missing or insufficient in number, making it difficult for traditional methods to adaptively select feature slices, resulting in limited classification accuracy and generalization capabilities. 3. Weak model interpretability: Radiologists often need to understand how the model makes phase decisions, for example, which blood vessels or organ regions the model focuses on. Without interpretability tools, the model will find it difficult to gain clinical trust. 4. Impact on downstream tasks: Accurate and efficient phase classification is the basis for organ segmentation, radiomics feature extraction, and disease prognosis analysis, but existing technologies have not yet fully integrated phase classification with the needs of downstream tasks.

[0007] In view of this, the present invention proposes a method for automatically distinguishing enhanced phases in CT scans based on the attention mechanism of multi-instance learning. The method aims to ensure efficient classification of enhanced phases in computed tomography (CT) while taking into account multi-instance learning, sampling strategy and interpretability to cope with the diverse scanning conditions and data imbalance problems in actual clinical practice, thereby laying a solid foundation for subsequent applications.

[0008] The present invention proposes a method for automatically distinguishing enhanced phases in CT scans based on a multi-instance learning attention mechanism, comprising: Acquire a file containing several CT scan records and the corresponding slice images of the CT scan records to form a VinDr-Multiphase dataset; extract the metadata of the DICOM format file and convert the pixel values ​​into HU units; at the same time, uniformly scale the images to a predetermined size and convert them into PNG or Numpy format; All scan data are divided into training set, validation set and test set; direct random sampling, indirect random sampling or interpolation strategy is used to obtain a fixed number of slices for each scan data; Each slice image is input into a pre-trained convolutional neural network to extract convolutional features or spatiotemporal features. Using the multi-instance attention aggregation mechanism, the features of each slice in the same scan are weighted summed to obtain the scan-level features. The model is trained and evaluated. After the model training and evaluation are completed, the model is used to predict the test set scan data to obtain the prediction results. The pixel-level contribution is generated through gradient backpropagation, and SaliencyMaps or heat maps superimposed with the original image are constructed to finally output the results.

[0009] Preferably, the metadata includes Pixel Data, Rescale Slope / Intercept, window width and / or window level.

[0010] Preferably, when all scan data are divided into a training set, a validation set, and a test set, the ratios of the training set, the validation set, and the test set are 70%, 15%, and 15%.

[0011] Preferably, the direct random sampling is to randomly select 64 slices; The indirect random sampling is to randomly select 30% of the slices and then cut or fill them to 64; The interpolation strategy is to generate intermediate slices using a cubic interpolation method.

[0012] Preferably, the convolutional neural network is a two-dimensional convolutional neural network.

[0013] Preferably, when using a multi-instance attention aggregation mechanism to perform weighted summation of slice features within the same scan to obtain scan-level features, the method includes: All slice features of the same scan are denoted as {h1, h2, ..., h i}, introduce attention weights {a1, ...,a i}, the scan-level feature H is obtained by weighted summation, and the calculation formula is H = ∑aᵢ·hᵢ.

[0014] Preferably, training and evaluating the model includes: using a cross entropy or weighted binary cross entropy loss function to train the model.

[0015] Preferably, the training and evaluation of the model includes: The Adam optimizer is used, and the training effect is evaluated by precision, recall, F1-score and AUC. A confusion matrix is ​​generated to analyze the classification performance.

[0016] Compared with the prior art, the present invention has the following beneficial effects: Efficiency: This method combines 2D CNN slice-level feature extraction with a multiple-instance learning (MIL) mechanism to effectively avoid the computational redundancy and inference bottlenecks faced by traditional 3D convolutional models when processing complete sets of CT volume data. Through optimized structure and lightweight model design, this method significantly reduces model parameters and GPU memory usage while ensuring classification accuracy, significantly improving processing speed and resource efficiency in large-scale clinical applications.

[0017] Flexibility: This method supports multiple slice sampling strategies, including direct random sampling, indirect proportional sampling, and interpolation-enhanced sampling. It automatically adjusts the input method based on the CT scan length, slice density, and phase completeness. It can handle practical issues such as missing and unevenly distributed enhancement phases, while maintaining high performance even with complex data structures.

[0018] Interpretability: By introducing an attention mechanism and saliency maps, the model can automatically identify and highlight slices or regions that are critical to classification during inference, such as enhanced features of the cardiac aorta, hepatic hilum, and renal artery. This allows doctors to intuitively observe the model's focus, thereby improving model verifiability and clinical trust, and promoting the transformation of AI from "black box" to "white box" applications.

[0019] Accuracy and Robustness: Through a multi-instance attention aggregation mechanism, this method effectively suppresses the interference of noisy slices and highlights the "high-information" slices that are critical for classification. The model demonstrates excellent generalization and classification robustness, especially when samples in other transitional categories are scarce and labels are unstable. The measured F1-score exceeds 0.98 on public datasets such as VinDr-Multiphase, demonstrating its stability and practicality in real-world medical scenarios.

[0020] Scalability for subsequent medical applications: Accurate enhanced phase classification results provide accurate and reliable input for downstream tasks such as organ segmentation, radiomics feature extraction and analysis, and disease prognosis prediction.

[0021] In summary, this invention provides an innovative and efficient method for automatically identifying the enhancement phase in contrast-enhanced CT scans, which has significant clinical value and practical application prospects. It not only improves the accuracy of medical image analysis but also provides physicians with more accurate diagnostic evidence, thereby enhancing diagnostic efficiency and the specificity and effectiveness of treatment plans. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 Schematic diagram of the data conversion process of the present invention; Figure 2 This is a schematic diagram of the working principle of multiple instance learning (MIL); Figure 3To verify the interpretability of the public dataset VinDr-Multiphase, we generated a Saliency Map visualization example diagram, which uses CT slices of the heart, such as the non-enhanced, venous phase, and arterial phase, and their heat map overlays to show the focus intensity of key parts; Figure 4 To verify the interpretability of the public dataset VinDr-Multiphase, we generated a Saliency Map visualization example diagram, which uses CT slices of the liver, such as non-enhanced, venous phase, and arterial phase, and their heat map overlays to show the focus intensity of key areas. Figure 5 To verify the interpretability of the public dataset VinDr-Multiphase, we generated a Saliency Map visualization example diagram, which uses CT slices of the kidney, such as non-enhanced, venous phase, and arterial phase, and their heat map overlays to show the focus intensity of key areas. Figure 6 To further verify interpretability on a new dataset, HCC-TACE-Seg, that the model had never seen before, we generated a Saliency Map visualization example diagram, overlaying CT slices of the heart's arterial and venous phases and their heat maps to show the intensity of attention on key areas. Figure 7 To further verify interpretability on a new dataset, HCC-TACE-Seg, that the model had never seen before, we generated a Saliency Map visualization example diagram, overlaying CT slices of the liver, such as the arterial and venous phases, with their heatmaps to show the intensity of attention in key areas. Figure 8 To further verify the interpretability of the new dataset HCC-TACE-Seg that the model has never seen before, a Saliency Map visualization example diagram was generated, which overlays CT slices of the kidney's arterial phase, venous phase, and other CT slices and their heat maps to show the focus intensity of key parts. DETAILED DESCRIPTION

[0023] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0024] This invention proposes a contrast-enhanced CT phase classification method that integrates two-dimensional / three-dimensional convolutional neural networks (2D / 3D CNN), multiple instance learning (MIL), and multiple random sampling and interpolation strategies. The specific implementation process includes the following steps: 1. Data Preprocessing DICOM parsing: Read the original CT scan DICOM file, extract key metadata (Pixel Data, Rescale Slope / Intercept, Window Width WL, Window Level WW, etc.), and calculate the HU value of each pixel.

[0025] Image conversion: Convert CT slices to PNG or Numpy format; PNG is suitable for 2D convolutional neural networks, and Numpy can preserve 3D volume data for use in 3D convolutional neural networks. Normalize the slice size (e.g., 128×128) according to actual needs. Figure 1 Schematic diagram of the data conversion process of the present invention.

[0026] 2. Data Selection Strategy Direct Random Sampling: Randomly select several slices (e.g., 64) from a scan to reduce the computational burden.

[0027] Indirect Random Sampling: First randomly extract some slices from the scan (a certain percentage, such as 30%), and then truncate or expand them to a uniform number.

[0028] Data Interpolation: Linear or cubic interpolation is performed on scans with insufficient number of slices or inconsistent resolution to enhance data continuity and facilitate the model to capture phase feature changes.

[0029] 3. Multiple Instance Learning and Attention Mechanism Treat a scan as a bag and each slice as an instance in the bag.

[0030] Feature extraction and aggregation: A lightweight 2D CNN architecture (such as EfficientNet and ResNet) is used to extract local features from each slice. These features are then aggregated into scan-level representations using two strategies: attention-based pooling, which introduces a learnable attention mechanism to enable the model to focus on key, more discriminative slices; and mean pooling, which averages the features across all slices to enhance model stability and reduce the impact of local noise.

[0031] Enhanced interpretability: By generating saliency maps, the key anatomical structures that the model focuses on are intuitively displayed, helping doctors better understand the model's decision-making process.

[0032] IV. Model Structure and Training Multi-structure model training and optimization: Model structure selection: 2D CNN: Based on pre-trained models such as ResNet and EfficientNet, it extracts slice features through convolution operations. It is suitable for fine-grained processing of slices, has fast inference speed, and is easy to deploy on a large scale and for real-time applications.

[0033] 3D CNN: Designed for volumetric data, it can learn changes in both spatial and temporal dimensions, making it suitable for applications requiring extremely high precision. Although it can process and model both spatial and temporal features simultaneously, its high parameter count and computational complexity mean it is not used in this project.

[0034] Training strategy: Loss function: Cross entropy or weighted cross entropy is used to solve the problem of class imbalance.

[0035] Optimizer: Use the Adam optimizer in combination with Cosine Annealing or Warm-up strategy to improve training efficiency.

[0036] Evaluation metrics: A multi-metric evaluation system includes precision, recall, and F1-score to ensure a comprehensive evaluation of model performance.

[0037] The above method of the present invention has the following advantages: First, it significantly improves the accuracy of phase classification. By cleverly integrating the strengths of two-dimensional and three-dimensional convolutional neural networks, this invention can more comprehensively and meticulously capture the various phase features in CT scan images. This comprehensive feature extraction capability enables the system to more accurately and meticulously classify different phases, significantly improving the accuracy of phase classification.

[0038] Second, the model's generalization capabilities are significantly enhanced. This invention combines multi-instance learning with an attention mechanism, enabling the model to focus more closely on key slice regions when processing CT scan data, effectively reducing the impact of local noise on classification results. This design ensures that the model maintains stable and reliable performance across a variety of patients and scanning conditions.

[0039] Third, computational efficiency is significantly improved. By introducing two strategies, direct random sampling and indirect random sampling, this paper significantly reduces the computational burden while ensuring that model performance is not affected. This optimization enables the model to process large amounts of CT scan data at a faster rate, improving overall computational efficiency.

[0040] Fourth, the model's interpretability is significantly enhanced. By generating a saliency map, doctors can visually visualize the key anatomical structures that the model focuses on when making decisions. This visualization not only helps doctors better understand the model's decision-making process but also helps improve diagnostic accuracy and reliability, making the model's clinical application more transparent and trustworthy.

[0041] The technical solution of the present invention is described below through specific embodiments. Example

[0042] 1. Environment and Data Preparation This example uses the VinDr-Multiphase Dataset, which contains 1,188 CT scans and 392,560 slice images. Each scan is labeled with a phase, including non-contrast enhancement, arterial phase, venous phase, and other phases.

[0043] The experimental environment is an Ubuntu server equipped with an NVIDIA RTX 4070 GPU. Python 3.8.10 and Torch 2.3.0+cu121 are used for model training and inference.

[0044] 2. Data Preprocessing Extract Pixel Data, Rescale Slope / Intercept, and Window Width / Window Level parameters from DICOM files and convert pixel values ​​into HU units; Appropriate truncation and linear scaling (window level WL = 40, window width WW = 400) were performed on the HU values ​​to enhance the contrast of the tissue of interest; The data is uniformly scaled to 128×128 PNG or Numpy format, and all scanned data are divided into training set, validation set, and test set in the ratio of 70%, 15%, and 15%.

[0045] 3. Sampling Strategy Selection Direct random sampling strategy: 64 slices were randomly selected from each scan; Indirect random sampling strategy: first randomly select 30% of the slices, and then crop or fill them to 64; Interpolation strategy: For scans with a small number of slices or requiring continuous timing information, a cubic interpolation method is used to generate intermediate slices.

[0046] Experimental results show that in the application scenario of attention aggregation, the combination of "indirect random sampling 30%" and cross-entropy loss function has the best effect; and in the "average aggregation" strategy, the "interpolation" method can also achieve higher prediction accuracy.

[0047] 4. Network Model and Attention Mechanism Two-dimensional Convolutional Neural Network (2D CNN): We utilize the EfficientNet-B2 model pre-trained on ImageNet, but remove its final classification layer to retain its feature extraction capabilities. For each slice, a series of convolution, activation, and pooling operations are performed to extract a 128-dimensional deep feature vector.

[0048] Multi-instance attention aggregation mechanism: All slice features of the same scan are recorded as {h1, h2, ..., h i}, introduce attention weights {a1, ..., ai}, and obtain the scan-level feature representation through weighted summation: H = ∑aᵢ·hᵢ. The attention module is a learnable structure that can dynamically highlight key anatomical regions (such as the abdominal aorta and hepatic hilum), taking into account both robustness and interpretability.

[0049] V. Training and Evaluation Loss function: Select Cross Entropy (CE) or Weighted Binary Cross Entropy (WBCE) to address the problem of class imbalance; experiments have found that CE usually achieves a higher F1-score.

[0050] Hyperparameters: Use the Adam optimizer, set the learning rate to 0.001, and adjust the batch size according to the graphics memory availability.

[0051] Evaluation metrics: Precision, recall, F1-score (weighted), and AUC are used. Confusion matrices and heat maps are also plotted to assess the interpretability of the model.

[0052] 6. Specific Implementation Process Scan import and format conversion: The system automatically reads CT scans and completes DICOM parsing, image standardization and format conversion (see Appendix Figure 1 ).

[0053] Sampling and interpolation processing: Select direct / indirect sampling or interpolation strategy according to the number of scan slices to form a consistent input sequence (see Appendix Figure 2 ).

[0054] Slice-level feature extraction (2D CNN): Each slice is fed into a 2D convolutional network to extract high-dimensional semantic features (see Appendix Figure 2 ).

[0055] Feature aggregation and scan-level classification (MIL + Attention): Aggregate all slice features and output the full scan phase classification probability through the attention mechanism (see figure attached) Figure 4 ), Softmax outputs 4 types of phase. (See Appendix Figure 2 ).

[0056] Result output and significance visualization: The model output is visualized to generate a Saliency Map, which is superimposed on the original slice to assist clinical interpretation. (See Appendix Figure 3-8 ).

[0057] VII. Summary and Applicability This example details the application of the method in a real-world multiphase recognition task on abdominal CT scans, demonstrating its high operability, reproducibility, and portability. Experimental results demonstrate that the method achieves excellent performance with an F1-score exceeding 0.98 on the VinDr-Multiphase dataset.

[0058] The above description is merely one preferred embodiment of the present invention. Within the scope of the present invention, technical alternatives (such as replacement of the CNN structure, adjustment of the sampling ratio, and expansion of the loss function) should also be considered equivalent embodiments and are all covered by the scope of protection of the present invention. The method and system provided by the present invention can be directly integrated into a medical imaging AI platform to achieve rapid, accurate, and interpretable classification of CT scan enhancement phases, providing a solid foundation for automatic segmentation, extraction and analysis of radiomics features, and prediction of disease prognosis.

[0059] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0061] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0062] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for automatically distinguishing enhanced phases in CT scans based on a multi-instance learning attention mechanism, characterized in that: include: Acquire several CT scan records and their corresponding slice images to form a VinDr-Multiphase dataset; extract metadata from DICOM format files and convert pixel values ​​to HU units; uniformly scale the images to a predetermined size and convert them to PNG or Numpy format; All scan data are divided into training set, validation set and test set; direct random sampling, indirect random sampling or interpolation strategy is used to obtain a fixed number of slices for each scan data; Input each slice image into the pre-trained convolutional neural network model to extract convolution features or spatiotemporal features; Using the multi-instance attention aggregation mechanism, the features of each slice in the same scan are weighted summed to obtain the scan-level features; The model is trained and evaluated. After the model training and evaluation are completed, the model is used to predict the test set scan data to obtain the prediction results. The pixel-level contribution is generated through gradient backpropagation, and SaliencyMaps or heat maps superimposed with the original image are constructed to finally output the results.

2. The method for automatically distinguishing enhanced phases in CT scans using an attention mechanism based on multiple instance learning according to claim 1, characterized in that: The metadata includes Pixel Data, Rescale Slope / Intercept, window width and / or window level.

3. The method for automatically distinguishing enhanced phases in CT scans using an attention mechanism based on multiple instance learning according to claim 1, characterized in that: When all scan data are divided into training set, validation set and test set, the ratios of training set, validation set and test set are 70%, 15% and 15% respectively.

4. The method for automatically distinguishing enhanced phases in CT scans using an attention mechanism based on multiple instance learning according to claim 1, characterized in that: The direct random sampling is to randomly select 64 slices; The indirect random sampling is to randomly select 30% of the slices and then cut or fill them to 64; The interpolation strategy is to generate intermediate slices using a cubic interpolation method.

5. The method for automatically distinguishing enhanced phases in CT scans using an attention mechanism based on multiple instance learning according to claim 1, characterized in that: The convolutional neural network is a two-dimensional convolutional neural network.

6. The method for automatically distinguishing enhanced phases in CT scans using an attention mechanism based on multiple instance learning according to claim 1, characterized in that: Using the multi-instance attention aggregation mechanism, the scan-level features are obtained by weighted summing the features of each slice in the same scan, including: All slice features of the same scan are denoted as {h1, h2, ..., h i }, introduce attention weights {a1, ..., a i }, the scan-level feature H is obtained by weighted summation, and the calculation formula is H = ∑aᵢ·hᵢ.

7. The method for automatically distinguishing enhanced phases in CT scans using an attention mechanism based on multiple instance learning according to claim 1, characterized in that: When training and evaluating the model, including: using cross entropy or weighted binary cross entropy loss function to train the model.

8. The method for automatically distinguishing enhanced phases in CT scans using an attention mechanism based on multiple instance learning according to claim 1, characterized in that: When training and evaluating a model, this includes: The Adam optimizer is used, and the training effect is evaluated by precision, recall, F1-score and AUC. A confusion matrix is ​​generated to analyze the classification performance.

Citation Information

Patent Citations

  • Automatic optimal phase recognition method for heart CT imaging

    CN110969633A

  • Medical image tumor localization method based on attention mechanism and multi-instance learning

    CN115602295A

  • Method for realizing precise diagnosis of three-level lymphatic structure of glioblastoma patient

    CN118675172A

  • Identification of a contrast phase depicted in a medical image

    US20220318567A1

  • Identifying calcification lesions in contrast enhanced images

    US20220405915A1

Cited By

  • Medical image sequence positioning method and system

    CN121544625A

  • Intelligent phase identification method and system for CT scanning based on organ segmentation

    CN122023393A