YOLO deep learning-based automatic grading system for grade of medial temporal lobe atrophy
Through the automatic scoring system of medial temporal lobe atrophy grade based on YOLO deep learning, the problem of inconsistent section selection and insufficient scoring algorithm for medial temporal lobe atrophy assessment in the prior art is solved, and efficient and accurate MTA scores are achieved, which are suitable for clinical high-throughput imaging analysis needs.
Patent Information
- Application Number
- CN202510077909.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The prior art has problems with inconsistent section selection, insufficient scoring algorithm robustness and lack of efficient clinical integration systems in the early diagnosis of Alzheimer's disease, especially in the evaluation of medial temporal atrophy (MTA).
The automatic scoring system of medial temporal lobe atrophy grade based on YOLO deep learning is adopted, including coronal slice screening module, improved YOLOv8 model building module, optimal coronal slice automatic selection module, and medial temporal lobe atrophy detection and automatic scoring module. The detection efficiency and accuracy of the model are improved through the EfficientViT structure and AdaptiveSlide Loss function.
The full process automation of MRI data import to automatic screening of optimal coronary slices, hippocampus detection of MTA score output is realized, which improves the accuracy and consistency of scores, reduces calculation overhead, and improves clinical work efficiency.
Smart Images

Figure CN120107158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of clinical diagnosis technology, and in particular to an automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning. Background Art
[0002] Alzheimer's disease (AD) is a common neurodegenerative disease, and its clinical features are mainly manifested as cognitive impairment, neuropsychiatric behavioral abnormalities, and decreased ability to live daily. With the aging of society, the prevalence of AD has increased year by year, which has placed a heavy burden on patients, their families, and society. Since there is currently no effective drug therapy to reverse or cure AD, early detection and accurate diagnosis of AD are particularly important. Assessment of brain structural changes based on medical imaging, such as magnetic resonance imaging (MRI), is one of the commonly used diagnostic and disease progression monitoring methods in clinical practice. Among them, quantitative or semi-quantitative assessment of medial temporal atrophy (MTA) plays an important role in the early diagnosis of AD.
[0003] In clinical diagnosis or scientific research applications, commonly used MTA evaluation methods include visual inspection and visual scoring based on standard scales (such as the MTA scoring standard proposed by Scheltens et al.). Although this type of scoring method based on manual observation is simple to operate and low in cost, it relies on the subjective experience of the evaluator, and often leads to inconsistent scoring between different physicians or evaluators (i.e., large subjective differences) and long time consumption. In addition, when faced with a large amount of MRI data, manual inspection of each image is inefficient and difficult to meet the needs of clinical high-throughput imaging analysis.
[0004] In recent years, with the rapid development of deep learning and computer vision technology, more and more studies have begun to explore the use of machine learning models for automatic segmentation, detection and disease diagnosis and prediction of medical images. Some literature and patents have proposed the application of convolutional neural networks (CNN) or other deep learning models to MRI structure segmentation and auxiliary diagnosis of mild cognitive impairment (MCI) or AD. However, in actual clinical scenarios, most automated methods still have the following shortcomings or difficulties:
[0005] 1) Inconsistent slice selection: MTA scoring usually requires the selection of the best slice in the coronal position for observation. Manual selection of the best slice is often time-consuming and susceptible to subjective influence. Existing automated methods may only target specific sequences or may not be accurate in slice selection, resulting in a decrease in the accuracy of subsequent scoring.
[0006] 2) The scoring algorithm is not robust enough: For MTA classification (0 to 4 levels), especially when the boundaries between adjacent levels (such as 1 and 2, 2 and 3) are blurred, it is difficult for the algorithm to accurately identify them; and when faced with unevenly distributed data sets, model training often suffers from overfitting or insensitivity to minority category recognition.
[0007] 3) Lack of efficient clinical integration system: Some automated scoring studies are only implemented in offline environments, and lack ease of use, visualization, and integration with clinical workflows, making it difficult to promote on a large scale. Summary of the invention
[0008] The purpose of the present invention is to provide an automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning, comprising: a coronal slice screening module, an improved YOLOv8 model building module, an optimal coronal slice automatic selection module, and a medial temporal lobe atrophy detection and automatic scoring module.
[0009] The coronal slice screening module screens out coronal slices from the brain MRI scan data input into the system and constructs a coronal slice set.
[0010] The improved YOLOv8 model construction module is used to construct two improved YOLOv8 models, and train them into a YOLOv8_slice model and a YOLOv8_MTA model respectively.
[0011] The optimal coronal slice automatic selection module inputs the coronal slices in the coronal slice set into the YOLOv8_slice model to select the optimal coronal slice.
[0012] The medial temporal lobe atrophy grade automatic scoring module inputs the best coronal slice into the YOLOv8_MTA model to determine the medial temporal lobe atrophy grade.
[0013] Furthermore, the structure of the brain MRI scan data input into the system includes an image matrix, an imaging direction, and a pixel spacing.
[0014] Furthermore, the coronal slices in the coronal slice set include uniform resolution.
[0015] Furthermore, the improved YOLOv8 model uses the EfficientViT structure to replace the CSPDarknet backbone network of the original YOLOv8 model, and adopts the AdaptiveSlide Loss function as the classification loss function.
[0016] The EfficientViT structure captures local and global features through a hierarchical multi-head self-attention mechanism and optimizes on weight quantization or image partitioning.
[0017] The AdaptiveSlide Loss function is as follows:
[0018]
[0019] In the formula, i is the sample number, N is the total number of samples, and x i represents the i-th sample. cls is the weighted classification loss function. SlideLoss(x i ) represents the classification loss function.
[0020] Among them, the weighting coefficient f(x i ) is as follows:
[0021]
[0022] In the formula, α represents the weight coefficient, μ represents the mean IoU, and δ represents the preset threshold.
[0023] Furthermore, the training steps of the YOLOv8_slice model include:
[0024] a1 obtains a historical coronal slice set, annotates the coronal slices in the historical coronal slice set, and marks the brainstem structure.
[0025] a2 uses the coronal slices in the historical coronal slice set as input and the target position and confidence of the brainstem structure as output to train the improved YOLOv8 model and obtain the YOLOv8_slice model.
[0026] The output of the YOLOv8_slice model As shown below:
[0027]
[0028] Where b i is the bounding box, c i For confidence.
[0029] Furthermore, the training steps of the YOLOv8_MTA model include:
[0030] b1 obtains a historical coronal slice set, and selects a coronal slice from the historical coronal slice set to reflect the medial temporal lobe atrophy score information.
[0031] b2 The brainstem structures are marked in the selected coronal slices, and the grade and location of atrophy of the medial temporal lobe are marked.
[0032] The positions include the left hippocampus and the right hippocampus.
[0033] b3 performs size standardization processing on the marked coronal slices to obtain processed coronal slices.
[0034] b4 Use the processed coronal slices to train the improved YOLOv8 model to obtain the YOLOv8_MTA model.
[0035] Further, the optimal coronal slice is as follows:
[0036]
[0037] In the formula, I * Indicates the best coronal slice. j represents the jth coronal slice in the coronal slice set. j )=1 indicates the jth coronal slice I j Has brain stem structure. j ) represents the detection confidence.
[0038] Furthermore, the output results of the YOLOv8_MTA model include bounding boxes, detection categories, and confidence levels.
[0039] Furthermore, when there are multiple output results for the brainstem structure, the one with the highest confidence is selected as the final result, as shown below:
[0040]
[0041] In the formula, They represent the final medial temporal lobe atrophy grade output results of the left hippocampus and the right hippocampus, respectively. k represents the optimal coronal slice I of the YOLOv8_MTA model. * Index of medial temporal lobe atrophy grade detected in . Respectively represent the medial temporal lobe atrophy level set detected by the YOLOv8_MTA model in the left hippocampus and the right hippocampus. k is the confidence level of the kth medial temporal lobe atrophy grade.
[0042] Furthermore, the automatic scoring system also includes a visualization platform.
[0043] The visualization platform includes a front-end interactive interface and a back-end service platform.
[0044] The front-end interactive interface is used to display the best coronal slice, the bounding box of the brainstem structure, and the medial temporal lobe assessment results.
[0045] The backend service platform is used to receive brain MRI scan data.
[0046] The technical effect of the present invention is unquestionable. The present invention provides an automated MTA scoring technology that takes into account accuracy, versatility, and ease of operation. By using a lightweight deep learning YOLO target detection framework and combining it with an appropriate loss function, the computational overhead is reduced while ensuring algorithm performance, and combined with a friendly front-end interactive interface, the efficiency and accuracy of clinical MTA scoring are improved.
[0047] The present invention realizes the automation of the whole process from MRI data import, automatic screening of optimal coronal slices, hippocampal area detection to MTA score output, and has high repeatability and portability. The present invention has the following advantages:
[0048] 1) Automation and high efficiency: Thanks to the advantages of the improved YOLOv8 in detection efficiency, the present invention can complete the coronal slice screening and MTA bilateral scoring of a single MRI data in a short time. The overall processing time is short, only about 1.9 seconds per case, and the MTA scoring results of the left and right hemispheres and the corresponding detection areas are displayed in real time, which greatly improves the clinical work efficiency and is suitable for clinical high-throughput image analysis needs.
[0049] 2) Improved accuracy and discrimination: The introduction of the EfficientViT backbone makes the model more advantageous in capturing detailed features; AdaptiveSlide Loss pays more attention to adjacent levels, greatly reducing the confusion rate of scoring. It shows excellent classification performance under various evaluation indicators, especially in the ability to distinguish adjacent levels, which is significantly improved through the adaptive weighted loss function.
[0050] 3) Friendly visualization and clinical usability: The built web platform allows clinicians to upload images and view results with one click. The system automatically completes slice selection and scoring, without the need for professional technical operations, saving time and effort, and facilitating deployment and promotion in hospitals or scientific research institutions. The MRI images with detection frames generated by the system intuitively display the atrophy of the hippocampus and related structures, assisting doctors in making quick decisions.
[0051] 4) High consistency: The high ICC value compared with the manual scoring results indicates that the system has a high degree of consistency and can effectively reduce human subjective errors.
[0052] 5) Convenient clinical integration: One-click operation is achieved through the Web front end, which can be easily integrated into the existing clinical information system to improve the convenience and popularity of practical applications.
[0053] In summary, the present invention provides an efficient, accurate and consistent automated medial temporal lobe atrophy scoring system, which has significant clinical application value and promotion prospects, and is expected to play an important role in the early diagnosis and monitoring of neurodegenerative diseases such as Alzheimer's disease. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flow chart of the present invention;
[0055] Figure 2 It is a structural diagram of the improved YOLOv8 model of the present invention;
[0056] Figure 3 is an example diagram of the system interface of the present invention, Figure 3 (a) is a schematic diagram of the file upload interface. Figure 3 (b) is a schematic diagram of the scoring result display interface. DETAILED DESCRIPTION
[0057] The present invention is further described below in conjunction with the embodiments, but it should not be understood that the above subject matter of the present invention is limited to the following embodiments. Without departing from the above technical ideas of the present invention, various substitutions and changes are made according to the common technical knowledge and customary means in the art, which should all be included in the protection scope of the present invention.
[0058] Embodiment 1:
[0059] See also Figures 1 to 3 , an automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning, including: a coronal slice screening module, an improved YOLOv8 model building module, an optimal coronal slice automatic selection module, and a medial temporal lobe atrophy detection and automatic scoring module.
[0060] The coronal slice screening module screens out coronal slices from the brain MRI scan data input into the system and constructs a coronal slice set.
[0061] The improved YOLOv8 model construction module is used to construct two improved YOLOv8 models, and train them into a YOLOv8_slice model and a YOLOv8_MTA model respectively.
[0062] The optimal coronal slice automatic selection module inputs the coronal slices in the coronal slice set into the YOLOv8_slice model to select the optimal coronal slice.
[0063] The medial temporal lobe atrophy grade automatic scoring module inputs the best coronal slice into the YOLOv8_MTA model to determine the medial temporal lobe atrophy grade.
[0064] Embodiment 2:
[0065] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content is shown in Example 1, further, the structure of the brain MRI scan data input into the system includes an image matrix, an imaging direction, and a pixel spacing.
[0066] Embodiment 3:
[0067] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 1 to 2, and further, the coronal slices in the coronal slice set contain uniform resolution.
[0068] Embodiment 4:
[0069] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 1 to 3. Furthermore, the improved YOLOv8 model uses the EfficientViT structure to replace the CSPDarknet backbone network of the original YOLOv8 model, and adopts the AdaptiveSlide Loss function as the classification loss function.
[0070] The EfficientViT structure captures local and global features through a hierarchical multi-head self-attention mechanism and optimizes on weight quantization or image partitioning.
[0071] The AdaptiveSlide Loss function is as follows:
[0072]
[0073] In the formula, i is the sample number, N is the total number of samples, and x i represents the i-th sample. cls is the weighted classification loss function. SlideLoss(x i ) represents the classification loss function.
[0074] Among them, the weighting coefficient f(x i ) is as follows:
[0075]
[0076] In the formula, α represents the weight coefficient, μ represents the mean IoU, and δ represents the preset threshold.
[0077] Embodiment 5:
[0078] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 1 to 4. Further, the training steps of the YOLOv8_slice model include:
[0079] a1 obtains a historical coronal slice set, annotates the coronal slices in the historical coronal slice set, and marks the brainstem structure.
[0080] a2 uses the coronal slices in the historical coronal slice set as input and the target position and confidence of the brainstem structure as output to train the improved YOLOv8 model and obtain the YOLOv8_slice model.
[0081] The output of the YOLOv8_slice model As shown below:
[0082]
[0083] Where b i is the bounding box, c i For confidence.
[0084] Embodiment 6:
[0085] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 1 to 5. Further, the training steps of the YOLOv8_MTA model include:
[0086] b1 obtains a historical coronal slice set, and selects a coronal slice from the historical coronal slice set to reflect the medial temporal lobe atrophy score information.
[0087] b2 The brainstem structures are marked in the selected coronal slices, and the grade and location of atrophy of the medial temporal lobe are marked.
[0088] The positions include the left hippocampus and the right hippocampus.
[0089] b3 performs size standardization processing on the marked coronal slices to obtain processed coronal slices.
[0090] b4 Use the processed coronal slices to train the improved YOLOv8 model to obtain the YOLOv8_MTA model.
[0091] Embodiment 7:
[0092] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 1 to 6. Further, the optimal coronal slice is as follows:
[0093]
[0094] In the formula, I * Indicates the best coronal slice. j represents the jth coronal slice in the coronal slice set. j )=1 indicates the jth coronal slice I j Has brain stem structure. j ) represents the detection confidence.
[0095] Embodiment 8:
[0096] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 1 to 7. Furthermore, the output result of the YOLOv8_MTA model includes a bounding box, a detection category, and a confidence level.
[0097] Embodiment 9:
[0098] An automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning, the main technical content of which is shown in any one of Examples 1 to 8. Further, when multiple output results appear in the brainstem structure, the one with the highest confidence is selected as the final result, as shown below:
[0099]
[0100] In the formula, They represent the final medial temporal lobe atrophy grade output results of the left hippocampus and the right hippocampus, respectively. k represents the optimal coronal slice I of the YOLOv8_MTA model. * Index of grade of medial temporal lobe atrophy detected in . Respectively represent the medial temporal lobe atrophy level set detected by the YOLOv8_MTA model in the left hippocampus and the right hippocampus. k is the confidence level of the kth medial temporal lobe atrophy grade.
[0101] Embodiment 10:
[0102] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 1 to 9. Furthermore, the automatic scoring system also includes a visualization platform.
[0103] The visualization platform includes a front-end interactive interface and a back-end service platform.
[0104] The front-end interactive interface is used to display the best coronal slice, the bounding box of the brainstem structure, and the medial temporal lobe assessment results.
[0105] The backend service platform is used to receive brain MRI scan data.
[0106] Embodiment 11:
[0107] See also Figures 1 to 3 , an automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning, including: a coronal slice screening module, an improved YOLOv8 model building module, an optimal coronal slice automatic selection module, and a medial temporal lobe atrophy detection and automatic scoring module.
[0108] The coronal slice screening module screens out coronal slices from the brain MRI scan data input into the system and constructs a coronal slice set.
[0109] DICOM files obtained from brain MRI scans were obtained from the image database or local file system. The Python pydicom tool library was used to parse the header metadata of each DICOM file (such as SeriesDescription, ImagePositionPatient, etc.) to form a data structure containing key information such as image matrix, imaging direction, pixel spacing, etc., thereby obtaining the three-dimensional matrix and sequence attributes of the original MR image, providing a basis for the subsequent screening of coronal slices.
[0110] Let D represent the set of all DICOM files that have been parsed, and let the keyword set K = {"Coronal", "COR", ...}. Then, traverse the SeriesDescription data field of each sequence in D. If any keyword k is matched, i ∈K, then this sequence is marked as the coronal sequence, and all slices in the coronal sequence (denoted as I 1 ,I 2 ,…,I n ) to make a unified size, and resample each image to 640×640. All slices that meet the coronal position are put into the candidate list for subsequent model detection to obtain the candidate coronal slice set {I i}, each slice contains uniform resolution.
[0111] The improved YOLOv8 model construction module is used to construct two improved YOLOv8 models, and train them into a YOLOv8_slice model and a YOLOv8_MTA model respectively.
[0112] The optimal coronal slice automatic selection module inputs the coronal slices in the coronal slice set into the YOLOv8_slice model to select the optimal coronal slice.
[0113] The medial temporal lobe atrophy grade automatic scoring module inputs the best coronal slice into the YOLOv8_MTA model to determine the medial temporal lobe atrophy grade.
[0114] Embodiment 12:
[0115] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content is shown in Example 11. Furthermore, the structure of the brain MRI scan data input into the system includes an image matrix, an imaging direction, and a pixel spacing.
[0116] Embodiment 13:
[0117] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 11 to 12. Furthermore, the coronal slices in the coronal slice set contain uniform resolution.
[0118] Embodiment 14:
[0119] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 11 to 13. Furthermore, the improved YOLOv8 model uses the EfficientViT structure to replace the CSPDarknet backbone network of the original YOLOv8 model, and adopts the AdaptiveSlide Loss function as the classification loss function.
[0120] The EfficientViT structure captures local and global features through a hierarchical multi-head self-attention mechanism and optimizes on weight quantization or image partitioning.
[0121] The default CSPDarknet backbone network in the original YOLOv8 is replaced with the lightweight Vision Transformer structure EfficientViT, which captures local and global features through a hierarchical multi-head self-attention mechanism and optimizes weight quantization or patch division to ensure high efficiency and accuracy in MRI medical image analysis. Then, the EfficientViT pre-trained weights are loaded, and the number of channels and feature map size of the YOLOv8 detection head are adjusted to adapt to the new backbone network, obtaining an efficient YOLOv8 backbone suitable for medical images, reducing inference time and enhancing the feature recognition capabilities of subtle structures (such as the hippocampus).
[0122] In the original YOLOv8, classification and regression losses usually use BCE (Binary Cross Entropy) and CIoU. In order to improve the recognition of MTA level boundaries, the present invention replaces the classification part with adaptive weighted Slide Loss (AdaptiveSlide Loss), and its core idea is as follows:
[0123] The mean IoU of all detection boxes and the true box in the current batch is μ, and a small threshold δ is set. If the IoU of a detection sample falls within the interval [μ-δ,μ+δ], it is considered a "difficult sample", that is, a sample that is difficult to distinguish from adjacent level samples. For samples with adjacent levels or a difference of 1 in MTA grading, increase the weight coefficient α, and set α=α 0 +λ, where λ is an additional increment. Finally, the adaptive weight f(x) is added to the Slide Loss to obtain the adaptive weighted Slide Loss. The AdaptiveSlideLoss function is as follows:
[0124]
[0125] In the formula, i is the sample number, N is the total number of samples, and x i represents the i-th sample. cls is the weighted classification loss function. SlideLoss(x i ) represents the classification loss function.
[0126] Among them, the weighting coefficient f(x i ) is as follows:
[0127]
[0128] In the formula, α represents the weight coefficient, μ represents the mean IoU, and δ represents the preset threshold.
[0129] During the training process, the default classification loss function is AdaptiveSlide Loss. In each iteration, the μ of this batch is dynamically calculated and the weight coefficient f(x i ), so when performing MTA scoring detection on situations where the boundaries between adjacent levels are easily confused, the model will tend to focus on these "difficult" samples and improve the accuracy of classification.
[0130] Embodiment 15:
[0131] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 11 to 14. Further, the training steps of the YOLOv8_slice model include:
[0132] a1 obtains a historical coronal slice set, annotates the coronal slices in the historical coronal slice set, and marks the brainstem structure.
[0133] a2 uses the coronal slices in the historical coronal slice set as input and the target position and confidence of the brainstem structure as output to train the improved YOLOv8 model and obtain the YOLOv8_slice model.
[0134] The output of the YOLOv8_slice model As shown below:
[0135]
[0136] Where b i is the bounding box, c i For confidence.
[0137] Embodiment 16:
[0138] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 11 to 15. Further, the training steps of the YOLOv8_MTA model include:
[0139] b1 obtains a historical coronal slice set, and selects a coronal slice from the historical coronal slice set to reflect the medial temporal lobe atrophy score information.
[0140] b2 The brainstem structures are marked in the selected coronal slices, and the grade and location of atrophy of the medial temporal lobe are marked.
[0141] The positions include the left hippocampus and the right hippocampus.
[0142] For each training slice, the bounding boxes and category labels (MTA levels) of the regions of interest such as the left and right hippocampus, temporal horn, and choroid fissure were marked, where ZERO represents MTA level 0, ONE represents MTA level 1, TWO represents MTA level 2, THREE represents MTA level 3, and FOUR represents MTA level 4.
[0143] b3 performs size standardization processing on the marked coronal slices to obtain processed coronal slices.
[0144] b4 Use the processed coronal slices to train the improved YOLOv8 model to obtain the YOLOv8_MTA model.
[0145] Embodiment 17:
[0146] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 11 to 16, and further, the optimal coronal slice is as follows:
[0147]
[0148] In the formula, I * Indicates the best coronal slice. j represents the jth coronal slice in the coronal slice set. j )=1 indicates the jth coronal slice I j Has brain stem structure. j ) represents the detection confidence.
[0149] Embodiment 18:
[0150] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 11 to 17. Furthermore, the output result of the YOLOv8_MTA model includes a bounding box, a detection category, and a confidence level.
[0151] Embodiment 19:
[0152] A medial temporal lobe atrophy grade automatic scoring system based on YOLO deep learning, the main technical content of which is shown in any one of Examples 11 to 18. Further, when multiple output results appear in the brainstem structure, the one with the highest confidence is selected as the final result, as shown below:
[0153]
[0154] In the formula, They represent the final medial temporal lobe atrophy grade output results of the left hippocampus and the right hippocampus, respectively. k represents the optimal coronal slice I of the YOLOv8_MTA model. * Index of medial temporal lobe atrophy grade detected in . Respectively represent the medial temporal lobe atrophy level set detected by the YOLOv8_MTA model in the left hippocampus and the right hippocampus. k is the confidence level of the kth medial temporal lobe atrophy grade.
[0155] Embodiment 20:
[0156] An automatic scoring system for medial temporal lobe atrophy grade based on YOLO deep learning, the main technical content of which is shown in any one of Examples 11 to 19. Furthermore, the automatic scoring system also includes a visualization platform.
[0157] The visualization platform includes a front-end interactive interface and a back-end service platform.
[0158] The front-end interactive interface is used to display the best coronal slice, the bounding box of the brainstem structure, and the medial temporal lobe assessment results.
[0159] Using HTML technology, we implement image visualization components, load the test results returned by the server, and display the medial temporal lobe evaluation results of the best slice. Users only need to upload MRI files on the front end to view the automatically selected best coronal slice, the bounding boxes of the left and right hippocampi, and the corresponding MTA grade scores. This one-stop clinical visualization platform makes it easy for doctors or researchers to complete a quick and objective MTA evaluation.
[0160] The backend service platform is used to receive brain MRI scan data.
[0161] The above two improved YOLOv8 reasoning processes (first YOLOv8_slice, then YOLOv8_MTA) are integrated into the Flask-based Web framework. The backend receives the DICOM files uploaded by the user, decodes them and performs batch reasoning, and finally returns the automatic scoring results to form API interfaces that can be called by the front end, such as " / upload", " / get_score", etc.
[0162] Embodiment 21:
[0163] See also Figures 1 to 3 , an automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning, the main technical contents include:
[0164] Step 1: DICOM data reading and coronal slice screening
[0165] 1) DICOM data analysis
[0166] DICOM files obtained from brain MRI scans were obtained from the image database or local file system. The Python pydicom tool library was used to parse the header metadata of each DICOM file (such as SeriesDescription, ImagePositionPatient, etc.) to form a data structure containing key information such as image matrix, imaging direction, pixel spacing, etc., thereby obtaining the three-dimensional matrix and sequence attributes of the original MR image, providing a basis for the subsequent screening of coronal slices.
[0167] 2) Coronal slice screening
[0168] Let D represent the set of all DICOM files that have been parsed, and let the keyword set K = {"Coronal", "COR", ...}. Then, traverse the SeriesDescription data field of each sequence in D. If any keyword k is matched, i∈K, then this sequence is marked as the coronal sequence, and all slices in the coronal sequence (denoted as I 1 ,I 2 ,…,I n ) to make a unified size, and resample each image to 640×640. All slices that meet the coronal position are put into the candidate list for subsequent model detection to obtain the candidate coronal slice set {I i}, each slice contains uniform resolution.
[0169] Step 2: Improve YOLOv8 model construction
[0170] In this paper, two improved YOLOv8 models (hereinafter referred to as YOLOv8_slice and YOLOv8_MTA) are trained and used for the two tasks of "automatic selection of the best coronal slice" and "MTA detection and automatic scoring". The core improvements of the two models include the introduction of EfficientViT as the backbone and the use of AdaptiveSlide Loss in the loss function to increase the focus on edge samples.
[0171] 1) Improved network backbone EfficientViT
[0172] The default CSPDarknet backbone network in the original YOLOv8 is replaced with the lightweight Vision Transformer structure EfficientViT, which captures local and global features through a hierarchical multi-head self-attention mechanism and optimizes weight quantization or patch division to ensure high efficiency and accuracy in MRI medical image analysis. Then, the EfficientViT pre-trained weights are loaded, and the number of channels and feature map size of the YOLOv8 detection head are adjusted to adapt to the new backbone network, obtaining an efficient YOLOv8 backbone suitable for medical images, reducing inference time and enhancing the feature recognition capabilities of subtle structures (such as the hippocampus).
[0173] 2) Adaptive Slide Loss
[0174] In the original YOLOv8, classification and regression losses usually use BCE (Binary Cross Entropy) and CIoU. In order to improve the recognition of MTA level boundaries, the present invention replaces the classification part with adaptive weighted Slide Loss (AdaptiveSlide Loss), and its core idea is as follows:
[0175] The mean IoU of all detection boxes and the true box in the current batch is μ, and a small threshold δ is set. If the IoU of a detection sample falls within the interval [μ-δ,μ+δ], it is considered a "difficult sample", that is, a sample that is difficult to distinguish from adjacent level samples. For samples with adjacent levels or a difference of 1 in MTA grading, increase the weight coefficient α, and set α=α 0 +λ, where λ is an additional increment. Finally, the adaptive weight f(x) is added to the Slide Loss to obtain the adaptive weighted Slide Loss, namely AdaptiveSlideLoss:
[0176]
[0177] Among them, f(x i )for:
[0178]
[0179] During the training process, the default classification loss function is AdaptiveSlide Loss. In each iteration, the μ of this batch is dynamically calculated and the weight coefficient f(x i ), so when performing MTA scoring detection on situations where the boundaries between adjacent levels are easily confused, the model will tend to focus on these "difficult" samples and improve the accuracy of classification.
[0180] Step 3: Automatic selection of the best coronal slice (YOLOv8_slice model)
[0181] 1) Detection target and training
[0182] The coronal slice set I obtained in step 1 i We annotate some of the data in the image, mark the key brainstem structures, and use the improved YOLOv8 model (including EfficientViT backbone and AdaptiveSlide Loss) for training. The detector outputs the target position and confidence for each candidate slice. where b i is the bounding box, c i After the training is completed, a model YOLOv8_slice is obtained that can automatically detect and judge whether "structures such as the brainstem are included and fully displayed".
[0183] 2) Model use stage
[0184] For all candidate slices I 1 ,...,I n They are sent to YOLOv8_slice respectively. If a brainstem target is detected, it is marked as D(I i )=1, otherwise D(Ii )=0, and record the highest detection confidence C(I i ). If it exists Make D(I j )=1, the slice with the highest confidence is selected as the optimal coronal slice, that is,
[0185]
[0186] If there is no I i Satisfy D(I i )=1, the middle slice can be used as the optimal coronal slice. After the MRI image of the new case is input, YOLOv8_slice is run and the calculation is completed, and the optimal slice I is finally output. * , used for subsequent medial temporal lobe atrophy detection scoring.
[0187] Step 4: Medial temporal lobe atrophy detection and automatic scoring (YOLOv8_MTA model)
[0188] In order to obtain the optimal slice I * Afterwards, another trained improved YOLOv8 model (YOLOv8_MTA) was used to identify features such as the hippocampus and evaluate the MTA grade.
[0189] 1) Detection target and training
[0190] From the existing MRI coronal slice library, select samples with clear hippocampus, medial temporal lobe related structure annotations and known MTA score information (0-4). For each training slice, mark the bounding box and category label (MTA level) of the left and right hippocampus, temporal horn, choroid fissure and other regions of interest, where ZERO represents MTA level 0, ONE represents MTA level 1, TWO represents MTA level 2, THREE represents MTA level 3, and FOUR represents MTA level 4. Referring to the previous steps, the image size is standardized to ensure consistency of model input. Use the above improved YOLOv8 model (including EfficientViT backbone and AdaptiveSlideLoss) for training. After training, a model YOLOv8_MTA that can automatically detect MTA scores is obtained.
[0191] 2) Model use stage
[0192] The optimal coronal slice I selected in step 3 * Input the YOLOv8_MTA model as the image data for inference. The model will output a series of candidate detection results, including the bounding box b k , detection category c k(such as "ZERO", "ONE", etc.) and confidence level p k When there may be multiple test results for the left or right hippocampus or medial temporal lobe, the one with the highest confidence can be selected as the final result:
[0193]
[0194] in and Respectively represent all candidate target sets detected by the model on the left and right.
[0195] Convert the selected left and right detection results into the corresponding MTA grade score S L With S R The MTA score, as the final output, automates the assessment of medial temporal lobe atrophy and can be used for the final visualization.
[0196] Step 5: Visualization platform construction
[0197] 1) Backend services
[0198] The above two improved YOLOv8 reasoning processes (first YOLOv8_slice, then YOLOv8_MTA) are integrated into the Flask-based Web framework. The backend receives the DICOM files uploaded by the user, decodes them and performs batch reasoning, and finally returns the automatic scoring results to form API interfaces that can be called by the front end, such as " / upload", " / get_score", etc.
[0199] 2) Front-end interactive interface
[0200] Using HTML technology, we implement image visualization components, load the test results returned by the server, and display the medial temporal lobe evaluation results of the best slice. Users only need to upload MRI files on the front end to view the automatically selected best coronal slice, the bounding boxes of the left and right hippocampi, and the corresponding MTA grade scores. This one-stop clinical visualization platform makes it easy for doctors or researchers to complete a quick and objective MTA evaluation.
[0201] Embodiment 22:
[0202] See also Figures 1 to 3 , an automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning, the main technical contents include:
[0203] 1) Experimental settings and evaluation indicators
[0204] In order to comprehensively evaluate the effectiveness of the method of the present invention, a total of 556 MRI data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database were selected. The experiment is divided into a training set, a validation set, and a test set. The specific distribution is shown in Table 1. The evaluation indicators used include classification sensitivity, specificity, accuracy, precision, and F1 score.
[0205] Table 1: Dataset distribution
[0206]
[0207] 2) Best slice selection performance
[0208] The present invention uses the improved YOLOv8_slice model to automatically select the best coronal slices and compares the selection results with those of professional evaluators. The experimental results show that the accuracy of automatic slice selection reaches 92.0%, which is significantly better than the traditional method. The specific data is shown in Table 2.
[0209] Table 2: Comparison of the best slice selection accuracy
[0210]
[0211] 3) MTA scoring performance
[0212] As shown in Table 3, on the ADNI test set, the automatic MTA scoring system of the present invention shows high accuracy and consistency, reaching 76.5% and 73.5% accuracy in the left and right halves, respectively, indicating that the improved model YOLOv8_MTA has good discrimination ability in classification tasks.
[0213] Table 3: MTA scoring performance
[0214]
[0215] 4) Ability to differentiate between different levels
[0216] In order to solve the difficulty of distinguishing adjacent levels (such as level 1 and level 2, level 2 and level 3) in MTA scoring, the present invention significantly improves the performance of the model in fine-grained classification by introducing an adaptive weighted loss function (AdaptiveSlide Loss).
[0217] As can be seen from Table 4, the classification accuracy of level 0 and level 4 is relatively high, especially the specificity has reached 1.000, indicating that the model has excellent recognition effect in the case of extreme atrophy and no atrophy. For level 1 and level 3, the sensitivity is 0.944 and 0.889 respectively, showing good discrimination ability. Level 2 has a relatively complex sample distribution and characteristics, and the sensitivity has decreased, but the specificity remains at a high level (0.947).
[0218] Table 4: MTA rating performance by level
[0219]
[0220] 5) Comparison with other methods
[0221] In order to verify the superiority of the method of the present invention, a comparative experiment was carried out with common detection and segmentation models such as YOLOv5, Faster R-CNN, Mask R-CNN and U-Net. The results are shown in Table 5.
[0222] Table 5: Performance comparison of different methods
[0223]
[0224] As can be seen from Table 5, the YOMTA system of the present invention is superior to other comparison methods in terms of slice selection accuracy and MTA scoring accuracy, and the average inference time is only 1.9 seconds per case, which is significantly lower than other methods. This shows that the present invention has higher computational efficiency while maintaining high accuracy, and is suitable for clinical high-throughput image analysis needs.
[0225] 6) Consistency evaluation
[0226] In order to evaluate the consistency between the automatic scoring system and manual scoring, a two-way random intraclass correlation coefficient (ICC) was used for statistical analysis, and the results are shown in Table 6.
[0227] As can be seen from Table 6, the ICC values between the left and right hemispheres and the manual scoring of the automatic scoring system of the present invention are 0.884 and 0.858 (single measurement), 0.939 and 0.924 (average measurement), respectively, which are all at the "highly consistent" level (ICC>0.75). This shows that the automatic scoring results of the present invention are highly consistent with the manual scoring and can effectively replace the traditional manual scoring method.
[0228] Table 6: ICC values of automatic scoring and manual scoring
[0229]
[0230] 7) Actual application effect
[0231] In actual clinical applications, the YOMTA system of the present invention realizes the automation of the entire process from uploading MRI data to outputting MTA scoring results through the integrated Flask front-end interface, which is specifically manifested as follows:
[0232] Easy operation: Medical staff only need to upload DICOM format MRI files through the web interface, and the system will automatically complete slice selection and scoring without the need for professional technical operations.
[0233] Real-time feedback: The entire scoring process is completed within 1.9 seconds, and the MTA scoring results of the left and right hemispheres and the corresponding detection areas are displayed in real time, greatly improving clinical work efficiency.
[0234] Data visualization: The system generates MRI images with detection frames, which intuitively display the atrophy of the hippocampus and related structures, helping doctors make quick decisions.
Claims
1. An automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning, characterized in that: include: Coronal slice screening module, improved YOLOv8 model building module, optimal coronal slice automatic selection module, medial temporal lobe atrophy detection and automatic scoring module; The coronal slice screening module screens coronal slices from the brain MRI scan data input into the system and constructs a coronal slice set; The improved YOLOv8 model construction module is used to construct two improved YOLOv8 models, and train them into a YOLOv8_slice model and a YOLOv8_MTA model respectively. The optimal coronal slice automatic selection module inputs the coronal slices in the coronal slice set into the YOLOv8_slice model to select the optimal coronal slice. The medial temporal lobe atrophy grade automatic scoring module inputs the best coronal slice into the YOLOv8_MTA model to determine the medial temporal lobe atrophy grade.
2. According to the YOLO deep learning-based medial temporal lobe atrophy automatic scoring system of claim 1, it is characterized in that: The structure of the brain MRI scan data input into the system includes an image matrix, an imaging direction, and a pixel spacing.
3. The automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning according to claim 1 is characterized in that: The coronal slices in the coronal slice set comprise a uniform resolution.
4. The automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning according to claim 1, characterized in that: The improved YOLOv8 model uses the EfficientViT structure to replace the CSPDarknet backbone network of the original YOLOv8 model, and adopts the AdaptiveSlide Loss function as the classification loss function; The EfficientViT structure captures local and global features through a hierarchical multi-head self-attention mechanism and optimizes weight quantization or image partitioning; The AdaptiveSlide Loss function is as follows: In the formula, i is the sample number, N is the total number of samples, and x i represents the i-th sample; L cls is the weighted classification loss function; SlideLoss(x i ) represents the classification loss function; Among them, the weighting coefficient f(x i ) is as follows: In the formula, α represents the weight coefficient; μ represents the mean IoU; δ represents the preset threshold.
5. The automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning according to claim 1, characterized in that: The training steps of the YOLOv8_slice model include: a1 obtains a historical coronal slice set, and annotates the coronal slices in the historical coronal slice set to mark the brainstem structure; a2 uses the coronal slices in the historical coronal slice set as input and the target position and confidence of the brainstem structure as output to train the improved YOLOv8 model to obtain the YOLOv8_slice model; Output of the YOLOv8_slice model As shown below: Where b i is the bounding box, c i For confidence.
6. The automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning according to claim 1, characterized in that: The training steps of the YOLOv8_MTA model include: b1 obtaining a historical coronal slice set, and selecting a coronal slice for reflecting the medial temporal lobe atrophy score information from the historical coronal slice set; b2 Mark the brainstem structure in the selected coronal slices and mark the atrophy grade and location of the medial temporal lobe; The locations include the left hippocampus and the right hippocampus; b3 performing size standardization processing on the marked coronal slices to obtain processed coronal slices; b4 Use the processed coronal slices to train the improved YOLOv8 model to obtain the YOLOv8_MTA model.
7. The automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning according to claim 1, characterized in that: The optimal coronal slices are as follows: In the formula, I * Indicates the best coronal slice; I j represents the jth coronal slice in the coronal slice set; D(I j )=1 indicates the jth coronal slice I j With brainstem structure; C(I j ) represents the detection confidence.
8. The automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning according to claim 1, characterized in that: The output results of the YOLOv8_MTA model include bounding boxes, detection categories, and confidence levels.
9. The automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning according to claim 8, characterized in that: When there are multiple output results for the brainstem structure, the one with the highest confidence is selected as the final result, as shown below: In the formula, They represent the final medial temporal lobe atrophy grade output results of the left hippocampus and the right hippocampus respectively; k represents the optimal coronal slice I of the YOLOv8_MTA model * The medial temporal lobe atrophy grade index detected in; Respectively represent the medial temporal lobe atrophy level set detected by the YOLOv8_MTA model in the left hippocampus and the right hippocampus; c k is the confidence level of the kth medial temporal lobe atrophy grade.
10. The automatic scoring system for medial temporal lobe atrophy based on YOLO deep learning according to claim 1, characterized in that: The automatic scoring system also includes a visualization platform; The visualization platform includes a front-end interactive interface and a back-end service platform; The front-end interactive interface is used to display the best coronal slice, the bounding box of the brainstem structure and the medial temporal lobe assessment result; The backend service platform is used to receive brain MRI scan data.
Citation Information
Patent Citations
Medical image display processing method, medical image display processing device, and program
CN106659424A
Alzheimer's disease detection method based on deep learning, and computer readable medium
CN113763343A
Cartridge and aerosol generating device comprising the same
KR1020250025195A
Efficient battery management system
KR1020250070684A
Method of providing diagnosis assistance information and method of performing the same
US20220207720A1
Cited By
Improved YOLO11-based water hyacinth target rapid detection method
CN120635395A