Ankylosing spondylitis multi-tag automatic diagnosis and evaluation system and method
By combining a deep learning multi-label classification system with prior attention mechanisms and transfer learning, the problems of accuracy and comprehensiveness in the diagnosis of ankylosing spondylitis have been solved. This system enables the simultaneous identification and efficient diagnosis of multiple joint lesions, is applicable to different medical environments, and improves diagnostic efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-03
AI Technical Summary
Existing diagnostic methods for ankylosing spondylitis rely on the experience of specialist physicians and expensive medical equipment, making them difficult to implement effectively, especially in low-resource areas. Furthermore, the lack of application of multi-label classification models leads to insufficient diagnostic accuracy and comprehensiveness.
A deep learning-based multi-label classification system, combined with prior attention mechanism and transfer learning, is used to perform multi-dimensional analysis of pelvic X-ray images through a multi-label deep learning model. Key areas are automatically identified and diagnostic reports are generated, supplemented by Grad-CAM and T-SNE algorithms for visualization interpretation.
It enables the simultaneous identification of multiple joint lesions, improves diagnostic efficiency and accuracy, is applicable to different medical environments, reduces the misdiagnosis rate and saves medical resources, and enhances the acceptability of AI in clinical practice.
Smart Images

Figure CN121789980A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to artificial intelligence technology, and more particularly to a multi-label automatic diagnosis and assessment system and method for ankylosing spondylitis based on deep learning. Background Technology
[0002] Ankylosing spondylitis is a chronic inflammatory disease that primarily affects the sacroiliac joints and spine, causing irreversible joint damage. Early diagnosis and assessment are crucial for effective management of the disease, but current diagnostic methods often rely on the experience of specialist physicians and expensive medical equipment (such as MRI). This limitation makes the diagnosis of ankylosing spondylitis even more challenging in low-resource areas due to insufficient medical resources.
[0003] In recent years, deep learning and artificial intelligence technologies have been widely applied to the automated diagnosis of medical images, demonstrating superior performance in the early identification and assisted diagnosis of various diseases, and providing new ideas for the intelligent diagnosis of musculoskeletal diseases such as ankylosing spondylitis. However, existing research mainly focuses on single tasks (such as single-label classification), lacking exploration of multi-label classification models, especially how to combine imaging data with clinical data for multi-dimensional analysis to improve the accuracy and comprehensiveness of diagnosis. Summary of the Invention
[0004] This invention provides a multi-label automated diagnosis and assessment system and method for ankylosing spondylitis, addressing the limitation of current model classification capabilities. The system mainly comprises the following six modules:
[0005] The data acquisition and preprocessing module is used to acquire X-ray images of the patient's pelvis and preprocess the image data.
[0006] The prior attention mechanism module is used to automatically identify key regions in the image by combining previously labeled information and to perform weighted processing on these regions.
[0007] The multi-label classification deep learning model module uses deep learning technology to perform multi-label classification on pelvic X-ray images.
[0008] The results fusion and evaluation module is used to fuse features from different image regions and generate a unified final diagnostic result using a weighted fusion strategy.
[0009] The model training and optimization module is used to train deep learning models and optimize model performance using transfer learning and data augmentation techniques.
[0010] The output module is used to generate a corresponding ankylosing spondylitis diagnosis report based on the model's diagnostic results, and to provide risk assessment results for joint damage.
[0011] In some embodiments, the data acquisition and preprocessing module collects image data from medical institutions from different sources.
[0012] In some embodiments, the image data is preprocessed, specifically including preprocessing steps such as noise removal, brightness and contrast normalization, and image size scaling of the original image.
[0013] In some embodiments, deep convolutional neural networks or object detection models are used to detect images, automatically identify and label key regions, and dynamically adjust region weights based on information previously labeled by experts.
[0014] In some embodiments, the multi-label classification deep learning model module employs the following techniques:
[0015] A multi-label classification model based on a combination of deep convolutional neural networks (CNN) and visual transformers (ViT) is used to jointly train the model on lesions of different joints using a multi-task learning approach.
[0016] In some embodiments, the model training and optimization module includes:
[0017] Transfer learning techniques are used to transfer deep learning models that have been pre-trained on large-scale datasets to medical image datasets for fine-tuning.
[0018] Furthermore, data augmentation techniques are used to increase the diversity of training samples, thereby enhancing the model's generalization ability and robustness.
[0019] In some embodiments, a model interpretation module is also included: the module uses algorithms such as Grad-CAM and T-SNE to perform visual analysis of the model's decision region, assisting clinicians in understanding the decision-making basis of the AI model.
[0020] In some embodiments, the image is adjusted to 224*224 pixels after preprocessing.
[0021] In some embodiments, different task branches of the model share the underlying network weights and achieve information interaction between multiple labels through a joint feature fusion layer.
[0022] On the other hand, embodiments of this application provide a diagnostic report generation method based on the aforementioned system, including:
[0023] Acquiring and preprocessing X-ray images of the patient's pelvis;
[0024] Multi-label classification training based on a deep learning model with prior attention mechanism;
[0025] A final diagnostic report is generated through feature fusion and weighted evaluation.
[0026] The model interpretation module generates visual interpretation results.
[0027] Compared with existing technologies, the advantages of this invention are: Through a multi-label deep learning model, it achieves simultaneous identification of multiple joint lesions, avoiding the inefficiency of traditional manual interpretation, significantly shortening diagnostic time, and greatly improving diagnostic efficiency; It introduces a prior attention mechanism, enabling the model to automatically focus on the lesion area, reducing non-lesion interference, and combines transfer learning and multi-center data training, resulting in higher accuracy and stability; By introducing algorithms such as Grad-CAM and T-SNE, this system can visually interpret the model's predictions, assisting doctors in understanding the basis of AI diagnosis and enhancing the acceptability and trustworthiness of AI in clinical practice; Furthermore, this system can be applied not only to imaging centers of large hospitals but also deployed in community medical institutions, primary hospitals, or mobile medical devices, maintaining high diagnostic performance even in low-resource environments, demonstrating wide applicability. In summary, this invention effectively implements artificial intelligence in the early screening and assisted diagnosis of ankylosing spondylitis, providing technical support for early intervention of the disease, and has significant advantages in reducing misdiagnosis rates, improving diagnostic efficiency, and saving medical resources, possessing high clinical and social application value. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0029] Figure 1 This is a flowchart of an AI-based multi-label automated diagnosis and assessment method for ankylosing spondylitis.
[0030] Figure 2 This is a flowchart of data acquisition and preprocessing.
[0031] Figure 3 This is a structural diagram of the prior attention mechanism module.
[0032] Figure 4 This is a schematic diagram illustrating the result fusion and training optimization of a multi-label classification model.
[0033] Figure 5 This is a schematic diagram showing the model output and results. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0035] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly, for example, as a fixed connection, a detachable connection, or an integral connection; a mechanical connection or an electrical connection; a direct connection or an indirect connection through an intermediate medium; or a connection within two elements. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0036] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.
[0037] Reference Figure 1 The embodiments of this application mainly include the following modules, and the functions of each module are as follows:
[0038] 1. Data Acquisition and Preprocessing Module:
[0039] This module is responsible for collecting patients' pelvic X-ray images and preprocessing the data, including background removal, image scaling, and standardization, to ensure data consistency and suitability for model training.
[0040] 2. Prior Attention Mechanism Module:
[0041] Based on expert-annotated information, this module focuses attention on key areas (such as the sacroiliac joint and hip joint) to help the model pay more attention to these areas during training, thereby improving diagnostic accuracy.
[0042] 3. Multi-label classification model module:
[0043] This module employs a multi-label classification model from deep learning technology, enabling simultaneous assessment of multiple tasks, including the diagnosis of ankylosing spondylitis, the evaluation of sacroiliac joint injuries, and the evaluation of hip joint injuries. Through algorithm optimization, the model can process multiple labels more efficiently, providing more comprehensive diagnostic results.
[0044] 4. Results Integration and Evaluation Module:
[0045] This module can fuse features extracted from multiple regions and use an adaptive weighting method to combine local and global features, thereby improving the model's accuracy in assessing various types of damage.
[0046] 5. Model Training and Optimization Module:
[0047] This invention employs various neural network architectures (such as ResNet, ViT, Poolformer, etc.) for training and optimizes the model through transfer learning techniques. Data augmentation and cross-validation are used to ensure the model's robustness and accuracy.
[0048] 6. Output module:
[0049] Based on the final results after fusion, this module automatically generates a standardized diagnostic report. The report includes the diagnosis of ankylosing spondylitis, the assessment of the severity of damage to each joint, and risk warnings. It can be exported in PDF or XML format for easy integration into electronic medical record systems and use in telemedicine.
[0050] Example 1: Data Acquisition and Preprocessing Module
[0051] Please see Figure 2This module aims to establish a high-quality input data source to ensure the stability and reliability of model training. In its implementation, this invention selected 6970 cases from image databases of multiple medical institutions, including 2716 cases of ankylosing spondylitis and 4254 cases of non-ankylosing spondylitis. The data were primarily in DICOM and PNG formats, and all data underwent anonymization to protect patient privacy. Due to differences in exposure parameters among different hospital equipment, the original images often exhibit inconsistencies in brightness, contrast, and noise. Therefore, this invention first performs the following preprocessing operations on the images:
[0052] (1) A cropping strategy is adopted to remove background areas unrelated to the pelvis and retain only the key pelvic structure areas. This reduces the interference of irrelevant noise and enhances the feature learning ability of the model, effectively reducing the complexity of the input image and helping the model converge faster.
[0053] (2) The AutoImageProcessor tool in the transformers library is used to complete the normalization process. The pixel values are standardized according to the mean and standard deviation of the pre-trained model to ensure that the data distribution is consistent under different devices and acquisition conditions. This significantly improves the numerical stability of the data and accelerates the convergence speed of the model training stage.
[0054] (3) The black border filling strategy is adopted to maintain the original aspect ratio and avoid stretching distortion. Then the image is uniformly adjusted to 224×224 pixels to be compatible with the input format of mainstream pre-trained models, which is conducive to transfer learning and reduces the uncertainty of model training.
[0055] After the above processing, a consistent image dataset was obtained and divided into a training set, an internal test set, and a multi-center external test set, with a ratio of 6365:437:168.
[0056] Example 2: Prior Attention Mechanism Module
[0057] Please see Figure 3 After the image data is prepared, the system needs to automatically identify key anatomical regions. This module, based on information annotated by radiology experts, uses a target detection network (such as the YOLOv11 algorithm) to automatically identify and locate key anatomical regions such as the sacroiliac joint and hip joint. The detection results are then used to generate corresponding region masks, which serve as input references for the attention module.
[0058] Then, a mask obtained from detection is superimposed on the original image to form an input image with "prior attention". This input is fed into the model simultaneously with the original image during the training phase, and multi-view feature learning is achieved through a weight-sharing mechanism. The model can automatically assign weight parameters according to the prior features of the input, applying higher weights to key regions to enhance the model's response to lesion sites.
[0059] Finally, the features extracted from the original input and prior input are weighted and fused to generate a more discriminative high-level semantic feature representation. This mechanism enables the model to learn both "global morphology" and "local lesion" information simultaneously during the training phase. Furthermore, it has been verified that adding the prior attention mechanism significantly improves the model's recognition sensitivity on multi-center test sets and outperforms the baseline model without the attention mechanism in the ankylosing spondylitis diagnosis task.
[0060] Example 3: Multi-label classification deep learning model module
[0061] Please see Figure 4 Based on the feature graphs output by the prior attention mechanism, the overall design goal of this module is to achieve multi-task joint learning, simultaneously processing multiple relevant labels in a single network, thereby improving diagnostic efficiency and generalization performance. It includes the following steps:
[0062] (1) Model Structure Design: This module adopts a hybrid architecture that integrates ResNet and Vision Transformer (ViT). ResNet is used to extract local spatial texture features (such as bone surface contours and joint boundaries), while ViT is responsible for capturing global dependencies and structural symmetry features. The two are fused at high-level feature layers to balance the feature representation of both local and global information.
[0063] (2) Multi-label task setting: The network outputs multiple classification results at the same time, including AS diagnosis label: to determine whether the patient has ankylosing spondylitis (positive / negative); sacroiliac joint injury label (RSJ): to grade and assess the degree of sacroiliac joint lesions; left and right hip joint injury labels (LHIP, RHIP): to assess the degree of hip joint injury respectively.
[0064] (3) Shared weights and feature fusion: Different task branches of the model share the underlying network weights and achieve information interaction between multiple labels through a joint feature fusion layer. This shared design can reduce parameter redundancy and improve the model's generalization ability.
[0065] (4) Loss function and optimization objective: In order to balance the learning weights among multi-label tasks, this invention adopts the binary cross-entropy loss function and introduces a weighted loss strategy to solve the imbalance problem caused by the difference in the number of samples in each task.
[0066] (5) Optimization Algorithm and Training Strategy: The model uses the Adam optimizer for gradient updates, with adaptive adjustment of the learning rate. To prevent overfitting and improve generalization, a three-fold cross-validation method is used during the training phase. Simultaneously, data augmentation (such as rotation, flipping, and pruning) and transfer learning strategies are employed to ensure stable convergence of the model under multi-center data conditions.
[0067] Example 4: Result Fusion and Evaluation Module
[0068] Please see Figure 4 This module's task is to integrate classification results from multiple anatomical regions to generate a unified diagnostic and damage assessment output for ankylosing spondylitis. Since different regions (such as the sacroiliac joint and left and right hip joints) exhibit different pathological manifestations, direct averaging would lead to an imbalance in weight distribution, affecting overall diagnostic stability. Therefore, this invention introduces an attention-weighted fusion strategy in the model's output stage, adaptively adjusting weights based on the importance of features from each region to achieve optimal fusion.
[0069] In addition, the system can optionally input patients' clinical indicators (CRP, ESR, HLA-B27, etc.) into the fusion layer to construct combined image-clinical features. This extension significantly improved the identification of early-stage ankylosing spondylitis cases in the validation set.
[0070] Example 5: Model Training and Optimization Module
[0071] Please see Figure 4 This module aims to improve the model's generalization ability and robustness. Before model training, all backbone networks are pre-trained on the ImageNet-1K dataset to achieve effective knowledge transfer; multi-label classification heads are randomly initialized to adapt to specific task requirements. During training, the learning rate is set to 0.00005, a warm-up strategy is adopted with a warm-up ratio of 0.1 to ensure stability in the early stages of training, and the AdamW optimizer is used to fine-tune the model parameters.
[0072] To further improve the reliability of model evaluation and reduce the impact of random fluctuations, a three-fold cross-validation strategy is adopted. The dataset is split according to the distribution of diagnostic labels to ensure that each fold reflects the statistical characteristics of the overall data. The final model performance is evaluated by combining the results of the three folds.
[0073] Based on the cross-validation results above, the optimal model structure and parameter configuration were selected, and the model was retrained using all the training data, keeping the training configuration unchanged. After training, the final model was validated on both the internal test set and a multi-center external test set.
[0074] All processes were completed on a high-performance computing platform. The software environment was based on Python and relied on key libraries such as torch 2.4.1, transformers 4.46.0, ultralytics 8.3.33, and scikit-learn 1.5.2 to ensure the stability of the training process and the reproducibility of the results.
[0075] Example 6: Output Module
[0076] Please see Figure 5 This module uses Grad-CAM to visualize the original and prior inputs label by label, revealing the image regions the model focuses on when making each label judgment. Simultaneously, T-SNE is used to reduce the fused features (before the classifier head) to 2D to observe the distribution and clustering of features among labels. The model output includes per-label predictions, confidence scores, Grad-CAM heatmaps, and comprehensive diagnostic conclusions based on multi-label fusion.
[0077] In summary, this invention has significant clinical implications and broad practical application prospects in the field of intelligent medical imaging diagnosis.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0079] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0080] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical coding feature maps; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-label automated diagnosis and assessment system for ankylosing spondylitis, characterized in that, include: The data acquisition and preprocessing module is used to acquire X-ray images of the patient's pelvis and preprocess the image data. The prior attention mechanism module is used to automatically identify key regions in the image by combining previously labeled information and to perform weighted processing on these regions. The multi-label classification deep learning model module uses deep learning technology to perform multi-label classification on pelvic X-ray images. The results fusion and evaluation module is used to fuse features from different image regions and generate a unified final diagnostic result using a weighted fusion strategy. The model training and optimization module is used to train deep learning models and optimize model performance using transfer learning and data augmentation techniques. The output module is used to generate a corresponding ankylosing spondylitis diagnosis report based on the model's diagnostic results, and to provide risk assessment results for joint damage.
2. The multi-label automatic diagnosis and assessment system for ankylosing spondylitis according to claim 1, characterized in that, The data acquisition and preprocessing module collects image data from medical institutions of different sources.
3. The multi-label automatic diagnosis and assessment system for ankylosing spondylitis according to claim 1, characterized in that, Image data preprocessing specifically includes: noise removal, brightness and contrast normalization, and image size scaling of the original image.
4. The multi-label automatic diagnosis and assessment system for ankylosing spondylitis according to claim 1, characterized in that, The system employs deep convolutional neural networks or object detection models to detect objects in images, automatically identify and label key regions, and dynamically adjust region weights based on previously labeled information from experts.
5. The multi-label automatic diagnosis and assessment system for ankylosing spondylitis according to claim 1, characterized in that, The multi-label classification deep learning model module employs the following techniques: A multi-label classification model based on a combination of deep convolutional neural networks (CNN) and visual transformers (ViT) is used to jointly train the model on lesions of different joints using a multi-task learning approach.
6. The multi-label automatic diagnosis and assessment system for ankylosing spondylitis according to claim 1, characterized in that, The model training and optimization module includes: Transfer learning techniques are used to transfer deep learning models that have been pre-trained on large-scale datasets to medical image datasets for fine-tuning. Furthermore, data augmentation techniques are used to increase the diversity of training samples, thereby enhancing the model's generalization ability and robustness.
7. The multi-label automatic diagnosis and assessment system for ankylosing spondylitis according to claim 1, characterized in that, It also includes a model interpretation module: The module uses algorithms such as Grad-CAM and T-SNE to perform visual analysis of the model's decision region, helping clinicians understand the decision-making basis of the AI model.
8. The multi-label automatic diagnosis and assessment system for ankylosing spondylitis according to claim 1, characterized in that, After preprocessing, the image is adjusted to 224*224 pixels.
9. The multi-label automatic diagnosis and assessment system for ankylosing spondylitis according to claim 1, characterized in that, Different task branches of the model share the underlying network weights and achieve information interaction between multiple labels through a joint feature fusion layer.
10. A diagnostic report generation method based on the system described in any one of claims 1-7: characterized in that, include: Acquiring and preprocessing X-ray images of the patient's pelvis; Multi-label classification training based on a deep learning model with prior attention mechanism; A final diagnostic report is generated through feature fusion and weighted evaluation. The model interpretation module generates visual interpretation results.