Multi-label medical image classification method
By combining image data and clinical risk data, using a univariate decision tree model to select important features and input them into the MedFusionNet model for classification training, the problem of low classification accuracy in the multi-label medical image classification is solved, and higher classification accuracy and better feature representation are achieved.
Patent Information
- Application Number
- CN202510270810.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
The existing multi-label medical image classification method based on CNN has the problem of low classification accuracy, and it is difficult to effectively capture and utilize the statistical dependence between labels.
A multi-label medical image classification method is adopted. By combining the image data in the training image set and the clinical risk data in the feature set, a univariate decision tree model is used to select the top-N feature set with the top-order ranking features from a single tag, and input it into the MedFusionNet model for classification training. The MedFusionNet model combines the DenseNet module, the Transformer module and the FPN module to extend multi-modal learning and enhance feature representation and fusion.
By extending the multi-modal learning of the MedFusionNet model, input features can be represented more abundantly, effectively solving the challenges of label dependence, data imbalance and modal integration, and improving the accuracy and interpretability of multi-label medical image classification.
Smart Images

Figure CN120107692A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical imaging technology, and in particular to a multi-label medical image classification method. Background Art
[0002] Multilabel image classification is a critical and challenging task in the field of medical imaging, where each image may be associated with multiple labels. Unlike single-label classification, where each image is assigned only one label, multi-label classification requires the model to understand and predict multiple simultaneous conditions or features from a single image. This complexity stems from the intricate dependencies between labels and the presence of data imbalance, which are common in medical datasets.
[0003] Traditional convolutional neural networks (CNNs) have been the core technology for image classification tasks due to their powerful feature extraction capabilities. However, they often fail to effectively capture and exploit statistical dependencies between labels, resulting in poor performance in multi-label classification tasks. Hybrid CNN-Transformer models often have limitations in interaction and information exchange, which restricts their classification capabilities.
[0004] In summary, the existing CNN-based multi-label medical image classification methods have the problem of low classification accuracy. Summary of the invention
[0005] In view of the above problems in the prior art, the present invention provides a multi-label medical image classification method, which solves the problem of low classification accuracy in the existing CNN-based multi-label medical image classification method.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A multi-label medical image classification method is provided, comprising the steps of:
[0008] S1. Obtain a training image set, a multi-label annotation set corresponding to the training image set, and a feature set, wherein each set of features in the feature set includes multiple features of clinical risk data of different categories.
[0009] S2. Input the feature set into the initial risk classification model. The initial risk classification model includes a univariate decision tree model. The univariate decision tree model preprocesses each feature of all feature sets, and ranks all features corresponding to each label in the multi-label annotation set by importance, selects N types of features with the highest importance for each label, and records multiple groups of N types of features corresponding to each label as top-N feature sets. The top-N feature sets are divided into top-N image feature sets and top-N non-image sets.
[0010] S3. Input the corresponding training image set, multi-label annotation set and top-N feature set into the MedFusionNet model for classification training; wherein the MedFusionNet model includes a DenseNet module, a Transformer module, an FPN module and a classifier. The DenseNet module is used to densely connect the local CNN features of each subset in the training image set and the top-N image set. The Transformer module is used to extract global features from each subset in the training image set and the top-N non-image set. The FPN module is used to extract fused features by fusing all local CNN features and global features of each subset through multi-scale and cross-mode fusion. The classifier is used to perform multi-label classification on the fused features of each subset.
[0011] S4. Evaluate and adjust the initial risk classification model and MedFusionNet model through the verification image set, multi-label verification set and feature verification set until the MedFusionNet model reaches the evaluation index.
[0012] S5. Obtain the patient's medical imaging pictures to be classified and clinical feature data, and obtain the multi-label classification results of the medical imaging pictures through the initial risk classification model and the MedFusionNet model.
[0013] The beneficial effects of this scheme are: combining the image data in the training image set with the clinical risk data in the feature set, and selecting the top-ranked features from a single label through a univariate decision tree model to form a top-N feature set, which is then input into the MedFusionNet model, expanding the multimodal learning of the MedFusionNet model, so that the input features can be represented more richly. The MedFusionNet model combines the DenseNet module, the Transformer module, and the FPN module. The dense connections in the DenseNet module ensure effective feature propagation and gradient flow, the Transformer module captures the dependencies between image regions, labels, and clinical risk data, and the FPN module provides cross-modal multi-scale feature representation and fusion. The proposed method for different data sets outperforms existing models in key indicators, effectively solving the challenges of label dependency, data imbalance, and modality integration, while enhancing interpretability and prediction capabilities and improving classification accuracy.
[0014] Furthermore, the preprocessing steps of the univariate decision tree model for the feature set include: binning, filling and CNN processing for the continuous variables, missing data and image data of the feature set. Binning can simplify the classification of continuous variables (such as age) and enhance the ability to capture nonlinear relationships. Filling and CNN processing combined with the dense connection of DenseNet can improve the efficiency of gradient propagation and reduce feature loss.
[0015] Furthermore, the decision-making method of the univariate decision tree model is: for multiple features of the same category, the best classification threshold of the feature is determined through grid search. When the value of a feature exceeds the best classification threshold determined by its category, the feature is determined to belong to the corresponding label, otherwise it is determined that the feature does not belong to the corresponding label. Determining the best classification threshold through grid search can improve the classification accuracy of a single label.
[0016] Furthermore, the initial risk classification model also includes a multivariate classification tree model and a risk assessment model for multivariate classification training of each group of N types of features through the top-N feature set and the corresponding multi-label annotation set; the multivariate classification tree model is used to perform multivariate classification of clinical feature data, and the risk assessment model is used to determine whether the patient corresponding to the clinical feature data has multiple diseases based on the multivariate classification results. The multivariate classification tree model and the risk assessment model can directly determine whether the patient's clinical risk data has the risk of multiple diseases. If the risk is high, the MedFusionNet model can be used for further analysis.
[0017] Furthermore, the training steps of the MedFusionNet model include:
[0018] S3.1. Initialize the parameters of the DenseNet module, Transformer module, FPN module and classifier;
[0019] S3.2. Divide the training image set, multi-label annotation set and top-N feature set into multiple rounds for training. Each round includes multiple training batches. For all data in each training batch, forward propagation, loss calculation, back propagation and parameter update are performed in sequence until all rounds are traversed.
[0020] Furthermore, the loss calculation function in step S3.2 is:
[0021]
[0022] Among them, L is the loss function; M is the total number of training images in the training image set; C is the total number of labels for each group of multi-labels in the multi-label annotation set; i and c are variables, y ic is the i-th group of multi-labels in the multi-label annotation set, The multi-label predicted for the i-th training image in the training image set predicted by the MedFusionNet model.
[0023] Furthermore, the evaluation indicators of the MedFusionNet model include execution time, classification accuracy, F score Value, Adaptive Rand Error, Adaptive Rand Precision, and Adaptive Rand Recall.
[0024] Furthermore, the calculation formula for classification accuracy is:
[0025]
[0026] Among them, C Acc is the classification accuracy; TP is the number of samples correctly predicted as positive; TN is the number of samples correctly predicted as negative; FP is the number of samples incorrectly predicted as positive; FN is the number of samples incorrectly predicted as negative.
[0027] Furthermore, F score The value is calculated as:
[0028]
[0029] Among them, F score It is an indicator used to measure the performance of the classification model; CR is the accuracy rate; CM is the completeness rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flowchart of a multi-label medical image classification method;
[0031] Figure 2 is a sample of the multi-label chest x-ray dataset;
[0032] Figure 3 is the accuracy on the NIH-Chest-Xray dataset, F score Schematic diagram of the comparison results with GMean;
[0033] Figure 4 It is the histogram of the distribution of each feature in the cervical cancer dataset;
[0034] Figure 5 This is the correlation heat map between various features in the cervical cancer dataset;
[0035] Figure 6 This is the confusion matrix result of MedFusionNet on the cervical cancer dataset;
[0036] Figure 7 is the accuracy of multiple models on the cervical cancer dataset, F score Schematic diagram of the comparison results with GMean; DETAILED DESCRIPTION
[0037] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0038] Example 1
[0039] Multi-label image classification faces challenges due to complex label dependencies, data imbalance, and the integration of multiple data modes. Traditional convolutional neural networks are difficult to capture the statistical dependencies between labels, and hybrid CNN-transformer models are often affected by limited interactions and information exchange, resulting in the existing CNN-based multi-label medical image classification methods having low classification accuracy. In order to solve this problem, this application provides a multi-label medical image classification method, which combines the image data in the training image set with the clinical risk data in the feature set, and selects the top-ranked features from a single label through a univariate decision tree model to form a top-N feature set, which is then input into the MedFusionNet model, expanding the multi-modal learning of the MedFusionNet model, so that the input features can be represented more richly. The MedFusionNet model combines the DenseNet module, the Transformer module, and the FPN module. The dense connections in the DenseNet module ensure effective feature propagation and gradient flow, the Transformer module captures the dependencies between image regions, labels, and clinical risk data, and the FPN module provides multi-scale feature representation and fusion across modes. The proposed methods for different datasets outperform existing models in key indicators, effectively solving the challenges of label dependency, data imbalance, and modality integration, while enhancing interpretability and predictive capabilities, achieving the technical effect of improving classification accuracy.
[0040] refer to Figure 1 , a multi-label medical image classification method, comprising the steps of:
[0041] S1. Obtain a training image set, a multi-label annotation set corresponding to the training image set, and a feature set, wherein each set of features in the feature set includes multiple features of clinical risk data of different categories.
[0042] Each set of clinical risk data includes one or more of clinical text data, metadata and image data, constituting a multimodal input. Clinical risk data may include age, sexual behavior, lifestyle, medical history, etc.
[0043] S2. Input the feature set into the initial risk classification model. The initial risk classification model includes a univariate decision tree model. The univariate decision tree model preprocesses each feature of all feature sets, and ranks the importance of all features corresponding to each label in the multi-label annotation set according to indicators such as information gain, Gini index or accuracy, and selects N types of features with the top importance ranking for each label. Multiple groups of N types of features corresponding to each label are recorded as top-N feature sets. The top-N feature sets are divided into top-N image feature sets and top-N non-image sets.
[0044] The preprocessing steps of the univariate decision tree model for the feature set include: binning, filling and CNN processing of the continuous variables, missing data and image data of the feature set. Missing data includes time series data and non-time series data. Time series data is filled by the forward filling method, and non-time series data is filled by the median filling method. Discretizing continuous variables (such as age) through binning can simplify classification and enhance the ability to capture nonlinear relationships. Filling and CNN processing combined with dense connections of DenseNet can improve the efficiency of gradient propagation and reduce feature loss.
[0045] The decision-making method of the univariate decision tree model is: for multiple features of the same category, the best classification threshold of the feature is determined through grid search. When the value of a feature exceeds the best classification threshold determined by its category, the feature is judged to belong to the corresponding label, otherwise it is judged that the feature does not belong to the corresponding label. Determining the best classification threshold through grid search can improve the classification accuracy of a single label.
[0046] S3. Input the corresponding training image set, multi-label annotation set and top-N feature set into the MedFusionNet model (Medical Fusion Network) for classification training; wherein the MedFusionNet model includes a DenseNet module, a Transformer module, an FPN module and a classifier. The DenseNet module is used to densely connect the local CNN features of each subset in the training image set and the top-N image set. The Transformer module is used to extract global features from each subset in the training image set and the top-N non-image set. The FPN module is used to extract fusion features by fusing all local CNN features and global features of each subset through multi-scale and cross-mode fusion. The classifier is used to perform multi-label classification on the fusion features of each subset.
[0047] In this embodiment, the full name of the DenseNet module is Densely Connected Convolutional Networks, which is a convolutional neural network architecture. Its core idea is to enhance feature reuse and information flow through dense connections (Dense Connection) based on CNN. The Transformer module is a deep learning model based on the self-attention mechanism (Self-Attention). The full name of the FPN module is Feature Pyramid Network, which is a feature pyramid network. By constructing a feature pyramid and fusing feature maps of different scales, the model's detection ability for multi-scale targets is improved. It combines the semantic information of high-level features and the detailed information of low-level features. The MedFusionNet model is also provided with cross-branch interaction modules—C2T (Convolutional to Transformer) and T2C (Transformer to Convolutional)—for promoting information exchange and interaction between CNN and Transformer modules in the DenseNet module.
[0048] S4. Evaluate and adjust the initial risk classification model and MedFusionNet model through the validation image set, multi-label validation set, and feature validation set until the MedFusionNet model reaches the evaluation index. The evaluation index of the MedFusionNet model includes execution time, classification accuracy, F score Value, Adaptive Rand Error, Adaptive Rand Precision, and Adaptive Rand Recall.
[0049] The evaluation indicators of the MedFusionNet model include execution time, classification accuracy, F score Value, Adaptive Rand Error, Adaptive Rand Precision, and Adaptive Rand Recall.
[0050] The calculation formula for classification accuracy is:
[0051]
[0052] Among them, C Acc is the classification accuracy; TP is the number of samples correctly predicted as positive; TN is the number of samples correctly predicted as negative; FP is the number of samples incorrectly predicted as positive; FN is the number of samples incorrectly predicted as negative.
[0053] F score The value is calculated as:
[0054]
[0055] Among them, F score It is an indicator used to measure the performance of the classification model; CR is the accuracy rate; CM is the completeness rate.
[0056] S5. Obtain the patient's medical imaging pictures to be classified and clinical feature data, and obtain the multi-label classification results of the medical imaging pictures through the initial risk classification model and the MedFusionNet model.
[0057] The initial risk classification model also includes a multivariate classification tree model and a risk assessment model for multivariate classification training of each group of N types of features through the top-N feature set and the corresponding multi-label annotation set. The multivariate classification tree model is used to perform multivariate classification of clinical feature data, and the risk assessment model is used to determine whether the patient corresponding to the clinical feature data has multiple diseases based on the multivariate classification results. The multivariate classification tree model and the risk assessment model can directly determine whether the patient's clinical risk data has the risk of multiple diseases. If the risk is high, the MedFusionNet model can be used for further analysis.
[0058] In this embodiment, the multivariate classification tree defines the thresholds and rules for classifying individuals into risk categories:
[0059] Low risk: No significant behavioral or diagnostic risk factors are present.
[0060] Moderate risk: Some risk factors are present, such as a history of smoking or a positive HPV diagnosis.
[0061] High risk: multiple risk factors or severe diagnoses, such as a cancer diagnosis or multiple sexually transmitted diseases. The model achieves robust performance indicators, making it suitable for identifying individuals at different levels of health risk and enabling targeted interventions.
[0062] Specifically, the training steps of the MedFusionNet model include:
[0063] S3.1. Initialize the parameters of the DenseNet module, Transformer module, FPN module and classifier, respectively. CNN ,θ Transformer ,θ Interaction ,θ Classify .
[0064] S3.2. Divide the training image set, multi-label annotation set and top-N feature set into multiple rounds for training. Each round includes multiple training batches. For all data in each training batch, forward propagation, loss calculation, back propagation and parameter update are performed in sequence until all rounds are traversed.
[0065] The loss is calculated as:
[0066]
[0067] Among them, L is the loss function; M is the total number of training images in the training image set; C is the total number of labels for each group of multi-labels in the multi-label annotation set; i and c are variables, y ic is the i-th group of multi-labels in the multi-label annotation set, The multi-label predicted for the i-th training image in the training image set predicted by the MedFusionNet model.
[0068] Among them, θ CNN ,θ Transformer ,θ Interaction ,θ Classify The relationship between back propagation and parameter update is:
[0069]
[0070]
[0071] Among them, α is the learning rate, which controls the step size along the negative gradient direction; and They are the gradients of the DenseNet module, Transformer module, FPN module and classifier respectively, and L is the loss function.
[0072] Example 2
[0073] This embodiment is further limited on the basis of Embodiment 1. The specific improvement lies in the specific use of data sets to evaluate the effectiveness of the multi-label medical image classification method. For other parts not mentioned, refer to Embodiment 1 or the prior art.
[0074] As a dataset for this implementation, a multi-label chest x-ray dataset is used, referring to Figure 2The multi-label chest x-ray dataset is the NIH ChestX-ray14 dataset. The NIH ChestX-ray14 dataset is an extended version of the ChestX-ray8 dataset, containing 112,120 frontal X-ray images from 30,805 different patients and annotated with 14 diseases. The size of each image is 1024×1024 pixels, and the data comes from a specific patient population. The 14 diseases included in the NIH ChestX-ray14 dataset are Atelectasis, Cardiomegaly, Consolidation, Edema, Effusion, Emphysema, Fibrosis, Hernia, Infiltration, Mass, Nodule, Pleural Thickening, Pneumonia, and Pneumothorax.
[0075] To evaluate the performance, the experiments were conducted using Python 3.7 and a 2.11 GHz system. Core TM i7-8650U and 16GB RAM. The MedFusionNet model was compared with 5 classification network models (RestNet50, DenseNet121, ConvNeXt, DeiT, CTransCNN) for multi-label medical image classification. As shown in Table 1, for the NIH-ChestXray dataset, the MedFusionNet model achieved the highest accuracy score of 95.35%, outperforming the other models. DenseNet121 had the lowest accuracy of 65.09%, followed by ConvNeXt at 72.34%. Robust analysis showed that models with different architectures exhibited different performance levels in different disease classification tasks.
[0076] Table 1 NIH ChestX-ray14 classification results
[0077]
[0078] F score It is the harmonic mean of precision and recall, an indicator used to measure the performance of classification models. It is used to comprehensively evaluate the performance of classification models, especially when the data is unbalanced. Gmean is the geometric mean.
[0079] Further, refer to Figure 3 , Figure 3The following is a diagram showing the accuracy, F-Score and GMean comparison results of RestNet50, DenseNet121, ConvNeXt, DeiT, CTransCNN and MedFusionNet on the NIH-Chest-Xray dataset. Figure 3 As shown in Figure 3, the MedFusionNet model performs better on the NIH-ChestXray dataset compared to other methods.
[0080] As another scheme of this embodiment, a cervical cancer data set is used. The cervical cancer data set contains features related to cervical cancer risk factors, wherein each row represents a data sample of a subject, and each column represents a feature. These features include age (Age), number of sexual behaviors (Number of diagnosis), cancer (Cancer), cervical intraepithelial neoplasia (CIN), human papillomavirus (HPV), Hinselmann examination, Schiller examination, cytology (Citology), and biopsy (Biopsy). The Hinselmann examination is a method of observing changes in the cervix by acetic acid staining, and the Schiller examination is a method of staining the cervix with iodine solution to detect abnormal cells. The cervical cancer data set consists of continuous values and discrete values. The data set has a total of 858 samples, each of which contains 36 features and a result to indicate whether an individual has cervical cancer. Figure 4 A histogram of each numerical feature is shown to show its distribution; Figure 5 It shows the correlation heat map of the numerical features in the dataset.
[0081] The proposed MedFusionNet model was compared with the cervical cancer dataset. Table 2 shows the comparison results of MedFusionNet with ResNet50, DenseNet121, ConvNeXt, DeiT and CTransCNN. The experimental results show that the proposed MedFusionNet model performs better than other comparison algorithms.
[0082] Table 1 Classification results of cervical cancer dataset
[0083]
[0084] Figure 6 The confusion matrix results of MedFusionNet on the cervical cancer dataset are shown. In addition, Figure 7The comparison results of accuracy, F-Score and GMean of ResNet50, DenseNet121, ConvNeXt, DeiT, CTransCNN and MedFusionNet on the cervical cancer dataset are shown. Figure 7 It is shown that the proposed MedFusionNet achieves better results on the cervical cancer dataset compared with other methods.
[0085] As a further solution of this embodiment, the Friedman test in the statistical analysis method is used to determine the performance of six machine learning models - ResNet50, DenseNet121, ConvNeXt, DeiT, CTransCNN and MedFusionNet in the classification task, and their accuracy, F score and Gmean.
[0086] Tables 3 and 4 show the results of Friedman rank sum test on the NIHChestX-ray dataset and cervical cancer dataset, respectively. The results in Tables 3 and 4 provide a detailed comparison of the performance of the models on two medical imaging datasets (NIHChestX-ray dataset and cancer dataset) based on rank sum. The rank sum method highlights the relative effectiveness of each model by assigning a rank based on performance, where a higher rank sum value indicates a more robust result. In both datasets, MedFusionNet achieved the highest rank sum value (24 for the NIHChestX-ray dataset and 30 for the cervical cancer dataset), indicating its superior performance in handling the complexity of medical imaging data. This suggests that MedFusionNet, with its advanced fusion and feature extraction capabilities, is particularly suitable for handling subtle tasks in medical datasets and may help achieve more accurate diagnoses. In contrast, models such as DeiT consistently perform poorly (4 for the NIHChestX-ray dataset and 5 for the cancer dataset), which may be limited in these applications. DenseNet121 and InceptionResNet show moderate to high rank sums, highlighting their relative advantages, but still falling short of the overall performance of MedFusionNet. The statistically significant differences observed in the rank sums emphasize the robustness of MedFusionNet across different datasets, making it a potential model of choice in the field of medical image analysis due to its consistently high level of performance in multiple complex medical imaging tasks.
[0087] Table 3 Rank sum of NIH ChestX-ray14 in each method
[0088] method Rank Sum ResNet50 12 DenseNet121 20 ConvNeXt 8 Dei 4 CTransCNN 16 MedFusionNet 24 P-value 0.0012497
[0089] Table 4. The rank sum of the cervical cancer dataset in each method
[0090] method Rank Sum ResNet50 15 DenseNet121 20 ConvNeXt 10 Dei 5 InceptionResNet 25 MedFusionNet 30 P-value 0.0001393
[0091] In summary, the beneficial effects of this scheme are summarized as follows:
[0092] 1. Risk stratification based on hybrid multivariate model: A two-stage risk stratification framework is introduced. First, univariate thresholds are applied to identify the Top-N risk features for each label, and then these features are integrated into a multivariate model using the MedFusionNet model. This process effectively bridges the gap between feature priority and comprehensive label relevance analysis, while also being able to handle multimodal inputs such as clinical text data and metadata.
[0093] 2. Parallel hybrid architecture for multi-label medical image classification: MedFusionNet, a novel parallel hybrid architecture that combines CNN and Transformer components for multimodal learning. The CNN branch uses dense connections (DenseNet) to ensure efficient feature propagation and alleviate the gradient vanishing problem. At the same time, the Transformer module branch uses the self-attention mechanism to capture the complex dependencies between image regions, labels, and modalities. In addition, the Feature Pyramid Network (FPN) is introduced for multi-scale feature representation and fusion, enabling MedFusionNet to effectively integrate images, text, and metadata, thereby enhancing the performance of multi-label medical classification.
[0094] 3. Enhanced cross-branch interaction: In order to improve the nonlinearity, representation ability and multimodal integration ability of the model, cross-branch interaction modules between CNN and Transformer are introduced. These modules promote information exchange between feature representation and modalities, thereby enhancing the exploration of implicit correlations between labels and improving overall classification accuracy.
[0095] 4. Extensive evaluation and performance analysis: MedFusionNet was comprehensively evaluated using two different datasets: NIH ChestX-ray14 and a self-built cervical cancer dataset with clinical text annotations. Experimental results show that MedFusionNet combined with a risk stratification framework outperforms existing state-of-the-art methods in multi-label classification tasks. This demonstrates the effectiveness of the invention in addressing the challenges of label dependency, data imbalance, and multimodal integration, while demonstrating strong generalization capabilities in diverse datasets and tasks.
[0096] Although the specific implementation of the invention is described in detail in conjunction with the drawings, it should not be understood as limiting the scope of protection of this patent. Within the scope described in the claims, various modifications and variations that can be made by those skilled in the art without creative work still fall within the scope of protection of this patent.
Claims
1. A multi-label medical image classification method, characterized in that: Includes steps: S1. Obtain a training image set, a multi-label annotation set corresponding to the training image set, and a feature set, wherein each set of features in the feature set includes multiple features of clinical risk data of different categories; S2. Input the feature set into the initial risk classification model, which includes a univariate decision tree model. The univariate decision tree model preprocesses each feature of all feature sets, and ranks the importance of all features corresponding to each label in the multi-label annotation set, selects N types of features with the highest importance for each label, and records multiple groups of N types of features corresponding to each label as top-N feature sets. The top-N feature sets are divided into top-N image feature sets and top-N non-image sets. S3, inputting the corresponding training image set, multi-label annotation set and top-N feature set into the MedFusionNet model for classification training; wherein the MedFusionNet model includes a DenseNet module, a Transformer module, an FPN module and a classifier, the DenseNet module is used to densely connect the local CNN features of each subset in the training image set and the top-N image set, the Transformer module is used to extract global features from each subset in the training image set and the top-N non-image set; the FPN module is used to extract fusion features by fusing all local CNN features and global features of each subset through multi-scale and cross-mode; The classifier is used to perform multi-label classification on the fused features of each subset; S4. Evaluate and adjust the initial risk classification model and MedFusionNet model through the verification image set, multi-label verification set and feature verification set until the MedFusionNet model reaches the evaluation index; S5. Obtain the patient's medical imaging pictures to be classified and clinical feature data, and obtain the multi-label classification results of the medical imaging pictures through the initial risk classification model and the MedFusionNet model.
2. The multi-label medical image classification method according to claim 1, characterized in that: The preprocessing steps of the univariate decision tree model for the feature set include: binning, filling and CNN processing of the continuous variables, missing data and image data of the feature set.
3. The multi-label medical image classification method according to claim 1, characterized in that: The decision-making method of the univariate decision tree model is: for multiple features of the same category, the optimal classification threshold of the feature is determined through grid search. When the value of a feature exceeds the optimal classification threshold determined by its category, the feature is judged to belong to the corresponding label, otherwise the feature is judged not to belong to the corresponding label.
4. The multi-label medical image classification method according to claim 1, characterized in that: The initial risk classification model also includes a multivariate classification tree model and a risk assessment model for performing multivariate classification training on each group of N types of features through a top-N feature set and a corresponding multi-label annotation set; the multivariate classification tree model is used to perform multivariate classification of clinical feature data, and the risk assessment model is used to determine whether a patient corresponding to the clinical feature data has multiple diseases based on the multivariate classification results.
5. The multi-label medical image classification method according to claim 1, characterized in that: The training steps of the MedFusionNet model include: S3.
1. Initialize the parameters of the DenseNet module, Transformer module, FPN module and classifier; S3.
2. Divide the training image set, multi-label annotation set and top-N feature set into multiple rounds for training. Each round includes multiple training batches. For all data in each training batch, forward propagation, loss calculation, back propagation and parameter update are performed in sequence until all rounds are traversed.
6. The multi-label medical image classification method according to claim 5, characterized in that: The loss calculation function in step S3.2 is: Among them, L is the loss function; M is the total number of training images in the training image set; C is the total number of labels for each group of multi-labels in the multi-label annotation set; i and c are variables, y ic is the i-th group of multi-labels in the multi-label annotation set, The multi-label predicted for the i-th training image in the training image set predicted by the MedFusionNet model.
7. The multi-label medical image classification method according to claim 1, characterized in that: The evaluation indicators of the MedFusionNet model include execution time, classification accuracy, F score Value, Adaptive Rand Error, Adaptive Rand Precision, and Adaptive Rand Recall.
8. The multi-label medical image classification method according to claim 7, characterized in that: The calculation formula for classification accuracy is: Among them, C Acc is the classification accuracy; TP is the number of samples correctly predicted as positive; TN is the number of samples correctly predicted as negative; FP is the number of samples incorrectly predicted as positive; FN is the number of samples incorrectly predicted as negative.
9. The multi-label medical image classification method according to claim 7, characterized in that: F score The value is calculated as: Among them, F score It is an indicator used to measure the performance of the classification model; CR is the accuracy rate; CM is the completeness rate.
Citation Information
Cited By
Thyroid nodule clinical feature filling and benign and malignant prediction optimization method and device
CN120727291A
Method and device for filling in clinical features of thyroid nodules and optimizing prediction of benignity and malignancy
CN120727291B