Multi-stage outcome prediction method based on multi-classification-head cascade of convolutional neural network

By employing a multi-class head cascade method using convolutional neural networks and Focal loss training, the problem of lacking temporal information in existing technologies is solved, enabling efficient prediction of multi-stage outcome probabilities from a single CT image and improving the clinical management capabilities of lung diseases.

CN120931992APending Publication Date: 2025-11-11XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510986439.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies lack in-depth modeling of the changing trends of patients' conditions across time scales in the clinical management of lung diseases, making it difficult to meet the needs of personalized follow-up and intervention. Furthermore, the lack of guidance from time-series information results in insufficient predictive ability for multi-stage outcomes.

Method used

A multi-classification head cascade method based on convolutional neural networks was designed. By setting three classification heads, the short-term, medium-term and long-term regression probabilities are predicted respectively. Focal loss is used to construct a loss function, and the convolutional neural network is trained to guide the temporal information and output multi-period regression probabilities.

Benefits of technology

It enables the simultaneous output of high-precision short-term, medium-term, and long-term outcome probabilities from a single acute-phase CT image, improving the accuracy and efficiency of multi-phase outcome prediction, and conforming to the gradual progression of short-term, medium-term, and long-term outcomes in clinical practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931992A_ABST
    Figure CN120931992A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and medical image processing, in particular to a multi-stage regression prediction method based on convolutional neural network multi-classification head cascading, which comprises the following steps: setting a first classification head for predicting a short-term regression probability; setting a second classification head used for predicting a medium-term outcome probability; and setting a third classification head used for predicting the long-term outcome probability. According to the method, three classification heads corresponding to short-term, middle-term and long-term prediction tasks are designed, and classification intermediate features of a previous period are fused in a cascade mode to serve as auxiliary input of a next period, so that guidance of time sequence information is realized, and the constructed convolutional neural network can be used for realizing accurate prediction of a CT image in a single acute period through a CT image in a single acute period. And high-precision rotation probability of three time points can be output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and medical image processing technology, specifically to a multi-stage outcome prediction method based on a convolutional neural network multi-classifier head cascade. Background Technology

[0002] In the clinical management of lung diseases (such as severe pneumonia and pulmonary fibrosis), physicians often rely on imaging information to determine the disease progression and develop follow-up plans. Traditional prediction methods are usually based on assessments at a single point in time and lack in-depth modeling of the changing trends of the patient's condition across time scales, making it difficult to meet the needs of personalized follow-up and intervention.

[0003] In recent years, deep learning has been widely used in medical image diagnosis. However, most current models focus on single-stage tasks, requiring prediction of short-term, medium-term, and long-term outcomes one by one. At the same time, the outcome states at different time points have a certain temporal continuity and logical consistency. Currently, there is a lack of guidance on temporal information, resulting in insufficient ability to predict the multi-stage outcomes of patients. Summary of the Invention

[0004] The purpose of this invention is to provide a multi-period outcome prediction method based on convolutional neural network multi-classification head cascade, in order to solve the technical problem that most current models focus on single-period tasks and lack guidance from time-series information.

[0005] To solve the above-mentioned technical problems, the present invention specifically provides the following technical solution: A multi-stage outcome prediction method based on convolutional neural network multi-classifier head concatenation includes the following steps: Acquire chest CT images of acute lung diseases and perform standardized resampling, cropping and intensity normalization on the chest CT images; Set up the first classification head to predict the short-term outcome probability based on high-dimensional feature images in chest CT images; Set up a second classifier to predict the intermediate outcome probability based on the high-dimensional feature images in chest CT images and the intermediate classification features of the first classifier; A third classifier is set up to predict the long-term outcome probability based on high-dimensional feature images in chest CT images, intermediate classification features of the first classifier, and intermediate classification features of the second classifier. A first, second, and third classifier head are trained using a loss function constructed with Focal loss to form a convolutional neural network for predicting short-term, medium-term, and long-term outcome probabilities in the chest CT images; the convolutional neural network is then used to output the short-term, medium-term, and long-term outcome probabilities of the patient based on the patient's chest CT images.

[0006] As a preferred embodiment of the present invention, the high-dimensional image features It was extracted from chest CT images by the encoder.

[0007] As a preferred embodiment of the present invention, the first classification head, the second classification head, and the third classification head are each composed of two linear layers and a softmax activation layer.

[0008] As a preferred embodiment of the present invention, the first classification head is used to classify high-dimensional image features. Given the input, output the short-term outcome probability; The structural expression of the first classification head is: ; ; In the formula, The short-term regression probability is the output of the first classifier. , These are the first and second linear layers in the classification head, respectively. High-dimensional image features Intermediate classification features generated through the first linear layer of the first classification head. This is the softmax activation layer.

[0009] As a preferred embodiment of the present invention, the second classification head is used to classify high-dimensional image features. and the intermediate classification features Compositional fusion characteristics Given the input, output the intermediate outcome probability; The structural expression for the second classification head is: ; ; In the formula, This represents the intermediate regression probability output by the second classification head. , These are the first and second linear layers in the classification head, respectively. For fusion features Intermediate classification features generated through the first linear layer of the second classification head. This is a softmax activation layer, and Concat is the feature concatenation operator.

[0010] As a preferred embodiment of the present invention, the third classification head is used to classify high-dimensional image features. The intermediate classification features and the intermediate classification features Compositional fusion characteristics Given the input, output the long-term outcome probability; The structural expression of the third classification head is: ; In the formula, This represents the long-term regress probability output by the third classification head. , These are the first and second linear layers in the classification head, respectively. This is a softmax activation layer, and Concat is the feature concatenation operator.

[0011] As a preferred embodiment of the present invention, the process of constructing the loss function includes: The short-term outcome prediction loss of the first classification head is established using Focal loss. ,in, In the formula, As the gold standard for short-term reversion, This represents the positive probability output by the first classification head. This represents the negative probability output by the first classification head. An adjustment factor to balance the contribution of difficult and easy samples to the loss; The short-term outcome prediction loss of the second classification head is built using Focal loss. ,in, In the formula, As the gold standard for mid-term reversion, The positive probability output by the second classification head. The negative probability output by the second classification head; Short-term outcome prediction loss for a third classification head is built using Focal loss. ,in, In the formula, As the gold standard for long-term outcomes, The positive probability output by the third classification head. The negative probability output by the third classification head; Short-term reversion to predict loss Interim outcome prediction loss and long-term outcome prediction loss The combination is the total loss. ,in, .

[0012] As a preferred embodiment of the present invention, the Adam optimizer is used in training the convolutional neural network to dynamically calculate the short-term, medium-term, and long-term total losses during the training process. Then, backpropagation is performed until the convolutional neural network achieves the best accuracy on the validation set, thus obtaining the optimal convolutional neural network.

[0013] Compared with the prior art, the present invention has the following advantages: This invention designs three classification heads corresponding to short-term (1 month), medium-term (3 months) and long-term (6 months) prediction tasks, and fuses intermediate classification features from the previous period in a cascade manner as auxiliary inputs for the next period to guide the temporal information. This enables the constructed convolutional neural network to output high-precision outcome probabilities for the three time points using a single acute-phase CT image. Attached Figure Description

[0014] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0015] Figure 1 The flowchart of the multi-stage regression prediction method based on convolutional neural network multi-class head cascade provided in the embodiments of the present invention is as follows: Figure 2 A diagram of a convolutional neural network structure provided in an embodiment of the present invention; Figure 3 Performance evaluation of multi-class head cascading provided in embodiments of the present invention; Figure 4 Effectiveness evaluation of convolutional neural networks provided in embodiments of the present invention; Figure 5 This is a structural diagram of Wo_Both provided in an embodiment of the present invention; Figure 6 This is a structural diagram of an encoder provided in an embodiment of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] like Figure 1 As shown, this invention provides a multi-period outcome prediction method based on a convolutional neural network multi-classification head cascade, comprising the following steps: Acquire chest CT images of acute lung diseases and perform standardized resampling, cropping and intensity normalization on the chest CT images; Set up the first classification head to predict the short-term outcome probability based on high-dimensional feature images in chest CT images; Set up a second classifier to predict the intermediate outcome probability based on the high-dimensional feature images in chest CT images and the intermediate classification features of the first classifier; A third classifier is set up to predict the long-term outcome probability based on high-dimensional feature images in chest CT images, intermediate classification features of the first classifier, and intermediate classification features of the second classifier. The first, second, and third classifiers were trained using a loss function constructed with Focal loss to form a convolutional neural network for predicting short-term, medium-term, and long-term outcome probabilities in chest CT images. Using a convolutional neural network, the short-term, medium-term, and long-term probabilities of a patient's outcome are output based on the patient's chest CT images.

[0018] High-dimensional image features Extracted from chest CT images by an encoder; the encoder is used to extract high-dimensional image features from chest CT images. The encoder uses a ResNet encoder, such as Figure 6 As shown, other types of encoders can also be used, among which the ResNet encoder has the best performance. It consists of an input convolutional layer, multiple residual blocks, an average pooling layer, and a linear transform layer. The structure expression of the encoder is as follows: ; In the formula, Features of high-dimensional images Chest CT image, It is a ResNet encoder.

[0019] In this invention, the first classification head is based solely on high-dimensional image features. The output transition probability is used. The second and third classifiers introduce the intermediate features of the first and the first two classifiers respectively and then combine them as auxiliary information to participate in the prediction. Through the additional design of the cascaded information transmission path between the three classifiers, the classifiers in later periods (medium and long term) can incorporate the information of the earlier periods when making predictions.

[0020] Compared to the first classification head, the second and third classification heads, corresponding to the March and June periods respectively, not only include the image features extracted by the encoder but also incorporate the intermediate classification features from their respective preceding classification heads. Taking the third classification head corresponding to the June period as an example, its input is a concatenation of image features and the intermediate classification features from the first and second classification heads. This cascading classification feature fusion strategy indirectly introduces short-, medium-, and long-term temporal prior information into the classification head's learning target mapping, thus allowing the longer-term classification head to refer to the shorter-term classification response when making decisions. This also aligns with the gradual progression of short-, medium-, and long-term outcomes in clinical practice.

[0021] The first, second, and third classification heads each consist of two linear layers and a softmax activation layer.

[0022] The first classification head is used to classify high-dimensional image features. Given the input, output the short-term outcome probability; The structural expression for the first classification head is: ; ; In the formula, The short-term regression probability is the output of the first classifier. , These are the first and second linear layers in the classification head, respectively. High-dimensional image features Intermediate classification features generated through the first linear layer of the first classification head. This is the softmax activation layer.

[0023] The second classification head is used to classify high-dimensional image features. and intermediate classification features Compositional fusion characteristics Given the input, output the intermediate outcome probability; The structural expression for the second classification head is: ; ; In the formula, This represents the intermediate regression probability output by the second classification head. , These are the first and second linear layers in the classification head, respectively. For fusion features Intermediate classification features generated through the first linear layer of the second classification head. This is a softmax activation layer, and Concat is the feature concatenation operator.

[0024] The third classification head is used to classify high-dimensional image features. Intermediate classification features and intermediate classification features Compositional fusion characteristics Given the input, output the long-term outcome probability; The structural expression for the third classification head is: ; In the formula, This represents the long-term regress probability output by the third classification head. , These are the first and second linear layers in the classification head, respectively. This is a softmax activation layer, and Concat is the feature concatenation operator.

[0025] Focal Loss is a loss function specifically designed to address class imbalance. By dynamically adjusting the weights of easily classified and difficult-to-classify samples, it makes the model focus more on difficult-to-classify samples, thereby improving the model's performance in object detection and rare class identification. This invention uses Focal Loss to construct the model's training loss, improving the model's predictive performance. This total loss... Composed of three sub-loss items The components correspond to the first, third, and sixth months, respectively, as shown in the following formula, where each sub-loss term includes the Focal Loss constraint for the corresponding period.

[0026] The process of constructing the loss function includes: The short-term outcome prediction loss of the first classification head is established using Focal loss. ,in, In the formula, As the gold standard for short-term reversion, This represents the positive probability output by the first classification head. This represents the negative probability output by the first classification head. An adjustment factor to balance the contribution of difficult and easy samples to the loss; The short-term outcome prediction loss of the second classification head is built using Focal loss. ,in, In the formula, As the gold standard for mid-term reversion, The positive probability output by the second classification head. The negative probability output by the second classification head; Short-term outcome prediction loss for a third classification head is built using Focal loss. ,in, In the formula, As the gold standard for long-term outcomes, The positive probability output by the third classification head. The negative probability output by the third classification head; Short-term reversion to predict loss Interim outcome prediction loss and long-term outcome prediction loss The combination is the total loss. ,in, .

[0027] To balance the contribution of difficult and easy samples to the loss, the adjustment factor is empirically set to 2, which gives the model the ability to adaptively focus on difficult samples during iterative updates.

[0028] Among them, when This indicates that the positive result has not turned into a positive outcome. This indicates that the negative result has been confirmed. middle Used to constrain the loss of positive samples in the current classification head. Used to constrain the loss of negative samples in the current classification head. .

[0029] This invention performs ablation analysis at the granular level of a multi-class head cascade architecture. At the architecture level, it explores whether multi-class head cascades truly contribute to prediction performance through implementation. Specifically, it implements additional unique settings: 1) no multi-class head cascade (Wo_Both), such as... Figure 5 As shown, 2) is a prediction model with multiple classifiers and cascaded heads (Ours), such as... Figure 2 As shown, the comparison is performed, and the results are as follows. Figure 3 As shown, the multi-class head cascade model prediction is significantly more accurate than the model with the Wo_Both setting. Adding the multi-class head cascade significantly improves the model's performance. It can be observed that the multi-class head cascade implementation can highlight more reasonable regions, enabling more reliable short-term, medium-term, and long-term outcome probability predictions.

[0030] like Figure 4 As shown, this invention compares four state-of-the-art methods with a multi-stage prediction model (Ours) using a multi-class head cascade to evaluate the effectiveness of the proposed multi-stage prediction model in predicting the probability of multi-stage outcomes from acute-phase CT images. The four methods include: 1) Mikhail et al. developed a U-Net-based multitask model for pneumonia classification, in which pneumonia segmentation is used as an auxiliary task to enhance feature capture (from CT-Based COVID-19 triage: Deep multitask learning improves joint identification and severity quantification.).

[0031] 2) Simon et al. input randomly selected 2D CT slices from different slices into the Inception-ResNet-v2 model to predict the prognosis of progressive fibrotic lung disease (from Deep Learning-based Outcome Prediction in Progressive Fibrotic Lung Disease Using High-Resolution Computed Tomography).

[0032] 3) Yun et al. used CNN to extract depth features from multiple CT slices of chronic obstructive pulmonary disease to predict 3-year and 5-year survival rates (derived from Deep radiomics-based survival prediction inpatients with chronic obstructive pulmonary disease).

[0033] 4) Wang et al. developed an attention alignment model to capture high-risk lesion changes from mammograms to predict the occurrence of breast cancer within 1 to 5 years (in International Conference on Medical Image Computing and Computer-Assisted Intervention.).

[0034] To ensure fair comparison, all methods were trained on the same settings on the dataset of this invention: input was restricted to acute-phase lung CT scans, and output was reformulated to predict short-, medium-, and long-term outcomes using a classification head. The performance of the four comparison methods and the method proposed in this invention was evaluated on the test set of this invention, as detailed below. Figure 4 The proposed method achieved best predictive performance across all timeframes, with an average AUC of 0.818 and an accuracy of 0.755. The methods of Simon et al. and Yun et al. significantly underperformed other methods, possibly because 2D slice input cannot provide a global representation of pneumonia lesions. Mikhail et al.'s multi-task framework achieved comparable results in short-term predictions, possibly attributed to improved lesion feature extraction through a segmentation auxiliary task. Similarly, the attention alignment mechanism designed by Wang et al., designed to focus on high-risk lesion changes, demonstrated strong performance in short-term predictions. However, the method of this invention outperformed the comparative methods in most short, medium, and long-term timeframes. Therefore, the effectiveness of the predictive model with a cascaded multi-classifier head architecture is further confirmed.

[0035] The Adam optimizer is used in training the convolutional neural network to dynamically calculate the short-term, medium-term, and long-term total loss during training. Then, backpropagation is performed until the convolutional neural network achieves the best accuracy on the validation set, thus obtaining the optimal convolutional neural network.

[0036] In constructing a convolutional neural network for predicting the short-term, medium-term, and long-term probabilities of patient outcomes, this invention selects three classification heads corresponding to short-term (1 month), medium-term (3 months), and long-term (6 months) prediction tasks during the model structure design phase. These classification heads are then cascaded and integrated with intermediate features from the previous period as auxiliary inputs for the next period. This approach guides the temporal information, enables the simultaneous execution of multi-period outcome prediction tasks, and allows for the direct acquisition of short-term, medium-term, and long-term outcome prediction probabilities from input CT images, eliminating the need for individual predictions and improving outcome prediction efficiency.

[0037] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.

Claims

1. A multi-stage outcome prediction method based on convolutional neural network multi-classification head cascade, characterized in that, Includes the following steps: Acquire chest CT images of acute lung diseases and perform standardized resampling, cropping and intensity normalization on the chest CT images; Set up the first classification head to predict the short-term outcome probability based on high-dimensional feature images in chest CT images; Set up a second classifier to predict the intermediate outcome probability based on the high-dimensional feature images in chest CT images and the intermediate classification features of the first classifier; A third classifier is set up to predict the long-term outcome probability based on high-dimensional feature images in chest CT images, intermediate classification features of the first classifier, and intermediate classification features of the second classifier. A convolutional neural network was formed by training a first, second, and third classifier head using a loss function constructed with Focal loss to predict short-term, medium-term, and long-term outcome probabilities in the chest CT images. Using the convolutional neural network, the short-term, medium-term, and long-term probabilities of a patient's outcome are output based on the patient's chest CT images.

2. The multi-stage outcome prediction method based on convolutional neural network multi-classification head concatenation as described in claim 1, characterized in that: The high-dimensional image features It was extracted from chest CT images by the encoder.

3. The multi-stage outcome prediction method based on convolutional neural network multi-classification head concatenation as described in claim 2, characterized in that: The first, second, and third classification heads each consist of two linear layers and a softmax activation layer.

4. The multi-stage outcome prediction method based on convolutional neural network multi-classification head concatenation as described in claim 3, characterized in that: The first classification head is used to classify high-dimensional image features. Given the input, output the short-term outcome probability; The structural expression of the first classification head is: ; ; In the formula, The short-term regression probability is the output of the first classifier. , These are the first and second linear layers in the classification head, respectively. High-dimensional image features Intermediate classification features generated through the first linear layer of the first classification head. This is the softmax activation layer.

5. The multi-stage outcome prediction method based on convolutional neural network multi-classification head concatenation as described in claim 4, characterized in that: The second classification head is used to classify high-dimensional image features. and the intermediate classification features Compositional fusion characteristics Given the input, output the intermediate outcome probability; The structural expression for the second classification head is: ; ; In the formula, This represents the intermediate regression probability output by the second classification head. , These are the first and second linear layers in the classification head, respectively. For fusion features Intermediate classification features generated through the first linear layer of the second classification head. This is a softmax activation layer, and Concat is the feature concatenation operator.

6. The multi-stage outcome prediction method based on convolutional neural network multi-class head concatenation as described in claim 5, characterized in that: The third classification head is used to classify high-dimensional image features. The intermediate classification features and the intermediate classification features Compositional fusion characteristics Given the input, output the long-term outcome probability; The structural expression of the third classification head is: ; In the formula, This represents the long-term regress probability output by the third classification head. , These are the first and second linear layers in the classification head, respectively. This is a softmax activation layer, and Concat is the feature concatenation operator.

7. The multi-stage outcome prediction method based on convolutional neural network multi-classification head concatenation as described in claim 6, characterized in that: The process of constructing the loss function includes: The short-term outcome prediction loss of the first classification head is established using Focal loss. ,in, In the formula, As the gold standard for short-term reversion, This represents the positive probability output by the first classification head. This represents the negative probability output by the first classification head. An adjustment factor to balance the contribution of difficult and easy samples to the loss; The short-term outcome prediction loss of the second classification head is built using Focal loss. ,in, In the formula, As the gold standard for mid-term reversion, The positive probability output by the second classification head. The negative probability output by the second classification head; Short-term outcome prediction loss for a third classification head is built using Focal loss. ,in, In the formula, As the gold standard for long-term outcomes, The positive probability output by the third classification head. The negative probability output by the third classification head; Short-term reversion to predict loss Interim outcome prediction loss and long-term outcome prediction loss The combination is the total loss. ,in, .

8. The multi-stage outcome prediction method based on convolutional neural network multi-class head concatenation as described in claim 7, characterized in that: The Adam optimizer is used in training the convolutional neural network to dynamically calculate the short-term, medium-term, and long-term total loss during training. Then, backpropagation is performed until the convolutional neural network achieves the best accuracy on the validation set, thus obtaining the optimal convolutional neural network.