Interstitial lung disease prediction method and device based on residual network and multiple instances, medium and program product
By applying residual network and multi-instance methods in ILD diagnosis, using the SPAIDNet framework to train HRCT images, the accuracy and consistency of ILD diagnosis in the prior art are solved, and the identification accuracy comparable to that of ILD experts is achieved.
Patent Information
- Application Number
- CN202510137178.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art has problems with accuracy and consistency in the diagnosis of interstitial lung disease (ILD), especially in the diagnosis of ILD subtypes common in Asian populations, and existing methods are difficult to provide accurate conclusions.
Using residual network and multi-instance method, the HRCT images are trained through the supervised convolutional neural network SPAIDNet framework to build a prediction model of common ILD subtypes, and using deep learning models to identify ILD subtypes such as HP, NSIP, OP, UIP, etc. to achieve accurate recognition of HRCT images.
This method can reliably identify ILD subtypes in HRCT images, and the recognition accuracy reaches the level of experienced ILD experts, overcoming the problem of inaccurate diagnosis in the prior art.
Smart Images

Figure CN120015293A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical treatment, and more specifically, to an interstitial lung disease prediction method, device, medium and program product based on residual network and multiple instances. Background Art
[0002] Interstitial lung disease (ILD) is a group of heterogeneous, life-threatening diseases characterized by progressive inflammation and / or fibrosis of the lungs, which may eventually lead to respiratory failure. Clinically, the main challenge of ILD is the difficulty in making a rapid and accurate diagnosis, which often leads to delays in therapeutic intervention. High-resolution computed tomography (HRCT) is one of the key technologies for the diagnosis and evaluation of ILD. Patients with different imaging patterns may have different prognoses, for example, patients with clear usual interstitial pneumonia (UIP) have a higher mortality rate. However, the classification of ILD on HRCT is often affected by observer subjectivity and relies heavily on the knowledge and experience of the radiologist. This difference may lead to inconsistent diagnoses, and diagnostic delays can seriously affect patient outcomes, as early recognition and treatment are essential to slow disease progression and improve quality of life.
[0003] Recent advances in artificial intelligence (AI) have shown potential in overcoming the limitations of human interpretation by providing objective, automated diagnostic assessments of lung images, improving diagnostic accuracy and consistency. Deep learning algorithms, especially convolutional neural networks (CNNs), have demonstrated impressive capabilities in medical imaging tasks. However, the performance of AI models can vary significantly across patient populations and clinical settings. Currently, there is an urgent need to develop specialized AI diagnostic models for ILD subtypes that are common in Asian populations. Summary of the invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a method, device, medium and program product for predicting interstitial lung disease based on residual network and multiple instances; the method of the present invention constructs a prediction model for common subtypes of interstitial lung disease by aggregating the residual network in the convolutional neural network and the histogram multiple instances, and recognizes HP, NSIP, OP and UIP in HRCT images through the deep learning model, and can reliably identify the corresponding interstitial pneumonia subtypes in HRCT images, and the recognition accuracy reaches the level of experienced interstitial lung disease experts.
[0005] The first aspect of the present application discloses a method for constructing an interstitial lung disease prediction model, the method comprising: 101. Obtain chest CT images and corresponding diagnostic result labels of training set samples; the CT images include images of at least any two of the following types of pneumonia: hypersensitivity pneumonitis, nonspecific interstitial pneumonia, organizing pneumonia, and common interstitial pneumonia; and the CT images are 2D slice images; 102, inputting the 2D slice image into a supervised convolutional neural network SPAIDNet framework for training, and calculating a first probability score of a single 2D slice image belonging to each pneumonia type; 103, collecting first probability scores of all 2D slice images in a single sample; and counting the number of 2D slice images corresponding to the same first probability score; 104. Input the number of 2D slice images, the first probability score of a single 2D slice image and the pneumonia type to which it belongs into a classifier to obtain a predicted classification result, compare it with the corresponding diagnosis result label, optimize the model according to the comparison result, and obtain a constructed prediction model.
[0006] In some embodiments, the method further includes processing the number of decimals of the first probability score in 102 to obtain second probability scores with the same number of decimals; the second probability scores are within a range between a minimum value and a maximum value; Optionally, the number of decimals is at least 1, preferably 2; Optionally, the standard for processing the decimal quantity includes any one of the following: rounding, directly discarding the Nth digit of the decimal, and N is a natural number greater than or equal to 1.
[0007] In some embodiments, between 103 and 104, the method further includes: normalizing the number of 2D slice images, the first probability score of a single 2D slice image, and the pneumonia type to which it belongs, and inputting the normalized features into a classifier; Optionally, the normalization processing method includes: dividing the difference between each true value of a feature and the average value of the feature by the standard deviation of the feature; Optionally, the CT image is a HRCT image.
[0008] In some embodiments, the supervised convolutional neural network SPAIDNet framework includes a pre-trained ResNet-18, a maximum pooling module, and a fully connected layer with a softmax activation function; Optionally, the ResNet-18 is an 18-layer residual framework deep learning model; the maximum pooling module performs feature aggregation; the combined features are input into a fully connected layer, and a probability score of a single sample suffering from each type of pneumonia (HP, NSIP, OP, UIP) is generated through a softmax activation function; Optionally, the supervised convolutional neural network SPAIDNet framework uses an SGD optimizer for training for M epochs, with an initial learning rate of a and a mini-batch size of b; Optionally, in each epoch, the algorithm will input all samples into the model in the set order for forward propagation, loss calculation, backpropagation and parameter update; Optionally, for disease types with small sample sizes, instance-balanced sampling is used to mitigate the impact of class imbalance, and class-balanced loss is selected as the loss function to train the model.
[0009] In some embodiments, the classifier includes any one or more of the following: K-nearest neighbor, decision tree, naive Bayes, logistic regression, support vector machine, random forest, gradient boosting tree, multilayer perceptron; Optionally, between 101 and 102, the method further includes: performing preprocessing including resampling, standardization, and cropping on the CT image to obtain a preprocessed 2D slice image.
[0010] The second aspect of the present application discloses a method for predicting interstitial lung disease based on a residual network and multiple instances, the method comprising: 201, obtaining a CT image of a subject to be identified; 202, inputting the CT image to be identified into the prediction model disclosed in the first aspect of the present application to obtain an auxiliary prediction classification result; Optionally, the CT image to be identified is a chest CT image; Optionally, the CT image to be identified is a HRCT image.
[0011] In some embodiments, the auxiliary prediction classification results include any one or more of the following: hypersensitivity pneumonitis (HP), nonspecific interstitial pneumonia (NSIP), organizing pneumonia (OP) and usual interstitial pneumonia (UIP).
[0012] The third aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the above method.
[0013] A fourth aspect of the present application discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0014] A fifth aspect of the present application discloses a computer program product, including a computer program, which implements the steps of the above method when executed by a processor.
[0015] This application has the following beneficial effects: 1. This application innovatively discloses a method for constructing an interstitial lung disease prediction model, which aggregates the residual network and histogram multiple instances of convolutional neural networks, uses the supervised convolutional neural network SPAIDNet framework to train HRCT image data sets containing HP, NSIP, OP and UIP, obtains a trained recognition model, and then uses the histogram likelihood to aggregate the category prediction probabilities of all 2D layers of the same set of HRCT to obtain 3D classification results. In clinical applications, the real data to be identified is identified through the above-mentioned trained model, and the classification of HP, NSIP, OP, and UIP on the corresponding HRCT image is determined. This method can reliably identify the corresponding interstitial pneumonia subtypes in HRCT images, and the recognition accuracy reaches the level of experienced interstitial lung disease experts. The method disclosed in the present invention uses histogram likelihood to aggregate the category prediction probabilities of all 2D layers of the same set of HRCT, and then uses a classifier for training, which effectively overcomes the problem of inaccurate diagnostic conclusions when diagnosing common subtypes of interstitial lung disease based on CT images in the prior art. Specifically, when diagnosing, the existing method needs to diagnose each section of the CT image that is split into 2D images separately, and then make a final diagnosis result based on the section results; but assuming that there are 200 images, when the diagnosis result obtained based on 100 images is result A, and the diagnosis result obtained based on the other 100 images is result B, it is difficult to obtain a more accurate conclusion at this time.
[0016] 2. This application innovatively adopts instance balanced sampling to alleviate the impact of class imbalance based on the sample size of the four types of pneumonia diseases, and selects class balanced loss as the loss function to train the model to ensure the accuracy of the application-side model in predicting HP, NSIP, OP and UIP results. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 is a schematic diagram of a method flow chart provided by the first aspect of an embodiment of the present invention; Figure 2 is a schematic diagram of a method flow chart provided by the second aspect of an embodiment of the present invention; Figure 3 is a schematic diagram of a system for constructing an interstitial lung disease prediction model provided by an embodiment of the present invention; Figure 4is a schematic diagram of an interstitial lung disease prediction system based on a residual network and multiple instances provided by an embodiment of the present invention; Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention; Figure 6 is a schematic diagram of the architecture of an exemplary computing device provided by an embodiment of the present invention; Figure 7 is a schematic diagram of a storage medium provided by an embodiment of the present invention; Figure 8 It is a block diagram of an intelligent classification system for common subtypes of interstitial lung disease based on residual network and histogram multi-instance aggregation provided by an embodiment of the present invention; Fig. 9 It is a classification block diagram of the SPAIDNet model provided by an embodiment of the present invention; Fig.10 is a model performance diagram provided by an embodiment of the present invention; wherein, Fig.10 A is the ROC curve of the internal validation set, Fig.10 B is the classification confusion matrix of the internal validation set, Fig.10 C is the ROC curve of the external test set 1, Fig.10 D is the classification confusion matrix of external test set 1, Fig.10 E is the ROC curve of external test set 2, Fig.10 F is the classification confusion matrix of external test set 2; Fig.11 is a representative Grad-CAM heat map of the HP, NSIP, OP and UIP models provided in the embodiments of the present invention; Fig.11 A is the original HRCT image of HP (middle), the corresponding attention heat map of HP (left), and the Merge image after pairing the original HRCT image of HP and the corresponding attention heat map. Fig.11 B is the NSIP original HRCT image (middle), the NSIP corresponding attention heat map (left), and the Merge image after pairing the NSIP original HRCT image and the corresponding attention heat map. Fig.11 C is the original HRCT image of OP (middle), the corresponding attention heat map of OP (left), and the Merge image after pairing the original HRCT image of OP and the corresponding attention heat map. Fig.11 D is the original HRCT image of UIP (middle), the corresponding attention heat map of UIP (left), and the Merge image after pairing the original HRCT image of UIP and the corresponding attention heat map; Fig.12 The ROC curves of different diseases identified by the SPAIDNet model and manually provided in the embodiments of the present invention; Fig.12A is the ROC curve of the junior radiologist in different cohorts. Fig.12 B is the ROC curve of the identification of junior radiologists in different cohorts under the assistance of AI. Fig.12 C is the ROC curve of the senior chest radiologist in different cohorts. Fig.12 D is the ROC curve of senior chest radiologists identifying in different cohorts with the assistance of AI. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0020] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0022] Figure 2 1 is a flow chart of a method for predicting interstitial lung disease based on a residual network and multiple instances provided by an embodiment of the present invention. Specifically, the method comprises the following steps: 201: Obtaining a CT image of a subject to be identified; In some embodiments, the term "subject" or "person to be tested" or "sample to be tested" as used herein refers to any animal (e.g., mammal), including but not limited to humans, non-human primates, rodents, etc., which will be the recipient of a specific treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Preferably, the subject is a human. In some embodiments, the sample to be tested is a patient clinically used for prognostic assessment.
[0023] 202: Inputting the CT image to be identified into the prediction model disclosed in the first aspect of the present application to obtain an auxiliary prediction classification result; Optionally, the CT image to be identified is a chest CT image; optionally, the CT image to be identified is a HRCT image.
[0024] In some embodiments, the auxiliary prediction classification results include any one or more of the following: hypersensitivity pneumonitis (HP), nonspecific interstitial pneumonia (NSIP), organizing pneumonia (OP) and usual interstitial pneumonia (UIP).
[0025] In some embodiments, the auxiliary prediction results include but are not limited to paper or electronic report forms. The results are only obtained by the intelligent machine based on the analysis of relevant data of the subject, and are only used as a reference for medical staff, and are not used as the final diagnosis result of the subject.
[0026] In some embodiments, Figure 1 As shown, a flow chart of a method for constructing an interstitial lung disease prediction model provided by an embodiment of the present invention is provided, and the construction method includes: 101. Obtain chest CT images and corresponding diagnostic result labels of training set samples; the CT images include images of at least any two of the following types of pneumonia: hypersensitivity pneumonitis (HP), nonspecific interstitial pneumonia (NSIP), organizing pneumonia (OP), and usual interstitial pneumonia (UIP); the CT images are 2D slice images; In some embodiments, between 101 and 102, the method further includes: performing preprocessing including resampling, standardization, and cropping on the CT image to obtain a preprocessed 2D slice image. Specifically, the preprocessing includes: resampling each HRCT image to 1*1*1mm and then adjusting the image to a lung window (-1500 HU to -600 HU), scanning each HRCT image for lung image segmentation (without retaining the pulmonary artery and heart); using a window center of 100 and a window width of 700, performing a window operation, adjusting the lung area to 224*224, and performing minimum-maximum normalization processing on the resized lung area.
[0027] 102, inputting the 2D slice image into a supervised convolutional neural network SPAIDNet framework for training, and calculating a first probability score of a single 2D slice image belonging to each pneumonia type; In some embodiments, the method further includes processing the number of decimals in the first probability score in 102 to obtain a second probability score with the same number of decimals; the second probability score is within a range between a minimum value and a maximum value; optionally, the number of decimals is at least 1, preferably 2; optionally, the standard for processing the number of decimals includes any one of the following: rounding, directly discarding the Nth digit of the decimal, where N is a natural number greater than or equal to 1.
[0028] In some embodiments, the supervised convolutional neural network SPAIDNet framework includes a pre-trained ResNet-18, a maximum pooling module, and a fully connected layer with a softmax activation function; Optionally, the ResNet-18 is an 18-layer residual framework deep learning model, whose weights are pre-trained by ImageNet and contain 10 million trainable parameters to extract instance features; the maximum pooling module performs feature aggregation; the combined features are input into a fully connected layer, and a probability score of a single sample suffering from each type of pneumonia (HP, NSIP, OP, UIP) is generated through a softmax activation function, and the fully connected layer is initialized using the Kaiming method; Optionally, the supervised convolutional neural network SPAIDNet framework uses an SGD optimizer to perform M epoch training, with an initial learning rate of a and a small batch size of b, wherein M is a natural number greater than 1, preferably 50; a is 0.001; and b is 32; Optionally, in each Epoch, the algorithm will input all samples into the model in a set order for forward propagation, loss calculation, backpropagation, and parameter update; Optionally, for disease types with a small number of samples, instance-balanced sampling is used to mitigate the impact of class imbalance, and class-balanced loss is selected as the loss function to train the model.
[0029] 103, collecting first probability scores of all 2D slice images in a single sample; and counting the number of 2D slice images corresponding to the same first probability score; In some embodiments, between 103 and 104, the method further includes: normalizing the number of 2D slice images, the first probability score of a single 2D slice image, and the pneumonia type to which it belongs, and inputting the normalized features into a classifier; optionally, the normalization method includes: dividing the difference between each true value of the feature and the average value of the feature by the standard deviation of the feature; 104. Input the number of 2D slice images, the first probability score of a single 2D slice image and the pneumonia type to which it belongs into a classifier to obtain a predicted classification result, compare it with the corresponding diagnosis result label, optimize the model according to the comparison result, and obtain a constructed prediction model.
[0030] Optionally, the CT image is a HRCT image.
[0031] In some embodiments, the classifier includes any one or more of the following: K-nearest neighbor, decision tree, naive Bayes, logistic regression, support vector machine, random forest, gradient boosting tree, multilayer perceptron.
[0032] Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention, such as Figure 5 As shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein the memories store computer-readable codes, and when the computer-readable codes are run by the one or more processors, the method described above may be executed.
[0033] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, operations and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X86 architecture or an ARM architecture.
[0034] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0035] For example, the method or device according to the embodiment of the present disclosure may also be implemented by Figure 6 The architecture of the computing device 3000 shown in FIG. Figure 6As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as ROM 3030 or hard disk 3070, may store various data or files used for processing and / or communication of the method provided by the present disclosure and program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 6 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 6 One or more components of a computing device are shown.
[0036] The embodiment of the present invention also provides a computer-readable storage medium, such as Figure 7 As shown, it is a schematic diagram of a storage medium 4000 provided in an embodiment of the present invention, and a computer readable instruction 4010 is stored on the computer storage medium 4020. When the computer readable instruction 4010 is executed by a processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0037] The embodiments of the present disclosure further provide a computer program product or system, including a computer program, which implements the steps of the above method when executed by a processor.
[0038] In some embodiments, this embodiment also discloses a system for constructing an interstitial lung disease prediction model, such as Figure 3 As shown, the system comprises: The first acquisition module 301 is used or configured to acquire chest CT images of training set samples and corresponding diagnostic result labels; the CT images include images of at least any two of the following pneumonia types: hypersensitivity pneumonitis, nonspecific interstitial pneumonia, organizing pneumonia, and common interstitial pneumonia; and the CT images are 2D slice images; A probability score calculation module 302 is used or configured to input the 2D slice image into a supervised convolutional neural network SPAIDNet framework for training, and calculate a first probability score of a single 2D slice image belonging to each pneumonia type; The probability score processing module 303 is used or configured to collect the first probability scores of all 2D slice images in a single sample; and count the number of 2D slice images corresponding to the same first probability score; The prediction model training module 304 is used for or configured to input the number of 2D slice images, the first probability score of a single 2D slice image and the pneumonia type to which it belongs into a classifier, obtain a prediction classification result, compare it with the corresponding diagnosis result label, optimize the model according to the comparison result, and obtain a constructed prediction model.
[0039] In some embodiments, this embodiment also discloses an interstitial lung disease prediction system based on a residual network and multiple instances, such as Figure 4 As shown, the system comprises: A second acquisition module 401 is used or configured to acquire a subject's to-be-identified CT image; A classification result output module 402 is used or configured to input the CT image to be identified into the prediction model disclosed in the first aspect of the present application to obtain an auxiliary prediction classification result; In some embodiments, this embodiment also discloses an intelligent classification and evaluation system for common subtypes of interstitial lung disease based on residual network and histogram multi-instance aggregation, such as Figure 5 As shown, the system comprises: A data acquisition module, which obtains HRCT image data sets containing HP, NSIP, OP and UIP for use in the deep learning model training module, the internal testing module and the external validation module, and obtains HRCT image data for data to be identified in the clinical application module; The data preprocessing module resamples, standardizes and cuts the data collected by the data acquisition module to a suitable size to obtain data that meets the requirements of subsequent modules; The deep learning model training module is implemented using the supervised convolutional neural network SPAIDNet framework. The SPAIDNet framework is trained with the preprocessed data to obtain a trained recognition model. Then, the histogram likelihood is used to aggregate the class prediction probabilities of all 2D layers of the same set of HRCT to obtain the 3D classification results. An internal testing module is used to perform internal testing on the trained recognition model according to the corresponding internal testing data set; The external verification module is used to identify the trained recognition model according to the corresponding external verification data set, and then perform statistical analysis on the recognition results and the manual recognition results to verify the accuracy of the recognition model; The clinical application module uses the recognition model to identify the real data to be identified and determines whether the corresponding HRCT image belongs to the HP, NSIP, OP, or UIP category.
[0040] In addition, the intelligent classification and evaluation system for common subtypes of interstitial lung disease based on residual network and histogram multi-instance aggregation according to the present invention may also have the following additional technical features: In some embodiments, the data acquisition module acquires 5213 HRCT image data sets including HP, NSIP, OP, and UIP; 3585 of them are used as training data sets for the deep learning model training module, 932 are used as test data sets for the internal testing module, and 696 are used as verification test sets for the external verification module.
[0041] In some embodiments, the inclusion criteria of the 5213 three-dimensional volume HRCT image dataset include: HP, NSIP, OP, and UIP that have been clearly diagnosed by multidisciplinary discussion.
[0042] In some embodiments, the processing content of the data preprocessing module includes: resampling each HRCT image to 1*1*1mm and then adjusting it to the lung window (-1500 HU to -600 HU), and then segmenting and cropping all images of the lung area.
[0043] In some of the embodiments, the window center is 100, the window width is 700, a window operation is performed, and then the lung area is adjusted to 224*224, and the resized lung area is subjected to minimum-maximum normalization processing.
[0044] In some of the embodiments, the supervised convolutional neural network SPAIDNet framework includes a pre-trained ResNet-18, a maximum pooling module, and a fully connected layer with a softmax activation function; The ResNet-18 contains 10 million trainable parameters to extract instance features; The maximum pooling module uses a maximum pooling operation to perform feature aggregation; The fully connected layer generates probability scores for classification through a softmax activation function.
[0045] In some of the embodiments, the deep learning model training module receives the 2D image data processed by the data preprocessing module, obtains a number of 2D slices, forms a slice bag, and then inputs the slice bag into the supervised convolutional neural network SPAIDNet framework for model training.
[0046] In some of the embodiments, the deep learning model training module uses a ten-fold cross-validation method to perform model training; in each fold, 80% of the training data is used as a development set for model training, and the remaining 20% is used as a tuning set for model selection; the model that obtains the best results on the tuning set is used for subsequent internal test sets and external validation.
[0047] In some of the embodiments, the deep learning model training module uses an optimizer to train the entire network for 50 epochs, with an initial learning rate of 0.001 and a small batch size of 32; instance balanced sampling is used to alleviate the impact of class imbalance, and class balanced loss is selected as the loss function to train the model; a simplified lon platform method based on adjusted set loss is used to adjust the learning rate.
[0048] In some of the embodiments, the system further comprises a visualization module which uses gradient weighted class activation mapping to visualize the decision-making process of the model.
[0049] The embodiments of the present invention are described in detail below through specific embodiments and application scenarios in conjunction with the accompanying drawings.
[0050] See also Figure 5 As shown, in some embodiments of the present invention, a common subtype intelligent classification evaluation system for interstitial lung disease based on residual network and histogram multi-instance aggregation is provided, which identifies HP, NSIP, OP and UIP for HRCT images based on a deep learning model. The present invention collects a number of image data and divides them into three parts, which are used for model training, internal testing and external testing, respectively, so as to obtain HP, NSIP, OP and UIP identification with satisfactory accuracy.
[0051] 1. Data collection: The present invention collected HRCT images and corresponding diagnostic information data sets of 5213 patients with HP, NSIP, OP and UIP from 3 hospitals (including 3433 male patients with an average age of 61.735 ± 10.841 years). Among them, 4517 patients from the first hospital from July 2016 to June 2023 were used as training sets and internal test sets; HRCT image data of 696 patients from the second and third hospitals from July 2016 to June 2023 were used as external validation test sets. The criteria for inclusion of the above datasets are as follows: multidisciplinary discussion has clearly diagnosed one of HP, NSIP, OP, and UIP, age is greater than 18 years old, and informed consent has been signed. HRCT with poor image quality or severe artifacts was excluded. HRCT without digital imaging or communication medical format (DICOM) was excluded. Patients diagnosed with other types of interstitial lung diseases, including sarcoidosis, connective tissue disease-related interstitial pneumonia, drug-related interstitial pneumonia, etc. were excluded.
[0052] II. HRCT Scanning Protocol: All patients underwent high-resolution computed tomography (HRCT) on multidetector-row CT systems, including GE Healthcare’s LightSpeed VCT / 64 (Chicago, IL, USA), Toshiba’s Aquilion ONETSX-301C / 320 (Tokyo, Japan), Philips’ iCT / 256 (Amsterdam, the Netherlands), and Siemens’ FLASH dual-source CT (Erlangen, Germany). Scanning was performed in the supine position in a single breath-hold scan.
[0053] Scanning and reconstruction parameters strictly followed the prescribed CT standards. The tube voltage was set in the range of 100–120 kV, the tube current range was between 100–300 mAs, and the gantry rotation speed was 0.8 seconds. The images were reconstructed step by step, the slice thickness reached 0.625–1.25 mm, and the scanning stage movement speed was 39.37 mm / s.
[0054] 3. Image preprocessing: All HRCT image datasets in DICOM format were preprocessed in several steps. Specifically, each HRCT was first resampled to a voxel representing 1*1*1mm, and then the image was adjusted to the lung window (-1500 HU to -600HU). The lung image segmentation of each HRCT scan was then performed using a deep learning method, which can be implemented using existing medical imaging solution software. The lung region was then extracted from each CT scan using the 3D bounding box extracted by the corresponding segmentation mask, excluding the pulmonary artery and heart. Finally, a window operation was performed with a window center of 100 and a window width of 700, and then the lung region was resized to 224*224, and the resized lung region was min-max normalized.
[0055] IV: Deep learning model and model training: The present invention adopts the supervised convolutional neural network SPAIDNet framework for fine-grained classification. Fig. 9 As shown in the figure, the pre-trained ResNet-18 is used, which is an 18-layer residual framework deep learning model with 10 million trainable parameters to extract instance features and uses the maximum pooling operation to aggregate features. The combined features are input into a fully connected layer, followed by a softmax activation function to generate probability scores for each type (HP, NSIP, OP, UIP). Since a patient has hundreds of layers of corresponding 2D lung slices, ResNet-18 will generate probability scores for these images separately. The present invention further uses likelihood histogram aggregation to summarize all possible diagnostic results for a patient to form 400 feature vectors. Finally, the present invention inputs these feature vectors into a support vector machine to finally give the diagnosis result for each patient.
[0056] Specifically, for HRCT images, the above preprocessing is first performed on them, and the processed images are segmented on the z-axis to obtain several image slices. Each image slice obtained is copied into a three-channel image as a two-dimensional instance. All instances are formed into a slice bag. Each three-channel image in the bag is passed to the SPAIDNet model for training.
[0057] Model training was performed using a ten-fold cross validation method. In each fold, the present invention split the patient-level training data, with approximately 80% of the training set data used as the training set for model construction and the remaining 20% used as the tuning set for model selection. After training, the model that can obtain the best results on all tuning sets is applied to the internal test set and the external validation set.
[0058] The weights of ResNet-18 in the SPAIDNet model are pre-trained by ImageNet, and the classification layer (fully connected layer) is initialized using the Kaiming method. The entire network is trained for 50 epochs using the SGD optimizer, with an initial learning rate of 0.001 and a mini-batch size of 32. An epoch refers to the training process of passing all samples in the training dataset once (and only once). In an epoch, the training algorithm inputs all samples into the model in a set order for forward propagation, loss calculation, backpropagation, and parameter update. The learning rate is an important hyperparameter that controls the step size of weight update. A smaller learning rate means that the model updates the weights more cautiously during training; the mini-batch size is set to 32, and each gradient update is based on a small number of samples. Considering the small sample size of OP in the training set, instance balanced sampling is used to mitigate the impact of class imbalance, and class balanced loss is selected as the loss function to train the model. Data augmentation is performed by a toolkit called batch generator. During the training phase, the performance of the model on the adjustment set is evaluated by accuracy.
[0059] In this paper, in order to aggregate the instance-level diagnostic probabilities into comprehensive information at the patient level, a histogram-based aggregation pipeline (SALH, Slide Aggregation via Likelihood Histograms) is designed. For each patient, the diagnostic probabilities of all HRCT slices on four types of diagnoses (such as UIP, HP, etc.) are calculated, and the probability distribution histogram is constructed based on this.
[0060] Specifically, when a HRCT image set containing 300 2D slices is input into the SPAIDNet model, the model will give 4 values for each 2D slice, which are interpreted as the probability of the slice belonging to one of the categories of HP, NSIP, OP, and UIP. For each category, the probability data of all 300 slices are summarized, and two decimal places are retained according to the rounding method. If it is less than 0.005, it is retained as 0.01, and if it is greater than 0.995, it is retained as 0.99. At this time, for each category, the probability of the 300 2D slices belonging to the category will be a value between 0.00-0.99. Then, the number of slices corresponding to each value is counted to form a probability distribution histogram. For each category, the histogram can be regarded as a 100-dimensional feature, which represents the number of slices within a certain probability range for that specific category, such as the number of slices corresponding to the UIP diagnosis within a certain probability range (such as 0.850-0.855). Since the present invention targets four common types of interstitial lung disease, a set of HRCT images will generate 4×100 histogram features, a total of 400 features, through the above-mentioned SPAIDNet and SALH pipelines, characterizing the distribution of different diagnostic categories in different probability ranges. These features are then normalized and used to train traditional machine learning classifiers to further optimize diagnostic prediction performance. The normalization method includes but is not limited to the difference between each true value of the feature and the mean value of the feature divided by the standard deviation of the feature.
[0061] V. Statistical analysis: In this part, the model was used to diagnose four groups of diseases; the performance of the model classification of HP, NSIP, OP and UIP was evaluated. Python3.7 was used for statistical analysis. The performance of HP, NSIP, OP and UIP classification was evaluated by area under the curve (AUC), sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV). Bootstrap (1000 times) was used to estimate the 95% confidence interval (CI) of each evaluation indicator. In addition, the DeLong test was used to compare the relationship between the AUC of the model and the manual recognition and judgment results of radiologists of different years of experience. The manual recognition and judgment was performed by radiologists of different years of experience on the display results of HRCT images.
[0062] The number of true positive, false positive, true negative, and false negative results of the classification performance by prediction results on the internal test and external validation sets are also described in a 4×4 contingency table representing the confusion matrix.
[0063] VI: Analysis of research results Table 1 shows the evaluation results of the classification model in the internal test set and the external validation set. According to the results, the SPAIDNet model prediction and the true label of the present invention have a good correlation with the four diseases. The AUC of HP, NSIP, OP and UIP are all close to 1.000, as shown in Table 1:
[0064]
[0065] The ROC curve (that is, the receiver operating characteristic curve) is as follows Fig.10 In addition, there is clear consistency between the model predictions and the ground truth in the internal validation set. Fig.10 (B) is the classification confusion matrix, showing the number of true positive results, false positive results, true negative results, and false negative results predicted by the model.
[0066] To enhance the interpretability of the model, Gradient Weighted Class Activation Mapping (Grad-CAM) is employed. Fig.11 Representative Grad-CAM heatmaps for the HP, NSIP, OP, and UIP models. For each sample, the original HRCT image (center) and the corresponding attention heatmap (left) are paired. The red in the heatmap highlights the activated areas associated with the predicted class, and the darker the color, the stronger the association. The orange arrow indicates the actual location of the object in the original image. The Grad-CAM heatmap shows that the model of the present invention is able to pay attention to the abnormalities of the image when making decisions.
[0067] Model performance in external validation set: The performance of SPAIDNet model is further evaluated in external validation set. The evaluation results are summarized in Fig.10 and Table 2. As can be seen from the figure and table, the performance of the SPAIDNet model in the external validation set has declined, but it still maintains a high level, as shown in Table 2:
[0068]
[0069] Furthermore, the ROC curves of the readers using the model of the present invention and two manual recognitions are shown in FIG. Fig.12As shown in Table 3. As can be seen from the figure and table, in the external validation cohorts I and II, the diagnostic accuracy of a junior radiologist with 3 years of experience was 0.73, and the macro-average AUC was 0.737 (0.691-0.784). However, its macro-average sensitivity and specificity were only 0.573 and 0.513, respectively. Another senior chest radiologist with 10 years of experience had a higher diagnostic accuracy of 0.786 and a macro-average AUC of 0.763 (0.713-0.812), but its macro-average sensitivity and specificity were still limited, at 0.611 and 0.561, respectively. In the diagnosis of HP and OP, the diagnostic performance of both radiologists was significantly improved after the auxiliary diagnosis of SPAIDNet. The accuracy of the junior radiologist increased to 0.907, the macro-average AUC was 0.817 (0.779-0.855), the sensitivity increased to 0.669, and the specificity remained unchanged. For senior radiologists, accuracy improved to 0.839, macro-average AUC was 0.787 (0.739-0.835), sensitivity improved to 0.642, and specificity did not change significantly. After adding AI results, the diagnostic ability of both radiologists improved in the diagnosis of four ILDs (OP: p<0.05, HP, NSIP, UIP: p<0.001), but no significant change was observed in the diagnosis of OP by senior chest radiologists. In addition, with AI assistance, junior radiologists performed better than senior radiologists who did not use AI assistance in the diagnosis of HP, NSIP, and UIP, and there was no statistically significant difference between the two in the diagnosis of OP. Finally, junior physicians performed better than senior physicians using AI assistance in the diagnosis of NSIP and UIP, and also performed better than junior physicians in the diagnosis of NSIP and UIP (p<0.05). The SPAIDNet model outperformed senior chest radiologists in the diagnosis of HP and OP, but achieved comparable performance in identifying NSIP and UIP. Table 3 is shown below:
[0070]
[0071] Junior: 5 years of experience as a radiologist; Senior: 15 years of experience as a respiratory imaging radiologist.
[0072] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0073] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0074] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0075] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0076] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0077] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0078] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. It should be understood by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for constructing an interstitial lung disease prediction model, characterized in that: The method comprises: 101, obtaining chest CT images and corresponding diagnostic result labels of training set samples; the CT images include at least any two of the following pneumonia types: HP, NSIP, OP, and UIP; the CT images are 2D slice images; 102, inputting the 2D slice image into a supervised convolutional neural network SPAIDNet framework for training, and calculating a first probability score of a single 2D slice image belonging to each pneumonia type; 103, collecting first probability scores of all 2D slice images in a single sample; and counting the number of 2D slice images corresponding to the same first probability score; 104. Input the number of 2D slice images, the first probability score of a single 2D slice image and the pneumonia type to which it belongs into a classifier to obtain a predicted classification result, compare it with the corresponding diagnosis result label, optimize the model according to the comparison result, and obtain a constructed prediction model.
2. The method for constructing an interstitial lung disease prediction model according to claim 1, characterized in that: The method further includes processing the number of decimals of the first probability scores in 102 to obtain second probability scores with the same number of decimals; the second probability scores are within a range between a minimum value and a maximum value; Optionally, the number of decimals is at least 1, preferably 2; Optionally, the standard for processing the decimal quantity includes any one of the following: rounding, directly discarding the Nth digit of the decimal, and N is a natural number greater than or equal to 1.
3. The method for constructing an interstitial lung disease prediction model according to claim 1, characterized in that: Between 103 and 104, the method further includes: normalizing the number of 2D slice images, the first probability score of a single 2D slice image, and the pneumonia type to which it belongs, and inputting the normalized features into a classifier; Optionally, the normalization processing method includes: dividing the difference between each true value of a feature and the average value of the feature by the standard deviation of the feature; Optionally, the CT image is a HRCT image.
4. The method for constructing an interstitial lung disease prediction model according to claim 1, characterized in that: The supervised convolutional neural network SPAIDNet framework includes a pre-trained ResNet-18, a maximum pooling module, and a fully connected layer with a softmax activation function; Optionally, the ResNet-18 is an 18-layer residual framework deep learning model; the maximum pooling module performs feature aggregation; The combined features are input into the fully connected layer, and the probability score of a single sample suffering from each type of pneumonia is generated through the softmax activation function; Optionally, the supervised convolutional neural network SPAIDNet framework uses an SGD optimizer for training for M epochs, with an initial learning rate of a and a mini-batch size of b; Optionally, in each epoch, the algorithm will input all samples into the model in the set order for forward propagation, loss calculation, backpropagation and parameter update; Optionally, for disease types with small sample sizes, instance-balanced sampling is used to mitigate the impact of class imbalance, and class-balanced loss is selected as the loss function to train the model.
5. The method for constructing an interstitial lung disease prediction model according to claim 1, characterized in that: The classifier includes any one or more of the following: K-nearest neighbor, decision tree, naive Bayes, logistic regression, support vector machine, random forest, gradient boosting tree, multi-layer perceptron; Optionally, between 101 and 102, the method further includes: performing preprocessing including resampling, standardization, and cropping on the CT image to obtain a preprocessed 2D slice image.
6. A method for predicting interstitial lung disease based on residual networks and multiple instances, characterized in that: The method comprises: 201, obtaining a CT image of a subject to be identified; 202, inputting the CT image to be identified into the prediction model according to any one of claims 1 to 5 to obtain an auxiliary prediction classification result; Optionally, the CT image to be identified is a chest CT image; Optionally, the CT image to be identified is a HRCT image.
7. The method for predicting interstitial lung disease based on residual network and multiple instances according to claim 6, characterized in that: The auxiliary prediction classification results include any one or more of the following: hypersensitivity pneumonitis, nonspecific interstitial pneumonia, organizing pneumonia and common interstitial pneumonia.
8. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Lung cancer pathological section classification method based on multi-scale pyramid convolutional neural network
CN110781953A
Method and system for classification and visualisation of 3D images
CN113728335A
Image processing method and device, electronic equipment and storage medium
CN118351061A
Intelligent classification method for infantile pneumonia based on multi-modal data
CN118503921A
Lung disease multi-subtype classification method, system and equipment based on chest image
CN119206380A