Classification method for CIP pneumonia based on cross-time-phase CT (Computed Tomography) image

Through the cross-phase CT image feature extraction IP-VIT network, the problems of insufficient utilization of CT image information in different phases and inaccurate capture of lesion features in the existing technology are solved, achieving higher CIP pneumonia diagnostic accuracy and clinical decision support.

CN120599322APending Publication Date: 2025-09-05CHINA JILIANG UNIV +1

Patent Information

Application Number
CN202510449247.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively utilizing CT image information at different phases in the diagnosis of CIP pneumonia, and traditional convolutional neural networks have difficulty accurately capturing the diversity and irregularity of CIP lesions, resulting in insufficient diagnostic accuracy.

Method used

The IP-VIT network based on feature extraction of cross-phase CT images is adopted. Through the image segmentation module, dynamic deformable convolution module, phase embedding layer, CP-SwinFormer structure, slice merging module and multi-layer perceptron classifier, combined with the dual-phase feature prompt fusion module, feature extraction and fusion of CT images of different phases are realized.

Benefits of technology

The classification accuracy of CT image data for CIP pneumonia has been improved, enabling a more comprehensive understanding of disease progression, adapting to the diverse characteristics of CIP lesions, and improving the computational efficiency and diagnostic accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599322A_ABST
    Figure CN120599322A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and medical image analysis, and discloses a cross-time-phase CT image-based classification method for CIP pneumonia, and the method comprises the steps: obtaining the CT image data of the chest of a CIP patient before and after ICIs treatment, carrying out the preprocessing of the CT image data, and inputting the preprocessed CT image data into an offline trained feature extraction IP-VIT network. The feature extraction IP-VIT network comprises an image division module, a dynamic deformable convolution module, a time phase embedding layer, a hierarchical structure of a CP-SwinFormer structure and a slice merging module, a multi-layer dual-time phase feature prompt fusion module and a multi-layer perceptron classifier. According to the feature extraction IP-VIT network based on the cross-time-phase CT image, the relevance between different time-phase CT images can be learned and utilized, so that the accuracy of CT image classification is more comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and medical image analysis, and specifically relates to a classification method for CIP pneumonia based on cross-phase CT images. Background Art

[0002] Immune checkpoint inhibitors (ICIs), a major breakthrough in tumor immunotherapy, have significantly improved the treatment of various malignancies. This type of targeted immunotherapy, by activating the body's inherent anti-tumor immune response by removing inhibitory signals from T cells in the tumor microenvironment, has achieved remarkable results in the treatment of various tumors, including lung cancer, melanoma, and renal cell carcinoma, extending the survival of some patients with advanced solid tumors by months or even years. However, this advanced immunotherapy strategy is not without cost, and immune-related adverse reactions have become a major challenge for clinicians. Among all adverse reactions, immune checkpoint inhibitor-associated pneumonitis (CIP) is of particular concern. Epidemiological studies have shown that the incidence of CIP ranges from approximately 2.7% to 9.3% across different tumor types and immunotherapy regimens, and the severity can range from mild inflammatory reactions to life-threatening acute respiratory failure.

[0003] Traditionally, the diagnosis of CIP relies primarily on clinical symptoms, imaging findings, and the exclusion of other diseases. However, due to its nonspecific clinical manifestations and overlapping imaging features with other lung diseases, diagnosis is difficult and prone to misdiagnosis or missed diagnosis. Medical imaging, especially computed tomography (CT), plays a key role in the diagnosis of CIP. Typical CT imaging manifestations of CIP include ground-glass opacities and lattice shadows. However, the identification and judgment of these imaging features still rely on the experience of radiologists, are highly subjective, and make it difficult to quantitatively analyze changes in lesions.

[0004] In recent years, deep learning technology has made significant progress in medical image analysis, enabling the automatic extraction of complex features from images, providing strong support for disease diagnosis and prognosis. However, the use of CT images for CIP diagnosis still faces several challenges. First, existing CIP-assisted diagnostic models underutilize cross-temporal information from clinical data. Patients may undergo multiple CT scans before and during ICI treatment. These images from different time periods contain important information about disease progression, but existing methods often analyze individual images independently, ignoring temporal variations. For example, patent publication number CN117668760A, "A Multimodal Classification Method for Immunosuppressant-Associated Pneumonia Based on Deep Learning," utilizes multimodal data from CT images of CIP pneumonia and electronic medical records. However, it still focuses on a single time point and fails to comprehensively utilize multiple CT images before and after ICI treatment. Second, the morphology and distribution of CIP lesions are diverse and irregular. Traditional convolutional neural networks have a fixed receptive field when processing lesions with complex geometric structures, making it difficult to accurately capture lesion features. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a classification method for CIP pneumonia based on cross-phase CT images, so as to more effectively utilize CT image information of different phases, accurately extract the morphological characteristics of lung lesions, and improve the accuracy of CT image data classification of CIP pneumonia.

[0006] In order to solve the above technical problems, the present invention provides a classification method for CIP pneumonia based on cross-phase CT images, which includes the following steps: obtaining chest CT image data of CIP patients before and after ICI treatment, then preprocessing the CT image data and inputting it into an offline trained feature extraction IP-VIT network. The feature extraction IP-VIT network includes an image segmentation module, a dynamic deformable convolution module, a phase embedding layer, a position encoding layer, a CP-SwinFormer structure, a slice merging module, a dual-phase feature prompt fusion module, and a multi-layer perceptron classifier; the preprocessed CT image data is processed by the feature extraction IP-VIT network as follows:

[0007] (1) The image is divided into a set of non-overlapping sub-slices after passing through the image partitioning module;

[0008] (2) The sub-slices are passed through a dynamic deformable convolution module to obtain a slice embedding sequence;

[0009] (3) The phase embedding layer generates a phase code for each sub-slice, and the position encoding layer adds an absolute position code to each sub-slice. The slice embedding sequence of each slice is added element-by-element to the corresponding phase code and absolute position code to obtain a token sequence;

[0010] (4) The token sequence is processed hierarchically by the CP-SwinFormer structure and the slice merging module to obtain a token sequence representing the dual-phase CT image features;

[0011] (5) The dual-phase CT image feature representation token sequence is passed through a two-layer dual-phase feature prompt fusion module to obtain the phase-fused token sequence;

[0012] (6) After the time-phase fusion, the classification head in the tokens sequence passes through a multi-layer perceptron classifier containing two linear layers, and finally outputs the probability value representing the CIP risk.

[0013] As an improvement of the classification method of CIP pneumonia based on cross-phase CT images of the present invention:

[0014] The preprocessing includes: spatially aligning the CT image data of the chest before and after ICIs treatment, then resampling the spatially aligned CT images to unify the resolution, normalizing and zero-centering the images, and selecting a group of slices in the center of the CT images as input to the feature extraction IP-VIT network.

[0015] As a further improvement of the classification method of CIP pneumonia based on cross-phase CT images of the present invention:

[0016] The phase embedding layer includes two learnable parameters that respectively represent the phase encoding of the CT images before and after ICIs treatment, and the parameters of the absolute position encoding and the phase encoding are the same as the dimensions of the slice embedding sequence;

[0017] A "[CLS] token" is added before all tokens in the token sequence as the classification header; if the pre-treatment CT image is missing, the phase mask at half of the positions corresponding to the pre-treatment CT image of ICIs in the token sequence is set to 1.

[0018] As a further improvement of the classification method of CIP pneumonia based on cross-phase CT images of the present invention:

[0019] The CP-SwinFormer structure is improved based on the Swin Transformer Block. The original window multi-head self-attention is replaced by cross-phase window multi-head self-attention, and the original shift window multi-head self-attention is replaced by cross-phase shift window multi-head self-attention. The cross-phase window multi-head self-attention and the cross-phase shift window multi-head self-attention merge the attention windows of the corresponding spatial positions of the dual-phase CT images.

[0020] The CP-SwinFormer structure consists of two consecutive layers of CP-SwinFormer modules in series. The first layer of CP-SwinFormer module uses cross-phase window multi-head self-attention to perform window self-attention operations, dividing the tokens sequence into multiple attention windows of fixed size according to the spatial position of the corresponding CT image. The tokens in each window only perform attention operations with the tokens in the same window; the window division method of the cross-phase shift window multi-head self-attention used in the second layer of CP-SwinFormer module is offset from the window division position of the cross-phase window multi-head self-attention in the first layer of CP-SwinFormer module, so that some areas of different windows in the cross-phase window multi-head self-attention can interact with the areas of adjacent windows.

[0021] As a further improvement of the classification method of CIP pneumonia based on cross-phase CT images of the present invention:

[0022] The slice merging module integrates 2×2×2 adjacent patches into a single patch in three spatial dimensions through convolution kernels, reduces the three spatial dimensions by half and expands the number of channels to twice the original number of channels.

[0023] As a further improvement of the classification method of CIP pneumonia based on cross-phase CT images of the present invention:

[0024] The operation of the dual-time phase feature prompt fusion module is:

[0025] (1) The input tokens sequence is split into two parts: pre-treatment slice tokens and post-treatment slice tokens;

[0026] (2) The slice tokens after treatment are passed through the MLP layer to output query hint tokens of length 2;

[0027] (3) The query hint tokens and the pre-treatment slice tokens are concatenated and sent to the Transformer layer, and the output results corresponding to the query hint tokens are extracted as fusion information tokens;

[0028] (4) The fusion information tokens are concatenated with the post-treatment slice tokens and the fusion prompt tokens and then sent to the Transformer layer again to complete the feature fusion. The fusion prompt tokens are generated by the trainable parameter quantity and have a length of 2.

[0029] As a further improvement of the classification method of CIP pneumonia based on cross-phase CT images of the present invention:

[0030] The offline training process of the feature extraction IP-VIT network is as follows:

[0031] (1) Collect chest CT imaging data before and after ICI treatment, including CT images of CIP patients and non-CIP patients, among which some patients only have CT imaging data after ICI treatment;

[0032] (2) performing the aforementioned preprocessing on the CT image data, and then performing a data augmentation operation to expand the number of CT images;

[0033] (3) The data set after data augmentation was randomly divided into 5 parts, one of which was selected as the test set in turn, and the remaining 4 parts were integrated and randomly divided into training set and validation set in a ratio of 9:1, and 5-fold cross-validation training was performed;

[0034] (4) The image features in the training set are input into the IP-VIT network, the cross entropy loss function is calculated through forward propagation, and then the gradient is calculated through back propagation. The Adam optimizer is used to update the training parameters based on the gradient. The training ends after the preset epochs are reached; then the evaluation indicators of the model are calculated using the test set.

[0035] The beneficial effects of the present invention are mainly reflected in:

[0036] 1. The feature extraction IP-VIT network based on cross-phase CT images of the present invention can learn and utilize the correlation between CT images of different phases, thereby more comprehensively understanding the progression of the disease and improving the accuracy of CT image classification. At the same time, the feature extraction IP-VIT network of the present invention can also solve the reasoning problem in the case of missing phase CT images through phase masking.

[0037] 2. The IP-VIT network, based on feature extraction from cross-temporal CT images, employs a dynamic deformable convolution module tailored to the diverse characteristics of CIP lesions. This adaptively adjusts the receptive field to better match the irregular morphology and multi-scale characteristics of CIP lesions, thereby more accurately capturing lesion features. The combination of large-scale slicing and the dynamic deformable convolution module improves the overall computational efficiency of the model.

[0038] 3. The IP-VIT network based on feature extraction of cross-temporal CT images in the present invention can effectively improve the accuracy of CT image classification of CIP and provide more powerful decision support for clinicians. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The specific embodiments of the present invention are further described in detail below with reference to the accompanying drawings.

[0040] Figure 1 This is the overall architecture diagram of the IP-VIT network for feature extraction in the present invention;

[0041] Figure 2 This is the structural diagram of the dynamic deformable convolution module;

[0042] Figure 3 Comparison of sampling points between standard convolution kernel and deformable convolution kernel;

[0043] Figure 4 The CrossPhase-SwinFormer structure diagram of the present invention ((a) is a two-layer CP-SwinFormer structure, (b) is a cross-phase window multi-head self-attention, and (c) is a cross-phase shift window multi-head self-attention);

[0044] Figure 5 This is a structural diagram of the dual-phase feature prompt fusion module of the present invention. DETAILED DESCRIPTION

[0045] The present invention is further described below with reference to specific embodiments, but the protection scope of the present invention is not limited thereto:

[0046] Example 1: A classification method for CIP pneumonia based on cross-phase CT images. A training dataset is constructed by collecting CT image data from a cooperative hospital. The feature extraction IP-VIT network of the present invention is trained offline to perform online inference on chest CT image data of CIP patients before and after ICI treatment. The specific process is as follows:

[0047] Step 1: Data collection and dataset construction

[0048] Data collection mainly came from the respiratory department of the cooperating hospital, including chest CT imaging data before and after treatment with immune checkpoint inhibitors (ICIs) after desensitization. The data included CT images of 356 CIP patients and 482 non-CIP patients (most of whom had lung diseases such as community-acquired pneumonia). 252 of the 356 CIP patients had dual-phase CT imaging data before and after ICI treatment, and 347 of the 482 non-CIP patients had dual-phase CT imaging data before and after ICI treatment. The remaining patients only provided CT imaging data after ICI treatment.

[0049] Step 2: CT image data preprocessing

[0050] In order to unify the input images and facilitate model processing, the input images are preprocessed through the following preprocessing process:

[0051] (1) Registration: 3D Slicer, an open-source medical image analysis and visualization platform software, was used to spatially align the CT images of the same patient before and after ICI treatment, so that the model could extract the image feature information for cross-temporal comparison.

[0052] (2) Resampling: Due to the differences in acquisition equipment and acquisition parameters of CT image data, all CT images after spatial alignment are resampled and the resolution is unified to [1mm, 1mm, 1mm].

[0053] (3) Normalization and centering: The pixel values ​​of the resampled CT images are normalized using the range [-1000, 400] as the upper and lower bounds and mapped to the range [0, 1]. After the mapping is completed, the zero-value centering operation is performed, and the average pixel value is subtracted from all pixel values.

[0054] (4) Slicing: After pixel value normalization and zero-value centering are completed, the image is sliced ​​according to 256×256 pixels, some peripheral slices are removed, and the 48 slices in the center of the CT image are uniformly selected as model input.

[0055] The dimensions of the preprocessed CT image data are 256×256×48 (width×height×number of slices).

[0056] Step 3: Construct feature extraction IP-VIT network

[0057] Construct a feature extraction IP-VIT network based on cross-phase CT images, which is improved based on the VIT (Vision Transformer) network and is used to extract the features of CT images of patients before and after ICIs treatment. The structure is as follows Figure 1 As shown. Specifically, it includes an image segmentation module, a dynamic deformable convolution module, a phase embedding layer, a position encoding layer, a CP-SwinFormer structure, a slice merging module, a dual-phase feature prompt fusion module and an MLP classifier. Compared with the original VIT (Vision Transformer) network, the feature extraction IP-VIT network of the present invention realizes effective feature extraction of CIP diversity lesions through the dynamic deformable convolution module, and at the same time solves the phase distinction and phase missing problems in the dual-phase CT image processing process by using phase embedding and phase masking. The CP-SwinFormer structure used in the feature extraction IP-VIT network applies an improved cross-phase window multi-head self-attention, which significantly enhances the interaction of CT image features of different phases in the feature extraction process, and improves the computational efficiency. In addition, through the dual-phase feature prompt fusion module, the feature extraction IP-VIT network realizes efficient fusion of dual-phase CT image features with an improved prompt learning method, greatly reducing the computational resource requirements for processing dual-phase CT images, and achieving excellent fusion effects.

[0058] (1) Image segmentation module

[0059] The dimensions of the preprocessed CT image data are 256×256×48 (width×height×number of slices). The feature extraction IP-VIT network adopts a large-size slice partitioning strategy. First, each 256×256 pixel slice is divided into non-overlapping sub-slices of 64×64 pixels. For each CT image data, a total of N=4×4×48=768 sub-slices are used.

[0060] (2) Dynamically deformable convolution module

[0061] The structure of the dynamic deformable convolution module is as follows Figure 2 As shown in Figure 1, two consecutive layers of deformable convolution kernels with a kernel size of 5×5, a stride of 3, and 6 output channels are used to extract local features in the slice. The first layer of deformable convolution is followed by a ReLU activation function, while the second layer of deformable convolution is not followed by an activation function. The deformable convolution kernel learns the offset Δp through two separately set convolution layers. n and the modulation factor Δm n , adaptively adjust the receptive field to match the morphological characteristics of CIP lesions, so that the sampling position can be flexibly adjusted according to the context during the convolution kernel operation, thereby sampling the image features more specifically. Figure 3 As shown in the figure, the sampling point position and sampling weight of the deformable convolution kernel can be adjusted accordingly according to the current sampling image. The calculation process of the deformable convolution kernel is as follows:

[0062]

[0063] Among them, y(p0) is the value of the output feature map at position p0; p n Represents the sampling points in the convolution kernel sampling grid; R is the sampling grid of the convolution kernel, for example, R = {(-1,-1), (-1,0), …, (1,1)} represents a 3 × 3 convolution kernel with a total of 9 sampling points;

[0064] w(p n ) is the convolution kernel at position p n The weight of x(p0+p n ) is the input feature map at position p0+p n value.

[0065] Deformable convolution via Δp n Adjust the sampling position of the convolution kernel. Before each convolution operation is performed, generate Δp through the convolution layer of the same size n , that is, the offset of each sampling point in the x-direction and y-direction, thereby adjusting the sampling position. Δm n It is the modulation factor of each sampling point. Before each convolution operation is performed, Δm is generated by a convolution layer of the same size. n, that is, the modulation factor of each sampling point, thereby achieving the weight adjustment for the sampling point. n = 0, the corresponding sampling point does not contribute any information. n =1, the sampling point contributes all its information.

[0066] After the dynamic deformable convolution module, each 64×64 pixel sub-slice is converted into a slice tensor of dimension 6×6×6 (width×height×number of channels), which is then flattened to a one-dimensional tensor of dimension 216 as a slice embedding, forming a slice embedding sequence.

[0067] In this example, the slice embedding sequence dimension is [216,768] ([number of channels, number of sub-slices]).

[0068] (3) Temporal embedding layer and position encoding layer

[0069] Because CT images before and after ICI treatment are processed simultaneously, the phase embedding layer generates a phase code for each sub-slice to distinguish between pre- and post-treatment CT images. During model construction, the model sets two learnable parameters to represent the phase codes of the CT images before and after ICI treatment. These parameters are updated during training and have the same dimensionality as the slice embedding sequence.

[0070] In addition, the position encoding layer adds an absolute position code to each sub-slice. For example, if the input CT image has dimensions of 256×256×48 and is divided into 768 sub-slices after slicing, the position encoding layer sets 768 learnable parameters, each corresponding to a specific position index. The absolute position encoding parameters have the same dimensions as the slice embedding sequence.

[0071] After the phase code and absolute position code are generated, the CT image slices before and after treatment are mapped to the phase code before and after treatment; the embedding vector of the i-th slice is mapped to the i-th learnable position code vector. The slice embedding of each slice is added element by element with the corresponding phase code and absolute position code to obtain a token, which constitutes the final token sequence as the input of the subsequent modules. Figure 1 As shown in Figure 2, the feature extraction IP-VIT network adds a "[CLS] token" before all tokens in the tokens sequence as a classification header used by subsequent classifiers.

[0072] For cases with missing pre-treatment CT images, the input image is padded with a 0-valued tensor, and the tokens at the corresponding positions are masked using a phase mask. The phase mask is a Boolean tensor with the same length as the token sequence. The masked tokens at the corresponding positions in the phase mask are set to 1. During subsequent self-attention operations, the attention weight corresponding to these positions is set to negative infinity, thus ignoring the corresponding tokens during the entire self-attention process.

[0073] If the pre-treatment CT image is missing, the phase mask at half of the positions corresponding to the ICIs pre-treatment CT image in the tokens sequence will be set to 1 to ensure that the missing information is ignored in the subsequent self-attention calculation.

[0074] (4) CP-SwinFormer structure

[0075] The tokens sequence after the dynamic deformable convolution module, position encoding layer and phase encoding layer is used as the input of the CP-SwinFormer (CrossPhase-SwinFormer) structure. Figure 4 As shown in (a), a basic CP-SwinFormer structure is formed by two consecutive layers of CP-SwinFormer modules in series. The CP-SwinFormer module is improved based on the Swin Transformer Block in the Swin Transformer network. The original window multi-head self-attention (W-MSA) in the Swin Transformer Block is improved and replaced with the cross-phase window multi-head self-attention (CPW-MSA) of the present invention, and the original shifted window multi-head self-attention (SW-MSA) is improved and replaced with the cross-phase shifted window multi-head self-attention (CPSW-MSA) of the present invention, thereby achieving the window range of the window multi-head self-attention (W-MSA) and the shifted window multi-head self-attention (SW-MSA) in the Swin Transformer Block to be extended to two windows corresponding to the same spatial position of the two phase images.

[0076] In the first layer of CP-SwinFormer module, CPW-MSA is used to perform window self-attention operations, dividing the token sequence into multiple attention windows of fixed size according to the spatial position of the corresponding CT image. The tokens in each window only perform attention operations with the tokens in the same window. The window division method of CPSW-MSA used in the second layer is offset from the window division position of CPW-MSA in the previous layer of CP-SwinFormer, so that certain areas of different windows in CPW-MSA can interact with the areas of adjacent windows. The window division of CPW-MSA and CPSW-MSA is as follows: Figure 4 As shown in (b) and (c).

[0077] Each CP-SwinFormer module contains a feed-forward network (FFN) and uses layer normalization and residual connections. The number of heads in the multi-head self-attention (MSA) layer is set to 8. The number of CP-SwinFormer layers in the network (i.e., the number of CP-SwinFormer modules) can be adjusted based on computing resource constraints and performance requirements. In the feature extraction IP-VIT network, L1 = L2, and both L1 and L2 are multiples of 2.

[0078] In this embodiment, L1=L2=2, that is, each CP-SwinFormer structure includes two layers of CP-SwinFormer modules connected in series. As the number of layers of CP-SwinFormer modules increases, the performance of the I feature extraction IP-VIT network has room for further improvement. For details, see Experiment 1.

[0079] (5) Slice merging module

[0080] The slice merging module is used to merge tokens from adjacent sub-slices in the spatial location of a 3D CT image. This module uses convolution kernels to combine 2×2×2 adjacent patches into a single patch in three spatial dimensions. This reduces the three spatial dimensions by half and doubles the number of channels to preserve information. After merging, the number of sub-slice tokens is reduced to 1 / 8, while the number of channels is doubled.

[0081] In this embodiment, the token sequence is as follows Figure 1 The secondary CP-SwinFormer structure and the hierarchical processing of the slice merging module are shown.

[0082] (6) Dual-phase feature prompt fusion module

[0083] After the hierarchical processing of the CP-SwinFormer structure and the slice merging module, the token sequence representing the dual-phase CT image features is obtained and input into the two-layer dual-phase feature prompt fusion module. The module performs inter-phase feature fusion through prompt learning. Each layer of the dual-phase feature prompt fusion module only has one Transformer layer, and its structure is as follows: Figure 5 shown.

[0084] The module splits the input bi-phase CT image feature representation token sequence into two parts according to their corresponding phases (pre-treatment slice tokens and post-treatment slice tokens, and CLS tokens and post-treatment tokens are placed in the same part). This module first passes the post-treatment slice token sequence through the MLP layer to output query hint tokens of length 2.

[0085] Query hint tokens are used to query the pre-treatment slice token sequence for relevant fusion information. These query hint tokens are concatenated with the pre-treatment slice tokens and fed into the module's Transformer layer. The output corresponding to the query hint tokens is extracted as fusion information tokens. The fusion information tokens carry relevant feature information from the pre-treatment slice tokens. They are concatenated with the post-treatment slice token sequence and the fusion hint tokens and fed back into the module's Transformer layer, participating in the attention operation to complete feature fusion. The fusion hint tokens used are generated using a learnable parameter and are 2 tokens long.

[0086] The Transformer layer in the dual-temporal feature cue fusion module undergoes two forward passes. The module contains a multi-head self-attention (MSA) layer and a feed-forward network (FFN), and uses layer normalization and residual connections. The number of heads in the MSA layer is set to 8, and the hidden layer dimension is 864.

[0087] (7) MLP classifier

[0088] After all levels of processing, the output embedding vector corresponding to the “[CLS] token” is taken as the classification head, and passed through a multi-layer perceptron (MLP) classifier containing two linear layers to finally output a probability value representing the CIP risk.

[0089] Step 4: Model training and evaluation

[0090] (1) To enhance the generalization and robustness of the model, data augmentation is applied to the CT image dataset after the CT image data preprocessing in step 2. The data augmentation strategy for CT images mainly includes spatial transformation and intensity perturbation. Spatial transformation includes random flipping and random angle rotation. Intensity perturbation is mainly achieved by randomly adding Gaussian noise to the image, with noise σ = 0.01, which expands the data size to 5 times the original data.

[0091] (2) The augmented dataset is randomly divided into 5 parts, and one part is selected as the test set in turn. The remaining 4 parts are integrated and randomly divided into training and validation sets in a ratio of 9:1, and 5-fold cross-validation training is performed.

[0092] (3) Offline training of feature extraction IP-VIT network.

[0093] The images in the training set are input into the feature extraction IP-VIT network. The loss function used is the cross entropy loss function. The AdamW optimizer is adopted, the initial learning rate is set to 0.0001, the weight decay coefficient is 0.01, and the batch size is set to 4. The loss function is calculated by forward propagation, and then the gradient is calculated by backpropagation. The Adam optimizer is used to update the training parameters based on the gradient. A total of 100 epochs are trained.

[0094] The testing process is to input the test set images into the trained model, calculate the evaluation indicators, and obtain an accuracy rate of 84.32 and a recall rate of 82.5, thereby obtaining a feature extraction IP-VIT network that can be used online.

[0095] Step 5: Use online

[0096] After completing model training and evaluation, the feature extraction IP-VIT network of the present invention can be deployed in a clinical environment for online use. The process is as follows:

[0097] (1) Data acquisition: Clinicians obtain chest CT imaging data of patients who need diagnosis before and after ICI treatment through the hospital's information system and imaging archives.

[0098] (2) Data upload: The clinician uploads the patient's CT image data to the host computer where the model of the present invention is deployed.

[0099] (3) Data preprocessing: In the host computer, the CT image data is preprocessed according to step 2, including CT image registration, resampling, normalization and centering, and slicing.

[0100] (4) Model reasoning: The preprocessed CT image data is used as the input of the feature extraction IP-VIT network to execute the reasoning process.

[0101] (5) Result output: The model outputs the probability value of CIP risk corresponding to the CT image for reference by clinicians.

[0102] experiment:

[0103] 1. Feature Extraction IP-VIT Network Comparison Experiment

[0104] To validate the effectiveness of the proposed CIP-assisted diagnosis model for cross-temporal CT images, an experimental evaluation was conducted based on clinical data. The model's performance was measured using the following metrics: accuracy, recall, F1-score, and area under the curve (AUC). The experimental dataset used was the same dataset constructed in Example 1. Table 1 shows the comparative classification performance results obtained using the dataset used in this experiment under 5-fold cross-validation.

[0105] Table 1

[0106] method Accuracy (%) Recall rate (%) F1 score AUC Document 1 72.7 73.9 0.718 0.796 Document 2 78.9 80.3 0.781 0.807 Document 3 81.0 82.0 0.798 0.837 <![CDATA[IP-VIT(L1=L2=2)]]> 84.3 82.5 0.828 0.849 <![CDATA[IP-VIT(L1=L2=4)]]> 84.9 83.9 0.832 0.861

[0107] Experimental results show that the feature extraction IP-VIT network based on cross-phase CT images designed by the present invention can effectively extract the morphological features of lung lesions. When the number of parameters is kept roughly equal, the dynamic deformable convolution module introduced in the feature extraction IP-VIT network (L1=L2=2) (i.e., the solution of Example 1 of the present invention) is better adapted to the irregular shape and multi-scale characteristics of CIP lesions than traditional convolutional neural networks, and can capture typical imaging manifestations such as ground-glass shadows and grid shadows. At the same time, the model is compared with the 3D Vision Transformer and Swin Transformer models, and accurately extracts cross-phase features of CT images through phase encoding and phase embedding. At the same time, the improved cross-phase window self-attention and the phase feature fusion mechanism based on prompt learning significantly enhance the phase feature interaction, achieving an accuracy of 84.3% in the CIP recognition task and an AUC of 0.849.

[0108] At the same time, experimental results show that as the model scale expands and the number of CP-SwinFormer layers increases, the performance of the IP-VIT network has room for further improvement. The classification accuracy of IP-VIT (L1=L2=4) in the CIP recognition task reaches 84.9%, and the AUC reaches 0.861.

[0109] 2. Ablation Experiment

[0110] This study, through a series of ablation experiments, aimed to evaluate the contribution of the four core components of the feature extraction IP-VIT network (dynamic deformable convolution module, temporal embedding layer, CP-SwinFormer structure, and dual-temporal feature prompt fusion module) to model performance. The experiments compared the overall model performance before and after using the MLP and dynamic deformable convolution module to obtain slice embedding sequences, before and after enabling the temporal embedding layer, and before and after using the CP-SwinFormer structure and the dual-temporal feature prompt fusion module. Under 5-fold cross-validation, based on the dataset of Example 1, the performance test results of the combined model using different functional modules are shown in Table 2.

[0111] Table 2

[0112]

[0113] The experimental results in Table 2 show that the dynamic deformable convolution module is significantly better than the traditional MLP. Compared with model ①, model ⑤ (i.e., the present invention) replaces the MLP with the dynamic deformable convolution module, and the model accuracy is improved by 6.2 percentage points, which confirms the excellent ability of the dynamic deformable convolution module in capturing the morphological characteristics of CIP lesions. This is due to the precise matching of its adaptive receptive field to irregular lesions.

[0114] The phase embedding layer has a significant impact on model performance. Compared with model ②, model ⑤ (i.e., the present invention) increases the accuracy by 2.5 percentage points by adding the phase embedding layer, emphasizing the key role of phase context information in distinguishing image features before and after treatment and capturing treatment responses.

[0115] The CP-SwinFormer structure is also crucial. Based on the dynamic deformable convolution module and the temporal embedding layer, model ⑤ (i.e., the present invention) introduces the CP-SwinFormer structure, which improves the accuracy by 3.8 percentage points compared with model ③.

[0116] The dual-phase feature cue fusion module also plays a crucial role in extracting the interaction features between the two phases. Compared with model 4, model ⑤ (the present invention) improved accuracy by 4.2 percentage points and the area under the curve (AUC) by 0.077. The complete feature extraction IP-VIT network achieved the highest classification accuracy of 84.3% and an AUC of 0.849, fully demonstrating the ability of the feature extraction IP-VIT network to identify CIP CT images.

[0117] Finally, it should be noted that the above examples are merely specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples and is subject to numerous variations. All variations that can be directly derived or conceived by a person of ordinary skill in the art from the disclosure of the present invention are considered to be within the scope of protection of the present invention.

[0118] Page 1 of 1 He K,Zhang X,Ren S,et al.Deep Residual Learning for ImageRecognition[C] / / Proceedings of the IEEE Conference on Computer Vision andPattern Recognition.2016:770-778.

[0119] Page 2 also Liu Z,Lin Y,Cao Y,et al.Swin Transformer:Hierarchical VisionTransformer Using Shifted Windows[C] / / Proceedings of the IEEE / CVFIinternational Conference on Computer Vision.2021:10012-10022.

[0120] Figure 3 in Sadeghi A,Sadeghi M,Sharifpour A,et al.Potential DiagnosticApplication of a NovelDeep Learning-Based Approach for COVID-19[J].Scientific Reports,2024,14(1):280.DOI:10.1038 / s41598-023-50742-9

Claims

1. A classification method for CIP pneumonia based on cross-phase CT images, characterized by: The process involves obtaining chest CT imaging data from CIP patients before and after ICI treatment. The CT imaging data is then preprocessed and fed into an offline-trained feature extraction IP-VIT network. The feature extraction IP-VIT network includes an image segmentation module, a dynamic deformable convolution module, a temporal embedding layer, a position encoding layer, a CP-SwinFormer structure, a slice merging module, a dual-temporal feature cue fusion module, and a multi-layer perceptron classifier. The preprocessed CT imaging data is then processed by the feature extraction IP-VIT network as follows: (1) The image is divided into a set of non-overlapping sub-slices after passing through the image partitioning module; (2) The sub-slices are passed through a dynamic deformable convolution module to obtain a slice embedding sequence; (3) The phase embedding layer generates a phase code for each sub-slice, and the position encoding layer adds an absolute position code to each sub-slice. The slice embedding sequence of each slice is added element-by-element to the corresponding phase code and absolute position code to obtain a token sequence; (4) The token sequence is processed hierarchically by the CP-SwinFormer structure and the slice merging module to obtain a token sequence representing the dual-phase CT image features; (5) The dual-phase CT image feature representation token sequence is passed through a two-layer dual-phase feature prompt fusion module to obtain the phase-fused token sequence; (6) After the time-phase fusion, the classification head in the tokens sequence passes through a multi-layer perceptron classifier containing two linear layers, and finally outputs the probability value representing the CIP risk.

2. The classification method for CIP pneumonia based on cross-phase CT images according to claim 1, characterized in that: The preprocessing includes: spatially aligning the CT image data of the chest before and after ICIs treatment, then resampling the spatially aligned CT images to unify the resolution, normalizing and zero-centering the images, and selecting a group of slices in the center of the CT images as input to the feature extraction IP-VIT network.

3. The classification method for CIP pneumonia based on cross-temporal CT images according to claim 2, characterized in that: The phase embedding layer includes two learnable parameters that respectively represent the phase encoding of the CT images before and after ICIs treatment, and the parameters of the absolute position encoding and the phase encoding are the same as the dimensions of the slice embedding sequence; A "[CLS] token" is added before all tokens in the token sequence as the classification header; if the pre-treatment CT image is missing, the phase mask at half of the positions corresponding to the pre-treatment CT image of ICIs in the token sequence is set to 1.

4. The classification method for CIP pneumonia based on cross-temporal CT images according to claim 3, characterized in that: The CP-SwinFormer structure is improved based on the Swin Transformer Block. The original window multi-head self-attention is replaced by cross-phase window multi-head self-attention, and the original shift window multi-head self-attention is replaced by cross-phase shift window multi-head self-attention. The cross-phase window multi-head self-attention and the cross-phase shift window multi-head self-attention merge the attention windows of the corresponding spatial positions of the dual-phase CT images. The CP-SwinFormer structure consists of two consecutive layers of CP-SwinFormer modules in series. The first layer of CP-SwinFormer module uses cross-phase window multi-head self-attention to perform window self-attention operations, dividing the tokens sequence into multiple attention windows of fixed size according to the spatial position of the corresponding CT image. The tokens in each window only perform attention operations with the tokens in the same window; the window division method of the cross-phase shift window multi-head self-attention used in the second layer of CP-SwinFormer module is offset from the window division position of the cross-phase window multi-head self-attention in the first layer of CP-SwinFormer module, so that some areas of different windows in the cross-phase window multi-head self-attention can interact with the areas of adjacent windows.

5. The classification method for CIP pneumonia based on cross-temporal CT images according to claim 4, characterized in that: The slice merging module integrates 2×2×2 adjacent patches into a single patch in three spatial dimensions through convolution kernels, reduces the three spatial dimensions by half and expands the number of channels to twice the original number of channels.

6. The classification method for CIP pneumonia based on cross-temporal CT images according to claim 5, characterized in that: The operation of the dual-time phase feature prompt fusion module is: (1) The input tokens sequence is split into two parts: pre-treatment slice tokens and post-treatment slice tokens; (2) The slice tokens after treatment are passed through the MLP layer to output query hint tokens of length 2; (3) The query hint tokens and the pre-treatment slice tokens are concatenated and sent to the Transformer layer, and the output results corresponding to the query hint tokens are extracted as fusion information tokens; (4) The fusion information tokens are concatenated with the post-treatment slice tokens and the fusion prompt tokens and then sent to the Transformer layer again to complete the feature fusion. The fusion prompt tokens are generated by the trainable parameter quantity and have a length of 2.

7. The classification method for CIP pneumonia based on cross-temporal CT images according to claim 6, characterized in that: The offline training process of the feature extraction IP-VIT network is as follows: (1) Collect chest CT imaging data before and after ICI treatment, including CT images of CIP patients and non-CIP patients, among which some patients only have CT imaging data after ICI treatment; (2) performing the aforementioned preprocessing on the CT image data, and then performing a data augmentation operation to expand the number of CT images; (3) The data set after data augmentation was randomly divided into 5 parts, one of which was selected as the test set in turn, and the remaining 4 parts were integrated and randomly divided into training set and validation set in a ratio of 9:1, and 5-fold cross-validation training was performed; (4) The image features in the training set are input into the IP-VIT network, the cross entropy loss function is calculated through forward propagation, and then the gradient is calculated through back propagation. The Adam optimizer is used to update the training parameters based on the gradient. The training ends after the preset epochs are reached; then the evaluation indicators of the model are calculated using the test set.

Citation Information

Patent Citations

  • Multi-mode deep learning classification method suitable for immunosuppressor-related pneumonia

    CN117668760A

Cited By

  • CNN-ViT-Meta fusion model-based pulmonary tuberculosis intelligent identification method

    CN121121406A