Classification Method and System for Invasive Evolution of Lung Adenocarcinoma Based on State Space Network
By adopting a classification method based on state space network in lung adenocarcinoma, teacher encoder, student encoder and fine-tuning encoder are constructed, the accuracy and efficiency of imaging characteristics classification of lung adenocarcinoma are solved, and more efficient classification effects are achieved and the risk of overtreatment is reduced.
Patent Information
- Application Number
- CN202410327064.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-03-21
AI Technical Summary
In the prior art, the classification accuracy of the IA type of GGO imaging characteristics of lung adenocarcinoma is poor and the classification efficiency is low, resulting in the risk of overtreatment in practical applications.
The invasive evolution classification method of lung adenocarcinoma based on state space network is adopted. By constructing teacher encoder, student encoder and fine-tuning encoder, a large sample data set is used to train teacher encoder and student encoder, and a small labeled data set is used to fine-tune the fine-tuning encoder to achieve more accurate and efficient classification of GGO imaging features.
It improves the accuracy and efficiency of IA type classification of GGO imaging characteristics, reduces the risk of overtreatment, and effectively overcomes the problems of small data volume and difficult label acquisition through self-supervised learning.
Smart Images

Figure CN118552758B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence medical technology, and relates to a classification method for lung glandular infiltration, and in particular to a classification method and system for the invasive evolution of lung adenocarcinoma based on a state space network. Background Art
[0002] Lung cancer is the malignant tumor with the highest mortality rate in the world. Lung cancer mainly includes small cell lung cancer and non-small cell lung cancer (accounting for the majority of lung cancers), and lung adenocarcinoma is the main pathological type of non-small cell lung cancer, accounting for about half of all lung cancers. In recent years, chest computed tomography (CT) has been widely used for lung cancer screening and early lung cancer diagnosis. Persistent ground-glass opacity (GGO) is considered to be a common imaging feature of early lung adenocarcinoma on chest CT. It is currently believed that for early lung adenocarcinoma with GGO manifestations, its evolution may have gone through different stages of invasiveness of non-invasive adenocarcinoma (atypical adenoid hyperplasia (AAH), carcinoma in situ (AIS) and microinvasive adenocarcinoma (MIA)) and invasive adenocarcinoma, while non-invasive adenocarcinoma (including AAH, AIS and MIA) does not require immediate surgical intervention.
[0003] In the existing technology, the prediction and classification task of GGO invasiveness of lung adenocarcinoma can be roughly divided into: multi-omics fusion method and deep neural network method, depending on whether or not the artificially designed imaging omics or other multi-omics features are integrated.
[0004] Multi-omics fusion methods generally fuse finely designed features to enhance the classification effect of the model. With the help of non-deep learning paradigms to assist the model, Xia (Xia et al., 2020) established a recurrent residual convolutional neural network of U-Net to segment GGO, and then established a DL scheme and a radiomics scheme to classify non-IA and IA. Hu (Hu, X et al., 2021) used a computer-aided diagnosis scheme (CADx) and fused deep learning and radiomics features to improve the classification performance of GGO IA / non-IA; Wang (Wang et al., 2021) established a cascade architecture with segmentation and classification networks on a priori mask images to provide a classification-guided attention mechanism for the original CT data to achieve better lung adenocarcinoma diagnosis performance; Gong (Gong et al., 2020) used preprocessing techniques such as resampling and residual learning architecture to predict the possibility of GGO being IA. Huang (Huang et al., 2022) combined CT images with general patient data and serum tumor markers to establish a multimodal data fusion model to distinguish IA from non-IA. Tung (Tung et al., 2023) proposed an effective CT lung invasiveness identifier through semi-automatic segmentation, which allows doctors to easily mark GGO objects in CT images and reduce annotation costs.
[0005] Pure deep neural network methods mainly focus on how to enhance the learning signal on a limited dataset to effectively train a prediction model. For the accurate classification of GGO, Ma (Ma et al., 2022) proposed a two-stage detection and classification model, using a class balance loss function to achieve accurate diagnosis; Yu (Yu et al., 2021) used a 3D multi-task deep learning network to determine the invasiveness of GGO, which helps to select patients with invasive lesions who need surgery and appropriate surgical methods; Ni (Ni et al., 2020) proposed a three-dimensional multi-receptive field automatic GGO adenocarcinoma invasiveness classification algorithm to provide guarantee for the subsequent treatment plan of patients; Jing (Jing et al., 2023) constructed an integrated multi-view 3D convolutional neural network to provide detailed predictive histological information for surgical treatment in terms of benign / malignant and pre-invasive / invasive GGO risk stratification. Fu (Fu et al., 2023) developed and validated a three-dimensional (3D) deep transfer learning model, providing a non-invasive, low-cost, fast and repeatable method for preoperative prediction of patients. PB-LNet (Zhang et al., 2023) applied class activation map to analyze GGO in preoperative CT and predict the IA pathological subtype without delineating nodules. Wang X (Wang et al., 2020) established three deep learning networks corresponding to the classification of AAH / AIS and MIA, the classification of MIA and IA, and the classification of AAH / AIS&MIA and IA respectively.
[0006] The invention patent application with the application number 202110363571.1 discloses a lung adenocarcinoma invasion degree prediction system based on a neural network, which includes the following five parts: the first part: the data input part; the second part: the data preprocessing part; the third part: the model construction part; the fourth part: the model training part; the fifth part: the prediction part; input the preprocessed data of the lesion characteristics and clinical characteristics of the lung adenocarcinoma patient to be predicted, and use the trained model for prediction, and output the prediction result: the invasion degree of the lung adenocarcinoma to be predicted is invasive adenocarcinoma or non-invasive adenocarcinoma (atypical adenomatous hyperplasia, carcinoma in situ or minimally invasive adenocarcinoma).
[0007] Although a lot of work has been done on the invasion degree of lung adenocarcinoma in the prior art, these above methods are still carried out based on a limited available dataset and fuse network features through limited prior knowledge, which cannot guarantee the classification performance of the network model in the real scenario. Moreover, the GGO imaging features between invasive adenocarcinoma (IA) and non-invasive adenocarcinoma (None-IA) are relatively difficult to distinguish, the classification accuracy of the IA type of GGO imaging features is poor, the classification efficiency is low, and there is a risk of over-treatment in clinical practice for lung cancer patients with GGO manifestations without being fully trained to guide the model. Summary of the Invention
[0008] The object of the present invention is to solve the technical problems of poor accuracy and low classification efficiency in the IA type classification of GGO imaging features in the prior art, and to provide a classification method and system for the invasive evolution of lung adenocarcinoma based on a state space network.
[0009] The present invention specifically adopts the following technical solutions to achieve the above object:
[0010] A classification method for the invasive evolution of lung adenocarcinoma based on a state space network, comprising the following steps:
[0011] Step S1, obtaining sample data;
[0012] Obtain sample data of lung CT images, where the sample data of lung CT images includes a large sample data set without GGO classification labels and a small data set with GGO classification labels;
[0013] Step S2, constructing a GGO invasiveness classification model;
[0014] Construct a GGO invasiveness classification model, where the GGO invasiveness classification model includes a teacher encoder, a student encoder, and a fine-tuning encoder, and the teacher encoder, the student encoder, and the fine-tuning encoder are all SSM state space networks;
[0015] Step S3, training the GGO invasiveness classification model;
[0016] Use the large sample data set without GGO classification labels to train the teacher encoder and the student encoder. The output of the teacher encoder is used as the pseudo label available to the student encoder, and the parameter update of the teacher encoder is guided by the exponential moving average of the output of the student encoder; select the expression of the student encoder as the backbone for the fine-tuning training of the fine-tuning encoder, and use the small data set with GGO classification labels to perform fine-tuning training on the fine-tuning encoder;
[0017] Step S4, real-time classification;
[0018] Obtain the lung CT image to be classified and input it into the trained fine-tuning encoder, and the fine-tuning encoder outputs the classification result.
[0019] Furthermore, in step S1, data preprocessing is performed on the sample data of lung CT images. The specific method is as follows:
[0020] Step S11, resample the CT image using a cubic spline image interpolation algorithm, and the voxel size is 1mm * 1mm * 1mm;
[0021] Step S12: Scan the resampled image with a window range of [-1200HU, 600HU], and normalize each CT value in the scan to the range of [0, 255].
[0022] Step S13: Use the scaling factor , map the gray values of the scanned image to the range [-1, 1] to obtain a normalized image, and use the normalized image as the input to the GGO invasiveness classification model.
[0023] Furthermore, in step S13, pass the normalized image through a patch embedding layer to obtain a patch sequence; for the patch sequence, establish two branches: a normal branch and a mask branch; among them, the mask branch randomly masks the patches in the patch sequence through a random mask of the local receptive field.
[0024] Furthermore, in step S2, the state space algorithm of the SSM state space network is as follows:
[0025]
[0026] Among them, represents the feature hidden state in the state space, represents a component of the output, represents the output, represents the input, represents calculating the multiplier for the last bit of the input sequence, where represents the sequence length; represents the evolution parameter, ; represents the projection parameter, ; represents the projection parameter, ; represents the input 's dimension.
[0027] Furthermore, in step S2, the SSM state space network includes a normalization module, a state space module, and a multi-layer perceptron; the state space module is defined as:
[0028]
[0029] Among them, represents the intermediate state of the state space module, represents the normalization layer, represents the state space calculation layer, represents the input of each state space module, represents the corresponding output, represents the multi-layer perceptron, Indicating the normalized , belongs to the sequence 1,..., M.
[0030] Furthermore, in step S3, when training the GGO invasiveness classification model, the fine-tuning encoder regularizes the fine-tuning process, and the specific regularization process is expressed as:
[0031]
[0032] where represents the entire GGO invasiveness classification model, represents the negative samples in the batch, represents the positive samples in a batch, represents the hyperparameter of the optimization function, represents any positive sample number, represents the number of any negative sample, represents the number of positive samples, represents the number of negative samples, represents a random positive sample, represents a random negative sample, represents the output of a certain positive sample passing through the GGO invasiveness classification model, represents the output of the negative sample;
[0033] The captured GGO invasiveness internal representation is used again to adjust the pre-trained student encoder, and the student encoder and the fine-tuning encoder are alternately trained to achieve the double-layer optimization between the teacher-student network and the fine-tuning network.
[0034] Furthermore, in step S3, when training the GGO invasiveness classification model, an iterative loss function is adopted, and the iterative loss function is specifically:
[0035]
[0036] where represents all label sequences, represents the i-th input, represents the i-th output, represents the total number of inputs, represents the entire GGO invasiveness classification model, represents the feature output of the student encoder, represents the feature output of the teacher encoder, represents the hyperparameter of the loss function, represents the cross-entropy classification loss function for the fine-tuning encoder, It represents the loss function inside the teacher encoder and the student encoder. It represents the final loss function.
[0037] A classification system for the invasive evolution of lung adenocarcinoma based on a state space network, comprising:
[0038] A sample data acquisition module for acquiring sample data of lung CT images, where the sample data of lung CT images includes a large sample data set without GGO classification labels and a small data set with GGO classification labels;
[0039] A GGO invasiveness classification model construction module for constructing a GGO invasiveness classification model, where the GGO invasiveness classification model includes a teacher encoder, a student encoder, and a fine-tuning encoder, and the teacher encoder, the student encoder, and the fine-tuning encoder are all SSM state space networks;
[0040] A GGO invasiveness classification model training module for training the teacher encoder and the student encoder using the large sample data set without GGO classification labels. The output of the teacher encoder serves as the pseudo-label available to the student encoder, and the parameter update of the teacher encoder is guided by the exponential moving average of the output of the student encoder; selecting the expression of the student encoder as the backbone for the fine-tuning training of the fine-tuning encoder, and performing fine-tuning training on the fine-tuning encoder using the small data set with GGO classification labels;
[0041] A real-time classification module for acquiring the lung CT image to be classified and inputting it into the trained fine-tuning encoder, and the fine-tuning encoder outputs the classification result.
[0042] A computer device, comprising a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the above method.
[0043] A computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the above method.
[0044] The beneficial effects of the present invention are as follows:
[0045] 1. In the present invention, in order to enable the model to extract global specific features, the SSM state space network is adopted, and the network adopts the mode of "state sequence mechanism + double-layer optimization mechanism fine-tuning", so that the trained GGO invasiveness classification model shows excellent robustness and generalization, realizes faster domain adaptation, and improves the classification accuracy and classification efficiency of the IA type of GGO imaging features.
[0046] 2. The present invention addresses the problem of over-treatment of invasive GGO, effectively alleviating the misclassification situation in practical applications, which is of great significance in clinical practice and can effectively reduce the risk of surgical infection.
[0047] 3. The present invention adopts self-supervised learning, effectively overcomes the problems of small GGO data volume and difficult label acquisition, combines the intrinsic heterogeneity characteristics of the data, effectively initializes the downstream GGO classification task, and enhances the model's ability to extract global features based on random masking.
[0048] 4. The present invention uses the state space network as the basic framework, achieves the model performance of the equivalent attention network through linear attention, and effectively reduces the number of model parameters and computational complexity, thereby efficiently identifying the categories of GGO.
[0049] 5. For the already trained model, it can quickly perform GGO classification, realize batch CT detection, can achieve unattended batch operation, and has a fast speed, which can be improved with the expansion of the device. It saves manpower and material resources in actual clinical applications, liberates doctors from work, and enables doctors to focus on the design of treatment plans for higher-level diagnoses.
[0050] 6. Promoting to the grass-roots level, the high efficiency and easy deployment of the model can solve the problem of insufficient professional medical resources in grass-roots hospitals, improve the diagnostic level of grass-roots hospitals, and reduce the situation of over-treatment in remote areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a schematic flowchart of the present invention;
[0052] Figure 2 is a schematic structural diagram of the GGO invasiveness classification model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0054] Therefore, based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0055] Embodiment 1
[0056] This embodiment provides a method for classifying the invasive evolution of lung adenocarcinoma based on a state space network, which is used to classify the GGO imaging features of lung CT images to obtain the classification result (invasive adenocarcinoma (IA) or non-invasive adenocarcinoma (None-IA)) of the image. As Figure 1 shown, the method specifically includes the following steps:
[0057] Step S1, obtain sample data;
[0058] Obtain lung CT image sample data, which includes a large sample dataset without GGO classification labels and a small dataset with GGO classification labels.
[0059] Among them, the large sample dataset comes from public datasets, including the LUNA 16 lung dataset and the PN9 public dataset, and the data in these two datasets do not carry label data. The small dataset comes from West China Hospital of Sichuan University. The data in this dataset all carry GGO classification labels, and these GGO classification labels are calibrated by different clinical experts through multiple rounds of annotation to ensure the accuracy and objectivity of each label. For controversial grade divisions, consultations and discussions will be carried out to decide whether to include them in the training database.
[0060] For the obtained sample data, data preprocessing is also performed. The specific method is as follows:
[0061] Step S11, use the cubic spline image interpolation algorithm to resample the CT image, and the voxel size is 1mm*1mm*1mm;
[0062] Step S12, scan the resampled image with a window range of [-1200HU, 600HU], and normalize each CT value in the scan to the range of [0, 255];
[0063] Step S13, use the scaling factor , map the gray value of the scanned image to the range [-1, 1] to obtain a normalized image, and use the normalized image as the input of the GGO invasiveness classification model.
[0064] In addition, the normalized image can also pass through a patch embedding layer to obtain a patch sequence; for the patch sequence, two branches, a normal branch and a mask branch, are established; among them, the mask branch randomly masks the patches in the patch sequence through random masking of the local receptive field to enhance the model's expression of locally highly heterogeneous features.
[0065] Step S2, construct a GGO invasiveness classification model;
[0066] Construct a GGO invasiveness classification model, which includes a teacher encoder, a student encoder, and a fine-tuning encoder. The teacher encoder, the student encoder, and the fine-tuning encoder are all SSM state space networks.
[0067] By constructing an invasiveness classification model with the ability to extract long-sequence features, as Figure 2 shown, it includes a teacher-student upstream pre-training network based on the state space architecture and masked autoencoding, and a downstream fine-tuning classifier that implements a two-layer optimization paradigm with the pre-training network.
[0068] The SSM state space network includes a normalization module, a state space module, and a multi-layer perceptron.
[0069] To enable the model to extract global specific features, the research intends to adopt a state sequence mechanism. The state space algorithm of the SSM state space network is:
[0070]
[0071] Among them, represents the feature hidden state in the state space, represents a component of the output, represents the output, represents the input, represents calculating the multiplier for the last bit of the input sequence, where represents the sequence length; represents the evolution parameter, ; represents the projection parameter, ; represents the projection parameter, ; represents the input dimension.
[0072] To better couple the state space blocks, in this embodiment, the three-dimensional image is divided into multiple input blocks and converted into one-dimensional vectors, and a part of the input blocks are randomly masked, and position embeddings for position sensing are added to further model the global vision.
[0073] Each state space module is defined as:
[0074]
[0075] Among them, represents the intermediate state of the state space module, represents the normalization layer, represents the state space calculation layer, represents the input of each state space module, represents Corresponding output, represents a multi-layer perceptron, represents the normalized , belonging to the sequence 1,..., M.
[0076] Step S3, training the GGO invasiveness classification model;
[0077] Use a large sample dataset without GGO classification labels to train the teacher encoder and the student encoder. The output of the teacher encoder serves as the pseudo-label available to the student encoder, and the parameter update of the teacher encoder is guided by the exponential moving average of the output of the student encoder. Select the expression of the student encoder as the backbone for the fine-tuning training of the fine-tuning encoder, and use a small dataset with GGO classification labels to fine-tune the fine-tuning encoder.
[0078] The output of the teacher encoder serves as the pseudo-label available to the student encoder, and the parameter update of the teacher encoder is guided by the exponential moving average (EMA) of the output of the student encoder. During the entire training process, multiple iterations are performed to calculate the update of the EMA, making its value converge to the average of the model parameters. There is a decay factor in the EMA calculation. The weight ratio of the earlier parameter values to the EMA gradually decreases, while the weight ratio of the newer parameter values gradually increases, enabling the EMA to effectively reflect the overall trend of the model parameters. The update rule of the EMA is given by given, where and represent the weight parameters of the teacher encoder and the student encoder respectively. During training, the hyperparameter increases from 0.996 to 1 according to the cosine schedule. In addition, the concentration and sharpening of the teacher encoder output are used to further promote convergence, aiming to make the teacher encoder and the student encoder produce consistent prediction results.
[0079] To further improve the robustness of the model, a fine-tuning encoder is used to perform binary classification of GGO instances. Compared with the teacher encoder with a larger input, the teacher encoder with a smaller input exhibits relatively fewer model parameters and computational complexity. In the fine-tuning task training, a lightweight student network framework is used as the feature encoder, showing excellent robustness and generalization, thus achieving faster domain adaptation. The pre-trained student encoder can indicate strong semantic correlations between domains. Therefore, the expression of the upper-layer student encoder is selected as the backbone for fine-tuning training to optimize the model expression ability in fine-tuning training.
[0080] To address the domain migration problem in fine-tuning tasks, this embodiment uses a two-layer optimization mechanism to regularize the fine-tuning process of the fine-tuning encoder. The two-layer optimization mechanism links the teacher encoder, the student encoder, and the fine-tuning encoder with a loss function This loss function is insensitive to imbalanced datasets. By improving the AUC score, the difference between positive and negative sample pairs is increased, and this difference is used as the training error. The fine-tuning encoder regularizes the fine-tuning process, and the specific regularization process is expressed as:
[0081]
[0082] where, represents the entire GGO invasiveness classification model, represents the negative samples in a batch, represents the positive samples in a batch, represents the hyperparameter of the optimization function, represents any positive sample number, represents the number of any negative sample, represents the number of positive samples, represents the number of negative samples, represents a random positive sample, represents a random negative sample, represents the output of a certain positive sample passing through the GGO invasiveness classification model, represents the output of the negative sample;
[0083] The captured GGO invasiveness internal representation is used again to adjust the pre-trained student encoder. The student encoder and the fine-tuning encoder are alternately trained to achieve two-layer optimization between the teacher-student network and the fine-tuning network.
[0084] When training the GGO invasiveness classification model, an iterative loss function is adopted. The iterative loss function is specifically:
[0085]
[0086] where, represents all label sequences, represents the i-th input, represents the i-th output, represents the total number of inputs, represents the entire GGO invasiveness classification model, represents the feature output of the student encoder, represents the feature output of the teacher encoder, represents the hyperparameter of the loss function, Represents the cross - entropy classification loss function for the fine - tuned encoder, Represents the loss function inside the teacher encoder and the student encoder, Represents the final loss function.
[0087] Use To standardize the output of the GGO invasiveness classification model. In Among them, Represents the mathematical abstraction of the entire fine - tuned network, Represents the input, Represents the output. The loss function here is based on cross - entropy. The cross - entropy loss function can measure the difference between the probability distribution of the model output and the true distribution. Therefore, during the training process, it can effectively guide the model output to be closer to the true distribution and improve the classification accuracy.
[0088] In Among them, the loss term obtained by comparing the global view (as the target feature) output by the teacher encoder with the local view processed by the student encoder is output to the fine - tuned encoder. Here, And Represent the feature outputs of different perspectives of the student encoder and the teacher encoder respectively, Represents the dot - product operator. This loss measures the difference between these two views. By minimizing the cross - entropy loss, the fine - tuned encoder learns these distributions, enabling the model to better distinguish GGO types. The loss function By normalizing Reduces the difference in the uneven distribution of data samples of the four subtypes of GGO.
[0089] The final comprehensive Takes the surrogate function As a penalty term and combines it with the cross - entropy loss to obtain the formula for the comprehensive loss function. To balance the weights between the penalty term and the cross - entropy, where Is set to 0.75.
[0090] Step S4, real - time classification;
[0091] Obtain the lung CT image to be classified and input it into the trained fine - tuned encoder. The fine - tuned encoder outputs the classification result, which includes invasive adenocarcinoma (IA) and non - invasive adenocarcinoma (None - IA), that is, a binary classification task.
[0092] Embodiment 2
[0093] This embodiment provides a lung adenocarcinoma invasive evolution classification system based on a state - space network. The system specifically includes:
[0094] A sample data acquisition module for acquiring lung CT image sample data, where the lung CT image sample data includes a large sample dataset without GGO classification labels and a small dataset with GGO classification labels.
[0095] Among them, the large sample dataset comes from public datasets, including the LUNA 16 lung dataset and the PN9 public dataset, and the data in these two datasets do not carry label data. The small dataset comes from West China Hospital of Sichuan University. The data in this dataset all carry GGO classification labels, and these GGO classification labels are calibrated by different clinical experts through multiple rounds of annotation to ensure the accuracy and objectivity of each label. For controversial grade classifications, consultations and discussions will be carried out to determine whether they are included in the training database.
[0096] For the acquired sample data, data preprocessing is also performed. The specific method is as follows:
[0097] Step S11, resample the CT image using the cubic spline image interpolation algorithm, with a voxel size of 1mm * 1mm * 1mm;
[0098] Step S12, scan the resampled image with a window range of [-1200HU, 600HU], and normalize each CT value in the scan to the range [0, 255];
[0099] Step S13, use a scaling factor , map the gray values of the scanned image to the range [-1, 1] to obtain a normalized image, and use the normalized image as the input of the GGO invasiveness classification model.
[0100] In addition, the normalized image can also pass through a patch embedding layer to obtain a patch sequence; for the patch sequence, two branches, a normal branch and a mask branch, are established; among them, the mask branch randomly masks the patches in the patch sequence through random masking of the local receptive field to enhance the model's expression of locally highly heterogeneous features.
[0101] A GGO invasiveness classification model construction module for constructing a GGO invasiveness classification model, where the GGO invasiveness classification model includes a teacher encoder, a student encoder, and a fine-tuning encoder, and the teacher encoder, student encoder, and fine-tuning encoder are all SSM state space networks.
[0102] By constructing an invasiveness classification model with the ability to extract long sequence features, as Figure 2 shown, it includes a teacher-student upstream pre-training network based on the state space architecture and masked autoencoding, and a downstream fine-tuning classifier that implements a two-layer optimization paradigm with the pre-training network.
[0103] The SSM state space network includes a normalization module, a state space module, and a multi-layer perceptron.
[0104] To enable the model to extract global specific features, the research intends to adopt a state sequence mechanism. The state space algorithm of the SSM state space network is as follows:
[0105]
[0106] Among them, represents the feature hidden state in the state space, represents a component of the output, represents the output, represents the input, represents calculating the multiplier for the last digit of the input sequence, where represents the sequence length; represents the evolution parameter, ; represents the projection parameter, ; represents the projection parameter, ; represents the input dimension.
[0107] To better couple the state space blocks, in this embodiment, the three-dimensional image is divided into multiple input blocks and converted into one-dimensional vectors. A part of the input blocks is randomly masked, and position embeddings for position sensing are added to further model global vision.
[0108] Each state space module is defined as:
[0109]
[0110] Among them, represents the intermediate state of the state space module, represents the normalization layer, represents the state space calculation layer, represents the input of each state space module, represents corresponding output, represents the multi-layer perceptron, represents the after normalization, belongs to the sequence 1,..., M.
[0111] The GGO invasiveness classification model training module is used to train the teacher encoder and the student encoder with a large sample dataset without GGO classification labels. The output of the teacher encoder serves as the pseudo-label available to the student encoder, and the update of the parameters of the teacher encoder is guided by the exponential moving average of the output of the student encoder. Select the expression of the student encoder as the backbone for the fine-tuning training of the fine-tuning encoder, and use a small dataset with GGO classification labels to perform fine-tuning training on the fine-tuning encoder.
[0112] The output of the teacher encoder serves as the pseudo-label available to the student encoder, and the update of the parameters of the teacher encoder is guided by the exponential moving average (EMA) of the output of the student encoder. During the entire training process, multiple iterations are performed to calculate the update of the EMA, making its value converge to the average of the model parameters. There is a decay factor in the EMA calculation. The weight ratio of the earlier parameter values to the EMA gradually decreases, while the weight ratio of the newer parameter values gradually increases, enabling the EMA to effectively reflect the overall trend of the model parameters. The update rule of the EMA is given by where and represent the weight parameters of the teacher encoder and the student encoder respectively. During training, the hyperparameter increases from 0.996 to 1 according to a cosine schedule. In addition, the concentration and sharpening of the output of the teacher encoder are used to further promote convergence, aiming to make the teacher encoder and the student encoder produce consistent prediction results.
[0113] To further improve the robustness of the model, a fine-tuning encoder is used to perform binary classification of GGO instances. Compared with the teacher encoder with a large input, the teacher encoder with a small input exhibits relatively fewer model parameters and computational complexity. In the fine-tuning task training, a lightweight student network framework is used as the feature encoder, showing excellent robustness and generalization, thus achieving faster domain adaptation. The pre-trained student encoder can demonstrate strong semantic correlations between domains. Therefore, the expression of the upper-layer student encoder is selected as the backbone for fine-tuning training to optimize the model expression ability in fine-tuning training.
[0114] To address the domain transfer problem in the fine-tuning task, this embodiment uses a two-layer optimization mechanism to regularize the fine-tuning process of the fine-tuning encoder. The two-layer optimization mechanism links the teacher encoder, the student encoder, and the fine-tuning encoder with a loss function . This loss function is insensitive to imbalanced datasets. By improving the AUC score, the difference between positive sample pairs and negative sample pairs is increased, and this difference is used as the training error. The fine-tuning encoder regularizes the fine-tuning process, and the specific regularization process is expressed as:
[0115]
[0116] Among them, represents the entire GGO invasiveness classification model, represents the negative samples in the batch, represents the positive samples in a batch, represents the hyperparameters of the optimization function, represents any positive sample number, represents the number of any negative sample, represents the number of positive samples, represents the number of negative samples, represents a random positive sample, represents a random negative sample, represents the output of a certain positive sample through the GGO invasiveness classification model, represents the output of the negative sample;
[0117] The captured GGO invasiveness internal representation is used again to adjust the pre-trained student encoder, and the student encoder and the fine-tuning encoder are alternately trained to achieve double-layer optimization between the teacher-student network and the fine-tuning network.
[0118] When training the GGO invasiveness classification model, an iterative loss function is adopted. The iterative loss function is specifically:
[0119]
[0120] Among them, represents all label sequences, represents the i-th input, represents the i-th output, represents the total number of inputs, represents the entire GGO invasiveness classification model, represents the feature output of the student encoder, represents the feature output of the teacher encoder, represents the hyperparameters of the loss function, represents the cross-entropy classification loss function for the fine-tuning encoder, represents the loss function inside the teacher encoder and the student encoder, represents the final loss function.
[0121] Use to standardize the output of the GGO invasiveness classification model. In Among them, represents the mathematical abstraction of the entire fine-tuning network, represents the input, It represents the output. The loss function here is based on cross-entropy. The cross-entropy loss function can measure the difference between the probability distribution of the model output and the true distribution. Therefore, during the training process, it can effectively guide the model output to be closer to the true distribution and improve the classification accuracy.
[0122] In , the global view output by the teacher encoder (as the target feature) is compared with the local view processed by the student encoder, and the obtained loss term is output to the fine-tuning encoder. Here, and respectively represent the feature outputs of different perspectives of the student encoder and the teacher encoder. represents the dot product operator. This loss measures the difference between these two views. By minimizing the cross-entropy loss, the fine-tuning encoder learns these distributions, enabling the model to better distinguish GGO types. The loss function relieves the difference in the uneven distribution of data samples of the four subtypes of GGO by normalizing .
[0123] The final comprehensive uses the surrogate function as a penalty term and combines it with the cross-entropy loss to obtain the formula for the comprehensive loss function. To balance the weights between the penalty term and the cross-entropy, is set to 0.75.
[0124] The real-time classification module is used to obtain the lung CT image to be classified and input it into the trained fine-tuning encoder, and the fine-tuning encoder outputs the classification result.
[0125] Example 3
[0126] A computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the lung adenocarcinoma invasive evolution classification method based on the state space network.
[0127] Among them, the computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.
[0128] The memory at least includes one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or D-interface display memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the memory may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is commonly used to store the operating system and various application software installed on the computer device, such as the program code of the lung adenocarcinoma invasive evolution classification method based on the state space network, etc. In addition, the memory can also be used to temporarily store various data that have been output or will be output.
[0129] In some embodiments, the processor may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the lung adenocarcinoma invasive evolution classification method based on the state space network.
[0130] Embodiment 4
[0131] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the lung adenocarcinoma invasive evolution classification method based on the state space network.
[0132] Wherein, the computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor to cause the at least one processor to execute the steps of the lung adenocarcinoma invasive evolution classification method based on the state space network as described above.
[0133] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method for classifying the invasive evolution of lung adenocarcinoma based on the state space network described in the embodiments of the present application.
Claims
1. A method for classifying the invasive evolution of lung adenocarcinoma based on a state space network, characterized in that: The following steps are involved: Step S1, obtaining sample data; Acquire lung CT image sample data, where the lung CT image sample data includes a large sample data set without GGO classification labels and a small data set with GGO classification labels; Step S2, constructing a GGO invasiveness classification model; Construct a GGO invasiveness classification model, which includes a teacher encoder, a student encoder, and a fine-tuning encoder, all of which are SSM state space networks; Step S3, training a GGO invasiveness classification model; A large sample dataset without GGO classification labels is used to train the teacher encoder and the student encoder. The output of the teacher encoder is used as a pseudo label that can be used by the student encoder, and the parameter update of the teacher encoder is guided by the exponential moving average of the student encoder output. The expression of the student encoder is selected as the backbone of the fine-tuning training of the fine-tuning encoder, and a small dataset with GGO classification labels is used to fine-tune the fine-tuning encoder. Step S4, real-time classification; Obtain a lung CT image to be classified and input it into a trained fine-tuning encoder, which then outputs a classification result. In step S3, when training the GGO invasiveness classification model, the fine-tuning encoder performs regularization constraints on the fine-tuning process. The specific regularization process is expressed as: in, Represents the entire GGO invasive classification model, Represents the negative samples in the batch, Represents the positive samples in a batch, represents the hyperparameters of the optimization function, represents any positive sample number, Represents the number of any negative sample, represents the number of positive samples, represents the number of negative samples, represents a random positive sample, represents a random negative sample, represents a positive sample output of the GGO invasiveness classification model. Represents the output of negative samples; The captured GGO aggressive intrinsic representation is used again to adjust the pre-trained student encoder. The student encoder and the fine-tuning encoder are trained alternately to achieve a two-layer optimization between the teacher-student network and the fine-tuning network.
2. The method for classifying the invasive evolution of lung adenocarcinoma based on a state space network according to claim 1, characterized in that: In step S1, data preprocessing is performed on the lung CT image sample data, and the specific method is as follows: Step S11, resampling the CT image using a cubic spline image interpolation algorithm, with a voxel size of 1 mm*1 mm*1 mm; Step S12, applying a window range of [-1200HU, 600HU] to scan the resampled image, and standardizing each CT value in the scan to the range of [0, 255]; Step S13, using the scaling factor , the grayscale value of the scanned image is mapped to the range [-1,1] to obtain a standardized image, and the standardized image is used as the input of the GGO invasiveness classification model.
3. The method for classifying the invasive evolution of lung adenocarcinoma based on a state space network according to claim 2, characterized in that: In step S13, the standardized image is passed through a patch embedding layer to obtain a patch sequence; for the patch sequence, two branches, a normal branch and a mask branch, are established; wherein the mask branch randomly masks the patches in the patch sequence by using a random mask of a local receptive field.
4. The method for classifying the invasive evolution of lung adenocarcinoma based on a state space network according to claim 1, characterized in that: In step S2, the state space algorithm of the SSM state space network is: in, represents the characteristic hidden state in the state space, represents a component of the output, Indicates the output, Indicates input, Indicates the calculation of the multiplier for the last digit of the input sequence, where Indicates the length of the sequence; represents the evolution parameter, ; represents the projection parameters, ; represents the projection parameters, ; Represents input Dimension.
5. The method for classifying the invasive evolution of lung adenocarcinoma based on a state space network according to claim 4, characterized in that: In step S2, the state space network contains multiple state space modules, each of which includes a normalization layer, a state space calculation layer and a multilayer perceptron; the state space module is defined as: in, represents the intermediate state of the state space module, represents the normalization layer, represents the state space computation layer, represents the input of each state-space module, express The corresponding output is, represents a multilayer perceptron, Represents the normalized , Belongs to the sequence 1,...,M.
6. The method for classifying the invasive evolution of lung adenocarcinoma based on a state space network according to claim 1, characterized in that: In step S3, when training the GGO invasiveness classification model, an iterative loss function is used, and the iterative loss function is specifically: in, represents all tag sequences, represents the i-th input, represents the i-th output, Represents the total number of inputs, A mathematical abstraction representing the entire GGO aggressiveness classification model, represents the feature output of the student encoder, represents the feature output of the teacher encoder, represents the hyperparameters of the loss function, represents the cross entropy classification loss function for the fine-tuned encoder, represents the loss function inside the teacher encoder and the student encoder, Represents the final loss function.
7. A state-space network-based classification system for the invasive evolution of lung adenocarcinoma, characterized in that: include: A sample data acquisition module is used to acquire lung CT image sample data, where the lung CT image sample data includes a large sample data set without GGO classification labels and a small data set with GGO classification labels; A GGO invasiveness classification model construction module is used to construct a GGO invasiveness classification model. The GGO invasiveness classification model includes a teacher encoder, a student encoder, and a fine-tuning encoder. The teacher encoder, the student encoder, and the fine-tuning encoder are all SSM state space networks. A GGO invasiveness classification model training module is used to train a teacher encoder and a student encoder using a large sample dataset without GGO classification labels, wherein the output of the teacher encoder is used as a pseudo label that can be used by the student encoder, and the parameter update of the teacher encoder is guided by the exponential moving average of the output of the student encoder; the expression of the student encoder is selected as the backbone of the fine-tuning training of the fine-tuning encoder, and the fine-tuning encoder is fine-tuned using a small dataset with GGO classification labels; A real-time classification module is used to obtain a lung CT image to be classified and input the trained fine-tuning encoder, which outputs the classification result; In the GGO invasive classification model training module, when training the GGO invasive classification model, the fine-tuning encoder regularizes the fine-tuning process. The specific regularization process is expressed as: in, Represents the entire GGO invasive classification model, Represents the negative samples in the batch, Represents the positive samples in a batch, represents the hyperparameters of the optimization function, represents any positive sample number, Represents the number of any negative sample, represents the number of positive samples, represents the number of negative samples, represents a random positive sample, represents a random negative sample, represents a positive sample output of the GGO invasiveness classification model. Represents the output of negative samples; The captured GGO aggressive intrinsic representation is used again to adjust the pre-trained student encoder. The student encoder and the fine-tuning encoder are trained alternately to achieve a two-layer optimization between the teacher-student network and the fine-tuning network.
8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Lung adenocarcinoma infiltration degree prediction system based on neural network
CN115169425A
Migration reinforcement learning system oriented to agent visual control
CN117725987A
System and method for training student friendly teacher model and student model
US20220207431A1